MLOps is the discipline that keeps machine learning running after the experiment ends: model deployment, model monitoring, model versioning, and the feature stores that feed both training and serving. Where an ML engineer stops at the checkpoint, an MLOps engineer takes the checkpoint to production and keeps it honest there. The tooling has consolidated onto Kubernetes. Kubeflow, created at Google in 2017 and an incubating CNCF project from 2023, graduated in August 2026 with more than 6,600 contributors across 1,000-plus organizations and over 33,000 GitHub stars . KServe, the inference layer that pairs with it, was accepted into CNCF at incubating maturity in 2025 .
Challenges in MLOps Recruiting
Machine learning operations grew its own platform layer on Kubernetes
The discipline now runs on a platform that did not exist when most MLOps veterans learned the trade. Kubeflow started as a bundle of components at Google in 2017 and arrived at CNCF graduation in August 2026 with 6,600-plus contributors, more than 1,000 participating organizations, and a roadmap pointed at LLM orchestration, fine-tuning, and agentic workloads . That growth tracks a real change in the job. Machine learning operations used to mean shell scripts, cron, and a hopeful notebook handoff. Now it means platform engineering on Kubernetes: operators, CRDs, GPU scheduling, training jobs that must restart cleanly from a checkpoint. A candidate whose experience predates the platform may hold a decade of depth without a year of the current stack, and a candidate who has only clicked through a hosted platform has never debugged a pod. Both populations sit in the same applicant pool, and neither one alone staffs a modern platform team.
Model deployment fragmented across serving runtimes
Serving stopped being a single deployment path. KServe, accepted into CNCF as an incubating project in 2025, exposes one inference API over many runtimes: scikit-learn and XGBoost for predictive work, vLLM for generative serving, with KV cache offloading, disaggregated serving, and event-driven autoscaling through KEDA . One job title now covers people who can tune a vLLM server's memory, people who understand request-based autoscaling, and people who have never left a batch job. The deployment question separates them fast. Ask about the last rollout the candidate owned, how traffic shifted, and what the rollback trigger was. People who have run model deployment under load answer with numbers and timestamps; people who have watched it answer with feature names. Predictive serving and generative serving also diverge sharply on the ground: batch latency curves against token-generation throughput, and experience with the first does not quietly transfer to the second.
Feature stores separate training features from serving features
The oldest MLOps failure has a name in the tooling now. Sculley and his co-authors catalogued it in 2015: machine learning systems accumulate debt through boundary erosion, entanglement, and data dependencies that conventional code review never sees . The feature store exists to fix one slice of that debt, the training/serving skew where features computed one way offline get computed another way online. Feature stores materialize point-in-time-correct values, version the logic that produced them, and hand the same vectors to training and inference. Candidates split sharply here. People who have run one argue about backfills, TTLs, and point-in-time joins; people who have not describe it as a cache with tables. For teams whose models quietly disagree between batch and online scoring, that difference is the whole job, and the interview must test it.
ML pipelines became the system of record for reproducibility
A model without its pipeline is a rumor. MLflow 3, released in June 2025, made the point architectural: the LoggedModel entity groups runs, traces, prompts, and evaluation metrics into one lineage, and runs capture environment metadata such as the Git commit hash that produced them . The hiring question follows from that design. Reproducibility in MLOps is not a PDF of hyperparameters; it is a rerunnable DAG whose inputs, images, and seeds are pinned, and the candidate who built one can explain what made a rerun diverge. The candidate who cannot is describing a notebook. ML pipelines treated as infrastructure, versioned and tested like any other service, change how a team hires: it needs people who think in artifacts and provenance, not people who think in scripts.
Model versioning answers who shipped which weights when
Model versioning sounds like a solved problem because code versioning was solved decades ago. The two share a verb and almost nothing else. A model version must bind weights, preprocessing, the feature schema it consumed, the evaluation that passed it, and the deployment state that released it, and all of that must survive an audit months later. MLflow's registry tracks stages and lineage for that lifecycle , and MLflow 3 extends the picture so metrics and parameters from every run sit next to the registered version . Candidates who have operated this talk about promotion gates, shadow traffic, and the one version that shipped with a stale feature schema. Candidates who have not talk about Git tags. The gap matters the first time a compliance question arrives and nobody can prove which artifact was serving on a given date.
Model monitoring is now a telemetry convention problem
Monitoring production models used to mean a dashboard someone eventually looked at. It is now standardization work. OpenTelemetry's GenAI semantic conventions define attributes for prompts, responses, token usage, and provider metadata, and a Python instrumentation library for OpenAI calls ships the first implementation, so cost, latency, and safety signals carry the same names across stacks . That matters for hiring because model monitoring increasingly means wiring drift detectors, payload loggers, and cost meters into a telemetry pipeline rather than writing one more script. Ask a candidate what a drift alert actually fired on, what they logged to diagnose it, and what it cost. People who have carried model monitoring through an incident describe distributions, thresholds, and the incident; people who have not describe tool names.
Machine learning infrastructure claims fail without a rollback story
Verification in this seat is about reversibility. A candidate claiming machine learning infrastructure ownership should be able to describe the worst deployment they shipped, how they knew it was bad, and the exact sequence that put the previous model back. If the answer is a redeploy of the old container with a start time attached, the candidate has operated. If the answer is a plan for what they would have done, they have not. The stakes justify the probe. A mis-hire in machine learning operations consumes the scarcest resource on a platform team, senior engineers' hours explaining pods and autoscalers to someone still learning them, while bad deployments keep shipping and models drift unattended . This is the hiring implication behind every earlier challenge: MLOps assessment belongs to people who can tell a broken rollout from a fixed one, and the question set is cheap enough to ask every shortlist candidate.
References
- CNCF Announces Kubeflow's Graduation, Solidifying a Standard for Cloud Native AI Operations — Cloud Native Computing Foundation (CNCF). (accessed 2026-09-28)
- KServe joins CNCF as an incubating project — Red Hat. (accessed 2026-09-28)
- Hidden Technical Debt in Machine Learning Systems — NeurIPS (NIPS 2015 Proceedings). (accessed 2026-09-28)
- MLflow 3 Release Notes — MLflow Project. (accessed 2026-09-28)
- ML Model Registry — MLflow Project. (accessed 2026-09-28)
- OpenTelemetry for generative AI — Cloud Native Computing Foundation (CNCF). (accessed 2026-09-28)
