About you and your background
01Tell me about your background and how you moved into ML platform work.
What they are checking: The interviewer wants to understand whether you came from infrastructure or from modeling and how that shapes your instincts.
Example answer
I started as a backend engineer on a data-heavy product, running the services and the batch jobs that fed them. When a data science team joined, their models kept failing in production for reasons that had nothing to do with modeling: a missing dependency, a feature computed differently in the notebook, a job that silently used last week's data. I ended up as the person who fixed those problems, and I realized it was a job. I moved into a platform role, where I built the training pipeline orchestration, a model registry and a deployment path with automated checks for a team of about a dozen data scientists. More recently I have worked on feature infrastructure and monitoring. I like the work because the leverage is high: one good platform decision removes a category of failures for every team that uses it.
02Describe the ML platform you know best and what you would change about it.
What they are checking: They want a clear mental model of a real platform and the judgment to see its weaknesses.
Example answer
The platform I know best was built around a workflow orchestrator for training pipelines, a model registry with metadata for every run, a feature store with batch and online paths, and a serving layer on Kubernetes with canary rollouts. It worked well for the standard case: a scheduled retrain, evaluation against a held-out set, promotion if metrics passed, and a gradual rollout with automatic rollback on latency or error budget breaches. What I would change is the parts that grew ad hoc. Monitoring was bolted on per model rather than defined as part of the deployment, so coverage was uneven. The feature store had no ownership model, so stale features accumulated. And local development was painful, which pushed people back toward notebooks. If I built it again, I would make monitoring and ownership mandatory at registration time and invest early in the local loop.
03What is your experience with the infrastructure side, such as Kubernetes, cloud services and infrastructure as code?
What they are checking: A practical depth check on the infrastructure skills the role will need every day.
Example answer
I have run production workloads on Kubernetes for several years: writing the manifests and Helm charts, setting resource requests and limits for GPU workloads, node pools for different accelerator types, and autoscaling for inference services. I manage cloud resources with Terraform and I treat the platform's own configuration as code with review and CI, since a platform team that clicks through consoles cannot expect its users to do better. I have worked on two major clouds and I am comfortable with their managed training and serving services, object storage, IAM and networking. I have also set up observability with metrics, logs and traces for model services, and cost allocation by team. Where I am weaker is deep networking and bare-metal GPU cluster management; I have worked alongside people who own those and I know enough to be a good partner.
04How do you work with data scientists who are not strong software engineers?
What they are checking: This reveals whether you design for your users or blame them, which decides whether a platform gets adopted.
Example answer
I treat it as a design constraint, not a complaint. Most data scientists are excellent at their job and do not want to become infrastructure engineers, so the platform should let them do the right thing with the least ceremony. In practice that means a project template with tests, packaging and a deploy path already wired; a local command that runs the same pipeline that runs in CI; and clear, short documentation with examples. I pair with people on their first deployment and I review pull requests with the aim of teaching one thing per review, not everything. I also make the platform give fast, specific feedback: a failing check that says exactly what to fix is worth more than a wiki page. When a team consistently struggles with something, I take that as a signal about the platform before I take it as a signal about the team.
05What is a reliability incident you handled, and what changed afterward?
What they are checking: They want to see calm incident handling and structural fixes that prevent the same class of failure.
Example answer
A nightly retrain for a ranking model started producing a model that scored every item nearly identically, and the automated evaluation still passed because the metric threshold was too loose. The degraded model reached a canary and click-through dropped visibly within an hour. I was on call. I rolled back to the previous model version, which took a few minutes because the registry kept the artifacts. Then I traced the cause to an upstream table whose partition had arrived empty, so the training set contained only a fraction of the usual data. The fixes were three: a data volume check before training, a stricter evaluation gate that compared against the current production model rather than a fixed threshold, and a canary metric on the business signal, not just latency. We wrote a postmortem and the pattern became a checklist item for every pipeline.
Platforms and operations
06Design a CI/CD pipeline for a model, from a merged change to a safe production rollout.
What they are checking: This tests whether you can describe every gate between code and production and what each one protects against.
Example answer
On a merged change, CI runs unit tests for the feature code and the training code, plus a smoke training run on a fixed sample. If the change affects the model, a full training job runs in the pipeline with pinned data and config versions, producing an artifact registered with its lineage. An evaluation stage compares the candidate against the current production model on a held-out set and on slices, failing if any guardrail metric regresses. Passing models are packaged into a versioned container with the exact preprocessing code. Deployment goes to a canary with a slice of live traffic, and business and system metrics are compared against the incumbent for a defined period. Promotion is automatic if the comparison passes and rollback is automatic on breach. Every step writes to an audit log, and a human approval gate can be inserted for high-risk models.
07What is a feature store, and when is it worth the complexity?
What they are checking: They want a precise definition and, more importantly, judgment about when not to build one.
Example answer
A feature store is a system that defines features once, computes them for training from historical data and for serving from fresh data, and guarantees the two paths produce the same values. It typically has an offline store for point-in-time correct training sets and an online store for low-latency lookup at inference time. It is worth the complexity when several models share features, when training-serving skew has actually caused incidents, when features need point-in-time correctness to avoid leakage, or when online features must be fresh within seconds. It is not worth it for a team with one or two batch models that score from a warehouse table; a well-organized set of tables and a shared transform library does the job. When I introduce one, I start with the offline path and a small number of high-value shared features, and add the online store only for models that need it.
08How do you monitor a model in production, and what do you alert on?
What they are checking: This checks that you separate system, data and model health and know which signals deserve a page.
Example answer
I monitor at three levels. System: latency percentiles, error rate, throughput, resource use and the health of upstream dependencies, alerted like any service. Data: the distribution of each input feature compared against the training distribution, null rates, schema changes and freshness of feature sources, with alerts on significant drift or a broken pipeline. Model: the distribution of predictions, calibration where labels arrive, and the business metric the model exists to move, such as conversion or defect catch rate, compared against a baseline. When labels are delayed, I track prediction drift as an early signal and compute accuracy retroactively when labels land. I alert on things that need a person now, such as a broken input or a sharp prediction shift, and I put slower signals like gradual drift on a dashboard reviewed weekly. Every alert links to a runbook, and every model has an owner.
09How do you make training runs reproducible?
What they are checking: They want the full list of what must be pinned, and honesty about what cannot be made fully deterministic.
Example answer
Reproducibility means that given a run id, I can recreate the model, or at least explain every difference. I pin the code at a commit, the dependencies with a lockfile and a container image, and the data with a versioned snapshot or an immutable query with a point-in-time filter. Configuration lives in a file that is logged with the run, not in command-line flags someone typed. Random seeds are set and recorded, with the caveat that GPU nondeterminism means results may differ slightly, so I document the expected variance. Every run logs its parameters, metrics, artifacts and lineage to a tracking system, and the model registry links back to the run. I test reproducibility periodically by re-running a registered model's training from its metadata and comparing the evaluation. The hardest part is data: if the source tables mutate in place, nothing else matters, so snapshots come first.
10How would you serve models from several teams on shared infrastructure without them interfering with each other?
What they are checking: This tests multi-tenancy thinking: isolation, quotas, per-model rollouts and safe defaults.
Example answer
Isolation and fairness are the two problems. For isolation, each team's models run in their own namespace with resource requests and limits, so a runaway model cannot starve another, and GPU workloads are scheduled onto node pools with quotas per team. Each model is a separate deployment with its own autoscaling rules and its own error budget. For fairness, I set quotas and priority classes, and I expose cost and utilization per team so the allocation is visible. Shared components, such as the gateway, the feature store and the model registry, are multi-tenant with authentication and per-team access control, and I load test them for noisy-neighbor behavior. Rollouts are per model, so one team's bad deploy cannot take down another's. The platform provides defaults that make the safe path the easy one, and teams that need something unusual get a bounded exception rather than a shared change.
Behavioral and teamwork
11Tell me about a time you migrated teams onto a new platform or tool.
What they are checking: They want to see how you drive adoption without stopping the teams you are migrating.
Example answer
We needed to move about ten data science teams from ad hoc cron jobs and notebooks onto a managed pipeline platform, without stopping their work. My role was to lead the migration. I started by migrating one friendly team end to end and writing the guide from what actually went wrong, not from the ideal path. Then I ran office hours twice a week and made a dashboard showing each team's migration status, which created some healthy pressure. The key decision was to support both systems in parallel for a quarter with a firm cutoff date announced early, rather than a hard switch. Two teams needed features the platform lacked and we built those before their deadline. By the cutoff, every team had moved, failed jobs dropped substantially, and the on-call load for the platform team went down because the failure modes became uniform.
12Describe a time you had to say no to a team's request.
What they are checking: This checks whether you can hold a line on safety or reliability while still solving the requester's problem.
Example answer
A team asked us to give their training jobs unrestricted access to the production database so they could pull fresh data on demand. It would have solved their immediate problem and created a security and reliability risk for everyone. I said no, but I did not leave it there. I asked what they actually needed, which was data no older than an hour, and I proposed a replicated table in the warehouse with an hourly refresh, which our data engineering team could set up in a week. I wrote down the reasoning, including the incidents that had happened elsewhere from direct production access, and shared it with their lead. They were annoyed for a week and then fine. The replicated table became a pattern other teams used. Saying no works when it comes with a real alternative and a clear reason.
13Tell me about a deployment that went wrong and what you did.
What they are checking: They want a fast, calm rollback story followed by a fix that closes the gap that let it happen.
Example answer
We rolled out a new version of the serving gateway that changed how request timeouts were handled. It passed staging. In production, one high-traffic model with slow preprocessing started timing out under load, and its fallback was misconfigured, so users saw errors. I was on the rollout. I rolled the gateway back within ten minutes, which stopped the errors, and then worked with the model's owner to understand why the new timeout behavior affected them and not others. The root cause was that staging had no model with comparable latency, so the behavior was never exercised. We fixed the fallback configuration, added a synthetic slow model to staging that mirrors the worst production latency, and made gateway changes roll out per model rather than globally. The postmortem was blameless and the changes closed the gap that let it happen.
14Describe working with security or compliance on an ML system.
What they are checking: This reveals whether you engage compliance early and build controls into the platform rather than around it.
Example answer
When a model began using customer data that fell under a stricter regulation, the compliance team needed evidence of where the data went, who could access it and how long it was retained. I set up a session early rather than treating them as a gate at the end. We mapped the lineage from source table through the feature pipeline, training artifacts, evaluation sets and logs, and found two places where raw fields were being logged unnecessarily. I removed those, added field-level tagging so sensitive columns were automatically excluded from logs and from the offline feature store, and set retention policies on training artifacts. I also gave the compliance team a read-only view of the model registry with lineage, which answered most of their questions without a meeting. The audit passed and the tagging system became the default for every new pipeline.
15Tell me about a time you reduced cost without hurting the teams using the platform.
What they are checking: They want to see measurement before action and changes that respect the users of the platform.
Example answer
GPU spend had grown steadily and leadership wanted it cut. Rather than imposing quotas, I started with visibility: a dashboard of utilization and cost per team and per job. It showed that a large share of GPU hours went to jobs running at low utilization, mostly because of oversized instance types and interactive notebooks left running. I added idle detection that shut down notebooks after a period of inactivity with a warning, made right-sized instance types the default in the job template, and introduced spot instances with checkpointing for training jobs that could tolerate interruption. I talked to each team before anything changed. Spend fell by roughly a third within two months, no team lost capacity they were actually using, and a couple of teams found their jobs ran faster because they stopped waiting for scarce large instances.
Your fit and the role
16Why platform work rather than building models yourself?
What they are checking: They want to confirm you are drawn to platform work for its own leverage, not as a fallback.
Example answer
Because I care more about whether models work in production than about which model wins, and I have found that the failures that matter most happen in the plumbing. On a platform team I can remove a category of failures for every team at once, which is more leverage than I would have on any single model. I also enjoy the customer relationship: my users are data scientists and ML engineers, they are demanding, and they tell you immediately when something is bad. I have built models and I am glad I did, because it means I understand what my users are trying to do and I do not build platforms that look tidy and are unusable. But the part of the job that energizes me is making the reliable path the easy path.
17What would you do first if you joined this team?
What they are checking: This checks for a listening-first onboarding plan and restraint about re-architecture.
Example answer
I would spend the first few weeks listening and measuring. I would talk to every team that uses the platform and ask what slows them down and what they distrust, and I would look at the data: failed jobs, time from merged change to production, incidents in the last six months and where on-call time goes. Then I would pick the single most painful, well-understood problem and fix it in a way that is visible, so the teams see the platform improving and I learn the codebase by changing it. I would avoid proposing a re-architecture in the first quarter; those usually come from people who have not yet learned why the current system is shaped the way it is. By the end of the first quarter I would have a written assessment and a prioritized plan the team and its users agree with.
18How do you decide between building a platform capability and buying a managed service?
What they are checking: They want a decision framework that weighs differentiation, hard requirements, exit cost and team capacity.
Example answer
I lean toward buying for anything that is not a differentiator and that a managed service does well: orchestration, experiment tracking, object storage, basic model serving. I build when the managed option cannot meet a hard requirement, such as data residency, latency, a specific hardware need or an integration with internal systems that the vendor cannot provide, or when the cost at our scale clearly favors it. I also consider the exit cost: a managed service behind a thin internal interface is easy to replace later, while one wired directly into every pipeline is not. And I weigh the team's capacity honestly; a platform team of four that builds its own feature store will spend a year maintaining it. My default is to buy with an abstraction, measure, and build only the pieces where the numbers or the requirements demand it.
19How do you measure whether a platform team is succeeding?
What they are checking: This reveals whether you judge the platform by its users' outcomes rather than by features shipped.
Example answer
By whether the teams using it ship faster and break less, measured rather than asserted. Concretely: time from a merged model change to production, the share of deploys that roll back, the number of incidents attributable to the pipeline or serving, and the time data scientists spend on infrastructure work, which I would survey quarterly. I would also watch adoption, since a platform nobody uses is a failure regardless of its quality, and platform cost per model or per prediction. Qualitative signals matter too: whether teams come to us early in a project or route around us, and whether our on-call load is falling. What I would not measure is the number of features shipped by the platform team, because it rewards building things nobody asked for. A good platform team is judged by its users' outcomes.
20Where do you want to be in a few years?
What they are checking: They are checking that your direction fits the growth available on this team.
Example answer
I want to lead the technical direction of an ML platform at a company where models are core to the product: setting the architecture, the reliability standards and the practices for how models are built and operated, and helping a team of platform engineers grow into that. That could be a staff engineer role or a lead role over a small team; I am open to either depending on what the organization needs. In the shorter term I want to go deeper on serving infrastructure for large models, since that is where cost and latency pressure is highest right now, and on evaluation systems for generative models, which most platforms handle poorly. This role fits because it is hands-on on a platform with real users and real production stakes, at a stage where the decisions made now will shape the next several years.
Questions worth asking them
- How many models are in production today, and how long does it take a change to reach them?
- What does the platform team own versus what the model teams own, and where is that boundary contested?
- What were the last two production incidents involving a model, and what changed afterward?
- Which parts of the stack are built in-house and which are managed services, and why?
- How is the platform team's success measured, and by whom?
How to prepare
- Prepare a system design answer for the full model lifecycle, from data validation through CI/CD, canary rollout, monitoring and rollback, and practice drawing it in twenty minutes.
- Have one incident story ready with the timeline, the root cause and the structural change that followed.
- Refresh Kubernetes, containers, infrastructure as code and one workflow orchestrator well enough to answer implementation questions, not just concepts.
- Be ready to explain training-serving skew, feature stores, drift detection and reproducibility with specific examples from your own work.