About you and your background
01Tell me about your background and the vision problems you have worked on.
What they are checking: The interviewer wants to map your experience across domains and see whether you have dealt with real sensors and environments.
Example answer
I studied electrical engineering and got into vision through a robotics lab project on visual odometry, which taught me camera geometry and how unforgiving real sensors are. My first job was at a retail analytics company, where I built people detection and tracking models that ran on in-store cameras, and I learned about domain shift the hard way when every new store had different lighting. After that I moved to a document understanding team, working on layout detection and OCR post-processing for scanned forms, which is a different set of problems: high resolution, small objects and structured output. Most recently I have worked on a defect inspection product for a manufacturing customer, where the data was scarce and imbalanced. Across those, the constant has been that the model is a fraction of the work; the data pipeline and the deployment environment decide whether it succeeds.
02Describe a vision model you took to production and what the deployment environment was.
What they are checking: They want to hear about the constraints of a real deployment, not just the training results.
Example answer
The people counting model ran on a small edge box attached to each store's cameras, with no reliable network connection, so the model had to run locally in real time on a low-power GPU. We trained a detector on a mix of public data and our own labeled frames, then distilled it into a lighter model and exported it to an optimized runtime with half-precision inference. The pipeline decoded video, ran detection every few frames with tracking in between, and aggregated counts locally before uploading summaries. The deployment lessons were about the environment: heat throttled the device in some stores, and we added a frame-rate fallback; firmware updates had to be staged; and we shipped a shadow mode that logged model output without acting on it for a week per store before enabling it. Accuracy in the lab was never the issue.
03How much of your work has been data and labeling versus modeling?
What they are checking: This reveals whether you understand where the leverage usually is in applied vision.
Example answer
Honestly, more than half data. On the inspection project the model choice took a week and the data took months: writing a labeling guide that distinguished cosmetic marks from real defects, running calibration sessions with annotators, measuring agreement, and building a review queue for disagreements. I built the tooling for active sampling, so annotators saw the frames the model was least sure about rather than random ones, which doubled the useful labels per hour. I also spent a lot of time on augmentation and on synthetic data for rare defect types, which helped the recall on those classes without inventing patterns the model then hallucinated in production. I still enjoy the modeling, but I have learned that an engineer who can raise label quality by a meaningful amount moves the product more than one who tunes an architecture.
04Which vision architectures have you used, and how do you choose between them?
What they are checking: They want a decision process driven by constraints and benchmarks rather than familiarity or fashion.
Example answer
I have used convolutional backbones for most production work, vision transformers where there was enough data or a good pretrained checkpoint, and one-stage detectors for real-time detection. For segmentation I have used encoder-decoder models and, more recently, promptable segmentation models for labeling assistance. My choice comes from the constraints, in this order: where the model runs, which sets the compute budget; how much labeled data exists, which decides whether a large pretrained transformer or a smaller convolutional model will generalize better; and the latency target. Then I benchmark two or three candidates on our validation set rather than trusting public leaderboards, because the domain gap is usually larger than the gap between architectures. I try to start from a pretrained checkpoint and only train from scratch when the domain is genuinely unlike natural images.
05What is a vision problem you found harder than expected?
What they are checking: This checks for humility and whether you learned something transferable from a hard problem.
Example answer
Reading gauges on industrial equipment from a fixed camera. It sounded like a solved problem: detect the dial, find the needle, compute the angle. In practice, glare moved with the time of day, condensation appeared in winter, the needle was thin enough that motion blur erased it, and the customer had a dozen gauge designs. A generic detector plus geometry worked in the lab and failed on site. What finally worked was much less elegant: per-gauge calibration at install time, a small keypoint model trained on augmented images that simulated glare and blur, and a confidence gate that flagged uncertain readings for a human rather than guessing. The lesson was that a fixed camera does not mean a fixed distribution, and that a system that knows when it cannot read is more valuable than one that reads slightly better on average.
Vision models and pipelines
06Explain how a modern object detector works, and how you would evaluate it.
What they are checking: A fundamentals check on detection architectures, losses and metrics, including what a single mAP number hides.
Example answer
A modern detector takes an image through a backbone that produces multi-scale feature maps, then a neck that fuses those scales, then a head that predicts, for each location or query, a class distribution and box coordinates. One-stage detectors predict directly from dense anchors or anchor-free points; transformer-based detectors use learned object queries matched to ground truth with a bipartite assignment, which removes the need for non-maximum suppression. Training uses a classification loss, often focal loss to handle background imbalance, and a box regression loss such as an IoU-based loss. Evaluation is mean average precision across classes at one or more IoU thresholds, but I also report per-class recall at the operating confidence, small-object performance, and latency on the target hardware, because a single mAP number can hide the failure the product actually cares about.
07Design a system to inspect products on a manufacturing line for defects using cameras.
What they are checking: This tests end-to-end design under physical constraints: optics, scarce labels, line speed, human review and drift.
Example answer
I would start on the line with the quality team to understand the defect types, their frequency and the cost of a miss versus a false reject, since those set the operating point. Cameras and lighting come first: fixed mounts, controlled illumination and a resolution that makes the smallest defect a few pixels wide. For a new line with few defect examples I would begin with an anomaly detection approach trained on good parts, which catches unknown defect types, and add a supervised classifier as labeled defects accumulate. Inference runs on an edge device at line speed with a hard latency budget. Every flagged part goes to a review station whose decisions feed back into training. I would monitor the reject rate, the reviewer overrule rate and image statistics to detect lighting or camera drift, and stage retrained models with a shadow period before they can reject parts.
08Your model performs well on the test set but fails on images from a new camera. What do you do?
What they are checking: They want a systematic approach to domain shift, from diagnosis through short-term and structural fixes.
Example answer
That is domain shift, and the fix starts with diagnosis. I would compare image statistics between the training set and the new camera: resolution, color balance, noise, compression, focal length and field of view. Often the new camera is sharper or has a different white balance, and a preprocessing step to normalize resolution and color, plus augmentation that spans the difference, recovers most of the gap. If that is not enough, I would label a small set from the new camera, first to measure the failure precisely and then to fine-tune. Test-time augmentation and calibrating the confidence threshold per camera help in the short term. Longer term I would add camera metadata to the pipeline so we track performance per device, and make the acceptance test for any new camera include a labeled evaluation set from that camera before rollout.
09How do you deploy a vision model on an edge device with limited compute?
What they are checking: This checks practical knowledge of quantization, operator support, distillation and on-device runtime concerns.
Example answer
The first step is to know the target precisely: the chip, the memory, the supported operators and the latency and power budget. Then I pick a model family that fits, usually a small efficient backbone, and train at the input resolution the device can afford rather than downscaling afterward. I export to the vendor runtime, applying post-training quantization to eight-bit integers and checking accuracy on the validation set; if it drops too far, I use quantization-aware training. Operator support is the usual surprise, so I profile the exported graph on the actual device early, not at the end. Knowledge distillation from a larger model recovers accuracy that the small model loses. Then I build the runtime around it: frame skipping, a fallback for thermal throttling, and a way to update models in the field with a rollback. I benchmark on the device under realistic load.
10When would you use a vision-language model instead of a task-specific model?
What they are checking: They want a clear decision rule that reflects current practice, including using both together.
Example answer
A vision-language model is my choice when the task is open-ended or changes often, when labeled data is scarce and the model's pretrained knowledge covers the domain, or when the output is naturally text, such as describing a scene or extracting fields from a varied set of documents. It is also useful for bootstrapping: generating candidate labels that humans correct. A task-specific model wins when the task is fixed and well defined, latency and cost matter, the domain is far from natural images, or the output must be precise geometry like boxes and masks at high recall. In practice I often use both: a vision-language model to explore and label, and a small specialized model in production. I would not put a large multimodal model on a real-time edge pipeline, and I would not train a detector for a task that will be redefined next month.
Behavioral and teamwork
11Tell me about a time labeling quality caused a problem and how you fixed it.
What they are checking: They want to see that you diagnose data problems with measurement and fix the process, not just the labels.
Example answer
On the inspection project, recall on one defect class was stuck below target and nothing in the model changed it. I suspected labels. I sampled a hundred images from that class and had two annotators relabel them blind, and agreement with the original labels was poor: the class definition in the guide was ambiguous between a scratch and a scuff. The task was to fix the labels without stalling training. I rewrote the guide with example images for the boundary cases, ran a one-hour calibration session, and had the ambiguous class relabeled with a review step. Then I added an ongoing agreement check on a sample of each batch. After relabeling, the same model architecture cleared the recall target, and the agreement metric became a standard part of our data pipeline reports.
12Describe working with hardware or firmware engineers on a camera pipeline.
What they are checking: This checks whether you can collaborate across the hardware boundary and make interfaces testable.
Example answer
For the edge counting product, the firmware team owned the camera pipeline and video decode, and I owned the model. Early on, my model received frames that had been resized and compressed differently from my training data, which cost accuracy, and neither side had noticed. I set up a weekly session with the firmware lead and we wrote a short interface contract: frame resolution, color format, timestamp source and what happened when frames dropped. I gave them a test harness with a set of reference frames and expected model outputs so they could verify any pipeline change did not shift results. They gave me a device to test on rather than a simulator. The collaboration worked because we treated the boundary as something to test, not assume, and both sides had a fast way to check it.
13Tell me about a time you had to cut scope to ship a vision feature.
What they are checking: They want to see judgment about what to cut, backed by evidence, and how you communicated it.
Example answer
We had committed to shipping a document capture feature that detected the page, corrected perspective, classified the document type and extracted fields, in eight weeks. Six weeks in, extraction accuracy on handwritten fields was far below target. I proposed cutting scope rather than slipping: ship detection, correction and classification, with extraction for printed fields only, and show handwritten fields to the user to type. I brought the accuracy numbers by field type and a mockup of the fallback to the product manager, and we agreed the reduced version still removed most of the manual effort. We shipped on time, users adopted it, and the handwriting data we collected from the fallback became the training set for the next version, which shipped a quarter later. Cutting scope early with evidence is easier than explaining a slip.
14Describe a disagreement about a metric or acceptance threshold.
What they are checking: This tests whether you can turn a subjective argument into a decision grounded in costs.
Example answer
The product manager wanted the defect detector to hit a very high recall before launch, since misses were expensive. I agreed misses mattered, but at that recall the false reject rate would have stopped the line several times an hour, and the customer would have turned the system off. We disagreed about the threshold for a week. I broke the deadlock by turning it into a cost question: I estimated the cost of a miss and the cost of a false reject from the customer's own numbers, plotted the expected cost across the operating curve and showed the threshold that minimized it. It was well below the recall the PM had asked for, but it kept the line running and still caught most defects. We launched there with a review station for borderline cases and agreed to revisit the threshold monthly as the model improved.
15Tell me about a time you improved a team's experimentation process.
What they are checking: They are looking for influence without authority and a change that stuck because it was useful.
Example answer
When I joined the vision team, experiments were run from personal scripts and results lived in chat messages, so nobody could reproduce a number from a month before. I did not have authority to mandate anything, so I started with my own experiments: a config-driven training script, a tracked run for every experiment with the dataset version, and an evaluation report generated the same way every time. I shared the template and offered to convert one project per teammate. Within two months the whole team was using it, mostly because comparing results became trivial. The concrete payoff came when a customer asked why a model had changed behavior and we could point to the exact dataset and config difference in a few minutes. Later we added a weekly review of tracked results, which replaced a lot of ad hoc debate.
Your fit and the role
16Why this company and this domain?
What they are checking: They want evidence that you understand the specific vision problem here and chose it on purpose.
Example answer
Because the vision problem here is central to the product rather than a feature bolted onto it, and because it is hard in the ways I find interesting: real hardware, real environments and a cost to being wrong. I have worked in retail and manufacturing settings where the camera is in a hostile place and the model has to earn trust from people who can switch it off, and your domain has the same shape. I also read what your team has published on its data pipeline, and the emphasis on labeling quality and per-device monitoring matches how I think the work should be done. Finally, the scale is right for me: enough deployments that engineering discipline matters, small enough that one engineer can still see the whole path from camera to decision.
17What would you want to learn in the first three months here?
What they are checking: This checks whether your onboarding plan prioritizes data, deployment and customers before modeling changes.
Example answer
Three things. First, your data: where it comes from, how it is labeled, what the annotators struggle with and what the evaluation sets actually cover. I would want to label a batch myself in the first week. Second, the deployment environment, including the hardware, the update mechanism and the monitoring, because that constrains every modeling decision. I would want to sit with whoever handles field issues and read the last few months of incidents. Third, the customers' definition of success, since accuracy metrics and what a customer notices are rarely the same thing. By the end of three months I would expect to have shipped one measurable improvement, to know which failure modes cost the most, and to have a view on where the next year of modeling and data investment should go.
18How do you feel about a role that includes on-call for a production vision system?
What they are checking: They want to know whether you will treat operational ownership as part of the job or as a burden.
Example answer
I am fine with it, and I think it makes for better engineers. A production vision system fails in ways that no test set predicts: a camera gets bumped, a lighting fixture is replaced, a firmware update changes the color pipeline. Being on the receiving end of those pages is the fastest way to learn what monitoring is missing and what the model should have been robust to. What I would want is a reasonable rotation, runbooks that exist, and a culture where each page turns into a fix or a monitoring improvement rather than a repeat. I have been on call before for the edge counting product and my main contribution was reducing the pages, by adding drift detection on image statistics so we caught camera changes before customers did. That is the attitude I would bring.
19How do you balance accuracy against latency and cost when the product manager wants all three?
What they are checking: This tests whether you can make tradeoffs explicit and collaborative rather than adversarial.
Example answer
I make the tradeoff explicit and measurable rather than arguing about it. I put the candidate models on a chart with accuracy on the validation set against latency and cost on the target hardware, and I mark the product's hard limits: the latency the pipeline cannot exceed, the budget per device. Usually one or two options are clearly dominated and the real choice is between two points on the frontier. Then I ask the product manager which failure hurts users more, a slower response or a wrong one, and I frame the choice in those terms. Often the answer is to buy accuracy in a different currency, for example a smaller model plus a confidence gate that sends hard cases to a slower path. When all three are demanded, the honest response is to show the frontier and ask which one moves.
20Where do you want your career to go?
What they are checking: They are checking that the role is a real step toward what you want, which predicts retention.
Example answer
I want to become the engineer who owns the full vision system for a product with real deployments: the data strategy, the model, the edge and cloud pipeline and the monitoring, and who can be trusted to say what is achievable before a customer commitment is made. In the near term that means going deeper on efficient inference and on data engineering for vision, since those are where I see the most leverage. I am also interested in how vision-language models change labeling and evaluation, and I want to be hands-on with that as it matures. I am not aiming for management soon; a staff-level technical path suits me. This role fits because it has ownership of a production system and a domain I would happily work in for years.
Questions worth asking them
- Where do the models run today, in the cloud or on devices, and what is the update path?
- How is labeling done, who does it, and how do you measure label quality?
- What does a typical production failure look like, and how quickly do you find out?
- How much of the team's time goes to data work versus modeling versus deployment?
- What hardware constraints or latency targets should I expect to design around?
How to prepare
- Refresh the fundamentals: convolutions, receptive fields, detection and segmentation losses, IoU and mean average precision, and how vision transformers differ.
- Prepare a data story: a labeling problem you found, how you measured it and what it did to the model.
- Be ready to design a camera-to-decision pipeline on a whiteboard, including lighting, resolution, edge versus cloud inference and drift monitoring.
- Practice an image processing coding problem in Python with numpy, since many loops include one under time pressure.