About you and your background
01Tell me how you got into language work and what you have shipped.
What they are checking: The interviewer wants to know whether your experience is production-grade or limited to demos and notebooks.
Example answer
I came in through search. My first job was on a support site search team, where I worked on query understanding: spelling correction, synonym expansion and a classifier that routed queries to the right product area. That taught me how much of language work is evaluation and data, not models. When transformer-based models became practical, I moved to a team building a customer support assistant, where I owned retrieval and the evaluation harness. Most recently I have been an LLM engineer on a document processing product, shipping extraction pipelines that turn contracts into structured fields with a review step for low-confidence output. Across all of it I have shipped things that real users depend on, and the recurring lesson is that the model is rarely the bottleneck; the data, the retrieval and the evaluation are.
02Which LLM-based system have you built that reached real users, and what broke?
What they are checking: They want evidence of production experience and a candid account of failure modes you have actually seen.
Example answer
The support assistant reached a few thousand users a day. It answered questions over a help center using retrieval plus a hosted model. The first thing that broke was retrieval on short queries: a one-word query like refunds pulled in every article containing the word and the model synthesized a confident answer from the wrong plan's policy. We fixed it with a hybrid retriever combining lexical and dense search, query rewriting for short inputs, and metadata filters on the user's plan. The second thing that broke was silent: a help center rewrite changed article structure, our chunking produced fragments, and quality dropped for a week before a user complained. That led to an automated evaluation set that ran on every index rebuild. Both failures were in the plumbing around the model, which is where I now look first.
03Have you fine-tuned a model, and how did you decide it was worth it over prompting?
What they are checking: This checks whether you make the prompt-versus-fine-tune decision from measurement rather than habit.
Example answer
Yes, twice, and both times the decision came from measurement rather than preference. For a classification task with a fixed label set and tens of thousands of labeled examples, a fine-tuned small model beat the best prompt on accuracy, ran far cheaper per request and had stable latency, so the choice was clear. For an extraction task the picture was different: prompting a large model with a few examples got us most of the way, and fine-tuning gained a few points on the validation set but locked us into a model version and added a training pipeline to maintain. We stayed with prompting there and invested in better examples. My rule is that fine-tuning is worth it when the task is narrow, the labels are plentiful and cost or latency matters; otherwise prompting plus retrieval is easier to keep current.
04How do you stay current when models and APIs change every few months?
What they are checking: They want a durable learning method that separates fast-changing details from lasting skills.
Example answer
I separate what changes from what does not. Model names, context limits and pricing change constantly, so I do not memorize them; I keep an evaluation harness that I can rerun against a new model in an afternoon, and I let the numbers tell me whether a change matters for my task. What changes slowly is the underlying craft: how to build a good eval set, how retrieval fails, how to structure output so it can be validated, how to handle a model that is confidently wrong. I read a few papers a month, mostly on evaluation and retrieval, and I follow release notes from the providers we use. I also try to build one small thing with each major new capability, because a day of hands-on work tells me more than a week of reading about it.
05What is your experience with classical NLP, and does it still matter?
What they are checking: This reveals depth beyond prompting and whether you know when a small classical tool beats a large model.
Example answer
I have built tokenizers, named entity models with conditional random fields, and a lot of TF-IDF and BM25 search before dense retrieval existed. It still matters in three places. First, lexical retrieval remains a strong component in hybrid search, especially for exact terms like product codes and names that embeddings blur. Second, the discipline of classical evaluation, with precision and recall on a held-out set, transfers directly to LLM work and is often missing on teams that came in through prompting. Third, a small classical model is frequently the right tool for a narrow task like language detection or routing, where an LLM call is slower and more expensive for no gain. Nobody needs to hand-build a parser today, but knowing what the old tools do well keeps me from reaching for the largest model by default.
Language models and retrieval
06Explain how attention works in a transformer and why it matters for long inputs.
What they are checking: A fundamentals check that also tests whether you understand the cost implications of context length.
Example answer
Attention lets each token in a sequence compute a weighted combination of every other token's representation. Each token produces a query, a key and a value vector; the score between a query and each key, scaled and passed through a softmax, becomes the weight on the corresponding value. Multiple heads learn different patterns in parallel, and stacking layers lets the model build up context-dependent meaning. The reason it matters for long inputs is cost: self-attention is quadratic in sequence length, both in compute and, without tricks, in memory for the key-value cache. That is why long-context serving relies on efficient attention kernels, cache management and sometimes sparse or windowed variants, and why a model with a large context window can still lose information from the middle of a long prompt. I treat context length as a budget, not a free resource.
07Design a retrieval-augmented generation system for answering questions over internal documents.
What they are checking: They want to see the whole pipeline, including access control, chunking, hybrid retrieval, grounding and evaluation.
Example answer
I would begin with the documents: their formats, how often they change, and who is allowed to see what, since access control has to be enforced at retrieval time, not by the model. Ingestion parses each document into chunks that respect structure, such as sections and tables, attaches metadata like source, date and permissions, and writes embeddings plus a lexical index. Retrieval is hybrid, with a reranker on the top candidates and filters from the user's permissions. The prompt includes the retrieved passages with citations and instructions to say when the answer is not in the sources. I would return citations so users can verify. Evaluation uses a set of real questions with graded answers, measuring retrieval recall separately from answer quality. Monitoring tracks retrieval scores, refusal rate and user feedback, and the index rebuilds on document change.
08How do you evaluate the quality of generated text when there is no single correct answer?
What they are checking: Evaluation is the hardest part of LLM work and they want a layered, calibrated approach rather than a single score.
Example answer
I use layers of evaluation and I am explicit about what each one measures. For tasks with reference answers I use task-specific automatic metrics, but for open-ended generation I rely mainly on rubric-based grading: a written rubric describing what a good answer must contain and must avoid, applied by human raters on a sample and by a model grader at scale, with the model grader calibrated against the human labels and rechecked whenever the prompt changes. Pairwise comparisons between two system versions are more reliable than absolute scores. I track specific failure types, like unsupported claims or missing citations, rather than a single quality number. And I keep a small set of hard cases that must pass before any release. The evaluation set grows from production failures, which keeps it honest about what users actually ask.
09A RAG system is giving confident wrong answers. How do you find out where it fails?
What they are checking: This checks whether you debug the pipeline stage by stage instead of guessing at the prompt.
Example answer
Confident wrong answers usually mean the model is generating from insufficient or incorrect context, so I trace the pipeline in stages. First I take a sample of the bad answers and check retrieval: were the right passages in the index at all, were they retrieved, and where did they rank. If they were missing, the problem is ingestion or chunking. If they were present but ranked low, it is the retriever or the query. If they were retrieved, I look at the prompt: is the context truncated, is the instruction to stay grounded clear, are conflicting passages included. Then I check whether the model ignores the context, which a grounding evaluation can quantify. I fix the earliest failing stage, add those cases to the evaluation set, and make sure the system can say it does not know rather than fill the gap.
10How do you reduce cost and latency of an LLM feature without hurting quality?
What they are checking: They want practical levers, such as caching, routing and distillation, tied to a measured quality floor.
Example answer
I measure first, because the cost is usually concentrated. Long system prompts and repeated context are common culprits, and prompt caching or trimming instructions can cut cost without touching quality. Then I look at routing: a smaller model handles the easy majority of requests and a larger one handles the cases a classifier or a confidence score flags as hard. For latency I stream output, cap the number of retrieved passages, and run retrieval and any tool calls in parallel where the dependencies allow. Caching responses for repeated queries helps for some products. Where the task is narrow, a fine-tuned small model can replace a large one entirely. Every change goes through the evaluation harness with a defined quality floor, and I watch production metrics after rollout, because offline evaluation misses some regressions.
Behavioral and teamwork
11Tell me about a time an LLM feature behaved unexpectedly in front of users.
What they are checking: They want to see how you contain a live failure and turn it into a durable improvement.
Example answer
We shipped a summarization feature for customer conversations, and within a day an account manager reported that a summary described a customer as threatening to cancel when the transcript showed a joke. The task was to contain it and understand the failure. I pulled the flagged examples and found the model over-weighted certain phrases when transcripts were long and the relevant context was far from the summary instruction. Short term, we added a confidence gate that hid summaries for long transcripts and surfaced a link to the transcript instead. Then I rebuilt the prompt to summarize in sections with quotes, added a grounding check that verified each claim against the transcript, and added twenty real transcripts with known sentiment to the evaluation set. The feature was fully restored in two weeks with a lower error rate than before launch.
12Describe a disagreement about whether to use an external model API or self-host.
What they are checking: This tests whether you resolve architecture debates with numbers and leave a path to revisit the decision.
Example answer
An engineering lead wanted to self-host an open-weights model to avoid vendor dependence; the product manager wanted the hosted API for speed. I was asked to recommend. Instead of picking a side, I ran both on our evaluation set with realistic traffic and priced them against our projected volume, including the engineering time to run inference infrastructure. The hosted model was better on quality and cheaper at our current volume; self-hosting would become cheaper at roughly ten times the traffic and gave us control over data residency, which mattered for one customer segment. I recommended the hosted API behind an abstraction layer so we could switch, with a scheduled review when volume or the data residency requirement changed. Both sides accepted it because the decision was tied to numbers and had a defined revisit point.
13Tell me about a time you had to explain the limits of a model to a non-technical stakeholder.
What they are checking: They are checking whether you can make model behavior concrete for people who make decisions about it.
Example answer
A sales leader wanted the assistant to answer pricing questions for enterprise deals, which involved negotiated contracts that were not in any document. I needed to explain why the model would produce plausible but wrong numbers. I avoided jargon and instead showed three real examples: the question, what the model said, and what the contract actually said. Then I explained that the model writes from what it is given plus general patterns, so where the source is absent, it fills in. I proposed a version that answered from the contract database when a deal was identified and otherwise handed off to a person with a prefilled message. The leader agreed once the failure was concrete rather than theoretical, and the handoff version shipped. Concrete examples convince people in a way that explanations of model behavior never do.
14Describe a project where evaluation was the hardest part.
What they are checking: This reveals whether you have built evaluation from scratch for a subjective task and what you learned from it.
Example answer
We built an assistant that drafted responses to inbound partnership emails. Everyone agreed the drafts were sometimes great and sometimes off, and nobody could say how often. Evaluation was hard because a good draft depended on context in the CRM and on tone, which reasonable people judged differently. I wrote a rubric with five criteria, had three people grade fifty drafts, measured agreement, and revised the rubric until agreement was acceptable. Then I calibrated a model grader against those labels so we could evaluate hundreds of drafts per change. The first surprising result was that most failures were factual, from stale CRM fields, not tone. We fixed the data path and the score rose more than any prompt change had produced. The rubric became the shared definition of quality across product and engineering.
15Tell me about a time you worked with a safety, legal or trust team on a language feature.
What they are checking: They want to see that you treat safety and data handling as design inputs rather than a gate at the end.
Example answer
When we added a feature that let users ask questions about their own uploaded documents, the trust and legal teams needed assurance that content from one customer could never surface for another and that the model would not produce harmful output from adversarial documents. I set up a working session early rather than at the end. We wrote down the guarantees: tenant isolation enforced in the retrieval query, no cross-tenant caching, and a content filter on both input and output with logging. I built a test suite of adversarial documents, including prompt injection attempts embedded in files, and shared results with them before release. They asked for one change, a visible notice when content was filtered, which was a good idea. The feature launched on time and the shared test suite runs on every release.
Your fit and the role
16Why this company, given how many teams are building with LLMs right now?
What they are checking: They want to hear that you can tell a durable product apart from a thin wrapper and chose deliberately.
Example answer
Most teams building with LLMs are wrapping an API around a use case that is not yet clear. What drew me here is that the product already has users and a defined job, and the language model is a means rather than the pitch. From what I can see, you have proprietary data that makes retrieval genuinely better than a generic assistant, and that is where an engineer can build a real advantage rather than a thin layer. I also noticed that the job description mentions evaluation and monitoring explicitly, which tells me the team has been through at least one painful release and learned from it. I want to work where quality is measured, because that is the only place the interesting engineering happens. And the problem domain is one I would be happy to spend years on.
17How would you approach your first month on an LLM feature that already exists?
What they are checking: This checks whether you would learn the system and its failures before changing it.
Example answer
I would resist changing anything for the first two weeks. I would read the prompts, the retrieval code and the evaluation set, run the evaluation myself, and read a few hundred real interactions, especially the ones users rated poorly. Then I would talk to support and to whoever owns the product metrics to learn which failures actually matter to the business. By week three I would have a ranked list of failure types with rough frequencies. My first change would be to the evaluation if it did not cover the top failure, because otherwise I could not tell whether my later changes helped. Then I would take the single most common failure and fix it end to end, with a before-and-after measurement. The goal for the month is to understand the system and to earn the right to propose bigger changes.
18How do you think about the balance between research and product engineering in this role?
What they are checking: They want to confirm your working mode matches a product team and that you can bound exploratory work.
Example answer
This role reads as product engineering with research awareness, and that is what I want. I keep up with papers on retrieval, evaluation and efficient inference because those change what is possible, but I measure myself on shipped features that hold up in production. Where I think research-style work belongs is in bounded experiments: a week to test whether a new reranker or a different chunking strategy moves the evaluation, with a decision at the end. What I would avoid is open-ended model exploration without a product question. If the team needs a deeper research track, for example on fine-tuning strategy, I would want that to have its own owner and a clear interface to the product work. I am comfortable in either mode but I want to know which one I am in on a given week.
19What kind of problems do you not want to work on?
What they are checking: A candid answer here shows self-knowledge and helps both sides avoid a bad match.
Example answer
I do not want to build a general-purpose chat interface with no defined task, because success cannot be measured and every conversation becomes an argument about taste. I am also not interested in work where the plan is to ship first and figure out evaluation later; I have seen how that ends. And I would rather not work on features whose main purpose is generating content that will be spammed at people, since I think that is bad for the field and I would not be proud of it. What I do want is a specific job to be done, real users, data that gives us an edge, and a team that treats evaluation as a first-class deliverable. From the description, this role is on the right side of each of those lines, and I would want to confirm that in the conversation.
20Where do you want to be in three years?
What they are checking: They are checking that your ambitions fit the growth this role and company can offer.
Example answer
In three years I want to be the engineer who defines how a company builds and evaluates language features: the retrieval architecture, the evaluation standards, the routing between models, and the guardrails. I expect the specific models to change several times in that period, so the durable skill is judgment about what to build and how to know it works. I would like to have led one large system from design to stable production and to have mentored a couple of engineers into owning parts of it. I am open to a staff engineer path or to leading a small team, depending on what the company needs. This role fits because it is hands-on now, on a product that will need exactly that kind of system-level ownership as it grows.
Questions worth asking them
- How do you evaluate a change to a prompt or a retrieval pipeline before it ships?
- Which models do you use today, and what would trigger a switch?
- What is the most common way the current LLM feature fails, and how do you find out?
- How do you handle customer data in prompts, logging and evaluation sets?
- Who owns the cost of inference, and how is it tracked?
How to prepare
- Build or refresh a small retrieval-augmented system with an evaluation set before the interview so you can talk about failures from experience.
- Be ready to explain attention, tokenization, context limits and the key-value cache clearly on a whiteboard.
- Prepare a story about evaluating generated output where you can describe the rubric, the raters and what changed as a result.
- Know the tradeoffs between prompting, retrieval and fine-tuning well enough to recommend one for a problem the interviewer describes.