A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

A strong answer does two things: it separates what the model knows from how it behaves, and it lets evidence make the call rather than preference. Answer in this order.

  1. Split the need in two. Say it plainly: “Retrieval changes what the model can see. Fine-tuning changes how it behaves. Prompting changes both, cheaply, within what fits in the context.” Then classify the customer’s failure: missing or stale facts, or the wrong format, label, tone or tool choice?
  2. Start with the cheapest thing that could pass. A clear instruction, a handful of worked examples, an output schema. Say you would build the evaluation set before anything else, because it is the only way to know whether the next step helped. Prompting is enough for knowledge too, when the whole corpus is small and stable: put it in the context, cache it, and build an index only when it stops fitting, changes faster than you can reload it, or needs per-user permissions.
  3. Add retrieval for knowledge. Facts that change, facts the customer owns, facts that must be cited or filtered by permission. Provenance and knowledge you can update without retraining were the stated motivations of retrieval-augmented generation (Lewis et al.); permission filtering is yours to add. Fine-tuning is the wrong tool here: it absorbs facts unreliably, can’t cite them, can’t drop them when they change, and ignores who is allowed to see what. The research backs the first point. Ovadia et al. compared the two for injecting knowledge and found retrieval did better than unsupervised fine-tuning, and Gekhman et al. found that as a model learns new facts through fine-tuning, it becomes more likely to hallucinate.
  4. Fine-tune for behavior a prompt can’t hold. A stable classification with labeled data, a strict format at high volume, or a narrow task you want a smaller, cheaper, faster model to do. Name the cost: labeled examples, a held-out split, a retraining path when the task drifts, and an upgrade tax. A fine-tune is pinned to one base model, so every model upgrade means retraining and re-evaluating, while a prompt and an index move to the next model for the cost of one evaluation run. Check, too, that the customer’s approved model and cloud allow tuning at all.
  5. Let the evaluation decide. Compare the options on the same set, with cost per call and latency beside accuracy.

Then give your example in that shape: the need, what you tried first, what moved on the evaluation, what you chose.

The model answer uses an invented insurer to show the shape. In the room, use a project you worked on. If you have never fine-tuned, say so, and walk through the decision you would make and what you would measure. An invented project falls apart at the first follow-up about its numbers. For the evaluation half of the answer, practice telling a customer how you know an AI system works. The syllabus for Production AI systems lists the Pro lesson on this decision.

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • Use your own example: what did you choose, and what evidence decided it?
  • The facts change weekly. What does that rule out?
  • The customer has a few hundred labeled examples. What can you do with them?

Where answers go wrong

  • Recommends fine-tuning to add knowledge.
  • Picks an option before building an evaluation set.
  • Compares accuracy and ignores cost per call and latency.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Compare with the model answer

Model answer

“I treat it as two questions. Is the model missing knowledge, or is it behaving wrong? Retrieval fixes the first, fine-tuning the second, and I start with prompting for both, because a prompt is the cheapest thing to change and to roll back. If the documents are few and stable, I don’t even build an index at first: they go in the context, cached.

An insurer’s claims team wanted an assistant that did three jobs: draft the acknowledgement letter to the policyholder, answer adjusters’ questions about policy wording, and tag each incoming claim note with one of the team’s internal loss categories. Those are different problems.

The letter stayed a prompt. A template, a handful of approved letters as examples and the claim’s own fields passed the claims team’s rubric on our test set, so there was nothing for retrieval or tuning to fix. When the prompt passes the evaluation, I stop.

The wording questions were a knowledge problem. The policy documents change with every product revision, answers have to point at the clause, and adjusters in one region may not see another region’s products. So that half is retrieval over clause-level chunks, region as a metadata filter, and a prompt that requires a citation for every statement. Fine-tuning fails on each point: the model would learn last quarter’s wording, couldn’t show where an answer came from, and every revision would mean retraining.

The tagging was a behavior problem. The categories were stable, the input was a short note, and the output was one label from a fixed list. Before any prompt, I built the evaluation from notes the claims team had already labeled, split into a small pool for examples and a held-out test set the prompt never saw. Then the prompt: the category definitions, a few examples, and the output constrained to the list. It kept confusing ‘escape of water’, which is sudden and covered, with ‘gradual seepage’, which is excluded: a line adjusters draw from experience, not from a definition you can write down. Before training anything, I tried the cheap middle steps: the most similar labeled notes as examples in the prompt, and a plain classifier on embeddings. Both helped, and neither closed the gap. That is where fine-tuning earns its cost: a narrow, stable task, labels that already exist, and a convention that is easier to show than to describe.

On the held-out set the prompt scored 0.84 macro-F1. A logistic regression on embeddings reached 0.86 for almost nothing, and retrieving the 8 most similar labeled notes into the prompt reached 0.88. A small fine-tuned model reached 0.93 at $0.0006 a claim against $0.004, with p95 latency of 180ms against 1.4s, so we shipped it, with a monthly check against fresh labels to catch drift. Macro-F1 hides the one pair that matters, so I also reported F1 on escape of water against gradual seepage alone, where the fine-tune gained most: 0.71 for the prompt, 0.89 tuned. The tag routes the claim; the adjuster still decides coverage. I also budgeted the upgrade tax: when the base model is retired, we retrain and rerun the same evaluation before switching.

If the facts changed weekly, that alone rules out fine-tuning for knowledge. It is retrieval, and the index refresh becomes the thing to engineer and monitor. And with only a few hundred labeled examples, they are the evaluation set first and a source of in-context examples second. Only after the prompt plateaus would I try parameter-efficient tuning such as LoRA (Hu et al.), with cross-validation, because a held-out split that small is noisy. With no labels at all, I’d have the large model with the best prompt label a few thousand notes, have adjusters check a sample, train the small model on those, and score it only against human labels.

So my default order is: evaluation, prompt, retrieval for facts, tuning for behavior, and I’d rather ship the prompt and pay per call than own a fine-tune I can’t move to next year’s model.”