A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

The customer isn’t asking about your tooling. They’re asking whether they can rely on the system for a decision they have to defend. At least one employer writes this into the job: OpenAI’s (Healthcare) posting in San Francisco, as of September 2026, lists defining evaluations and that measure quality against customer-specific acceptance thresholds as a responsibility. Source 1Forward Deployed Engineer (FDE), Healthcare - SFPublisherOpenAI (Ashby)Source typecompany job posting

Answer in this order, one or two sentences each:

  1. The verdict, then a check. Lead with the one-sentence verdict: where it meets the bar and where it’s weak. Then confirm the decision it serves instead of telling them what it is.
  2. The decision and the costly error. “Your team uses it to answer employees’ benefits questions. The expensive mistake is a confident wrong answer about coverage.”
  3. The evaluation set. Where the questions came from, how many, how they match the real mix, and who wrote the correct answers.
  4. The metric and the bar, agreed before you built. Tie each metric to the costly error, and say who signed off on the threshold.
  5. The baseline. What the system is compared against: the process it replaces, or the simplest alternative.
  6. The result, by segment, with counts. Say the count (1 of 130), not just the rate (0.8%), and say whether the sample is big enough to show the bar is met. Then say where it misses and what you did about the misses.
  7. Monitoring. What you watch in production, how you would notice a drop, and what happens then.
  8. The limits. What the evidence does not cover.

Say the numbers out loud, and say where they came from. Then invite the next question: “Want to see the misses?” Showing your failures builds more trust than a clean average.

Rehearse the three follow-ups below: who labeled the data, how you checked the judge, and what changed after launch. Each one tests whether the evidence is real. For the judge, compare its agreement with your labelers to their agreement with each other, and to the error rate it has to measure: a grader that disagrees more often than the system errs can’t measure that error. Then count how often it passes a wrong answer, because raw agreement hides that when most answers are right.

Naming the eval framework, the dashboard and the tracing tool is the pitfall below. Tools describe how you look, not what you found.

Practice it with a customer

In Pro, the claim intake case asks you to get a model an insurer already bought into production safely, and to decide what evidence the claims team needs before any draft reaches their system.

GlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineerGlossaryLaunch criteriaThresholds agreed with a customer before building that decide whether a system goes live.More on Launch criteria

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • How did you build the evaluation set, and who labeled it?
  • Who or what judged the outputs, and how did you check the judge?
  • What changed after launch, and how did you notice?

Where answers go wrong

  • Describes tools and dashboards instead of evidence tied to the customer’s decision.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Compare with the model answer

Model answer

The setting: an assistant that answers employees’ benefits questions for a manufacturer’s HR team. Their HR director asks how we know it works.

“Short answer: on your own past questions it was right more often than the ticket queue, it made a single confident mistake about coverage, and it’s weak on leave for part-time staff. Is the decision whether your people can use this instead of opening a ticket with your benefits team?”

She says it is.

“Then the mistake that matters is a confident wrong answer about coverage, because someone might skip care or pay a bill they shouldn’t.

“So before we built anything, your benefits team gave us real questions from last year’s tickets, sampled to match the actual mix: coverage, enrollment, leave and a set of questions whose correct answer is ‘contact HR’. Two of your benefits specialists wrote the correct answer for each. Where they disagreed, your benefits manager decided, and we kept her reasoning as the grading guide.

“We agreed each bar with you in writing, before we built, and we compared against your ticket queue’s answers on the same questions. In one sentence: on this set it meets every bar we agreed, it’s only slightly ahead of your ticket queue, the coverage number needs more answers before it proves anything, and one segment falls short. Here are the numbers:”

400 real questions. Correct
answers by 2 specialists,
disagreements settled by
your benefits manager

Correct overall
  assistant      364/400   91%
  ticket queue   352/400   88%
  agreed bar               >= 85%
Confident wrong on coverage
  assistant        1/130   0.8%
  ticket queue     2/130   1.5%
  agreed bar               <= 1%
Correctly says 'ask HR'
  assistant       43/50    86%
  agreed bar               >= 80%
Weakest: part-time leave
  assistant       26/35    74%

Specialist vs specialist
  agreement                93%
Grader vs settled labels,
100 held-out answers
  agreement       95/100
  wrong answers   14
  passed anyway    2

“Overall, the lead over the ticket queue is small, and on one set this size I wouldn’t claim the assistant is better; what matters is that it clears the bar. A single wrong coverage answer is too little evidence to show we’re under that bar: one error in 130 is consistent with a true rate as high as about 4%, at 95% confidence. To show we’re under 1% with no errors, I’d need about 300 coverage answers. So for the first month a specialist reviews every coverage answer in production, and the set grows from real traffic until we have them.

“The weak spot is leave rules for part-time staff, because the policy PDF contradicts the handbook. Until your team decides which one is right, those questions route to a person.

“Grading: exact facts, like deductibles and dates, are checked against the policy by code. A model grades whether the rest of each answer is correct. Before trusting it, we checked it against your specialists’ labels on 100 answers it had never seen; that’s the last block of the numbers. Agreement isn’t the number I watch. Of the 14 wrong answers in that held-out 100, the grader passed 2, and how often it lets a wrong answer through is what matters, because a grader that passes wrong answers hides exactly the failures you care about. The grader also disagrees with your specialists more often than the assistant makes coverage errors, so it can’t measure that number. Your specialists graded every coverage answer in the set by hand.

“The model version and the prompt are pinned. Before any change, ours or the model provider’s, the full set reruns and must clear every bar, and your team sees the results before we switch. The same version doesn’t give the same answer every time, so we ran the set 3 times: 9 questions changed verdict between runs. The numbers above are the worst run, and those 9 questions are in the weekly review sample.

“After launch, every answer logs its sources. Your team reviews a weekly sample, and any answer an employee flags goes back into the evaluation set. If a topic falls below the bar, it routes to a person the same day while we find the cause, whether that’s a change of ours or a change in your documents. A few weeks in, the review caught wrong answers after the dental plan changed carrier mid-year. We now pull the list of superseded documents from your benefits system each night, and any answer that cites one is blocked and routed to a person.

“What this doesn’t tell you: how it does on questions nobody has asked yet, and on next year’s plan until we rerun the set against it. That rerun is part of your renewal checklist.”