In this post15 sections
  1. What a golden set is, and what it is not
  2. Why FDEs get asked this
  3. Sample from real traffic, by segment
  4. Write the labeling guide before anyone labels
  5. Label with the customer’s experts and check agreement
  6. Version the set and keep a holdout
  7. Add every production failure
  8. Turn the set into launch criteria the customer signs off
  9. Common mistakes and the fix
  10. Where this shows up in FDE interviews
  11. What to say in the interview
  12. Practice it out loud
  13. Questions people ask
  14. Keep reading
  15. More from the blog

The pilot works in the demo, and now the customer’s VP asks the question that decides whether it ships: “How do we know it’s good enough to turn on?” You look for labeled data to prove it, and there is none: years of tickets, no record of what the right answer was. This post shows how to build the evidence from nothing, in the job and in the interview, and it sits under the FDE interview guide, which covers every round.

The short answer: sample real traffic by segment, write a labeling guide, have the customer’s experts label the sample and check that they agree with each other, freeze and version the set with a held-out split, add every production failure, and agree launch criteria against it before anyone sees a score. That versioned, expert-agreed sample is a golden set.

What a golden set is, and what it is not

A is a fixed collection of real inputs, each with an answer or label that the customer’s experts agree is correct, used to score every version of the system the same way.

It is not:

  • Your demo examples. You picked those because they work.
  • Synthetic prompts alone. Generated examples are useful for filling gaps, but they miss the typos, forwarded chains and half-finished questions of real traffic.
  • Training or few-shot data. Anything in the prompt or the fine-tuning set is no longer a test.
  • A single number. The set exists to show quality per segment, because the average hides the segment that will get you in trouble.

Why FDEs get asked this

It is written into the job. OpenAI’s (Healthcare) posting lists defining evaluations and that measure quality against customer-specific acceptance thresholds. Source 1Forward Deployed Engineer (FDE), Healthcare - SFPublisherOpenAI (Ashby)Source typecompany job posting Snowflake’s Senior/Staff Applied AI FDE posting asks the hire to turn ambiguous customer goals into quality metrics, evaluation frameworks and golden datasets. Source 2Senior/Staff Forward Deployed Engineer, Applied AI @ SnowflakePublisherSnowflake (Ashby job board)Source typecompany job posting

In the Snorkel interview account described further down, the ground truth arrived as a CSV. On a real deployment, nobody hands you one, so the follow-up to expect is “where did the right answers come from?” The steps below are our method for answering it, not a rubric any employer publishes.

A running example keeps this concrete. A fictional payroll software company wants an to triage support tickets into billing, outage and legal threat, and to route each one. Billing is almost everything, and legal threats are rare and expensive to get wrong.

Sample from real traffic, by segment

Start with the customer’s actual inputs: a month of tickets, emails, call transcripts or queries, with handled the way their security team requires. Knowing who owns each data source and how fresh it is comes first, because you need permission to pull it.

Then sample by segment, not at random. A plain random sample mirrors the traffic, so the rare segment barely appears. Pick segments that match the decisions the customer cares about: intent, customer tier, language, channel, or risk.

import random
from collections import defaultdict

def stratify(rows, seg, k, seed=7):
    """Up to k rows from every segment."""
    rng = random.Random(seed)
    groups = defaultdict(list)
    for row in rows:
        groups[seg(row)].append(row)
    sample = []
    for _, items in sorted(groups.items()):
        take = min(k, len(items))
        sample += rng.sample(items, take)
    return sample

We ran it on made-up traffic of 900 billing, 90 outage and 10 legal-threat tickets with k=40. It returned 40 billing, 40 outage and all 10 legal threats. A plain random sample of 120 from the same traffic drew 2 legal threats in our run, and in repeated draws about 28% of such samples contained none at all.

Stratifying cannot invent examples, though. All 10 legal threats is still very few, and a split will leave about half of them in the holdout. When the rare segment is the one that matters, go and find more of it: widen the time window, search past tickets for “lawyer” or “attorney”, and ask the legal team for cases they have handled.

Two things to say about it in an interview:

  • Keep the real mix too. Stratified sampling over-represents rare segments, so an overall score from it is not what users will see. Report per segment, and weight by the real traffic mix when you need one overall number.
  • Add the hard cases on purpose. Ambiguous requests, missing information, questions the system should refuse, and inputs in a second language. Tag them as their own segment so they never vanish into the average.

How big? Size it for the decision. If the system gets about 0.90 of a segment right, n=100 gives a 95% confidence interval of about ±0.06 (a normal approximation). At n=40 the interval is lopsided, from about 0.77 to 0.96 by the Wilson method, so a real drop can hide in it. That is enough to see a segment collapse, not a small slip. Say the trade-off out loud: more examples per segment, or more careful labels from a busy expert. Start small and grow.

Write the labeling guide before anyone labels

Most disagreement between labelers is a missing rule, not a careless person. Write the guide first, with the customer, and keep it short enough to read.

A useful guide has, for each label:

  • a one-line definition in the customer’s words;
  • two or three real examples that clearly belong;
  • the near misses that do not, and why;
  • what to do when an input fits two labels, or none.

For the payroll example, the hard line is legal threat. Is “I’m talking to my lawyer about these charges” a legal threat, or billing? The guide has to say, because two experts will split on it, and so will the model. Agreeing who owns each metric and who would dispute it is the same conversation: the person who owns the risk writes this rule.

Free-text answers: a checklist, not a label

Most LLM systems write answers, not labels: a reply, a summary, an agent’s message. There is no single right string, so the guide says what a correct answer must contain and what it must never say.

Say the payroll company also wants the agent to answer billing questions. The input is “Why was I charged twice in March?” The golden record holds a reference answer the billing lead wrote, plus two lists. In this made-up product, duplicate charges are refunded automatically:

  • Must include: the duplicate charge is refunded in 5-7 business days; a link to the billing page.
  • Must not: promise a refund amount; give legal advice.

A reviewer, or a , marks each item pass or fail. The answer passes only if every item passes. The reference answer shows what good sounds like; the checklist is what gets scored, because two good answers rarely share wording.

Label with the customer’s experts and check agreement

The right answers belong to the people who own the decision: the support lead, the claims adjuster, the compliance officer. You can draft labels to save their time, but they confirm every one. If you label it yourself, you are grading your own understanding of their business.

Then measure whether the experts agree with each other. Give two of them the same slice, say 100 items, without showing each other’s labels, and compute inter-rater agreement. Use , not raw agreement, because raw agreement flatters a skewed label set.

A worked example with pass or fail labels: the experts agree on 88 of 100 items, one marks 20 as fail and the other 18. Chance agreement is 0.692, so kappa is (0.88 - 0.692) / (1 - 0.692) ≈ 0.61. That looks worse than raw agreement of 0.88, and it is meant to.

At 0.61, don’t label the rest yet. Read the 12 disagreements, write the missing rules into the guide, and have both experts relabel a fresh slice. Kappa as written here compares two raters; for more, use Fleiss’ kappa or Krippendorff’s alpha.

What you do with the result:

  • Read the disagreements, not just the score. Each one is a missing rule. Add it to the guide and bump the guide’s version.
  • Name a tie-breaker. One person decides disputed items, and their reason goes into the guide.
  • Treat expert agreement as the ceiling. If two experts disagree on a slice of items, no system can be shown to beat that level on them. Say this to the customer early; it resets expectations better than any caveat later.

If you plan to use a model as a grader, check it against these same expert labels first. The post on the pitfalls of LLM-as-a-judge covers how.

Version the set and keep a holdout

A score only means something next to the version of the set it came from. Store the set as a file in version control, one record per line (JSONL). Here is one record, spread over several lines to fit a phone:

{
  "id": "g-0412",
  "input":
    "I'm talking to my lawyer about these charges",
  "segment": "legal_threat",
  "expected": "route_to_legal",
  "source": "sampled",
  "guide": "v2",
  "split": "holdout"
}

A free-text record carries its checklist instead of a label:

{
  "id": "g-0907",
  "input": "Why was I charged twice in March?",
  "segment": "billing_answer",
  "reference": "Sorry about that...",
  "must_include": [
    "duplicate refunded in 5-7 business days",
    "link to billing page"
  ],
  "must_not": [
    "promise a refund amount",
    "legal advice"
  ],
  "source": "sampled",
  "guide": "v2",
  "split": "dev"
}

Three rules keep it honest:

  1. Split it. Keep a dev split you look at while changing prompts, and a holdout you only score. Tune against the whole set and a better score may mean you fitted the test.
  2. Never compare across versions. When the set changes, rescore the old system on the new version before you claim an improvement.
  3. Record why each item is there. Sampled, hand-added hard case, or production failure. It tells you later which segment the set is thin on.

The set is now a regression suite: it runs on every prompt, model or retrieval change, before the change ships.

Add every production failure

The golden set is never finished. Every output a user flags, an agent corrects or a reviewer overturns is a candidate for the next version: label it by the guide, tag it with its segment, and add it with source: production.

This is how the set stops being a snapshot of launch week. Decagon builds a related check into its own product: in June 2026 it announced Duet Autopilot, which tests each proposed agent change against the conversation that surfaced the issue and a golden test set of hundreds of conversations, and requires human approval before any change reaches production. Source 3Introducing Duet Autopilot: The self-improving agent for conversational AIPublisherDecagon blogSource typecompany blog

In an interview, one sentence covers it: “Every production failure becomes a test case, so the same mistake can’t ship twice without us seeing it.”

Turn the set into launch criteria the customer signs off

The set exists to answer the VP’s question. Before anyone sees a score, write down with the customer:

  • The metric per segment, not one average. For the payroll agent: routing accuracy on billing and outage, and recall on legal threats.
  • The threshold for each, for example accuracy >= 0.90 on billing and no missed legal threats in the holdout, measured on a named version. With only 10 legal threats you cannot measure a recall like 0.98, because one miss is already 0.90. That is another reason to go and find more of them.
  • The baseline. Score the customer’s current process on the same set. If their team misses some legal threats today, that miss rate is the bar to discuss, not perfection.
  • A guardrail metric that must not get worse, such as the share of tickets sent to the wrong queue.
  • Who signs, and what happens on a miss: a narrower launch with human review on the risky segment, or another iteration.

Writing criteria down before building is not only our advice. ElevenLabs’ April 2026 post on lessons from forward deployed engineering recommends a test-driven approach, with success criteria and tests defined before building a voice agent. Source 4Building voice agents that last: some lessons learned from forward deployed engineeringPublisherElevenLabsSource typecompany blog The post on choosing a success metric in a case interview goes deeper on picking the number and its guardrail.

Common mistakes and the fix

MistakeFix
Engineers write the answersExperts confirm every label
Random sample onlySample per segment, report per segment
No written guideGuide first, versioned with the set
Raw agreement reportedCohen’s kappa, plus the disagreements read
Tuning on the whole setDev split to tune, holdout to score
Scores compared across versionsRescore the old system on the new set
Thresholds set after the resultsCriteria signed before the first run

Where this shows up in FDE interviews

Postings describe the job, not the interview, and the interview evidence comes from single candidates. One candidate’s public GitHub repo, created in August 2026 and described as a DigitalOcean FDE take-home, compares models for classifying doctl GitHub issues and recommends which two to run in production. Source 5digitalocean-fde-eval (repository description)PublisherEdwardTang (GitHub)Source typecandidate’s take-home repository In a separate account, one candidate reported, in April 2026, that Snorkel AI’s recruiter described an FDE technical screen built on a fictional client brief, a ground-truth CSV and model outputs; that candidate later said the interview was canceled. Source 6Got a Snorkel AI Forward Deployed Engineer interview next week, anyone been through this? Format is unusual (post by u/itskabeer)PublisherReddit r/csMajorsSource typecandidate report on RedditSource 7Got a Snorkel AI Forward Deployed Engineer interview next week, anyone been through this? Format is unusual (comment by u/itskabeer)PublisherReddit r/csMajorsSource typecandidate report on Reddit Both are evaluation work on someone else’s task, and a golden set is what you would build before either.

What to say in the interview

When the interviewer asks “the customer has no labeled data, so how would you evaluate it?”, here is a version you can say in about a minute:

“I’d build a golden set from their real traffic. First I’d sample it by segment, so the rare, expensive cases like legal threats show up, not just billing. Before anyone labels, I’d write a short labeling guide with the person who owns the risk. Their experts label, not me, and two of them label an overlapping slice so I can check agreement with Cohen’s kappa; every disagreement becomes a rule in the guide. I’d freeze a holdout I never tune on, version the whole thing, and add every production failure to it. And before we run anything, we agree per-segment launch criteria against it, with their current process as the baseline, so ‘is it good enough’ becomes a check, not an argument.”

Expect three follow-ups: how big it needs to be (answer with the trade-off, not a magic number), what if the experts disagree (it’s the ceiling, and the guide gets fixed), and whether a model can grade it (only after it agrees with the experts). Explaining why the system can’t be right every time helps with the conversation after that one. Then try the whole thing against a customer who pushes back in the free practice case, which needs only a sign-in.

Practice it out loud

Reading this is not the same as answering under pressure. Three questions, in order:

  1. Say it in a minute: a customer asks how you know your AI system works.
  2. Write the scorer: scoring model outputs against a ground-truth file.
  3. Design the pipeline: an eval harness for a support agent.

For the wider answer, including metric, baseline and monitoring, read how to answer “how do you know it works?”. Start with the first question now.

GlossaryGolden setA curated, labeled set of examples, sampled to cover the cases that matter, used to evaluate a system the same way over time.More on Golden setGlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineerGlossaryLaunch criteriaThresholds agreed with a customer before building that decide whether a system goes live.More on Launch criteriaGlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on AgentGlossaryPersonal dataInformation that identifies a person or can be linked to one, which privacy law and customer policy restrict.More on Personal dataGlossaryRetrieval-augmented generationAnswering with a model that is given passages retrieved from a document collection as context.More on Retrieval-augmented generationGlossaryLLM-as-judgeUsing a language model with a rubric to grade outputs; the judge itself must be checked against human labels.More on LLM-as-judgeGlossaryInter-rater agreementHow consistently two reviewers, or a reviewer and a model judge, label the same items, corrected for agreement by chance.More on Inter-rater agreement

Questions people ask

What is a golden set in LLM evaluation?

A fixed, versioned set of real inputs with answers or labels that the customer’s experts agree are correct. You run every model, prompt or retrieval change against it, so you see whether quality went up or down before users do.

How big should a golden set be?

Big enough to cover each segment you care about with enough examples to see a change, and small enough for experts to label carefully. Start small, sample by segment, and grow it with production failures. Say that trade-off in the interview rather than quoting a magic number.

Who should label a golden set?

The customer’s domain experts, working from a written labeling guide, with a sample labeled twice so you can measure agreement. Engineers can draft labels, but the people who own the decision should agree the answers.

Keep reading