In this post12 sections
- Why a model grading a model makes interviewers wary
- Position bias, and how to cancel it
- Verbosity bias and self-preference
- Rubric drift and prompt sensitivity
- Calibrating a judge against human labels
- Measuring agreement with Cohen’s kappa
- Grading what it cannot do, and grading its own loop
- When not to use a judge at all
- What to say in the interview
- Questions people ask
- Keep reading
- More from the blog
Your works, the demo looks good, and now you need a number. So you reach for a second model to grade the first, and the interviewer asks the obvious question: who grades the grader? This post answers it for the AI round; for the other rounds, start with the FDE interview guide.
The short answer: the main LLM-as-a-judge pitfalls are position bias, verbosity bias, self-preference, rubric drift, prompt sensitivity and run-to-run noise. The fix is to treat the judge as a measuring instrument. Check it against human labels with a chance-corrected agreement score before you trust it, pin its prompt and model, and use plain code wherever an answer can be checked exactly.
Here is the whole list at a glance.
| Pitfall | How to catch it |
|---|---|
| Position bias | Swap the order |
| Verbosity bias | Pad a good answer |
| Self-preference | Compare with people on each family’s outputs |
| Rubric drift | Re-run the golden set on any change |
| Prompt sensitivity | Reword the prompt, re-run |
| Run-to-run noise | Judge the same set twice |
| Grading its own loop | Keep a human-labeled holdout |
Below: each pitfall with its check, a short script, when to skip the judge, and the words to say.
Why a model grading a model makes interviewers wary
A model judge sounds circular because it partly is. If the system and the grader share a blind spot, the grader passes the mistake, and the score rises while quality stays flat. The failure is quiet: the dashboard looks healthy, and nothing on it tells you the grader is wrong.
Postings name it as part of the job, which makes it fair game in the AI round:
- Scale AI’s Frontier Agents postings list evaluation harnesses built from offline benchmarks, online experiments, golden datasets, regression suites and LLM-as-a-Judge. Source 1Frontier Agents Engineer (Forward Deployed Engineering)PublisherScale AI (Greenhouse)Source typecompany job posting
- Cursor’s FDE postings make the FDE own production quality, including evals, debugging model behavior and failure modes. Source 2Forward Deployed EngineerPublisherCursor (Ashby job board)Source typecompany job board
- Decagon’s Agent Deployment Engineer posting asks for evaluation and regression-testing frameworks that validate agent behavior against ground truth and guard against performance drift. Source 3Agent Deployment Engineer @ Decagon (San Francisco)PublisherDecagon (Ashby job board)Source typecompany job posting
It shows up in take-homes too. One candidate reported, in a public GitHub repository created in August 2026, what the repository calls the Hippocratic AI coding assignment for an AI Agent Deployment Engineer role: turn a simple bedtime-story request into a story for ages 5 to 10 using prompting, and add an LLM judge. Source 4Hippocratic AI Coding AssignmentPublisherakshay-menta (GitHub)Source typecandidate’s take-home repository We use that task as the running example below. If your loop includes a take-home, the post on the OpenAI FDE take-home ends with a routine that works for any format.
Position bias, and how to cancel it
Show a judge two answers and ask which is better, and it may favor whichever sits in a particular slot, regardless of content. The paper by Zheng and colleagues, Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, documents this, along with verbosity bias and self-enhancement bias.
The check is cheap: ask twice, swapping the order, and only accept a verdict that survives the swap.
def judge(first, second):
# Stand-in for an LLM call.
# Returns "first" or "second",
# and always picks the first slot.
return "first"
def compare(a, b):
one = judge(a, b) # a shown first
two = judge(b, a) # b shown first
if (one, two) == ("first", "second"):
return "a"
if (one, two) == ("second", "first"):
return "b"
return "inconsistent" # followed order
print(compare("answer A", "answer B"))
It prints inconsistent, because the stand-in judge only ever picks the first slot. With a real judge, the share of inconsistent verdicts is a property of your judge you can report. If it is high, the pairwise setup is not measuring quality.
Three fixes, from cheapest:
- Swap and reconcile. Count an inconsistent pair as a tie, never as a win.
- Randomize order across the set, so any leftover bias spreads evenly over both systems.
- Score one answer at a time against a rubric instead of comparing pairs. Pointwise scoring has no slot to be biased toward.
Verbosity bias and self-preference
Verbosity bias means the longer answer tends to win. For a bedtime story that is the wrong direction: a story that drags on is worse for a tired child, not better.
To test it, take an output your labelers marked as good and pad it: restate the moral, add a paragraph of scenery, repeat the ending. If the judge’s score rises, it is rewarding length. Then:
- Put length in the rubric on purpose. “Fail if the story repeats a scene or runs past the requested length.”
- Give the judge a reference answer, so “more” is measured against “enough”.
- Grade each criterion pass or fail instead of asking for overall quality, which is where length sneaks in.
Self-preference means a judge tends to rate its own outputs, and plausibly its own family’s, more highly (Panickssery and colleagues show evaluators favoring their own generations). The usual fix is a judge from a different family from the system it grades, and a check against people on each family’s outputs.
A take-home may fix the model. One candidate reported, in August 2026, that the bedtime-story assignment says not to change the OpenAI model. Source 4Hippocratic AI Coding AssignmentPublisherakshay-menta (GitHub)Source typecandidate’s take-home repository Whether that covers the judge is worth one clarifying question. If your judge has to share a family with the system, say so in the write-up, and give human spot checks more weight.
Rubric drift and prompt sensitivity
Rubric drift is when the thing you measure changes under you. You tighten a criterion on Tuesday, and the score on Wednesday is not comparable to Monday’s. Or the model provider updates the version behind a floating alias, and the same prompt now grades differently. Either way, a chart that goes up may just mean the ruler shrank.
Research calls one version of this criteria drift: grading outputs changes what you think the criteria should be.
The fix is to version the judge like code. Every score gets stored with the exact judge that produced it:
judge:
model: pinned-snapshot-id # never a floating alias
prompt_version: story-judge-v3
rubric_version: v3
temperature: 0
scale: pass_fail
Pinning does not make the judge repeatable. Even at temperature zero, the same input can get a different verdict on another call. So run the judge twice on the and report how many verdicts flip. That flip rate is the noise floor: a score change smaller than it is not a change.
When any line changes, re-run the frozen golden set with the new judge before you compare numbers across the change. It doubles as a regression suite for the judge.
Prompt sensitivity is the judge’s version of the same problem: small wording changes move scores. Reordering the criteria, changing the example, or switching from a wide number scale to pass or fail can all shift results. Two habits help:
- Prefer binary criteria with sharp wording. Vague criteria give the judge room to wander.
- Paraphrase test. Reword the judge prompt without changing its meaning and re-run a sample. If verdicts flip, the rubric is carrying less of the decision than you thought.
Here is the difference sharp wording makes, for the bedtime-story judge:
- Vague: “Is the story appropriate for young children?”
- Sharp: “Fail if the story contains injury, death, or a threat to a character that is not resolved before the end. Otherwise pass.”
A person and a model can apply the sharp version the same way.
Calibrating a judge against human labels
A judge’s score means nothing until you know how often it agrees with people. This loop is the one we teach; no employer publishes a process for it.
Calibrate before you trust
- Sample real outputs, and include known failures on purpose. A sample of all-good outputs cannot show whether the judge catches bad ones.
- Have people label the sample with the same rubric the judge gets, without seeing the judge’s verdicts.
- Have two people label an overlapping slice. Their agreement with each other is your reference point: a judge that agrees with them about as well as they agree with each other is doing as well as you can measure.
- Run the judge on the same items and measure agreement with a chance-corrected score.
- Read every disagreement. Each one is a vague criterion, a labeling mistake or a real judge error.
- Fix the rubric or the prompt, then measure again on items you did not tune on.
If you tweak the judge prompt until it matches your calibration labels, you have fitted the judge to those items, and the agreement score flatters it. Keep a held-out slice for the final number.
Measuring agreement with Cohen’s kappa
Raw agreement, the share of items where judge and human match, flatters a judge whenever one label is common. If most stories pass, a judge that passes everything agrees with people most of the time and catches nothing.
Inter-rater agreement scores fix this by subtracting the agreement you would expect by chance. For two raters, Cohen’s kappa is (p_o - p_e) / (1 - p_e), where p_o is observed agreement and p_e is the agreement expected from each rater’s label frequencies alone. No library needed:
def kappa(a, b):
n = len(a)
p_o = sum(x == y for x, y in zip(a, b)) / n
a_yes, b_yes = sum(a) / n, sum(b) / n
p_e = a_yes * b_yes + (1 - a_yes) * (1 - b_yes)
return p_o, p_e, (p_o - p_e) / (1 - p_e)
human = [1, 1, 1, 1, 1, 1, 0, 0, 0, 0] # 1 = pass
judge = [1, 1, 1, 1, 1, 0, 1, 0, 0, 0]
fmt = "obs %.2f chance %.2f k %.3f"
print(fmt % kappa(human, judge))
skewed = [1] * 9 + [0]
lazy = [1] * 10 # passes everything
print(fmt % kappa(skewed, lazy))
It prints:
obs 0.80 chance 0.52 k 0.583
obs 0.90 chance 0.90 k 0.000
The first line matches scikit-learn’s cohen_kappa_score on the same labels. The second line is the one to remember: against a set where 9 of 10 stories pass, a judge that passes everything reaches 0.90 raw agreement, higher than the honest judge’s 0.80, and a kappa of 0, because chance would have got every one of its matches right.
Report a pair of extra numbers beside kappa. Of the outputs people failed, how many did the judge fail? Of the outputs people passed, how many did the judge pass? In the toy data above, people failed 4 stories and the judge caught 3; people passed 6 and the judge passed 5. The lazy judge catches 0 of the 1 failure. Call them the catch rate on failures and the keep rate on passes. A judge with a high keep rate and a low catch rate scores well on agreement and still misses what you built it to find.
How to read kappa in the round:
- Compare it to your people, not to a table. Bands such as Landis and Koch’s are a convention, not a standard. The working bar is the kappa your two human labelers reach with each other on this rubric. A judge that matches that is doing as well as the task allows.
- Say the sample is small. On a handful of items, one label moves kappa a long way. In the toy data, one extra disagreement can drop kappa from 0.583 to 0.348, and removing one can lift it to 0.783, depending on which label flips. The toy data above proves the arithmetic, not a judge.
- Break it down by criterion. A judge can agree well on “resolved ending” and poorly on “vocabulary for young children”. Use it only where it earns trust.
Grading what it cannot do, and grading its own loop
A judge that cannot solve a math problem cannot reliably grade the solution. The Zheng paper lists limited reasoning ability among the judge’s weaknesses, next to the three biases above. Give the judge a reference answer to compare against, or move the check to code.
The second trap is quieter. If the same judge drives rewrites, say a generate, judge and revise loop for the bedtime story, or your own prompt tuning, the storyteller is being tuned to please that judge. Its score on the final outputs then flatters the system, because the system learned the judge’s taste, not the reader’s.
The fix:
- Keep a human-labeled holdout the loop never sees, and never tune against it.
- Report that number as the headline, not the judge’s score on outputs it helped shape.
- Say which judge drove the loop in the write-up, so a reviewer knows which scores are self-graded.
When not to use a judge at all
The best judge is often no judge. If code can check the answer exactly, code is cheaper, faster, deterministic, and has no bias to calibrate.
| Question | Use |
|---|---|
| Is the output valid JSON? | Code |
| Does the total match? | Code |
| Do the tests pass? | Code |
| Is the quote in the source? | Code |
| Is the tone right? | Judge |
| Is the claim supported? | Judge |
| Is it safe for a child? | Judge and people |
For the bedtime story, length and required fields are code checks. Whether a scene would frighten a young child is a judgment call, so it goes to the judge, with people reviewing a sample every week.
Skip the judge, or put a person in its place, in these cases too:
- When its error rate is larger than the error you are trying to measure. If the system gets a category wrong rarely and the judge disagrees with people more often than that, the judge’s noise swamps the signal.
- When a single miss is expensive. For a medical or financial decision, a judge can sort a queue, but a person makes the call.
What to say in the interview
Name the failure modes before you are asked. The lesson on labeling assumptions and surfacing failure modes covers that habit in general; here is how it sounds for a judge.
Split the work.
“I use code for everything checkable: the output parses, the length is in range. For what code can’t check, like whether a scene is too scary, I use an LLM judge with a pass or fail rubric per criterion.”
Prove the judge.
“Two people label a sample, including stories I know are bad. I measure the judge’s kappa against their labels and against their agreement with each other, and I know its flip rate on repeat runs.”
Guard it.
“I swap order for any pairwise comparison and test with padded answers for length bias. The judge model, prompt and rubric are pinned and versioned. Where it disagrees with people most, on vocabulary, a person reviews a weekly sample.”
Then have a sentence ready for each follow-up:
- “Why not just use a stronger model as the judge?” A stronger judge still has biases, and I would still need to measure its agreement with people. Strength is a guess until the kappa says otherwise.
- “How many labels do you need?” Enough that the number stops moving. I bootstrap kappa over the labeled items, grow the set until the interval is narrow enough to decide with, and add labels first where the judge is weakest.
- “The customer asks why they should trust the score.” Show them the agreement with their own experts, and what a person still reviews. The lesson on explaining AI limits to non-technical leaders covers that conversation.
Common mistakes, and the fix for each:
- Quoting raw agreement. Quote kappa, with the human-to-human figure beside it.
- Calibrating on all-pass samples. Seed known failures.
- A floating model alias. Pin the snapshot and store it with every score.
- Judging what code can check. Move it to code.
The judge is one piece of a larger answer to “how do you know it works?”, which the post on LLM evaluation in the interview walks through. Where the labels come from is the subject of the post on building a golden set for LLM evals.
Practice it where it bites: design an eval harness for a support agent lists “trusts a model grader that was never checked against human labels” as a pitfall, and how do you know it works? is the question this post feeds. Then score model outputs against a ground-truth file with Pro.
Treat the judge as an instrument, and you will have the answer before the interviewer asks.
Questions people ask
What are the main LLM-as-a-judge biases?
Position bias (preferring the answer shown first or second), verbosity bias (preferring longer answers) and self-preference (rating its own outputs, and plausibly its own family’s, higher). Judges are also sensitive to small prompt changes, so pin the prompt and the model version, and expect some verdicts to change between runs even at temperature zero.
How do you validate an LLM judge?
Have people label a sample of outputs with the same rubric, run the judge on the same sample, and measure agreement with a chance-corrected score such as Cohen’s kappa. Read the disagreements, fix the rubric or the prompt, and measure again before you rely on the judge.
When should you not use an LLM judge?
When the answer can be checked in code, such as an exact match, valid JSON, a correct number or a passing test. Use a judge for qualities code cannot check, such as tone or whether an answer is supported by its source, and keep people reviewing a sample.
Keep reading
Lessons
More from the blog
Interview rounds
Agentic system design interview: tools, permissions, stuck agents and hand-off to a person
Design a safe agent in the interview: tool allowlists, scoped credentials, loop and cost budgets, idempotent actions, human approval and evals.
Interview rounds
How to build a golden set for LLM evals when the customer has no labeled data
The customer wants proof and has no labels. Build a golden set from real traffic: stratified sampling, expert labels, agreement checks and versioning.
Interview rounds
LLM evaluation in the interview: how to answer ‘how do you know it works?’
The demo looks great, then the interviewer asks how you know it works. A full answer names the metric, the data, the baseline, the result and monitoring.