LLM-as-judge means using a language model, given a rubric and often a reference answer, to score outputs or pick the better of two. The judge is a model with its own measurable errors: Zheng et al. found position bias, verbosity bias and self-enhancement bias in model judges, so a judge’s scores mean little until its agreement with human labels has been measured.
In FDE interviews
Scale AI’s Frontier Agents postings ask the engineer to deploy evaluation harnesses that use LLM-as-a-Judge alongside offline benchmarks, online experiments, golden datasets and regression suites. Source 1Frontier Agents Engineer (Forward Deployed Engineering)PublisherScale AI (Greenhouse)Source typecompany job posting One candidate reported, in a repo created in August 2026, that the Hippocratic AI take-home for an AI Deployment Engineer role asked for a program that turns a bedtime-story request into a story suitable for ages 5 to 10 and adds an LLM judge. Source 2Hippocratic AI Coding AssignmentPublisherakshay-menta (GitHub)Source typecandidate’s take-home repository A judge there has to check fit for the audience, not only whether the story is good.
A strong candidate writes a rubric with a pass or fail criterion per dimension and tests the judge by swapping answer order and padding answers. Before anyone relies on its scores, they report how often it catches known failures and how often it passes known-good outputs, or its kappa with expert labels; raw agreement flatters a judge that mostly sees passes. They use a judge from a different model family from the one it grades, and keep people reviewing a sample after launch.
Related: golden set, regression suite.