A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

A score is useful only if its reader knows what to fix next. Say the script’s two jobs first: get the denominator right, and sort every error into exactly one bucket someone can act on.

  1. Pin the join and the denominator. Key both files by ID. A ground-truth row with no prediction is a missing error, not a skipped row. A prediction with no ground truth, or a duplicate ID, stops the run. Carry segment columns from the ground truth (source system, document type) through the join. Say: “The denominator is the ground truth, always.”
  2. Normalize on purpose, and say what. Case, surrounding whitespace, number formats like 1,200 versus 1200. Keep raw values.
  3. Pick metrics from the task. For labels: accuracy plus per-class precision and recall, because accuracy hides a rare class the model never predicts (scikit-learn’s classification report shows the shape). For extracted fields: exact match per field, not per record.
  4. Classify errors with ordered rules. missing, then unparseable, then per item: for labels, wrong_label as an expected-and-got pair; for fields, field_missing and wrong_value, counted per field name. The first matching rule wins, so counts sum to the error total; check that they do. Beside them, count format_only: values that match only after normalization. They score as correct, and the count shows how much the score leans on your normalizing.
  5. Print something a person can act on. Counts by type, largest first; top confusion pairs; the breakdown by segment, with the row count beside every rate so nobody acts on a rate computed over a handful of rows; example IDs per type. Write the same result as JSON for comparing runs. Run it first on a fixture where you know every number.

The trap is one accuracy figure over the rows that parsed. It rewards the model for failing loudly on hard cases.

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • Accuracy rose after a prompt change, but one class now gets mislabeled far more often. What does your report show, and what do you do?
  • The ground truth itself is wrong on some rows. How would your script help you find them?
  • How do you compare this run with last week’s run on the same file?

Where answers go wrong

  • Computes accuracy over only the rows that parsed, so a model that returns garbage on hard inputs scores higher than one that tries.
  • Reports one headline number with no breakdown, so nobody knows which fix to try first.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Model answer

“I’ll assume two JSON Lines files. Each ground-truth row has an id, the expected answer and any metadata columns, such as the source system.