A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

There is no exact answer here, so the interviewer is watching how you manage being wrong. Name the pipeline in one breath (normalize, block, score, decide, cluster), then spend your time on the two places where it fails: pairs you never compare, and pairs you merge wrongly.

  1. Ask what a wrong merge costs. If the output merges billing accounts, two people fused into one is far worse than a duplicate left behind. If it is a mailing list, the reverse. Ask how many records there are and whether anyone can label a sample; the thresholds come from that sample, not from your intuition.
  2. Normalize, and keep the fields that separate people. Fold case and accents, drop punctuation, expand St and Apt, sort name tokens so “Smith, John” meets “John Smith”, and compare a ZIP+4 as its five-digit ZIP. Keep the generational suffix, the house number and the unit apart: they are what tells a father from a son and a neighbor from a duplicate.
  3. Block before you score. Comparing every pair is quadratic. Generate candidate pairs from several cheap keys (postcode plus name prefix, postcode plus house number, phone, email), because a pair that shares no key is never compared. Cap giant blocks and report them.
  4. Score with reasons. Per-field similarity plus exact contact matches, returned with the reasons, so a reviewer sees why.
  5. Two thresholds, three outcomes. Match, review, no match. Some conflicts force review whatever the score: a different suffix, a different unit.
  6. Cut the chains. Matches chain: A to B to C can join two records that never matched. Don’t throw the whole cluster away; join two groups only when every pair across them matches on its own, and send the edge that would have joined them to a person. Cap cluster size, so one shared phone can’t build a cluster of hundreds.

The trap is merging rows in place. Output a mapping from record to cluster with the reasons, so every merge can be explained and undone.

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • Two records for the same person share no blocking key. How would you estimate how many pairs like that you are missing?
  • The billing team says a wrong merge is far worse than a missed one. What do you change, and what does it cost the review queue?
  • New records arrive every day. How do you match them without rerunning the whole table?
  • Record A matches B and B matches C, but A and C look nothing alike. What does your output say?

Where answers go wrong

  • Compares every pair of records, which finishes on the sample and never finishes on the customer’s table.
  • Uses one threshold and merges rows in place, so a wrong merge of two real customers can be neither found nor undone.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Model answer

“I’ll assume this feeds a merge of customer accounts, so a wrong merge costs more than a missed one, and the table is too big for all pairs. Standard library only, so it runs here; with the customer I’d use rapidfuzz for speed and keep the same structure.”