A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
The question hands you the diagnosis: recall is fine, ranking is not. A strong answer names that, looks at the misses, fixes the ordering, and proves it on an evaluation set, not on the query that prompted it.
- Restate it as metrics. “Recall at twenty is good; precision at the top is the problem, so I’d watch hit rate at three and mean reciprocal rank.” If the evidence is anecdotes, build a labeled set first: real queries, each paired with the document that answers it. Label documents, not chunk IDs, so the set survives a chunking change.
- Read the misses before building anything. Old and new versions of the same document competing, boilerplate such as disclaimers crowding the top, chunks that separate an answer from its heading, exact identifiers such as SKUs or error codes that embeddings match poorly. Each has a cheaper fix than a new model. For the last one, adding keyword search and merging the two ranked lists with reciprocal rank fusion (Cormack, Clarke and Buettcher) addresses it.
- Add a between retrieval and the prompt. Retrieve a wide candidate set cheaply, score each query and passage pair with a cross-encoder, keep the top three. Say why it works: a bi-encoder embeds query and passage separately, while a cross-encoder reads them together and can see which part of the passage answers which part of the query (Nogueira and Cho, Passage Re-ranking with BERT).
- Measure before and after on the same set. The same quality metrics, plus p95 latency and cost per query. Set the bar for shipping before you run it.
- Say what it can’t fix. A reranker only reorders what retrieval returned. If the right passage is missing from the candidates, the work is recall: chunking, , query rewriting.
Putting all twenty chunks into the prompt instead costs more, runs slower, and models use the middle of a long context less reliably than its start and end (Liu et al., Lost in the Middle), so the right passage is present and still ignored.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- How do you measure the change before and after the reranker?
- The reranker adds a noticeable delay to every query. Is it worth it here?
- What if the right document is not in the top twenty either?
Where answers go wrong
- Increases the number of chunks in the prompt.
- Adds a reranker before reading a single miss.
- Labels gold chunks, so the evaluation breaks when chunking changes.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
“Top twenty but not top three means retrieval found it and ranking buried it. That’s a precision problem at the head of the list. Before touching anything, I want numbers: a set of real queries, each labeled with the document that answers it, so I can report hit rate at three and mean reciprocal rank now and after the change. I label documents, not chunk IDs, so the set survives a chunking change.
Before building anything, I’d read a sample of misses. If an old and a new version of the same document compete, a filter on the current version fixes it for free. If boilerplate such as disclaimers crowds the top, I’d drop it at indexing. If the misses are product codes or error strings, I’d add keyword search and fuse the two lists before . If they’re passages cut off from their section heading, I’d fix chunking to carry the heading.
If the misses are real ranking errors, the fix is a second-stage reranker. The change is small in code. Retrieve a wider candidate set from the existing index, score each pair with a cross-encoder and keep the best three. Then score before and after on the labeled set, and list which queries the change fixed and which it broke:
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")
def retrieve(query: str, n_candidates: int = 30,
k: int = 3) -> list[Chunk]:
# existing vector search, already filtered to what
# this user may open
candidates = index.search(query, k=n_candidates)
pairs = [(query, c.text) for c in candidates]
scores = reranker.predict(pairs)
ranked = sorted(zip(candidates, scores),
key=lambda p: p[1], reverse=True)
return [c for c, _ in ranked[:k]]
def evaluate(labeled: list[tuple[str, str]], search,
k: int = 3) -> dict[str, float]:
"""search(query) returns chunks in ranked order.
Each query has one gold document id."""
hits, rr = 0, 0.0
for query, gold_id in labeled:
ids = [c.doc_id for c in search(query)]
hits += gold_id in ids[:k]
if gold_id in ids:
rr += 1 / (ids.index(gold_id) + 1)
n = len(labeled)
return {f"hit@{k}": hits / n, "mrr": rr / n}
def fixed_and_broken(labeled, old, new, k: int = 3):
"""Queries moved into, and out of, the top k."""
fixed, broken = [], []
for query, gold_id in labeled:
was = gold_id in [c.doc_id for c in old(query)][:k]
now = gold_id in [c.doc_id for c in new(query)][:k]
if now and not was:
fixed.append(query)
if was and not now:
broken.append(query)
return fixed, broken
before = evaluate(labeled, lambda q: index.search(q, k=30))
after = evaluate(
labeled, lambda q: retrieve(q, n_candidates=30, k=30))
fixed, broken = fixed_and_broken(
labeled, lambda q: index.search(q, k=30), retrieve)
The cross-encoder is better at ordering because it reads the query and passage together instead of comparing two separately computed vectors, which is also why it’s too slow to run over the whole corpus. The Sentence Transformers cross-encoder docs cover the API. Both runs score the same ranked depth, so the only difference is the ordering. This model was trained on web search queries and reads at most 512 tokens of query and passage together, so I’d check chunk lengths first, since a long chunk is scored on its opening. I’d also try a hosted or domain-tuned reranker on the same set before choosing one, and sweep n_candidates, because a wider set helps recall into the reranker but costs latency linearly.
Three other options go on the same scoreboard. If the misses cluster in our own jargon, the cheaper fix may be fine-tuning the model, using the wrong passages that outranked the right ones as hard negatives, rather than adding a stage. When latency allows, an LLM can rerank the candidates listwise, reading several in one prompt and returning their order, the approach in Sun et al., RankGPT. And I’d keep a score floor, calibrated on the labeled set, so that when nothing clears it the assistant gets fewer than three passages, not three weak ones, and can say it doesn’t know.
Permissions come first. Filtering to what the user may open happens inside the search, before the reranker, so it never sees a document the user can’t open, and n_candidates is counted after that filter. Filter after the search instead, and a user with narrow access could lose most of the candidates to the filter, leaving a thin list to rerank.
If the reranker adds 400ms per query, whether that’s worth it depends on the . For a chat assistant where the model then spends seconds generating, it’s worth it if hit rate at three rises by at least 5 points and p95 time to first token stays inside the product’s budget with the reranker’s time added. For search-as-you-type, it isn’t. Ways to cut the delay: rerank fewer candidates, use a smaller cross-encoder, run it on a GPU and batch the pairs, or cache scores for repeated queries.
I’d write that 5-point bar down before running it, and I’d report both directions. On 200 labeled queries, 5 points is 10 queries, so I count how many queries the reranker fixed and how many it broke, and I read the broken ones before shipping. A net gain that hides 15 regressions on one query type is not a win.
And if the right passage isn’t in the top twenty either, the reranker can’t help, because it never sees it. Then it’s a recall problem: I’d measure recall at fifty, then work on chunking, and query rewriting.
What I wouldn’t do is pass all twenty chunks to the model. It raises cost and latency, and long contexts bury the passage in the middle where models use it least reliably.”