A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

A retriever with no ranking test is a function that returns something. As of September 2026, Cohere’s Agentic Platform posting asks for experience building and deploying and agentic applications, and the ability to build evaluation frameworks that measure accuracy, safety and latency. Source 1Forward Deployed Engineer, Agentic PlatformPublisherCohere (Ashby job board)Source typecompany job posting This question tests the retrieval half of that, and the habit of measuring it, at the smallest scale there is.

Aim to show three things in order: that you know what TF-IDF actually computes, that you retrieve passages rather than files, and that you can show the ranking is sensible. Structure the answer the same way.

  1. Pin the unit of retrieval first. Say: “A passage is what I return and what I score, so document frequency is counted over passages, not files.” Then pick a chunker you can defend: split on blank lines, merge short paragraphs such as headings into the next one until a passage has a minimum number of words, window only long paragraphs with overlap, and keep the file path and chunk position with each passage.
  2. Write the weighting down before coding it. State your term frequency (raw count, or 1 + log(tf) so a word repeated ten times doesn’t dominate), your inverse document frequency (smoothed, and say what the smoothing does), and cosine similarity over length-normalized vectors. Name each choice; that is what makes the code reviewable.
  3. Tokenize plainly and say what you gave up. Lowercase, split on non-alphanumerics, a short stopword list. No stemming, so “refund” and “refunds” don’t match; say that out loud.
  4. Score only candidates. An inverted index from term to passages means a query touches only passages that share a term with it. Return the top k with a heap.
  5. Prove it with a small relevance test. Build a fixture where the right answer is obvious and only IDF can produce it, and assert the order. Then name how you’d measure it for real: a labeled set of queries and hit rate at k.

The trap is spending the time on parsing edge cases and never testing relevance. The code can be flawless and the ranking still wrong, and only a test with a known right answer shows it.

GlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineerGlossaryRetrieval-augmented generationAnswering with a model that is given passages retrieved from a document collection as context.More on Retrieval-augmented generationGlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on Agent

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • A word appears in every passage. What weight does it get, and is that what you want?
  • How would you know a change to the chunker made retrieval better or worse?
  • The folder grows to millions of passages. What breaks first?

Where answers go wrong

  • Scores whole files instead of passages, or compares raw term counts without normalizing for length, so the longest file wins every query.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Compare with the model answer

Model answer

“I’ll index passages, not files, because that’s what the caller gets back. Tokens are lowercased alphanumeric runs minus a few stopwords; no stemming, so plural and singular are different terms, which I’d revisit. Term frequency is sublinear, 1 + ln(tf). IDF is ln((1 + N) / (1 + df)) + 1. The outer plus-one means a term in every passage gets weight 1.0, the lowest this formula gives, instead of zero, so it can still break ties. The inner plus-ones act as if an extra passage contained every term, which damps extreme weights on a tiny corpus; they would also avoid dividing by zero for a term no passage contains, but I drop those anyway. Vectors are L2-normalized, so the dot product is cosine similarity and a long passage doesn’t win by being long.”

Sublinear term frequency and cosine normalization are as in Manning, Raghavan and Schütze, Introduction to Information Retrieval, chapter 6, whose IDF is the unsmoothed log(N / df). The smoothed form is the one scikit-learn uses with smooth_idf=True (TfidfTransformer documentation).

“Cosine also rewards short passages: a passage that is only the heading ‘Refunds’ would score perfectly against a refunds query and answer nothing. So the chunker merges short paragraphs forward into the next one.”

import heapq
import math
import re
from collections import Counter
from pathlib import Path

# Letters and digits in any script
TOKEN = re.compile(r"[^\W_]+")
STOP = {"the", "a", "an", "and", "or", "of", "to", "in",
        "is", "for", "on", "with"}


def tokens(text: str) -> list[str]:
    return [t for t in TOKEN.findall(text.lower())
            if t not in STOP]


def chunk(text: str, min_words: int = 25,
          max_words: int = 120, overlap: int = 30) -> list[str]:
    """Merge short paragraphs (headings, sign-offs) forward;
    window long ones."""
    out: list[str] = []
    buf: list[str] = []
    for para in re.split(r"\n\s*\n", text):
        buf += para.split()
        if len(buf) < min_words:
            # A bare heading waits for its paragraph
            continue
        step = max_words - overlap
        last = max(len(buf) - overlap, 1)
        for start in range(0, last, step):
            out.append(" ".join(buf[start:start + max_words]))
        buf = []
    if buf:
        # A short tail joins the last passage
        if out:
            out[-1] += " " + " ".join(buf)
        else:
            out.append(" ".join(buf))
    return out


class Index:
    def __init__(self, folder: Path):
        self.passages: list[tuple[str, int, str]] = []
        for path in sorted(folder.rglob("*.txt")):
            text = path.read_text(encoding="utf-8",
                                  errors="replace")
            # Relative path, so two faq.txt files stay distinct
            rel = str(path.relative_to(folder))
            self.passages += [(rel, i, c)
                              for i, c in enumerate(chunk(text))]

        counts = [Counter(tokens(text))
                  for _, _, text in self.passages]
        n = len(counts)
        df = Counter(term for tf in counts for term in tf)
        self.idf = {t: math.log((1 + n) / (1 + d)) + 1
                    for t, d in df.items()}

        self.vectors = [self._weigh(tf) for tf in counts]
        self.postings: dict[str, list[int]] = {}
        for i, vec in enumerate(self.vectors):
            for term in vec:
                self.postings.setdefault(term, []).append(i)

    def _weigh(self, tf: Counter) -> dict[str, float]:
        vec = {t: (1 + math.log(c)) * self.idf[t]
               for t, c in tf.items() if t in self.idf}
        norm = math.sqrt(sum(w * w for w in vec.values()))
        if not norm:
            return {}
        return {t: w / norm for t, w in vec.items()}

    def search(self, query: str, k: int = 5):
        """Top k as (score, path, chunk index, text)."""
        # Terms the corpus has never seen drop out here
        q = self._weigh(Counter(tokens(query)))
        candidates = {i for t in q for i in self.postings[t]}
        scored = (
            (sum(w * self.vectors[i].get(t, 0.0)
                 for t, w in q.items()), i)
            for i in candidates
        )
        top = heapq.nlargest(k, scored)
        return [(round(s, 4), *self.passages[i]) for s, i in top]

“Each result is score, path, chunk index and text. The path and chunk index are what a generated answer cites, so a reader can open the exact passage. Query terms the corpus has never seen have no IDF, so they drop out, and a query made only of them returns an empty list rather than raising. If the files had marked headings, I’d also prefix every window of a long section with its heading. The test I care most about is a ranking test where I know the answer:”

def test_rare_term_outranks_repeated_common_term(tmp_path):
    files = {
        "refunds.txt": "Refunds go back to the original card "
                       "after billing approves them.",
        "team.txt": "For billing, ask billing support "
                    "or the billing team.",
        "invoices.txt": "Billing sends invoices monthly.",
        "plans.txt": "Billing plans renew each year.",
    }
    for name, text in files.items():
        (tmp_path / name).write_text(text)
    top = Index(tmp_path).search("billing refunds", k=2)
    assert [path for _, path, _, _ in top] == [
        "refunds.txt", "team.txt"]
    # A wide margin, so a small tokenizer change can't flip it
    assert top[0][0] > 1.5 * top[1][0]

“‘billing’ is in every file, so IDF shrinks it, and the one file with ‘refunds’ wins even though team.txt says ‘billing’ three times. Run on that folder, the search returns:”

>>> for hit in index.search("billing refunds", k=2):
...     print(hit)
...
(0.3922, 'refunds.txt', 0, 'Refunds go back to the original card after billing approves them.')
(0.2472, 'team.txt', 0, 'For billing, ask billing support or the billing team.')
>>> index.search("chargeback")
[]

“Remove IDF and team.txt wins, 0.5454 to 0.4714, which is the regression this catches. The margin assertion is there because a test that passes by a hair can flip on a harmless stopword change, and then it no longer proves anything.

For real use I’d collect a few dozen questions from the people who’ll ask them, label the passage that answers each, and track hit rate at k every time I touch the chunker or the tokenizer. At larger scale the first thing to break is memory, because every vector sits in a Python dict; I’d move postings to disk or into a search engine with BM25, which saturates term frequency and makes length normalization tunable, and I’d reach for embeddings when queries and passages use different words for the same thing.

In a customer deployment, each passage carries its source file’s access list, and I filter candidates by the caller’s permissions before scoring, never after, so a restricted passage can’t shape the results or leak through a snippet. And when a model answers from these passages, I pass each one with its path and chunk index and require the answer to cite them, then measure retrieval (hit rate at k) and the answer’s faithfulness to its passages separately, so I know which half failed.”

Next, in Pro

In Pro, the BM25 question implements BM25, runs it beside TF-IDF on the same queries, and shows where the rankings differ and why.