A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

A strong answer picks chunk boundaries from the documents and the questions, not from a default. A contract and a support ticket differ in exactly the ways that matter: a contract is long, numbered, cross-referenced and amended; a ticket is short, conversational, and useful mainly for its resolution.

Decide from the queries up:

  1. Start from the queries. Ask what people will ask of each corpus. Contract users ask clause questions (“what is the termination notice period in the Acme master agreement?”). Support users ask “has anyone seen this error before, and what fixed it?”
  2. Parse before you chunk. Recover the structure first: clause numbers, headings, tables and exhibits in contracts; messages, quoted replies and signatures in tickets. If the parse is wrong, no chunk size will fix it.
  3. Set boundaries from that structure. Contracts split at clauses, with the heading path attached. Tickets usually stay whole: one problem and its resolution.
  4. Attach the metadata retrieval needs, and use it. For contracts: counterparty, document type (master agreement or amendment), effective date and clause number. For tickets: product, version and date. Filter on it before vector search.
  5. Test retrieval on its own. Build a small labeled query set per corpus, with labels that point at the source (a clause or a ticket), not at a chunk, and compare candidate strategies on hit rate at the same context budget before any generation step is involved.

Say the test out loud: “I’d try clause-level against fixed windows on the same labeled questions and pick the one that retrieves the right clause more often.” Then name what you’d look at when it fails.

The pitfall below, one fixed token size for both corpora chosen without a test, splits a clause from its exceptions and buries a ticket’s fix under the thread above it.

Parsing, chunking and embeddings are taught in Production AI systems, in Pro. For the coding side of this answer, try Write a text chunker that respects headings and never splits a table: write it yourself, then check the follow-ups.

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • Which chunk boundaries would you use for a contract, and why?
  • Retrieval recall drops on tickets after your change. How do you find out why?
  • Tables span pages in the contracts. What do you do?

Where answers go wrong

  • Picks a fixed token size for both corpora without testing retrieval.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Compare with the model answer

Model answer

“I won’t pick a chunk size first. I’d start from what people ask each corpus, pick the unit that answers it, and then prove the choice on a labeled question set before any generation is involved.”

Contract users ask about a clause; support users ask whether anyone has seen this error and what fixed it. So the unit is the clause for one and the problem-and-fix for the other.

Contracts. I’d split at numbered clauses, because that’s the unit people ask about and cite. Each chunk carries its heading path, so “Notice” under “Termination” isn’t confused with “Notice” under “Payment”. A clause over 800 tokens splits at its sub-clauses, and each piece repeats the path. Definitions get special handling: when a chunk uses a defined term like “Confidential Information”, I attach that definition’s clause ID in metadata and fetch it alongside the chunk. Amendments are separate documents with a supersedes link, so retrieval can prefer the governing version. At query time I filter to the named counterparty and the governing version before vector search, and run BM25 alongside it, because “section 12.3” and “Confidential Information” are exact-match queries that embeddings handle poorly. Contracts are need-to-know, so each chunk also carries the document’s access groups from the contract system, and the query filters on the user’s groups before vector search, the same way it filters on counterparty. One deal team must never retrieve another counterparty’s terms, and that is the first thing the customer’s security review will ask about.

import re
from dataclasses import dataclass

CLAUSE = re.compile(r"^(\d+(?:\.\d+)*)\.?\s+([A-Z][^\n]{0,80})$", re.M)

@dataclass
class Chunk:
    text: str
    meta: dict

def is_next(prev: list[int], num: list[int]) -> bool:
    # After 7.2 a heading is 7.2.1, 7.3 or 8. Anything else, such as a list
    # restarting at "1." inside 7.2, is body text, not a heading.
    if not prev or num == prev + [1]:
        return True
    return any(num == prev[:i] + [prev[i] + 1] for i in range(len(prev)))

def chunk_contract(text: str, meta: dict) -> list[Chunk]:
    # meta carries title, counterparty, doc_type, effective_date and acl (the
    # groups allowed to read this document, from the contract system); every
    # chunk inherits it, so the query can filter on the user's groups.
    # Assumes the parser put each clause heading on its own line and indented
    # list items. An unindented "2." inside 1.1 passes is_next as clause 2 and
    # swallows everything after it. A heading run into its text ("12.1 Notice.
    # Either party may ...") is missed if the line is long, and taken whole
    # as the heading if it is short. A numbering gap, such as a clause
    # deleted by amendment or a "3. [Reserved]" line the pattern doesn't
    # match, stops is_next from matching, so every clause after the gap merges
    # into the one before it: log the gaps rather than trust the output.
    # Splitting over-long clauses at sub-clauses is left out here.
    marks, prev = [], []
    for m in CLAUSE.finditer(text):
        num = [int(x) for x in m[1].split(".")]
        if is_next(prev, num):
            marks.append(m)
            prev = num
    path: dict[int, str] = {}
    chunks = []
    for i, m in enumerate(marks):
        end = marks[i + 1].start() if i + 1 < len(marks) else len(text)
        depth = m[1].count(".")
        path = {d: t for d, t in path.items() if d < depth}
        path[depth] = f"{m[1]} {m[2].strip()}"
        body = text[m.end():end].strip()
        if body:  # a heading with only sub-clauses under it adds nothing alone
            header = " > ".join(path[d] for d in sorted(path))
            chunks.append(Chunk(f"{meta['title']} | {header}\n{body}",
                                {**meta, "clause": m[1]}))
    return chunks

Tickets. Most tickets fit in one chunk, and splitting them separates the symptom from the fix. I strip quoted replies, signatures and auto-responses at parse time. Queries look like symptoms, so I embed the subject and first message and carry the resolution as payload. A first message can be as vague as “it’s broken”, so at index time a model writes a one-line problem statement (product, error, symptom) from the whole thread, and I embed that alongside the subject. It’s the same idea as Anthropic’s Contextual Retrieval write-up, which prepends model-written context to each chunk before indexing it. Unresolved tickets stay out of the index, because their last message is a “thanks” or an auto-close notice, not a fix. Very long threads get one chunk per problem-and-answer pair, with the ticket ID in metadata so results can be grouped by ticket.

def chunk_ticket(t: dict) -> Chunk | None:
    if t.get("status") != "solved" or not t.get("resolution"):
        return None
    # "problem" is a one-line statement a model wrote at index time
    # (product, error, symptom), for first messages like "it's broken"
    text = f"{t['subject']}\n{t['problem']}\n{t['messages'][0]['body']}"
    meta = {k: t.get(k) for k in ("id", "product", "version", "closed_at")}
    return Chunk(text, {**meta, "resolution": t["resolution"]})

The test. Before choosing, I’d have the customer’s contract managers and support leads write real questions and mark the clause or ticket that answers each one. The labels point at the source, not at a chunk, so the same set scores every strategy: search returns the source keys its top chunks cover, such as acme-msa#12.3 or a ticket ID, and a fixed window that spans two clauses reports both. For contracts I score recall of every labeled clause, since an answer that misses the exception is wrong. I compare strategies at the same context budget, not the same k, or large windows win by covering more text. So search fills a token budget from the top of its ranking.

def hit_rate(labeled: list[tuple[str, set[str]]], search, budget: int = 2000) -> float:
    # search(q, budget) returns the source keys covered by the top-ranked
    # chunks that fit in `budget` tokens, so every strategy gets equal context
    return sum(bool(search(q, budget) & gold) for q, gold in labeled) / len(labeled)

def full_recall(labeled: list[tuple[str, set[str]]], search, budget: int = 2000) -> float:
    return sum(gold <= search(q, budget) for q, gold in labeled) / len(labeled)

Tables that span pages. The parser stitches them back together, using the repeated header row to detect a continuation, then I chunk the table as a unit, or in row groups that each repeat the header. A pricing row without its column names can’t answer anything.

If ticket recall drops after a change, I’d diff the misses: for each query that used to hit, which ticket ranks first now, and where did the gold ticket go? The culprits I check first are a parser change that dropped resolutions (which now drops those tickets from the index), a new split cutting symptom from fix, or version metadata filtering out the right ticket.

Chunk size follows the unit of meaning; how much context I pass the model is a separate lever. I pass few passages and put the strongest first, because the Lost in the Middle study by Liu and colleagues found that the models it tested used information in the middle of a long context less reliably than information at its start or end. Newer models may do better, so I’d check on the customer’s model rather than assume it.