In this post12 sections
- What a strong RAG design answer has to cover
- The whole answer as a numbered flow
- Start with the documents, the users and the permissions
- Ingestion and chunking
- Retrieval: hybrid search, then reranking
- Permissions filtered at retrieval time, not in the prompt
- Answers with citations the user can check
- Evaluation, cost and latency budgets
- The first version you would ship, and the follow-ups to expect
- Questions people ask
- Keep reading
- More from the blog
You drilled the URL shortener and the news feed. For an AI role, “design a system” is a prompt worth knowing cold, because several AI FDE postings ask for exactly that work. Source 1Forward Deployed Engineer, Agentic PlatformPublisherCohere (Ashby job board)Source typecompany job postingSource 2Software Engineer, Forward Deployed AI Solutions @ RampPublisherRamp (Ashby job board)Source typecompany job postingSource 3AI Engineer – Forward Deployed Engineering (AI FDE)PublisherDatabricks (careers site / Greenhouse)Source typecompany job posting This post is one worked answer; the FDE interview guide covers every other round.
The short answer: a strong RAG system design answer starts from the customer’s documents, users and permissions, not from the vector database. Then it walks the pipeline in order: ingestion and chunking, with a , a permission filter inside retrieval, answers with citations the user can check, evaluation, and cost and latency budgets. It ends by naming what the first version ships and what it leaves out.
What a strong RAG design answer has to cover
The prompt is short on purpose. “Design a RAG system over our internal documents” leaves out who the users are, which documents they may see and what a wrong answer costs. Treat filling those gaps as the job: anyone can draw the RAG loop of retrieve, then generate; what makes an answer strong is everything around it.
As of September 2026, Cohere, Ramp, Databricks and Scale AI FDE postings name RAG or retrieval systems in the experience they ask for, and Glean’s lists retrieval systems among its nice-to-haves. Source 1Forward Deployed Engineer, Agentic PlatformPublisherCohere (Ashby job board)Source typecompany job postingSource 2Software Engineer, Forward Deployed AI Solutions @ RampPublisherRamp (Ashby job board)Source typecompany job postingSource 3AI Engineer – Forward Deployed Engineering (AI FDE)PublisherDatabricks (careers site / Greenhouse)Source typecompany job postingSource 4Frontier Agents Engineer (Forward Deployed Engineering)PublisherScale AI (Greenhouse)Source typecompany job postingSource 5Founding Forward Deployed EngineerPublisherGlean (Greenhouse job board)Source typecompany job posting Scale AI’s Frontier Agents role also deploys evaluation harnesses built on golden datasets and regression suites. Source 4Frontier Agents Engineer (Forward Deployed Engineering)PublisherScale AI (Greenhouse)Source typecompany job posting
One candidate reported, in May 2026, that the role-related knowledge round of their Google FDE (GenAI) loop covered multi-agent systems, RAG architecture and GenAI implementation. Source 6Rejected by Google (L3 FDE) after "killing" the domain round and solving the coding prompt. Feeling blindsided. (post by u/Superb_Pen9988)PublisherReddit r/FAANGrecruitingSource typecandidate report on Reddit That is a single report, not a pattern. The flow below is how we teach the answer; it is our method, not a rubric any interviewer has published.
The whole answer as a numbered flow
In our method, you say this plan up front, so the interviewer can steer:
- Scope. The documents, the users, the permissions, and what a wrong answer costs.
- Ingest. Parse by structure, keep metadata and the access list on every chunk, update incrementally.
- Chunk. Split on the document’s own units: sections, clauses, a ticket and its fix.
- Retrieve. Keyword and vector search together, merged, filtered by the user’s permissions.
- Rerank. Reorder a wide candidate set with a cross-encoder and keep a few.
- Answer. Generate only from those passages, with a citation per claim, or say the documents don’t answer it.
- Evaluate. Score retrieval and answers separately on a labeled set, before launch and on every change.
- Budget. Latency per stage and cost per query, stated as numbers you will check.
Start with the documents, the users and the permissions
Take a concrete prompt: an insurer wants its claims handlers and underwriters to query policy wordings, underwriting guidelines and internal memos. Before you draw a box, ask these, and say what each answer changes:
- “What documents, in what formats, and how often do they change?” Scanned PDFs mean OCR and a parse-quality check. Versioned wordings mean effective dates on every chunk, or you answer from last year’s policy.
- “Who asks, and what do they do with the answer?” A handler who pays a claim on it needs citations and refusals.
- “Can everyone see every document?” At an insurer, no: pricing memos and underwriting guidelines stay with underwriting. That decides where the permission check lives.
- “What does a wrong answer cost?” If it costs money or a complaint, the system refuses when it cannot cite.
- “Where must it run?” If documents may not leave the customer’s cloud, the models, the index and the logs move inside their trust boundary.
When time is short, state assumptions instead: “I’ll assume staff users, a document system with access lists, some restricted content, and that a wrong answer has a cost. Stop me if any of that is wrong.”
The free lesson Enterprise system design is not ‘design a social network’ drills this opening, and labeling assumptions and surfacing failure modes covers writing each assumption down so you can revisit it. The lessons on assumptions and on explaining AI limits come with Pro, which starts with a 7-day free trial. The pricing page has the details.
Ingestion and chunking
Say four things.
Parse before you chunk. Recover headings, section numbers and tables. No chunk size fixes a bad parse.
Split on the document’s own units. A policy wording splits at its numbered sections, and each chunk carries its heading path, such as Exclusions > 4.2 Flood, so an exclusion never floats free of its section. A short memo can stay whole. Chunking contracts compared with support tickets works through two corpora, with code.
Keep metadata on every chunk. Document ID, version, effective date, section, a source link, and the access list copied from the source system at ingest. That list makes the permission filter possible.
Update incrementally. Key the pipeline on document ID and a content hash, so a rerun skips unchanged files and a changed file replaces its old chunks. Write a new version’s chunks first, then flip the document’s active version, so no query sees half of each. A document whose access list hasn’t synced yet is indexed as visible to no one, never as open.
Embed chunks and queries with the same embedding model. Changing that model means re-embedding the whole index, so name the cost early.
The mistake: “I’d chunk at 512 tokens with some overlap” as your whole ingestion story. That is a default, not a design.
Retrieval: hybrid search, then reranking
Vector search matches meaning, and it matches exact strings poorly. Insurance documents are full of exact strings: policy numbers, form codes, clause references. So run keyword search (BM25) alongside vector search, the combination called hybrid search, and merge the two ranked lists.
Reciprocal rank fusion is a simple way to merge them. It uses ranks only, so it ignores the two systems’ scores, which aren’t comparable (Cormack, Clarke and Buettcher):
def rrf(lists, k=60):
s = {}
for ranked in lists:
for r, d in enumerate(ranked, 1):
s[d] = s.get(d, 0) + 1/(k + r)
return sorted(s, key=s.get,
reverse=True)
kw = ["w4.2", "g9", "m17"]
vec = ["m17", "w4.2", "faq3"]
print(rrf([kw, vec]))
# ['w4.2', 'm17', 'g9', 'faq3']
Here w4.2 is the policy wording’s section 4.2, g9 a guideline, m17 a memo. A passage both lists rank well rises to the top. The constant k=60 is the paper’s value.
Then add a reranker. Retrieve a wide candidate set cheaply, score each query and passage pair with a cross-encoder that reads the two together, and keep the best few. It is too slow for the whole corpus, so it only sees candidates.
Say its limit before you’re asked: a reranker only reorders what retrieval found. If the right passage never made the candidate set, the fix is recall. The question retrieval finds the right document, but ranks it too low is the follow-up to rehearse, with the evaluation code.
The words to use: “Hybrid retrieval for recall, a reranker for precision at the top, and I’ll measure each on its own.”
Permissions filtered at retrieval time, not in the prompt
In our method, this is the section to get exactly right. The wrong design retrieves from everything and tells the model “don’t show underwriting content to claims staff”. The model has now read the restricted text, and a cleverly worded question, or an instruction planted in another retrieved document, can get it out.
The right design filters inside the search, before ranking:
- Take the user’s identity and groups from the verified sign-in token, never from the request body.
- Pass them into the search as a filter, so it returns only chunks whose access list includes the user or one of their groups, and has no deny entry for them:
allowed & ({user.id} | user.groups)non-empty. Mirror the source system’s own rules, including per-user shares and denies. - Rank and rerank only what passed. The model never sees text the user couldn’t open.
- Check again when a citation is opened, with the user’s own session in the source system.
Filtering after ranking has two bugs: the top results may all be restricted, leaving nothing, and the gap can reveal that a restricted document exists.
Then name the delay: access lists sync on a schedule, so revoked access lingers until the next sync, and you agree that delay with the customer’s security team. If you cache answers, put the permission set in the cache key, or one person’s answer reaches another.
Vendors state the same principle for their own AI features: Rippling’s March 2026 announcement says Rippling AI inherits the user’s existing roles and access permissions, so a manager sees their team and an HR admin sees more. Source 7Rippling AI: Built to Do the Work, Not Just Talk About ItPublisherRipplingSource typecompany blog
The sentence never to say
“The system prompt tells the model which documents the user may see.” An instruction can be talked out of its job. A filter in the retrieval query cannot.
Answers with citations the user can check
The customer’s real question is “can I trust this?”, and a citation lets a user check. A customer quote on Microsoft’s Frontier Company page credits Microsoft FDEs with co-engineering LSEG’s Workspace AI Search, which combines RAG, and frontier LLMs “to deliver grounded, cited answers”. Source 8Microsoft Frontier CompanyPublisherMicrosoftSource typecompany website
Make citations data your code checks, not prose the model writes:
- Put each retrieved passage in the prompt under a short ID your code owns, and keep the mapping in code.
- Ask for structured output: each statement, the passage IDs that support it, and a short quote.
- Check every response before display. The ID was retrieved for this request, the quote appears in that passage, and the passage supports the statement, which takes an entailment model or an LLM judge calibrated against human labels.
- Render titles and links from your source store, never from the model’s text.
- When nothing relevant comes back, say so and show the closest passages.
Raise the hard case yourself: the passage says the policy covers water damage from a burst pipe, and the answer says it covers flooding. The ID and quote checks pass; only the support check catches it. Making an assistant cite its sources, and checking they’re real has a validator to talk through.
Evaluation, cost and latency budgets
“How do you know it works?” is coming, so answer it inside the design. Score retrieval and answers separately: a bad answer from good passages is a different bug from good reasoning over the wrong passages.
- The set. A golden set of real questions written with handlers and underwriters, each labeled with the section that answers it, not a chunk ID, so it survives a chunking change. Add questions the documents can’t answer, and questions from users who lack access to the answer.
- Retrieval. Is the right section in the top results? Report hit rate at
kbefore and after every change. - Answers. Is the answer correct, is each claim supported by its cited passage, and did it refuse the questions it couldn’t answer?
- Permissions. Did any restricted passage reach a user without access? Run this as a test on every change, not as a metric: any leak blocks the release.
- After launch. Log questions, retrieved IDs and citations, let users flag bad answers, and review a weekly sample.
How to answer ‘how do you know it works?’ goes deeper on this half. Explaining the errors that remain to a claims director is its own skill, covered in explaining AI limits to non-technical leaders.
Latency. Give a latency budget per stage. These are numbers you propose and then measure, not facts, and the two “shown” rows are running totals:
| Stage | p95 budget |
|---|---|
| Embed the query | 50 ms |
| Filtered hybrid search | 150 ms |
| Rerank candidates | 250 ms |
| Passages shown | 450 ms |
| Model, full answer | 3 s |
| Citation check | 500 ms |
| Checked answer shown | 4 s |
The reranker is worth its time in a chat assistant, not in search as you type. Checking citations before display means the answer appears in one piece rather than streaming, so show the retrieved passages at about half a second while the answer and its check run. If the customer wants streaming, stream it and mark statements as checked once each passes.
Cost. Write the formula, not a guess. Cost per query is the sum of:
- the query;
- reranking the candidates;
- input tokens times the input price;
- output tokens times the output price;
- the citation check: an entailment model per statement, or a judge call with the statements and their passages.
Input tokens are the instructions, the passages and the question, so the biggest lever is how many passages you send. Others: cache the fixed part of the prompt where the provider supports it, send simple questions to a smaller model, and cache repeated answers with the permission set in the key. A small entailment model keeps the check cheap; save the LLM judge for the statements the small model is unsure about.
The first version you would ship, and the follow-ups to expect
Close by scoping. For the insurer: the claims-handling guidelines only, the claims and underwriting access split, hybrid retrieval with a reranker, cited answers that refuse when they can’t cite, the , and a pilot with a single claims team.
Then say what it leaves out, and why:
- Scanned archives, until parse quality is measured on a sample.
- Agents that take actions, until the answers are trusted.
- Fine-tuning, because new documents mean a re-index for retrieval but a new training run for a tuned model, and a tuned model can’t cite its source. If the interviewer pushes, fine-tuning versus RAG in the interview takes that question on its own.
Follow-ups to expect, with a first line to say:
- “Thousands of new documents land overnight.” “Ingestion is a queue keyed by document and hash, embedding has its own concurrency cap, and new chunks go live when the active version flips.”
- “A user saw a document they shouldn’t have.” “First I turn off that path. Then I check the access list sync, the filter and the cache key, in that order, against the audit log.”
- “It has to run inside our cloud.” “Then the index, both models and the logs sit inside the boundary, and I name everything that crosses it.” The court retrieval question is that version in full, and the enterprise system design walkthrough draws its boundary and rollback.
- “Why not a long-context model and no retrieval?” “The corpus doesn’t fit, and if it did, every query would pay for all of it. I’d still have to build a per-user context to respect permissions, which is retrieval with extra steps, and I’d still need to show which passage the answer came from.”
Before your design round
- Say the numbered flow from memory, step by step.
- Ask the scoping questions and say what each answer changes.
- Explain in two sentences why the permission filter goes inside the search.
- Give a per stage and the cost formula.
- Name your first version and what it leaves out.
Now rehearse. Set a timer and design this out loud with the eight-step flow: a state court wants its clerks to ask questions of procedure manuals and standing orders, and nothing may leave the court’s own cloud. Then compare yours with the model answer to the court retrieval question.
To practice scoping against a customer who answers back, run the free practice case; it needs a sign-in.
Questions people ask
What should a RAG system design answer cover?
Ingestion and chunking, retrieval (hybrid keyword and vector search with a reranker is a sound default), permission filtering, answer generation with citations, evaluation, and cost and latency. Start from the customer’s documents and users, and say what the first version leaves out.
Should document permissions be enforced in the prompt?
No. Filter by the user’s permissions when you retrieve, so text the user cannot see never reaches the model. An instruction in the prompt can be ignored or overridden by other text; a filter in the retrieval query cannot.
How do you evaluate a RAG system?
Score retrieval and answers separately. For retrieval, check whether the right passage appears in the top results on a labeled set of real questions. For answers, check correctness and whether each claim is supported by its cited passage, and keep tracking both after launch.
Keep reading
Lessons
Questions
- How do you choose a chunking strategy for contracts compared with support tickets?
- Retrieval finds the right document in the top twenty but not the top three. What do you do?
- How do you make an assistant cite its sources, and how do you check the citations are real?
- Design a retrieval assistant over a state court system’s procedure manuals, deployed inside the court’s own cloud tenancy, with no case records leaving it.
More from the blog
Interview rounds
Agentic system design interview: tools, permissions, stuck agents and hand-off to a person
Design a safe agent in the interview: tool allowlists, scoped credentials, loop and cost budgets, idempotent actions, human approval and evals.
Interview rounds
Enterprise system design interview: a worked FDE answer, from constraints to rollback
You prepared to shard a timeline and got asked to deploy inside a customer’s cloud. A worked enterprise design, from constraints to rollback.
Interview rounds
Fine-tuning vs RAG vs prompting: the decision an interviewer wants to hear you make
When a customer says ‘just fine-tune it’: how to choose between fine-tuning, RAG and prompting, with worked scenarios and the words to explain your call.