A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
The prompt hands you the hard part in its last clause. Anyone can draw ingest, embed, retrieve, generate; a stronger answer draws the boundary first and then shows where permissions and evidence live.
- Clarify what shapes the design. Who asks (clerks, judges’ staff, self-represented litigants at a help desk)? Are some manuals restricted, such as procedures for sealed or juvenile matters? Which cloud, and has the court’s security office approved a model service inside it? Say the point that changes the design: a clerk will paste case details into the question, so the query is a case record too, and every place it goes ( call, model call, logs) has to sit inside the boundary.
- Draw the , then place every box. Identity, API, index, both models, logs: all inside. Name anything that crosses it, and the direction it crosses.
- Ingest by structure. Parse manuals by section, and keep the section number, effective date, source link and access list on every chunk.
- Enforce permissions in the retrieval query. Filter by the user’s groups inside the search, before ranking, using groups from the verified identity token. The model never sees a chunk the user couldn’t open.
- Answer with citations, or refuse. Every claim points to a section a clerk can open. No supporting passage, no answer.
- Evaluate, operate and audit. A question set written with clerks, measured on every change; operations that never need to see a query; an audit log kept under the court’s retention rules.
- Say what ships first. One manual set, one permission split, a pilot with a handful of clerks.
Tripwire: A RAG diagram with no boundary
If your diagram would fit any company’s chatbot, you have not answered this question. Expect to be asked where the permission check is, and “the prompt tells the model” is the wrong answer.
After this one, answer the SSO design question, which covers where the groups your permission filter trusts come from and how fast a leaver loses them, and the citations question for the step where every claim has to point to a section a clerk can open.
Practice it with a customer
In Pro, the claim intake case puts you in front of an insurer’s security architect who wants to know what the model’s service account can touch and where medical details end up, with a launch date already promised.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- Where exactly is each document permission checked?
- The court adds several thousand documents overnight. What happens to the index?
- The clerks’ supervisor asks how you know answers are right. What do you show?
Where answers go wrong
- Draws a generic RAG diagram with no trust boundary or permission check.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
“Before drawing, I’d ask a few things. Who uses it: I’ll assume court staff, not the public. Are any manuals restricted: I’ll assume most are open to all staff and some, like sealed-matter procedures, are limited to certain roles. And what models can run inside the tenancy: either the cloud provider’s managed model service reached over a private endpoint such as AWS PrivateLink, or open-weights models on GPU instances the court runs. And which cloud environment: a government region such as AWS GovCloud (US) or Azure Government? I’d also ask whether the security office treats any of this under the FBI’s CJIS Security Policy, which counts a court as a criminal justice agency and applies to every contractor with access to criminal justice information. If it does, that decides who on our side may ever touch the system, and it’s one reason our team has no standing access. I’d treat the question text as a case record, because clerks will paste case details into it.
“That makes the managed option a real decision. PrivateLink keeps traffic off the public internet, but the service runs outside the court’s network, so with that option the query and prompt do leave the court’s account. I’d put it in front of the security office in writing: region pinned, no prompt retention, including for abuse monitoring, and no training on inputs. If they say no, it’s open weights on the court’s own GPUs.”
“Here is the boundary in one breath. Inside: identity, the document system and its ingest worker, the API, the search index and the audit log, and both models when they run on the court’s own GPUs. Crossing in: signed container images and model weights. Crossing out, and only with the managed option: the query, the prompt and the retrieved chunks, over the private endpoint, to the provider’s service in the court’s region.”
“Ingestion. The document system stays the source of truth. The worker parses each manual by section, and each chunk carries its manual, section number, effective date, a link back and the document’s access list. Manuals are full of rule numbers and form codes, which embeddings match poorly, so the index is hybrid: keyword and vector, merged with reciprocal rank fusion. The query is embedded by the same model as the chunks, so that call carries case details too and sits inside the same boundary.
“Size. A quick estimate before any design choice that depends on scale:”
~50,000 manual sections
-> ~200,000 chunks
x 1,024 dims x 4 bytes
= ~800 MB of vectors:
one node, in memory
Open weights, if the court
says no to the managed option:
70B model at 16-bit
= ~140 GB of weights
-> two 80 GB GPUs at least
8B model at 16-bit
= ~16 GB -> one GPU
“This is small. The hard parts are the boundary and the permissions, not scale.
“Permissions: copied at ingest, checked at three points. At query time, as a filter inside the search, using groups the API takes from the validated token, never from the request body, so ranking only ever sees permitted chunks; filtering after ranking can return nothing, or leak that a restricted section exists. Before generation, for restricted chunks, against the live access list. When a citation is opened, by the document system with the user’s own session. Access lists sync on a schedule, so revocation has a delay I’d state and agree with the court.
“Generation. The prompt holds only permitted chunks, and the answer must cite a section for each claim. A post-check confirms every cited section was in the retrieved set. With nothing relevant retrieved, it says so and links the manual’s index. On a clerk’s screen, from an invented manual:”
Q: Can a fee waiver be filed after
judgment in small claims?
A: Yes, within 30 days of entry of
judgment. [Civil Procedures
Manual 4.12, effective March 2026]
Q: What sentence will this judge
give?
A: The procedure manuals don't cover
this. See the manual index.
“The overnight batch. Ingestion is a queue keyed by document ID and content hash, so reruns are idempotent and unchanged documents are skipped. A chunk with no access list is indexed as restricted to nobody until its list syncs, never as open by default. The runs on its own queue with a concurrency cap, so a large overnight drop finishes late rather than slowing the clerks’ morning questions, and the morning report says how many documents are still pending. A new version’s chunks are written first, then the document’s active version flips, so a question never gets half of each. Superseded versions stay, filtered by effective date. Parse failures go to a report for the manual’s owner.
“How the supervisor knows it’s right. Before launch, senior clerks write questions in the form clerks actually type them, with made-up case details in place of real ones, each paired with the section that answers it, plus some the manuals can’t answer. I measure whether the right section is retrieved, whether citations support the sentences they’re attached to, and whether it refuses the unanswerable ones, and rerun it on every change. In production, every answer logs its citations; clerks flag bad ones, and a clerk reviews a sample each week. I’d show the supervisor that dashboard, including the failures.
“Operating it without seeing it. No SaaS error tracker or APM , because stack traces and support bundles carry query text. Metrics and traces go to the court’s own monitoring, and query text is never in them; only the audit log holds it. Our team ships signed images into the court’s registry and has no standing access. When something breaks, a court engineer runs the diagnostics under a break-glass procedure and shares redacted output, and every such session is logged.
“Audit and retention. The log holds who asked what and which chunks were used. Because queries contain case details, how long we keep them is the court’s records officer’s call, not mine.
“First release: the civil procedure manuals, the open-versus-restricted split, the evaluation set, and a pilot with one clerk’s office.”