Retrieval-augmented generation (RAG) answers a question by first retrieving passages from a document collection, then giving them to a language model as context so it answers from them rather than from memory alone. The term comes from Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, which paired a dense retriever with a generator; it now covers any retrieve-then-prompt pipeline.
In an FDE interview
Ramp’s AI Solutions and Cohere’s Agentic Platform FDE postings, as of September 2026, both ask for experience building RAG systems. Source 1Software Engineer, Forward Deployed AI Solutions @ RampPublisherRamp (Ashby job board)Source typecompany job postingSource 2Forward Deployed Engineer, Agentic PlatformPublisherCohere (Ashby job board)Source typecompany job posting Before drawing the pipeline, ask which questions RAG should answer at all: a count or a total over records (“how many claims were denied last quarter”) goes to a SQL query, not to retrieved passages. Ask how the customer marks the current version of a document, because retrieving last year’s policy yields a confident wrong answer.
A strong candidate then splits the system into parsing and chunking, retrieval, and generation, and measures retrieval on its own (did the right passage reach the top results for a labeled set of real questions?) before tuning the prompt, because a passage that was never retrieved cannot be cited. They filter by the user’s access rights at retrieval time, show which passage each claim came from, and say what the system does when nothing relevant is found. With a non-technical stakeholder: “it looks up your documents first and answers only from what it found, so when it is wrong, the first thing we check is whether it found the right page.”
TF-IDF retrieval over a folder builds the retrieval half.