Set the total from what the user needs, then split it among the steps a request passes through, for example retrieval 150ms + rerank 100ms + model 1200ms + overhead 50ms <= 1500ms at p95. Budget the tail (p95 or p99) as well as the median, because a request that fans out waits for its slowest call: with 100 parallel calls that each miss their p99 1% of the time, 1 - 0.99^100 ≈ 0.63 of requests hit at least one slow call (Dean and Barroso, The Tail at Scale).
For a model call, split the step again: time to first token grows with the prompt, and the rest is output tokens times the time per token, so capping or shortening the output is often the cheapest win. In an , steps run in series and their count varies, so cap the steps or the budget is fiction. Summed per-step p95s usually overstate the end-to-end p95; measure the end-to-end tail, and use per-step traces to find which step to cut.
As of September 2026, Cursor’s postings list latency and cost trade-offs among the production quality an FDE owns, and Cohere’s Agentic Platform FDE posting asks for evaluation frameworks that measure agent latency alongside accuracy and safety. Source 1Forward Deployed EngineerPublisherCursor (Ashby job board)Source typecompany job boardSource 2Forward Deployed Engineer, Agentic PlatformPublisherCohere (Ashby job board)Source typecompany job posting When a customer says the assistant is slow, ask for a per-step trace before guessing and look at the model calls first. Then offer moves with their costs (shorten the output, stream the first tokens, cache the stable prompt prefix, route easy requests to a smaller model, send less retrieved context) and let the customer choose between faster and more accurate with numbers from their own system.
Related: service level objective, reranking.