Store two values per key, the token count and the time of the last update, and compute the refill lazily on each request: tokens = min(capacity, tokens + (now - last) * rate); last = now. A request takes its cost in tokens and is rejected or delayed when there are not enough, so bursts up to capacity pass while the long-run rate stays at rate (Stripe, Scaling your API with rate limiters, and the Redis Lua script it links).

One candidate reported, in a repository created in September 2026, that the Quilr AI take-home for a Solutions Engineer / role included token rate limiting with model failover. Source 1Quilr FDE take-homePublisherPalmCoast (GitHub)Source typecandidate’s take-home repository In front of a model API, count model tokens as well as requests, because providers limit both (OpenAI, Rate limits), so keep two buckets per key: one for requests and one for tokens. A request’s token cost is unknown until it finishes, so debit the prompt tokens plus the request’s output cap (max_tokens or its equivalent) up front and refund the difference when the response reports actual usage.

In one process, read time from an injected monotonic clock so tests are exact and wall-clock jumps cannot mint tokens. Across instances, use an atomic Redis Lua script keyed per tenant that reads the Redis server’s TIME, so instances with skewed clocks agree. Reject with HTTP 429 and Retry-After = ceil((cost - tokens) / rate) (RFC 6585, section 4). A request whose cost exceeds capacity can never pass, so reject it outright instead of telling the client to retry. Contrast it with sliding-window limiters: the bucket permits bursts up to its capacity, while a sliding log counts exactly but stores every timestamp, and a sliding-window counter approximates the log with two counters per key. Write one in the token bucket question.

GlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineer