A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

Two things are graded: transitions right at the edges, and proof without sleeping. Draw the state machine first.

  1. Pin the contract. “call(fn) runs fn or raises CircuitOpen at once. I’ll inject the clock.” Then ask what counts as a failure. Timeouts, connection errors, 5xx and 429, the dependency shedding load, do. Other 4xx responses are the caller’s mistake, not the dependency’s. Martin Fowler’s description makes the general point: not all errors should trip the circuit.
  2. Draw three states and every arrow. Closed to open when the failure rule fires. Code consecutive failures; name the production version: a failure rate over a window, with a minimum call count so a quiet night’s few failures can’t trip it. Open to half-open once the timeout passes, checked on the next call, never by a timer thread. Half-open admits one probe: success closes the breaker and clears the counts; failure reopens with a fresh timeout.
  3. Say what everyone else gets. Every other call fails fast while the probe runs; admitting them is the thundering herd the breaker exists to prevent.
  4. Guard the edges in code. Report the probe’s result in finally, so an exception or cancellation can’t strand the breaker in half-open. A hung probe never reaches finally, so give it a deadline: the first call after now + probe_timeout counts it failed and reopens. Number each state change and ignore results from calls started under an earlier number. Hold the lock to admit and to report, never while fn runs.
  5. Test by moving the fake clock. Trip at exactly the threshold, fail fast without calling fn, admit one probe after the timeout, reopen on a failed or hung probe, close on a good one.

The trap is treating open-to-half-open as a sleep or a background timer. Both make the class untestable and racy. Practice it in the drill.

GlossaryCircuit breakerA pattern that stops calling a failing dependency for a while, then probes it, so failures do not cascade.More on Circuit breaker

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • A probe is still running when a second request arrives in half-open. What does that request get?
  • A call that started while the breaker was closed fails after it has opened. Does that failure count?
  • The dependency returns a burst of HTTP 404 responses for records that do not exist. Should the breaker open?
  • You keep one breaker per dependency, and the service runs on many instances, each with its own. Is that a problem, and would you share the state?

Where answers go wrong

  • Letting every waiting request through in half-open, so the recovering service takes the full load at once and falls over again.
  • A probe that raises or is canceled without reporting back, which leaves the breaker stuck in half-open and refusing everything.
  • Counting client errors as failures, so one caller’s bad requests cut every caller off from a healthy service.
  • Holding the lock while the wrapped call runs, so every request to the dependency waits in single file behind the slowest one.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Model answer

“Here’s the contract I’ll build to: call(fn) either runs fn or raises CircuitOpen at once, with a retry_after. It opens after failure_threshold consecutive dependency failures, and after reset_timeout it lets exactly one probe through.