A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
The difference from shipping to your own production is that the customer owns the environment, the change window and the consequences. Structure the answer around that.
- Ask what makes it risky and who holds the keys. Is it code, configuration, a schema change, a data , a prompt or model? Do you deploy, or does their platform team? What change process do they run, and can you see their metrics?
- Make it reversible before you ship it. Split the irreversible parts out: expand the schema first, deploy code that works with both shapes, contract later, a parallel change. Put the new behavior behind a flag. Then rehearse the rollback in a copy of their environment; a rollback nobody has run is a hope.
- Define healthy before you start. The signals, the thresholds, the baseline to compare against, who is watching and what triggers an abort without a meeting.
- Stage it. Shadow traffic or a dark launch, then one site or a small group of users, then more, with time to settle between steps. Compare canary against control, as the SRE Workbook’s chapter on canarying describes.
- Communicate as part of the change. A change request in their format: what, when, risk, how you’ll know, how you’ll undo it and how long that takes, and who to call. Get the rollback approved with the change, so undoing it never waits for their process. Their on-call is on the call. Afterwards, a short note of what happened.
Say the rollback’s limits out loud: “redeploy the old version” doesn’t undo a migration, and turning a flag off doesn’t fix data the new code already wrote.
Change windows, canaries and flags, and rolling back code, config, prompts and data in a customer’s environment are on the syllabus of Enterprise system design for FDEs, in Pro. Its first lesson, enterprise design is different, is free.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- The change includes a schema migration. What does rollback mean now?
- You have no access to the customer’s metrics. How do you know the canary is healthy?
- Halfway through, the customer’s change window closes. What do you do?
Where answers go wrong
- A rollback plan of “redeploy the previous version” for a change that also migrates data or rewrites configuration, which the redeploy does not undo.
- Planning the rollout as if it were your own production, with no mention of the customer’s change process, their on-call or when they hear from you.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
“I’ll use a concrete change: we run a document-processing service inside a hospital network’s Kubernetes cluster, and we’re replacing the PDF parser and adding a column to the extracted-records table in their Postgres. Their platform team applies our manifests, and changes go through their weekly change board.”
“First, what’s risky. The parser can extract different text from the same document, which changes everything downstream. The column is a schema change on a database I don’t own. And I can’t see their dashboards directly. So my plan is built around three questions: can I undo it, will I know it’s bad, and does the customer know what’s happening.”
“Reversible. I split it into three releases. Each can be rolled back on its own, and none of them deletes data.”
“Release one adds a nullable parser_version column, so every record the new code writes says which parser produced it. Old code ignores the column, as long as nothing reads SELECT * into a fixed struct or copies rows with INSERT ... SELECT * into a table that lacks the column, which I check in their consumers first. The column is there so a bad rollout is a query plus a reprocess, not a hunt. Adding a nullable column is instant, but it needs a brief ACCESS EXCLUSIVE lock, which conflicts with every other lock on the table, even the one a plain SELECT takes (PostgreSQL table-level locks). While the ALTER waits for that lock, every query that arrives after it waits too. So I set a lock timeout and retry, rather than let it queue behind one of their long reports and block every read and write on the table:”
SET lock_timeout = '3s';
-- NULL means the old code wrote the row
ALTER TABLE extracted_records ADD COLUMN parser_version text;
-- after a rollback: every row the new parser wrote
SELECT document_id
FROM extracted_records
WHERE parser_version = 'v2';
“Release two ships the new parser behind a flag, parser.version, defaulting to the old one. It can also run the new parser in shadow, storing its output in a separate table for comparison. That table is a second copy of patient data, so it’s in the change request with who can read it and when it’s dropped, which is at the end of the rollout, and their privacy officer signs off before the shadow week starts.”
“Release three, weeks later, drops the old parser. Before the change board, I rehearse the rollback of release two on our staging copy of their cluster, with their Postgres version, and time it.”
“Healthy, defined up front. For the parser: documents processed per hour, failure rate, and field-level agreement between old and new output on the shadow sample, with a threshold I set from a replay of 2,000 of last month’s documents. For the service: error rate and p99 latency against the previous week.”
“Here’s how I set the threshold. Rerunning the old parser on the same 2,000 documents agrees with itself on 99.2% of fields, because its OCR step doesn’t return exactly the same text on every run over a scanned page; that’s the noise floor. The new parser agrees with the old on 97.6%. I read the gap of 1.6 points by hand, and most of it is the new parser getting dates right that the old one missed. So the abort line is 96%, and on the fields billing reads, any disagreement above the noise floor aborts on its own.”
“Abort criteria are written down: agreement below 96%, disagreement on billing fields above the noise floor, or a failure rate above 2x last week’s for 30 minutes. The change request gets the rollback approved in advance as part of the same change, and the flag lives somewhere their on-call can flip without a new ticket, an admin endpoint or a pre-approved ConfigMap edit, so rolling back never waits for next week’s board. Since I can’t see their metrics, we ship a status endpoint with the abort signals on it, which their on-call and I both watch during each step. Where their policy allows, it also pushes those counts to a dashboard I can see. A daily report covers the days between steps.”
“Staged. Shadow mode for a week: the new parser runs on every document, and only the old one’s output is used. Then the flag goes on for one department’s documents, then half, then all, with at least a day between steps, each step inside their change window.”
“Communication. The change request lists the three releases, the flag, the abort criteria, the rollback steps with the time I measured, and my phone number. Their on-call and I are both online for each step. After each step, their platform lead gets a short message: what changed, what we saw, what’s next.”
“If it goes wrong, flipping the flag stops new damage, but it doesn’t fix records the new parser already wrote. That’s why every record carries parser_version: I can list exactly what the new parser touched, rerun those documents through the old parser, and overwrite the extracted fields. That reprocess job is rehearsed on staging alongside the flag flip, and the change request names which downstream teams read those records and get told if we run it.”
“If I abort, their platform lead and those teams get a message like this within minutes, with the real numbers in it:”
We turned the new parser off at ten past two, half an hour into today’s step, and the old parser is running again. No data was lost. The new parser wrote 312 records while it was on; we’re reprocessing them now, and they’ll be correct by five this afternoon. Next update in an hour.