A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
Treat a prompt as a deployable artifact, like a container image: immutable versions, and a separate pointer that says which version runs where. Say that sentence early, because every later answer follows from it.
- Clarify the shape. How many prompts, how many customers, who edits them (engineers, or customer-facing staff without deploy access), and how fast a bad prompt must come out of production.
- The version record. Template text, model and parameters, the input variables it expects, a content hash, its parent, author, change note and the evaluation run that approved it. Versions are never edited.
- Overrides. Decide how a customer’s difference is expressed. Slots in a shared template (tone, policy text, product names) let global fixes flow through; a full fork does not. Allow forks only when slots cannot express the difference, and pin every fork to the base version it came from.
- Release. Moving an environment’s pointer, per customer, after the customer’s evaluation set passes; canary on one customer first.
- Resolution at runtime. Prompt, customer and environment resolve to one version, cached briefly, and every model call logs the versions it used: base, override and a hash of the rendered text.
- Rollback. One pointer move, audited, with a compatibility check on the variables and the model.
The trap is tying a prompt change to an application deploy. Git is a fine place to write prompts; the answer fails when the only way to release or roll back one is to redeploy the app, and when a customer’s wording means a code branch. Say what a rollback is, in one command, without waiting to be asked.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- What does a prompt version record contain?
- One customer has an override. How does a global change reach them?
- Roll back one prompt in production. Walk through it.
Where answers go wrong
- Keeps prompts in code with no per-customer override or rollback.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
I’d treat each prompt like a container image: immutable versions, and a separate pointer per customer and environment that says which version runs. Releasing is moving a pointer, and so is rolling back.
Clarify. I’d ask who edits prompts, whether customers need their own wording, and whether the model and parameters change alongside the text. I’ll assume solutions engineers edit without deploying code, several customers need overrides, and model settings travel with the prompt.
Storage. Two tables, in the same Postgres as the application. The first holds every version ever approved; the second holds the pointers.
create table prompt_version (
prompt_id text not null,
version int not null,
-- '*' for the base prompt
customer_id text not null,
-- a fork: the base version it was copied from
base_version int,
-- null for a slot-only override
template text,
-- slot values filled into the template
slots jsonb not null,
model text not null,
-- temperature, max_tokens
params jsonb not null,
-- variables the template expects
input_schema jsonb not null,
-- hash of what this row stores: template (if any),
-- slots, model and params
content_hash text not null,
author text not null,
note text not null,
-- the evaluation run that approved it
eval_run_id text,
created_at timestamptz not null default now(),
primary key (prompt_id, customer_id, version)
);
create table prompt_release (
prompt_id text not null,
-- '*' is the base
customer_id text not null,
env text not null,
version int not null,
-- slot-only override: hold this base version;
-- null follows the base release
pinned_base_version int,
-- always '*', so the database can check the pin
base_id text generated always as ('*') stored,
moved_by text not null,
moved_at timestamptz not null default now(),
primary key (prompt_id, customer_id, env),
foreign key (prompt_id, customer_id, version)
references prompt_version,
foreign key (prompt_id, base_id, pinned_base_version)
references prompt_version
);
Rows in prompt_version are insert-only: the application’s database role has no UPDATE or DELETE grant on that table. Authors can edit in git; CI runs the evaluation set and inserts an immutable prompt_version row. content_hash covers what the row itself stores, so for a slot-only override it is the slots and settings, not a template; which base they rendered with is recorded per call, below. Releasing is a separate pointer move: an update to prompt_release, which also writes to an audit log. The foreign keys mean a mistyped rollback to a version that does not exist fails at the database, and so does a pin to a base version that does not exist.
Resolution. At runtime, resolve(prompt_id, customer_id, env) looks for a release under the customer’s id, then under '*'. It is cached with ttl=30s, so a rollback reaches every worker within that window without a deploy. A conversation resolves once and keeps that version for its remaining turns, so one transcript never mixes two prompts; an urgent rollback can override the pin. resolve returns the base version, the override version and a hash of the rendered template, and every model call logs all three. A slot-only override is built from two rows, so one version number cannot say which change broke it.
Overrides. Most customers differ in slots, not structure, so most overrides are slot-only rows: they store slots and render with the template of the base version released in that environment. A global fix reaches them on the next base release, and so would a global break, so a base release runs the evaluation set of every customer with a slot-only override and refuses if the new template needs a slot a customer has not filled. A customer whose set fails stays on the old base, through pinned_base_version on its release row, until someone fixes it. A full fork records base_version. When the base moves, a job lists every fork now behind it, runs each customer’s evaluation set against a rebased candidate, and opens a review for the fork’s owner. A fork is never silently rebased.
Release. A new version runs the prompt’s evaluation set, and for overrides the customer’s set too. Passing makes it eligible. A second, named person approves every release to prod. A rollback returns to a version that already passed, so it moves first and that person reviews it after. Each customer can see a change log of their own prompts, because a contract may require notice of changes. The first move goes to one canary customer; for 24h I watch its handoff rate, thumbs-down rate and a graded sample of live answers against the previous version, then release to the rest.
Rollback, walked through. An alert shows wrong answers for one customer. The call logs show northwind’s triage override at version 42 since the last release, on an unchanged base. I run promptctl rollback triage --customer northwind --env prod --to 41. The tool checks that every variable version 41 requires is one the running code still sends. If the code has since renamed or dropped one, it refuses and tells me to roll the code back too; a variable the code added and 41 never uses is harmless. It also checks that 41’s model is still served. If the provider has retired that model, a rollback to 41 is really a new version, so it goes through evaluation like any other. It moves the pointer and writes the audit row. Within the cache window, calls log override version 41. If the log had shown the base changing instead, not northwind’s slots, the override pointer would be the wrong lever: I set northwind’s pinned_base_version to the previous base and roll the base back for everyone only if other customers show the same failure. Then I add the failing conversation to the evaluation set, so override version 43 cannot repeat it.
First slice. The two tables, resolve with its cache, the versions in every log line, and promptctl rollback, for the one prompt with the most overrides. The editing UI and fork rebasing come later.