In this post12 sections
- The short answer
- What voice agent take-homes have asked for
- Why a chat habit fails on a phone call
- Latency: set a budget for every turn
- Turn-taking, interruptions and silence
- When a tool call fails mid-call
- Hand-off to a person
- The metrics dashboard and the demo
- Practice the customer, not just the build
- Questions people ask
- Keep reading
- More from the blog
The take-home arrived and it says “voice ”. You have built chat agents before, but a phone call gives you nowhere to hide: the caller cannot scroll back, will not wait through silence, and talks over your agent mid-sentence. This post covers what voice agent take-homes have asked for and what your submission has to get right; the FDE interview guide covers the rest of the loop around it.
The short answer
The voice take-homes candidates have posted ask you to build, or prompt, an agent that does one phone task for a customer scenario, and then to show that it works. One candidate reported a HappyRobot brief that asks for an inbound carrier-sales agent, a self-built metrics dashboard, an email to the prospect and a 5-minute video. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository A second candidate posted a voice brief that asks only for a system prompt and a Slack message to the customer. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository
Whatever the brief, a call exposes what a chat demo hides. Show a , sensible turn-taking, a plan for failed tool calls, a hand-off to a person, and metrics the customer would read.
What voice agent take-homes have asked for
These are the briefs we could find in public, all from candidates’ repos or, in one case, an account whose company link is inferred. They are examples, not a census.
Inbound carrier sales on HappyRobot. One candidate, IanDalton on GitHub, posted the brief “ Technical Challenge: Inbound Carrier Sales” in a June 2026 repo: you present a proof of concept built on the HappyRobot platform to a customer played by the interviewer. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository In IanDalton’s brief, the agent verifies a carrier’s MC number with the FMCSA API, pitches matching loads, negotiates up to three counter-offers, extracts the offer data and classifies outcome and sentiment. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository IanDalton’s brief adds: “Don’t use the platform analytics and rather build something yourself as we want to assess your product vision and build capabilities.” Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository
A May 2026 HappyRobot take-home repo by timbruckdorfer describes a similar build, “wrapped in a custom metrics dashboard that a freight broker would actually want to use.” Source 3HappyRobot FDE Take-Home — Inbound Carrier SalesPublishertimbruckdorfer (GitHub)Source typecandidate’s take-home repository
A timed voice prompt test. The GitHub user hobbes3 posted their submitted assessment in a January 2026 repo named Bland-FDE-assessment, headed “Take-Home Assignment: Voice AI Prompt Engineering”. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository hobbes3’s brief sets a 2-hour maximum and asks for the system prompt for a voice agent that calls dental patients who missed appointments. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository In hobbes3’s brief, the fictional scenario says 40% of patients claim they never made the original appointment. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository hobbes3’s brief also warns: “Don’t use ChatGPT to write this - we can tell.” Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository A brief like this is a prompt test more than a build, and our post on prompt engineering interview tests covers how to work one under the clock.
An HVAC booking agent. A September 2026 repo by othmane-zizi-pro, described as a “Clerk FDE take-home”, builds a voice agent that books AC repair for Ridgeway HVAC against a mocked two-endpoint API. Source 4ridgeway-hvac-voice-agent (repository description)Publisherothmane-zizi-pro (GitHub)Source typecandidate’s take-home repository The repo does not say which company named Clerk set it. othmane-zizi-pro explains the mock: “The API is mocked here because production sits on the customer’s VPN.” Source 4ridgeway-hvac-voice-agent (repository description)Publisherothmane-zizi-pro (GitHub)Source typecandidate’s take-home repository That one line is the FDE job in miniature.
A law-firm intake agent. A July 2026 repo by AryaanSheth, described as “General Magic FDE Take Home Law Firm VIP”, builds a voice-driven client-intake agent for a law firm; the README does not reproduce the prompt. Source 5Voice-Intake-Pipeline (repository description)PublisherAryaanSheth (GitHub)Source typecandidate’s take-home repository
A three-hour voice option. A GitHub account named brianbrainforge-wq published an “AI FDE Challenge” offering a choice: a voice agent that answers questions from the live brainforge.ai website, or an evaluation framework for a software engineering agent. Source 6AI FDE ChallengePublisherbrianbrainforge-wq (GitHub account)Source typecompany websiteSource 7Shared challenge briefPublisherbrianbrainforge-wq (GitHub account)Source typecompany website Its link to Brainforge is inferred from the account’s name, not confirmed. Its suggested split gives you up to three hours and sets aside 20 minutes for failure testing and 10 minutes for limitations. Source 6AI FDE ChallengePublisherbrianbrainforge-wq (GitHub account)Source typecompany websiteSource 7Shared challenge briefPublisherbrianbrainforge-wq (GitHub account)Source typecompany website For how to spend a time box like that, see our post on scoping a take-home to its time box.
All of these were posted by candidates, except the Brainforge one:
| Brief | You hand in | Time box |
|---|---|---|
| HappyRobot, via IanDalton Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository | Agent, dashboard, email, five-minute video | Not stated |
| Dental prompt, via hobbes3 Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository | Prompt, Slack message | Two hours max |
| Brainforge, via brianbrainforge-wq Source 6AI FDE ChallengePublisherbrianbrainforge-wq (GitHub account)Source typecompany websiteSource 7Shared challenge briefPublisherbrianbrainforge-wq (GitHub account)Source typecompany website | Voice agent or eval framework | Up to three hours |
| HVAC, via othmane-zizi-pro Source 4ridgeway-hvac-voice-agent (repository description)Publisherothmane-zizi-pro (GitHub)Source typecandidate’s take-home repository | Build only; prompt not posted | Not posted |
| Law firm, via AryaanSheth Source 5Voice-Intake-Pipeline (repository description)PublisherAryaanSheth (GitHub)Source typecandidate’s take-home repository | Build only; prompt not posted | Not posted |
Voice also turns up in live rounds. One candidate reported, in a September 2026 Medium post, an FDE intern interview at Nugget, Zomato’s voice AI team in India, where the engineering manager round covered agentic workflows, voice AI and . Source 8My Interview Experience: FDE Intern Role at Nugget (Zomato’s Voice AI Team) -EternalPublisherMedium (Divyanshu Rajput)Source typecandidate’s personal write-up That was an interview, not a take-home, and the Medium post says one scenario went “deep on STT, VAD, TTS models, routing logic, and how all of it ties together.” Source 8My Interview Experience: FDE Intern Role at Nugget (Zomato’s Voice AI Team) -EternalPublisherMedium (Divyanshu Rajput)Source typecandidate’s personal write-up Learn those terms before any voice brief: speech-to-text (STT), voice activity detection (VAD), text-to-speech (TTS), and routing: how the agent decides which flow, tool or person handles each request. If the voice round is a design interview, see the agentic system design interview post.
The job itself points the same way. As of September 2026, Decagon’s Field Engineer posting covers custom integrations against enterprise telephony and CCaaS (contact-center-as-a-service) platforms. Source 9Field Engineer @ Decagon (New York City)PublisherDecagon (Ashby job board)Source typecompany job posting ElevenLabs says its FDE team helped Deutsche Telekom deploy an agent for live translation and assistance within a phone call. Source 10Alex Holt appointed as Field CTO at ElevenLabsPublisherElevenLabsSource typecompany blog
Why a chat habit fails on a phone call
On a call, a bulleted list, a link and three options become a wall of speech the caller has to remember. Four habits to drop:
- Long turns. Keep each agent turn to one idea and at most one question. The caller cannot skim.
- Formatting. Markdown, lists and URLs get read aloud or skipped. Write the prompt for the ear.
- Trusting the transcript. Speech-to-text mishears IDs and numbers. Read back anything the agent will act on, such as an MC number, a date or an address.
- Numbers as text. A rate the model writes as digits may be spoken oddly. Tell the model how to say money, dates and IDs, and test it.
Here is the difference on the carrier-sales call:
Chat habit:
"Great! Here are 3 loads that match:
1) Dallas to Atlanta, $1,850 ...
2) ... Which would you like?"
Voice:
"I've got a dry van out of Dallas going
to Atlanta, picking up tomorrow morning.
It pays eighteen fifty. Want to hear
more on that one?"
The voice version pitches one load, says the price the way a broker would, and ends on a yes-or-no question. In a prompt-only test, a reviewer reads your prompt imagining it spoken.
Worked example: the missed-appointment objection
The hardest caller in hobbes3’s brief is the patient who says they never made the appointment. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository Here is our example of a prompt section for that moment, written for the ear:
If the patient says they never booked:
- Do not argue or cite records.
- Say: "That's alright, it may have
been booked for you. Would you like
to find a time that suits you now?"
- If they are still upset, apologize
once, offer to remove them from
reminders, and end the call kindly.
- Never repeat the missed date twice.
Arguing over records wins the point and loses the reschedule, which is the whole goal of the call.
Latency: set a budget for every turn
On a call, a slow reply is dead air, and callers fill dead air by talking over your agent’s reply. Write down a budget for one turn and measure against it.
| Stage | Example budget |
|---|---|
| Detect end of speech | 300 ms |
| Final transcript | 150 ms |
| Model first token | 400 ms |
| First audio out | 150 ms |
| Turn total | 1000 ms |
These are our example values, not a standard. The point is that every stage has a number, and your submission logs where each turn spent its time.
- Stream everything. Start speech synthesis on the first sentence of the model’s reply, not the whole reply.
- Keep the hot-path prompt short. It costs time on every turn.
- Speak before slow tools. If a lookup may take a while, the agent says “Let me pull that up” first, so silence has a reason.
- Log p50 and p95 per stage. A good average hides the one turn that left a caller in silence.
In the review, say it plainly: “My budget was one second a turn. Model first token was the slowest stage, so I cut the prompt and moved load details into a tool call.”
Turn-taking, interruptions and silence
The caller interrupts. This is barge-in. When the caller starts speaking over the agent, stop playback as soon as you are sure it is a turn, not a backchannel. Then fix the part that is easy to miss: the conversation history must record only what the caller heard. Otherwise the model believes it said things the caller never heard.
def heard_text(words, played_ms):
# words: [(word, end_ms)], grouped
# from the TTS timestamps
said = [w for w, end in words
if end <= played_ms]
return " ".join(said + ["[cut off]"])
Store the result of heard_text as the agent’s turn, not the text the model generated. Many TTS APIs return timestamps per character, so group them into words first. Measure played_ms at playback, not at send: the TTS runs ahead of the caller’s ear. OpenAI’s Realtime API conversation.item.truncate event exists for exactly this: audio already sent to the client but not yet played.
The caller says “mm-hm”. A short acknowledgment is not an interruption. Require a minimum number of words (Pipecat ships this as MinWordsUserTurnStartStrategy), or ignore a short list of backchannels like “mm-hm”, “yeah” and “okay”, before you stop playback, or your agent will keep cutting itself off.
The caller goes quiet. Set a silence timeout, for example silence_timeout = 6s. On the first timeout, reprompt with a short question. On the second, close politely and log the call as abandoned, not as failed.
A useful mini-scenario for your test set: the agent is reading a load’s details, the carrier cuts in with “what’s the weight?”, and the agent should answer the weight, not restart the pitch.
When a tool call fails mid-call
Every brief above depends on an API. When one fails mid-call, the agent goes silent, or worse, carries on as if it succeeded and tells a carrier they are verified when they are not.
Wrap every tool with a timeout, one retry with , and an explicit failure result the prompt knows how to handle:
import asyncio, random
async def verify_carrier(mc, lookup,
timeout=1.5):
for attempt in range(2):
try:
return await asyncio.wait_for(
lookup(mc), timeout)
except Exception:
if attempt == 0:
await asyncio.sleep(
random.uniform(0.1, 0.4))
return {"status": "unverified"}
We ran it against four fake lookups: one that hangs once and then answers, one that refuses connections, one that always hangs, and one that fails with an HTTP 500. It returns the record, then unverified three times. Worst case is just over three seconds, so the agent says “Let me pull up your MC number” before it calls this. Our question on retrying with exponential backoff and jitter covers why the jitter matters.
Then decide, in the prompt, what the agent says when the result is unverified. For example: “I can’t confirm your authority right now. I can have a broker call you back within the hour. What’s the best number?” Never let the model improvise a verification.
Name these failure modes in your write-up too:
- Double booking. If the HVAC booking call times out, it may still have succeeded. If the booking API accepts an , send one so a retry cannot create a second appointment; if it does not, look up the caller’s bookings before you retry, and say so in the write-up.
- Wrong input. A misheard MC number returns “not found”. The agent should read the number back and ask again before treating the carrier as ineligible.
Our lesson on labeling assumptions and surfacing failure modes shows how to present this list.
Hand-off to a person
Decide the hand-off triggers before you build:
- The caller asks for a person. Honor it the first time.
- The agent has failed to understand the same thing twice.
- The request is outside what the agent may do: a counter-offer beyond the negotiation limit, or legal advice in a law-firm intake.
- Safety. An HVAC caller who mentions a gas smell gets a scripted safety message and a person, not a booking slot.
Make it a warm transfer: pass the person a short summary so the caller does not repeat themselves. For example: “Carrier verified, interested in the Dallas load, countered twice, wants more than we can offer.” If no one is available, take a callback number. Show one hand-off in the demo: it proves you thought about the customer’s staff, not just the model.
The metrics dashboard and the demo
IanDalton’s HappyRobot brief asks for a dashboard you build yourself, to assess “product vision”. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository So build it for the person who runs the desk, not for yourself. Our lesson on stakeholders and success metrics covers how to choose them. A broker’s view might show:
- Calls by outcome: booked, no matching load, carrier not eligible, handed to a person, abandoned.
- Negotiation: agreed rate against listed rate, and how many rounds it took.
- Sentiment by outcome, so a bad trend shows up early.
- Health: turn latency at p95, tool error rate and hand-off rate.
The last group is what an engineer owns. Our question on defining SLOs for an AI assistant turns it into targets you can defend.
Before you build, write the test calls. ElevenLabs’ post on lessons from forward deployed engineering recommends defining success criteria and tests before building a voice agent. Source 11Building voice agents that last: some lessons learned from forward deployed engineeringPublisherElevenLabsSource typecompany blog
Test calls to run before you record
- The happy path, start to finish
- A caller who interrupts mid-pitch
- A caller who says “mm-hm” while the agent talks
- A caller who goes silent
- A misheard ID that needs a read-back
- The main API timing out
- A request for a person
- A request the agent must refuse
Run each one, keep the recordings, and show the results. Our question on designing an eval harness for a support agent shows how to turn these calls into a repeatable suite, and how do you know your AI system works goes deeper on evaluation.
For the video, lead with a real call, then show the dashboard, then one failure handled well. Say what you would do next with another week. Start small: a walking skeleton that takes one call end to end beats a clever negotiation engine that has never answered the phone.
Practice the customer, not just the build
IanDalton’s HappyRobot brief ends with you in front of a customer played by the interviewer. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository The free practice case rehearses exactly that: ten minutes with an AI customer who has a vague building-permits problem, then a score that quotes what you said. Sign in free and run it before your call.
Questions people ask
What voice AI agent take-home briefs have candidates reported?
The briefs candidates have posted vary. One candidate posted a HappyRobot FDE challenge that asks for an inbound voice agent that verifies a carrier, pitches loads, negotiates and classifies the outcome, plus a metrics dashboard and a short video. A second candidate posted a voice prompt-engineering brief that asks only for a system prompt and a message to the customer.Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repositorySource 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository
Do I need my own telephony for a voice agent take-home, or did the brief one candidate posted supply a platform?
Read the brief. In the HappyRobot challenge one candidate posted, you build the agent on HappyRobot’s own platform, and the brief asks you to build your own metrics view rather than use the platform’s analytics.Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository
How do I test a voice agent before I submit it?
Write the test calls before you build, covering the happy path, an interruption, a caller who goes silent, a failed API call and a request for a person. ElevenLabs’ post on its forward deployed work recommends defining success criteria and tests before building a voice agent.Source 11Building voice agents that last: some lessons learned from forward deployed engineeringPublisherElevenLabsSource typecompany blog
Keep reading
Lessons
Questions
- Define service level objectives for an AI assistant a customer depends on. What do you measure?
- Design an evaluation harness that runs before every change to a customer-facing support agent.
- A customer asks how you know your AI system works. Answer them.
- Write a retry wrapper with exponential backoff and full jitter. Which errors should it not retry?
More from the blog
Interview rounds
Forward deployed engineer take-homes: what real prompts ask for and how to scope yours
What published and candidate-reported FDE take-home prompts ask for, the deliverables they share, and how to scope yours to the time box.
Interview rounds
Prompt engineering interview tests: what they look like and how to work under the clock
What timed prompt-engineering tests for FDE and applied AI roles look like, and a test-first way to work against the clock without guessing.
Interview rounds
The take-home walkthrough video: a script that shows your judgment in a few minutes
Several FDE take-homes ask for a recorded walkthrough. A shot-by-shot script, a README template and the recording mistakes that hide your judgment.