In this post12 sections
  1. The short answer
  2. What voice agent take-homes have asked for
  3. Why a chat habit fails on a phone call
  4. Latency: set a budget for every turn
  5. Turn-taking, interruptions and silence
  6. When a tool call fails mid-call
  7. Hand-off to a person
  8. The metrics dashboard and the demo
  9. Practice the customer, not just the build
  10. Questions people ask
  11. Keep reading
  12. More from the blog

The take-home arrived and it says “voice ”. You have built chat agents before, but a phone call gives you nowhere to hide: the caller cannot scroll back, will not wait through silence, and talks over your agent mid-sentence. This post covers what voice agent take-homes have asked for and what your submission has to get right; the FDE interview guide covers the rest of the loop around it.

The short answer

The voice take-homes candidates have posted ask you to build, or prompt, an agent that does one phone task for a customer scenario, and then to show that it works. One candidate reported a HappyRobot brief that asks for an inbound carrier-sales agent, a self-built metrics dashboard, an email to the prospect and a 5-minute video. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository A second candidate posted a voice brief that asks only for a system prompt and a Slack message to the customer. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository

Whatever the brief, a call exposes what a chat demo hides. Show a , sensible turn-taking, a plan for failed tool calls, a hand-off to a person, and metrics the customer would read.

What voice agent take-homes have asked for

These are the briefs we could find in public, all from candidates’ repos or, in one case, an account whose company link is inferred. They are examples, not a census.

Inbound carrier sales on HappyRobot. One candidate, IanDalton on GitHub, posted the brief “ Technical Challenge: Inbound Carrier Sales” in a June 2026 repo: you present a proof of concept built on the HappyRobot platform to a customer played by the interviewer. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository In IanDalton’s brief, the agent verifies a carrier’s MC number with the FMCSA API, pitches matching loads, negotiates up to three counter-offers, extracts the offer data and classifies outcome and sentiment. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository IanDalton’s brief adds: “Don’t use the platform analytics and rather build something yourself as we want to assess your product vision and build capabilities.” Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository

A May 2026 HappyRobot take-home repo by timbruckdorfer describes a similar build, “wrapped in a custom metrics dashboard that a freight broker would actually want to use.” Source 3HappyRobot FDE Take-Home — Inbound Carrier SalesPublishertimbruckdorfer (GitHub)Source typecandidate’s take-home repository

A timed voice prompt test. The GitHub user hobbes3 posted their submitted assessment in a January 2026 repo named Bland-FDE-assessment, headed “Take-Home Assignment: Voice AI Prompt Engineering”. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository hobbes3’s brief sets a 2-hour maximum and asks for the system prompt for a voice agent that calls dental patients who missed appointments. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository In hobbes3’s brief, the fictional scenario says 40% of patients claim they never made the original appointment. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository hobbes3’s brief also warns: “Don’t use ChatGPT to write this - we can tell.” Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository A brief like this is a prompt test more than a build, and our post on prompt engineering interview tests covers how to work one under the clock.

An HVAC booking agent. A September 2026 repo by othmane-zizi-pro, described as a “Clerk FDE take-home”, builds a voice agent that books AC repair for Ridgeway HVAC against a mocked two-endpoint API. Source 4ridgeway-hvac-voice-agent (repository description)Publisherothmane-zizi-pro (GitHub)Source typecandidate’s take-home repository The repo does not say which company named Clerk set it. othmane-zizi-pro explains the mock: “The API is mocked here because production sits on the customer’s VPN.” Source 4ridgeway-hvac-voice-agent (repository description)Publisherothmane-zizi-pro (GitHub)Source typecandidate’s take-home repository That one line is the FDE job in miniature.

A law-firm intake agent. A July 2026 repo by AryaanSheth, described as “General Magic FDE Take Home Law Firm VIP”, builds a voice-driven client-intake agent for a law firm; the README does not reproduce the prompt. Source 5Voice-Intake-Pipeline (repository description)PublisherAryaanSheth (GitHub)Source typecandidate’s take-home repository

A three-hour voice option. A GitHub account named brianbrainforge-wq published an “AI FDE Challenge” offering a choice: a voice agent that answers questions from the live brainforge.ai website, or an evaluation framework for a software engineering agent. Source 6AI FDE ChallengePublisherbrianbrainforge-wq (GitHub account)Source typecompany websiteSource 7Shared challenge briefPublisherbrianbrainforge-wq (GitHub account)Source typecompany website Its link to Brainforge is inferred from the account’s name, not confirmed. Its suggested split gives you up to three hours and sets aside 20 minutes for failure testing and 10 minutes for limitations. Source 6AI FDE ChallengePublisherbrianbrainforge-wq (GitHub account)Source typecompany websiteSource 7Shared challenge briefPublisherbrianbrainforge-wq (GitHub account)Source typecompany website For how to spend a time box like that, see our post on scoping a take-home to its time box.

All of these were posted by candidates, except the Brainforge one:

BriefYou hand inTime box
HappyRobot, via IanDalton Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repositoryAgent, dashboard, email, five-minute videoNot stated
Dental prompt, via hobbes3 Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repositoryPrompt, Slack messageTwo hours max
Brainforge, via brianbrainforge-wq Source 6AI FDE ChallengePublisherbrianbrainforge-wq (GitHub account)Source typecompany websiteSource 7Shared challenge briefPublisherbrianbrainforge-wq (GitHub account)Source typecompany websiteVoice agent or eval frameworkUp to three hours
HVAC, via othmane-zizi-pro Source 4ridgeway-hvac-voice-agent (repository description)Publisherothmane-zizi-pro (GitHub)Source typecandidate’s take-home repositoryBuild only; prompt not postedNot posted
Law firm, via AryaanSheth Source 5Voice-Intake-Pipeline (repository description)PublisherAryaanSheth (GitHub)Source typecandidate’s take-home repositoryBuild only; prompt not postedNot posted

Voice also turns up in live rounds. One candidate reported, in a September 2026 Medium post, an FDE intern interview at Nugget, Zomato’s voice AI team in India, where the engineering manager round covered agentic workflows, voice AI and . Source 8My Interview Experience: FDE Intern Role at Nugget (Zomato’s Voice AI Team) -EternalPublisherMedium (Divyanshu Rajput)Source typecandidate’s personal write-up That was an interview, not a take-home, and the Medium post says one scenario went “deep on STT, VAD, TTS models, routing logic, and how all of it ties together.” Source 8My Interview Experience: FDE Intern Role at Nugget (Zomato’s Voice AI Team) -EternalPublisherMedium (Divyanshu Rajput)Source typecandidate’s personal write-up Learn those terms before any voice brief: speech-to-text (STT), voice activity detection (VAD), text-to-speech (TTS), and routing: how the agent decides which flow, tool or person handles each request. If the voice round is a design interview, see the agentic system design interview post.

The job itself points the same way. As of September 2026, Decagon’s Field Engineer posting covers custom integrations against enterprise telephony and CCaaS (contact-center-as-a-service) platforms. Source 9Field Engineer @ Decagon (New York City)PublisherDecagon (Ashby job board)Source typecompany job posting ElevenLabs says its FDE team helped Deutsche Telekom deploy an agent for live translation and assistance within a phone call. Source 10Alex Holt appointed as Field CTO at ElevenLabsPublisherElevenLabsSource typecompany blog

Why a chat habit fails on a phone call

On a call, a bulleted list, a link and three options become a wall of speech the caller has to remember. Four habits to drop:

  • Long turns. Keep each agent turn to one idea and at most one question. The caller cannot skim.
  • Formatting. Markdown, lists and URLs get read aloud or skipped. Write the prompt for the ear.
  • Trusting the transcript. Speech-to-text mishears IDs and numbers. Read back anything the agent will act on, such as an MC number, a date or an address.
  • Numbers as text. A rate the model writes as digits may be spoken oddly. Tell the model how to say money, dates and IDs, and test it.

Here is the difference on the carrier-sales call:

Chat habit:
  "Great! Here are 3 loads that match:
   1) Dallas to Atlanta, $1,850 ...
   2) ... Which would you like?"

Voice:
  "I've got a dry van out of Dallas going
   to Atlanta, picking up tomorrow morning.
   It pays eighteen fifty. Want to hear
   more on that one?"

The voice version pitches one load, says the price the way a broker would, and ends on a yes-or-no question. In a prompt-only test, a reviewer reads your prompt imagining it spoken.

Worked example: the missed-appointment objection

The hardest caller in hobbes3’s brief is the patient who says they never made the appointment. Source 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository Here is our example of a prompt section for that moment, written for the ear:

If the patient says they never booked:
- Do not argue or cite records.
- Say: "That's alright, it may have
  been booked for you. Would you like
  to find a time that suits you now?"
- If they are still upset, apologize
  once, offer to remove them from
  reminders, and end the call kindly.
- Never repeat the missed date twice.

Arguing over records wins the point and loses the reschedule, which is the whole goal of the call.

Latency: set a budget for every turn

On a call, a slow reply is dead air, and callers fill dead air by talking over your agent’s reply. Write down a budget for one turn and measure against it.

StageExample budget
Detect end of speech300 ms
Final transcript150 ms
Model first token400 ms
First audio out150 ms
Turn total1000 ms

These are our example values, not a standard. The point is that every stage has a number, and your submission logs where each turn spent its time.

  • Stream everything. Start speech synthesis on the first sentence of the model’s reply, not the whole reply.
  • Keep the hot-path prompt short. It costs time on every turn.
  • Speak before slow tools. If a lookup may take a while, the agent says “Let me pull that up” first, so silence has a reason.
  • Log p50 and p95 per stage. A good average hides the one turn that left a caller in silence.

In the review, say it plainly: “My budget was one second a turn. Model first token was the slowest stage, so I cut the prompt and moved load details into a tool call.”

Turn-taking, interruptions and silence

The caller interrupts. This is barge-in. When the caller starts speaking over the agent, stop playback as soon as you are sure it is a turn, not a backchannel. Then fix the part that is easy to miss: the conversation history must record only what the caller heard. Otherwise the model believes it said things the caller never heard.

def heard_text(words, played_ms):
    # words: [(word, end_ms)], grouped
    # from the TTS timestamps
    said = [w for w, end in words
            if end <= played_ms]
    return " ".join(said + ["[cut off]"])

Store the result of heard_text as the agent’s turn, not the text the model generated. Many TTS APIs return timestamps per character, so group them into words first. Measure played_ms at playback, not at send: the TTS runs ahead of the caller’s ear. OpenAI’s Realtime API conversation.item.truncate event exists for exactly this: audio already sent to the client but not yet played.

The caller says “mm-hm”. A short acknowledgment is not an interruption. Require a minimum number of words (Pipecat ships this as MinWordsUserTurnStartStrategy), or ignore a short list of backchannels like “mm-hm”, “yeah” and “okay”, before you stop playback, or your agent will keep cutting itself off.

The caller goes quiet. Set a silence timeout, for example silence_timeout = 6s. On the first timeout, reprompt with a short question. On the second, close politely and log the call as abandoned, not as failed.

A useful mini-scenario for your test set: the agent is reading a load’s details, the carrier cuts in with “what’s the weight?”, and the agent should answer the weight, not restart the pitch.

When a tool call fails mid-call

Every brief above depends on an API. When one fails mid-call, the agent goes silent, or worse, carries on as if it succeeded and tells a carrier they are verified when they are not.

Wrap every tool with a timeout, one retry with , and an explicit failure result the prompt knows how to handle:

import asyncio, random

async def verify_carrier(mc, lookup,
                         timeout=1.5):
    for attempt in range(2):
        try:
            return await asyncio.wait_for(
                lookup(mc), timeout)
        except Exception:
            if attempt == 0:
                await asyncio.sleep(
                    random.uniform(0.1, 0.4))
    return {"status": "unverified"}

We ran it against four fake lookups: one that hangs once and then answers, one that refuses connections, one that always hangs, and one that fails with an HTTP 500. It returns the record, then unverified three times. Worst case is just over three seconds, so the agent says “Let me pull up your MC number” before it calls this. Our question on retrying with exponential backoff and jitter covers why the jitter matters.

Then decide, in the prompt, what the agent says when the result is unverified. For example: “I can’t confirm your authority right now. I can have a broker call you back within the hour. What’s the best number?” Never let the model improvise a verification.

Name these failure modes in your write-up too:

  • Double booking. If the HVAC booking call times out, it may still have succeeded. If the booking API accepts an , send one so a retry cannot create a second appointment; if it does not, look up the caller’s bookings before you retry, and say so in the write-up.
  • Wrong input. A misheard MC number returns “not found”. The agent should read the number back and ask again before treating the carrier as ineligible.

Our lesson on labeling assumptions and surfacing failure modes shows how to present this list.

Hand-off to a person

Decide the hand-off triggers before you build:

  • The caller asks for a person. Honor it the first time.
  • The agent has failed to understand the same thing twice.
  • The request is outside what the agent may do: a counter-offer beyond the negotiation limit, or legal advice in a law-firm intake.
  • Safety. An HVAC caller who mentions a gas smell gets a scripted safety message and a person, not a booking slot.

Make it a warm transfer: pass the person a short summary so the caller does not repeat themselves. For example: “Carrier verified, interested in the Dallas load, countered twice, wants more than we can offer.” If no one is available, take a callback number. Show one hand-off in the demo: it proves you thought about the customer’s staff, not just the model.

The metrics dashboard and the demo

IanDalton’s HappyRobot brief asks for a dashboard you build yourself, to assess “product vision”. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository So build it for the person who runs the desk, not for yourself. Our lesson on stakeholders and success metrics covers how to choose them. A broker’s view might show:

  • Calls by outcome: booked, no matching load, carrier not eligible, handed to a person, abandoned.
  • Negotiation: agreed rate against listed rate, and how many rounds it took.
  • Sentiment by outcome, so a bad trend shows up early.
  • Health: turn latency at p95, tool error rate and hand-off rate.

The last group is what an engineer owns. Our question on defining SLOs for an AI assistant turns it into targets you can defend.

Before you build, write the test calls. ElevenLabs’ post on lessons from forward deployed engineering recommends defining success criteria and tests before building a voice agent. Source 11Building voice agents that last: some lessons learned from forward deployed engineeringPublisherElevenLabsSource typecompany blog

Test calls to run before you record

  • The happy path, start to finish
  • A caller who interrupts mid-pitch
  • A caller who says “mm-hm” while the agent talks
  • A caller who goes silent
  • A misheard ID that needs a read-back
  • The main API timing out
  • A request for a person
  • A request the agent must refuse

Run each one, keep the recordings, and show the results. Our question on designing an eval harness for a support agent shows how to turn these calls into a repeatable suite, and how do you know your AI system works goes deeper on evaluation.

For the video, lead with a real call, then show the dashboard, then one failure handled well. Say what you would do next with another week. Start small: a walking skeleton that takes one call end to end beats a clever negotiation engine that has never answered the phone.

Practice the customer, not just the build

IanDalton’s HappyRobot brief ends with you in front of a customer played by the interviewer. Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository The free practice case rehearses exactly that: ten minutes with an AI customer who has a vague building-permits problem, then a score that quotes what you said. Sign in free and run it before your call.

GlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on AgentGlossaryLatency budgetThe total time a request may take, divided among the steps it passes through, measured at the slow tail as well as the median.More on Latency budgetGlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineerGlossaryRetrieval-augmented generationAnswering with a model that is given passages retrieved from a document collection as context.More on Retrieval-augmented generationGlossaryExponential backoff with jitterRetrying with growing, randomized delays so that clients recovering from the same failure do not retry in lockstep.More on Exponential backoff with jitterGlossaryIdempotency keyA client-supplied identifier that lets a server apply a repeated request once, making retries safe.More on Idempotency key

Questions people ask

What voice AI agent take-home briefs have candidates reported?

The briefs candidates have posted vary. One candidate posted a HappyRobot FDE challenge that asks for an inbound voice agent that verifies a carrier, pitches loads, negotiates and classifies the outcome, plus a metrics dashboard and a short video. A second candidate posted a voice prompt-engineering brief that asks only for a system prompt and a message to the customer.Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repositorySource 2Take-Home Assignment: Voice AI Prompt Engineering (PDF)Publisherhobbes3 (GitHub)Source typecandidate’s take-home repository

Do I need my own telephony for a voice agent take-home, or did the brief one candidate posted supply a platform?

Read the brief. In the HappyRobot challenge one candidate posted, you build the agent on HappyRobot’s own platform, and the brief asks you to build your own metrics view rather than use the platform’s analytics.Source 1FDE Technical Challenge: Inbound Carrier Sales (PDF)PublisherIanDalton (GitHub)Source typecandidate’s take-home repository

How do I test a voice agent before I submit it?

Write the test calls before you build, covering the happy path, an interruption, a caller who goes silent, a failed API call and a request for a person. ElevenLabs’ post on its forward deployed work recommends defining success criteria and tests before building a voice agent.Source 11Building voice agents that last: some lessons learned from forward deployed engineeringPublisherElevenLabsSource typecompany blog

Keep reading