Table of Contents
In SQM Group's benchmark, 19% of customers who call a contact center are transferred to another agent. Those transferred callers rate their satisfaction 12% lower than callers who never change hands.
A cold transfer makes the caller repeat a name, an account number, and the whole story to someone starting from zero. A warm transfer keeps the story attached to the call. When the first agent is software, the story already exists as structured data, so losing it is an architecture choice.
This guide covers four escalation triggers, six context-package fields for a warm transfer payload, and the difference between SIP REFER and a conference bridge.
Key takeaways
A warm transfer preserves context only when you pass structured data and keep the required call leg alive.
- Transferred calls score lower on CSAT; what travels with the call decides how much of that drop you absorb.
- Your voice agent signals escalation through a function call, while the telephony platform executes the warm transfer and manages dialing and the bridge through hangup.
- Hand off verified identity, intent, entities, actions taken, an escalation code, and a two-sentence summary as structured fields.
- A conference bridge keeps the caller's leg and media stream alive while the stream-bearing leg remains attached. SIP REFER requests the transfer, and a later BYE ends the agent's leg.
Cold transfers vs. warm transfers: What the data shows
The second agent either inherits the caller's story or restarts it, and that difference is what separates a warm transfer from a cold one.
Cold transfer mechanics
When the AI agent's call leg drops before a human answers, everything the caller said travels nowhere unless your application persists it and surfaces it to the human. A voice agent that ends its session and lets the carrier route the call without surfacing stored context has done exactly that.
Warm transfer mechanics
In a warm transfer, you bridge the human into the live call, brief them while the caller listens or waits on hold, and then drop off. With an AI as the first agent, the briefing gets better, because the agent can read its own history for verified identity and intent, along with any completed actions. Your receiving agent can open with "I see you're calling about the April invoice" rather than "How can I help?"
Why the second agent starts slower
If your second agent starts from zero, they have to rediscover the problem before solving it, and that rediscovery delays the work. A warm transfer with a structured payload deletes the rediscovery step, so the human starts working on the problem.
When a voice agent should escalate to a human
A voice agent should escalate to a human whenever it hits deterministic handoff triggers in your orchestration code, saving the model's judgment for genuinely ambiguous cases. Deepgram's prompting guidance makes the same point that every agent needs a way to hand off to a human.
Explicit requests and repair loops
When a caller asks for a person, transfer them without a confidence check in between. Repair loops need a hard count, and Deepgram's example transfers after two unsuccessful attempts, per the same prompting guide mentioned above. Keep that counter in your server so the threshold holds even when the model improvises.
Requests outside the agent's scope
A refund above the agent's approval limit, a contract change, or a formal complaint shouldn't wait for the repair-loop threshold. Map each intent to a target department at classification time and let the department argument carry it.
Sensitive or high-risk situations
You should route emergency symptoms and fraud reports to a human on first mention, whatever the model's confidence. The same Deepgram prompting doc makes the same call for medical emergencies, triggering an immediate transfer the moment a caller describes chest pain, difficulty breathing, or another emergency symptom.
Business-defined escalation rules
If the LLM is the only thing standing between a caller and a transfer, your escalation policy changes every time the model version does. Encode business rules as deterministic checks in the orchestration layer, and treat the model's transfer signal as one input among several.
What belongs in the context package
The payload that keeps a warm transfer warm carries verified identity, intent, entities, actions taken, an escalation code, and a concise summary. If the receiving agent has to re-ask for identity, intent, or the disputed invoice, they've restarted the call.
Caller identity and intent
Verification status should travel with the call, so the human never re-asks for a date of birth the agent already checked. Encode identity and intent as structured text, or persist them before the session closes. The Amazon Connect integration guide makes the same point that context from the AI conversation can be stored in a CRM or ticketing system before the transfer.
Facts, entities, and actions already taken
If the agent already looked up an order or issued a credit, the human needs that result. Preserve unique IDs, complete arguments and responses, chronological order, and only the actions relevant to this conversation. Pull order numbers and account IDs out of structured arguments rather than re-extracting them from transcript text.
Deepgram's Voice Agent API keeps conversation history, injected context, and prior function results in the agent's working state. Your handler can read them directly instead of reconstructing them from audio.
Escalation reason and conversation summary
You want an enumerated code for why the handoff happened, plus a concise summary a human can absorb while the bridge connects. Codes such as explicit_request and repair_loop_exceeded support routing and reporting. Use policy_fraud for fraud-policy escalations, since free-text reasons such as "customer seemed upset" don't. Generate the summary with a secondary LLM call at transfer time, and keep it out of the system prompt.
Transfer architecture: Who does what
Different systems handle the decision and execution of a warm transfer. Deepgram's Voice Agent API supports the voice agent, while your server and telephony platform make the handoff happen. Keeping those roles separate lets you swap carriers or contact center platforms without touching the agent's prompt. The Amazon Connect integration is the reference pattern for this split, where the voice agent decides and packages, and Amazon Connect performs the actual call routing.
The voice agent's role
Use the model to signal when a transfer should occur by exposing a transfer_to_human function the agent can call. Your application receives the function-call event, executes the transfer against your telephony platform, and returns a function result so the agent can narrate what happened. Check the current feature overview before binding your orchestration logic to an agent event.
The telephony layer's role
Your server calls the routing API, and the contact center platform behind it owns the complete handoff. The Amazon Connect guide walks through the chain. The agent detects escalation, the Bot Media Gateway initiates the transfer, the contact enters an Amazon Connect queue, and a human answers. The same split holds on Twilio's conference platform, where the server adds participants to a Twilio conference through the Participants API.
Direct transfer vs. bridged or conference handoff
A direct transfer hands off the call and steps away. Your system tells the caller's phone to connect to the human agent, and once that connection succeeds, the AI agent's leg of the call closes.
Direct transfers stay simple and cheap this way, since you stop paying for the AI agent's connection the moment the human picks up. But it also means anything tied to that leg, like a live transcript, ends with it.
A bridged or conference handoff keeps everyone on the line together for a moment. The human joins the call alongside the caller and the AI agent, and the three of you can hear each other briefly. Then the agent drops off once the human is up to speed.
This costs more, since you're paying for an extra leg during the overlap, and takes more setup. But it keeps the media path open for things like live coaching, transcription, or quality review during the handoff.
Managing the handoff in real time
What the caller hears while the handoff happens decides whether it feels smooth or like a dropped line. Narrate the transfer as it happens, and keep hold audio and the agent's final turns working together so the caller is never met with silence.
Narrating the transfer to the caller
Tell the caller who's coming and why before the dial begins. On the bridge side, Twilio's conference coach attribute lets one participant hear and speak to a single other participant without anyone else hearing. Use it to brief the human privately before releasing the caller from hold. The human then repeats one payload detail back, so the caller hears that the context arrived.
Avoiding dead air
While the bridge connects, play hold audio so the caller knows the call is still active. On Twilio you control this per participant through the Participants API's Hold and HoldUrl parameters. A held participant hears your hold audio and none of the other parties. Prefer spoken updates over music alone, and keep the agent on the bridge until the human confirms they're ready.
When the transfer times out or nobody answers
Amazon Connect waits a fixed 20 seconds for an agent to pick up, then marks the attempt missed and tries the next available agent. Twilio's ring time defaults to 30 seconds but you can set it anywhere from 5 to 600. It also tells your server how the call ended, so you can tell a real answer from a voicemail pickup.
Either way, if nobody picks up in time, have the agent offer the caller a callback and confirm the number out loud. Save that commitment along with the context package to your CRM.
Building a transfer_to_human handoff
Keep execution in a server-side handler beside your telephony client. Trigger it after you've fixed your escalation rules, payload fields, and bridge mechanism.
The function schema
Register transfer_to_human on your agent so the model can request an escalation the same way it would call any other tool. The Voice Agent API forwards the arguments to your server, where your handler validates them, calls the telephony transfer API, and returns a result the agent can speak.
{
"name": "transfer_to_human",
"description": "Escalate the caller to a human agent. Call this when the caller explicitly requests a person, after two failed repair attempts, on out-of-scope requests, or on emergency or fraud signals.",
"parameters": {
"type": "object",
"properties": {
"department": {
"type": "string",
"enum": ["billing", "tech_support", "fraud", "sales", "general"],
"description": "Target queue for the transfer."
},
"escalation_reason": {
"type": "string",
"enum": ["explicit_request", "repair_loop_exceeded", "out_of_scope", "policy_fraud", "policy_billing_dispute", "safety"],
"description": "Enumerated reason code used for routing and reporting."
},
"summary": {
"type": "string",
"description": "Two-sentence summary the human reads before speaking."
}
},
"required": ["department", "escalation_reason", "summary"]
}
}
The handoff payload
Your handler enriches the function-call arguments with verified identity and prior function results already in the agent's working context. It then posts the full object to the contact center platform as contact attributes or a ticket before calling the transfer API. Attach the full turn-by-turn transcript to the CRM record separately; this briefing object is what the human reads before speaking.
{
"caller": { "verified": true, "method": "dob_plus_zip", "customer_id": "C-88213" },
"intent": "billing_dispute",
"entities": { "invoice_id": "INV-2026-04-117", "amount": "84.20" },
"actions_taken": [
{ "name": "lookup_invoice", "response": "Invoice found, status: charged twice" }
],
"escalation_reason": "policy_billing_dispute",
"summary": "Caller was charged twice for April. Agent confirmed the duplicate but couldn't refund above the policy limit."
}
Run it against real escalations
Before production, replay recordings of your own escalations through the agent. Confirm each request arrived after the caller's final word, and each payload named the entity the human needs first. When both hold, wire the handler to your platform's transfer API and switch one queue to the warm transfer path.
Want to test the flow before switching a queue? Create a Deepgram account, grab your $200 free credits, and run your real escalation recordings through it.
FAQ
What if the model requests a department that doesn't exist?
Validate the department argument against your live queue list in the handler, and substitute a default queue on a miss. Tell the caller where they're actually headed.
How should I handle consent and redact sensitive context?
Before posting anything, apply the applicable consent and redaction rules to the context package. Pass only the identity fields, entities, actions, and summary the receiving agent needs.
Keep the full transcript separate from the briefing object, which lets you limit what appears in the agent's immediate screen while preserving the conversation record under your own storage rules.
How do I prevent duplicate transfers?
Give each transfer request a stable identifier tied to the call and escalation attempt. Store the result before retrying the routing API, then return that result when the same request arrives again.
Track states such as requested, dialing, bridged, completed, and failed. A retry can then resume the current attempt instead of adding another human participant or creating a second callback commitment.
How do I preserve context when the warm transfer becomes a callback?
Key the callback record to the call's stable transfer identifier, and associate it with the stored briefing object. Use that identifier to surface the record before the human starts dialing.
The opening should reference the caller's issue and the promised next step. Continue from the stored identity, intent, and completed actions rather than collecting them from zero.
How do I measure whether context actually reached the human?
Track how often the receiving agent's first question asks for something already in the payload; QA sampling or your call analytics can flag it. A rising rate means the screen pop or whisper isn't landing, whatever your transfer API returned.










