Inhoud samenvatten met:
A caller adds a detail while your AI assistant is already answering: “Thursday works. But only after 3 pm.” A good conversation takes that detail into account and moves forward. That is what Full Duplex (Beta) in Famulor is designed for: the assistant can listen and speak at the same time, adapting its ongoing response to new information.
The update adds a new conversation mode to Realtime. You choose where to use it. Existing assistants keep their engines and settings. This guide compares Full Duplex with Pipeline, Half Cascade and standard Realtime, explains the relevant model benchmarks, and shows how to evaluate the options for your own callers.
Product and source information checked on September 13, 2026. Figures below come from published provider evaluations and external benchmarks. Business dialogues are illustrative examples.
Hear a Full Duplex conversation
This Famulor demo presents a workplace conversation with follow-up questions and information lookups, including tasks and product updates. It illustrates the flow between dialogue and retrieval. It is not a benchmark of inbound customer-call success. The conversation is in German.
Famulor demo: follow-up questions and information lookups in conversation. YouTube
Four conversation paths, side by side
Pipeline, Half Cascade and Realtime describe different ways of turning speech into an answer. In Famulor, Full Duplex is a variant of the Realtime engine. Comparing four conversation paths makes the practical choices easier to understand.
Four conversation paths in Famulor | |||
Conversation path | Processing | Voice | Typical priority |
|---|---|---|---|
Pipeline | Audio → STT → text model → TTS | Separate TTS voice | Structured workflows and independent component selection |
Half Cascade | Audio → Realtime model → text → TTS | Separate TTS voice | Audio input with an independent output voice |
Standard Realtime | Audio → native speech model → audio | Native model voice | Direct speech-to-speech conversations |
Realtime: Full Duplex (Beta) | Receive and produce audio, including simultaneously | Compatible native conversation voice | Dialogues with additions and dynamic turn-taking |
Pipeline: A speech recognition service turns audio into text. A language model processes the text, and a separate text-to-speech service voices the answer. You can choose recognition, reasoning and voice separately. Streaming can make this architecture responsive; its name alone does not determine a fixed response time.
Half Cascade: A Realtime model processes incoming audio and produces response text. A separate TTS service speaks that text. This combines direct audio input with an independently selected output voice. It does not add the separate upstream STT stage found in a classic pipeline.
Standard Realtime: The model processes and generates audio natively. Listening, response generation and voice are more closely connected. Pause handling and interruption behavior depend on the specific model and configuration.
Realtime with Full Duplex: Incoming and outgoing speech can be processed together as part of the ongoing dialogue. The assistant can hear an addition while it is speaking. This mode uses a compatible native conversation voice.
Interrupting playback and listening while speaking are different capabilities
Many voice assistants already support barge-in: they stop playback when the caller starts speaking. That is useful. But “yes, exactly” and “no, the other address” require different reactions. One is a brief acknowledgment; the other changes the information the assistant should use.
Full Duplex targets a more flexible exchange. The assistant needs to interpret a contribution in context and respond appropriately. Test this with phrases your callers actually use: brief acknowledgments, longer additions, self-corrected details and speech with background noise.
What matters to the caller is the whole experience. Did the assistant understand? Did they have to repeat themselves? Was the outcome correct? A fast first syllable does not answer all three questions.
Keep the conversation moving while information is retrieved
A useful conversation sometimes needs an action in the background. A calendar must return available slots, a CRM must locate a contact, or a knowledge base must supply the relevant information. The assistant should explain what it is doing and continue with the actual result.
For example, the assistant says, “I’ll check the available appointments.” The caller adds, “Afternoons only, please.” That preference should shape the search. An appointment should only be described as booked after the connected calendar confirms it. Natural dialogue and correct actions need to work together.
When configuring an action, describe its purpose, required inputs and expected result. Specify when the assistant should ask a follow-up question. For calendar tools, distinguish looking up availability from creating a confirmed booking. A friendly acknowledgment should never stand in for a successful system response.
GPT-Live and Realtime: published benchmark results
OpenAI introduced GPT-Live-1 on September 10, 2026. These figures come from the charts in the original GPT-Live-1 announcement. They compare specific model configurations.
OpenAI model benchmarks: higher percentages and lower latency are better | |||
Benchmark | GPT-Live-1 | GPT-Realtime-2.1 | GPT-Realtime-2 |
|---|---|---|---|
Tau3 (Voice), Pass@1 | 86.2 % | 45.7 % | 42.4 % |
Full Duplex Bench v1.5: interactivity | 80.1 % | 45.4 % | 47.8 % |
Full Duplex Bench v1: turn-taking latency | 0.798 s | 1.410 s | 1.630 s |
Full Duplex Bench v3: Tool Calling | 87 % | 60 % | 58 % |

Famulor chart using published OpenAI results. Two separate metrics, not an overall score.
Methodology: Tau3 (Voice) uses equal domain weighting. The GPT-Live configuration uses the Astra backend at medium for that result; the tool-calling evaluation uses Terra at low. These describe the published evaluations, not selectable Famulor configurations. The rows represent different tests and benchmark versions.
Two questions matter independently: how well does a model handle a dynamic conversation, and how reliably does it complete a task? Interactivity and tool calling provide different signals. These results are neither guaranteed Famulor response times nor forecasts for your own call outcomes.
A separate market view of native Realtime models
Coval evaluates speech models using its own conversation setup. Its published comparison provides another perspective on native audio models. Because the procedure differs, its numbers cannot be directly combined with OpenAI’s Full Duplex Bench results.
External Realtime comparison with a separate methodology | |||
Model | Voice-to-voice latency ↓ | Instruction adherence ↑ | Samples |
|---|---|---|---|
GPT Realtime 2 | 1332 ms | 50 % | 448 |
Gemini 3.1 Flash Live (Preview) | 1340 ms | 56 % | 445 |
Source: Coval: Gemini Live and comparison models, rolling 30-day figures; last listed run August 27, 2026, retrieved September 13. The Gemini row specifically covers 3.1 Flash Live Preview. It does not measure the Gemini 2.5 Native Audio variants also listed in Famulor.
A model revision can change a result. Check the complete model name, test region, channel and date. Comparing a browser microphone with a telephone call can also introduce differences caused by audio transport rather than the model.
TTS benchmarks: how quickly does the voice start?
Pipeline and Half Cascade let you select speech output separately. Time to First Audio, or TTFA, helps assess that choice: how long does it take from the TTS request to the first audible output? The following selection comes from Coval’s published measurements.
Selected TTS models: time to first audio and word error rate | ||
Model | TTFA ↓ | WER ↓ |
|---|---|---|
Inworld TTS 2 | 177 ms | 4.6 % |
Eleven Flash v2.5 | 252 ms | 6.7 % |
Soniox TTS v2 | 255 ms | 4.0 % |
Cartesia Sonic 3.5 | 274 ms | 5.9 % |
Deepgram Aura 2 | 314 ms | 5.3 % |
Rime Coda | 316 ms | 5.0 % |
Fish Audio S2.1 Pro | 355 ms | 4.7 % |
Eleven v3 Conversational | 380 ms | 4.3 % |
Cartesia Sonic 3.6 | 451 ms | 5.3 % |
GPT-4o mini TTS | 1026 ms | 4.8 % |

Famulor chart based on Coval TTS measurements. The time shown covers the measured TTS request.
Source: Coval TTS model comparison. Rolling 30-day averages, retrieved September 13, 2026; last listed run September 10. WER means word error rate. Lower is better for both columns. Neither column rates personal voice preference or how natural a complete conversation sounds.
Coval defines TTFA as time to audible output, including leading silence. Its TTFA methodology and TTS dataset description document a fixed prompt set and provider APIs in a US region. Those conditions differ from an end-to-end Famulor call.
Provider figures also need context. ElevenLabs describes roughly 75 milliseconds of model inference for Flash under suitable conditions. Its latency documentation separates that measurement from additional network and application time. It therefore measures something different from the 252 milliseconds in the Coval table.
Choose a voice with a short listening test based on your business: company names, street names, dates, times, email addresses and a longer explanation. Check whether the output stays clear as the content becomes more complex. Listen for number grouping and the pronunciation of abbreviations.
Full Duplex uses a different voice choice. Its native conversation voice is part of the mode. Selecting a fast standalone TTS provider does not turn that provider into a Full Duplex voice. If a particular TTS voice or brand voice is essential throughout the conversation, compare Pipeline and Half Cascade as well.
Reasoning model benchmarks: solving tasks and using tools
A voice assistant needs to connect information, follow rules and take the right actions. Text and tool benchmarks help assess those abilities. They measure different capabilities from speech latency or voice quality.
A comparison from the Gemini 3.5 Flash model card
Text and tool benchmarks; higher is better | ||
Model | MCP Atlas | Toolathlon |
|---|---|---|
Gemini 3.5 Flash | 83.6 % | 56.5 % |
GPT-5.5 | 75.3 % | 55.6 % |
Gemini 3 Flash | 62.0 % | 49.4 % |
Source: Google DeepMind: Gemini 3.5 Flash Model Card. Higher is better. These are provider evaluations for the stated model and test configurations, not measurements of business phone calls.
Large and small models on the same task benchmark
Tau2-Bench Telecom by model and reasoning effort | ||
Model | Reasoning effort | Result ↑ |
|---|---|---|
GPT-5.4 | xhigh | 98.9 % |
GPT-5.4 mini | xhigh | 93.4 % |
GPT-5.4 nano | xhigh | 92.5 % |
GPT-5 mini | high | 74.1 % |
Source: OpenAI: GPT-5.4 mini and nano, Tau2-Bench Telecom. Higher is better. The reasoning effort is part of each result; it is not a recommendation to use that setting for every time-sensitive voice response.
Use these numbers to shortlist candidates. Then evaluate your real tasks: requesting a missing part of an address, distinguishing similar products, or withholding an action when a required input is absent. A model that leads a broad benchmark may not offer the best combination of speed, cost and quality for a short reception dialogue.
Pipeline response time has several components: detecting the end of speech, transcription, response generation, any tool lookup and speech output. Some work can overlap. Adding figures from unrelated leaderboards does not produce a reliable end-to-end benchmark.
Six use cases for more natural calls
These scenarios reflect common Famulor workflows for reception, lead qualification, appointment booking and support. The dialogues are invented examples, not published customer recordings.
1. Reception: clarify the request mid-conversation
While the assistant explains the next steps, the caller adds, “This is about maintenance, not a new installation.” The agent incorporates the clarification and asks for the relevant details. The team receives a correctly categorized request. Full Duplex is worth testing when additions are common; Pipeline remains a useful baseline for a structured intake process.
2. Booking: incorporate a new time constraint
“Tuesday could work.” As the reply begins, the caller adds, “Only after school, though.” The agent asks for the specific time and checks availability again. It confirms the appointment after a successful calendar action. Evaluate whether corrected preferences lead to bookings that actually match the caller’s needs.
3. Sales: adjust qualification when the scope changes
A prospect mentions one location, then adds three more. That changes the relevant solution and the questions to ask next. The agent should combine the new details and send structured information to the CRM. Both conversation flow and data accuracy matter in this workflow.
4. Support: handle a corrected reference number
“The order ends in 42. Sorry, 24.” The agent continues with the corrected number and confirms sensitive matches before taking action. Full Duplex can help with self-corrections, while the connected search still needs to return an unambiguous match or trigger a follow-up question.
5. Field service: capture extra access details
A customer describes a fault and adds, “You can only get in through the courtyard.” The agent includes the detail in the service request. Clear business rules determine what information to capture and when to hand over to the team. The conversation mode supports those rules; it does not define them.
6. Brand voice: choose the right balance
A company uses a particular TTS voice for customer contact. Half Cascade can combine direct audio input with that separate speech output. If frequent additions shape the calls, also test Full Duplex with a native voice. Use identical content and assess clarity, brand recognition and conversational flow separately.
Set up Full Duplex in Famulor
Open the relevant assistant. In an eligible workspace with beta features enabled, select Realtime → Full Duplex (Beta).
Listen to the native voice previews. Choose a compatible voice using vocabulary from your workflow. Full Duplex previews use the actual conversation voice.
Review greetings and announcements. Check the opening, consent message and tool announcements. An uploaded greeting keeps its recorded voice.
Define actions clearly. Specify required inputs, follow-up questions and unsuccessful outcomes. Spoken confirmations should reflect successful tool results.
Test through the intended contact channel. Include interruptions, brief acknowledgments, corrected details, silence and the end of the conversation. Reuse the scenarios you run with your existing assistant.
The update also improves startup with longer assistant instructions, recognition of configured actions and consistent recording of conversation usage. Hanging up after the farewell no longer opens an unnecessary new voice connection.
Natural announcements and required wording
The compatible native voice is used for Full Duplex conversations and the announcements intended for that mode. Ordinary announcements may be rephrased in conversation. Required consent announcements are checked and played completely before consent is accepted. Test this flow with your actual wording and enabled actions.
The editor preserves a fallback voice and explains conflicts with output filters before switching. Review the editor’s guidance before saving the configuration so the new conversation voice fits your existing workflow.
Build a useful comparison for your business
A meaningful comparison uses the same tasks under similar conditions. For example, start with 20 to 30 representative dialogue cases, repeat each one, and alternate the order of the configurations. This is a suggested starting point for your evaluation, not a test series conducted for this article.
Record the conditions: language, phone or web channel, region, background noise, voice, model version, prompt and connected tools.
Measure conversation flow: time to an audible response, reactions to corrections, unnecessary interruptions and repeated questions. Separate median behavior from slower cases.
Check the outcome: was the correct appointment created? Are the CRM fields accurate? Did the assistant explain a failed action honestly?
Listen to the output: assess clarity, pronunciation and brand fit. A percentage cannot replace this listening test.
Document cases where a simpler setup performs better, too. A short intake call may have different priorities from an advisory conversation with frequent additions. Choose the configuration that fits the workflow and the result you need.
Frequently asked questions
Who can select Full Duplex (Beta)?
The option appears in eligible workspaces with beta features enabled, suitable Realtime access and a supported region. At the time of writing, Global and US workspaces are supported; the option is not yet available for EU workspaces. Your editor shows the options available to your workspace.
Will existing assistants switch automatically?
No. Existing assistants retain their engines and settings. You choose which assistants use the new mode.
Can I keep my existing TTS voice?
The ongoing Full Duplex conversation uses a compatible native voice. Uploaded greetings keep their recorded voices. If a particular standalone TTS voice must speak the entire dialogue, compare Pipeline and Half Cascade.
Is Full Duplex always faster?
The published tests show advantages for certain configurations. Actual response time also depends on your channel, region, prompt, tools and conversation. Use benchmarks to shortlist options and your own workflow to decide.
What does the new mode cost?
Check the current terms and usage information in your Famulor workspace. External model prices and benchmark timings do not provide a reliable cost calculation for a complete Famulor call.
Start with a workflow you know well
Choose an assistant with a clear task: capturing a request, finding an appointment or classifying a support issue. Compare its existing conversation path with Full Duplex using the same scenarios. That will show where listening and speaking together helps your callers.
The Famulor Full Duplex introduction video. YouTube
Open Famulor and check Full Duplex (Beta) in your assistant. Find more product updates on the Famulor blog.
Auteur bij Famulor




