Volver al BlogProduct Update

Full Duplex in Famulor: GPT-Live and Voice AI Compared

Listen and speak at once with Full Duplex in Famulor. Compare Pipeline, Half Cascade and Realtime, with model benchmarks, TTS data and use cases.

Sarah MüllerSeptember 13, 202614 min de lectura
Full Duplex in Famulor: GPT-Live and Voice AI Compared

Resumir contenido con:

A caller adds a detail while your AI assistant is already answering: “Thursday works. But only after 3 pm.” A good conversation takes that detail into account and moves forward. That is what Full Duplex (Beta) in Famulor is designed for: the assistant can listen and speak at the same time, adapting its ongoing response to new information.

The update adds a new conversation mode to Realtime. You choose where to use it. Existing assistants keep their engines and settings. This guide compares Full Duplex with Pipeline, Half Cascade and standard Realtime, explains the relevant model benchmarks, and shows how to evaluate the options for your own callers.

Product and source information checked on September 13, 2026. Figures below come from published provider evaluations and external benchmarks. Business dialogues are illustrative examples.

Hear a Full Duplex conversation

This Famulor demo presents a workplace conversation with follow-up questions and information lookups, including tasks and product updates. It illustrates the flow between dialogue and retrieval. It is not a benchmark of inbound customer-call success. The conversation is in German.

Famulor demo: follow-up questions and information lookups in conversation. YouTube

Four conversation paths, side by side

Pipeline, Half Cascade and Realtime describe different ways of turning speech into an answer. In Famulor, Full Duplex is a variant of the Realtime engine. Comparing four conversation paths makes the practical choices easier to understand.

Four conversation paths in Famulor

Conversation path

Processing

Voice

Typical priority

Pipeline

Audio → STT → text model → TTS

Separate TTS voice

Structured workflows and independent component selection

Half Cascade

Audio → Realtime model → text → TTS

Separate TTS voice

Audio input with an independent output voice

Standard Realtime

Audio → native speech model → audio

Native model voice

Direct speech-to-speech conversations

Realtime: Full Duplex (Beta)

Receive and produce audio, including simultaneously

Compatible native conversation voice

Dialogues with additions and dynamic turn-taking

Pipeline: A speech recognition service turns audio into text. A language model processes the text, and a separate text-to-speech service voices the answer. You can choose recognition, reasoning and voice separately. Streaming can make this architecture responsive; its name alone does not determine a fixed response time.

Half Cascade: A Realtime model processes incoming audio and produces response text. A separate TTS service speaks that text. This combines direct audio input with an independently selected output voice. It does not add the separate upstream STT stage found in a classic pipeline.

Standard Realtime: The model processes and generates audio natively. Listening, response generation and voice are more closely connected. Pause handling and interruption behavior depend on the specific model and configuration.

Realtime with Full Duplex: Incoming and outgoing speech can be processed together as part of the ongoing dialogue. The assistant can hear an addition while it is speaking. This mode uses a compatible native conversation voice.

Interrupting playback and listening while speaking are different capabilities

Many voice assistants already support barge-in: they stop playback when the caller starts speaking. That is useful. But “yes, exactly” and “no, the other address” require different reactions. One is a brief acknowledgment; the other changes the information the assistant should use.

Full Duplex targets a more flexible exchange. The assistant needs to interpret a contribution in context and respond appropriately. Test this with phrases your callers actually use: brief acknowledgments, longer additions, self-corrected details and speech with background noise.

What matters to the caller is the whole experience. Did the assistant understand? Did they have to repeat themselves? Was the outcome correct? A fast first syllable does not answer all three questions.

Keep the conversation moving while information is retrieved

A useful conversation sometimes needs an action in the background. A calendar must return available slots, a CRM must locate a contact, or a knowledge base must supply the relevant information. The assistant should explain what it is doing and continue with the actual result.

For example, the assistant says, “I’ll check the available appointments.” The caller adds, “Afternoons only, please.” That preference should shape the search. An appointment should only be described as booked after the connected calendar confirms it. Natural dialogue and correct actions need to work together.

When configuring an action, describe its purpose, required inputs and expected result. Specify when the assistant should ask a follow-up question. For calendar tools, distinguish looking up availability from creating a confirmed booking. A friendly acknowledgment should never stand in for a successful system response.

GPT-Live and Realtime: published benchmark results

OpenAI introduced GPT-Live-1 on September 10, 2026. These figures come from the charts in the original GPT-Live-1 announcement. They compare specific model configurations.

OpenAI model benchmarks: higher percentages and lower latency are better

Benchmark

GPT-Live-1

GPT-Realtime-2.1

GPT-Realtime-2

Tau3 (Voice), Pass@1

86.2 %

45.7 %

42.4 %

Full Duplex Bench v1.5: interactivity

80.1 %

45.4 %

47.8 %

Full Duplex Bench v1: turn-taking latency

0.798 s

1.410 s

1.630 s

Full Duplex Bench v3: Tool Calling

87 %

60 %

58 %

Bar charts comparing GPT-Live-1, GPT-Realtime-2.1 and GPT-Realtime-2 for interactivity and turn-taking latency. All values are in the preceding table.

Famulor chart using published OpenAI results. Two separate metrics, not an overall score.

Methodology: Tau3 (Voice) uses equal domain weighting. The GPT-Live configuration uses the Astra backend at medium for that result; the tool-calling evaluation uses Terra at low. These describe the published evaluations, not selectable Famulor configurations. The rows represent different tests and benchmark versions.

Two questions matter independently: how well does a model handle a dynamic conversation, and how reliably does it complete a task? Interactivity and tool calling provide different signals. These results are neither guaranteed Famulor response times nor forecasts for your own call outcomes.

A separate market view of native Realtime models

Coval evaluates speech models using its own conversation setup. Its published comparison provides another perspective on native audio models. Because the procedure differs, its numbers cannot be directly combined with OpenAI’s Full Duplex Bench results.

External Realtime comparison with a separate methodology

Model

Voice-to-voice latency ↓

Instruction adherence ↑

Samples

GPT Realtime 2

1332 ms

50 %

448

Gemini 3.1 Flash Live (Preview)

1340 ms

56 %

445

Source: Coval: Gemini Live and comparison models, rolling 30-day figures; last listed run August 27, 2026, retrieved September 13. The Gemini row specifically covers 3.1 Flash Live Preview. It does not measure the Gemini 2.5 Native Audio variants also listed in Famulor.

A model revision can change a result. Check the complete model name, test region, channel and date. Comparing a browser microphone with a telephone call can also introduce differences caused by audio transport rather than the model.

TTS benchmarks: how quickly does the voice start?

Pipeline and Half Cascade let you select speech output separately. Time to First Audio, or TTFA, helps assess that choice: how long does it take from the TTS request to the first audible output? The following selection comes from Coval’s published measurements.

Selected TTS models: time to first audio and word error rate

Model

TTFA ↓

WER ↓

Inworld TTS 2

177 ms

4.6 %

Eleven Flash v2.5

252 ms

6.7 %

Soniox TTS v2

255 ms

4.0 %

Cartesia Sonic 3.5

274 ms

5.9 %

Deepgram Aura 2

314 ms

5.3 %

Rime Coda

316 ms

5.0 %

Fish Audio S2.1 Pro

355 ms

4.7 %

Eleven v3 Conversational

380 ms

4.3 %

Cartesia Sonic 3.6

451 ms

5.3 %

GPT-4o mini TTS

1026 ms

4.8 %

TTS bar chart showing 10 models and time to first audible audio from 177 to 1026 milliseconds. The table contains every value.

Famulor chart based on Coval TTS measurements. The time shown covers the measured TTS request.

Source: Coval TTS model comparison. Rolling 30-day averages, retrieved September 13, 2026; last listed run September 10. WER means word error rate. Lower is better for both columns. Neither column rates personal voice preference or how natural a complete conversation sounds.

Coval defines TTFA as time to audible output, including leading silence. Its TTFA methodology and TTS dataset description document a fixed prompt set and provider APIs in a US region. Those conditions differ from an end-to-end Famulor call.

Provider figures also need context. ElevenLabs describes roughly 75 milliseconds of model inference for Flash under suitable conditions. Its latency documentation separates that measurement from additional network and application time. It therefore measures something different from the 252 milliseconds in the Coval table.

Choose a voice with a short listening test based on your business: company names, street names, dates, times, email addresses and a longer explanation. Check whether the output stays clear as the content becomes more complex. Listen for number grouping and the pronunciation of abbreviations.

Full Duplex uses a different voice choice. Its native conversation voice is part of the mode. Selecting a fast standalone TTS provider does not turn that provider into a Full Duplex voice. If a particular TTS voice or brand voice is essential throughout the conversation, compare Pipeline and Half Cascade as well.

Reasoning model benchmarks: solving tasks and using tools

A voice assistant needs to connect information, follow rules and take the right actions. Text and tool benchmarks help assess those abilities. They measure different capabilities from speech latency or voice quality.

A comparison from the Gemini 3.5 Flash model card

Text and tool benchmarks; higher is better

Model

MCP Atlas

Toolathlon

Gemini 3.5 Flash

83.6 %

56.5 %

GPT-5.5

75.3 %

55.6 %

Gemini 3 Flash

62.0 %

49.4 %

Source: Google DeepMind: Gemini 3.5 Flash Model Card. Higher is better. These are provider evaluations for the stated model and test configurations, not measurements of business phone calls.

Large and small models on the same task benchmark

Tau2-Bench Telecom by model and reasoning effort

Model

Reasoning effort

Result ↑

GPT-5.4

xhigh

98.9 %

GPT-5.4 mini

xhigh

93.4 %

GPT-5.4 nano

xhigh

92.5 %

GPT-5 mini

high

74.1 %

Source: OpenAI: GPT-5.4 mini and nano, Tau2-Bench Telecom. Higher is better. The reasoning effort is part of each result; it is not a recommendation to use that setting for every time-sensitive voice response.

Use these numbers to shortlist candidates. Then evaluate your real tasks: requesting a missing part of an address, distinguishing similar products, or withholding an action when a required input is absent. A model that leads a broad benchmark may not offer the best combination of speed, cost and quality for a short reception dialogue.

Pipeline response time has several components: detecting the end of speech, transcription, response generation, any tool lookup and speech output. Some work can overlap. Adding figures from unrelated leaderboards does not produce a reliable end-to-end benchmark.

Six use cases for more natural calls

These scenarios reflect common Famulor workflows for reception, lead qualification, appointment booking and support. The dialogues are invented examples, not published customer recordings.

1. Reception: clarify the request mid-conversation

While the assistant explains the next steps, the caller adds, “This is about maintenance, not a new installation.” The agent incorporates the clarification and asks for the relevant details. The team receives a correctly categorized request. Full Duplex is worth testing when additions are common; Pipeline remains a useful baseline for a structured intake process.

2. Booking: incorporate a new time constraint

“Tuesday could work.” As the reply begins, the caller adds, “Only after school, though.” The agent asks for the specific time and checks availability again. It confirms the appointment after a successful calendar action. Evaluate whether corrected preferences lead to bookings that actually match the caller’s needs.

3. Sales: adjust qualification when the scope changes

A prospect mentions one location, then adds three more. That changes the relevant solution and the questions to ask next. The agent should combine the new details and send structured information to the CRM. Both conversation flow and data accuracy matter in this workflow.

4. Support: handle a corrected reference number

“The order ends in 42. Sorry, 24.” The agent continues with the corrected number and confirms sensitive matches before taking action. Full Duplex can help with self-corrections, while the connected search still needs to return an unambiguous match or trigger a follow-up question.

5. Field service: capture extra access details

A customer describes a fault and adds, “You can only get in through the courtyard.” The agent includes the detail in the service request. Clear business rules determine what information to capture and when to hand over to the team. The conversation mode supports those rules; it does not define them.

6. Brand voice: choose the right balance

A company uses a particular TTS voice for customer contact. Half Cascade can combine direct audio input with that separate speech output. If frequent additions shape the calls, also test Full Duplex with a native voice. Use identical content and assess clarity, brand recognition and conversational flow separately.

Set up Full Duplex in Famulor

  1. Open the relevant assistant. In an eligible workspace with beta features enabled, select Realtime → Full Duplex (Beta).

  2. Listen to the native voice previews. Choose a compatible voice using vocabulary from your workflow. Full Duplex previews use the actual conversation voice.

  3. Review greetings and announcements. Check the opening, consent message and tool announcements. An uploaded greeting keeps its recorded voice.

  4. Define actions clearly. Specify required inputs, follow-up questions and unsuccessful outcomes. Spoken confirmations should reflect successful tool results.

  5. Test through the intended contact channel. Include interruptions, brief acknowledgments, corrected details, silence and the end of the conversation. Reuse the scenarios you run with your existing assistant.

The update also improves startup with longer assistant instructions, recognition of configured actions and consistent recording of conversation usage. Hanging up after the farewell no longer opens an unnecessary new voice connection.

Natural announcements and required wording

The compatible native voice is used for Full Duplex conversations and the announcements intended for that mode. Ordinary announcements may be rephrased in conversation. Required consent announcements are checked and played completely before consent is accepted. Test this flow with your actual wording and enabled actions.

The editor preserves a fallback voice and explains conflicts with output filters before switching. Review the editor’s guidance before saving the configuration so the new conversation voice fits your existing workflow.

Build a useful comparison for your business

A meaningful comparison uses the same tasks under similar conditions. For example, start with 20 to 30 representative dialogue cases, repeat each one, and alternate the order of the configurations. This is a suggested starting point for your evaluation, not a test series conducted for this article.

  • Record the conditions: language, phone or web channel, region, background noise, voice, model version, prompt and connected tools.

  • Measure conversation flow: time to an audible response, reactions to corrections, unnecessary interruptions and repeated questions. Separate median behavior from slower cases.

  • Check the outcome: was the correct appointment created? Are the CRM fields accurate? Did the assistant explain a failed action honestly?

  • Listen to the output: assess clarity, pronunciation and brand fit. A percentage cannot replace this listening test.

Document cases where a simpler setup performs better, too. A short intake call may have different priorities from an advisory conversation with frequent additions. Choose the configuration that fits the workflow and the result you need.

Frequently asked questions

Who can select Full Duplex (Beta)?

The option appears in eligible workspaces with beta features enabled, suitable Realtime access and a supported region. At the time of writing, Global and US workspaces are supported; the option is not yet available for EU workspaces. Your editor shows the options available to your workspace.

Will existing assistants switch automatically?

No. Existing assistants retain their engines and settings. You choose which assistants use the new mode.

Can I keep my existing TTS voice?

The ongoing Full Duplex conversation uses a compatible native voice. Uploaded greetings keep their recorded voices. If a particular standalone TTS voice must speak the entire dialogue, compare Pipeline and Half Cascade.

Is Full Duplex always faster?

The published tests show advantages for certain configurations. Actual response time also depends on your channel, region, prompt, tools and conversation. Use benchmarks to shortlist options and your own workflow to decide.

What does the new mode cost?

Check the current terms and usage information in your Famulor workspace. External model prices and benchmark timings do not provide a reliable cost calculation for a complete Famulor call.

Start with a workflow you know well

Choose an assistant with a clear task: capturing a request, finding an appointment or classifying a support issue. Compare its existing conversation path with Full Duplex using the same scenarios. That will show where listening and speaking together helps your callers.

The Famulor Full Duplex introduction video. YouTube

Open Famulor and check Full Duplex (Beta) in your assistant. Find more product updates on the Famulor blog.

SM
Sarah Müller

Autor en Famulor

Asistente telefónico IA

Todo incluido, un plan. prueba Famulor

IA de voz, flujos de trabajo e integraciones en una plataforma.

Llamada entrante de Famulor AI en un smartphone
Newsletter

Responde primero. Crece rápido.

Suscríbase para recibir las últimas noticias, actualizaciones de productos y contenido de IA seleccionado.