Inhoud samenvatten met:
A convincing cloned voice needs three separate approvals: Does it sound like the authorized target speaker, does the delivery sound natural, and is the audio technically clean? Hume’s Voice Replication Benchmark, published on September 10, 2026, assesses those qualities independently and adds an objective similarity measure. Its results are not a universal vendor verdict. They do show why one overall score is insufficient for a production decision.
For teams that want to clone an authorized voice in Famulor, the sequence is straightforward: verify permission and consent before upload, test realistic phone phrases, and approve the voice only after it passes through the actual call path. This guide focuses on speaker fidelity—not a general TTS comparison.
Key Takeaways
- Naturalness and speaker identity are not the same quality.
- Audio quality needs an independent assessment.
- Objective similarity complements human listening; it does not replace it.
- Without documented authorization and consent, there is no test and no rollout.
Why one quality score is not enough for a cloned voice
An aggregate rating hides why a voice succeeds or fails. The Hume Voice Replication Leaderboard, released on September 10, 2026, makes those differences visible. Hume tested 11 models with the same 25 reference voices and seven prompts. The set included five standard voices, five expressive references, and 15 voices selected for native and non-native English accent variation. Hume documents additional methodology and artifacts in the official leaderboard README. Hume both publishes the benchmark and sells voice-AI and evaluation services, so this is a primary source for Hume’s own test—not independent market validation.
Three paid, independent raters—blind to model identity—rated every generated clip for speaker identity, audio quality, and naturalness on one-to-five scales. Hume also calculated cosine similarity between TitaNet speaker embeddings. The three reference groups received equal weight in the aggregate results.
The useful finding is not the ranking but the separation between measurements. In this specific benchmark, Cartesia’s sonic-3.6-beta received the highest naturalness rating at 4.36 while ranking eighth for the human same-speaker rating at 3.63. VoxCPM2 led the same-speaker rating at 4.21 but ranked fifth for naturalness at 3.96. Hume explicitly says this does not demonstrate an inevitable trade-off. No model led every metric and every reference category.
These figures are Hume results under Hume’s published protocol. They are not Famulor benchmarks or performance figures for a voice created in Famulor. Versions, configurations, and access paths do not consistently match Famulor’s catalog. Hume also evaluated isolated TTS clips, not a phone channel or complete voice agent. A leaderboard position therefore cannot be transferred to your workspace or call workflow.
The production lesson: A polished demo clip does not prove that a person remains recognizable. Strong resemblance says nothing about clicks, clipped endings, or unnatural pauses.
Which three questions should you answer separately?
At minimum, an approval record should contain three parallel judgments. If you average them too early, one strength can conceal a critical weakness.
1. Does the output sound like the authorized target speaker?
Speaker identity means perceived similarity between reference and output: timbre, characteristic stress, rhythm, and pronunciation. A clone can capture the general sound yet miss the target person—or drift on numbers and longer explanations.
The boundary matters: speaker similarity is not identity verification. A high human or machine similarity rating neither authenticates a person nor proves that they consented. Security-sensitive identity checks require separate controls; a cloned voice should not be treated as an authentication factor.
2. Does the output sound natural?
Naturalness concerns delivery: pace, pauses, prosody, emphasis, and intonation. Does the voice remain plausible during a polite refusal, a long sentence, or an unfamiliar name?
Natural output can still sound like the wrong person. Conversely, a clip may capture the target speaker’s timbre while sounding artificial because of rigid pauses or misplaced emphasis. “Sounds human” belongs in a different field from “sounds like our authorized spokesperson.”
3. Is the audio technically clean?
Audio quality covers distortion, metallic artifacts, volume changes, hiss, noise, and clipped endings. A clean file can carry the wrong voice; a recognizable voice can still be unfit because of artifacts.
Assess this dimension first in preview and later through the phone channel. That comparison helps distinguish a problem already present in the synthesized output from one that becomes noticeable along the transmission path.
Where does objective speaker similarity help?
An embedding-based similarity measure can screen clips consistently and flag outliers. Hume nevertheless treats objective similarity and human judgments as complementary: their rankings did not always align.
Do not infer a universal threshold from that finding. The production question remains whether authorized listeners recognize the voice under the intended operating conditions and whether the same output also sounds natural and clean. A machine score can structure the test; it cannot own the approval decision.
Why must testing start with authorization and consent?
A real person’s voice should not be cloned first and assessed for lawful use afterward. Authorization is a precondition: if it cannot be documented, the process stops before upload.
Famulor’s current guidance on how to document consent for voice cloning requires clear consent before cloning a real person’s voice. Its checklist describes consent as voluntary, informed, specific, documented, and revocable. Famulor also states that the page provides general EU legal context and is not legal advice.
An internal approval record should capture:
- who provides the voice and who authorizes its use;
- the specific purpose, channels, and languages covered;
- duration, authorized access, and the withdrawal, deactivation, and deletion process;
- who owns initial approval and later changes;
- how callers will be informed about the AI or synthetic voice when the specific deployment requires disclosure.
This checklist is a governance tool, not a guarantee of legal compliance. Requirements vary by country, sector, people affected, and use case. Seek qualified legal advice when the commercial deployment is uncertain. Famulor’s documentation describes completed cloned voices as private to the workspace, but workspace privacy does not answer separate questions about access, retention, purpose limitation, and misuse prevention.
How do you build a production-oriented voice cloning test?
A useful test uses a small, fixed set representing the hardest parts of the future call. This lets you compare changes instead of accumulating unrelated impressions.
1. Control the reference material
Use authorized, single-speaker recordings only. Famulor’s current documentation on models and voices recommends clear, high-quality reference audio with steady, natural delivery and no background noise. Also check for room echo, clipping, changing microphone distance, and other voices. Those traits can affect the clone or make evaluation ambiguous.
Do not put a static minimum duration or file limit in your internal playbook. Famulor says that current requirements shown in the interface—or, for API use, returned by GET /api/v1/voices/clone/capability—are authoritative. Availability and capacity also depend on the active workspace and plan. Check them when you run the test.
2. Build a fixed set from the real call flow
For an authorized company voice used in customer service, a test set could include:
| Test case | Example phrase | What to listen for |
|---|---|---|
| Greeting | “Hello, this is Exampleworks’ digital service assistant.” | recognition, welcoming entry, brand tone |
| Appointment | “Your appointment is Thursday, September 24, at 2:30 p.m.” | date, time, pauses, numbers |
| Name and address | “The delivery is for Ms. Nguyen at 18B Goethe Street.” | proper names, consonants, letter-number shifts |
| Service update | “The replacement part has been ordered. We’ll contact you as soon as a delivery date is confirmed.” | continuity, factual emphasis, longer phrasing |
| Polite boundary | “I can’t confirm that for you. I can transfer you to our service team.” | empathy without exaggeration, clear handoff |
| Edge case | an uncommon abbreviation or domain term | pronunciation, stability, error note |
Add only conditions that actually occur in the deployment. If calls require multiple languages, test every required language independently. Hume’s reference design examined English varieties and non-native English accents; it does not establish German or broadly multilingual production quality.
3. Listen blind and rate each dimension separately
Play the reference and generated clip in sequence, but hide model and provider names during rating. Record three separate fields for every clip: speaker identity, naturalness, and audio quality. Add a brief, observable error note such as “final syllable clipped,” “digits too fast,” or “timbre similar, emphasis uncharacteristic.”
An internal “approve / revise / reject” scheme may be sufficient. If you use numbers, define what each level means before listening. Hume’s one-to-five scale is part of its research method, not a prescribed Famulor production standard. Do not invent a universal pass score. The right acceptance criterion is use-case-specific and consistent across test variants.
4. Change one variable at a time
Depending on the selected voice and model, the Famulor editor may expose compatible controls for rate, volume, stability, similarity, expressiveness, or style. Not every control applies to every combination. Select the voice first, change one available setting, and replay the same test set.
Change the reference, voice, language, and controls together, and the cause of improvement becomes unclear. Record the variant, date, configuration, and rated clips.
What must you test through the actual phone path?
A clean browser clip does not prove that the voice and its recognizability will survive a phone call. The device, background noise, network, and telephony path can all change perception. At least one end-to-end call therefore belongs in the approval process.
First play the same sentence in preview, then through the intended phone path. Ask:
- Does the target speaker remain recognizable on an ordinary mobile phone?
- Are quiet consonants or word endings lost?
- Do volume jumps, distortion, or extra noise appear?
- Are numbers, names, and addresses still intelligible?
- Do longer answers and transfer phrases remain natural?
- Are critical cases stable amid realistic background noise?
Latency, interruptions, tool calls, and conversation logic also belong in a go-live process, but they are not this article’s focus. Use the broader guide to test the full voice agent before go-live for those layers. Famulor’s general TTS rollout test goes deeper on names, numbers, and the call channel. The question here stays narrower: Is this authorized voice still the right, natural, and technically clean voice in a real phone call?
Repeat critical cases after changing the reference material, voice, language, or relevant model. A change that improves one sentence can still make another worse.
A realistic use case: an authorized company voice for service updates
Consider a midsize service company whose recognizable spokesperson voluntarily provides their voice for narrowly defined appointment confirmations and status updates. The AI phone agent is not meant to imitate that person without limits. It produces short, factual messages within a documented purpose.
Before upload, the team records the purpose, channels, duration, withdrawal process, and accountable owners. It then uses approved reference material to create a private cloned voice, if that feature is available in its active Famulor workspace. In the editor, the team tests the greeting, date and time, customer names, order numbers, one longer status sentence, and a human-transfer phrase. Brand, service, and operations reviewers listen blind and rate the three quality dimensions independently.
Only then does the team make test calls to several internal devices. Suppose the speaker is recognizable in short sentences, but longer explanations develop an uncharacteristic intonation. That voice should not receive blanket approval. The team can restrict use to stable, tested phrases, change one compatible control, or reject the variant. Consent and any required disclosure remain mandatory regardless of the quality result.
This fictional example does not claim higher conversion, trust, or customer satisfaction. Those outcomes would need their own measurement design and contextual safeguards.
How do you approve a voice without benchmark hype?
Approve the cloned voice as a tested profile, not as an unexplained aggregate score. A compact decision matrix keeps a polished first impression from masking other risks:
| Review area | Decision question | Approval logic |
|---|---|---|
| Authorization and consent | Is the intended use documented and approved before upload? | Binary gate: without “yes,” no test or rollout |
| Speaker identity | Do authorized listeners recognize the target speaker in critical phrases? | record separately; never average with naturalness |
| Naturalness | Do pace, pauses, prosody, and emphasis fit the use case? | record separately |
| Audio quality | Is the output free of disruptive artifacts and clipping? | record separately |
| Phone path | Do all three qualities survive a real call? | required for production telephony |
| Edge cases | Have names, numbers, domain terms, and long sentences been tested? | fix open issues or narrow the approved scope |
| Withdrawal and disclosure | Are deactivation, deletion handling, and required notices ready? | launch gate, not part of a quality average |
Weighting depends on purpose. Recognition is central for an authorized company voice. For a generic service voice, intelligibility may matter more—in which case cloning may be the wrong choice. Consent is never averaged with quality scores.
Always check the live options in your Famulor workspace. Hume offers a measurement idea, not a shopping list: neither a “winner” nor one score determines whether a voice meets your purpose in a real call.
What is community signal, and what is evidence?
A weak practitioner signal—not performance evidence: A Hiya job description in September’s Ask HN thread places speaker verification and deepfake detection alongside SIP/RTP, STT/TTS, and conversational AI in a proposed carrier voice stack (Ask HN, September 2026). This single-employer hiring signal suggests that authenticity is being treated as an engineering concern beside telephony. It does not establish industry consensus, production deployment, or the effectiveness of any solution.
The distinction matters. Hume measured under a published protocol. Hacker News can surface useful questions, but it cannot substitute for performance, security, or legal evidence.
Limits: what this test does not prove
Even a careful test remains specific to the profile you evaluated. Results depend on the target speaker, reference recording, text, language, emotional delivery, model version, settings, and phone channel. Hume’s 25 voices, seven prompts, and three raters per clip likewise cannot establish performance for every speaker and deployment.
In particular, the test does not prove:
- that the voice performs equally well in an untested language;
- that similar sound authenticates a person;
- that human listeners would be deceived in every context;
- that the deployment is lawful or automatically compliant;
- that callers prefer or trust the voice;
- that one provider or model is generally superior;
- that approval remains valid after later changes.
Plan regression tests for the phrases that matter most. When the model, voice, language, or relevant reference material changes, run those cases again. That turns a one-time listening exercise into an auditable approval process.
Conclusion: approve a cloned voice as a profile, not a single score
A natural voice is not automatically the right voice. For a production AI phone agent, evaluate speaker identity, naturalness, and audio quality separately—first in controlled clips and then through the real call path. Objective similarity can support that process, but it cannot replace human listening or a documented decision.
Most importantly, consent sits outside the quality calculation. Without documented authorization, a defined purpose, and a withdrawal and disclosure process, there is no approval. Combine that gate with a fixed test set, blind rating, and regression checks. The result will not be a universal quality seal. It will be something more useful: a defensible decision for a clearly bounded use case.
Test your authorized voice in the real call path
Create only a voice you are authorized to use. Then test the same realistic phrases in Famulor’s preview and phone-call path, rating identity, naturalness, and audio quality separately. Start with Famulor’s voice cloning workflow.
Voice cloning evaluation FAQ
Is a natural-sounding AI voice automatically a good voice clone?
No. Naturalness describes how human the delivery sounds. Speaker identity describes whether the output sounds like the intended person. They require separate listening judgments.
What is the difference between speaker similarity and speaker verification?
Speaker similarity describes how alike a reference and output sound. Speaker verification is a security process used to check identity. A similarity score does not authenticate a person.
Can an objective similarity score replace human listening tests?
No. Hume’s benchmark treats embedding similarity and human ratings as complementary signals because their rankings can diverge in relevant cases.
How much audio does Famulor require to clone a voice?
Do not rely on a static number. The current number of samples, file size, and maximum duration are shown through Famulor’s cloning capability or interface and may depend on the active workspace.
Does the voice owner need to consent?
Use only voices you are authorized to clone and deploy. Document purpose and scope before upload. Specific legal requirements depend on the country and use case; this article is not legal advice.
How should I test a cloned voice for phone calls?
Play the same realistic phrases in preview and through the actual phone path. Rate identity, naturalness, and audio quality separately, and repeat critical cases after every relevant change.
Auteur bij Famulor




