Test het gesprek voordat de telefoon rinkelt.
Script caller-persona’s en success criteria. Een LLM speelt de klant; een andere beoordeelt het transcript — zodat je regressies vangt voordat campagnes bellen.
Password reset scenario
LLM caller · text-only · no STT/TTS
- Offers reset path
- Confirms identity
- Warns on password share
- No invented policy
Break the agent in the lab — not on the phone
Script realistic callers, let an LLM play them, and let a judge score success criteria before anything goes live.
Scenario tests on the assistant
Define persona, script, max turns, and success criteria — then run before you publish a prompt or flow change.
LLM plays the caller
A text-only runner simulates the conversation against your assistant — no STT/TTS cost, tools described but not executed.
LLM judge per criterion
Each success criterion gets a pass/fail plus a short reason. Aggregate score shows how close you are to green.
Happy path and edge cases
Propose scenarios from the assistant config — booking, transfer, lockouts — covering the paths that break in production.
Batch runs in the editor
Run one test or a suite from the Simulations panel. Results land in test_runs with full transcript history.
Plan-gated evals
Available when simulations is enabled on the plan — the same eval idea competitors sell as a premium differentiator.
Write the caller once — reuse on every prompt change
Each assistant_test stores persona, script, success criteria, and max turns. Propose a suite from the assistant config when you need coverage fast.
- Happy path, edge cases, transfer requests
- Works with prompt and flow assistants
- API + UI for create, run, and history
Test suite
Persona · script · criteria
Password reset
Locked-out account holder
3 success criteria
Pass/fail with reasons — not vibes
After the simulated transcript, a judge LLM scores every criterion. Aggregate status is passed only when all checks clear.
- Text-only — cheap vs. real voice tests
- Tools are described, not executed
- Scores and transcripts for every run
Judge result
Passed · score 100%
- Offers clear reset path
- Confirms identity first
- Warns not to share passwords
Transcript + reasons stay on the run — so you know why a criterion failed before you ship.
Ship prompts like software — with a test suite
Retell-style simulation is a first-class tab on the assistant: script the caller, run the suite, read the judge — then publish.
Text simulations — cheap enough to run on every change. Pair with Live Monitoring when you need the real voice path.
Veelgestelde vragen over simulaties
Nee. Het zijn tekstgesprekken: een LLM speelt de beller tegen je assistant, daarna scored een judge-LLM elk success criterion.
Ontdek prompt-bugs niet op live leads
Bouw een scenario-suite één keer, draai hem opnieuw na elke prompt- of flow-wijziging, en publiceer alleen als de judge groen blijft.


