Testez les agents avant la mise en ligne.
Écrivez des callers réalistes, un LLM les joue, un judge note les critères de succès.
Scenario tests on the assistant
Define persona, script, max turns, and success criteria, then run before you publish a prompt or flow change.
LLM plays the caller
A text-only runner simulates the conversation against your assistant. No STT/TTS cost, tools described but not executed.
LLM judge per criterion
Each success criterion gets a pass/fail plus a short reason. Aggregate score shows how close you are to green.
Happy path and edge cases
Propose scenarios from the assistant config, like booking, transfer or lockouts, covering the paths that break in production.
Batch runs in the editor
Run one test or a suite from the Simulations panel. Results land in test_runs with full transcript history.
Plan-gated evals
Available when simulations is enabled on the plan: the same eval idea competitors sell as a premium differentiator.
Write the caller once, reuse on every prompt change
Each assistant_test stores persona, script, success criteria, and max turns. Propose a suite from the assistant config when you need coverage fast.
- Happy path, edge cases, transfer requests
- Works with prompt and flow assistants
- API + UI for create, run, and history
Test suite
Persona · script · criteria
Password reset
Locked-out account holder
3 success criteria
Pass/fail with reasons, not vibes
After the simulated transcript, a judge LLM scores every criterion. Aggregate status is passed only when all checks clear.
- Text-only: cheap vs. real voice tests
- Tools are described, not executed
- Scores and transcripts for every run
Judge result
Passed · score 100%
- Offers clear reset path
- Confirms identity first
- Warns not to share passwords
Transcript + reasons stay on the run, so you know why a criterion failed before you ship.
Livrez les prompts comme du code. Avec une suite.
La simulation est un onglet de premier plan sur l'assistant. Scriptez l'appelant, lancez la suite, lisez le judge, puis publiez.
Les simulations texte tournent à chaque modification, à faible coût. Vérifiez le parcours vocal réel avec Live Monitoring.
Questions fréquentes sur les simulations
Non. Ce sont des conversations texte : un LLM joue l’appelant contre votre assistant, puis un LLM judge score chaque critère de succès.
Testez les prompts avant que les campagnes composent.
Construisez une suite de scénarios une fois, relancez-la après chaque changement de prompt ou de flux, et ne publiez que lorsque le judge reste vert.