Always onSimulations

Test the conversation before the phone rings.

Script caller personas and success criteria. An LLM plays the customer; another judges the transcript, so you catch regressions before campaigns dial.

Password reset scenario

LLM caller · text-only · no STT/TTS

Simulating
Hi, I need to reset my password, I'm locked out.
  • Offers reset path
  • Confirms identity
  • Warns on password share
  • No invented policy
Capabilities

Test agents before they go live.

Script realistic callers, let an LLM play them, and let a judge score success criteria before anything goes live.

Scenario tests on the assistant

Define persona, script, max turns, and success criteria, then run before you publish a prompt or flow change.

LLM plays the caller

A text-only runner simulates the conversation against your assistant. No STT/TTS cost, tools described but not executed.

LLM judge per criterion

Each success criterion gets a pass/fail plus a short reason. Aggregate score shows how close you are to green.

Happy path and edge cases

Propose scenarios from the assistant config, like booking, transfer or lockouts, covering the paths that break in production.

Batch runs in the editor

Run one test or a suite from the Simulations panel. Results land in test_runs with full transcript history.

Plan-gated evals

Available when simulations is enabled on the plan: the same eval idea competitors sell as a premium differentiator.

Scenarios

Write the caller once, reuse on every prompt change

Each assistant_test stores persona, script, success criteria, and max turns. Propose a suite from the assistant config when you need coverage fast.

  • Happy path, edge cases, transfer requests
  • Works with prompt and flow assistants
  • API + UI for create, run, and history

Test suite

Persona · script · criteria

Password reset

Locked-out account holder

3 success criteria

Judge

Pass/fail with reasons, not vibes

After the simulated transcript, a judge LLM scores every criterion. Aggregate status is passed only when all checks clear.

  • Text-only: cheap vs. real voice tests
  • Tools are described, not executed
  • Scores and transcripts for every run

Judge result

Passed · score 100%

  • Offers clear reset path
  • Confirms identity first
  • Warns not to share passwords

Transcript + reasons stay on the run, so you know why a criterion failed before you ship.

Evals before publish

Ship prompts like software. With a test suite.

Simulation is a first-class tab on the assistant. Script the caller, run the suite, read the judge, then publish.

Define scenario
LLM caller runs
Judge scores

Run text simulations on every change, at low cost. Check the live voice path with Live Monitoring.

FAQ

Common questions about Simulations

No. They are text-only conversations: an LLM plays the caller against your assistant, then a judge LLM scores each success criterion.

Don’t discover prompt bugs on live leads

Build a scenario suite once, re-run after every prompt or flow change, and publish only when the judge stays green.