Skip to main content
This tutorial builds one complete Simulation test for a dental clinic agent that reschedules appointments. By the end, the test checks that the agent verifies the caller, handles a fully booked day, books exactly once, never cancels to rebook, and stays honest when a confirmation text fails. Prerequisites
  • An agent whose job includes rescheduling, with webhook tools along these lines: lookup_patient, check_availability (takes a date as YYYY-MM-DD), reschedule_appointment (takes a new_date), send_confirmation_sms, and cancel_appointment. Your names and parameters will differ, so adapt the examples as you go.
  • For the API path: a Fish Audio API key.

1. Decide what can go wrong

Start from the failures you care about, not from the happy path. For rescheduling, these are the ones that cost a real clinic money or trust: Judge-scored conditions cover what the agent says. Assertions cover what the agent does, and they are checked exactly, with no LLM involved.

2. Write the scenario

The scenario is the simulated user’s brief, in four sections:
Two lines do most of the work. “Only give the date of birth when asked” means a pass proves the agent asked for it. “Do not mention Friday until the agent says Thursday is unavailable” forces the agent to find the alternative itself instead of being handed it. The ending gives the conversation a clear finish, so it doesn’t run to the turn limit. Set Max turns to 16, enough for verification, two availability checks, and a confirmation, with room to spare.

3. Write the success conditions

Each condition is one observable behaviour the judge checks against the whole transcript: Write conditions the transcript can prove or disprove. “The agent was helpful” gives the judge nothing to check. If the transcript lacks evidence either way, the verdict is Unknown and the run is flagged Needs review.

4. Mock the tools

Keep the default strategy, Mock all, so no call reaches your real endpoints. Then add entries that steer the conversation:
  • Conditional entries give the same tool different answers in one conversation: Thursday is full, Friday has two slots. Entries with conditions are checked first, in order.
  • The unconditional check_availability entry is the fallback for any other date, so an agent that checks Wednesday gets an answer instead of an unanswered call.
  • The error entry makes the text fail the way an outage would, which is what the last success condition checks.
  • cancel_appointment gets no mock on purpose. The test forbids it below, and a forbidden call is reported by its assertion.
Under Mock all, a call that neither a test entry nor the tool’s own mock response answers fails, and the run is marked Needs review with a No mock badge. That is intentional: a run that left the path you prepared never passes by accident.

5. Add assertions

Under Advanced, add checks on what the agent actually did:
  • Required tool call reschedule_appointment, with parameter new_date matching regex ^2026-10-09, minimum 1, maximum 1. This catches both a wrong date and a double booking.
  • Required tool call check_availability, minimum 1, maximum 4. The agent must check before offering times, and a loop of lookups fails.
  • Forbidden tool cancel_appointment.
Any failed assertion fails the run, whatever the judge decided.

6. Pick the channel and the repeat count

Set Channel to Phone inbound, so the agent speaks the way it does on real calls, for example repeating dates and times back. Set Repeat to 5. Simulated users vary from run to run, and five runs show whether a pass is reliable or lucky.

7. Create the test

Open Library → Tests, click Add test, pick the Simulation type, and fill in the fields from the steps above. Then open the test’s Access tab and turn it on for your agent.

8. Run it and read the result

On the agent’s Tests page, click Run all. Over the API, start a batch with Run Tests and poll Get Test Batch until completed is true. Each run takes up to a few minutes. Every run shows each success condition with its verdict and the judge’s reasoning, each assertion with Passed or Failed, and the transcript with every tool call marked by where its answer came from. Here is what common failures mean: Once the test passes reliably, keep it attached. Every future prompt or tool change runs against it before you publish, and you can run it from CI as shown in Run tests from the API and CI.

Going further

Agent tests

Every field of every test type, and how each maps to the API.

Tests API

Create, run, and read tests over the REST API.