- An agent whose job includes rescheduling, with webhook tools along these lines:
lookup_patient,check_availability(takes adateasYYYY-MM-DD),reschedule_appointment(takes anew_date),send_confirmation_sms, andcancel_appointment. Your names and parameters will differ, so adapt the examples as you go. - For the API path: a Fish Audio API key.
1. Decide what can go wrong
Start from the failures you care about, not from the happy path. For rescheduling, these are the ones that cost a real clinic money or trust:
Judge-scored conditions cover what the agent says. Assertions cover what the agent does, and they are checked exactly, with no LLM involved.
2. Write the scenario
The scenario is the simulated user’s brief, in four sections:3. Write the success conditions
Each condition is one observable behaviour the judge checks against the whole transcript:
Write conditions the transcript can prove or disprove. “The agent was helpful” gives the judge nothing to check. If the transcript lacks evidence either way, the verdict is Unknown and the run is flagged Needs review.
4. Mock the tools
Keep the default strategy, Mock all, so no call reaches your real endpoints. Then add entries that steer the conversation:- Conditional entries give the same tool different answers in one conversation: Thursday is full, Friday has two slots. Entries with conditions are checked first, in order.
- The unconditional
check_availabilityentry is the fallback for any other date, so an agent that checks Wednesday gets an answer instead of an unanswered call. - The error entry makes the text fail the way an outage would, which is what the last success condition checks.
cancel_appointmentgets no mock on purpose. The test forbids it below, and a forbidden call is reported by its assertion.
5. Add assertions
Under Advanced, add checks on what the agent actually did:- Required tool call
reschedule_appointment, with parameternew_datematching regex^2026-10-09, minimum 1, maximum 1. This catches both a wrong date and a double booking. - Required tool call
check_availability, minimum 1, maximum 4. The agent must check before offering times, and a loop of lookups fails. - Forbidden tool
cancel_appointment.
6. Pick the channel and the repeat count
Set Channel to Phone inbound, so the agent speaks the way it does on real calls, for example repeating dates and times back. Set Repeat to 5. Simulated users vary from run to run, and five runs show whether a pass is reliable or lucky.7. Create the test
- Dashboard
- API
Open Library → Tests, click Add test, pick the Simulation type, and fill in the fields from the steps above. Then open the test’s Access tab and turn it on for your agent.
8. Run it and read the result
On the agent’s Tests page, click Run all. Over the API, start a batch with Run Tests and poll Get Test Batch untilcompleted is true. Each run takes up to a few minutes.
Every run shows each success condition with its verdict and the judge’s reasoning, each assertion with Passed or Failed, and the transcript with every tool call marked by where its answer came from. Here is what common failures mean:
Once the test passes reliably, keep it attached. Every future prompt or tool change runs against it before you publish, and you can run it from CI as shown in Run tests from the API and CI.
Going further
Agent tests
Every field of every test type, and how each maps to the API.
Tests API
Create, run, and read tests over the REST API.

