- Catch regressions before they ship — run the same suite of cases after every prompt change or flow edit and watch the pass/fail diff.
- Probe edge cases at scale — angry caller, mumbling caller, caller who switches language mid-call, caller who already knows the answer.
- Audit behaviour — prove to a customer or auditor that the agent hits required compliance phrases on every call.
- Fail PRs in CI — block a deploy when the score drops below threshold.
The four pieces
The relationships:
Designing a persona
Personas are reusable — write them carefully. A persona has:identity_prompt— second-person (“You are a 34-year-old tenant…”). Include the role, the situation, what brought them to the call.behavior_traits— JSON knobs the persona LLM is told to honour:patience,verbosity,cooperation,interruption_tendency,goal. Free-form — agents will read whatever keys you provide.language—heoren. Set this to match the agent under test.
- Keep
identity_promptunder ~120 words. Long prompts make the persona LLM rigid and break the illusion. - Put adjustable axes in
behavior_traits, not in the prompt. That way you can have one identity (“frustrated tenant”) and several variants (“patient version”, “impatient version”, “screaming version”) without copy-pasting. - Don’t pre-write the script for the persona. Describe their goal and let them improvise.
Writing success criteria
Every case has an array of assertions evaluated after the run. Each assertion has atype, a weight, and type-specific fields. The case passes when weighted_score >= pass_threshold (default 80).
The case-level
max_turns field (default 20) is a hard safety cap on the conversation, not a success criterion — if the conversation runs that long the run terminates with termination_reason: max_turns.
Weights default to 1.0. A case with three weight-1 assertions and one weight-3 assertion has a max score of 6; the heavy assertion is worth half the case on its own.
Polling contract: status vs queue_status
Two lifecycle fields exist on every run; they update at different moments:
status— what the customer-facing UI shows. Flips tocompleted/failed/cancelledthe moment the conversation ends.queue_status— internal worker pipeline state (pending→claimed→done). Only afterqueue_status === "done"are the scoring + billing fields written (score,pass_fail,evaluation,total_cost_cents).
queue_status === "done" before reading scoring or cost fields. The transient window between status === "completed" and queue_status === "done" is typically under 5 seconds but can be longer under contention. Reading score early returns null, not 0 — which can silently fail CI checks if you aren’t careful.
POST /agent-eval/cases/:id/run already blocks up to 60 seconds for short single-case runs and returns 200 only when fully done (queue_status === "done"). It returns 202 and the latest state if it times out. Suite-run is always async (202) — poll the aggregate.
CI integration
The simplest pattern: run the suite, poll until done, check pass rate. Eval runs do not emit webhooks today — polling is the only mechanism.Billing
Eval runs are charged against your normal credit balance. The current public rate card:
These are user-facing prices (covering the AI cost plus margin), not provider costs. A typical 10-turn case lands between 0.05 depending on prompt length. A 50-case regression suite for under a dollar is normal.
The
total_cost_cents field on each run is the exact amount debited. After-the-fact roll-ups are available via GET /billing/consumption?product=eval_run.
Failure modes worth understanding
status: failed— the worker hit an unrecoverable error (model 5xx, malformed flow_config, persona LLM refused). Theerrorfield has the diagnostic. The run is partially billed for any turns produced before the error.status: completed+pass_fail: false— the run executed cleanly but didn’t meet the pass threshold. This is the case you want to investigate.termination_reason: max_turns— the conversation hit the case’smax_turnscap. Often a sign the persona is adversarial enough that the agent never reaches a closing node, or the case’s success criteria are unreachable.
GET /agent-eval/runs/:id/turns) and walk them top to bottom. For flow agents, the flow_event rows mirror the flow’s routing decisions and are usually where bugs hide.