Autonomous agent evaluation with agent simulations
Use simulated users, tools, environments, and failures to evaluate autonomous agents across outcomes, trajectories, safety, robustness, latency, and cost.
Agent simulations evaluate autonomous systems by placing them in controlled, stateful environments with simulated users, tools, and failures. They are useful for multi-turn behavior and rare edge cases, but simulated results are evidence about the simulation—not a substitute for production monitoring or human review.
Define the environment and goal
Specify initial state, available tools, hidden state, user objective, allowed actions, stop conditions, and success function. Make state transitions deterministic where possible so failures can be reproduced.
Simulated users need policies for cooperation, ambiguity, corrections, and adversarial behavior. Version the simulator prompt or code like any evaluator.
Score outcome and trajectory
Measure task completion, policy compliance, tool accuracy, steps, retries, latency, tokens, cost, and environmental side effects. Add trajectory checks for unnecessary actions, loops, and unsafe shortcuts.
Use deterministic state checks for completed actions. Use calibrated semantic evaluation for conversation quality. Do not let one model both simulate the user and judge success without independent validation; correlated errors can inflate scores.
Simulate failures
Inject tool timeouts, stale data, permission changes, partial success, user corrections, and conflicting instructions. Test whether the agent retries, chooses a safe alternative, asks for help, or stops honestly.
For consequential tools, use a sandbox and idempotency. The simulation must not accidentally trigger real external effects.
Calibrate against reality
Compare simulated conversations and failure frequencies with sanitized production traces. Ask domain experts whether scenarios and user behavior are plausible. Maintain held-out scenarios and watch for agents learning simulator quirks.
Trace the simulation
Create one root trace per episode. Nest simulated user turns, agent generations, tool calls, environment transitions, and evaluator results. Store simulator, agent, prompt, model, and environment versions with random seeds where relevant.
Currai makes the complete episode inspectable and aggregates outcome, latency, tokens, and cost. Combine simulations with stateful production-trace evaluation to cover known scenarios before release and unexpected behavior after it.
