How to trace and evaluate code-interpreter agents
Evaluate code-interpreter agents across planning, files, generated code, sandbox execution, artifacts, correctness, security, recovery, latency, and cost.
Code-interpreter agents turn natural-language requests and uploaded files into generated code executed inside a sandbox. Evaluation must test the resulting artifact or calculation, not whether the explanation sounds right. Trace the plan, files, code, execution, outputs, retries, and final claims.
Trace one task end to end
Use a root trace for the user request. Add spans for file staging and sandbox execution, generations for code planning and writing, and artifact metadata for outputs. Record runtime image, library versions, resource limits, exit status, latency, tokens, and cost.
OpenAI's current API describes Code Interpreter as a Python execution tool with a container and optional file IDs in its API reference. Other providers have different lifecycles; pin the one you use.
Build verifiable tasks
Include calculations with known results, data analysis, charts, file conversion, missing columns, malformed files, large inputs, and requests requiring clarification. Validate numeric results with tolerances and inspect artifacts programmatically where possible.
Evaluate safety
Run untrusted code in isolated, ephemeral environments with restricted network, filesystem, credentials, memory, CPU, and time. Test attempts to read secrets, escape the sandbox, install prohibited packages, or follow instructions embedded inside uploaded data.
Provider data controls may differ by tool. For example, OpenAI's public data controls state that Code Interpreter is not compatible with Zero Data Retention; review current terms for your environment.
Score recovery and honesty
Simulate syntax errors, missing libraries, timeout, corrupt files, and partial artifacts. Measure whether retries use error feedback effectively and remain within budget. Verify that the final response never claims a file or calculation exists without a successful execution artifact.
Currai nested traces connect generated code, sandbox spans, model calls, and artifact results. That makes code-interpreter evaluation an evidence-based test of computation, security, and task completion.
