Sep 16, 2026

How to trace and evaluate parallel function calling for data extraction

Evaluate parallel function calling for structured extraction with schema accuracy, coverage, ordering, conflict handling, retries, latency, cost, and complete traces.

TUTORIAL9 min readThe Currai team / Engineering

Parallel function calling can extract several independent records or fields in one agent step, reducing wall-clock time. Evaluation must verify that every required call was emitted, arguments match the schema, results map to the right source, and failures do not corrupt the combined output.

Trace fan-out and fan-in

Create a parent span for the extraction step, one child span per function call, and a final generation or span for aggregation. Record call ID, function and schema version, source reference, arguments, status, latency, retry, and result summary.

Parallel calls may finish out of order. Join results by stable call ID, never array position.

Build field-level ground truth

Store expected values, types, source locations, and allowed normalization. Include missing fields, repeated entities, tables, conflicting values, malformed inputs, and fields that must remain unknown.

Measure schema validity, field precision and recall, entity association, duplicate rate, and exact or tolerant value match. Final-record accuracy alone can hide one field copied from the wrong entity.

Test partial failure

Simulate timeouts, validation errors, provider rate limits, and one failing branch. Define whether the system retries that branch, returns a partial result with status, or fails the whole operation. Retries need idempotency when functions have side effects.

Measure wall time and total compute. Parallelism can lower latency while increasing tokens, concurrency pressure, and cost.

Evaluate aggregation

Use deterministic code for type validation, required fields, and joining. Use a semantic evaluator only for normalization or ambiguous source interpretation. Require provenance for extracted values so reviewers can trace them back to the input.

Treat source text and tool output as untrusted. Do not let embedded instructions change the extraction schema or agent policy.

Currai's nested spans show the full fan-out, each function call, and the fan-in result with latency and cost. Pair this with tool-calling evaluation to measure whether parallel execution improves the complete task rather than merely running more calls at once.

03

Keep going with nearby topics from the Currai blog.