Oct 2, 2026

Auto-evaluating Anthropic long-context models: a 100K-token test

Design a reproducible long-context evaluation for Anthropic models using retrieval placement, distractors, synthesis, citations, latency, token usage, and cost.

DEEP DIVE9 min readThe Currai team / Research

A 100K-token context test should measure whether a model finds, combines, and cites relevant evidence across a long input—not merely whether the API accepts the tokens. Evaluate retrieval position, distractor resistance, multi-document synthesis, unsupported claims, latency, and cost under a pinned Anthropic model version.

Context limits and model behavior change. Verify the current Anthropic model documentation before running or publishing results.

Build controlled long contexts

Create documents with known facts, realistic distractors, conflicting versions, and answerable plus unanswerable questions. Vary relevant evidence near the beginning, middle, and end. Include questions requiring evidence from several locations.

Do not repeat one synthetic sentence thousands of times; that measures a narrow needle task. Mix structured tables, prose, headings, and duplicated language.

Define evaluation dimensions

Measure answer correctness, evidence recall, citation precision, synthesis, abstention, and robustness to distractors. Record time to first token, total latency, input/output tokens, cache behavior, and cost.

Use deterministic checks for known facts and citation IDs. Use a calibrated judge for synthesis quality. Review a sample manually, especially disagreements.

Compare context strategies

Run identical questions with full context, retrieval-augmented context, hierarchical summaries, and hybrid approaches. Long context is not automatically better: irrelevant tokens can increase cost and reduce focus.

Pin prompt, model snapshot, ordering, and settings. Repeat stochastic trials and report confidence rather than one cherry-picked run.

Trace every experiment

Store dataset version, context construction method, evidence locations, model, prompt version, usage, latency, and evaluation results. Avoid sending sensitive production documents into a benchmark without appropriate controls.

Currai generations record the prompt metadata, model, tokens, latency, and cost; evaluation scores remain connected to the exact run. That makes a long-context claim reproducible and comparable when Anthropic releases a new model or your context strategy changes.

03

Keep going with nearby topics from the Currai blog.