Robots Atlas>ROBOTS ATLAS
Evaluation

AA-LCR (Artificial Analysis Long Context Reasoning)

2026ActivePublished: 23 September 2026Updated: 23 September 2026Published
Key innovation
Does not test needle-in-a-haystack retrieval, but forces integration of information from multiple places in a 100K-token input and reasoning toward a conclusion.
Category
Evaluation
Abstraction level
System
Operation level
InferenceModel
Use cases
Assessing the real usefulness of a long context windowComparing models for work on lengthy documentsA component of the Artificial Analysis Intelligence IndexSelecting a model for documentation and report analysis

How it works

Each question is embedded in a document set totalling roughly 100,000 tokens under the cl100k_base tokenizer. The answer is not in one place — the model must locate several relevant passages scattered across the input, relate them to one another and derive a conclusion. Answers are graded by a GPT-based equality checker that compares the model's output against a reference answer; scoring is pass@1, so only the first attempt counts. The set contains 100 questions, and the result enters the Artificial Analysis Intelligence Index with a 5% weighting.

Problem solved

Advertised context windows grow faster than models' real ability to use them. Popular needle-in-a-haystack tests check only the retrieval of a single fact, which a model can do without understanding the whole. AA-LCR forces the integration of scattered premises — the thing that actually separates a useful long context from a nominal one.

Components

~100K-token document setInput material

Length measured with the cl100k_base tokenizer; the premises needed for the answer are deliberately scattered across multiple passages.

100 integration-requiring questionsTest items

Questions constructed so that a single passage is not enough to answer correctly.

GPT-based equality checkerAnswer grading

Compares the model's answer with a reference; pass@1 scoring without repeats.

Official

Implementation

Implementation pitfalls
Tokenizer dependenceLow

Input length is counted with cl100k_base; models using a different tokenizer will see a different count of their own tokens for the same text.

Fix:Treat "100K tokens" as a measure of text length, not as the context-window load of a specific model.
Model-based judgeMedium

Grading by a GPT checker introduces its own noise and may treat substantively correct but differently phrased answers inconsistently.

Fix:Interpret small point differences between models cautiously, within the judge's noise band.

Evolution

2026
Version v1.1 with a 5% weighting in the Intelligence Index
Inflection point

AA-LCR v1.1 comprises 100 questions and enters the Artificial Analysis Intelligence Index with a 5% weighting.