Robots Atlas>ROBOTS ATLAS
Evaluation

Coding Agent Index

2026ActivePublished: 23 September 2026Updated: 23 September 2026Published
Key innovation
Reduces coding-agent evaluation to a single number by averaging three benchmarks that measure different facets of software engineering โ€” implementation, terminal workflow and repository understanding.
Category
Evaluation
Abstraction level
System
Operation level
Agent runtimeSystem
Use cases
Comparing coding agents with a single metricRanking models for work on codeSelecting a model for an agentic harnessTracking progress in models' engineering capability

How it works

The agent is run on the three component benchmarks. DeepSWE v1.1 covers implementation ability, Terminal-Bench 4.0 covers work in a terminal environment, and SWE-Atlas-QnA covers code repository understanding. The partial results are averaged without weighting โ€” all three components enter the index with equal share. The simplicity of the aggregation is deliberate: it lets anyone reproduce the index value from the published component results and see which component drives a model's change in ranking.

Problem solved

Coding agents are assessed with many benchmarks, each measuring a different slice of the work; results are often contradictory and hard to compare across models. The index solves this organizationally โ€” it provides a single, explicitly defined number whose composition can be inspected.

Components

DeepSWE v1.1Implementation component

Measures the agent's ability to carry out software engineering programming tasks.

Terminal-Bench 4.0Terminal component

66 terminal-based tasks spanning software engineering, system administration and data processing; pass/fail scoring via a test suite, pass@1 averaged over 3 repeats.

SWE-Atlas-QnARepository-understanding component

Tests how well the agent understands the structure and content of a code repository, not merely whether it can write correct fragments.

Implementation

Implementation pitfalls
Averaging masks the capability profileMedium

Two models with the same index can have entirely different profiles โ€” one strong in the terminal and weak at repository understanding, the other the reverse.

Fix:Always read the component results alongside the index, especially when selecting a model for a specific use case.
A version change shifts the scaleMedium

The index is versioned (v1.5), and swapping a component benchmark shifts the values โ€” results from different versions are not directly comparable.

Fix:Compare only results from the same index version, and cite the version number.

Evolution

2026
Index version v1.5
Inflection point

Artificial Analysis describes the Coding Agent Index as the average of DeepSWE v1.1, Terminal-Bench 4.0 and SWE-Atlas-QnA.