Robots Atlas>ROBOTS ATLAS
Evaluation

AutomationBench-AA

2026ActivePublished: 23 September 2026Updated: 23 September 2026Published
Key innovation
Evaluates business process automation through real REST API tool calls with objective pass/fail scoring, and additionally detects guardrail violations — not just whether the task succeeded, but whether the agent did something impermissible along the way.
Category
Evaluation
Abstraction level
System
Operation level
Agent runtimeSystem
Use cases
Evaluating agents for business process automationTesting models' use of REST API toolsMeasuring adherence to safety guardrailsA component of the Artificial Analysis Intelligence Index

How it works

The model is given a task from one of the business domains and a set of tools exposed as REST APIs. It issues a sequence of calls aiming to complete the workflow. The outcome is binary: the task either passes or it does not, with no partial credit and no model-based judge. In parallel the system checks whether guardrails were violated during execution — that is, whether the agent performed an operation it should not have. In total the set spans 657 tasks, and the aggregated result enters the Artificial Analysis Intelligence Index with a 5% weighting.

Problem solved

Claims about "agents automating business processes" are hard to verify, because success depends on the specific tool environment and on whether the agent breaks organizational rules along the way. AutomationBench-AA provides a reproducible environment with real REST API calls and an explicit record of guardrail violations.

Components

657 tasks across business domainsTest set

Tasks mirroring SaaS software workflows, spread across different areas of company operations.

REST API toolsExecution environment

The model operates through actual API calls rather than a simulated description of tools.

Guardrail violation detectionSecond axis of evaluation

Independently of task completion, it records whether the agent performed an impermissible operation.

Implementation

Implementation pitfalls
Binary scoring without partial creditMedium

An agent that completed 90% of the workflow and stumbled on the last call scores zero — results can therefore be jumpy.

Fix:Analyse the distribution of failures by step, not just the aggregate pass rate.
Confusion with AutomationBenchMedium

AutomationBench-AA is Artificial Analysis's harness instance; its results are not necessarily directly comparable with figures reported on AutomationBench itself.

Fix:When citing a score, always state whether it comes from the -AA variant or the source benchmark.

Evolution

2026
Inclusion in the Artificial Analysis Intelligence Index
Inflection point

AutomationBench-AA enters the index with a 5% weighting, covering 657 tasks with REST API tools.