Robots Atlas>ROBOTS ATLAS
Evaluation

WebArena-Infinity

2026ActiveUpdated: 17 September 2026Published
Key innovation
Automatic, scalable generation of realistic web environments with verifiable tasks — instead of a single, hand-built test suite.
Category
Evaluation
Abstraction level
System
Operation level
Evaluation (runtime)Agent runtime
Use cases
Evaluating browser agentsTraining general-purpose web agentsContinuous benchmarking (leak-resistant)Verifying improvement in self-improvement loops (e.g. DarwinX)

How it works

The engine starts from real-world artifacts; coding agents then build a working web environment while browser agents test it; a generate–test–audit–refine loop produces tasks with automatically verifiable success conditions. This yields dozens of environments and thousands of tasks instead of a single fixed suite.

Problem solved

Static, hand-authored web-agent benchmarks saturate quickly and do not scale; WAI automates the production of many environments and tasks with verifiable success, reducing data leakage and enabling continuous evaluation.

Evolution

Original paper · 2026 · web-arena-x (March 2026) · Shuyan Zhou
WebArena-Infinity: Generating browser environments with verifiable tasks at scale
Shuyan Zhou
2023
WebArena (CMU)

A static, hand-built benchmark of realistic web environments (Shuyan Zhou et al., NeurIPS 2024).

2026
WebArena-Infinity
Inflection point

Automatic, scalable generation of environments and verifiable tasks (March 2026).

2026
Used in DarwinX

DarwinX reports pass@1 rising to 93.0% (audit-clean) on the official 10-application / 1,260-task suite.