Research franchise
AI Agent Reliability Benchmark
A measurement of how completion rates fall as task chains lengthen, run against reproducible environments rather than vendor demonstrations.
- Research question
- How does an agent's completion rate change as the number of required sequential steps increases?
- Cadence
- Twice yearly
- Led by
- Martin Anderson, edited by Aleksi Virtanen
- Next release
- January 2027 data release
Editions
Methodology release
AI Agent Reliability Report — January 2027
The benchmark is built around one variable: chain length. Tasks are constructed in graded tiers requiring four, eight, twelve and twenty sequential tool calls, with the same underlying domain at every tier.
Published 2026-09-16 · Methodology v1.0