FC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 QFC Health Index1,428.60+0.42%Novo NordiskDKK 812.4+1.10%Intuitive SurgicalUSD 546.9−0.30%EU AI Act — Art. 6in forceM+7FDA 510(k) AI clearances (YTD)312+18 w/wNHS AI Diagnostic Fund£123mcommittedKarolinska trials open48+2Reimbursement CPT codes (AI)17+1 Q
Friday, 18 September 2026 · Oslo · London · New York

Research franchise

AI Agent Reliability Benchmark

A measurement of how completion rates fall as task chains lengthen, run against reproducible environments rather than vendor demonstrations.

Research question
How does an agent's completion rate change as the number of required sequential steps increases?
Cadence
Twice yearly
Led by
Martin Anderson, edited by Aleksi Virtanen
Next release
January 2027 data release

Editions

  • Methodology release

    AI Agent Reliability Report — January 2027

    The benchmark is built around one variable: chain length. Tasks are constructed in graded tiers requiring four, eight, twelve and twenty sequential tool calls, with the same underlying domain at every tier.

    Published 2026-09-16 · Methodology v1.0