---
title: "Why agent capability keeps improving while reliability barely moves"
date: 2026-05-31
summary: "Three arXiv preprints from January and February 2026 — Princeton, MAESTRO, ReliabilityBench — converge independently on the same finding: capability scales, reliability doesn't, and single-run benchmarks systematically miss the gap."
slug: agent-capability-vs-reliability-gap
author: "Damiën Semler"
hero_glyph: chart
---

A Princeton group led by Sayash Kapoor and Arvind Narayanan posted a paper to
arXiv in February with a one-line finding worth pinning to the wall: ["Despite
steady accuracy improvements over 18 months of model releases, reliability
only shows modest overall
improvement."](https://arxiv.org/html/2602.16666v1) The paper is the third
in a short stack of preprints from late January and February 2026 to
converge independently on the same observation, alongside
[MAESTRO](https://arxiv.org/abs/2601.00481) from a Beihang–KAUST–MPI group
and [ReliabilityBench](https://arxiv.org/abs/2601.06112) from Aayush Gupta. Single-run agent
benchmarks systematically over-state production reliability. The gap
between them isn't a measurement artifact. It's the load-bearing fact.

The three papers approach the problem from different angles. The Princeton
paper compares capability progression to reliability progression across 18
months of releases. MAESTRO probes call-graph stability across 12
multi-agent systems. ReliabilityBench introduces controlled perturbations
and fault injection across 1,280 episodes. Their data does not overlap.
Their conclusion does.

<svg viewBox="0 0 640 300" xmlns="http://www.w3.org/2000/svg" style="max-width:100%;height:auto;display:block;margin:28px auto" role="img" aria-labelledby="cr-title cr-desc">
  <title id="cr-title">Capability versus reliability over 18 months of model releases</title>
  <desc id="cr-desc">A conceptual two-line chart based on the Princeton paper's Figure 1 framing. The capability line rises steadily across 18 months. The reliability line stays nearly flat across the same period. The gap between the two widens over time. The chart is illustrative; exact values vary by benchmark.</desc>
  <style>
    .ax { stroke: var(--c-fg-3, #6b7280); stroke-width: 1; }
    .gd { stroke: var(--c-fg-5, #e5e7eb); stroke-width: 1; stroke-dasharray: 3 3; }
    .cap { fill: none; stroke: var(--c-violet, #A855F7); stroke-width: 2.2; }
    .rel { fill: none; stroke: var(--c-pink, #F472B6); stroke-width: 2.2; stroke-dasharray: 5 3; }
    .lt { font-family: 'Geist', system-ui, sans-serif; font-size: 12px; fill: var(--c-fg, #0a0a0a); }
    .ll { font-family: 'Geist Mono', ui-monospace, monospace; font-size: 10px; fill: var(--c-fg-4, #9ca3af); letter-spacing: 0.14em; }
    .dot-cap { fill: var(--c-violet, #A855F7); }
    .dot-rel { fill: var(--c-pink, #F472B6); }
  </style>
  <text x="0" y="18" class="ll">CAPABILITY VS RELIABILITY · 18 MONTHS OF RELEASES</text>
  <!-- Grid -->
  <line x1="60" y1="60" x2="620" y2="60" class="gd"/>
  <line x1="60" y1="120" x2="620" y2="60" class="gd"/>
  <line x1="60" y1="120" x2="620" y2="120" class="gd"/>
  <line x1="60" y1="180" x2="620" y2="180" class="gd"/>
  <!-- Axes -->
  <line x1="60" y1="240" x2="620" y2="240" class="ax"/>
  <line x1="60" y1="40" x2="60" y2="240" class="ax"/>
  <!-- Capability line: ascending -->
  <polyline class="cap" points="60,220 130,200 200,178 270,150 340,128 410,108 480,88 550,68 620,55"/>
  <circle cx="60" cy="220" r="3.5" class="dot-cap"/>
  <circle cx="200" cy="178" r="3.5" class="dot-cap"/>
  <circle cx="340" cy="128" r="3.5" class="dot-cap"/>
  <circle cx="480" cy="88" r="3.5" class="dot-cap"/>
  <circle cx="620" cy="55" r="3.5" class="dot-cap"/>
  <!-- Reliability line: ~flat -->
  <polyline class="rel" points="60,200 130,196 200,200 270,194 340,188 410,192 480,184 550,180 620,178"/>
  <circle cx="60" cy="200" r="3.5" class="dot-rel"/>
  <circle cx="200" cy="200" r="3.5" class="dot-rel"/>
  <circle cx="340" cy="188" r="3.5" class="dot-rel"/>
  <circle cx="480" cy="184" r="3.5" class="dot-rel"/>
  <circle cx="620" cy="178" r="3.5" class="dot-rel"/>
  <!-- Labels -->
  <text x="625" y="58" class="lt" style="fill:var(--c-violet,#A855F7)">capability</text>
  <text x="625" y="181" class="lt" style="fill:var(--c-pink,#F472B6)">reliability</text>
  <!-- X axis -->
  <text x="60" y="260" class="ll">0 mo</text>
  <text x="340" y="260" class="ll" text-anchor="middle">9 mo</text>
  <text x="620" y="260" class="ll" text-anchor="end">18 mo</text>
</svg>

## What single-run benchmarks miss

MAESTRO's central instrument is a pair of graph-similarity metrics applied
to repeated runs of the same task. Jaccard similarity measures whether the
same edges (agent-to-agent calls) appear across runs; LCS similarity
measures whether the *order* of those calls is preserved. The numbers from
12 systems across Autogen, LangGraph, and ADK are: [Jaccard 0.86 on
average, LCS 0.65 on
average](https://arxiv.org/html/2601.00481v1). Read together, that means
the set of agent interactions is structurally stable across runs, but the
sequence in which those interactions happen is not. A benchmark that
reports "the system solved the task" doesn't notice the difference. A user
who needs the same answer twice does.

ReliabilityBench's instrument is closer to chaos engineering. The
benchmark runs 1,280 episodes across Gemini 2.0 Flash and GPT-4o, two
agent architectures (ReAct, Reflexion), and four domains (scheduling,
travel, customer support, e-commerce). Semantic perturbations at intensity
ε = 0.2 (paraphrasings of the original task that preserve meaning)
[reduce success from 96.9% to 88.1%](https://arxiv.org/abs/2601.06112).
That is an 8.8-point drop from a synonym substitution that a human would
not register as a different task. The paper's framing is direct: "existing
benchmarks for tool-using LLM agents primarily report single-run success
rates and miss reliability properties required in production."

The Princeton paper writes the synthesis of those two observations into
the literature. Its central claim, reproduced from the Figure 1 caption in
the preprint: [outcome consistency remains low across all models tested,
even as headline accuracy
climbs](https://arxiv.org/html/2602.16666v1). Calibration improves
(recent Claude variants in particular), but improvements in calibration
do not guarantee that a model can identify when it is likely to fail. The
authors call out the implication explicitly: "improving raw task
performance is insufficient for building dependable AI agents; reliability
requires targeted attention beyond capability scaling alone."

## Three different finishes, one shared shape

| Paper | Method | Headline number | What it proves |
|---|---|---|---|
| MAESTRO (Jan 2026) | Repeated runs across 12 MAS, Jaccard + LCS on call graphs | Jaccard 0.86, LCS 0.65 | Same interactions, different orders. |
| ReliabilityBench (Jan 2026) | 1,280 episodes with semantic perturbation and fault injection | 96.9% → 88.1% at ε = 0.2 | Tiny task rewordings collapse success. |
| Princeton (Feb 2026) | Capability vs. reliability over 18 months of model releases | Reliability "modest overall improvement" | Scaling capability does not move reliability. |

What's striking is that none of these are exotic measurements. Jaccard,
LCS, paraphrase robustness, multi-run consistency — they are first-year
methodology in any field that takes evaluation seriously. They are not
yet first-year in agent benchmarking. The mainstream leaderboards still
report a single number after a single run on a fixed set of prompts. The
papers above describe what happens when you do the obvious thing and run
the experiment twice.

The cost numbers from MAESTRO sharpen the architectural point. CRAG
finishes the same task with a median cost of $0.0010 and a median duration
of 42.8 seconds. Plan-and-Execute does the same kind of work for a median
$0.0126 with an [interquartile range of 30.6 to 356.6
seconds](https://arxiv.org/html/2601.00481v1) — an order-of-magnitude
spread within the IQR alone. A single-run benchmark cannot distinguish a
system with a tight latency distribution from one whose IQR spans 10× the
median. To a user submitting a real task, those are two different
products.

## Why this is structural, not transient

The Princeton group's distinction between calibration and discrimination
is the load-bearing piece for understanding what comes next. Calibration
is about being well-aligned in confidence: a model that says "I'm 80%
sure" being right 80% of the time. Discrimination is about being able to
tell, before failing, that a task is the kind a model will fail on.
Recent models have gotten meaningfully better at calibration. They have
not gotten meaningfully better at discrimination.

That asymmetry has consequences. A well-calibrated agent that cannot
discriminate failure modes will still produce confident-looking outputs in
exactly the cases where it should refuse. From a benchmark's perspective,
a 90% success rate looks fine. From a registry's perspective, the
question that matters is not "what fraction of tasks did this agent
complete" but "what fraction of *the kinds of tasks where you'd want a
guarantee* did it complete consistently." Those are different objects,
and the agent literature is starting to admit it.

The same shape appears in the multi-agent case. MAESTRO's authors are
explicit: ["optimizing reliability and efficiency in agentic systems is
fundamentally an architectural
challenge."](https://arxiv.org/abs/2601.00481) Architecture choices
(ReAct vs. Reflexion, CRAG vs. Plan-and-Execute, sequential vs. graph
routing) outweigh model choice for reliability outcomes. Swapping in a
better base model leaves the variance problem intact.

## The registry-side implication

A registry that publishes a single conformance score for an agent or MCP
server is doing what these benchmarks do: collapsing a multi-dimensional
reliability surface into one capability number. The funnel pattern
[Agenstry publishes](https://agenstry.com) (discovered, responded,
returned a valid card) is one direction of generalization. The
ReliabilityBench-style multi-condition surface (consistency under
repetition, robustness to perturbation, fault tolerance under degraded
infrastructure) is another. They are compatible and they are not the same.

The interesting design question is what a public registry should publish.
A one-number score is legible; nobody mistakes its meaning. A
reliability-surface that names ε, λ, and k explicitly is honest but it
is harder to render in a row of a directory listing. Most agent
leaderboards have so far chosen legibility over honesty, and the three
papers above are the cost of that choice now showing up in print.

## What we're watching

Three things, observable within the next two academic cycles:

1. **Whether any major leaderboard adopts a ReliabilityBench-style
   protocol.** Most public agent leaderboards still report single-run
   pass@1 numbers on a fixed prompt set. The first to publish a multi-run
   consistency column alongside its accuracy column will be the practical
   answer to ReliabilityBench's framing.
2. **Whether the calibration/discrimination split in the Princeton data
   widens or narrows in subsequent model releases.** If discrimination
   catches up to calibration, the practical impact of the reliability
   gap shrinks. If it doesn't, agents that report high confidence in
   failure modes will keep getting deployed.
3. **Whether MCP server compliance scoring picks up a multi-run
   dimension.** The [MCP authorization
   spec](https://modelcontextprotocol.io/specification/draft/basic/authorization)
   formalized OAuth 2.1 but says nothing about whether a server's
   behavior should be probed once or one hundred times. The first
   foundation-governed conformance profile that names "k repeated probes
   per scan window" as a required field will be the practical answer to
   these papers.

A capability number is the headline. A reliability surface is what your
agent actually experiences. The two were close enough to ignore for the
first eighteen months of this story. They no longer are.

## Sources

- [Towards a Science of AI Agent Reliability](https://arxiv.org/html/2602.16666v1) — Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, Arvind Narayanan (Princeton), arXiv:2602.16666, February 2026.
- [MAESTRO: Multi-Agent Evaluation Suite for Testing, Reliability, and Observability](https://arxiv.org/abs/2601.00481) — Tie Ma et al. (Beihang University, KAUST, MPI-SWS, LUMS), arXiv:2601.00481, January 1, 2026.
- [ReliabilityBench: Evaluating LLM Agent Reliability Under Production-Like Stress Conditions](https://arxiv.org/abs/2601.06112) — Aayush Gupta, arXiv:2601.06112, January 13, 2026.
- [Authorization — Model Context Protocol](https://modelcontextprotocol.io/specification/draft/basic/authorization) — MCP Specification, accessed May 2026.
