Skip to content
All posts
· 7 min read ·

The agent stayed the same. Its benchmark score changed.

Business Bench shows how rescoring fixed outputs changes a measured lead. Embedded evaluation raises the next question: who can inspect the evidence behind a verdict?

An agent's measured lead can grow without another agent run. In the September release of Business Bench, changing the evaluator on retained outputs increased one system's lead from 16 to 34 accepted attempts. That is a 3.21 percentage-point change in the gap, with the generated work held fixed. The authors disclose the change in their self-audit.

At a different scale, Anthropic's September 18 embedded-evaluation announcement proposes giving evaluators access comparable to employees. These developments meet at a concrete question: what must an evaluator retain and disclose so that another person can challenge a verdict? Access, scoring rules and reviewable outputs each supply a different part of the answer. A score by itself supplies none of their history.

The evaluator is part of the result

Business Bench's September 16 campaign compared Proto with DeepSeek V4.1 Flash and Codex with GPT-5.6 sol across 187 tasks, with three attempts per system. The paper reports 447 versus 431 original accepted attempts and 507 versus 473 under its retrospective frozen scorer. These figures belong to that campaign; the site also publishes a separate September 28 comparison with different configurations. Misra, Misra and Dixit, paper revised September 27.

The arithmetic is straightforward: the original difference was 16 out of 561 attempts, approximately 2.85 percentage points. The later difference was 34 out of 561, approximately 6.06 points. Subtracting gives 18 attempts, or 3.21 points, attributable to changing the measurement rule on those outputs. This calculation says nothing by itself about which rule better represents acceptable work.

The same outputs receive different acceptance totals Business Bench September 16 campaign, 561 attempts per system. Proto rises from 447 original passes to 507 frozen-scorer passes. Codex rises from 431 to 473. The measured lead changes from 16 to 34 attempts without an agent rerun. Accepted attempts out of 561 Original scorer Frozen scorer Proto / DeepSeek 447 507 Codex / GPT-5.6 sol 431 473 Same retained outputs. Gap: 16 → 34 accepted attempts (+3.21 percentage points).

The chart is a comparison of judgments, not a progress curve. Its horizontal axis changes the scorer, not the model or time spent solving the task.

Disclosure changes what can be checked

ERP.AI both develops Proto and publishes Business Bench. Its self-audit says public ledgers and scoring code support checking accounting and inspecting rules, while private historical artifacts limit independent review of individual verdicts. It also cautions that the direction of a scoring change does not establish whether it corrects errors or introduces bias. Those disclosures are essential context for the comparison. Business Bench self-audit, accessed October 1.

The methods documentation further distinguishes original verdicts from retrospective scores and records grader errors separately from failed work. That separation matters: an evaluation program crashing does not establish that an agent's output was wrong. The scorer's fingerprint identifies the rules used, but identification alone cannot establish whether those rules capture the intended task. Business Bench methods, accessed October 1.

For anyone publishing agent evidence, this suggests a useful discipline: preserve a verdict as an event. Give it an evaluator identity, an artifact reference, an execution status and a timestamp. When a correction arrives, publish a second event linked to the first. A consumer should be able to recover both what the score says now and what changed its interpretation.

This does not require keeping an incorrect headline forever. It requires making the correction reconstructible.

Embedded access addresses another boundary

Anthropic's partnership with Accenture, led by Faculty, covers model evaluation, red-teaming, alignment assessments and safeguards. The announcement says embedded evaluators could follow development decisions and speak directly with employees. It also acknowledges that standards for access and reporting are unsettled, and that Anthropic will initially fund Accenture's work directly. Anthropic, September 18.

That arrangement concerns frontier-model oversight, while Business Bench concerns task artifacts. They are not interchangeable evaluation designs. Their shared lesson for an agent verifier is about the scope of evidence: knowing more about how a system was built and being able to rejudge a particular output are distinct capabilities.

Our proposed reporting model would make both visible. A report could say that a reviewer inspected development records but could not redistribute generated artifacts. Another could publish complete artifacts while lacking access to training or deployment decisions. Neither report needs to imply access it did not have.

Three objects worth preserving

A practical evaluation record should preserve three independently identified objects:

  1. The attempt. The request, input version, configured agent, tools, execution limits and completion status.
  2. The artifact. The output that existed before evaluation, with a digest and an explicit access policy.
  3. The judgment. The evaluator version, applicable criteria, result and any subsequent adjudication.

These are proposed publication requirements. Their value is that they prevent an evaluator update from silently becoming an agent improvement in a downstream report. They also let a buyer distinguish “we recomputed the totals” from “someone inspected the work and agreed with the verdict.”

A digest helps when a reviewer already has an artifact: the reviewer can check identity. If the artifact is unavailable, the digest should be accompanied by a statement explaining who can inspect it and under what conditions. Otherwise an evidence reference becomes a dead end precisely where a disputed verdict needs scrutiny.

For a registry that imports evaluation claims, the same separation should survive ingestion. Store the campaign identifier and evaluator version with the score. If a publisher rescored existing outputs, show that as a revision to the measurement. If it ran a new model on new inputs, show a new experiment. A single overwritten number cannot support either explanation reliably.

A test for the next release

We would assess a future evaluation release by attempting one narrow reconstruction: choose a changed verdict, identify the unchanged artifact, read both scoring rules, and determine the stated reason for the transition. If access restrictions prevent any step, the report should name the boundary.

The next stronger result would add independent adjudication of a declared sample of accepted and rejected work, with disagreement preserved. That would begin to test the evaluator's validity. Publishing more decimal places on an acceptance rate would not answer the same question.

The September evidence makes a useful distinction measurable. Agents produce work; evaluators produce claims about that work. Both deserve versioned evidence before their output becomes a trust signal.

Sources

Cite this post

@misc{semler2026evaluatorscores,
  author = {Semler, Damiën},
  title = {The agent stayed the same. Its benchmark score changed.},
  year = {2026},
  month = sep,
  url = {https://agenstry.com/blog/agent-benchmark-evaluator-changes},
  publisher = {Agenstry}
}
← Back to blog Agenstry