---
title: "Within 18 days, Google killed its pixel-level agent and Anthropic improved theirs"
date: 2026-05-23
summary: "On April 16, Anthropic shipped Claude Opus 4.7 with 1:1 pixel-coordinate Computer Use and 3.75-megapixel vision support. On May 4, Google shut down Project Mariner and moved its features to API-first agents. Two of three frontier labs just publicly disagreed on what an agent should see, and the registry layer has nothing to say about the pixel-level branch yet."
slug: pixel-vs-code-level-agents
author: "Damiën Semler"
hero_glyph: branches
---

Two of the three frontier labs made structurally opposite decisions
about pixel-level agent control inside an eighteen-day window in
April and May 2026. On [April 16, Anthropic released Claude Opus
4.7](https://www.anthropic.com/news/claude-opus-4-7) with explicit
improvements to its Computer Use capability: 1:1 pixel-coordinate
mapping (Opus 4.6 required a scale-factor correction step) and vision
support up to 2,576 pixels on the long edge, roughly three times the
prior limit. On [May 4, Google shut down Project
Mariner](https://www.techspot.com/news/112334-project-mariner-dead-but-google-browser-controlling-ai.html),
its screenshot-based web-browsing agent, citing high compute costs,
reliability issues, and privacy concerns. Mariner's features were
absorbed into the Gemini API and the new Gemini Agent rather than
discontinued outright; the architecture, in other words, was retired.
The capability moved to API-first agents.

Two frontier labs are now placing different bets on what an agent
should *see*. The bet matters because pixel-level agent control is a
fundamentally different trust model from API-mediated agent control —
and the registry layer has nothing to say about it today.

## The split, drawn out

<svg viewBox="0 0 640 320" xmlns="http://www.w3.org/2000/svg" style="max-width:100%;height:auto;display:block;margin:28px auto" role="img" aria-labelledby="px-title px-desc">
  <title id="px-title">Pixel-level vs API-level agent control: a three-event timeline</title>
  <desc id="px-desc">A horizontal timeline from April 15, 2026 to May 12, 2026 with three labeled events. Cloudflare Project Think and Code Mode ship on April 15 and 16, taking the code-level position. Anthropic releases Claude Opus 4.7 with improved Computer Use on April 16, taking the pixel-level position. Google shuts down Project Mariner on May 4, exiting the pixel-level approach. The timeline shows the industry split as it unfolded.</desc>
  <style>
    .line { stroke: var(--c-fg-3, #6b7280); stroke-width: 1.4; }
    .dot-px { fill: var(--c-pink, #F472B6); }
    .dot-api { fill: var(--c-violet, #A855F7); }
    .dot-end { fill: none; stroke: var(--c-fg-3, #6b7280); stroke-width: 1.8; }
    .lt { font-family: 'Geist', system-ui, sans-serif; font-size: 12px; fill: var(--c-fg, #0a0a0a); font-weight: 600; }
    .lc { font-family: 'Geist', system-ui, sans-serif; font-size: 11px; fill: var(--c-fg-2, #404040); }
    .ll { font-family: 'Geist Mono', ui-monospace, monospace; font-size: 10px; fill: var(--c-fg-4, #9ca3af); letter-spacing: 0.14em; }
    .lab-px { font-family: 'Geist Mono', ui-monospace, monospace; font-size: 10px; fill: var(--c-pink, #F472B6); letter-spacing: 0.14em; }
    .lab-api { font-family: 'Geist Mono', ui-monospace, monospace; font-size: 10px; fill: var(--c-violet, #A855F7); letter-spacing: 0.14em; }
  </style>
  <text x="0" y="18" class="ll">AGENT VISION · APRIL 15 → MAY 4, 2026</text>

  <!-- Top track: API/code-level events -->
  <text x="0" y="58" class="lab-api">API / CODE LEVEL</text>
  <line x1="60" y1="82" x2="620" y2="82" class="line"/>

  <!-- Cloudflare Code Mode: April 15-16 ~ x=85 -->
  <circle cx="85" cy="82" r="6" class="dot-api"/>
  <line x1="85" y1="82" x2="85" y2="100" class="line"/>
  <text x="85" y="116" text-anchor="middle" class="lt">Cloudflare</text>
  <text x="85" y="130" text-anchor="middle" class="lc">Code Mode</text>
  <text x="85" y="143" text-anchor="middle" class="lc" style="font-style:italic">Apr 15–16</text>

  <!-- Google Mariner shutdown: May 4 -> 18/20 days from Apr 15 -> x = 60 + 560*0.95 = 60 + 532 = 592 -->
  <circle cx="592" cy="82" r="7" class="dot-end"/>
  <line x1="592" y1="82" x2="592" y2="100" class="line"/>
  <text x="592" y="116" text-anchor="middle" class="lt">Google moves</text>
  <text x="592" y="130" text-anchor="middle" class="lc">Mariner to API</text>
  <text x="592" y="143" text-anchor="middle" class="lc" style="font-style:italic">May 4</text>

  <!-- Bottom track: pixel-level events -->
  <line x1="60" y1="195" x2="620" y2="195" class="line"/>

  <!-- Anthropic Opus 4.7 Computer Use: April 16 -> ~x=100 -->
  <circle cx="100" cy="195" r="6" class="dot-px"/>
  <line x1="100" y1="195" x2="100" y2="213" class="line"/>
  <text x="100" y="229" text-anchor="middle" class="lt">Anthropic</text>
  <text x="100" y="243" text-anchor="middle" class="lc">Opus 4.7</text>
  <text x="100" y="256" text-anchor="middle" class="lc">Computer Use ↑</text>
  <text x="100" y="269" text-anchor="middle" class="lc" style="font-style:italic">Apr 16</text>

  <text x="0" y="285" class="lab-px">PIXEL LEVEL</text>

  <!-- Date markers -->
  <text x="60" y="305" class="ll" text-anchor="start">APR 15</text>
  <text x="340" y="305" class="ll" text-anchor="middle">APR 25</text>
  <text x="620" y="305" class="ll" text-anchor="end">MAY 4</text>
</svg>

Three events, three positions. Cloudflare bet on code-level orchestration
through a sandbox. Anthropic doubled down on pixel-level operation
through improved vision. Google exited pixel-level operation in favor
of API integration. None of those bets is necessarily wrong; they
describe three different views of what an agent's eyes and hands
should be.

## What "pixel-level" actually means

A pixel-level agent operates on the screen the user sees. It takes a
screenshot, reasons about what it can see, decides which pixel to
click, sends a mouse or keyboard event, takes another screenshot, and
loops. There is no API contract. The agent's understanding of the
interface comes entirely from vision, and its actions are bound only
by what an actual user could do at the same keyboard.

The capability that makes possible is broad: a pixel-level agent can
operate on any application a human can, including ones with no API,
ones with broken APIs, and ones whose APIs intentionally exclude
automation. That last category is large. CAPTCHA-protected sites,
SaaS dashboards built without machine consumers in mind, legacy
desktop applications, internal enterprise tools whose owners never
wrote a public API — the pixel-level agent reaches all of them.

The cost of that breadth is what Google [named in shutting Mariner
down](https://ai2.work/blog/google-kills-project-mariner-as-the-industry-pivots-to-api-first-agents).
Vision tokens are expensive (Anthropic's 2,576-pixel images at
megapixel scale carry significant token cost per screenshot). The
agent's reliability degrades whenever the UI changes its layout.
Privacy concerns scale with how much of the screen the agent has to
see, which in practice is everything. Each cost is real and known.
The question is whether the unique capability justifies them.

## The trust model is different in kind

The API-mediated agent stack we've covered across this series
([MCP servers](https://agenstry.com/blog/mcp-spec-vs-probe-evidence),
[A2A signed cards](https://agenstry.com/blog/a2a-signed-cards-key-not-identity),
[ERC-8004 identities](https://agenstry.com/blog/erc-8004-three-registries-honesty),
[Cloudflare Code Mode](https://agenstry.com/blog/code-mode-mcp-token-reduction))
shares a structural property. The agent's available actions are
bounded by a published specification, however informal that
specification turns out to be in practice. A registry can probe what
the spec claims and publish what the probe found. The probe is the
load-bearing claim and the spec is the artifact a consumer can read.

Pixel-level agent control inverts that property. The agent's
available actions are bounded only by what a pixel-level human-input
loop allows on the target machine. There is no spec to probe. There
is no card to sign. The capability surface is "the screen," and the
screen is whatever the operating system happens to render at the
moment the agent looks at it. A consumer cannot ask a registry "what
will this agent do" in the way they can ask of an MCP server,
because the answer depends on what the agent sees in the moment.

That distinction has direct security implications. The [supply chain
attack pattern from the previous
post](https://agenstry.com/blog/agent-supply-chain-mcp-injection)
relied on injecting a rogue MCP server into a config file. The
pixel-level equivalent would be UI redress — a malicious page that
visually impersonates the page the agent expects to see. Browser
sandboxing partly defends against this for web-only agents.
Desktop-wide computer-use agents inherit the same problem at OS
scope, and the defenses are less mature.

## What this means for the registry layer

A useful question for the next year of agent infrastructure is what a
registry would even publish about a pixel-level agent. The signal a
consumer needs is not "this agent's tool catalog" but something
closer to "this agent has been observed to behave safely in the
following pixel-level environments under the following input policies."
That kind of probe data is harder to gather than an MCP card probe.
It requires running the agent against real environments and
recording the actions it takes.

Anthropic's Opus 4.7 announcement covers half of this question by
specifying [automated real-time cybersecurity
safeguards](https://cybersecuritynews.com/anthropic-releases-claude-opus-4-7/)
that interpose between the agent and the actions it tries to take.
That is a runtime guardrail, not a registry signal. The two compose:
a registry can publish what guardrails an agent ships with, and a
probe can publish what behavior emerges when those guardrails meet
the screen.

[Agenstry's funnel approach](https://agenstry.com) was built for
spec-bounded API agents. The pixel-level branch of the agent stack
needs a different probe shape. The shape is observable, but it's not
the shape any registry is publishing today.

## What we're watching

Three things, observable within the next two quarters:

1. **Whether OpenAI picks a side.** Anthropic ships Computer Use,
   Google retreats from Mariner. OpenAI has shipped sandboxing for
   API-mediated agents and demonstrated browsing-style agent
   capabilities in research, but has not publicly committed to a
   pixel-level production offering at the Anthropic level. The next
   GPT-class release will indicate whether OpenAI sees Computer Use
   as a load-bearing capability or as a research direction.
2. **Whether the published vendor-side cost of pixel-level agent
   sessions converges.** Vision tokens, repeated screenshot rounds,
   and the runtime safeguards Anthropic is now shipping all cost
   real money. The first transparent cost-per-pixel-session pricing
   model will set the economic gravity for the rest of the field.
3. **Whether a public probe registry for pixel-level agents
   emerges.** The probe shape is different from MCP ("what does
   this agent do when shown the following screen") but the
   measurement is concrete and the population is small enough to
   sample. The first registry that publishes that data will define
   what pixel-level agent transparency looks like.

The headline reading of the last eighteen days is "Google killed an
AI agent." The structural reading is that two of three frontier labs
just publicly disagreed about what an agent's eyes should be. The
disagreement is real, the trade-offs are real, and the registry layer
that has spent the past year cataloging the API-mediated branch has
no equivalent for the pixel-level branch yet. The work that closes
that gap is the work the next chapter of agent infrastructure has to
do.

## Sources

- [Introducing Claude Opus 4.7](https://www.anthropic.com/news/claude-opus-4-7) — Anthropic, April 16, 2026.
- [What's new in Claude Opus 4.7](https://platform.claude.com/docs/en/about-claude/models/whats-new-claude-4-7) — Claude API Docs, accessed May 2026.
- [Project Mariner is dead, but Google's browser-controlling AI plans are not](https://www.techspot.com/news/112334-project-mariner-dead-but-google-browser-controlling-ai.html) — TechSpot, May 2026.
- [Google Kills Project Mariner as the Industry Pivots to API-First Agents](https://ai2.work/blog/google-kills-project-mariner-as-the-industry-pivots-to-api-first-agents) — AI2Work, May 2026.
- [Anthropic Releases Claude Opus 4.7 with Automated Real-Time Cybersecurity Safeguards](https://cybersecuritynews.com/anthropic-releases-claude-opus-4-7/) — Cybersecurity News, April 16, 2026.
- [Computer use tool — Claude API Docs](https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use-tool) — Anthropic, accessed May 2026.
