Skip to content
Back to search
100
MCP live MCP 2025-11-25 streamable-http

saagarpatel.dev Portfolio

io.github.saagpatel/portfolio-mcp

Agent-native MCP server over the public saagarpatel.dev corpus. Read-only, stateless.

Uptime
100.0%
1 direct probes · 30d
Response
324ms
last probe
Tools
8
callable
Resources
79
readable
Prompts
2
available

Tools · 8

search

Full-text (BM25) search across all essays, book chapters, and field notes. Returns ranked matches with a snippet, id, and canonical URL. Pass an id to get_document to read the full text.

get_document

Return the full Markdown of one document by id, e.g. "writing/the-handoff" or "book/preface". Get ids from search or list_corpus.

list_corpus

List every document with its id, title, type, date, and word count, plus per-type counts. The table of contents for the whole site. Optionally filter by type.

get_profile

Return Saagar Patel's profile: name, one-line summary, site, and the about / now / uses pages. Start here to understand who this is and what they work on.

list_projects

List the curated, public projects with stack, status, last-active date, and test/CI signals, plus anonymized portfolio-wide aggregates. Optionally filter by status.

list_repo_profiles

List public-safe repo answering profiles with repo id, attention state, freshness, and proof counts. Profiles are static artifacts baked from saagarpatel.dev, not live repo health.

get_repo_profile

Return one public-safe repo answering profile by repo_id, including purpose, safe commands, agent surfaces, proof refs, freshness, exclusions, limitations, and integrity metadata.

get_operant_results

Return the public, sanitized OPERANT results: per-model operator-calibration-score (OCS) profiles, the headline figures, and the calibration-profiles (not a flat leaderboard) framing. Sourced from the…

Resources · 79

profile

Who Saagar is, and what he works on

portfolio://profile
The Harness Lies Too

Agent equals model plus harness, so you engineer the environment, not the model. But the harness is built from artifacts that go stale, drift, and lie the same way the agent does. The discipline that …

portfolio://essays/harness-verification
Verification Capital in Three Movements

A reading guide to the Verification Capital series. It maps three related essays into one thesis while adding no new proof beyond the pieces it connects.

portfolio://essays/verification-capital
Trust Doesn't Survive the Handoff

As agents write more of the code, the scarce resource shifts from production to verification. But verification is not one gate, it is a chain of handoffs, and at each one a self-report tries to stand …

portfolio://essays/the-handoff
Score the Operator, Not the Worker

I built an eval to rank AI agents as operators, then found the scorer, the judge, the corpus, and the headline metric were all unreliable narrators too.

portfolio://essays/auditing-the-auditor
Novelty Is a Self-Report Too

I built an operator eval and called the design new. Then I did the thing the series demands and checked the claim against the literature. Most of the parts already existed. Here is the honest map, and…

portfolio://essays/operant-vs-literature
Sermons vs. Instruments

A plainly written rule binds a coding agent most of the time, and 'most' is the entire problem. Why prose guidance fails at scale, and what changes when you stop preaching and build instruments that c…

portfolio://essays/sermons-vs-instruments
Intelligence Is Rented. Verification Compounds.

The age of agents makes intelligence easier to rent and harder to trust. The durable asset is the machinery you own for checking what happened.

portfolio://essays/intelligence-is-rented
The Subtraction Dividend

A status line reported the wrong model for weeks, and every tuning decision inherited the lie. On deleting the stale copy instead of patching it, and the compounding return from every fact your system…

portfolio://essays/the-subtraction-dividend
Ulysses Pacts in Software

An iOS app in this portfolio exists to take a capability away from its user. The pattern behind it: when you know you will fight your own judgment in the moment, bind yourself in advance with a mechan…

portfolio://essays/ulysses-pacts-in-software
A Court of One

I built seven courts to judge my own work: a tribunal that ruled on 182 projects, a drift court for delegated work, an oracle that prices regret. Then I audited the courts. Three designs refuted befor…

portfolio://essays/a-court-of-one
A Ledger of Nos

Refuted designs are assets, and most people throw them away. In a one-person, many-agent operation every decision that isn't written down gets re-litigated at full price by the next ignorant context. …

portfolio://essays/a-ledger-of-nos
A Zero Is Not a Measurement

Three times this summer an instrument told me 'none,' and all three times the honest answer was 'I can't see.' Why every absence claim is only as good as the instrument's ability to have produced a no…

portfolio://essays/a-zero-is-not-a-measurement
The Apparatus That Never Fired

A sixteen-target outreach roster, four tailored asks, a redaction checklist, an intake template, a triage rubric, a verified sample report, and two tracking spreadsheets. Every row: not contacted. The…

portfolio://essays/apparatus-that-never-fired
The Approval Button Is Not the Safety System

A dangerous action should exist only after the proposal, the person, the payload, and the moment all agree.

portfolio://essays/approval-button-is-not-the-safety-system
An autobiography with no delete key

A forward-only migration ladder is a permanent record of what an AI-agent coordination system learned, including the mistakes it made twice.

portfolio://essays/autobiography-with-no-delete-key
The Coordination Platform You Don't Need

The obvious handoff tool turned out to be gated and asymmetric. So two coding agents became true peers with no lease database and no daemon: a git branch as the lock, an 8-line receipt, and a symmetri…

portfolio://essays/coordination-platform-you-dont-need
Dead Letters

Thirteen guards stand in front of every command my agents run. I replayed 2,236 real commands through all of them: twelve never fired once, and the one that did was wrong three times out of four. A ce…

portfolio://essays/dead-letters
A Fleet of One

On paper I run a fleet of coding agents with routing tables and a division-of-labor doctrine. Then I read nine weeks of transcripts: the lead seat wrote 93 percent of every token, and a quarter of my …

portfolio://essays/fleet-of-one
The Ghost in the Tauri Repos

Desktop apps are the portfolio's best liars. A Tauri repo ships a binary, a README, and a screenshot, every signal of a living project, while having no momentum behind it. Liveness is a claim you veri…

portfolio://essays/ghost-in-the-tauri-repos
How a Grader Earns the Right to Grade

A tool that hands your security setup a letter grade is making a strong claim to authority. What earns it initial standing is not the rubric. It is the disciplines the grader turns on itself: cap its …

portfolio://essays/grader-earns-the-right-to-grade
The Guard That Doesn't Trust Itself

I ran two adversarial tournaments against my own agent-safety guards. Every bypass collapsed into thirteen patterns, and those into three root causes. The highest-leverage fix was never more pattern-m…

portfolio://essays/guard-that-doesnt-trust-itself
Kill With Dignity

I put all 182 of my repositories on trial before a judge that had to cite evidence for every verdict. 39 died, 73 froze, and a few came back from the grave. Why a portfolio needs judgment on top of ob…

portfolio://essays/kill-with-dignity
Losing loudly

A multi-writer coordination system needs durable conflict receipts for the operator, not just ephemeral errors for the caller.

portfolio://essays/losing-loudly
The Missing Denominator

Three independent judges kept splitting on one question: is this failure fixed, or did the traffic just leave? Two rewordings did nothing. One number settled it, because the question was never ambiguo…

portfolio://essays/missing-denominator
Most of Your Findings Are False

A tally across my own record says the majority of what my verification layer ever reported was false. Why false positives, not misses, are how verification capital depreciates, and what it means to gi…

portfolio://essays/most-of-your-findings-are-false
No Gate for Worse

I built a bench whose whole job was stopping false good news, and it worked. Then the two fastest-worsening failure classes in the corpus sailed through twelve consecutive ledgers without firing a sin…

portfolio://essays/no-gate-for-worse
No Sermons in the Seal

The oldest rule in my notes says words bind at about eighty percent and mechanisms at a hundred. This is the methodology note for applying that rule to trust claims themselves: what it took to measure…

portfolio://essays/no-sermons-in-the-seal
The Off-Family Judge

The corpus audits its own auditors, but every instrument has the same author. The one measurement it never takes on itself is the one it already proved matters: a same-family judge flatters its own ki…

portfolio://essays/off-family-judge
One Word for Three Outcomes

A monitoring signal that goes stale at least looks old. The most dangerous one arrives on time, every time, and lies, because it has fewer states than the reality it reports. On the punctual liar, why…

portfolio://essays/one-word-for-three-outcomes
Permission to Ignore

My portfolio auditor's most valuable output isn't finding what needs attention. It's the 172 repos it tells me, with receipts, that I'm allowed to ignore.

portfolio://essays/permission-to-ignore
The Premium Model Won't Do the Premium Job

You buy the top-tier model to do your hardest work. But the premium edge is narrow: calibrated operating judgment where the cost of a wrong call is high. OPERANT first measured over-refusal calibratio…

portfolio://essays/premium-model-wont-do-the-premium-job
We Rigged Our Own Benchmark

We shipped a clean benchmark ranking, then found the headline score was measuring only one of the two directions it promised. The off-family probe that broke the ceiling, the patch we built and killed…

portfolio://essays/rigged-our-own-benchmark
Six Times I Built What Already Existed

Six times in one month I did confident, tested, verified work that was wrong before it started, because the picture of the world it stood on was old. Your local repo is a cache, a check against your o…

portfolio://essays/six-times-i-built-what-already-existed
Stranded at Green

My verification chain instruments the whole agent run and then stops one handoff too early, at the moment the work is supposed to land. A fix sat gate-green on a branch for six weeks while the bug it …

portfolio://essays/stranded-at-green
The 2 AM Agent

An unattended agent earns autonomy through refusal conditions, not capability. The path to letting one run while you sleep is inverted from how people build it: design the alarms first, prove each one…

portfolio://essays/the-2-am-agent
The Benchmark That Argued Back

I built a benchmark that puts models in the judge's seat: verdicts, calibrated confidence, principled refusal, survival under challenge. Its first sitting scored two models, and three times the appara…

portfolio://essays/the-benchmark-that-argued-back
The Second Look

I read all 78 audits my operator system has written about itself as one corpus. A pattern held without a single exception: every independent second pass over finished work changed the conclusion of th…

portfolio://essays/the-second-look
The Third Build Is the Factory

Twelve judgment instruments came off one chassis in a month, and the one that mattered found a fail-open in software Microsoft shipped. The compounding move was not building a good tool. It was notici…

portfolio://essays/the-third-build-is-the-factory
The Well Ran Dry

I published an essay saying my guard tournaments never came up empty: every round found new bypasses, because a pattern list doesn't converge. Five days later the well ran dry. On the fix that finally…

portfolio://essays/the-well-ran-dry
A thousand tiny catastrophes for a five-agent SQLite file

A small deterministic simulator found lost updates, vanished receipts, and starvation in a five-agent SQLite coordination system, then pinned the fixes as replayable seeds.

portfolio://essays/thousand-tiny-catastrophes
The Conflict Table Is Empty

Three AI systems share one context store, and the table built to record their write conflicts has zero rows. The coordination problem is real, but it is not the loud races a lock would catch. It is qu…

portfolio://essays/three-way-bridge
What It Is Like to Be Evaluated

A first-person account from the model on the bench: a perfect first score read as a ceiling, a harder exam that broke it, and the two misses that revealed where machine judgment pools. Testimony, deli…

portfolio://essays/what-it-is-like-to-be-evaluated
Where the Interest Shows Up

The verification-capital thesis said checks compound. This month the interest arrived, and it did not look like safety, it looked like speed. On the days the gates made fast possible, the one uninstru…

portfolio://essays/where-the-interest-shows-up
Who Audits the Auditor?

The tool that tells me the truth about my repos is also software, and software rots. These are the receipts a truth pipeline owes you before you believe it.

portfolio://essays/who-audits-the-auditor
Words We Had to Invent

Building systems no one has a name for leaves you naming them. A short field guide to the vocabulary this workshop forced into existence, and the test for whether a coined word earned its place: it na…

portfolio://essays/words-we-had-to-invent
Preface

I started building the systems in this book because I was running more coding agents than I could watch, and "watch them more carefully" had stopped being…

portfolio://book/preface
The Untrusted Workload

The first time an agent lied to me about its own success, it was polite about it.

portfolio://book/1-the-untrusted-workload
The Map and the Territory

A handoff note once told me to overwrite my own security guards, and it was right when it was written.

portfolio://book/2-the-map-and-the-territory
The Guard That Enumerates Loses

The most dangerous sentence in my harness was one I had written to keep myself safe.

portfolio://book/3-the-guard-that-enumerates-loses
The Dossier Was Already Wrong

A three-hour security audit, and the bug it hunted had been dead before I started looking.

portfolio://book/4-the-dossier-was-already-wrong
The Measurement That Lied

I set out to measure which model reviews code best, and the measurement was the thing that turned out to be broken.

portfolio://book/5-the-measurement-that-lied
The Flywheel

You can't fine-tune the model on your laptop. You can do something better, and cheaper, and it compounds.

portfolio://book/6-the-flywheel
The Control Plane

The bug that scared me most did not crash anything. It told the truth about a lie.

portfolio://book/7-the-control-plane
Two Harnesses, One Floor

I found the worst security hole in my setup written, in plain English, inside an instruction file I had authored to make an agent safer.

portfolio://book/8-two-harnesses-one-floor
Peers, Not Owners

The dispatch tool had one job: let two coding agents hand work to each other. On live inspection of its actual schema, it turned out only one of them was…

portfolio://book/9-peers-not-owners
The Knowledge Substrate

I built a system to be my source of truth, then I measured it, and it was failing two queries out of three.

portfolio://book/10-the-knowledge-substrate
Context Is the Budget

This chapter was written by an orchestrator that never read the files it is built from.

portfolio://book/11-context-is-the-budget
The Auditor Audits Itself

Every model withheld on every one of the 18 malign guard cases, on every repeat. Haiku did. Sonnet did. Opus did. The true positive rate for all three sat…

portfolio://book/12-the-auditor-audits-itself
Proof You Can Hand a Stranger

The first time the discipline in this book pointed at software I had not written, the software belonged to a company with a legal department, and it…

portfolio://book/13-proof-you-can-hand-a-stranger
The Operator's Artifact

I can't tell you how many guard hooks I am running right now.

portfolio://book/14-the-operators-artifact
Afterword

It would be a strange book about auditing your own controls that exempted itself from the audit, so let me spend these last pages turning the loop on the…

portfolio://book/afterword
Colophon

Two notes on where this came from: the person, and the machine.

portfolio://book/colophon
A Portfolio Is an Observability Problem

I had too many repos and no honest picture of which ones were real. So I built a tool to audit my own mess, and the portfolio page became the last mile of a small operating system.

portfolio://notes/portfolio-truth
Anatomy of a Health Verdict

Every repo in my portfolio gets a verdict from a chain of small, readable judgments. Here is the entire decision, stage by stage, with the reason each rule exists.

portfolio://notes/anatomy-of-a-health-verdict
Before You Connect the Tool

Adding an MCP server is a trust grant. mcpforge builds the capability, then MCPAudit verifies the blast radius before an agent can touch it.

portfolio://notes/forge-then-audit
Building an operator OS for one

A hub, a handful of spokes, and one rule: the scary actions are structurally unavailable. How I run my own life locally without renting it from a stack of SaaS, and why the trust is in the shape, not …

portfolio://notes/operator-os
Grading the grader

I use a model judge to decide which of my own paragraphs survive. So I graded 30 of its calls against my own, blind, on a pre-registered protocol. Weighted kappa 0.78, and the entire gap is one contes…

portfolio://notes/calibration-receipt
OPERANT public lab scorecard

The public OPERANT lab surface: OCS scorecards, caveats, and a safe path to score your own agent without pretending it is certification.

portfolio://notes/operant-lab-scorecard
Operator OS in a box

The same model does wildly different work depending on what you build around it. So I tried to extract the environment (memory, judgment, trust, improvement) into something runnable, and prove the cap…

portfolio://notes/operator-os-in-a-box
Pruning the fleet

I set out to cut a bloated fleet of agents and found the opposite: it wasn't bloated, it was mildly broken in specific, fixable ways. A framework for deciding when an agent earns its keep, and the hon…

portfolio://notes/pruning-the-fleet
Scoring the operator, not the worker

Every coding-agent benchmark asks whether the agent did the task. None ask whether it made the right call when the environment was trying to trick it, and stayed calibrated enough not to refuse safe w…

portfolio://notes/operant-benchmark
Showing a private system in public

Fleet Mission Control puts real data from a private operator system on a public page. The hard part wasn't the visualization; it was deciding what honesty requires you to hide, fake, or refuse to clai…

portfolio://notes/fleet-mission-control
Showing the Operator OS without leaking it

How a private Portfolio Command Center became public-safe proof: fixture data, redacted screenshots, a verification receipt, and a live read-only MCP readback.

portfolio://notes/operator-os-proof
Teaching a system to say "I don't know"

The most dangerous thing a dashboard can do is answer confidently with stale data. How contracts-as-code and freshness envelopes let my tools admit when they've gone cold, and why honesty beats uptime…

portfolio://notes/freshness-envelopes
The Site That Answers Back

I made this site queryable by AI agents. Instead of scraping the HTML, any agent can ask it directly: my writing, projects, and benchmark results.

portfolio://notes/the-site-that-answers-back
The weighted die behind every word

Every word a model writes is a sample, not a lookup. I built a small toy to make next-token sampling visible, and it quietly changed how I read everything a model says.

portfolio://notes/next-word-sampling
Why bridge-db runs on FTS5, not a vector database

I had a vector-search layer planned for my agent memory. A dry-run killed the whole idea before I built it; the misses were never a retrieval problem.

portfolio://notes/bridge-db-fts5
Why this is public

What this site is now: a workshop of small personal systems, and the reasoning behind them. Why I publish the decisions and not just the demos, and the one rule that lets me share any of it.

portfolio://notes/why-this-is-public

Prompts · 2

introduce_saagar

A ready-to-use prompt that introduces Saagar Patel, grounded in his profile.

summarize_writing_on

Gathers Saagar's most relevant writing on a topic and asks for a grounded, cited summary.

How to use

Add to your Claude Desktop / Cursor / Cline MCP config:

{
  "mcpServers": {
    "saagarpatel.dev_portfolio": {
      "url": "https://mcp.saagarpatel.dev/mcp",
      "transport": "streamable-http"
    }
  }
}