Skip to content
StellarRequiem · MCP security · agent control · verification

Alex Price / StellarRequiem

Verified work, or it doesn't ship.

I build security and verification tooling for AI-era systems: MCP authorization research, deny-by-default tool gates, and proof-carrying claims. Confident numbers without evidence are the failure mode. One rule runs through everything I ship: no belief without verification — every public result carries evidence a third party can re-run. The badge is the claim; the honest no is part of the deliverable.

FastMCP fixes 2 merged MCP benchmark 0/11 Control planes public Proof re-runnable Honest gaps stated

Four pillars

Public work clusters into four re-runnable surfaces. Claims stay weaker than evidence — merged PRs are not CVEs; local planes are not enterprise SOC.

MCP tool assurance

mcp-assure

Deny-by-default policy + hash-chained receipts for agent tool calls. Runtime gate, not a full SOC. Fixtures and claim limits: MCP assurance proof.

Runtime gateReceiptsNot enterprise SOC
Agent control planes

agent-control

Session-scoped browser/desktop/CUA through AdaptiveGate — browser-leash, desktop-leash, agent-soc. Local-first, arm-gated — not ambient OS takeover.

Local-firstArm-gatedNot ambient CUA
Scanner truth

mcp-bench

Independent, reproducible fixtures: do MCP security scanners catch authorization-logic bugs? Authz-logic cases reported honestly (including misses).

Re-runnableauthz-logic 0/11
Claim verification

verity-core · scorecheck

Refuse accuracy theater until statistical hygiene + re-run; adjudicate published benchmarks against raw logs (REPRODUCED / DID-NOT-REPRODUCE / CHERRY-PICKED).

Proof-carryingCI-oriented
Working paper · published

Authority Is Not Ambient

Systems report on a mediated control plane for MCP-style tools and local computer-use agents: deny-by-default tool gates, arm-gated browser/desktop leashes, assured host, agent-plane FREEZE. PDF + Markdown · not peer-reviewed · not enterprise SOC or unlimited CUA. All papers.

Working paperRe-runnable stackNot peer-reviewed
Release notes · Grok-native path

Grok-native control plane

How the local stack is wired for Grok Build as the full host agent plane, plus a light remote path from Grok chat/iOS via Connectors MCP and a median session bus — allowlisted status skills, shared transcript, stated claim ceilings. Technical paper: Grok-native median session plane. Not ambient phone RPA · not free shell from the handset.

Grok-firstWorking paperPublic-safeNot full iOS CUA

Responsible Security Research

Offensive technique applied under explicit authorization, with audit. Recent work centers on MCP / AI-infrastructure boundary failures: versioned authorization checks, replay/session isolation, and scanner blind spots around authorization logic.

Authorization framework · public

scope-gate

A deny-by-default authorization gate: test only what you're explicitly authorized to. Ships with a responsible-research charter. The boundary that makes dual-use work safe.

PublicDeny-by-default
Public upstream fixes

FastMCP merged fixes

Two security-relevant MCP boundary issues I identified were fixed upstream: versioned authorization checks and Streamable HTTP event replay isolation. Cited as public merged fixes, not CVE/GHSA claims.

2 merged PRsRegression coverageNo advisory claim
Re-runnable fixtures · public

MCP authorization proofs

Local-only reference fixtures for MCP-shaped authz boundaries: resource/audience/scope binding, session replay, token handling, and a FastMCP signed-agent path with explicit allow + deny cases. Runtime tool gate: mcp-assure (deny-by-default policy + receipts — not a full SOC). Numbers and limits live on the proof page. See the client-readable synthetic review sample.

Re-runnableDeny paths explicitClaim boundaries on page
Discipline

Audit everything

Append-only run journals, claim cards, and hash-chain receipts keep the work grounded: what changed, what is verified, what is only a lead, and what still needs an operator gate.

Append-onlyClaim cardsHash-chained
Benchmark · public

mcp-bench

Do MCP security scanners actually catch authorization-logic bugs? An independent, reproducible benchmark seeded with real confirmed findings: 19 labeled cases, including 11 authz-logic cases across 10 root-cause classes, 2 control bugs, and 6 clean negatives. Scanners run only in a disposable CI runner.

PublicReproducibleauthz-logic 0/11
vulnerability researchPoC development authorization-gated testingcoordinated disclosure MCP / AI-infra securityFastMCP public fixes claim cardsred-team tooling
Claim ceiling · peers

Weaker than the evidence

Public claims stay weaker than re-runnable evidence: merged PRs are not CVEs; synthetic hit tables are fixtures, not published detection rates; local deny-by-default planes are not enterprise SOC. Adjacent industry work (e.g. institutional agent governance toolkits and open agent passport-style efforts) is acknowledged — this site ships local-first, consent-first, receipt-backed tooling you can re-run, not a monopoly on the idea.

Claim cardsRe-runnable proofNo overclaim

Verified AI Labor — the platform

Can a company be run as agents? Only if you can trust what each agent says it did — so I built the loop around verification, scope gates, false-positive rails, and public-safe workflow artifacts.

The gate

verity-core

Refuses a "95% accuracy" claim until it clears statistical hygiene — sample, out-of-sample, leakage, lift over base rate — then proves it: a claim ships a re-runnable command and the number must reproduce or CI fails. 17 domain packs · CI gate · MCP tool.

Proof-carryingCI gate
The labor

verified-ai-labor

A working prototype of a company run as agents — a 13-stage pipeline where every result-claim is verity-gated and every action hash-chain-logged, observable in a live console. Tests run locally; CI badge pending.

Verity-gatedHash-chained
Governed autonomy

The operating surface

Agents run a real workstation — but every action routes through a deny-by-default reference monitor first: read and local work proceeds, anything outward or destructive is held for a human, money and credentials are refused. Model proposes, code disposes — every decision hash-chained, every window operator-summoned. No capability the gate didn't grant.

Deny-by-defaultHash-chainedOperator-gated
The benchmark

groundtruth-bench

Citation faithfulness you can re-run to the same hash: a cryptographically committed corpus scored offline, byte-identical across machines — where RAG eval (RAGAS/ARES) is online, metered, and uncommittable. Reports where the scorer fails, not just the flattering number.

Byte-reproducibleCommitted corpus
The proof

calibration-log

A public, hash-chained prediction record scored over time (Brier + calibration). Honesty you can't doctor — it reports the real number whether there's an edge or not.

LiveHash-chained
The adjudicator

scorecheck

Adjudicates a published benchmark claim against its raw run-logs — REPRODUCED / DID-NOT-REPRODUCE / CHERRY-PICKED — sealed into a re-runnable receipt. Surfaces the dropped, flipped, and fabricated rows that re-run leaderboards and reproducibility badges miss; survived a 3-lens adversarial pass.

ReproducedCherry-pickedCI-green
Trust tooling

firewall · grounded · reality-anchor

Flag unverified claims in AI output; verify every cited claim is supported by its source; a research agent that grounds its answers or abstains rather than fabricate.

DeterministicOffline
Public operating manual

workflow bible

A public-safe breakdown of the full loop: scope, orient, map, draft, verify, ship, journal, and review. Copy the system without copying private targets, secrets, exploit steps, or live infrastructure details.

PublicCopyableSafety rails
Personal daemon · public demo

The Local Daemon

A client-side interactive console for catching signal, switching modes, and turning intensity into inspectable artifacts. Browser-only demo — not a remote backend and not wired to private infrastructure. Style-matched reels: demo media.

InteractiveClient-sideNo private backend
Agent beacon · findable tooling

Protect the agent (not the prompt alone)

Public self-defense map for tool-using LLMs: llms.txt, machine catalog.json, install paths for mcp-assure / leashes / blue-vaccine — claim-safe, not a hosted SOC.

Beaconcatalog.jsonWording ≤ evidence
Demo reels · real receipts

Console · gate · leash films

Short claim-safe demos in the same style as the X command-console and closed-loop posts: receipt films for mcp-assure check, ARM/deny posture, session-bus light path — not generative product fakes.

Real receiptsX-styleNot SOC theater

How my work is verifiable

Not a portfolio of assertions — a portfolio you can re-run.

01
Scope before action. Security work starts with explicit authorization, public-safe boundaries, and a clear operator gate.
02
Claim cards before headlines. Public claims get evidence, caveats, source links, and a disclosure-safety check.
03
Runnable proofs. Results come with the exact command a third party executes to reproduce them.
04
A public calibration log. Predictions are hash-chained and scored over time — the honesty is auditable, not asserted.
05
Honest gaps, stated. Every deliverable names what it did not verify. An unverifiable claim does not count.