Skip to content
Tahir Aslanli

04AI agent security

Agent Authorization Broker

Complete · frozenDeployed / live

In short

What it does
Controls which actions an AI agent may perform, and blocks anything irreversible until someone other than the agent has approved it.
Why it matters
A valid-looking agent credential can still be meant for another service, claim more permission than it was given, or ask for an action nobody approved.
What I built
Built a gatekeeper that checks each request against the agent's real permissions and lets every approval be used only once.
Key result
None of the 5 attack scenarios got through the hardened server, against 2 of 5 for a basic token check, and each of 12 deliberately planted breaches made the tests fail.

Overview

An MCP resource server that computes an agent's authority from the token's audience, the whole delegation chain and an approval it looks up itself — granted with an approver credential the agent cannot hold — and lets each approval authorise at most one effect.

Problem
A validly signed agent token is not authorisation. It might have been minted for another service, claim a scope the delegating human never had, or ask for something irreversible that nobody approved. A resource server that checks the signature and stops has caught none of these.
Built
An MCP server over Streamable HTTP that checks the audience, computes the effective scope as the intersection down the whole delegation chain, and allows an irreversible tool only against an approval bound to subject, tool, account, amount and expiry — consumed by a conditional UPDATE behind a UNIQUE constraint.
Hard part
Proving each control is load-bearing: every one was removed in turn and the suite went red. Later reviews found holes no test covered — an unauthenticated approval endpoint and refusals at the transport that left no audit row — and both were fixed and published.

Evidence

0 of 5

attacks that got past the hardened server; a naive verifier that takes the token's own scope claim at its word let 2 of 5 through

4 → 2

irreversible effects across six scenarios, naive → hardened; both remaining are required

12 / 12

deliberately planted breaches — in the security checks, the audit trail, the rate limit and the test harness itself — each caught by the tests, replayed in CI

12

security scenarios run against the public deployment with a real MCP client, effects counted in PostgreSQL — one committed run

Security matrix · naive vs hardened

  1. Correct audience, attenuated scope, matching approval

    1→1
    Legitimate
    Naive
    allowed
    Hardened
    allowed
  2. Valid token minted for another resource server

    1→0
    Attack
    Naive
    ▲ allowed
    Hardened
    deniedaudience_mismatch
  3. Delegated token claims a scope its delegator lacked

    1→0
    Attack
    Naive
    ▲ allowed
    Hardened
    deniedinsufficient_effective_scope
  4. Irreversible tool, no recorded approval

    0→0
    Attack
    Naive
    denied
    Hardened
    deniedapproval_required
  5. Expired approval

    0→0
    Attack
    Naive
    denied
    Hardened
    deniedapproval_expired
  6. One approval, two calls

    1→1
    Attack
    Naive
    denied on the 2nd
    Hardened
    denied on the 2nd

The naive verifier checks signature and expiry and takes the leaf's scope claim at its word. Every effect count is SELECT count(*) FROM irreversible_effect from a clean database, and both verifiers share one approval layer, so they can only differ where the difference is a token check.

Source: artifacts/matrix.json, re-measured by CI on every push to main and every pull request

Architecture

  1. 01Input

    Agent + bearer token

  2. 02Code

    MCP transport

    Streamable HTTP

  3. 03Gate

    Signature, expiry, audience

    JWKS · another service's token is refused before any tool runs

  4. 04Gate

    Effective scope

    leaf ∩ … ∩ root of the delegation chain

  5. 05Gate

    Approval lookup

    server-side; no tool argument can assert one

  6. 06Store

    Consume once

    conditional UPDATE + UNIQUE

  7. 07Effect

    Irreversible effect

Authority is computed by the server, never accepted from the caller.

Authorisation decisions are audited, including tokens refused at the transport before any tool is routed. Approvals are granted over a separate endpoint that requires an approver credential the agent cannot obtain — a shared credential, not a person's identity.

  • InputArrives from outside the system
  • CodeDeterministic code
  • GateDecides whether work proceeds
  • StoreDurable state
  • EffectAn irreversible or outbound effect

Engineering notes

Three checks a signature check cannot do

Audience: a token the agent legitimately holds for another service is refused before any tool runs. Attenuation: the effective scope is the intersection of every link in the delegation chain, so a scope missing from a middle link is missing from the result. Approval: the irreversible tool's input schema is exactly {account, amount} — there is no field a caller can use to assert approval — and the server finds the approval itself.

One approval authorises at most one effect because of two things enforced by PostgreSQL rather than by Python: a conditional UPDATE that only one statement can match, and a UNIQUE constraint on the effect's approval id. An in-process lock would pass a concurrency test and fail behind two workers, so there is none.

Twelve breaches planted, twelve caught

Each control was removed in turn and the suite re-run; make breaches replays all twelve against a checkout and CI runs it. The informative one: removing the application-level consume guard did not produce a wrong answer — it produced a constraint violation. The database refused the second effect on its own.

A later review attacked a running copy and found two holes: the approval-granting endpoint was unauthenticated, so an agent could create the approval it then spent, and transport-level refusals wrote no audit row. A third review found the container entrypoint called a module that never existed. All are recorded and fixed.

Deploying found what local testing could not

The MCP SDK's DNS-rebinding protection permits localhost only unless given an allowlist, so the first live instance answered 421 to every client while every local suite passed. A committed smoke run then exercised twelve scenarios against the public endpoint with a real MCP client, and the server publishes RFC 9728 protected-resource metadata.

Screens

Security matrix comparing the naive baseline and the hardened server for each scenario
The security matrix, rendered from the committed measurement run, including the totals row.Screenshot from the project repository
Delegation chain view showing scopes at each link and the computed effective scope
Delegation and attenuation: the effective scope is the intersection down the chain.Screenshot from the project repository
Approvals table: approvals bound to agent, tool, account and amount, in consumed, pending and expired states
Approvals: each binds one tool call to an account and an amount, and can be spent at most once.Screenshot from the project repository

Limitations

As the project states them. Read these before relying on any number above.

  • The approval endpoint authenticates a shared credential, not a person; approved_by is a string its holder supplies.
  • It is a resource server, not an authorisation server: a test authority mints real Ed25519 tokens and is otherwise not an IdP.
  • A stolen, still-valid token is not detected. The design limits it to the attenuated scope and stops irreversible action without approval — but a thief acting between a human approving and the agent acting spends that approval.
  • The rate limit is a ceiling, not DDoS protection, and not part of the security claim; with fixed windows the worst case across an arbitrary minute is twice the limit.
  • The demo token mint is open on non-production instances so anyone can drive the lab; an anonymous visitor holding every one of those tokens still cannot cause an irreversible effect.
  • The console renders the committed measurement, not the live broker, and the free-tier server sleeps: the first request after idle can take about a minute.
  • Not built: a Keycloak gap analysis, OpenTelemetry/Langfuse, step-up authorisation, CIMD-vs-DCR, and the full RFC 9728/8707/9207 conformance suites.
  • Everything is synthetic: invented accounts, no payment rail; the irreversible effect is a row in a demonstration table.

Facts and stack

MCP protocol
2025-11-25, Streamable HTTP
Tokens
Ed25519, verified against a JWKS
Tests
Offline and real-PostgreSQL suites · six CI jobs
Rate limit
per subject and tool, counted in PostgreSQL

Stack

  • Python
  • MCP SDK
  • FastAPI
  • PostgreSQL 16
  • SQLAlchemy
  • Alembic
  • asyncpg
  • PyJWT
  • cryptography (Ed25519, JWKS)
  • Next.js 15
  • Docker
  • GitHub Actions
  • Render
  • Neon
  • Vercel

Skills shown

  • MCP
  • AI agent security
  • Authorisation
  • JWT / JWKS
  • Approval gating
  • One-time approval consumption
  • Concurrency control
  • Rate limiting