Supervision and enforcement for autonomous coding agents

Agents that don’t report don’t merge.

Madebook briefs every coding agent from what is currently true — the architecture, the decisions already made, the contract for the work — and posts one required status on every pull request. Work that went around it cannot be merged. Work that went through it arrives with a record of who checked what.

One status, madebook/supervised, required in branch protection. That is the entire enforcement: nothing to install in CI, nothing Madebook merges or blocks itself. GitHub (one-click App), GitLab and Azure DevOps; Claude Code, Codex, Cursor and Copilot through the MCP — with hooks that report every edit so the protocol never depends on the model remembering.

The enforcement

Every rule an agent is given is advisory — until the pull request cannot merge without it.

A coding agent that never calls the supervision layer is never told anything. That is true of every tool in this category. Madebook’s answer is not a bigger prompt; it is a commit status the host already knows how to require.

  1. 01

    Turn on Supervised merges for the repository.

    One switch in Madebook. Every open pull request is judged at once.

  2. 02

    Require the status in branch protection.

    One-time, by a repository admin on GitHub. From then on it is binding.

  3. 03

    Agents report through the MCP as they work.

    The moment a session claims its pull request, the light turns green. A critical control firing turns it red.

Checks · pull request #418

One required status. Nothing to install in CI. Nothing Madebook merges or blocks itself.

✓ci/build

Build succeeded.

Yours. Madebook does not replace CI; it stands next to it.

✓madebook/supervised

Supervised — Agent 02 (session #12) reported this work through Madebook. Evidence gate: 1 requirement still open — browser.

A session claimed it and no critical control fired. What is still unproven rides in the text — it does not turn the light red.

✕madebook/supervised

Not supervised — no work session reported this pull request through Madebook.

The work went around the tool. With the status required, this pull request cannot be merged.

Green means a supervised session reported this work and no critical control fired. Red means it went around the tool. Nothing else changes the colour — a status that goes red for a dozen reasons is one people learn to override.

The problem

Running several agents is a distributed-systems problem inside your team.

Not literally distributed computing, but the same shape. Each agent has its own snapshot of reality, taken at a different moment, and no mechanism reconciles them. The failures that follow are quiet, expensive, and invisible to every tool that watches git.

Four agents · one repository · one morning

3 working from a stale picture
Claude Codebriefed 10:01

“Tokens expire after 60 minutes”

current at the time
Codexbriefed 10:05

“Tokens expire after 24 hours”

superseded
Cursorbriefed 10:09

“State via Redux”

superseded
Copilotbriefed 10:12

“Secrets from .env”

superseded

10:07 — a person decides: externally emailed security tokens expire after 30 minutes. Three of these four sessions were briefed before that existed. No file overlaps. No signature moved. Nothing will conflict at merge.

No overlapping file

Different directories, different worktrees. Git has nothing to say.

No changed signature

A decision can make code wrong without changing a single line of it.

No failing test

Both branches work. Both suites pass. The architecture still forked.

What it actually looks like

Three ways this costs you a week, none of which involve a merge conflict.

01

A decision the agent never saw

Developer 1’s agent
implements password reset. Picks a 60-minute token expiry.
Developer 2’s agent
implements invitations. Picks 24 hours.
A person, 10:07
decides: externally emailed security tokens expire after 30 minutes.

Developer 2’s session was briefed at 10:05. It is now building against an answer the team has replaced. There is no overlapping file, no changed signature and nothing to conflict at merge.

What Madebook does

Decision 81 was made after this session was briefed. It affects this mission. Re-brief.

While the work is happening — not as a comment on the pull request, by which point the branch is finished and the context that produced it is gone.

02

Architecture quietly forking

Developer A’s agent
adds Redux and starts migrating state.
Developer B’s agent
keeps building with React Query, as it always has.
Both branches
work. Both test suites pass. Neither touches the other’s files.

Nothing is broken, so nothing objects — and the codebase now has two state management approaches, chosen independently by two machines that were never told the question had been settled.

What Madebook does

The architecture constitution says React Query. Stop Agent A now, not at the pull request.

While the work is happening — not as a comment on the pull request, by which point the branch is finished and the context that produced it is gone.

03

A rule the agent had no way to know

The agent
needs an environment variable, so it adds .env loading. Reasonable.
Governance
says secrets come from Key Vault. Written down, never in the prompt.
The diff
is clean, idiomatic, well-tested, and wrong for this company.

Give the agent the rule before it writes the code and the problem does not happen. Miss it, and the diff itself shows what was done — not what was claimed.

What Madebook does

Safe to remediate: tell the agent to fix it. Needs a person: stop that piece of work and ask.

While the work is happening — not as a comment on the pull request, by which point the branch is finished and the context that produced it is gone.

How it works

Before the work, during the work, and after it.

Before

Every agent starts from what is currently true

Architecture, decisions already made, standards, security rules, the mission contract and the current state of the repository — assembled into the brief before the agent writes anything. Constraint applied at the start costs nothing. Applied at the pull request, it costs the whole branch.

A precision forming jig shaping material as it passes through, rather than inspecting it afterwards.
During

When one agent invalidates another, the other one is told

The detection is a join, not a judgement: what a session moved, against what another session is standing on. Deterministic, explainable, and cheap enough to run on every report. When something a session depends on moves, that session is re-briefed — or held — while there is still time for it to matter.

Four offset translucent planes, the last one drifted out of alignment and lit amber.
After

What actually happened, separately from what was claimed

An agent may report what it did. It may not certify it. Reported, observed and attested are three different things and are stored as three different things — so “the tests passed” from an agent and “the tests passed” from CI never quietly become the same sentence.

A machined part measured in a calibration instrument beside a ghosted amber copy that does not match.

For the engineer running the agent

Your agent stops re-deriving the codebase every session.

The first twenty minutes of most agent sessions are spent rediscovering what the team already knows: which state library, which authorization pattern, what was decided last month and why. Madebook compiles that into the brief before the first token, from a confirmed record rather than from whatever the codebase happened to do most often.

It is the same brief the supervision runs on — so the agent that is faster for having it is also the agent whose pull request goes green. Seven tools, not eighteen; nothing to poll; a person’s answer to a question arrives on the next re-brief with the instruction attached.

Brief · Agent 02 · claims-search

Compiled before the first token. Frozen as a snapshot, with a manifest of what went in and why.

~7,500 tokens
  1. 01

    Role and policies in force

    What the organization allows: providers, clients, secrets handling, the review gates.

  2. 02

    Architecture rules

    Derived from the indexed code — state management, authorization pattern, data access — with do-not-use lists. Confirmed by a person before an agent is held to them.

  3. 03

    Decisions already made

    The ledger, scoped to the paths this session is touching. "Externally emailed tokens expire after 30 minutes" reaches the agent before it picks 60.

  4. 04

    The mission contract

    What correct means: success, preserve, forbidden, security, proof. Machine-checked clauses are marked as such.

  5. 05

    Project context

    What the system is, the constraints only a human knows, the standards — confirmed, never inferred and quietly enforced.

  6. 06

    Other live sessions

    The symbols every other agent on the project is moving right now, so a signature change is reported before it becomes somebody else’s surprise.

Seven tools, one per thing an agent does in a session: open, re-brief, plan, report, ask, finished, close. The re-brief is also how a person’s answer reaches a working agent — once, with the instruction attached. Nothing polls.

In the product

One queue of things that need a person, ordered by what it costs to ignore them.

Not a feed. Everything here is computed from something recorded — a symbol that moved, a clause that was broken, a session briefed before a decision existed.

Attention

What requires your brain right now — computed, ordered, and nothing else.

5 intervene

Agent 02 changed ClaimsService.GetDocument; Agent 07 is building against it

The shape of ClaimsService.GetDocument changed — it now returns ClaimDocumentResult instead of Stream, and takes a CancellationToken. The other session was given this symbol when it was briefed.

Agent 07 has been working for 41 minutes and does not know this changed.

Re-brief the affected agentLet it continuecross_agent_collision

M-002 breaks 2 clauses of its own contract

Must preserve: “ClaimsService.GetDocument keeps its current signature.” Forbidden: “No new state management library” — @reduxjs/toolkit was added to the dependency manifest.

See the contractAmend the contractcontract_violation

And a record that distinguishes what was claimed from what was seen.

Change record · M-002

Counted by who saw it. Never “ready to merge”.

Tests passclaimed

reported by the agent

a claim, recorded as a claim

Tests passobserved

observed in CI

independently seen

Build succeedsobserved

observed in CI

independently seen

Contract clauses heldobserved

derived from the diff

independently seen

Browser verificationnot run

nobody

no evidence either way

Security reviewnot run

nobody

no evidence either way

The agent said its tests passed. CI says they did too — and those are two different facts, stored separately. Where nobody looked, the record says nobody looked.

What it refuses to do

A supervision layer is only worth having if it is honest when it does not know.

Every one of these makes Madebook look worse in a demo and makes it worth trusting on a Tuesday afternoon six months in.

It never says “ready to merge.”

It says who checked what, and what nobody checked.

Silence is never green.

A check that did not run reads as absent, not as passing.

The diff outranks the agent.

Where the report and the diff disagree, the diff wins.

Nothing is invented.

A field with no evidence behind it stays empty rather than plausible.

If your team runs more than one agent at a time, this is already happening.

Madebook installs above whatever your engineers already use — Claude Code, Codex, Cursor, Copilot — through the MCP. The agents keep their tools. The team gets one place where what is currently true is written down.

Request access

Tell us how many agents run on your codebase in a normal week and what broke last time two of them disagreed.

[email protected]

Telemetry is metadata only by default — counts, names and symbols. Not your prompts, and not your source.