Changes and review capsules

What a change records, how risk is scored, what a review capsule contains, and the focused diff — with the source-retention rules that govern it.

Workspace → Changes.

A change is what a session reported changing: files, symbols and dependencies, with its risk, its evidence and its review capsule.

A change is a peer of a session, not a page inside one. One session produces zero, one or several changes, and several sessions can contribute to the same pull request. Watching an agent work — why is it stuck, what was it told — and reviewing what it produced — what needs me, worst first — are two different jobs asked by two different people.

Changes are sorted by risk, not by time.

What a change holds

Paths, line counts, and the symbols touched. Not the diff body.

Holding a copy of a customer's source is a liability that buys nothing for those questions.

Field Notes
Sequence Monotonic per session, so "the third report from Agent 04" is addressable.
Base and head commit Load-bearing rather than decorative — the focused diff needs both.
Files changed, lines added, lines removed
Status in_progress, proposed, merged or discarded.
Origin reported, or retrospective.

A retrospective change is one Madebook read from a pull request that had already merged, with no session, no brief and no contract behind it. Its capsule shows surface and risk only, and says so: nothing about how the work was supervised can be claimed for work nobody supervised.

Symbols: claimed versus read

A symbol is a named thing a change touched — a class, an interface, an exported function, an endpoint, a table.

Each one carries where it came from:

Source Meaning
derived Read from the diff. Evidence.
reported_and_derived Both agree. Still evidence.
reported The agent claimed a symbol the diff does not show. Kept — it may be in a language the parser cannot read — but marked, because it is not evidence.

The interface shows this as a badge: read from the diff, or the agent says so.

Where both exist, the derived record wins on the facts — the change kind and whether the public shape moved.

contract_changed is the load-bearing flag. A body edit affects nobody; a signature change breaks everyone downstream.

Symbols are reconciled by path and name, not name alone. Keying on name alone collapsed every Create and run and index across every controller, and dropped real collisions.

Risk

A deterministic, fully attributed score from 0 to 100. Every point traces back to a factor with evidence. "Why is this risky?" is always the list of factors.

Band Score Gate it requires
low 0–24 auto merge
medium 25–59 a review capsule
high 60–84 human review
critical 85–100 human review plus a security review

What scores points

Sensitive surfaces — a deliberately narrow list, because every false positive here spends a human's attention on something safe:

What it touches Points
Authorization logic 24
Authentication or session handling 22
Credential or cryptographic handling 20
Money movement 18
Database schema change 16
Environment configuration 15
Request-level protection (CORS, CSRF, rate limiting) 14
The request pipeline (middleware, interceptors, guards) 12
File handling 10

Only the highest-scoring category counts at full points; each additional distinct category adds 35% of its own. Summing every match would make a change that touches one auth file across four folders look like a catastrophe.

Everything else

Factor Points
Public contract changes 12 + 5 each, capped at 28
Blast radius 16 above 60 files, 10 above 25, 5 above 8
Size 10 above 2,000 lines of churn, 7 above 800, 4 above 250
Architecture deviations 15 each, capped at 30
Contract violations 18 each, capped at 35
Collisions 10 each, capped at 20
Policy violations 40 for a critical one, else 15 each capped at 30
Failing evidence +25
No evidence at all +18
Evidence recorded but none passing +12

Size is a weak signal on purpose: a 2,000-line generated-client update is safer than a nine-line change to a token check.

Absence of proof raises risk. "We did not check" must cost something, or it becomes the default.

Policy gates do not inflate the score

A policy that requires a gate raises the gate directly and adds zero points. The factor is recorded so you can see it:

{policy} requires this gate for changes like this one, whatever the score.

The score is a measurement. An earlier design inflated it to force an outcome, and that taught people the number lies.

The review capsule

Computed first, prose second.

The deterministic half is built from the change surface, the decisions the agent reported, the risk assessment and its factors, the evidence and the gate status, and an inspection list of at most eight places to look.

The optional AI half is a behaviour summary, clearly labelled, and only produced when you ask for it. Firing a model on every risk calculation would spend an organization's quota on changes nobody is reviewing yet.

The capsule carries its own mission, session, person, size and risk, and it links back to the mission, the session and the pull request.

The raw diff is always one click away. Madebook prioritises review; it does not hide source.

An inspection is raised with a kind and a severity. An architecture deviation, for example, becomes a new_pattern inspection at high:

Introduces Redux, which departs from what this codebase already does. Worth a decision rather than an accident.

The focused diff

Madebook explains; the pull request proves.

Diff hunks are fetched from the provider only when a person opens one specific place to look, one path at a time, and cached against the change's base and head commit.

Three rules govern this:

  1. Diffs are never mirrored on report.
  2. A whole change is never shipped to the browser.
  3. This is not a code-review surface. GitHub already is one, and the only advantage a focused view has is that it is short.

The diff budget is per file. Each file's patch is capped and the file count is bounded, and what was not read is reported — truncated and files_omitted — and said on the session's timeline. An earlier design spent one pool in file order, so a large first file consumed it and every file after came back empty, with nothing saying why.

When there is no diff, it says which kind of nothing

Every reason is distinct, because an empty diff panel that means "we could not fetch it" and an empty diff panel that means "nothing changed" must never look the same.

  • "No path was asked for."
  • The retention message, when your organization holds no source.
  • "This change set is not attached to a repository, so there is nothing to fetch the diff from."
  • "The repository this change set names no longer exists."
  • "No base and head commit were reported for this change, and it is not linked to a pull request. Ask the session to report head_commit so the exact diff can be shown."
  • "The provider returned no changes for {path} between {base} and {head}."
  • "No diff hunks came back for {path}."
  • The connection error itself.

And when the provider cut a file short: "The provider truncated this file. What is shown is the beginning of it, not all of it."

Asking for a fix

Remediate on a change opens a request the agent will act on.

Do not interrupt a human because an AI made a mistake; interrupt a human when the AI cannot safely resolve the situation itself. A secret in a config file, a forbidden dependency, a logging line that leaks a document id — none of those need a decision.

Two required fields:

Field Hint
What is wrong In your words. It is shown to the agent and kept with the change.
What compliant looks like Write it as an instruction the agent can act on directly — name the approved pattern, not just the problem.

Plus optional files, and a severity:

Required blocks the affected work until it is fixed. Advisory is carried as a warning.

A required remediation makes the agent's next action REMEDIATE. It is delivered once, on the agent's next brief.

Refused on a change with no session behind it:

This change is not attached to a session, so there is nobody to ask. Request the change in the pull request instead.

The reply tells you whether anything will read it: "{label} will be given this on its next brief." or "{label} has ended, so nothing will read this. Start a session on the mission to have it acted on."

Reviewing

A review needs capsule.review — owner, admin or member. A security review is separately gated to owner and admin, because a policy can require one specifically.