Evidence and gates
Reported, observed and attested — three different things stored three different ways — and the gate that decides whether a change is ready.
An agent may report what it did. It may not certify it.
Reported, observed and attested are three different things, and Madebook stores them as three different things — so "the tests passed" from an agent and "the tests passed" from CI never quietly become the same sentence.
The three sources
| Source | Who | How the interface shows it |
|---|---|---|
reported |
An agent said so. | the agent says so |
observed |
Madebook or CI saw it happen. | CI saw it |
attested |
A person signed for it. | a person signed |
An agent can only ever report. Whatever an agent sends, the server rewrites the source to reported before it is stored. A session claiming "observed" or "attested" is corrected, not believed. That includes what a linked CodeBook reports about its own checks — stored as reported, labelled provisional.
Every observed and attested record is also tied to the repository, pull request, commit, issuer and attempt it came from, with a standing of current, superseded or revoked. A push supersedes results on the old commit; a re-run supersedes the earlier run; a request for changes or a dismissal withdraws an approval. The compliance check counts only current, independent evidence on the head commit.
observed comes from a check run arriving from your code host. attested comes from a person — and a pull request approval is recorded as one, because an approval is the human review.
What an evidence record holds
Kind, label, and a status of passed, failed, partial, not_run or unavailable. Plus optional counts — passed, failed, total — an artifact URL, a detail, and the reason it is unavailable.
unavailable is never left blank. If nothing is given, it records "No reason given." rather than silence.
Evidence is upserted by label, not appended — a re-run that passes replaces the failure it replaced.
The evidence gate
A gate is a set of requirements, unioned from three sources:
- The mission contract — every
proofclause. - Fired policies — each
require_*action becomes a requirement. - The risk band —
human_reviewrequires a human review;security_reviewrequires a security review and a human review.
A policy's requirement names itself: "Payments and billing requires human review."
The set is recomputed and replaced on every verification. It is never stale.
What satisfies a requirement
Every requirement must be matched by an evidence record that passed.
And for the two review kinds — human_review and security_review — only evidence whose source is attested counts. A person has to have signed.
Test evidence may be reported, and then the gate marks it provisional. A reviewer sees "the agent says the tests passed" rather than "the tests passed."
What the agent is told
When an agent calls madebook_implementation_complete, it gets back:
| Field | What it means |
|---|---|
ready_for_pull_request |
The gate is satisfied, nothing is provisional, and nothing is blocking. |
still_required |
The missing requirements, named, with where each came from. |
resting_on_your_own_word |
Requirements satisfied only by the agent's own report. |
evidence_gate |
The whole gate. |
Plus a constant note:
Implementation is your claim. The evidence gate, the pull request and a person decide whether the mission is complete.
How a check run becomes evidence
When a check run arrives, Madebook maps its name to an evidence kind. First match wins, and the list is ordered most specific first:
| Name contains | Kind |
|---|---|
| playwright, cypress, selenium, puppeteer, browser, visual, screenshot | browser |
| e2e, end-to-end, integration, acceptance, contract test | integration_test |
| codeql, snyk, trivy, semgrep, sast, dependabot, dependency review, security scan, gitleaks, trufflehog | security_scan |
| lint, eslint, rubocop, clippy, format, prettier, style | lint |
| tsc, type check, typecheck, mypy | type_check |
| unit test, jest, vitest, pytest, rspec, xunit, nunit, go test, spec | unit_test |
| test | unit_test |
| anything else | build |
build is the deliberate catch-all. No evidence requirement asks for it, so an unrecognised check can never close a gate by accident.
A guessed kind closes a gate and reads exactly like proof.
And separately:
A scanner is not a review. CodeQL passing is worth showing, and it is recorded, and
security_reviewis satisfied only by evidence a person attested.
The check's conclusion maps onto the evidence status: success is passed; failure, timed out and startup failure are failed; neutral and skipped are partial; cancelled, stale and action-required are unavailable; anything not yet completed is not_run.
Reviews are narrower still. Only approved counts — a comment is a conversation, and changes-requested is the opposite of an attestation.
Wiring it up
Until a repository is wired for CI, every test result Madebook holds is the agent's own account of its own work. Two feeds change that — a signed webhook, and an outbound pull. Both are covered in The pull request check.
In the attention queue
A mission whose contract is satisfied but whose evidence is not raises an item, but only once the mission is being verified.
| Condition | Disposition | Urgency |
|---|---|---|
| A review requirement is missing | review | 55 |
| Only test evidence is missing | inform | 35 |
The body says which: "A review requirement is satisfied only by evidence a person attests. Record yours on the mission." or "This is not blocked on you — it is blocked on proof. It will not reach Ready to ship until something demonstrates it works."
A mission that reaches the other side — contract satisfied, gate met — raises ready_to_ship at disposition review, urgency 55. Good news can wait; a session doing invalid work cannot.
How long evidence is kept
By plan: 14 days on Free, 180 days on Pro, 1095 days on Business.