Operational readiness
For whoever runs a Madebook deployment — the /ready probe and what it checks, the Admin Readiness page, the diagnostics token, backup and restore, the restore drill, and the manual steps an operator has to take.
This page is for the operator of a hosted Madebook, not for the organizations using it. The full runbook — alert thresholds, retention, the step-by-step restore — ships with the application as docs/OPERATIONS.md. This is the shape of it.
Two probes
| Question | Touches | Answers | |
|---|---|---|---|
GET /api/health |
Is the process up? | nothing | 200 while the process runs |
GET /api/ready |
Can this instance do its job? | every database, migrations, the job queue, Bookbag SSO, the last deploy, the audit trail, audit streams, enforcement, linked products and the app registry | 200 {"ready":true} or 503 {"ready":false} — nothing else to an anonymous caller |
Point container liveness at /health (restarting a healthy process because its database blinked makes an outage worse) and traffic routing at /ready.
Readiness is evaluated at most once every 5 seconds whoever asks, and each check has a 3-second deadline: a dependency that hangs is a failure that says did not answer, not a probe that times out with no reason.
What /ready checks
| Check | Fails when | Warns when |
|---|---|---|
| each context database | it does not answer | — |
| migrations | a migration this build ships is not applied | the database has a migration the build does not (the image was rolled back) |
| jobs | runnable jobs waited more than 15 minutes, or a running job stopped heartbeating | jobs failed permanently in the last hour |
| sso | Bookbag refuses the app token, is unreachable, or publishes no usable OIDC configuration | — |
| deploy | — | the last container start recorded a failed step |
| audit | the spool cannot be read | events are waiting in the spool, or were lost since boot |
| audit_streams | — | an enabled stream failed its last delivery |
| enforcement | — | a repository set to enforce reads misconfigured, unavailable or stale |
| peer_deliveries | — | messages owed to CodeBook or OpenBook gave up, or the delivery sweep is not running |
| peer_directory | — | the sibling products are not being read from Bookbag's app registry: Bookbag is unreachable and the last answer is being served, the server runs on local overrides or on nothing, Bookbag lists this app as disabled, or the last answer is well past its 60-second freshness |
Only a fail returns 503. Warnings still count as ready — taking an instance out of rotation because one customer's audit stream is failing would turn that into "the product is down".
Reading the detail
The full evaluation — every check, its status, reason, latency and detail — goes to:
- a platform admin, at Admin → Readiness (re-read every 30 seconds, with Check again);
- a monitor holding
MADEBOOK_DIAGNOSTICS_TOKEN, as the body ofGET /readywithAuthorization: Bearer <token>.
The token is off unless set, ignored if shorter than 32 characters, compared in constant time, and never shown in the app. An API token or a session is not a diagnostics token.
Linked products and the app registry
Who CodeBook and OpenBook are — their API addresses — the secret the three products share, and Madebook's own public API address all come from Bookbag's app registry, not from the host's environment. An address is fixed at Bookbag (Platform admin → Apps), once, for every product. See Linked products.
The boot log says where the answer came from: peer directory: N peers from SSO, or peer directory: SSO unreachable, using cached (or none). If Bookbag has not answered within 5 seconds the server starts anyway and logs that no sibling call is accepted until it does.
The peer_directory check on Admin → Readiness shows the source (sso, cached, overrides or none), the siblings and their addresses, this API's own address, whether a secret rotation is inside its one-hour grace and when it happened, and any local override in effect — never a secret. Only a warning, never a failure: a Madebook with no siblings still governs its own workspaces.
| Source | What it means |
|---|---|
sso |
Read from Bookbag. What production should show. |
cached |
Bookbag is unreachable and the last good answer is being served. Fine for a while; see the sso check. |
overrides |
The host has PEER_* or MADEBOOK_PUBLIC_URL set. These are for local development and should not be set in production. |
none |
Bookbag has never answered and there are no overrides. No sibling can call in and nothing can be called back until it answers. |
MADEBOOK_PUBLIC_URL, where it is still set, is read in either form: a bare origin (https://madebook.acme.example) has /api appended; a value with a path (https://madebook.acme.example/api) is used as given. Both build the same GitHub webhook and callback URLs.
Backup and restore
npm run backupdumps each of the twelve context databases withpg_dump -Fcinto a dated directory with a checksummed manifest — per-file hashes, the migrations read back out of each archive, the git commit, and a fingerprint of the encryption key. A partial backup is marked incomplete and exits 1; restore refuses it.npm run restore <dir> --confirm <prefix>verifies before touching anything: completeness and every checksum, that the code is not older than the backup, the encryption-key fingerprint, the confirm prefix, and that no other session is connected. Then each database is restored in one transaction, and migrations the backup predates are applied forward.npm run restore:checkboots only the databases and exits 1 unless readiness passes, every organization's audit chain verifies end to end, and every stored secret decrypts with this environment's key. It lists what catches up afterwards and can queue the pull request re-syncs.npm run drill:restorerehearses the whole cycle on a throwaway PostgreSQL 16 cluster it creates and deletes: seed, seal, back up, prove seven bad restores are refused, drop all twelve databases, restore, compare every row by hash, run the check, boot the API and get200from/ready. 39 checks. Run it after any change to backup, restore or migrations, and at least quarterly.
What is not in the backup: MADEBOOK_ENCRYPTION_KEY — every stored secret and the private half of every export signing key is encrypted with it, and a restore without it restores ciphertext nothing can read. Keep the key in your secret store, separately from the backups. Organizations, people, AI configuration and file storage live at Bookbag and are backed up there.
The operator's manual steps
Nothing does these for you:
- Put the audit spool on a persistent volume. Set
MADEBOOK_AUDIT_SPOOL(defaultlog/audit-spool.jsonl) to a path that survives a redeploy, or events the database refused during an outage are lost with the container. - Set
MADEBOOK_DIAGNOSTICS_TOKEN(openssl rand -hex 32) and point the uptime monitor at/api/readywith it, so an alert says which dependency failed. - Schedule
backup.sh— nightly, plus before every deploy that ships a migration — off the database host, encrypted at rest, with an alert when the newest manifest is older than 26 hours or sayscomplete: false. - Keep the encryption key outside the backup.
- Take the
PEER_*overrides off production once Bookbag's app registry answers: removePEER_SECRETfrom all three products in one step and redeploy them together (a secret overridden on one product and read from Bookbag on another makes peer traffic fail silently), then remove the address overrides. Rotate the shared secret at Bookbag only after all three read it from there. - Record every restore in the incident log and tell affected organizations' admins: after a restore the audit ids restart below what a streaming destination already received, so the seal chain visibly forks at the restore point. That fork is the tamper evidence doing its job, and it needs an explanation on file rather than an investigation.
Things to know before relying on it
- The twelve dumps are twelve points in time, seconds apart, unless the API and worker are stopped first. For anything you intend to restore deliberately, quiesce first or take a server-level snapshot.
- The alert thresholds in the runbook are judgement calls.
- The restore drill has been run on macOS only so far.