Codebase intelligence

The five tabs on a repository — analysis, architecture, files, dependencies and findings — what each is derived from, and how much to trust it.

Open a repository from Workspace → Repositories and you get five tabs.

The tabs are ordered by how much they can be trusted, which is also the order they become available: the file index and the dependency map are facts, the rule-based findings are reproducible, and the written analysis is a model's interpretation.

Above them sit five figures: Files indexed (with the vendored count excluded), Languages, Dependencies (with how many are outdated), Open findings, and Test files per source — which is labelled "a file-count ratio, not coverage", because coverage is a measurement from a test run and this is not one.

If the written analysis describes an older commit, a callout says which two commits and offers a re-index.

Analysis

The written interpretation. It carries the model's name, or a mock badge, a confidence figure, when it was produced, and a link reading "See what it was given" that opens the AI run with the full manifest of everything that went into the prompt.

A callout headed "What it could not see" carries the model's own confidence note.

Five sections:

  • If you are joining this repository — the first thing to read
  • What this application does
  • How data moves through it
  • External services it depends on
  • Where you are most likely to break something

Caps on what it will produce: 4,000 characters each for the summary and the onboarding brief, 2,000 for the data flow, 1,500 for external services, 2,000 for danger zones.

Without an analysis, the tab says exactly why:

Index status The reason shown
Never indexed "This repository has not been indexed yet."
Queued "Indexing is queued and will start shortly."
Indexing "Indexing is in progress. The file index, dependencies and rule-based findings appear first; the written analysis follows."
Indexed "The repository is indexed — its files, dependencies and rule-based findings are available. The written analysis has not been produced yet, usually because no AI provider is configured or the organization has not approved an AI provider."
Failed The stored error.

Architecture

A map of components and how they connect. Between 1 and 25 nodes, and at most 50 edges.

Node kinds: frontend, backend, api, database, queue, cache, auth, external_service, job, storage, gateway. Edge kinds: http, grpc, sql, queue, event, file, auth, depends_on.

Every node carries the file path it was inferred from. That is a design rule, not a nicety: a model that must name the file it inferred a component from invents far fewer components.

Three guard rails run over whatever the model produced, before anything is stored:

  • Duplicate node keys are dropped.
  • An edge pointing at a node that does not exist is dropped.
  • A finding whose path is not in the file index has its path removed — a finding pointing at a file that does not exist is worse than one with no path.

It is drawn as grouped cards with their relationships listed, not as a force-directed graph. A graph of eight nodes looks impressive and tells you less than a list.

Produced without a model, the map carries a callout: "This map is placeholder content. It was generated without a model."

Files

A directory browser over the index.

Paths and metadata only — Madebook does not mirror source into its own database.

Directories show their file count and languages. Files show their name, an entry point badge where it applies, their role when it is not plain source, their language and their size.

Clicking a file fetches it from the host on demand and shows it — "never mirrored into the Madebook database."

Roles a file can carry: source, test, config, docs, build, generated, vendor, asset.

Dependencies

Every package from every manifest, with its declared version, its ecosystem, and a note.

Version age is checked against a curated list of well-known packages, deterministically — no model is involved. This is not a vulnerability scan, and Madebook does not maintain a CVE database.

A package two or more major versions behind gets a {n} behind badge — red at three or more, amber at two.

If nothing was parsed:

Madebook reads package.json, *.csproj, requirements.txt, go.mod, pom.xml and Gemfile. None were found, or the repository has not been indexed.

Findings

These are shapes worth a human look. The rule-based ones are reproducible; the AI ones are hypotheses. Madebook is not a security scanner and does not claim to be one.

Each finding shows a severity, a category, and a source badge reading rule-based or AI hypothesis. Findings are ordered: open before reviewed, then by severity (critical, high, medium, low, info), then rule-based before AI.

Severities are info, low, medium, high, critical. Categories are dependency, secret, security, duplication, size, testing, separation_of_concerns and documentation.

The rule-based checks, exactly

Secret patterns — all critical, all at 85% confidence, one finding per pattern per file:

Name What it looks for
Connection string with an inline password
AWS access key id AKIA or ASIA plus 16 characters
Private key block
GitHub token ghp_, gho_, ghu_, ghs_, ghr_, github_pat_
Provider API key sk-ant-, sk-proj-, sk-, xoxb-, xoxp-, SG.
Azure storage account key

Every one of these ends with the same sentence: "The value itself is not stored by Madebook."

Exempt from the scan: anything ending .example, .sample or .template, .env.example, anything under a test directory, and anything under a docs directory.

Structural checks

Check Threshold Severity
Very large file over 40,000 bytes medium, or high above 120 KB
No tests alongside a module 4 or more files in it, none of them tests medium
Very little test code ratio under 0.1 with more than 20 source files high
No documentation files at all low

Dependency checks — packages two or more majors behind become one grouped finding, high if any is three or more behind. Packages that have been superseded — moment, request, node-sass, tslint, enzyme, Newtonsoft.Json, AutoMapper — produce up to five low findings.

Acting on a finding

Two buttons on an open finding.

Track it records it as a known problem on the workspace context and marks the finding confirmed. It maps the finding's category onto a context area: dependency issues to dependencies, secrets and security to security, duplication, size and separation-of-concerns to architecture, testing to testing, documentation to process. An info severity becomes low; everything else keeps its severity.

Recorded as a known problem on the workspace

Dismiss closes it without recording anything. A finding can be open, confirmed, dismissed or fixed, each with an optional note.

Your verdict survives a re-index. Rule-based findings are re-derived every run, but a finding you have already ruled on keeps its status, its reviewer and its note.

Codebase Activity

Workspace → Codebase Activity (/system) is computed from the index's modules and the change sets that touched them — which modules are being worked on, by whom, at what risk. There is no hand-curated map of areas; a repository does not declare one, so Madebook does not invent one.