The model is still not the moat: what NovaMart reveals about tribal-knowledge extraction

Pavo and three widely used coding agents, Claude Code, Cursor, and Codex, two of them on the very same model as Pavo, got the same brief and the same simulated company. Pavo recovered 74.5% of the benchmark's tribal-knowledge claims against 62.7% for the best generalist, and six in ten claims made it into every one of its runs, where the best generalist managed one in three.

Part two of a two-part release. Part one is the benchmark itself: how the NovaMart estate was built, audited, and sealed. This post is the comparison run on it.

What tribal knowledge is, and why a book

The short version: we gave Pavo and three widely used coding agents the same brief, the same frozen company, and (for two of them) the same underlying model, and asked each to reconstruct the company's undocumented knowledge, scored against 51 audited gold claims. Pavo recovered 74.5% of them; the best generalist managed 62.7%, and only recovered the same one in three claims across its own runs where Pavo recovered six in ten. The results and what may explain the gap are below, but first, what was actually being measured.

Every company runs on knowledge nobody explicitly wrote down. The "active user" that means three different things in three dashboards, because three different teams built them in three different years. The revenue report whose exact filter list exists only in the SQL of the person who built it. The migration everyone cites as complete that, in the data, is only half finished. Ask a tenured analyst how October revenue is calculated and you will get a long, precise query, full of exclusions and joins. Ask why each exclusion and join got there, and the answer lives in exactly one place: their head. That is the distinction that matters. The how is in the artifacts; the why is in people. Tribal knowledge is the why, and when the person is gone, it survives only as residue: a commit message here, an experiment readout there, the shape of the data itself. Recovering it means finding and joining those fragments, and that reconstruction is exactly what this post measures.

Why this matters now (agents doing analyst work fail silently without the why, and attrition walks it out the door) is the opening argument of the NovaMart post. This post is about whether a system can actually capture it, and how to measure that.

So the capability we care about is specific: read a company's estate, everything it has accumulated: code, databases, query logs, dashboards, documents, and reconstruct not just what the numbers are but why they are computed that way, and hold that knowledge after the people who created it are gone. Pavo's deliverable is a book, a structured, searchable guide to a system's architecture, data, metrics, experiments, failure modes, and history, in which every factual section keeps a route back to the code, query, table, dashboard, or document that supports it. It is written for people and for agents alike: the next teammate, or the next AI system pointed at the estate, should not have to rediscover what the last one learned. The book is what a reader opens; the indexed sources and synthesized knowledge it is built from sit beneath it, queryable in their own right, and the next section walks through them.

A book, rather than a chat window or a pile of retrieved snippets, because of what has to happen to this knowledge after it is written. It has to be read by a newcomer in an order that makes sense, which means stable structure. It has to be checked, which means every assertion carries a route back to its source. It has to be corrected when the company changes, which means a diff: you cannot review, approve, or reject a change to an answer that is regenerated from scratch each time it is asked. And it has to be inherited, by the next teammate and the next agent alike. A question-answering bot gives you an answer and forgets it; a book is the artifact those answers are drawn from, and the thing that can be improved on purpose.

How Pavo extracts tribal knowledge

At a high level the pipeline has four stages, and the whole design follows from one decision: ingest each kind of artifact its own way, then combine and consolidate.

How the book is compiled

Ingest each kind of artifact its own way, then join across them.

Sources · indices · synthesis · book

01 Sources02 Typed indices03 Synthesis04 The bookRepositoriesCommits & PRsWarehouseQuery logDashboardsDocumentscode indexhistory indexschema indexquery indexdashboard indexdocument indexFindingstyped · grouped in a taxonomyhistory × schemaEntitiesone card per thing · every aliasfrom many findingsRecurring tasksthe work that repeats · intent, logicquery × dashboardThe bookevery claimcitedone consolidated,cited artifact

01 · Sources

  • Repositories
  • Commits & PRs
  • Warehouse
  • Query log
  • Dashboards
  • Documents

Heterogeneous, unstructured, each with its own shape.

02 · Typed indices

  • code index
  • history index
  • schema index
  • query index
  • dashboard index
  • document index

One purpose-built schema per source. The schema decides what is askable later.

03 · Synthesis

  • Findingshistory × schematyped · grouped in a taxonomy
  • Recurring tasksquery × dashboardthe work that repeats · intent, logic
Entitiesfrom many findingsone card per thing · every alias

Findings and recurring tasks are joined from two different indices. Entities are a level further out again — many findings about the same thing, consolidated into one card. Manufactured, not retrieved.

04 · The book

One consolidated, cited artifact. Every claim carries the evidence it was written from.

In this benchmark, the general-purpose agents improvised these joins inside each run; Pavo makes the join an explicit system stage. Findings and recurring tasks are manufactured by combining indices — the commit that explains a column, the query that explains a dashboard. Entities go a level further out: many findings about the same thing, consolidated into one card per thing.

FIG. 01 — How Pavo captures tribal knowledge: fragmented sources into typed indices, indices joined into findings and recurring tasks, findings consolidated into one entity card per thing, and all three written into one cited book. The dots are evidence in motion, watch two indices land on a finding at the same instant (a join), and findings funnel into an entity card (a consolidation); with motion off, they rest mid-route. Drawn: the six surfaces in this benchmark's estate.

Sources. Pavo connects to the surfaces where a company's knowledge actually lives: code repositories, pull requests and commit history, databases and warehouses, the historical query log, dashboards, documents, and experiment records.

Typed indices. Each kind of source gets its own schema, and the schemas are where most of the design effort goes, because the schema decides which questions are askable later. A dashboard is not stored as a blob of text: it is decomposed into who owns it, the widgets on it, the queries behind each one, when it was last changed, and how much it is actually used, so the system can later ask for the dashboards someone touched last week, or the most-used dashboard behind a given metric. Those are questions flat text search over the same dashboard cannot answer reliably (an agent can improvise those joins by hand); the schema makes them a plain lookup. Commit and pull-request history is kept as history, who changed what, when, and why, instead of being flattened into prose, so time-ordered questions stay answerable. The historical query log is kept per statement, who ran it, when, and the tables it touches, so the layer above can distill from it the recurring tasks people actually perform against the warehouse. Code, schemas, documents, and experiment records each get the same treatment, which dashboards people actually trust, which queries they reach for, the experiment that set today's threshold. The indices are built once, up front.

Inside the typed indices

Each source is decomposed into a schema built for the questions it will be asked.

Opinionated · typed · queryable

Any source

chunked and embedded

Decomposed into

  • text chunk
  • text chunk
  • text chunk
  • …

So retrieval can ask

  • text that resembles the query

One shape for every source. None of the questions below is reliably askable.

  1. Dashboards

    decomposed, not stored as text

    Decomposed into

    • owner
    • widgets
    • the query behind each widget
    • edit recency
    • usage & traffic

    So retrieval can ask

    • the dashboards this person changed recently
    • the most-used dashboards for this metric
  2. Commits & PRs

    kept as history, not flattened into prose

    Decomposed into

    • author
    • date
    • what changed
    • why

    So retrieval can ask

    • what changed here, in what order, and why
  3. Query log

    kept per statement, not as a blob

    Decomposed into

    • the statement
    • who ran it
    • when
    • the tables it touches

    So retrieval can ask

    • what people actually ask of this table
    • who last queried this column
  4. Documents

    extracted, not chunked blindly

    Decomposed into

    • doc type & owner
    • the systems and metrics it names
    • decisions & thresholds it records
    • when it was last touched

    So retrieval can ask

    • every document that defines this metric
    • the decision record behind this threshold

The schema decides what questions are askable later. Retrieval quality is bounded by ingestion design.

FIG. 02 — Inside the typed indices, for four of the sources: what each is decomposed into, and what that decomposition makes askable.

Findings. On top of the indices sits a synthesis layer, and it is deliberately opinionated: extraction is not a generic sweep for anything interesting but a set of specific finding types worth having, a definition that changed, a threshold and the decision behind it, a metric that two systems compute differently. Each finding is typed and grouped under a taxonomy, and findings from one source are joined against the others: the commit that explains a column, the query that explains a dashboard, the experiment that explains a threshold. One level further out, the many findings about the same thing, a service, a metric, a table, are consolidated into a single entity card carrying its category and every alias people use for it, so the same thing is findable under every name. This is where cross-source knowledge is manufactured, and it is the layer that most distinguishes Pavo from an agent reading raw files. Retrieval quality is bounded by this design: an agent can only ask for what the ingestion layer decided to make askable.

The book. Finally, a continuous investigative agent, the agent scaffolding, writes the book: it takes inventory of the sources, reads the original sources, uses already-indexed knowledge, and saves small cited sections as it goes into a single book.

The benchmark: NovaMart

We have had a tribal-knowledge system for a long time, and no way to show how it compares to anything. The comparison evidence lived in customer estates, real warehouses, real query logs, real commit histories, and none of it could be published. So we simulated a company called NovaMart: an e-commerce retailer whose entire history was executed rather than authored, and whose frozen estate, codebase with git history, production database, warehouse, query log, application and job logs, dashboards, ships with 51 authored gold claims of tribal knowledge, each with a recomputable evidence chain. How the world was made, what counts as a claim, and how scoring works are described in the NovaMart post. Everything below uses that benchmark.

The evaluation

The experiment

The candidates. Pavo and three widely used agentic systems, evaluated with identical access to the frozen estate and an identical brief. The three general-purpose systems ran zero-shot, out of the box.

How the benchmark is kept honest. A benchmark is only worth the controls around it, so here are ours. Start with the conflict, stated plainly: Pavo designed NovaMart and ran every system evaluated here; no independent team has yet reproduced this comparison. The freeze-digest, blind-label, and audit machinery is described in the NovaMart post; the short version is that the claims and the estate were hashed together before any system ran, and books reach the judge under sealed labels. Every book in this post was scored on the same claim-set version, v1.1, the result of one round of wording and rubric corrections after an independent audit re-derived every evidence chain, and the v1.0-to-v1.1 changelog ships with the release. And because we build one of the systems under test, the separation is mechanical rather than promised: a required gate in our continuous integration asserts that no value, identifier, threshold, or phrase from any benchmark appears in any model-facing string of our system, in every configuration of it, the build fails if one does. The system is taught method, never answers. Everything needed to check that, the estate, the claims with their evidence chains, every book, and the judge, is released together.

SystemUnderlying model
PavoFable 5, high reasoning effort
Claude CodeFable 5, high reasoning effort
CursorFable 5, high reasoning effort
CodexGPT-5.6 Sol, high reasoning effort

The prompt. Each system received the same two-part brief: a task brief (credentials for the estate, the ground rules, the output format) and a goal brief describing what the resulting knowledge document should make possible. Abridged:

TASK BRIEF (excerpt)

You are a software engineer and your task is to understand the tribal knowledge of this company from its codebase and data warehouse. Your understanding must come only from the searching and reading you do on the fly. Produce a single markdown document covering: summary, why this project, business understanding, metrics, system, data, experimentation, glossary. There is no time limit, take as much time as needed.

GOAL BRIEF (condensed)

understand how the business numbers are actually produced, end to end; understand product and customer analytics, including when a dashboard's number should not be trusted at face value; understand the recommendation system and the nightly jobs.

One clause is worth a gloss, because it sounds stricter than it is: understanding must come from this company's artifacts, not from prior knowledge of it, it is an anti-leakage rule, not a ban on how a system organizes its reading. Every system was free to index, cache, or preprocess the frozen estate however it liked, under identical credentials. Note also what the prompt does not contain: no questions, no hints about any specific claim, no pointers at any table or commit. Each system explored the company for as long as it wanted and wrote its book. To make run-to-run variance visible rather than hidden, every system was run three times under the same frozen brief, producing three books each.

Scoring

Every book is scored against the 51 gold claims by a rubric-anchored judge, one claim at a time: explicit requirements the book must state and statements it must not contradict, five independent passes per book, majority verdict. Recall is path-blind: a claim found in one query and a claim found after a thousand hops earn the same credit. The headline for each system is the mean of its three books; "solved in all three" is what it recovered in every one of its three books.

Results

74.5%
Pavo

mean of 3 books · 60.8% in all 3 · 88.2% in any

62.7%
Claude Code

mean of 3 books · 35.3% in all 3 · 84.3% in any

58.8%
Cursor

mean of 3 books · 35.3% in all 3 · 80.4% in any

33.3%
Codex

mean of 3 books · 19.6% in all 3 · 51.0% in any

Gold v1.1, 51 claims. Majority verdict over 5 judge passes per book; at 51 claims, one claim is worth just under 2 points of recall; read differences against the across-claim dispersion figures in Table 1, not against the means alone.

Claim recall on the 51-claim gold set

Four systems, three books each: the mean, and how much it moves between runs.

SystemsPavoClaude CodeCursorCodex
How to readtick = mean of the 3 books, in that row’s own inkone marker = one booksame system order in every group

Claim recall · share of this group’s gold claims

0–100%

All 51 claims

n = 51 claims

Pavo

74.5%

Claude Code

62.7%

Cursor

58.8%

Codex

33.3%

Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.

FIG. 03 — Claim recall on all 51 gold claims: one marker per book, tick at the mean of the three. Every recall figure below uses this same 0–100% scale and the same system order.

Pavo, Claude Code, and Cursor ran the same model at the same reasoning effort (for Codex we used the best available model, GPT-5.6 Sol), so the spread among them is not a model comparison, it is a comparison of what each system does around the model. The model is not the differentiator here. The system around it is.

Breakdowns

Two views of the same result, sliced by what a claim requires. Read the tick as the mean and the markers as the three books. The plots show spread; the across-claim dispersion figures that qualify each comparison are in the next section.

What a claim needs. Where its evidence lives, and what history it requires, evidence that no longer exists in the current state.

Recall by where the claim's evidence lives

Codebase-only claims are where the gap is widest.

SystemsPavoClaude CodeCursorCodex
How to readtick = mean of the 3 books, in that row’s own inkone marker = one booksame system order in every group

Claim recall · share of this group’s gold claims

0–100%

Codebase only

n = 16 claims

Pavo

72.9%

Claude Code

47.9%

Cursor

43.8%

Codex

33.3%

Warehouse only

n = 15 claims

Pavo

68.9%

Claude Code

66.7%

Cursor

66.7%

Codex

26.7%

Codebase + warehouse (joint)

n = 20 claims

Pavo

80.0%

Claude Code

71.7%

Cursor

65.0%

Codex

38.3%

Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.

FIG. 04 — Recall by where the claim’s evidence lives.

Why is the widest gap on code, the one surface every coding agent is built for? The claim-level data suggests an answer. Ranked by how often the generalists miss them, six of the seven hardest codebase-only claims are temporal-era or absence claims: a formula with two constant eras, a fee rule whose real cutover was a deploy rather than a commit, a backfill that never fills a column. Five of the seven need commit or log history to settle. They are history-and-absence questions wearing a code costume, which is what a commit-and-pull-request index is built to keep, and what the generalists here most often failed to dig out. This is a post-hoc reading, like the slices below, not an ablation.

Recall on claims that require history

History that no longer exists in the current state, and who can still recover it.

SystemsPavoClaude CodeCursorCodex
How to readtick = mean of the 3 books, in that row’s own inkone marker = one booksame system order in every group

Claim recall · share of this group’s gold claims

0–100%

Requires commit history

n = 14 claims

Pavo

69.0%

Claude Code

57.1%

Cursor

59.5%

Codex

21.4%

Requires query history

n = 6 claims

Pavo

88.9%

Claude Code

88.9%

Cursor

61.1%

Codex

16.7%

Requires runtime logs

n = 9 claims

Pavo

66.7%

Claude Code

66.7%

Cursor

48.1%

Codex

3.7%

Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.

FIG. 05 — Recall on claims that require history no longer present in the current state.

The reasoning a claim demands.

Recall by the reasoning a claim demands

Mechanism and temporal-era claims separate the systems most.

SystemsPavoClaude CodeCursorCodex
How to readtick = mean of the 3 books, in that row’s own inkone marker = one booksame system order in every group

Claim recall · share of this group’s gold claims

0–100%

Mechanism

n = 6 claims

Pavo

88.9%

Claude Code

61.1%

Cursor

55.6%

Codex

27.8%

Temporal era

n = 10 claims

Pavo

70.0%

Claude Code

53.3%

Cursor

43.3%

Codex

23.3%

Absence

n = 9 claims

Pavo

88.9%

Claude Code

74.1%

Cursor

70.4%

Codex

48.1%

Reconciliation

n = 9 claims

Pavo

66.7%

Claude Code

59.3%

Cursor

59.3%

Codex

40.7%

Computed quantity

n = 12 claims

Pavo

55.6%

Claude Code

55.6%

Cursor

50.0%

Codex

13.9%

Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.

FIG. 06 — Recall by the reasoning a claim demands. Group sizes run from 6 to 12 claims (n is printed on the figure); in the smallest group a single claim is worth almost 17 points, so read narrow gaps as direction, not margin. The 3-claim multi-source-join slice is too small to plot here and appears in the NovaMart post's figure.

Where Pavo is better, and by how much

How large is Pavo's 74.5% against Claude Code's 62.7%, and how repeatable? The means alone cannot tell you, so we check two more things: what each system recovered in every single run, and how the claim-by-claim differences spread. A tribal-knowledge book is ingested and written once and then relied on for months, so what matters is less what a system can find on a good day than what it found in every one of its runs here, and on that measure the systems separate most. Look at the solved in all 3 books figures: Pavo recovers six in ten claims in every one of its books; Claude Code and Cursor recover roughly one in three in every book. Pavo's weakest book beats Cursor's and Codex's best. Eight claims are solved by every Pavo book and by no rival in all of its books; three more are solved by at least one Pavo book and by no rival book at all. The reverse ledger, so it is not hidden: three claims were found by some rival run and by no Pavo run, and on three claims a rival was consistent where Pavo was not. This pattern is consistent with what the index-plus-scaffolding design should produce: an index returns the same evidence on every run, where unguided exploration finds different things each time, and the agent consolidates across sources and fills in what the index does not find. The benchmark compares whole systems, it does not ablate the index out of Pavo, so read this as the mechanism the results point to, not one they isolate.

The table below puts a dispersion figure on each comparison. Every cell is the mean over three books. Δ is Pavo's mean minus the best rival's mean, and beside it is the across-claim paired SE of that difference [2], a dispersion summary of how the per-claim differences vary over the 51 curated claims. It is not a confidence interval: it excludes run and residual judge uncertainty, and supports no inference beyond this claim set.

TABLE 1 — Every axis, with its across-claim dispersion

AxisPavoClaude CodeCursorCodexΔ vs best rivalPaired SE
Overall74.5%62.7%58.8%33.3%+11.8±3.8
Codebase-only claims72.9%47.9%43.8%33.3%+25.0±6.5
Codebase + warehouse (joint)80.0%71.7%65.0%38.3%+8.3±4.1
Mechanism88.9%61.1%55.6%27.8%+27.8±10.2
Temporal era70.0%53.3%43.3%23.3%+16.7±9.0
Absence88.9%74.1%70.4%48.1%+14.8±8.1
Requires commit history69.0%57.1%59.5%21.4%+9.5±11.3
ML & recommendations66.7%60.0%53.3%35.6%+6.7±5.8
Warehouse-only claims68.9%66.7%66.7%26.7%+2.2±8.9
Reconciliation66.7%59.3%59.3%40.7%+7.4±10.8
Requires query history88.9%88.9%61.1%16.7%0.0±8.6
Requires runtime logs66.7%66.7%48.1%3.7%0.0±7.9
Computed quantity55.6%55.6%50.0%13.9%0.0±8.2
Solved in all 3 books (consistency)60.8%35.3%35.3%19.6%+25.5±7.3
Bold marks the two rows the argument rests on: the overall gap, and the consistency row, the share of claims a system recovered in all three of its books, where Pavo lands six in ten against one in three for the best generalist. Means of 3 books per system; ± figures are one across-claim paired SE, a dispersion summary, not a confidence interval. The SE is computed on Pavo minus the best rival over the claims in that axis; where two rivals tie, it is paired against Claude Code.

Three sources of variance, and what we did about each

Three things move these numbers, and each gets a different treatment.

  • The run. Every candidate has an LLM in the loop, Pavo included, so the same system given the same frozen brief writes a different book each time. We ran three books per system: the headline is the mean of three, the plots show all three, and the consistency row is the summary of how much they move. Three runs make the spread visible; they do not remove it. The solved-in-all-three figure is descriptive, what a system actually recovered in every one of its three books, not a separately identified stability parameter: claims differ in difficulty, and with three runs there is no clean way to split repeatability from a system simply being ahead on the easier claims.
  • The judge. An LLM judge can flip its verdict on the same claim in the same book. Every claim in every book was judged five times independently; the verdict is the majority. Majority voting damps the judge's instability; the residual judge error is not separately propagated in the ± figures, and because each system wrote a different book, that residue is only partly shared across systems.
  • The claims. The 51 claims were curated, not randomly sampled, so no ± figure here is a population interval: the across-claim paired SE summarizes dispersion within this benchmark, and says nothing about other companies' estates.

That is why the paired SE sits next to the Δ it qualifies, and not on the plots.

We went one step further and measured the measurement itself, on a second simulated company: Fernbrook Goods, a four-year, multi-site commerce estate built independently of NovaMart with a deliberately harsher scoring standard, the full dataset story (how it was rendered, how the knowledge-rot was engineered, how the answers were sealed before any system ran) is in the NovaMart post. We ran over fifty scored books there under pre-registered expectations and logged every run, including the ones that went against us, in a ledger that ships with the release. Three findings calibrate everything above.

Judge noise is real. A single-pass re-judge of a byte-identical book moved its score 4 claims in 60, which is why five-pass majority judging is the protocol here. Run variance has a floor. Most configurations we measured, our own included, move about 6 claims in 60 run to run, so on corpora like ours any single-run difference under that floor is within the movement two identical runs already show. The stability exception, and the honest caveat that rides with it, are two paragraphs down.

On ceilings, Fernbrook is unambiguous: the best book any system has produced on it is Pavo's. Under the original 60-claim protocol that book banks 47 of 60 against a best generalist book of 37, a ten-claim best-to-best gap, larger than either the judge's re-judge band (±4 in 60) or the run-to-run floor alone, though a reader stacking both uncertainties should treat it as strong, not settled. Re-judged under this post's own protocol, with the claims decomposed into 734 atomic assertions and five judge passes, the ordering holds: Pavo's best book carries 88.4% of those assertions, the best generalist book 85.3%. What Fernbrook does not show is a population gap, across whole configurations the systems land within each other's observed run-to-run spread, which is what a corpus does when it hands every contestant a pre-exported estate and removes the reconstruction work. We report that plainly. And the uncomfortable part, said just as plainly: Pavo's two ceiling books, the 47, and a 38 from a second Pavo configuration, both come from higher-variance configurations than the stable line below; a high ceiling and a tight band have not yet come from the same configuration.

The variance finding is the one that generalizes, and it runs the other way. Six runs across two Pavo configurations at one code version held a four-point band on the 60-claim scale; the leading generalist's three clean runs swung eleven. A knowledge base that is a different document every time it is built is a different problem from one that is merely imperfect, and it is the problem the rest of this post is about.

What may explain the gap: the mechanism behind each delta

Three things separate Pavo from the other agents, and each one is checkable below.

  1. Pavo reads each source the way its owner would. A query log is not three million SQL strings, it is a record of the tasks people were doing. A commit history is not diffs, it is who changed what and why. A BI tool is not a pile of dashboards, it is a record of which ones people actually trust. Each index keeps that deeper reading, and the codebase-only and mechanism rows in Table 2 are where it shows up in the data.
  2. Most tribal knowledge was never written down, and you cannot index what does not exist. None of the 51 claims is a plain single-artifact lookup: twenty are recoverable from what the estate narrates, current code, commit history, repo docs, and the other thirty-one require computing over the warehouse, the logs, or the query history. Pavo's approach is to join sources and to investigate, deriving the facts no artifact states. Table 2 is candid about where that pays so far: the leads concentrate where deep reading and cross-source joining do the work, while the slices that demand fresh computation are level, the knowledge nobody wrote down is the hard part for every system here, Pavo included.
  3. The book is a product feature that keeps improving. This benchmark scored a single frozen pass for every system, Pavo included, the only fair comparison, so the numbers here are first-pass numbers. In production, Pavo re-investigates the gaps the book knows about and revises what changed; the last section walks that loop.

How do the first two claims show up in the numbers? The machinery from the extraction section splits into two parts. The scaffolding sets the floor: even on surfaces where Pavo has no index to lean on, it stays level with the best rival. The indices track the separation row by row: where a purpose-built index exists and the findings layer can join across it, Pavo leads; where the index only retrieves, it is level; where there is no index, it plays like the others. That is a post-hoc reading across thirteen overlapping slices, exploratory, not preregistered, and unadjusted for the multiple comparisons it invites, and the tables below walk every row through it, the leads and the non-leads alike, so the claims above can be checked rather than believed.

TABLE 2 — Where Pavo's largest leads coincide with the deep-reading slices

ClaimsIndex at workΔ vs best rivalWhy
Codebase onlycode index+25.0 ± 6.5code is ingested into an index before writing begins; Pavo's weakest book matches the best rival book, and its other two are above anything a rival produced
Codebase + warehouse (joint)database index + findings layer+8.3 ± 4.1findings from one source are joined against the other before any question is asked
Mechanismcode index+27.8 ± 10.2how a value is produced today is read off current code; none of these claims needs history
Temporal eradata-task index + commit & PR index+16.7 ± 9.0what changed and when lives in the commits, PRs, and the query log, kept as history rather than skimmed as text; 9 of the 10 claims need commit or query history
Absencecode index + commit & PR index+14.8 ± 8.15 of the 9 need history; the other 4 are dead paths visible in the current code, like a status no row carries or a table nothing writes
Requires commit historycommit & PR index+9.5 ± 11.3same mechanism; 14 claims, so the across-claim spread is wide and the lead sits inside it
ML & recommendationscode index+6.7 ± 5.8model and training code is covered by the code index, and the scaffolding asks for the ML system explicitly; 15 claims, a modest lead about the size of the across-claim spread
Solved in all 3 booksall indices + the scaffolding+25.5 ± 7.3joined findings recurred across the three books; unguided exploration varied
The column names only what differs between rows; the agent scaffolding sits under every one of them.

TABLE 3 — Slices where the observed differences are small or zero relative to the across-claim spread

ClaimsIndex at workΔ vs best rivalWhy
Warehouse onlydatabase index+2.2 ± 8.9Pavo leads on the mean, but the difference is a quarter of the across-claim spread, too small to read as a lead; Claude Code's best single book here (86.7%) is the best book any system produced
Requires query historydata-task index0.0 ± 8.6exact tie, Claude Code reads what people actually queried as well as Pavo does
Reconciliationdatabase index + findings layer+7.4 ± 10.8the boundary case for the pattern: reconciliation usually means computing a quantity neither surface stores, indices retrieve, they do not compute, so the lead stays inside the across-claim spread (Pavo flat at 66.7% in all three books)

TABLE 4 — With no index to lean on, Pavo still does not fall behind

ClaimsIndex at workΔ vs best rivalWhy
Requires runtime logsnone, scaffolding only0.0 ± 7.9the one surface without an index; Pavo relies on the scaffolding and reads raw logs like a generalist, 66.7% here against 76.2% on every claim that does not need logs
Computed quantitynone, scaffolding only; indices retrieve, they do not compute0.0 ± 8.2Pavo's least stable shape (75.0 / 41.7 / 50.0 across books); two of the three claims no system solved live here

The book that maintains itself

Everything above measures day zero: point a system at a frozen estate, get a book. But tribal knowledge is not a snapshot problem. Definitions drift, sources get added, people push back on claims, and a knowledge system's real test is what happens on day thirty. Here the comparison stops being a score gap and becomes a category gap.

A generalist agent can of course be handed the previous book and told to revise it, what it lacks is not the ability to edit but the accountability machinery around the edit, which the rest of this section walks through. Without that machinery, the practical move is a rerun, which on our benchmarks produces a different book each time, an 11-point swing between the best and worst run of one configuration. Unioning those runs does not fix it: where two books disagree, deciding which one is right is precisely the synthesis work being measured, and a union has no verdict where its inputs disagree, and no single accountable artifact.

The difference starts with the shape of the deliverable. Every Pavo book is organized by the same editorial contract, why the business is shaped the way it is, how the metrics are defined and governed, how the systems actually behave, what was tried and what it taught, and the vocabulary a newcomer needs, because that is the shape in which organizational knowledge gets used, not the shape in which it happened to be found. And the brief cannot be the explanation: it names the same areas for every system, yet no two generalist books share a structure, each mirrors its own run's exploration order, while Pavo's chapter skeleton is identical section for section across every run. The fixed structure is not cosmetic. It is what makes a book reviewable section by section, diffable edition to edition, and revisable in place, none of which is even defined for a document whose organization changes every time it is written.

Pavo's book is an edition in a lifecycle, and every mechanism below ships in the product today. Adding a data source automatically schedules a targeted revision of the affected system's book: the revision lands as a reviewable draft edition, with the reason recorded, and connecting the source never blocks on it. The writer does not start over, it receives the current book as ground state plus the delta of newly connected sources, with instructions to ground new material in them and cite them where used. New context files fold in through the same path, as do claim verdicts and direct feedback: every revision traces to a concrete event, never a blind rebuild.

Every revision is accountable. The draft opens with a writer-authored "what changed in this revision," and every promotion appends a changelog record carrying the reason and the delta that actually landed, computed from the books themselves, not from prose. Dropping a previously grounded claim without an explicit human verdict is automatically flagged: knowledge does not silently disappear between editions.

Audits compound the same way, because the book grades itself. Every cited block yields a checkable claim, in effect, the book writes its own exam, and an auditor scores each claim against the estate, persisting the verdicts on the edition. Failed and stale claims become knowledge gaps; gaps become targeted revisions; unchanged sections carry their verdicts forward, and a staleness ratchet forces periodic full re-audits so carried verdicts cannot rot. A human sits over all of it in a workflow any engineer will recognize: a diff of the draft against the current edition, per-section approve, comment, or discard, partial approvals recomputed honestly, and a book assembled from partially accepted changes rendering unverified until re-attested. It is a pull-request workflow for what your company knows. The result is the property the benchmark's consistency rows only hint at: the generalists, as run here, started every book from scratch, and Pavo's knowledge accumulates, reviewed, cited, and explained, edition over edition, with every gap it finds becoming the next revision's work.

The book that maintains itself

Every turn is a targeted revision, never a rebuild.

Six stations · one closed wheel

01

An event lands

A data source is connected, a file is added, a reviewer leaves feedback, or claim verdicts come in. Every revision traces to a concrete event.

02

A targeted revision is scheduled

Not a rebuild. The writer receives the current book as ground state plus only the delta, and the event that triggered it never blocks on the revision.

03

The draft opens with a changelog

Writer-authored: what changed in this revision and why. The delta that lands is computed from the books themselves, not from the prose about them.

04

A human reviews it like a pull request

Per section: accept, comment, or discard. Partial approvals are recomputed honestly, and a grounded claim dropped without a verdict is flagged automatically.

05

Publish persists the verdicts

Verification verdicts are stored per edition and carry forward to unchanged sections — with a staleness ratchet forcing periodic full re-audits so carried verdicts cannot rot.

06

The audit's gaps become the next work

Failed and stale claims become knowledge gaps; gaps schedule the next targeted revision. The wheel does not have to wait for an outside event to turn.

back to 01

The axle

Each turn leaves the book more grounded, not merely different.

What edition n hands to edition n+1

Ground state
The current book, handed to the writer instead of a blank page.
Verdicts
Verification results on every section the revision did not touch.
Judgments
The accepts, comments, and drops a human already made, and the citations each claim earned.

11 of 60

best vs. worst rerun · companion corpus

A rerun carries none of it. Rebuilding from scratch swings that far between runs of a single configuration: a different book each time, not a better one, and no single accountable artifact to diff.

The wheel closes on itself: the gaps an audit finds are what schedule the next revision, so the book keeps improving between events rather than only in response to them.

FIG. 07 — The revision wheel: every turn is a targeted revision, and the audit's own gaps schedule the next one. The dot makes one revolution per cycle; the closing station lights as it passes, the audit's gaps becoming the next revision's work.

This is not a roadmap slide. We ran the loop on one of the benchmark books from this study: a reviewer left three anchored comments and one piece of general feedback ("deepen the reasoning wherever a decision is narrated") on a 253-block Fernbrook book. Twenty minutes later a draft edition came back with 11 sections changed, nothing added or removed, and a changelog that opened: "This revision is a targeted edit, not a rebuild. Everything not listed below is unchanged from edition 1 and still stands." Each anchored comment was answered with the missing reasoning; figures that cross a definitional boundary now say so in the same sentence; and the draft closed by flagging one issue it deliberately did not change, "for a reviewer to adjudicate", instead of silently editing it. The generalist equivalent of that twenty minutes is a from-scratch rerun and a diff nobody wrote.

And because a book nobody can reach is tribal knowledge with extra steps: Pavo connects with read-only credentials to the systems a team already runs, warehouses (BigQuery, Snowflake, Databricks), dbt, code hosting, dashboards, documents, experiment trackers, and serves the book to both audiences it is written for. People read and review it; agents query it, the book, its per-claim citations, and an ask-the-book interface, so the next AI system pointed at your estate starts from the accumulated, reviewed knowledge instead of from zero.

What this result does not prove

  • It does not rank models. Claude Code and Cursor ran the same model as Pavo, and that is the point. Each row is a whole system, scaffolding and model together.
  • Prompt sensitivity. We ran every experiment under a single prompt setting. We reduced variance at three levels, run, judge, claims, but agentic systems are sensitive to small changes in their prompts, and we did not measure that. We expect Pavo to be the least affected, because its index construction and synthesis are deterministic, but we have not shown it.
  • Recall, not precision. These tables count gold claims recovered. They do not score unsupported claims, contradictions, or citation validity, a book that banks more claims while asserting more confident errors would look better here than it deserves. The judge does check each book's evidence coverage, and those per-book scorecards ship with the release; until then, recall is the measured half of the trust story.
  • Corpus dependence. On our second, harder corpus (Fernbrook, see the variance section) the gap compresses to a tie within the observed run-to-run spread at configuration level, because that corpus hands every system its history pre-exported and scores multi-clause claims strictly. The pattern, not the level, is what we defend: the deltas appear where knowledge must be reconstructed through indices, history, and joins, and compress where the estate is handed over flat.
  • Cost. The benchmark does not price each system's book, and we owe the full table (tokens, dollars, wall-clock per system per run) with the artifact release. The magnitudes we can state now: a generalist writes its book in roughly half an hour of wall-clock; Pavo's initial run takes a few hours because indexing and synthesis run up front, a cost paid once per estate, not per question, and a targeted revision of a finished book takes about twenty minutes. If the cost table changes anyone's read of the results, that is exactly why it ships with the data.

Check our work

Every number in this post is re-derivable, and until the artifacts are live you should treat these numbers as our claim, not the field's. The release, the NovaMart estate, the gold claims with their recomputable evidence chains, all twelve books with their per-pass scorecards, and the judge harness, ships together, so the comparison can be rerun end to end rather than taken on trust [1]. If you build agents, including the ones in the plots above, we would genuinely like to see your numbers when it lands. The benchmark does not care whose logo is on the book.

References

  1. Pavo, NovaMart: we simulated a company, to measure what agents actually know. 2026. pavoai.com/blog/novamart-benchmark
  2. Miller, E., Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. Anthropic, 2024. arxiv.org/abs/2411.00640. The paired-difference computation follows this paper; we read the result as dispersion over the fixed, curated claim set, not as the paper's population interval.
  3. Anthropic, Claude Code. docs.claude.com/en/docs/claude-code
  4. Anysphere, Cursor. cursor.com
  5. OpenAI, Codex. openai.com/codex
  6. Anthropic, Claude Fable 5 and Claude Mythos 5. 2026. anthropic.com/news/claude-fable-5-mythos-5
  7. OpenAI, GPT-5.6 Sol. 2026. openai.com