The model is still not the moat: what NovaMart reveals about tribal-knowledge extraction
Pavo and three widely used coding agents, Claude Code, Cursor, and Codex, two of them on the very same model as Pavo, got the same brief and the same simulated company. Pavo recovered 74.5% of the benchmark's tribal-knowledge claims against 62.7% for the best generalist, and six in ten claims made it into every one of its runs, where the best generalist managed one in three.
Part two of a two-part release. Part one is the benchmark itself: how the NovaMart estate was built, audited, and sealed. This post is the comparison run on it.
What tribal knowledge is, and why a book
The short version: we gave Pavo and three widely used coding agents the same brief, the same frozen company, and (for two of them) the same underlying model, and asked each to reconstruct the company's undocumented knowledge, scored against 51 audited gold claims. Pavo recovered 74.5% of them; the best generalist managed 62.7%, and only recovered the same one in three claims across its own runs where Pavo recovered six in ten. The results and what may explain the gap are below, but first, what was actually being measured.
Every company runs on knowledge nobody explicitly wrote down. The "active user" that means three different things in three dashboards, because three different teams built them in three different years. The revenue report whose exact filter list exists only in the SQL of the person who built it. The migration everyone cites as complete that, in the data, is only half finished. Ask a tenured analyst how October revenue is calculated and you will get a long, precise query, full of exclusions and joins. Ask why each exclusion and join got there, and the answer lives in exactly one place: their head. That is the distinction that matters. The how is in the artifacts; the why is in people. Tribal knowledge is the why, and when the person is gone, it survives only as residue: a commit message here, an experiment readout there, the shape of the data itself. Recovering it means finding and joining those fragments, and that reconstruction is exactly what this post measures.
Why this matters now (agents doing analyst work fail silently without the why, and attrition walks it out the door) is the opening argument of the NovaMart post. This post is about whether a system can actually capture it, and how to measure that.
So the capability we care about is specific: read a company's estate, everything it has accumulated: code, databases, query logs, dashboards, documents, and reconstruct not just what the numbers are but why they are computed that way, and hold that knowledge after the people who created it are gone. Pavo's deliverable is a book, a structured, searchable guide to a system's architecture, data, metrics, experiments, failure modes, and history, in which every factual section keeps a route back to the code, query, table, dashboard, or document that supports it. It is written for people and for agents alike: the next teammate, or the next AI system pointed at the estate, should not have to rediscover what the last one learned. The book is what a reader opens; the indexed sources and synthesized knowledge it is built from sit beneath it, queryable in their own right, and the next section walks through them.
A book, rather than a chat window or a pile of retrieved snippets, because of what has to happen to this knowledge after it is written. It has to be read by a newcomer in an order that makes sense, which means stable structure. It has to be checked, which means every assertion carries a route back to its source. It has to be corrected when the company changes, which means a diff: you cannot review, approve, or reject a change to an answer that is regenerated from scratch each time it is asked. And it has to be inherited, by the next teammate and the next agent alike. A question-answering bot gives you an answer and forgets it; a book is the artifact those answers are drawn from, and the thing that can be improved on purpose.
How Pavo extracts tribal knowledge
At a high level the pipeline has four stages, and the whole design follows from one decision: ingest each kind of artifact its own way, then combine and consolidate.
How the book is compiled
Ingest each kind of artifact its own way, then join across them.
Sources · indices · synthesis · book
01 · Sources
- Repositories
- Commits & PRs
- Warehouse
- Query log
- Dashboards
- Documents
Heterogeneous, unstructured, each with its own shape.
02 · Typed indices
- code index
- history index
- schema index
- query index
- dashboard index
- document index
One purpose-built schema per source. The schema decides what is askable later.
03 · Synthesis
- Findingshistory × schematyped · grouped in a taxonomy
- Recurring tasksquery × dashboardthe work that repeats · intent, logic
Findings and recurring tasks are joined from two different indices. Entities are a level further out again — many findings about the same thing, consolidated into one card. Manufactured, not retrieved.
04 · The book
One consolidated, cited artifact. Every claim carries the evidence it was written from.
In this benchmark, the general-purpose agents improvised these joins inside each run; Pavo makes the join an explicit system stage. Findings and recurring tasks are manufactured by combining indices — the commit that explains a column, the query that explains a dashboard. Entities go a level further out: many findings about the same thing, consolidated into one card per thing.
Sources. Pavo connects to the surfaces where a company's knowledge actually lives: code repositories, pull requests and commit history, databases and warehouses, the historical query log, dashboards, documents, and experiment records.
Typed indices. Each kind of source gets its own schema, and the schemas are where most of the design effort goes, because the schema decides which questions are askable later. A dashboard is not stored as a blob of text: it is decomposed into who owns it, the widgets on it, the queries behind each one, when it was last changed, and how much it is actually used, so the system can later ask for the dashboards someone touched last week, or the most-used dashboard behind a given metric. Those are questions flat text search over the same dashboard cannot answer reliably (an agent can improvise those joins by hand); the schema makes them a plain lookup. Commit and pull-request history is kept as history, who changed what, when, and why, instead of being flattened into prose, so time-ordered questions stay answerable. The historical query log is kept per statement, who ran it, when, and the tables it touches, so the layer above can distill from it the recurring tasks people actually perform against the warehouse. Code, schemas, documents, and experiment records each get the same treatment, which dashboards people actually trust, which queries they reach for, the experiment that set today's threshold. The indices are built once, up front.
Inside the typed indices
Each source is decomposed into a schema built for the questions it will be asked.
Opinionated · typed · queryable
Any source
chunked and embedded
Decomposed into
- text chunk
- text chunk
- text chunk
- …
So retrieval can ask
- text that resembles the query
One shape for every source. None of the questions below is reliably askable.
Dashboards
decomposed, not stored as text
Decomposed into
- owner
- widgets
- the query behind each widget
- edit recency
- usage & traffic
So retrieval can ask
- the dashboards this person changed recently
- the most-used dashboards for this metric
Commits & PRs
kept as history, not flattened into prose
Decomposed into
- author
- date
- what changed
- why
So retrieval can ask
- what changed here, in what order, and why
Query log
kept per statement, not as a blob
Decomposed into
- the statement
- who ran it
- when
- the tables it touches
So retrieval can ask
- what people actually ask of this table
- who last queried this column
Documents
extracted, not chunked blindly
Decomposed into
- doc type & owner
- the systems and metrics it names
- decisions & thresholds it records
- when it was last touched
So retrieval can ask
- every document that defines this metric
- the decision record behind this threshold
The schema decides what questions are askable later. Retrieval quality is bounded by ingestion design.
Findings. On top of the indices sits a synthesis layer, and it is deliberately opinionated: extraction is not a generic sweep for anything interesting but a set of specific finding types worth having, a definition that changed, a threshold and the decision behind it, a metric that two systems compute differently. Each finding is typed and grouped under a taxonomy, and findings from one source are joined against the others: the commit that explains a column, the query that explains a dashboard, the experiment that explains a threshold. One level further out, the many findings about the same thing, a service, a metric, a table, are consolidated into a single entity card carrying its category and every alias people use for it, so the same thing is findable under every name. This is where cross-source knowledge is manufactured, and it is the layer that most distinguishes Pavo from an agent reading raw files. Retrieval quality is bounded by this design: an agent can only ask for what the ingestion layer decided to make askable.
The book. Finally, a continuous investigative agent, the agent scaffolding, writes the book: it takes inventory of the sources, reads the original sources, uses already-indexed knowledge, and saves small cited sections as it goes into a single book.
The benchmark: NovaMart
We have had a tribal-knowledge system for a long time, and no way to show how it compares to anything. The comparison evidence lived in customer estates, real warehouses, real query logs, real commit histories, and none of it could be published. So we simulated a company called NovaMart: an e-commerce retailer whose entire history was executed rather than authored, and whose frozen estate, codebase with git history, production database, warehouse, query log, application and job logs, dashboards, ships with 51 authored gold claims of tribal knowledge, each with a recomputable evidence chain. How the world was made, what counts as a claim, and how scoring works are described in the NovaMart post. Everything below uses that benchmark.
The evaluation
The experiment
The candidates. Pavo and three widely used agentic systems, evaluated with identical access to the frozen estate and an identical brief. The three general-purpose systems ran zero-shot, out of the box.
How the benchmark is kept honest. A benchmark is only worth the controls around it, so here are ours. Start with the conflict, stated plainly: Pavo designed NovaMart and ran every system evaluated here; no independent team has yet reproduced this comparison. The freeze-digest, blind-label, and audit machinery is described in the NovaMart post; the short version is that the claims and the estate were hashed together before any system ran, and books reach the judge under sealed labels. Every book in this post was scored on the same claim-set version, v1.1, the result of one round of wording and rubric corrections after an independent audit re-derived every evidence chain, and the v1.0-to-v1.1 changelog ships with the release. And because we build one of the systems under test, the separation is mechanical rather than promised: a required gate in our continuous integration asserts that no value, identifier, threshold, or phrase from any benchmark appears in any model-facing string of our system, in every configuration of it, the build fails if one does. The system is taught method, never answers. Everything needed to check that, the estate, the claims with their evidence chains, every book, and the judge, is released together.
| System | Underlying model |
|---|---|
| Pavo | Fable 5, high reasoning effort |
| Claude Code | Fable 5, high reasoning effort |
| Cursor | Fable 5, high reasoning effort |
| Codex | GPT-5.6 Sol, high reasoning effort |
The prompt. Each system received the same two-part brief: a task brief (credentials for the estate, the ground rules, the output format) and a goal brief describing what the resulting knowledge document should make possible. Abridged:
TASK BRIEF (excerpt)
You are a software engineer and your task is to understand the tribal knowledge of this company from its codebase and data warehouse. Your understanding must come only from the searching and reading you do on the fly. Produce a single markdown document covering: summary, why this project, business understanding, metrics, system, data, experimentation, glossary. There is no time limit, take as much time as needed.
GOAL BRIEF (condensed)
understand how the business numbers are actually produced, end to end; understand product and customer analytics, including when a dashboard's number should not be trusted at face value; understand the recommendation system and the nightly jobs.One clause is worth a gloss, because it sounds stricter than it is: understanding must come from this company's artifacts, not from prior knowledge of it, it is an anti-leakage rule, not a ban on how a system organizes its reading. Every system was free to index, cache, or preprocess the frozen estate however it liked, under identical credentials. Note also what the prompt does not contain: no questions, no hints about any specific claim, no pointers at any table or commit. Each system explored the company for as long as it wanted and wrote its book. To make run-to-run variance visible rather than hidden, every system was run three times under the same frozen brief, producing three books each.
Scoring
Every book is scored against the 51 gold claims by a rubric-anchored judge, one claim at a time: explicit requirements the book must state and statements it must not contradict, five independent passes per book, majority verdict. Recall is path-blind: a claim found in one query and a claim found after a thousand hops earn the same credit. The headline for each system is the mean of its three books; "solved in all three" is what it recovered in every one of its three books.
Results
- 74.5%
- Pavo
- 62.7%
- Claude Code
- 58.8%
- Cursor
- 33.3%
- Codex
mean of 3 books · 60.8% in all 3 · 88.2% in any
mean of 3 books · 35.3% in all 3 · 84.3% in any
mean of 3 books · 35.3% in all 3 · 80.4% in any
mean of 3 books · 19.6% in all 3 · 51.0% in any
Claim recall on the 51-claim gold set
Four systems, three books each: the mean, and how much it moves between runs.
Claim recall · share of this group’s gold claims
0–100%
All 51 claims
n = 51 claims
Pavo
74.5%
Claude Code
62.7%
Cursor
58.8%
Codex
33.3%
Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.
Pavo, Claude Code, and Cursor ran the same model at the same reasoning effort (for Codex we used the best available model, GPT-5.6 Sol), so the spread among them is not a model comparison, it is a comparison of what each system does around the model. The model is not the differentiator here. The system around it is.
Breakdowns
Two views of the same result, sliced by what a claim requires. Read the tick as the mean and the markers as the three books. The plots show spread; the across-claim dispersion figures that qualify each comparison are in the next section.
What a claim needs. Where its evidence lives, and what history it requires, evidence that no longer exists in the current state.
Recall by where the claim's evidence lives
Codebase-only claims are where the gap is widest.
Claim recall · share of this group’s gold claims
0–100%
Codebase only
n = 16 claims
Pavo
72.9%
Claude Code
47.9%
Cursor
43.8%
Codex
33.3%
Warehouse only
n = 15 claims
Pavo
68.9%
Claude Code
66.7%
Cursor
66.7%
Codex
26.7%
Codebase + warehouse (joint)
n = 20 claims
Pavo
80.0%
Claude Code
71.7%
Cursor
65.0%
Codex
38.3%
Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.
Why is the widest gap on code, the one surface every coding agent is built for? The claim-level data suggests an answer. Ranked by how often the generalists miss them, six of the seven hardest codebase-only claims are temporal-era or absence claims: a formula with two constant eras, a fee rule whose real cutover was a deploy rather than a commit, a backfill that never fills a column. Five of the seven need commit or log history to settle. They are history-and-absence questions wearing a code costume, which is what a commit-and-pull-request index is built to keep, and what the generalists here most often failed to dig out. This is a post-hoc reading, like the slices below, not an ablation.
Recall on claims that require history
History that no longer exists in the current state, and who can still recover it.
Claim recall · share of this group’s gold claims
0–100%
Requires commit history
n = 14 claims
Pavo
69.0%
Claude Code
57.1%
Cursor
59.5%
Codex
21.4%
Requires query history
n = 6 claims
Pavo
88.9%
Claude Code
88.9%
Cursor
61.1%
Codex
16.7%
Requires runtime logs
n = 9 claims
Pavo
66.7%
Claude Code
66.7%
Cursor
48.1%
Codex
3.7%
Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.
The reasoning a claim demands.
Recall by the reasoning a claim demands
Mechanism and temporal-era claims separate the systems most.
Claim recall · share of this group’s gold claims
0–100%
Mechanism
n = 6 claims
Pavo
88.9%
Claude Code
61.1%
Cursor
55.6%
Codex
27.8%
Temporal era
n = 10 claims
Pavo
70.0%
Claude Code
53.3%
Cursor
43.3%
Codex
23.3%
Absence
n = 9 claims
Pavo
88.9%
Claude Code
74.1%
Cursor
70.4%
Codex
48.1%
Reconciliation
n = 9 claims
Pavo
66.7%
Claude Code
59.3%
Cursor
59.3%
Codex
40.7%
Computed quantity
n = 12 claims
Pavo
55.6%
Claude Code
55.6%
Cursor
50.0%
Codex
13.9%
Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.
Where Pavo is better, and by how much
How large is Pavo's 74.5% against Claude Code's 62.7%, and how repeatable? The means alone cannot tell you, so we check two more things: what each system recovered in every single run, and how the claim-by-claim differences spread. A tribal-knowledge book is ingested and written once and then relied on for months, so what matters is less what a system can find on a good day than what it found in every one of its runs here, and on that measure the systems separate most. Look at the solved in all 3 books figures: Pavo recovers six in ten claims in every one of its books; Claude Code and Cursor recover roughly one in three in every book. Pavo's weakest book beats Cursor's and Codex's best. Eight claims are solved by every Pavo book and by no rival in all of its books; three more are solved by at least one Pavo book and by no rival book at all. The reverse ledger, so it is not hidden: three claims were found by some rival run and by no Pavo run, and on three claims a rival was consistent where Pavo was not. This pattern is consistent with what the index-plus-scaffolding design should produce: an index returns the same evidence on every run, where unguided exploration finds different things each time, and the agent consolidates across sources and fills in what the index does not find. The benchmark compares whole systems, it does not ablate the index out of Pavo, so read this as the mechanism the results point to, not one they isolate.
The table below puts a dispersion figure on each comparison. Every cell is the mean over three books. Δ is Pavo's mean minus the best rival's mean, and beside it is the across-claim paired SE of that difference [2], a dispersion summary of how the per-claim differences vary over the 51 curated claims. It is not a confidence interval: it excludes run and residual judge uncertainty, and supports no inference beyond this claim set.
TABLE 1 — Every axis, with its across-claim dispersion
| Axis | Pavo | Claude Code | Cursor | Codex | Δ vs best rival | Paired SE |
|---|---|---|---|---|---|---|
| Overall | 74.5% | 62.7% | 58.8% | 33.3% | +11.8 | ±3.8 |
| Codebase-only claims | 72.9% | 47.9% | 43.8% | 33.3% | +25.0 | ±6.5 |
| Codebase + warehouse (joint) | 80.0% | 71.7% | 65.0% | 38.3% | +8.3 | ±4.1 |
| Mechanism | 88.9% | 61.1% | 55.6% | 27.8% | +27.8 | ±10.2 |
| Temporal era | 70.0% | 53.3% | 43.3% | 23.3% | +16.7 | ±9.0 |
| Absence | 88.9% | 74.1% | 70.4% | 48.1% | +14.8 | ±8.1 |
| Requires commit history | 69.0% | 57.1% | 59.5% | 21.4% | +9.5 | ±11.3 |
| ML & recommendations | 66.7% | 60.0% | 53.3% | 35.6% | +6.7 | ±5.8 |
| Warehouse-only claims | 68.9% | 66.7% | 66.7% | 26.7% | +2.2 | ±8.9 |
| Reconciliation | 66.7% | 59.3% | 59.3% | 40.7% | +7.4 | ±10.8 |
| Requires query history | 88.9% | 88.9% | 61.1% | 16.7% | 0.0 | ±8.6 |
| Requires runtime logs | 66.7% | 66.7% | 48.1% | 3.7% | 0.0 | ±7.9 |
| Computed quantity | 55.6% | 55.6% | 50.0% | 13.9% | 0.0 | ±8.2 |
| Solved in all 3 books (consistency) | 60.8% | 35.3% | 35.3% | 19.6% | +25.5 | ±7.3 |
Three sources of variance, and what we did about each
Three things move these numbers, and each gets a different treatment.
- The run. Every candidate has an LLM in the loop, Pavo included, so the same system given the same frozen brief writes a different book each time. We ran three books per system: the headline is the mean of three, the plots show all three, and the consistency row is the summary of how much they move. Three runs make the spread visible; they do not remove it. The solved-in-all-three figure is descriptive, what a system actually recovered in every one of its three books, not a separately identified stability parameter: claims differ in difficulty, and with three runs there is no clean way to split repeatability from a system simply being ahead on the easier claims.
- The judge. An LLM judge can flip its verdict on the same claim in the same book. Every claim in every book was judged five times independently; the verdict is the majority. Majority voting damps the judge's instability; the residual judge error is not separately propagated in the ± figures, and because each system wrote a different book, that residue is only partly shared across systems.
- The claims. The 51 claims were curated, not randomly sampled, so no ± figure here is a population interval: the across-claim paired SE summarizes dispersion within this benchmark, and says nothing about other companies' estates.
That is why the paired SE sits next to the Δ it qualifies, and not on the plots.
We went one step further and measured the measurement itself, on a second simulated company: Fernbrook Goods, a four-year, multi-site commerce estate built independently of NovaMart with a deliberately harsher scoring standard, the full dataset story (how it was rendered, how the knowledge-rot was engineered, how the answers were sealed before any system ran) is in the NovaMart post. We ran over fifty scored books there under pre-registered expectations and logged every run, including the ones that went against us, in a ledger that ships with the release. Three findings calibrate everything above.
Judge noise is real. A single-pass re-judge of a byte-identical book moved its score 4 claims in 60, which is why five-pass majority judging is the protocol here. Run variance has a floor. Most configurations we measured, our own included, move about 6 claims in 60 run to run, so on corpora like ours any single-run difference under that floor is within the movement two identical runs already show. The stability exception, and the honest caveat that rides with it, are two paragraphs down.
On ceilings, Fernbrook is unambiguous: the best book any system has produced on it is Pavo's. Under the original 60-claim protocol that book banks 47 of 60 against a best generalist book of 37, a ten-claim best-to-best gap, larger than either the judge's re-judge band (±4 in 60) or the run-to-run floor alone, though a reader stacking both uncertainties should treat it as strong, not settled. Re-judged under this post's own protocol, with the claims decomposed into 734 atomic assertions and five judge passes, the ordering holds: Pavo's best book carries 88.4% of those assertions, the best generalist book 85.3%. What Fernbrook does not show is a population gap, across whole configurations the systems land within each other's observed run-to-run spread, which is what a corpus does when it hands every contestant a pre-exported estate and removes the reconstruction work. We report that plainly. And the uncomfortable part, said just as plainly: Pavo's two ceiling books, the 47, and a 38 from a second Pavo configuration, both come from higher-variance configurations than the stable line below; a high ceiling and a tight band have not yet come from the same configuration.
The variance finding is the one that generalizes, and it runs the other way. Six runs across two Pavo configurations at one code version held a four-point band on the 60-claim scale; the leading generalist's three clean runs swung eleven. A knowledge base that is a different document every time it is built is a different problem from one that is merely imperfect, and it is the problem the rest of this post is about.
What may explain the gap: the mechanism behind each delta
Three things separate Pavo from the other agents, and each one is checkable below.
- Pavo reads each source the way its owner would. A query log is not three million SQL strings, it is a record of the tasks people were doing. A commit history is not diffs, it is who changed what and why. A BI tool is not a pile of dashboards, it is a record of which ones people actually trust. Each index keeps that deeper reading, and the codebase-only and mechanism rows in Table 2 are where it shows up in the data.
- Most tribal knowledge was never written down, and you cannot index what does not exist. None of the 51 claims is a plain single-artifact lookup: twenty are recoverable from what the estate narrates, current code, commit history, repo docs, and the other thirty-one require computing over the warehouse, the logs, or the query history. Pavo's approach is to join sources and to investigate, deriving the facts no artifact states. Table 2 is candid about where that pays so far: the leads concentrate where deep reading and cross-source joining do the work, while the slices that demand fresh computation are level, the knowledge nobody wrote down is the hard part for every system here, Pavo included.
- The book is a product feature that keeps improving. This benchmark scored a single frozen pass for every system, Pavo included, the only fair comparison, so the numbers here are first-pass numbers. In production, Pavo re-investigates the gaps the book knows about and revises what changed; the last section walks that loop.
How do the first two claims show up in the numbers? The machinery from the extraction section splits into two parts. The scaffolding sets the floor: even on surfaces where Pavo has no index to lean on, it stays level with the best rival. The indices track the separation row by row: where a purpose-built index exists and the findings layer can join across it, Pavo leads; where the index only retrieves, it is level; where there is no index, it plays like the others. That is a post-hoc reading across thirteen overlapping slices, exploratory, not preregistered, and unadjusted for the multiple comparisons it invites, and the tables below walk every row through it, the leads and the non-leads alike, so the claims above can be checked rather than believed.
TABLE 2 — Where Pavo's largest leads coincide with the deep-reading slices
| Claims | Index at work | Δ vs best rival | Why |
|---|---|---|---|
| Codebase only | code index | +25.0 ± 6.5 | code is ingested into an index before writing begins; Pavo's weakest book matches the best rival book, and its other two are above anything a rival produced |
| Codebase + warehouse (joint) | database index + findings layer | +8.3 ± 4.1 | findings from one source are joined against the other before any question is asked |
| Mechanism | code index | +27.8 ± 10.2 | how a value is produced today is read off current code; none of these claims needs history |
| Temporal era | data-task index + commit & PR index | +16.7 ± 9.0 | what changed and when lives in the commits, PRs, and the query log, kept as history rather than skimmed as text; 9 of the 10 claims need commit or query history |
| Absence | code index + commit & PR index | +14.8 ± 8.1 | 5 of the 9 need history; the other 4 are dead paths visible in the current code, like a status no row carries or a table nothing writes |
| Requires commit history | commit & PR index | +9.5 ± 11.3 | same mechanism; 14 claims, so the across-claim spread is wide and the lead sits inside it |
| ML & recommendations | code index | +6.7 ± 5.8 | model and training code is covered by the code index, and the scaffolding asks for the ML system explicitly; 15 claims, a modest lead about the size of the across-claim spread |
| Solved in all 3 books | all indices + the scaffolding | +25.5 ± 7.3 | joined findings recurred across the three books; unguided exploration varied |
TABLE 3 — Slices where the observed differences are small or zero relative to the across-claim spread
| Claims | Index at work | Δ vs best rival | Why |
|---|---|---|---|
| Warehouse only | database index | +2.2 ± 8.9 | Pavo leads on the mean, but the difference is a quarter of the across-claim spread, too small to read as a lead; Claude Code's best single book here (86.7%) is the best book any system produced |
| Requires query history | data-task index | 0.0 ± 8.6 | exact tie, Claude Code reads what people actually queried as well as Pavo does |
| Reconciliation | database index + findings layer | +7.4 ± 10.8 | the boundary case for the pattern: reconciliation usually means computing a quantity neither surface stores, indices retrieve, they do not compute, so the lead stays inside the across-claim spread (Pavo flat at 66.7% in all three books) |
TABLE 4 — With no index to lean on, Pavo still does not fall behind
| Claims | Index at work | Δ vs best rival | Why |
|---|---|---|---|
| Requires runtime logs | none, scaffolding only | 0.0 ± 7.9 | the one surface without an index; Pavo relies on the scaffolding and reads raw logs like a generalist, 66.7% here against 76.2% on every claim that does not need logs |
| Computed quantity | none, scaffolding only; indices retrieve, they do not compute | 0.0 ± 8.2 | Pavo's least stable shape (75.0 / 41.7 / 50.0 across books); two of the three claims no system solved live here |
The book that maintains itself
Everything above measures day zero: point a system at a frozen estate, get a book. But tribal knowledge is not a snapshot problem. Definitions drift, sources get added, people push back on claims, and a knowledge system's real test is what happens on day thirty. Here the comparison stops being a score gap and becomes a category gap.
A generalist agent can of course be handed the previous book and told to revise it, what it lacks is not the ability to edit but the accountability machinery around the edit, which the rest of this section walks through. Without that machinery, the practical move is a rerun, which on our benchmarks produces a different book each time, an 11-point swing between the best and worst run of one configuration. Unioning those runs does not fix it: where two books disagree, deciding which one is right is precisely the synthesis work being measured, and a union has no verdict where its inputs disagree, and no single accountable artifact.
The difference starts with the shape of the deliverable. Every Pavo book is organized by the same editorial contract, why the business is shaped the way it is, how the metrics are defined and governed, how the systems actually behave, what was tried and what it taught, and the vocabulary a newcomer needs, because that is the shape in which organizational knowledge gets used, not the shape in which it happened to be found. And the brief cannot be the explanation: it names the same areas for every system, yet no two generalist books share a structure, each mirrors its own run's exploration order, while Pavo's chapter skeleton is identical section for section across every run. The fixed structure is not cosmetic. It is what makes a book reviewable section by section, diffable edition to edition, and revisable in place, none of which is even defined for a document whose organization changes every time it is written.
Pavo's book is an edition in a lifecycle, and every mechanism below ships in the product today. Adding a data source automatically schedules a targeted revision of the affected system's book: the revision lands as a reviewable draft edition, with the reason recorded, and connecting the source never blocks on it. The writer does not start over, it receives the current book as ground state plus the delta of newly connected sources, with instructions to ground new material in them and cite them where used. New context files fold in through the same path, as do claim verdicts and direct feedback: every revision traces to a concrete event, never a blind rebuild.
Every revision is accountable. The draft opens with a writer-authored "what changed in this revision," and every promotion appends a changelog record carrying the reason and the delta that actually landed, computed from the books themselves, not from prose. Dropping a previously grounded claim without an explicit human verdict is automatically flagged: knowledge does not silently disappear between editions.
Audits compound the same way, because the book grades itself. Every cited block yields a checkable claim, in effect, the book writes its own exam, and an auditor scores each claim against the estate, persisting the verdicts on the edition. Failed and stale claims become knowledge gaps; gaps become targeted revisions; unchanged sections carry their verdicts forward, and a staleness ratchet forces periodic full re-audits so carried verdicts cannot rot. A human sits over all of it in a workflow any engineer will recognize: a diff of the draft against the current edition, per-section approve, comment, or discard, partial approvals recomputed honestly, and a book assembled from partially accepted changes rendering unverified until re-attested. It is a pull-request workflow for what your company knows. The result is the property the benchmark's consistency rows only hint at: the generalists, as run here, started every book from scratch, and Pavo's knowledge accumulates, reviewed, cited, and explained, edition over edition, with every gap it finds becoming the next revision's work.
The book that maintains itself
Every turn is a targeted revision, never a rebuild.
Six stations · one closed wheel
01
An event lands
A data source is connected, a file is added, a reviewer leaves feedback, or claim verdicts come in. Every revision traces to a concrete event.
02
A targeted revision is scheduled
Not a rebuild. The writer receives the current book as ground state plus only the delta, and the event that triggered it never blocks on the revision.
03
The draft opens with a changelog
Writer-authored: what changed in this revision and why. The delta that lands is computed from the books themselves, not from the prose about them.
04
A human reviews it like a pull request
Per section: accept, comment, or discard. Partial approvals are recomputed honestly, and a grounded claim dropped without a verdict is flagged automatically.
05
Publish persists the verdicts
Verification verdicts are stored per edition and carry forward to unchanged sections — with a staleness ratchet forcing periodic full re-audits so carried verdicts cannot rot.
06
The audit's gaps become the next work
Failed and stale claims become knowledge gaps; gaps schedule the next targeted revision. The wheel does not have to wait for an outside event to turn.
back to 01
The axle
Each turn leaves the book more grounded, not merely different.
What edition n hands to edition n+1
- Ground state
- The current book, handed to the writer instead of a blank page.
- Verdicts
- Verification results on every section the revision did not touch.
- Judgments
- The accepts, comments, and drops a human already made, and the citations each claim earned.
11 of 60
best vs. worst rerun · companion corpus
A rerun carries none of it. Rebuilding from scratch swings that far between runs of a single configuration: a different book each time, not a better one, and no single accountable artifact to diff.
The wheel closes on itself: the gaps an audit finds are what schedule the next revision, so the book keeps improving between events rather than only in response to them.
This is not a roadmap slide. We ran the loop on one of the benchmark books from this study: a reviewer left three anchored comments and one piece of general feedback ("deepen the reasoning wherever a decision is narrated") on a 253-block Fernbrook book. Twenty minutes later a draft edition came back with 11 sections changed, nothing added or removed, and a changelog that opened: "This revision is a targeted edit, not a rebuild. Everything not listed below is unchanged from edition 1 and still stands." Each anchored comment was answered with the missing reasoning; figures that cross a definitional boundary now say so in the same sentence; and the draft closed by flagging one issue it deliberately did not change, "for a reviewer to adjudicate", instead of silently editing it. The generalist equivalent of that twenty minutes is a from-scratch rerun and a diff nobody wrote.
And because a book nobody can reach is tribal knowledge with extra steps: Pavo connects with read-only credentials to the systems a team already runs, warehouses (BigQuery, Snowflake, Databricks), dbt, code hosting, dashboards, documents, experiment trackers, and serves the book to both audiences it is written for. People read and review it; agents query it, the book, its per-claim citations, and an ask-the-book interface, so the next AI system pointed at your estate starts from the accumulated, reviewed knowledge instead of from zero.
What this result does not prove
- It does not rank models. Claude Code and Cursor ran the same model as Pavo, and that is the point. Each row is a whole system, scaffolding and model together.
- Prompt sensitivity. We ran every experiment under a single prompt setting. We reduced variance at three levels, run, judge, claims, but agentic systems are sensitive to small changes in their prompts, and we did not measure that. We expect Pavo to be the least affected, because its index construction and synthesis are deterministic, but we have not shown it.
- Recall, not precision. These tables count gold claims recovered. They do not score unsupported claims, contradictions, or citation validity, a book that banks more claims while asserting more confident errors would look better here than it deserves. The judge does check each book's evidence coverage, and those per-book scorecards ship with the release; until then, recall is the measured half of the trust story.
- Corpus dependence. On our second, harder corpus (Fernbrook, see the variance section) the gap compresses to a tie within the observed run-to-run spread at configuration level, because that corpus hands every system its history pre-exported and scores multi-clause claims strictly. The pattern, not the level, is what we defend: the deltas appear where knowledge must be reconstructed through indices, history, and joins, and compress where the estate is handed over flat.
- Cost. The benchmark does not price each system's book, and we owe the full table (tokens, dollars, wall-clock per system per run) with the artifact release. The magnitudes we can state now: a generalist writes its book in roughly half an hour of wall-clock; Pavo's initial run takes a few hours because indexing and synthesis run up front, a cost paid once per estate, not per question, and a targeted revision of a finished book takes about twenty minutes. If the cost table changes anyone's read of the results, that is exactly why it ships with the data.
Check our work
Every number in this post is re-derivable, and until the artifacts are live you should treat these numbers as our claim, not the field's. The release, the NovaMart estate, the gold claims with their recomputable evidence chains, all twelve books with their per-pass scorecards, and the judge harness, ships together, so the comparison can be rerun end to end rather than taken on trust [1]. If you build agents, including the ones in the plots above, we would genuinely like to see your numbers when it lands. The benchmark does not care whose logo is on the book.
References
- Pavo, NovaMart: we simulated a company, to measure what agents actually know. 2026. pavoai.com/blog/novamart-benchmark
- Miller, E., Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. Anthropic, 2024. arxiv.org/abs/2411.00640. The paired-difference computation follows this paper; we read the result as dispersion over the fixed, curated claim set, not as the paper's population interval.
- Anthropic, Claude Code. docs.claude.com/en/docs/claude-code
- Anysphere, Cursor. cursor.com
- OpenAI, Codex. openai.com/codex
- Anthropic, Claude Fable 5 and Claude Mythos 5. 2026. anthropic.com/news/claude-fable-5-mythos-5
- OpenAI, GPT-5.6 Sol. 2026. openai.com
