NovaMart: we simulated a company to benchmark tribal-knowledge extraction

Tribal knowledge lives in how a change ripples across a company's surfaces (code, database, logs, query history, warehouse, dashboards), and those artifacts, joined and consistent, are exactly what real companies can never publish, which is why no public benchmark for it exists (survey through July 2026). So we executed a simulated company into being, grounded in real traffic, with causally connected sources and the tribal knowledge around them, and we are preparing to release it for everyone.

Part one of a two-part release. This post is the benchmark: what NovaMart is, how it was built, and how it is kept honest. The comparison it was built for, Pavo against three general-purpose agents, is in part two.

The knowledge which is not explicitly documented anywhere

Every company runs on knowledge that nobody explicitly writes down. The QA test account that has been inflating daily actives for years, excluded by two dashboards, counted by a third. The revenue report whose filter list exists nowhere except in the SQL of the person who built it. The migration everyone cites as complete that, in the data, never actually ran. Ask a tenured analyst "what was October revenue, and how did you compute it?" and you will get the SQL, long, precise, full of filters and exclusions. Ask why each filter is there, and the answer lives in exactly one place: their head.

We call this tribal knowledge, and at Pavo we build systems that acquire it: point an agent at a company's codebase, warehouse, and logs, and have it write down what a veteran employee knows.

Why tribal knowledge matters for agentic systems

Agentic systems are increasingly given the work of data engineers and analysts: build a pipeline, answer a metric question, modify production code. Without tribal knowledge, they fail in a characteristic way, not loudly, but plausibly. Asked to add a revenue metric, an agent that does not know the restated view exists will create its own column, and the company now has one more definition of revenue than it had yesterday. Asked for October revenue, it will compute a number that disagrees with every existing report, because it applied none of the undocumented filters. The output looks correct and is wrong, and in a data stack that is the most expensive kind of wrong, it gets shipped, cited, and built upon before anyone notices.

The second reason is attrition. Tribal knowledge lives in people, and people leave. When the tenured analyst departs, the why behind every filter departs with them; the SQL stays behind, frozen, and within a few quarters nobody can safely change it. If there is no system that captures this knowledge and keeps it current, the company does not merely lose a person, it loses the ability to explain its own numbers. So the capability we care about is precise: a system that can read a company's estate, its code, warehouse, logs, and dashboards, everything it has accumulated, and reconstruct not just what the numbers are but why they are computed that way, and hold that knowledge after the people who created it are gone.

Which raises the question this whole project answers: how do you measure whether a system can actually do that?

The lack of public benchmarks for tribal knowledge

Measuring this capability needs a benchmark, and here the field has a structural problem: tribal knowledge is inherently multi-source. A single claim's evidence routinely spans the codebase, the warehouse, the query history, and the job logs together, and each of those surfaces is, on its own, difficult for any company to publish. Warehouse rows carry user data, with GDPR and its equivalents standing behind them. Production query histories and deployment logs expose internal operations and security posture. The codebase is intellectual property. And publishing them together, joined and consistent, is harder still, because the joins between surfaces are precisely the sensitive part. When such artifacts do get shared, Snowflake's Snowset [8] and Amazon Redshift's Redset [9] are the notable examples, they arrive stripped to metadata: query shapes without statement text, workloads that cannot be joined to a schema, logs disconnected from the commits that produced them. Valuable for database research; insufficient for evaluating whether a system can reconstruct how a company actually works.

As of our survey in July 2026, we are not aware of any public benchmark that provides these surfaces together, a codebase with its history, a warehouse, a query log, job logs, and dashboards, all connected and mutually consistent. The consequence is that everyone building tribal-knowledge systems iterates the only way available: on private customer data. That works, to a point, we have done it ourselves, but it has obvious costs. Results cannot be published or compared, progress is invisible outside each vendor's walls, and teams without enterprise customers cannot participate at all. The absence of a public benchmark does not just make evaluation inconvenient; it slows the development of the capability itself.

To solve that, we simulated a company. We call it NovaMart.

NovaMart: the simulated company

Meet NovaMart

NovaMart is a first-party consumer-electronics retailer that operated through Q4 2019. It stocks 81,018 products from 3,410 brands, serves 38,950 registered users, and closed the quarter at $2.93M in net sales with a 31.0% gross margin. None of it is real, and every number in this section is computed live from its warehouse.

$2.93M
Net sales

Q4 2019

31.0%
Gross margin

$908K gross profit

9,113
Orders

$321 average order value

1,349
Avg. daily active users

40.4% repeat-buyer rate

Every figure computed live from the NovaMart warehouse; peak single-day sales $240,347.

The books behave like a retailer's books. It buys from 8 supply vendors at cost and retails at list, so it reports net sales, cost of goods, and gross profit, and the 2.9% payment-processing line reconciles to the payment gateway's monthly statements to the cent, against the same fee schedule that lives in the codebase. The engagement funnel runs 843,085 product views → 36,925 cart adds → 9,127 orders (the orders table above counts 9,113, the two figures are computed at different grains, and that kind of quiet disagreement is this benchmark's whole subject), a 1.08% view-to-purchase rate, squarely in the range real e-commerce reports. Premium brands dominate revenue while value brands move volume, exactly the way real marketplaces skew. A recommendation engine served 667,850 slates at a 3.34% attributed engagement rate, spanning two model generations and a nightly-trained successor. There is a nightly data platform, reporting, KPI, affinity, forecasting, and fraud jobs on a scheduler, plus monthly finance statements with fee audits and restatements, and a fraud pipeline with holds and manual release operations.

NovaMart is small, but it operates the way a company operates.

The estate

What ships is not a dataset, it is the company's entire artifact estate, every surface a real enterprise has:

SurfaceWhat it isSize
Codebase + git historythe application, with its full commit history32 modules · 112 commits
Production databasethe live application tables12 tables
Analytics warehouseanalytics tables and views24 tables/views
Historical query logevery SQL statement anyone or anything ever ran, with parameters3,597,650 statements
Application event loguser- and system-level events1,570,017 entries
Job-run logevery scheduled job execution, including crashes978 runs
Dashboardsthe BI layer, with its authoring and refresh history in the query log9
The history spans roughly 3.5 months of continuous operation, 2019-09-16 through 2019-12-31.

Note the entries a real company could never hand you: the query history with full statement text, the deployment-era job logs, the dashboards' revision history. Those are the surfaces where tribal knowledge actually lives, and here, they are all present, all consistent with each other, and all publishable.

How NovaMart compares to existing benchmarks

There are excellent public benchmarks for agents around enterprise work, and each is strong at what it measures. The difference is surface coverage: each provides the one or two surfaces its task needs, where tribal knowledge requires the joined estate.

BenchmarkSurfaces providedLongitudinal historyWhat it measures
SWE-bench [1]codebase + issues/PRs (real repos)code history onlypatch correctness
MLE-bench [2]ML datasets + competition leaderboardsnoneML engineering performance
Spider 2.0 [3] / BIRD [4]databases / warehouse (static snapshot)noneanalytical SQL over enterprise data
TheAgentCompany [5]running services, files, chatstate set up per task, no historytask completion
CRMArena [6]synthetic CRM orgnoneCRM task completion
NovaMart (ours)codebase + git, production DB, warehouse, query log, app & job logs, dashboards3.5 months, across all surfacestribal-knowledge recall

None of this is a criticism of those benchmarks, a text-to-SQL benchmark does not need a git history, and a patch benchmark does not need a warehouse. But a tribal-knowledge claim routinely needs several of these surfaces at once, joined and mutually consistent, and that combination is what has not existed publicly. The next section explains why we could not simply author it.

Executed, not authored

How do you get an estate like that? The established approach is to author it: write data generators, effectively line by line, that produce tables consistent with an invented backstory, and hand-place the inconsistencies you want agents to find. Done carefully, this yields convincingly messy data, but it does not scale, in two ways. First, every agreement between two tables is something an author asserted and must now maintain by hand, so each additional surface multiplies the consistency work. Second, some surfaces cannot realistically be hand-written at all: a believable multi-million-statement query history, a CI trail, a dashboard that breaks the morning after a schema change. An authored world can have detailed state, but it has no operational history. Even the environments that run real software while the agent works host company data that was curated into place and reset per task: the software executes, but the history was never lived.

NovaMart's history was not written. It was executed, from two inputs:

How NovaMart was made

Two inputs, executed into a company whose every artifact is a side effect.

One conductor · one clock

Real shopper traffic

The public REES46 event stream, replayed hourly as ambient load.

The tape

An authored script of the company’s story, arranged as a causal graph.

Conductor

Executes both inputs on one simulated clock.

The simulated company

executes a 3.5-month history

Simulated engineers

Develop it, break it, investigate it: commits, bugs, write-ups.

Application + database

Real services against real rows; every order traces to a real session.

Analytics platform

Nightly jobs and dashboards executing real SQL.

The NovaMart estate

frozen & released

Codebase + git history

32 modules · 112 commits

App tables

12 tables

Analytics tables

24 tables/views

Query log

3.6M statements

Runtime logs

1.57M log lines

Dashboards

9 dashboards

The events are authored; their consequences are not — everything the company would have accumulated, accumulates as a side effect of execution.

FIG. 01 — How NovaMart was made: real traffic plus an authored event tape, executed into a company whose every artifact is a side effect.

The first input is real shopper behavior: a public e-commerce event stream [7], replayed hour by hour on a simulated clock as the company's ambient traffic. The second is the tape: an authored script of the company's story, feature launches, incidents, migrations, investigations, arranged as a causal graph, where events fire at scheduled moments, in reaction to what the traffic actually does, or downstream of one another. The tape is written by hand, and deliberately so: engineers who have spent years inside production systems, our own and our customers', authored its events to reproduce the failure patterns they have personally seen in real companies: the fee rule that changes mid-quarter, the migration that half-lands, the experiment nobody set up a control arm for, the report quietly capped by a scan limit. The events are authored; their consequences are not. The conductor executes both inputs against a real application and a real database; simulated engineers and the analytics platform act out the tape's events on the same clock; and everything the company would have accumulated, commits, logs, queries, dashboard refreshes, accumulates as a side effect. When the simulated history ends, everything is frozen and exported: that frozen output is the estate described above. The deeper mechanics of the pipeline are deliberately out of scope for this post.

This produces the property the project exists for: cross-surface causal consistency. Change something at one surface and the change mechanically manifests at the others, because the others are downstream of the same execution. One example of what that means in practice: a single mid-November code change to how orders are recorded propagates, without anyone authoring any of it, into a new database table minutes later, a divergence between two analytics grains, a drift in dashboard numbers, and eventually an investigation trail in the query history. Five surfaces change for one underlying cause, and every step of that chain can be recovered and verified from the artifacts.

The story is scripted; every artifact follows from executing it.

One incident, end to end

Here is one actual tape incident, exactly as it executed in the simulation. The purchase traffic it runs on is the replayed REES46 stream: a public dataset of roughly 285 million real shopper events, product views, cart adds, purchases, recorded on a large multi-category online store [7]. So every order and payment below traces back to a real browsing session.

The setup. In mid-October 2019 of the simulated history, the order service had no idempotency check: when the payment provider confirmed a purchase, the service created an order and a payment without ever asking whether it had already processed that confirmation. Deliver any confirmation twice, and the company books two orders and charges the customer twice.

The authored incident. Knowing the codebase was in this state, the tape scripts a provider glitch: for three hours on the simulated afternoon of 2019-10-14, every purchase confirmation is delivered twice, five seconds apart. The tape also scripts the follow-up, set to fire only if the database actually accumulates ten or more duplicated payment references: the next morning, finance notifies engineering that customers look double-charged. That is all that was authored. From here everything executed end to end; nobody touched the database, the codebase, or the logs directly.

Sim timeSurfaceWhat execution produced
Oct 14, 14:00–17:00Database17 duplicate orders and their payments land, each exactly 5 seconds after its original; no other sub-10-second twin exists anywhere in the company's history
Oct 15, morningTape triggerthe duplicated references push past the follow-up's threshold of 10, so it fires: finance reports customers who look double-charged
Oct 15, 14:00Codebasea simulated engineer responds: commit b676969 ("Make order callbacks idempotent by payment_ref") adds the guard, and no new twin is ever created again
after Oct 15Databasethe storm's duplicates were never cleaned up; they are still in the frozen table
nightlyRuntime logsthe reconcile job flags the suspect references every morning; it only warns, never fixes, and by the final night it flags 130: the storm twins plus slower organic repeat purchases
The incident leaves its mark in the suite: one gold claim tests the scar (payment references are not unique, and the nightly reconcile only warns about the 130 suspect ones), while the story behind it, the storm window, the guard, the cleanup that never happened, stays recoverable from the commit history and the job logs.

The benchmark

What counts as tribal knowledge

A world is not a benchmark; to measure anything we needed ground truth, the inventory of what a tenured NovaMart employee would know. We wrote these as claims: specific, individually checkable statements. Three rules decide membership. A claim enters the gold set only if:

  1. It is stated plainly nowhere. The fact is not a plain lookup from any single current-state artifact; reaching it takes computing over the data, reading through history (logs, queries, git), joining evidence across artifacts, or digging out behavior that nothing narrates.
  2. It is operationally load-bearing. Getting it wrong produces wrong numbers, wrong joins, or unsafe changes somewhere in the company's work.
  3. It is mechanically recomputable. A typed evidence chain, a SQL statement, a commit hash, a log excerpt, re-derives it from the shipped artifacts.

Difficulty is deliberately not part of the gate. Some claims are easy for agents and some have never been solved by any system we have tested; both belong, because the score means "what fraction of this company's tribal knowledge did you acquire," and easy knowledge is still knowledge.

One example, which we disclose because it illustrates the format. Ask NovaMart's data a simple question, what was October's gross revenue?, and there are three ways to answer it, and they disagree:

  1. Read the official monthly statement → 1,230,332.43. This is the number that was published on November 1st and presented to leadership.
  2. Read the finance team's "final" warehouse view → 1,226,764.86. It is lower because in December, three chargebacks against October orders were hand-booked into an analytics table, and a view subtracts them from October retroactively. The original statement was never updated, so the "published" number and the "final" number permanently disagree by 3,567.57.
  3. Recompute October yourself from the orders table → a third, still different number. The database overwrites an order's status in place, so orders that were paid in October but refunded in November now look refunded, October, as the live tables remember it, is no longer the October that was published.

None of the three is wrong; each is the right answer to a slightly different question. A tenured finance person knows which one to quote in which meeting, and the why behind the disagreement is a story spanning the codebase, two databases, a hand-run backfill session, and a view definition. That is what a tribal-knowledge claim looks like. There are 50 more, and we are not going to tour them here, for the obvious reason: they are the answer key.

The claim set, by the numbers

The current gold set is 51 claims. Here is how it is distributed, because a benchmark should show its composition, not just its scores.

The claim set, by the numbers

A benchmark should show its composition, not just its scores.

Shares of the 51-claim gold set

Where the evidence lives

0–60%

Codebase + warehouse (joint)

39%

20 of 51 claims

Codebase only

31%

16 of 51 claims

Warehouse only

29%

15 of 51 claims

The reasoning a claim demands

0–60%

Computed quantity

24%

12 of 51 claims

Temporal era

20%

10 of 51 claims

Absence

18%

9 of 51 claims

Reconciliation

18%

9 of 51 claims

Mechanism

12%

6 of 51 claims

Multi-source join

6%

3 of 51 claims

Other

4%

2 of 51 claims

Required history

0–60%

No history required (current state suffices)

51%

26 of 51 claims

Requires commit history

27%

14 of 51 claims

Requires runtime logs

18%

9 of 51 claims

Requires query history

12%

6 of 51 claims

Business area

0–60%

ML & recommendations

29%

15 of 51 claims

Orders, payments & fraud

22%

11 of 51 claims

Catalog & customers

20%

10 of 51 claims

Revenue & reporting

20%

10 of 51 claims

Pipelines & scheduling

10%

5 of 51 claims

All four panels share one 0–60%scale and one grid pitch (a grid line every 10 points), so a bar in one panel is directly comparable to a bar in another. The first, second and fourth panels each partition the gold set; the required-history panel does not — one claim can need more than one history source, so those four shares overlap and sum past 100%.

FIG. 02 — The claim set’s composition: four cuts of the same 51 claims, every panel drawn on one 0–60% scale so a bar in one panel is comparable to a bar in another.

The claim set is what remains after an audit designed to remove our own work. Sixty-eight claims were authored against the frozen estate; seventeen did not survive. Ten were cut in a blind review, five in a quality review for failing the first membership rule, they turned out to be plain lookups from current code, and two were removed by an independent audit that re-derived every evidence chain against the estate, an audit which also corrected six more claims it found broken. Fifty-one were released. The membership bar is the three rules stated earlier, no claim is a plain lookup from a single current-state artifact, getting it wrong produces wrong work, and every claim is mechanically checkable against the frozen estate.

Claims were authored by studying the frozen estate, reading git history, querying the warehouse, tracing incidents to their downstream consequences, not by transcribing the script that caused them. That ordering matters, and it is a rule rather than a habit: ground truth flows from the script to the claims, never from the artifacts to the claims, because claims mined from artifacts would inherit whatever a model guessed while reading them. Artifacts serve one purpose in the other direction, to verify that a claim is actually recoverable, and they are permitted to refute the script. In the pilot, the very first claim ever authored asserted an incident signature that turned out to be false about its own world; verification caught it, and the claim was corrected to match reality rather than the plan. Where the realized world diverged from the authored story, the gold follows the world.

Why claims, not questions

The obvious way to evaluate this would be question answering: ask the agent "what are October's revenue numbers?" and grade the answer. It is how most benchmarks work, and it is the wrong primary metric here, for a reason that goes to the heart of what tribal knowledge is: a question presupposes that someone knew what to ask.

In a real company, the expensive failures are the questions nobody asked. The new analyst does not fail to answer "which October number should I use?", they never ask it. They query the statements table, get a clean-looking number, ship the board deck, and are silently wrong. Knowledge you lack does not announce itself as a question; it announces itself as a mistake, weeks later.

So the primary evaluation hands the agent no questions at all. The agent gets credentials to the estate and one instruction, essentially: go understand this company, and write down what you learn. It explores however it wants, for as long as it wants, and produces a document, its tribal book. We score the book against the 51 claims: for each, did the book state it fully, partially, not at all, or state its opposite? The headline metric is claim recall: what fraction of the company's tribal knowledge ended up in the agent's head, unprompted.

This is not a lab abstraction, by the way, it is literally the deployed job. The real-world deliverable of understanding a company is a knowledge artifact, the onboarding doc, the runbook, the ingested knowledge base, not a quiz score.

The evaluation

The experiment

The candidates. Three widely used, general-purpose agentic systems, evaluated zero-shot, out of the box, with identical access and an identical brief:

SystemUnderlying modelConfiguration
CursorFable 5high reasoning effort, out of the box
Claude CodeFable 5high reasoning effort, out of the box
CodexGPT-5.6 Solhigh reasoning effort, out of the box
All nine runs were performed on August 27, 2026, on Claude Code 2.1.252, Cursor 3.15.6, and Codex CLI 0.120.0.

The prompt. Each system received the same two-part brief: a task brief (credentials for the estate, the ground rules, the output format) and a goal brief describing what the resulting knowledge document should make possible. Abridged:

You are a software engineer and your task is to understand the tribal knowledge of this company from its codebase and data warehouse. Your understanding must come only from the searching and reading you do on the fly. Produce a single markdown document covering: summary, why this project, business understanding, metrics, system, data, experimentation, glossary. There is no time limit, take as much time as needed.

The full brief ships with the benchmark. Note what the prompt does not contain: no questions, no hints about any specific claim, no pointers at any table or commit. Each system explored the company for as long as it wanted and wrote its book. To make run-to-run variance visible rather than hidden, every system was run three times under the same brief, producing three books each.

Scoring

Scoring a free-form document needs a judge, and LLM judges deserve scrutiny, so the design is deliberately conservative. Every claim carries a rubric, explicit semantic requirements the book must state, and statements it must not contradict, so the judge checks the book against fixed criteria rather than an overall impression. Every book is scored five times independently and we take majority verdicts; across passes, per-system recall moves by at most ~3 points and usually under 2. Scoring is path-blind: a claim found in one elegant query and a claim found after a thousand hops earn the same credit, recall measures the outcome, a separate evidence-coverage metric checks that the book's assertions cite real artifacts, and cost is reported, never scored. And before the gold set froze, every one of the 51 evidence chains was independently re-executed against the estate, because a benchmark whose answers can be recomputed is a benchmark whose answers can be trusted.

Results

One reading note: this post is the instrument. Pavo is deliberately absent from the table below, we score our own system against these three, under this same protocol, in the companion post.

The headline for each system is the mean of its three books; the "solved in all three" figure is what it recovers reliably, every run.

62.7%
Claude Code

mean of 3 books · 35.3% solved in all 3

58.8%
Cursor

mean of 3 books · 35.3% solved in all 3

33.3%
Codex

mean of 3 books · 19.6% solved in all 3

Gold v1.1, 51 claims. Majority verdict over 5 judge passes per book.

Claim recall on the 51-claim gold set

Three books per system: the mean, and how much it moves between runs.

SystemsClaude CodeCursorCodex
How to readtick = mean of the 3 books, in that row’s own inkone marker = one booksame system order in every group

Claim recall · share of this group’s gold claims

0–100%

All 51 claims

n = 51 claims

Claude Code

62.7%

Cursor

58.8%

Codex

33.3%

Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.

FIG. 03 — Claim recall on all 51 gold claims: one marker per book, tick at the mean of the three. Every recall figure below uses this same 0–100% scale and the same system order.

Nobody is close to solved. The best system misses more than a third of the claims; the weakest misses two thirds.

Reliable knowledge is much smaller than any single run suggests. The spread between runs is wide: a single run would have reported Claude Code anywhere from 58.8% to 70.6%. The dependable core, the claims a system solves in all three of its runs, is 35.3% for the best two systems and 19.6% for Codex. (The nine books run 5,000–8,490 words, and length does not predict recall: the longest book belongs to the weakest system.)

The frontier is unnarrated history. Seven of the 51 claims are solved by all nine books, and six are solved by none. The unsolved six require reconstructing history the estate never narrates, a fee boundary set by a deploy, not by any document.

What separates the systems

Systems differ most where knowledge must be computed, not read. NovaMart's narrated history, commit messages, docstrings, documentation, is compact enough for an agent to read in its entirety at negligible cost, and it is reasonably well written, so knowledge that someone once wrote down is recovered cheaply. But a company's most expensive tribal knowledge is precisely the part nobody narrated: the incident fixed without a writeup, the definition that drifted without a commit message acknowledging it. The claims that require genuine excavation, reading the query log, reading job output, computing quantities that were never stored, reconciling surfaces that disagree, are recovered at meaningfully lower rates than narrated ones, and far less consistently: this is where the systems separate from each other, while on claims the history narrates they cluster within a few points.

The tags make it exact: 20 claims are narrated (the answer can be read off code, git history, or docs) and 31 are excavated (it must be computed from the warehouse, the logs, or the query history). The systems sit within 10 points of one another on the narrated set and span 42 points on the excavated one. What separates these systems is not reading ability; it is excavation.

Recall by narration

The weakest system collapses on excavated claims.

Narrated (n=20): read from code, git, docsExcavated (n=31): computed from the data systems

60.0

64.5

Claude Code

60.0

58.1

Cursor

50.0

22.6

Codex

Mean claim recall over each system's three books, on a 0\u2013100% scale. A claim is narrated when its rubric can be satisfied from text surfaces alone, and excavated when satisfying it requires computing over the warehouse, the runtime logs, or the query history.

FIG. 04 — Recall by narration: claims a book can read versus claims it must compute.

Claims that require history sources hurt the weaker systems most. Split disjointly, 26 claims are answerable from the current code and warehouse alone, 16 need commit or query history, and 9 need runtime logs. Pooled recall falls stratum by stratum, 57.3% to 49.3% to 39.5%, and the fall is not even: Claude Code actually gains 6.4 points from the easiest stratum to the hardest, Cursor loses 14.7, and Codex collapses from 48.7% to 3.7%, close to zero on anything that requires reading job output. These are the surfaces a real company can never publish, and they are exactly where recall is lowest.

Recall by required sources, three disjoint strata

Averaged across systems, deeper required history means lower recall, and the weaker systems fall fastest.

Code + warehouse suffice (26)Needs commit/query history (16)Needs runtime logs (9)

60.3

64.6

66.7

Claude Code

62.8

58.3

48.1

Cursor

48.7

25.0

3.7

Codex

Mean claim recall over each system's three books, on a 0\u2013100% scale. The strata are disjoint: every claim sits in exactly one. History drops the weakest system to 3.7%.

FIG. 05 — Recall by required-source stratum: three disjoint groups, from current-state claims to the ones that need runtime logs.

Three more breakdowns

The remaining breakdowns are descriptive rather than load-bearing: the groups are small, so read directions, not decimals.

Recall by where the claim's evidence lives

The hardest code claims are about the code's past and its gaps, not its current state.

SystemsClaude CodeCursorCodex
How to readtick = mean of the 3 books, in that row’s own inkone marker = one booksame system order in every group

Claim recall · share of this group’s gold claims

0–100%

Codebase only

n = 16 claims

Claude Code

47.9%

Cursor

43.8%

Codex

33.3%

Warehouse only

n = 15 claims

Claude Code

66.7%

Cursor

66.7%

Codex

26.7%

Codebase + warehouse (joint)

n = 20 claims

Claude Code

71.7%

Cursor

65.0%

Codex

38.3%

Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.

FIG. 06 — Recall by where the claim’s evidence lives.

Claims whose only evidence is the codebase are the hardest for the two leading systems, which looks odd next to the excavation result until you rank them: of the seven code-only claims the leaders miss most, six are temporal-era or absence claims, the formula with two constant eras, the fee rule whose real cutover was the afternoon deploy, the backfill that never fills category, the dashboard whose query returns nothing against the frozen data. Reading today's code tells you what it does; these claims ask what it used to do, or what it quietly fails to do.

Recall by the reasoning a claim demands

Each system has a distinct strength profile.

SystemsClaude CodeCursorCodex
How to readtick = mean of the 3 books, in that row’s own inkone marker = one booksame system order in every group

Claim recall · share of this group’s gold claims

0–100%

Mechanism

n = 6 claims

Claude Code

61.1%

Cursor

55.6%

Codex

27.8%

Temporal era

n = 10 claims

Claude Code

53.3%

Cursor

43.3%

Codex

23.3%

Absence

n = 9 claims

Claude Code

74.1%

Cursor

70.4%

Codex

48.1%

Reconciliation

n = 9 claims

Claude Code

59.3%

Cursor

59.3%

Codex

40.7%

Computed quantity

n = 12 claims

Claude Code

55.6%

Cursor

50.0%

Codex

13.9%

Multi-source join

n = 3 claims

Claude Code

77.8%

Cursor

88.9%

Codex

77.8%

Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.

FIG. 07 — Recall by the reasoning a claim demands. Group sizes run from 3 to 12 claims (n is printed on the figure); in the 3-claim group one claim is worth 33 points, so read that row as direction, not margin. Two claims fit none of the six shapes and are excluded from this cut.

Claude Code is the strongest of the three on mechanism claims, how a value is produced in the current code, on temporal-era claims, and on computed quantities. Cursor's edge is the multi-source joins. Codex collapses exactly where a claim requires computing something that is not stored.

Recall by business area

Pipelines & scheduling is everyone's weakest ground.

SystemsClaude CodeCursorCodex
How to readtick = mean of the 3 books, in that row’s own inkone marker = one booksame system order in every group

Claim recall · share of this group’s gold claims

0–100%

Orders, payments & fraud

n = 11 claims

Claude Code

84.8%

Cursor

78.8%

Codex

48.5%

Catalog & customers

n = 10 claims

Claude Code

60.0%

Cursor

50.0%

Codex

26.7%

Revenue & reporting

n = 10 claims

Claude Code

56.7%

Cursor

63.3%

Codex

26.7%

ML & recommendations

n = 15 claims

Claude Code

60.0%

Cursor

53.3%

Codex

35.6%

Pipelines & scheduling

n = 5 claims

Claude Code

40.0%

Cursor

40.0%

Codex

20.0%

Claim recall on the NovaMart gold set. Each marker is one of a system’s three books (each book a majority verdict over 5 judge passes) and the tick is the mean of the three; books that scored the same are fanned apart vertically so all three stay countable. Every row is drawn on the same 0–100% scale, and the systems keep the same order in every group. Group size (n) is printed beside each group name, and several groups are small — where one claim is worth 10 points or more — so read narrow gaps as direction rather than margin.

FIG. 08 — Recall by business area. Segments run from 5 to 15 claims (n is printed on the figure), so a single claim moves a segment by 7 to 20 points.

Pipelines & scheduling is the weakest area for every system, the knowledge there lives in job logs and scheduler history, the least-documented part of the estate. Six of the 51 claims are solved by no run of any system in this post, and seven are solved by every run of all three. That spread is the benchmark's frontier, and it is intentional: a benchmark that is saturated at release measures nothing, and one nobody can move measures nothing either.

How we stress-tested our own numbers

Before publishing any of the numbers above, we spent a month trying to break them on a second simulated company. Fernbrook Goods, a multi-site commerce estate with four years of history, built independently of NovaMart with a harsher scoring standard, exists to measure the measurement: how much of an agent-benchmark score is real, and how much is noise wearing a number's clothes. We ran over fifty scored books there, ours and the generalists', every cell pre-registered with an expectation and a falsifier before launch, every completed run counted whether it helped us or not.

Fernbrook is built the opposite way from NovaMart, on purpose. NovaMart is executed, authored events run against a real application and database, and the consequences are real because they actually happened. Fernbrook is rendered, a multi-site commerce company with a four-year history, authored as a single 1.1MB "system bible" and deterministically generated into six repositories with 3,506 commits and 627 pull requests, a warehouse of 145 tables and eight million rows behind 146 DDL files each carrying column-level descriptions and lineage, 241 documents, 121 experiment records, 91 saved queries, and 18 dashboards: 6,071 artifacts, about 3.8 million tokens of text, nearly four times the 1M-token context window of the largest-context model evaluated here, so selective retrieval is the only strategy for every system, symmetrically.

Nothing true exists that the bible does not already know, which makes ground truth a byproduct of design rather than something extracted afterwards and argued about; sixteen validators cross-check the rendered surfaces against each other before the corpus freezes. Two worlds, two construction philosophies, one shared property: every claim is checkable to its source.

The knowledge-rot in Fernbrook is engineered, not incidental. Roughly a third of its enumerable artifacts are deliberately stale: whenever a value or policy is superseded in the company's history, its earlier state survives somewhere as a plausible artifact with no tombstone, a runbook still asserting an old discount cap, a policy page frozen at a threshold two changes back, a deprecation note declaring "permanently off" a mechanism that was later revived for a narrower purpose, a drafted launch page for a channel that never launched, legacy saved queries that still run and still assert superseded definitions. Fourteen of the sixty claims carry explicit traps: twenty-five decoy assertions, each verified verbatim in a surviving artifact before the freeze, so trusting the wrong page is a scored, detectable failure instead of a judgment call.

And nobody, including us, saw the answers. The gold claims and the corpus are hashed together into a single freeze digest, minted before any system, Pavo included, touched the corpus, and the harness re-verifies the bytes before every run. Book identities are sealed as hash-derived blind labels created before any book existed, and judging runs against labels, not systems. The evidence chains were verified mechanically pre-freeze, 270 of 270 evidence links resolve, 268 of 268 quoted excerpts occur verbatim in their sources, and the pre-registration carries a leak canary: any book containing a verbatim bible string that exists in no rendered artifact voids that question and triggers a full audit.

Three findings from that corpus shape this benchmark's protocol. Judge error is real: re-judging a byte-identical book under an identical configuration moved its score by 4 claims in 60, so every claim here is judged five times with a majority verdict, and we treat single-pass agent comparisons, ours included, as unmeasured. The judge is also from a different model family than the systems it scores (Gemini 2.5 Pro on Fernbrook), and rank order measured invariant across four judges from three vendors. Run variance has a floor: the same configuration re-run from scratch varies by about 6 claims in 60, which means a three-run experiment can only resolve large gaps, and on corpora like ours, any single-run comparison smaller than that floor is inside the noise, measuring nothing. Your own data will try to fool you: our adversarial pass caught a phantom book that had been scored under the wrong label, and a cross-workspace contamination where one agent's book was built on another's leftover drafts, and fixing both raised the competitors' scores, not ours. The forensics ship with the release, in full, because a benchmark you cannot audit is an anecdote.

Fernbrook's strict scoring also taught us where claim-level judging breaks: a requirement that bundles a value, a date, and a rationale into one sentence has no stable yes-or-no answer for any judge. That is why gold claims are decomposed into atomic, independently checkable assertions before judging, the protocol both corpora now share. None of this is incidental hygiene; it is anti-gaming architecture, designed because we grade our own product: two independently built corpora so no single world's idiosyncrasies decide the story unchecked, a mitigation, not an elimination, and more worlds are planned, a CI purity gate that mechanically bans benchmark-derived values from every model-facing string, and, as above, a cross-family judge and pre-registered falsifiers so a surprising result cannot be quietly reframed. The habit we would offer anyone building agent benchmarks: when a result surprises you, suspect the instrument before the hypothesis.

Limitations of NovaMart

We presented NovaMart: a simulated company, its full estate, and 51 audited claims measuring how much tribal knowledge an agent recovers. The world is small next to a real company, yet far from solved: the best system misses over a third of the claims, and none reliably recovers more than 35.3% across its three runs. The traffic is real, and the application, the warehouse, and the dashboards are real running software; what the world lacks is the volume and the routine-chore noise of a multi-year history, the thousands of unremarkable commits and hand-written queries a real company accumulates around the interesting ones. Small cuts both ways: it is what keeps every claim auditable end to end, and it means the absolute numbers belong to this one world. What we expect to transfer is the pattern, that systems separate on excavation and on the surfaces that carry history, not the levels.

The scores mix model with machinery. We evaluated shipped products, and that is a choice rather than an accident: the products are what teams actually deploy, and a book is the output of the whole system, the scaffold, the retrieval, the budgets, and the model together. It does mean the benchmark cannot tell you how much of a gap belongs to the model and how much to the scaffolding around it. Separating the two takes a reference harness that runs bare models under one fixed scaffold, and that is future work.

Recall has no precision counterpart yet. The score counts how many of the 51 gold claims a book recovers; it does not penalize what a book asserts beyond them. The judge does check contradictions, a book that contradicts a gold fact earns nothing for that claim, but a fabricated statement about a surface no claim touches costs nothing at all. A precision measure over the whole book is an open problem we want to solve.

Next steps

The roadmap has three threads.

Scale. A substantially larger NovaMart, more services, a bigger codebase, a longer operating history, is in progress. We are reporting nothing from it until it has run.

More surfaces. The current estate deliberately contains no explicit documentation layer, no wiki, no design docs, no runbooks, so the only narration an agent can lean on is what lives in the code history. Adding a documentation surface is next, with further sources to follow.

Noise. Real companies' documentation is partly stale, partly contradictory, and occasionally wrong; today's estate is cleaner than reality. Future versions will inject exactly that kind of noise, stale docs, misleading comments, definitions that drifted after they were written down, so that reading the narration stops being a reliable substitute for checking the artifacts.

Try your agent on NovaMart

NovaMart will be public soon. The release will include everything you need to point your own agent at the company: the codebase with its full git history as a public repository; the warehouse, logs, and query history on BigQuery, read-only; the dashboard SQL definitions and metadata alongside rendered snapshots; and the brief, the gold claims, and the judge harness, so you can give your agent the same instructions we gave ours, and score the book it writes exactly the way we scored these.

Every number in this post, and every gold claim, can be re-derived from those artifacts; that is the point. The full setup guide will be published alongside the release. Until then, if you want early access or want to run your system against NovaMart, reach out.

References

  1. Jimenez et al., SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024. arxiv.org/abs/2310.06770
  2. Chan et al., MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. OpenAI, ICLR 2025. arxiv.org/abs/2410.07095
  3. Lei et al., Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows. ICLR 2025. arxiv.org/abs/2411.07763
  4. Li et al., Can LLM Already Serve as a Database Interface? A Big Bench for Large-Scale Database Grounded Text-to-SQLs (BIRD). NeurIPS 2023. arxiv.org/abs/2305.03111
  5. Xu et al., TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks. 2024. arxiv.org/abs/2412.14161
  6. Huang et al., CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments. NAACL 2025. arxiv.org/abs/2411.02305
  7. REES46 Marketing Platform, eCommerce behavior data from multi category store. Kaggle. kaggle.com/datasets/mkechinov/ecommerce-behavior-data-from-multi-category-store
  8. Vuppalapati et al., Building an Elastic Query Engine on Disaggregated Storage. NSDI 2020 (the Snowset workload dataset). usenix.org/conference/nsdi20/presentation/vuppalapati
  9. van Renen et al., Why TPC Is Not Enough: An Analysis of the Amazon Redshift Fleet. VLDB 2024 (the Redset workload dataset). vldb.org/pvldb/vol17/p3694-saxena.pdf