Frontier Benchmarks are the wayto evaluate production AI
Pavo learns from production to discover what matters, build validated metrics and calibrated judges, and keep every release benchmark aligned with real outcomes.

Meridian supportEvals
- loyalty-tier-mismatchPromoted
Assistant · 130 tok
“This may be because you failed to plan ahead before booking.”
Assigned blame to the customer.
Regressions block the deploy.
INFO eval.runner, suite rebuilt from production · 84/84 pass
14:23:01A new standard for production evaluation. Built with the judgment of an applied scientist. Designed to improve with every release.
Model the system first. Pavo reads your traces, code and metrics to learn how it behaves, and how it should.
Measure what predicts outcomes. Pavo finds failure modes and validates candidate metrics against expert labels and outcomes.
Assemble the benchmark. Pavo builds golden datasets and hard slices, then calibrates judges to your rubric.
Gate every release. Each build runs the suite; new failures expand the set and judge drift triggers recalibration.
Every disagreement improves the judge. Every regression expands the dataset. Every release makes the benchmark smarter.
The benchmark stays calibrated as production shifts.
Pavo investigates, labels, calibrates and maintains. Your team reviews the evidence, adds judgment where it matters and decides what ships.
Building benchmarks by hand
With Pavo
Works with your existing stack
Connect Langfuse, Braintrust, Datadog, OpenTelemetry or your warehouse with a native connector.
- Langfuse
- Braintrust
- Datadog
- OpenTelemetry
- Your warehouse
+ 20 more
- Golden dataset128 cases
- Judge agreement94%
- Gate checks84/84
- Langfuse
- Braintrust
- Datadog
- OpenTelemetry
- Your warehouse
+ 20 more
- Golden dataset128 cases
- Judge agreement94%
- Gate checks84/84
A benchmark, decomposed
Metrics are earned, most candidates are rejected, and the reasons stay visible.
An objective broken into weighted behaviours, judged, calibrated against people, and validated against production.
- Failure modes · last 7 days1,204 tracesPromoted to regression test.Refund policy · stale quote41Promoted to regression test.Tool retry loop28Multi-item cart drop17Locale fallback9
4 new modes · 2 promoted to regression tests
“Pavo finds failure modes our eval suite never anticipated, groups the related traces, and turns the hardest cases into regression tests before our next release.”
Principal PM - Candidate c47 vs baselineBaseline v14Task success0.860.91+5Grounding0.940.95+1Tool calls0.900.90±0Escalations0.120.08−4
Better on 3 of 4, cleared to ship.
“Before Pavo, every model or prompt change came down to opinions. Now we compare it against a trusted baseline and know whether it is actually ready to ship.”
Product Engineer - Claimed vs actualSaysDidRefund issuedClaimed: succeeded.Actually: succeeded.Address updatedClaimed: succeeded.Actually: failed.Order cancelledClaimed: succeeded.Actually: succeeded.Ticket escalatedClaimed: succeeded.Actually: failed.
2 silent failures, the agent reported success on both.
“Our agent could report success even when the action failed underneath. Pavo checks what the agent claims against what actually happened and catches failures that were previously invisible.”
VP Engineering - Trajectory · run 8825 turns · 2 tools1User intent parsed-2search_orderstool3Clarify line item-4issue_refundtool5State verified-
Task completed, verified against state, not the agent’s summary.
“Single-turn evals missed most of what mattered. Pavo evaluates the entire trajectory, every turn, tool call, and state change, to determine whether the task was actually completed.”
FDE
Case 01 · Production failures
4 new modes · 2 promoted to regression tests
“Pavo finds failure modes our eval suite never anticipated, groups the related traces, and turns the hardest cases into regression tests before our next release.”
Research
Peer-reviewed at NeurIPS, ICLR, KDD, WWW, and WSDM. The measurement methods Pavo ships are the ones we publish.
FAQs
Living Benchmarks are evals that continuously learn from production. They combine datasets, metrics, calibrated judges, hard cases and release criteria, all grounded in real agent behavior and business outcomes.
As your agents and users evolve, Living Benchmarks detect new failures, drift and stale tests, then update themselves so your team always knows what “better” means.
Observability tools show you what happened. Traditional evaluation tools help you run tests your team has already defined.
Pavo goes further. It discovers what should be measured, builds and calibrates the metrics, judges and datasets, connects them to real outcomes and keeps the benchmark current as production changes. Pavo works alongside your existing observability stack.
Pavo is model- and framework-agnostic. It evaluates the behavior captured in your traces, so teams can use it across different models, agent frameworks and architectures.
Pavo works with traces from systems such as OpenTelemetry, Langfuse, LangSmith, Braintrust, Arize, Datadog and your existing data warehouse.
If your agent already emits traces, you can connect Pavo without migrating your stack or instrumenting the application again. Pavo begins learning your system, mining production cases and building the first benchmark from there.
The time to a decision-ready benchmark depends on the complexity of the agent, the available production history and the outcomes being measured.
Pavo is designed to protect sensitive production data throughout ingestion, analysis and evaluation. The controls below are the ones teams ask about most:
- Encryption in transit and at rest
- Data retention and deletion
- PII redaction
- Customer-data isolation
- Whether customer data trains shared models
- SOC 2 and other certifications
- Cloud, VPC and on-premise deployment options
Talk to us for the specifics on any of these, including our current certifications and deployment options.
Pavo pricing is based on the scale of your production system and the capabilities your team needs. We consider factors such as trace volume, evaluation volume, benchmark complexity and deployment requirements.
Talk to us and we’ll recommend the right plan for your system.