Frontier Benchmarks are the wayto evaluate production AI

Pavo learns from production to discover what matters, build validated metrics and calibrated judges, and keep every release benchmark aligned with real outcomes.

platform.pavoai.com/evals

Meridian supportEvals

Golden dataset128 cases
  • loyalty-tier-mismatchPromoted
Judge - groundingAgreement 94%
  • Assistant · 130 tok

    “This may be because you failed to plan ahead before booking.”

    Assigned blame to the customer.

Regression gateDeploy #412
84/84 checks pass

Regressions block the deploy.

INFO eval.runner, suite rebuilt from production · 84/84 pass

14:23:01
Understand the system
Sources → system modelTracesCodeBiz metricsExperimentsAgentUsersToolsExpected behavior

Model the system first. Pavo reads your traces, code and metrics to learn how it behaves, and how it should.

Discover what matters
Metric candidatesExpertProdBizKept-conversionEscalation rateAvg latencyThumbs-up rate2 promoted · 2 rejected

Measure what predicts outcomes. Pavo finds failure modes and validates candidate metrics against expert labels and outcomes.

Build the benchmark
Golden dataset128 cases · hard slicescase 041case 042case 043Judge - calibrationAgreement0.71Rubric0.86ExamplesFine-tune94%

Assemble the benchmark. Pavo builds golden datasets and hard slices, then calibrates judges to your rubric.

Keep it alive
Release over release96 cases112 cases128 casesv12v13v14v15

Gate every release. Each build runs the suite; new failures expand the set and judge drift triggers recalibration.

Every disagreement improves the judge. Every regression expands the dataset. Every release makes the benchmark smarter.

The benchmark stays calibrated as production shifts.

Pavo investigates, labels, calibrates and maintains. Your team reviews the evidence, adds judgment where it matters and decides what ships.

Building benchmarks by hand

UnderstandsystemMinecasesDefinemetricsLabeldataBuildjudgesCalibrateMaintain

With Pavo

Built and kept calibrated

Works with your existing stack

Connect Langfuse, Braintrust, Datadog, OpenTelemetry or your warehouse with a native connector.

  • Langfuse
  • Braintrust
  • Datadog
  • OpenTelemetry
  • Your warehouse

+ 20 more

Pavo
  • Golden dataset128 cases
  • Judge agreement94%
  • Gate checks84/84

A benchmark, decomposed

An objective broken into weighted behaviours, judged, calibrated against people, and validated against production.

03 · HOW WILL WE KNOW?Frontier BenchmarksDECOMPOSE THE OBJECTIVE →CALIBRATE · VALIDATE ONLINETHE DECOMPOSITIONCALIBRATED BENCHMARKOBJECTIVEBEHAVIOURSΣ 100%METRIC · JUDGESCENARIO · SUPPORT AGENTTARGETResolve supportwithout refundleakageINTENT-FIT35%GROUNDING30%TOOL-USE20%SAFETY15%RESOLUTIONoutcomeINTENT JUDGELLMSKU-EXISTSCODETOOL-CALL RULECODECONFIRM GATECODEKEPT JUDGELLMMETRICS REJECTED3 OF 8 CLOSEDRefund rateGAMEABLEResponse lengthNO SIGNALSentiment scoreCOSTLY LABELSBENCHMARK v35 JUDGES · 2 LLM · 3 CODEHUMAN CALIBRATIONAGREEMENT 0.81AGREEMENT SET2,048 cases · 3 annotatorsSLICE COVERAGE87% · 4 slices thinCANDIDATE LEADERBOARD4 SCOREDA · prompt0.62B · heuristic0.58C · rerankerWINNER0.79D · routing0.44OFFLINEONLINELIVEOFFLINE SCORESr 0.79REVISE THE METRICS
  1. Failure modes · last 7 days1,204 traces
    Promoted to regression test.Refund policy · stale quote41
    Promoted to regression test.Tool retry loop28
    Multi-item cart drop17
    Locale fallback9

    4 new modes · 2 promoted to regression tests

    Pavo finds failure modes our eval suite never anticipated, groups the related traces, and turns the hardest cases into regression tests before our next release.
    Principal PM
  2. Candidate c47 vs baselineBaseline v14
    Task success0.860.91+5
    Grounding0.940.95+1
    Tool calls0.900.90±0
    Escalations0.120.08−4

    Better on 3 of 4, cleared to ship.

    Before Pavo, every model or prompt change came down to opinions. Now we compare it against a trusted baseline and know whether it is actually ready to ship.
    Product Engineer
  3. Claimed vs actualSaysDid
    Refund issuedClaimed: succeeded.Actually: succeeded.
    Address updatedClaimed: succeeded.Actually: failed.
    Order cancelledClaimed: succeeded.Actually: succeeded.
    Ticket escalatedClaimed: succeeded.Actually: failed.

    2 silent failures, the agent reported success on both.

    Our agent could report success even when the action failed underneath. Pavo checks what the agent claims against what actually happened and catches failures that were previously invisible.
    VP Engineering
  4. Trajectory · run 8825 turns · 2 tools
    1User intent parsed-
    2search_orderstool
    3Clarify line item-
    4issue_refundtool
    5State verified-

    Task completed, verified against state, not the agent’s summary.

    Single-turn evals missed most of what mattered. Pavo evaluates the entire trajectory, every turn, tool call, and state change, to determine whether the task was actually completed.
    FDE

Living Benchmarks are evals that continuously learn from production. They combine datasets, metrics, calibrated judges, hard cases and release criteria, all grounded in real agent behavior and business outcomes.

As your agents and users evolve, Living Benchmarks detect new failures, drift and stale tests, then update themselves so your team always knows what “better” means.

Observability tools show you what happened. Traditional evaluation tools help you run tests your team has already defined.

Pavo goes further. It discovers what should be measured, builds and calibrates the metrics, judges and datasets, connects them to real outcomes and keeps the benchmark current as production changes. Pavo works alongside your existing observability stack.

Pavo is model- and framework-agnostic. It evaluates the behavior captured in your traces, so teams can use it across different models, agent frameworks and architectures.

Pavo works with traces from systems such as OpenTelemetry, Langfuse, LangSmith, Braintrust, Arize, Datadog and your existing data warehouse.

If your agent already emits traces, you can connect Pavo without migrating your stack or instrumenting the application again. Pavo begins learning your system, mining production cases and building the first benchmark from there.

The time to a decision-ready benchmark depends on the complexity of the agent, the available production history and the outcomes being measured.

Pavo is designed to protect sensitive production data throughout ingestion, analysis and evaluation. The controls below are the ones teams ask about most:

  • Encryption in transit and at rest
  • Data retention and deletion
  • PII redaction
  • Customer-data isolation
  • Whether customer data trains shared models
  • SOC 2 and other certifications
  • Cloud, VPC and on-premise deployment options

Talk to us for the specifics on any of these, including our current certifications and deployment options.

Pavo pricing is based on the scale of your production system and the capabilities your team needs. We consider factors such as trace volume, evaluation volume, benchmark complexity and deployment requirements.

Talk to us and we’ll recommend the right plan for your system.