A real evidence page (illustrative data)

This is what a defensible
model switch looks like.

Every ReasonRank recommendation ships with a page like this — the verdict, the statistics behind it, and the per-case receipts your staff engineer will ask for.

All figures on this page are illustrative — your models, cases, and traffic produce your own numbers.

support-triage-agent · recommendation #rec_4f2a · dataset v4 · ~175k calls/moillustrative

Verdict: switch

GPT-5.6 Sol → GPT-5.6 Luna

$2,890/mo saved

$0.0210/call in prod$0.0045/call recommended

95% CI on quality delta

−0.6% … +2.3% · clears −2% tolerance

quality holds: worst case −0.6%, may improve

paired over 24 test cases · production baseline: 1,204 pre-switch calls · switch detected in live traffic

Illustrative figures — a representative evidence page, not a live recommendation.

per-case scores · 9 of 24 shown · dataset v4 (pinned)illustrative

Every case, both models, side by side

paired observations
Test caseSol scoreLuna scoreSol costLuna cost
refund-request-multi-order0.960.94$0.0223$0.0047
angry-customer-escalation0.980.92$0.0246$0.0052
billing-dispute-partial-creditcandidate wins0.910.93$0.0231$0.0049
password-reset-loop1.001.00$0.0118$0.0026
shipping-delay-international0.950.90$0.0209$0.0044
duplicate-charge-report0.930.92$0.0187$0.0041
feature-request-misfiled0.880.85$0.0164$0.0036
cancellation-save-flow0.970.91$0.0258$0.0055
gdpr-data-export-request0.920.90$0.0201$0.0043

Illustrative scores. In a real run, every score links to the graded output and the evidence behind it.

What the auditors check

The three questions this page answers before anyone asks.

Method & sample, in the open

Paired non-inferiority test with a cluster bootstrap over test cases — never over correlated repeats. The exact case count, interval, and tolerance sit next to the verdict, so the sample size can’t hide.

Judge trust status

The judge behind these scores was measured against the team’s own labels before its scores counted — and it’s re-checked on a schedule. A judge that stops agreeing with humans is benched, visibly.

Dataset pin

This run is pinned to an immutable snapshot of the eval suite — v4. “Which cases did that number come from” has exactly one answer, and dataset diffs separate suite drift from model drift.

Want this for
your agents?

Every recommendation ReasonRank makes ships with an evidence page like this one — on your models, your cases, and your production traffic. The page above is illustrative; yours won't be.