A real evidence page (illustrative data)
This is what a defensible
model switch looks like.
Every ReasonRank recommendation ships with a page like this — the verdict, the statistics behind it, and the per-case receipts your staff engineer will ask for.
All figures on this page are illustrative — your models, cases, and traffic produce your own numbers.
Verdict: switch
GPT-5.6 Sol → GPT-5.6 Luna
$2,890/mo saved
$0.0210/call in prod→$0.0045/call recommended
95% CI on quality delta
−0.6% … +2.3% · clears −2% tolerance
quality holds: worst case −0.6%, may improve
paired over 24 test cases · production baseline: 1,204 pre-switch calls · switch detected in live traffic
Illustrative figures — a representative evidence page, not a live recommendation.
Every case, both models, side by side
paired observations| Test case | Sol score | Luna score | Sol cost | Luna cost |
|---|---|---|---|---|
| refund-request-multi-order | 0.96 | 0.94 | $0.0223 | $0.0047 |
| angry-customer-escalation | 0.98 | 0.92 | $0.0246 | $0.0052 |
| billing-dispute-partial-creditcandidate wins | 0.91 | 0.93 | $0.0231 | $0.0049 |
| password-reset-loop | 1.00 | 1.00 | $0.0118 | $0.0026 |
| shipping-delay-international | 0.95 | 0.90 | $0.0209 | $0.0044 |
| duplicate-charge-report | 0.93 | 0.92 | $0.0187 | $0.0041 |
| feature-request-misfiled | 0.88 | 0.85 | $0.0164 | $0.0036 |
| cancellation-save-flow | 0.97 | 0.91 | $0.0258 | $0.0055 |
| gdpr-data-export-request | 0.92 | 0.90 | $0.0201 | $0.0043 |
Illustrative scores. In a real run, every score links to the graded output and the evidence behind it.
What the auditors check
The three questions this page answers before anyone asks.
Method & sample, in the open
Paired non-inferiority test with a cluster bootstrap over test cases — never over correlated repeats. The exact case count, interval, and tolerance sit next to the verdict, so the sample size can’t hide.
Judge trust status
The judge behind these scores was measured against the team’s own labels before its scores counted — and it’s re-checked on a schedule. A judge that stops agreeing with humans is benched, visibly.
Dataset pin
This run is pinned to an immutable snapshot of the eval suite — v4. “Which cases did that number come from” has exactly one answer, and dataset diffs separate suite drift from model drift.
Want this for
your agents?
Every recommendation ReasonRank makes ships with an evidence page like this one — on your models, your cases, and your production traffic. The page above is illustrative; yours won't be.