Research · benchmarks · evidence
Compare
Build snapshot · Not livePut two to four runs side by side — after the comparability preflight, never before it. An invalid comparison is blocked by default, and the numbers are not rendered at all until a reader explicitly overrides it and accepts every labelled mismatch — including a reader who arrived on a link that asked for the override, which is a request to them, never an act performed on their behalf.
Build snapshot Not live Captured 2026-08-10T07:58:14.329Z Source benchmarks/reports Ingested 2026-08-10T07:58:14.316Z 44 JSON files seen
Why this page refuses more than it shows
Two runs are comparable only when every identity field their contracts require is present in both and equal: same suite and contract version, same metric name,
unit, aggregation and direction, same evidence tier, same workload and its hash, same dataset,
split and preprocessing where the suite is dataset-sensitive, same harness revision, same
environment profile, same threshold version, same warmup and measurement protocol, and the same
model-serving seam for a served model. Hardware is the one axis a reader may deliberately vary,
and varying it licenses nothing else.
In this corpus no two artefacts satisfy that bar — dataset hashes, workload hashes and harness revisions are absent almost everywhere — so every comparison here is blocked. That is a fact about the evidence, not a limitation of this page.
In this corpus no two artefacts satisfy that bar — dataset hashes, workload hashes and harness revisions are absent almost everywhere — so every comparison here is blocked. That is a fact about the evidence, not a limitation of this page.
Select runs
0 / 4Choose between 2 and 4 runs. This page never selects a run for you and never adds one to your selection.
No comparison yet
selectionA comparison needs at least 2 runs; 0 selected. Choose the runs explicitly — this page never picks them for you.
Select between 2 and 4 runs explicitly. Nothing is selected for you.