Research note · Experimental evidence
Executable experiment-contract tooling
Before the Benchmark: Make the Claimed Win Reviewable
An experiment is not reviewable after the evidence has been discarded.
The familiar failure
A good-looking chart is not a reviewable result.
Teams often begin with a real question, change a configuration, keep the most favorable output, and only later ask whether the comparison can be trusted. By then, the baseline may be unclear, the seed plan may be unrecoverable, failed trials may be absent, and the raw outputs needed to inspect a score may be gone.
The problem is not simply that the result may be wrong. It is that no one can determine what the result means without reconstructing the experiment from memory. A claimed win needs a declared comparison before it needs an audience.
A practical response
Preflight the experiment. Retain the evidence. Assess the claim.
- Preflight the contract. State the question, baseline, candidate, datasets or splits, metrics, seed plan, stopping rule, and required evidence before results exist.
- Retain each trial. Preserve the trial outputs, provenance, identifiers, failed attempts, and the inputs needed to check that the evidence belongs to the declared comparison.
- Assess the supplied record. Check whether the retained evidence is complete and consistent enough to support the claimed scope. If it is not, say so without converting missing evidence into a performance conclusion.
The executable artifact
ACE turns that discipline into a checkable workflow.
ACE, the Assured Comparison Engine , is an open-source Python package for preflighting experiment contracts and assessing retained JSON or CSV trial outputs. Its v0.1.1 release produces an ACCEPTED, REJECTED, or INCONCLUSIVE decision pack from the evidence supplied to it.
ACE checks declared identity, provenance, split, seed, baseline, metric, failed-trial, and required-statistical evidence. Essential gaps or mismatches fail closed. That is deliberately narrower than a benchmark runner and more useful than a post-hoc checklist that cannot see the underlying record.
First applied case
Token-Bleed shows why the verdict must be allowed to disappoint the headline.
The Token-Bleed benchmark completed its retained 20-seed local run across three catalog sizes and mapped the evidence into ACE. The complete record passed its evidence and provenance checks. ACE nevertheless returned REJECTED because the governed route missed its prespecified holdout F1 floor.
The business lesson is specific. Selective context used 16.9x, 25.2x, and 26.6x fewer prompt tokens than full-context stuffing at 300, 1,500, and 3,000 catalog objects. But at the 3,000-object holdout, governed and full-context routes were effectively tied on F1, while a simple lexical prefilter scored substantially higher than the governed route. The result supports a cost-reduction mechanism, not a universal accuracy claim for governed routing.
This is not a verdict on the sponsored McKnight study that inspired the benchmark’s structure. It is a bounded result for one synthetic task, one local model, and one route implementation. It does establish the practical discipline: compare a sophisticated route with a cheap baseline, retain the record, and let the prespecified test decide what can be said.
Read the R2.1 technical result and the preregistered R3 protocol .
Source record
Use the package and release record for implementation.
Cite this PubPoint record
Saleme, Michael K. (2026). Before the Benchmark: Make the Claimed Win Reviewable [Research note]. PubPoint. https://pubpoint.com/publications/before-the-benchmark/
Published August 15, 2026. Updated August 16, 2026 with the completed Token-Bleed R2.1 result and ACE v0.1.1. Material revisions are recorded here.
Related research