n=3 · σ 0.01
Measurement — INSTRUMENTO-MENTE
We scored 0/25 five times. None of them was the model.
Five generations of harness. Exit 0 every time. Five different causes, and not one of them was the thing we were trying to measure.
Measured, repeated and archived[MEDIDO]
Predicted or single-run, not yet archived[PREVISTO]
Retracted — the previous number stays visible[RETRATADO]
solid ink = repeated and archived · light = single run · struck = retracted
On the fifth attempt we ran the benchmark’s own gold patch — the diff it declares correct — through the exact same generate/apply/test path as every model arm. Apply rate: 92%. Resolved: 0/25. File: brutos/swe_gold_sanity_20260804/swe_gold_sanity.json (n=25). After the classifier correction, the same n=25 is in swe_gold2.json: still 0/25 resolved and 92% apply; infra_error 21 and test_fail 4, versus test_fail 23 and infra_error 2 before.
The test judge was handing Django-style test identifiers to pytest. Pytest replied no tests ran in 0.00s. The harness counted that as the model failing.
That day taught us nothing about capability. It taught us the thesis of this note.
The instrument lies more often than the model — and almost nobody audits the instrument.
The number we had to correct downward
n=3 · σ 0.01
We had been saying 10.1×. Then we measured it properly — same batch, three repetitions, the whole machine to ourselves — and it came out lower.
TP8 no speculation 183.6 → TP2 + DFlash-64 1,795.1 tok/s (9.78×); the TP2 arm alone is 112.7 tok/s, so speculation by itself is 15.5× (same session, n=3)measured
Four replicas, conc=32 per replica: aggregate 29.503 ± 104 tok/s, n=3, overlap 118 s. The archive has no interference percentage.measured
The gain is 10.1× over baselineretracted
Solid ink means repeated and archived. Light means model prediction, not measured. Struck through means we published it and later took it back.
Why it moved
The first run gave 200.7 for the baseline. That reading is in brutos/single10x-20260825-0011/resumo.json: TP8 baseline 200.7 tok/s, n=3, σ 2.0, in a window without cron suspension; the DFlash arm was not measured (failed the content gate). It was retracted in favor of window 0035 (cron suspended, isolated node, both arms in the same batch): 183.6 tok/s. We do not treat the cause of the 200.7 → 183.6 difference as proven.
TP8 no speculation 183.6 → TP2 + DFlash-64 1,795.1 tok/s (9.78×); the TP2 arm alone is 112.7 tok/s, so speculation by itself is 15.5× (same session, n=3). File: brutos/partp2-20260826-1950/resumo.json.
That is the whole point. A bad measurement does not announce itself. It returns a plausible number, on time, with exit 0.