OpenAI stopped reporting SWE-bench Verified (February 23, 2026) after an audit found gains “increasingly reflect how much the model was exposed to the benchmark at training time,” and retracted its SWE-bench Pro recommendation (July 8, 2026) after flagging roughly 27–34% of public tasks as broken. A benchmark whose hypotheses, metrics, and analysis are written down after the runs is an instrument for confirming whatever its author hoped.
Designed for two-stage preregistration
Disclosed S-REF development pilots predate the freeze and cannot support scientific claims. Hypotheses, matrix, metrics, and analysis rules freeze before any live-sandbox observation; adapter and baseline operationalizations, artifact hashes, and environment pins freeze before confirmatory execution. Frozen content is append-only; amendments are numbered and stamped.
Draft methodology — no scientific results yet
Sealed private holdout
The Stage-B design requires the held-out fault-seed split to be sealed as complete, labeled artifacts, with ciphertext and plaintext root hashes committed before confirmatory execution. That human-gated sealing act has not occurred. Once sealed, the split will govern the publish threshold and kill criterion without entering the repository or an agent context.
Draft methodology — no scientific results yet
Statistics that can say “no”
At least 5 seeds per cell (10 planned); confidence intervals and effect sizes, never point estimates; every executed cell reported; INVALID runs retained and marked, never deleted; second-machine reproduction and independent recomputation before any public claim. BetterBench self-scoring, Datasheets, and Croissant metadata ship with results.
Draft methodology — no scientific results yet