IrrevonBench

The methodology is the product. Scientific results don’t exist yet.

Developmental pilots disclosed · no scientific results

The preregistration is a DRAFT — nothing is frozen. Synthetic S-REF harness and fault-smoke pilots have occurred, including a 488-effect attribution-hardening pilot; they are permanently non-confirmatory engineering evidence. No live-sandbox observation or confirmatory run has occurred. Stage A must precede live-sandbox work and Stage B must precede confirmatory execution; both freezes are human acts. There are no scientific numbers on this page.

Evidence: Verified evidence · Design specificationTechnical provenance

The reproduction contract is published before the first live-sandbox or confirmatory run: benchmark reproduction guide →

Why this shape

Benchmark credibility failed publicly in 2026

OpenAI stopped reporting SWE-bench Verified (February 23, 2026) after an audit found gains “increasingly reflect how much the model was exposed to the benchmark at training time,” and retracted its SWE-bench Pro recommendation (July 8, 2026) after flagging roughly 27–34% of public tasks as broken. A benchmark whose hypotheses, metrics, and analysis are written down after the runs is an instrument for confirming whatever its author hoped.

Evidence: Verified evidenceTechnical provenance

Designed for two-stage preregistration

Disclosed S-REF development pilots predate the freeze and cannot support scientific claims. Hypotheses, matrix, metrics, and analysis rules freeze before any live-sandbox observation; adapter and baseline operationalizations, artifact hashes, and environment pins freeze before confirmatory execution. Frozen content is append-only; amendments are numbered and stamped.

Draft methodology — no scientific results yet

Sealed private holdout

The Stage-B design requires the held-out fault-seed split to be sealed as complete, labeled artifacts, with ciphertext and plaintext root hashes committed before confirmatory execution. That human-gated sealing act has not occurred. Once sealed, the split will govern the publish threshold and kill criterion without entering the repository or an agent context.

Draft methodology — no scientific results yet

Statistics that can say “no”

At least 5 seeds per cell (10 planned); confidence intervals and effect sizes, never point estimates; every executed cell reported; INVALID runs retained and marked, never deleted; second-machine reproduction and independent recomputation before any public claim. BetterBench self-scoring, Datasheets, and Croissant metadata ship with results.

Draft methodology — no scientific results yet

Evidence: Design specificationTechnical provenance

The baseline ladder

“Never weakened so the proposed system wins” is a design rule, not a slogan

Every rung is pre-specified in the current draft. The preselected primary comparator is the composite B5+B3+B6 — the strongest realistic conventional stack — with B5 reported alongside; superiority must reject against both. The ladder is rendered whole, residual failure modes included.

Draft methodology — no scientific results yet

Scroll table horizontally

#BaselineSolvesResidual failure mode
B0No protectionnothingall duplicates / orphans / lost
B1Model-argument hashingidentical-retry dedupbreaks under re-synthesis (different args)
B2Agent-generated idempotency keyssome retrieskey changes when the model regenerates it
B3Stable workflow-issued operation IDscaller-side identitydestination still may not honor the key (C2)
B4Destination-native idempotencyC1 duplicatesunavailable on C2/C3
B5Durable runtime + native idempotencyC1 + own-state exactly-oncepunts external APIs to idempotency keys; no C2 adjudication
B6Provider-native status check / reconcilesome detectionnot general; varies by provider
B7Model-assisted semantic matchingflags likely dupesprobabilistic; unsafe as sole authority
RProposed system (Irrevon)C2 detection + safe re-dispatch + orphan/lost surfacingnone beyond destination capability (C3 unsolvable)

Evidence: Design specificationTechnical provenance

The kill criterion

The result that ends the project is written down first

Verbatim from the draft preregistration (§1, implementing master doc §8.6 as amended by AM-14):

“the kill fires iff, for every primary metric on the confirmatory stratum, either (a) the preselected composite comparator is statistically equivalent to R — TOST against the §5.3 margin, never CI overlap — or (b) it is statistically better than R; and (c) no confirmatory cell shows the comparator worse than R by at least the worst-cell gate (§5.3) — a pooled equivalence coexisting with a large local regression is reported as inconclusive with a localized signal, not as a kill and not as a win. If the kill fires, Irrevon is unnecessary and the project is reframed as a teaching artifact (master doc §14). The benchmark is explicitly not designed so the system must win; clause (b) exists so a baseline that beats R kills at least as surely as one that ties it.”

The pre-committed C1 null

On destinations with dependable native idempotency, Irrevon is expected to show no advantage on duplicate rate. That null is a registered hypothesis (H0-C1) and will be reported as prominently as any win.

Evidence: Testable hypothesisTechnical provenance

The C3 impossibility demonstration

On opaque destinations every method — including Irrevon — must fail to detect lost and orphaned effects. The boundary is documented, not excused.

Evidence: Design specificationTechnical provenance

Evidence: Testable hypothesisTechnical provenance

Metrics glossary

Denominators fixed by the oracle, not by the arm

An arm cannot change the denominator it is judged by, and no metric rewards adopting Irrevon’s internal data model.

Draft methodology — no scientific results yet

Scroll table horizontally

MetricDefinition
Duplicate-effect ratesurplus destination effects per true intent ÷ intents eligible for dispatch (oracle-fixed denominator; never computed from finding counts)
Orphaned-effect ratedestination effects with no corresponding true fixture intent ÷ total destination effects at read-back
Lost-legitimate-effect ratelegitimate intended effects absent at read-back ÷ legitimate intended effects
False-suppression ratelegitimate effects blocked ÷ legitimate effects
Unresolved-ambiguous rateAMBIGUOUS unresolved at run end ÷ dispatched (descriptive, arm-conditional)
Precision / recallvs the oracle, per classification (descriptive)
Human-review rate & mean timeescalations ÷ effects; wall-clock per escalation
Time-to-detectdispatch → detection interval
Compensation correctnesscorrect compensations ÷ compensations attempted (N/A when none)

Evidence: Design specificationTechnical provenance

Research integrity

Read the plan, not a summary of it

The draft, in the open

The full preregistration — hypotheses, strata, analysis plan, invalid-run rules, holdout protocol — is a public document in the repository, labeled DRAFT until a human freezes it.

benchmark-preregistration.md

Evidence status, everywhere

Every public claim distinguishes verified evidence, evidence-backed analysis, testable hypotheses, design specifications, and open questions. Each claim links to the detailed source record.

Evidence: Verified evidenceTechnical provenance

Publication path — a plan

arXiv preprint → agents/reliability workshop → NeurIPS Datasets & Benchmarks track, only if the preregistered criteria and independent recomputation hold. No preprint exists today.

Evidence: Design specificationTechnical provenance