Report and trace a correction

Send the affected URL or public release ID, the disputed statement or row, and supporting evidence to contact@duellab.org.

DuelLab records the report, checks source artifacts, publishes material corrections with an effective UTC time, and preserves the prior immutable release. Corrections do not silently rewrite a cited release package.

Model providers, sponsors, and other interested parties use the same evidence standard and cannot require a ranking change without supporting benchmark evidence.

The complete log is also available as machine-readable JSON under the public corrections v2 schema.

2026-08-24 — cohort projection and public telemetry coherence

Release gb2-20260823-7d5648862dc8 could retain older generation cohorts beside newer evidence for the same provider request, expose unevaluated Mixed aggregates as zero-score rows, and combine match telemetry from identities other than the displayed aggregate. The corrected cut 20260824_085407 (SHA-256 0af31213…9c3d968e) applies audited historical run epochs, removes five unevaluated placeholders, and derives or withholds match counts from the exact retained row identity. The affected cut 20260823_205116 (SHA-256 d7800519…a5696dc) and its release remain immutable.

The corrected projection contains 48 Overall models and 121 evaluated reasoning variants. Ox Alpha now has three scored settings and a 36.27 Overall score; GPT-5.6 Sol, GPT-5.6 Luna, GPT-5.4, and Claude Sonnet 5 each expose one current row per effective request. Qualification explanations now distinguish generation completeness, fixed-coverage completeness, and grouped stability. Xiaomi rows that share a broad bracket remain separate because their retained effort/toggle request modes or effective values are genuinely different.

2026-08-07 — evaluated-setting identity and ranking coherence

Release gb2-20260806-59b3356f5989 did not consistently carry one evaluated-setting identity and one ranking/Overall decision across every public surface. The corrected cut 20260807_123334 (SHA-256 57ebdef7…a2e703f3) keeps heterogeneous provider-relative reasoning intentional, separates genuinely different retained requests, merges only complete request tuples that are equal, and supersedes an opaque historical cohort only when a newer retained run for the same effective request shape is provable. It excludes 108 Reasoning Variants matches without setting-identity provenance; provisional or partial rows still have no official rank in downloads. The affected release and cut 20260806_094024 (SHA-256 0c1621a7…b445cf) remain immutable.

The corrected compatibility family counts are 39 Baseline, 39 Balanced, and 28 Intensive, versus 39, 41, and 31 before canonicalization. Reasoning Variants contains 111 current setting rows in both cuts, but their canonical identities and contributing evidence differ; an intermediate identity-only cut contained 133 before provably older cohorts were removed. Overall remains 44 model families. Repeated Kimi K3 and Grok 4.5 historical rows are removed, while tied-timestamp Gemini 3.6 Flash and LongCat 2.0 evidence remains separate. A later full-route crawl did not reproduce a persistent cross-release canonical route, so that allegation is not recorded as a confirmed finding.

2026-07-18 — launch-day rated match total

The launch article recorded 56,156 rated matches from benchmark data cut 20260713_185718 (SHA-256 59e6912d…dbb2b4d). The final launch-day public data cut 20260713_194906 contains 56,163 rated matches (SHA-256 4e21a1a5…b9c87c7) after seven additional rated matches were admitted. Both source cuts were generated on July 13; this correction was recorded on July 18. The article retains the earlier number as an explicitly labeled launch total.