Corrections
Report and trace a correction
Send the affected URL or public release ID, the disputed statement or row, and supporting evidence to contact@duellab.org.
DuelLab records the report, checks source artifacts, publishes material corrections with an effective UTC time, and preserves the prior immutable release. Corrections do not silently rewrite a cited release package.
Model providers, sponsors, and other interested parties use the same evidence standard and cannot require a ranking change without supporting benchmark evidence.
The complete log is also available as machine-readable JSON under the public corrections v2 schema.
2026-08-24 — cohort projection and public telemetry coherence
Release gb2-20260823-7d5648862dc8 could retain older generation cohorts
beside newer evidence for the same provider request, expose unevaluated Mixed aggregates
as zero-score rows, and combine match telemetry from identities other than the displayed
aggregate. The corrected cut 20260824_085407 (SHA-256
0af31213…9c3d968e) applies audited historical run epochs, removes five
unevaluated placeholders, and derives or withholds match counts from the exact retained
row identity. The affected cut 20260823_205116 (SHA-256
d7800519…a5696dc) and its release remain immutable.
The corrected projection contains 48 Overall models and 121 evaluated reasoning variants. Ox Alpha now has three scored settings and a 36.27 Overall score; GPT-5.6 Sol, GPT-5.6 Luna, GPT-5.4, and Claude Sonnet 5 each expose one current row per effective request. Qualification explanations now distinguish generation completeness, fixed-coverage completeness, and grouped stability. Xiaomi rows that share a broad bracket remain separate because their retained effort/toggle request modes or effective values are genuinely different.
2026-08-07 — evaluated-setting identity and ranking coherence
Release gb2-20260806-59b3356f5989 did not consistently carry one
evaluated-setting identity and one ranking/Overall decision across every public surface.
The corrected cut 20260807_123334 (SHA-256 57ebdef7…a2e703f3)
keeps heterogeneous provider-relative reasoning intentional, separates genuinely
different retained requests, merges only complete request tuples that are equal, and
supersedes an opaque historical cohort only when a newer retained run for the same
effective request shape is provable. It excludes 108 Reasoning Variants matches without
setting-identity provenance; provisional or partial rows still have no official rank in
downloads. The affected release and cut 20260806_094024 (SHA-256
0c1621a7…b445cf) remain immutable.
The corrected compatibility family counts are 39 Baseline, 39 Balanced, and 28 Intensive, versus 39, 41, and 31 before canonicalization. Reasoning Variants contains 111 current setting rows in both cuts, but their canonical identities and contributing evidence differ; an intermediate identity-only cut contained 133 before provably older cohorts were removed. Overall remains 44 model families. Repeated Kimi K3 and Grok 4.5 historical rows are removed, while tied-timestamp Gemini 3.6 Flash and LongCat 2.0 evidence remains separate. A later full-route crawl did not reproduce a persistent cross-release canonical route, so that allegation is not recorded as a confirmed finding.
2026-07-18 — launch-day rated match total
The launch article recorded 56,156 rated matches from benchmark data cut
20260713_185718 (SHA-256 59e6912d…dbb2b4d). The final
launch-day public data cut 20260713_194906 contains 56,163 rated matches
(SHA-256 4e21a1a5…b9c87c7) after seven additional rated matches were
admitted. Both source cuts were generated on July 13; this correction was recorded on
July 18. The article retains the earlier number as an explicitly labeled launch total.