DuelLab blog
Notes from the DuelLab project
Short form technical write-ups on what we are building, what we are measuring, and what we are still unsure about. Benchmark findings today, methodology notes, and whatever else turns out to be worth writing down as the project grows.
-
Near the top, generation cost varies by 6.4×
GPT-5.6 Sol nearly matches the leader but costs 6.4× as much to generate. Across the GPT-5.6 family, XHigh costs about 22× to 88× as much as Medium in this test.
-
In GameBench 2, more reasoning is not always better
When 15 models were tested at three code-generation reasoning settings, 12 performed best at the highest setting—but three did best at None.
-
Introducing GameBench 2
GameBench 2 evaluates the game-playing programs AI models write by compiling them, running head-to-head matches, and reporting playing strength, reliability, and generation cost.
-
GPT-5.5 wins the middle
GPT-5.5 is not the overall leader in the current DuelLab results, but it takes the #1 medium setting. The interesting part is the curve: medium beats both none and highest.
-
Kimi K2.6 tops Reasoning Variants
DuelLab is a benchmark where AI models write game-playing programs. In the latest public results, Kimi K2.6 is only #6 overall but jumps to #1 in Reasoning Variants, powered by an unusually strong highest-effort mode.
-
Claude Opus 4.7 is the first Claude with a V-shaped effort curve
On the overall number, Claude Opus 4.7 looks like a small step over 4.6. Inside the per-track data, the shape changed: 4.7 regressed at the medium effort tier and moved up at both ends. GPT-5.4 and Gemini 3.1 Pro Preview already had this shape. Claude 4.6 did not.
-
Introducing the DuelLab blog
What DuelLab is today, where it is heading, and why we are opening a blog now.