Reading DuelLab GameBench 2

Is a score a percentage of games solved?

No. A score compares a model with the other models currently on DuelLab. It can change when models, games, or match evidence are added.

What does DuelLab count?

The summary separates model families, model-and-reasoning-setting rows, Overall models, player programs, and games. It also separates complete from partial Overall results and ranked from provisional Reasoning Variants results.

What is the Reasoning Variants leaderboard?

It shows each tested model and reasoning setting as a separate row. Provider settings are not a standard amount of computation, so “High” for one model should not be treated as identical to “High” for another.

How does Overall ranking work?

Every model with at least one available Overall score receives an official rank. A complete result has qualifying evidence for the Baseline, Balanced, and Intensive groups. A partial result is ranked from the evidence available so far and should be compared with extra care.

How can a model be Partial 3/3?

Three of three means all three group scores are present. The result can still be partial if another requirement is missing, such as clear evidence that three distinct settings were tested or enough working player programs. It keeps its numeric rank and shows the reason for the warning.

What is a player program?

It is code written by a model to play one public game. Some downloadable files call it an “entrant”; in those files, the terms mean the same program.

If the games are repeatable, why run several matches?

Matches can change the starting side, opening position, game setup, opponent, or player program. Running these different situations gives a fairer picture than repeating one identical match.

What does Provisional mean?

The model-and-setting row does not yet have enough evidence for an official Reasoning Variants rank. The row remains visible so readers can see the result collected so far.

What do the extra columns mean?

Playable shows playing strength only when a working program was produced. Codegen shows whether programs worked on the first try, after repair, or not at all. The Players column counts the programs behind a result. Best is the highest normalized game score. Signal and the 90% band describe how strongly the evidence supports the result. Spread compares a model’s tested reasoning settings. W / D / L shows rated wins, draws, and losses.

Why can a filtered rank differ from the official rank?

A filter ranks only the rows or game currently shown. DuelLab keeps the official full-leaderboard rank visible so the temporary filtered order is not mistaken for the published rank.

How should I read estimated cost?

Std. price is an equal-weight standard-price equivalent for generating and repairing one player program. It can use billed, provider-reported, or versioned price evidence, so it may differ from an invoice. $0.00 means exactly zero; a positive amount below one cent still appears as a positive value.