Benchmark menu

Normalized score across public games (0–100)

Showing top 24 of 95 benchmarked reasoning variants (updates when chart loads)

Scale: relative 0-100

  1. 87.3
    GPT-5.6 Sol
  2. 83.3
    Claude Fable 5
  3. 76.0
    Claude Opus 5
  4. 73.5
    Claude Opus 5
  5. 72.7
    GPT-5.4
  6. 68.5
    Claude Opus 5
  7. 67.5
    GPT-5.5
  8. 66.6
    Claude Opus 4.8
  9. 66.5
    GPT-5.6 Sol
  10. 66.1
    Claude Fable 5
  11. 63.2
    Gemini 3.6 Flash
  12. 62.1
    Grok 4.5
  13. 57.5
    GPT-5.6 Terra
  14. 57.2
    Grok 4.5
  15. 55.8
    Gemini 3.6 Flash
  16. 55.8
    GPT-5.6 Sol
  17. 55.4
    Gemini 3.5 Flash
  18. 55.0
    GPT-5.4
  19. 54.3
    Claude Sonnet 5
  20. 53.4
    Kimi K3
  21. 53.1
    GPT-5.5
  22. 52.8
    Muse Spark 1.2
  23. 52.8
    GPT-5.6 Luna
  24. 49.0
    Claude Opus 4.5
View
Reasoning

Detailed leaderboard

How to read

We ask each AI model to write a player program for every game. DuelLab runs those programs against each other. Higher scores mean stronger, better-supported performance against the current models.

Score is the normalized 0–100 result across public games. Playable averages only games where the model produced a working player program. Codegen shows the share that worked first try / worked after repair or resend / failed. Players counts the player programs behind a row, while Best is its highest single-game score. W / D / L shows rated wins, draws, and losses. The 90% band, Signal, and reasoning spread help show how much evidence supports the score and how much results vary.

Partial n/3 is an official Overall rank with incomplete evidence; n/3 shows how many reasoning settings have scores, so compare it with care. Provisional reasoning variants do not receive an official rank. Under-covered means evidence is sparse. Capped draws reached the shared move limit and still count as draws. Filters can change the visible order, but the Official or Base column keeps the published rank. Std. price is an estimate; hover over a value for its source and pricing details.

Columns

Choose the optional columns shown in this table.

Reasoning Variants leaderboard for DuelLab Benchmark
Rank Model Reasoning Score Playable W / D / L Best Std. price Codegen By game
1GPT-5.6 SolXHigh87.387.3585 / 31 / 70100 × 2$0.80List-price estimate88%12%0%n 8
2Claude Fable 5XHigh83.383.3441 / 72 / 91100 × 2$1.58List-price estimate63%37%0%n 8
3Claude Opus 5XHigh76.080.0586 / 52 / 140100$1.38List-price estimate81%0%19%n 16
4Claude Opus 5Medium73.576.0785 / 160 / 18790.2$0.45List-price estimate88%0%12%n 16
5GPT-5.4XHigh72.775.2481 / 84 / 7495.1$1.52List-price estimate88%0%12%n 8
6Claude Opus 5None68.570.8809 / 261 / 21591.0$0.29List-price estimate88%0%12%n 16
7GPT-5.5XHigh67.569.8322 / 107 / 11593.8$2.74List-price estimate88%0%12%n 8
8Claude Opus 4.8XHigh66.668.9355 / 71 / 14679.5$0.93List-price estimate75%13%12%n 8
9GPT-5.6 SolMedium66.566.5443 / 98 / 14990.1$0.23List-price estimate100%0%0%n 8
10Claude Fable 5Medium66.168.3350 / 57 / 11777.4$0.50List-price estimate88%0%12%n 8
11Gemini 3.6 FlashHigh63.267.91003 / 285 / 43691.6$0.48Provider-reported50%25%25%n 16
12Grok 4.5High62.162.1660 / 233 / 32791.3$0.07Provider-reported75%25%0%n 8
13GPT-5.6 TerraXHigh57.557.5280 / 137 / 18576.4$8.17List-price estimate63%37%0%n 8
14Grok 4.5Medium57.257.2366 / 170 / 20887.2$0.06Provider-reported88%12%0%n 8
15Gemini 3.6 FlashMedium55.857.7942 / 443 / 42890.6$0.26Provider-reported88%0%12%n 16
16GPT-5.6 SolNone55.855.8365 / 161 / 17278.7$0.13List-price estimate100%0%0%n 8
17Gemini 3.5 FlashMedium55.455.4285 / 195 / 15684.5$0.35Provider-reported0%100%0%n 8
18GPT-5.4Medium55.055.0349 / 211 / 170100$0.28List-price estimate100%0%0%n 8
19Claude Sonnet 5XHigh54.354.3326 / 160 / 19883.6$0.35List-price estimate100%0%0%n 8
20Kimi K3Low53.453.4581 / 396 / 36198.9$0.08Provider-reported88%12%0%n 8
21GPT-5.5Medium53.153.1265 / 224 / 18773.9$0.44List-price estimate88%12%0%n 8
22Muse Spark 1.2Medium52.854.6306 / 135 / 15679.7$0.05Provider-reported88%0%12%n 8
23GPT-5.6 LunaXHigh52.852.8306 / 200 / 17583.1$0.03List-price estimate100%0%0%n 8
24Claude Opus 4.5None49.049.0254 / 237 / 18770.1$0.12List-price estimate100%0%0%n 8
25Kimi K3High47.647.6416 / 376 / 38568.5$0.29Provider-reported88%12%0%n 8
26Claude Opus 4.8None47.549.2246 / 147 / 16173.8$0.16List-price estimate88%0%12%n 8
27Claude Opus 4.8Medium47.447.4238 / 266 / 16680.1$0.22List-price estimate100%0%0%n 8
28Gemini 3.5 FlashHigh47.350.8165 / 133 / 16880.8$0.49Provider-reported0%75%25%n 8
29GPT-5.4 MiniMedium46.846.8287 / 219 / 21758.3$0.25List-price estimate88%12%0%n 8
30O3Medium46.750.2238 / 143 / 18562.6$0.05List-price estimate75%0%25%n 8
31Grok 4.5Low46.446.4470 / 383 / 37285.5$0.07Provider-reported88%12%0%n 8
32DeepSeek V4 ProXHigh46.349.8188 / 125 / 20177.2$0.08Provider-reported13%62%25%n 8
33GPT 5High46.246.2217 / 193 / 23067.7$0.28List-price estimate75%25%0%n 8
34Ox AlphaMax44.344.3191 / 61 / 8152.1$0.00Provider-reported88%12%0%n 25
35Qwen3.7 PlusEnabled43.546.7202 / 176 / 18872.5$0.05Provider-reported13%62%25%n 8
36GPT-5.4None43.543.5264 / 216 / 24075.2$0.10List-price estimate75%25%0%n 8
37GPT-5.6 TerraMedium43.443.4229 / 259 / 21260.6$0.17List-price estimate63%37%0%n 8
38GLM-5.2Medium43.243.2161 / 195 / 13866.3$0.07Provider-reported0%100%0%n 8
39O3High42.848.81295 / 282 / 122456.0$0.09Combined: Recorded attempt + Standard list-price estimate59%0%41%n 32
40GPT 5Medium42.844.3200 / 156 / 20462.9$0.18List-price estimate75%13%12%n 8
41GLM-5.2None42.842.8325 / 327 / 34659.6$0.02Provider-reported56%44%0%n 16
42Qwen3.7 MaxEnabled42.744.1209 / 240 / 20379.1$0.09Provider-reported13%75%12%n 8
43Kimi K2.7 CodeEnabled42.445.6216 / 163 / 19363.1$0.16Provider-reported0%75%25%n 8
44Qwen3.7 MaxNone42.143.5184 / 190 / 19468.9$0.02Provider-reported88%0%12%n 8
45GPT-5.6 TerraNone41.943.2189 / 265 / 20655.2$0.07List-price estimate78%11%11%n 9
46GPT 4.1None41.541.5199 / 228 / 22463.3$0.06List-price estimate63%37%0%n 8
47Claude Sonnet 5None41.241.2213 / 190 / 27257.1$0.06List-price estimate100%0%0%n 8
48Claude Sonnet 5Medium40.940.9209 / 200 / 26652.4$0.07List-price estimate100%0%0%n 8
49DeepSeek V4 ProNone40.842.1200 / 207 / 18766.6$0.02Provider-reported50%38%12%n 8
50Gemini 3.5 Flash LiteMinimal40.641.2662 / 730 / 84573.0$0.02Provider-reported56%38%6%n 16
51Muse Spark 1.2Minimal40.240.2226 / 163 / 29959.4$0.02Provider-reported88%12%0%n 8
52GPT-5.6 LunaMedium40.240.2203 / 217 / 26861.4$0.0063List-price estimate88%12%0%n 8
53GPT-5.4 NanoMedium39.742.6410 / 200 / 41663.1$0.02List-price estimate31%44%25%n 16
54Grok Build 0.1Enabled39.539.5209 / 276 / 24168.0$0.06Provider-reported0%100%0%n 8
55GPT-5.5None39.240.5137 / 318 / 15754.0$0.14List-price estimate88%0%12%n 8
56DeepSeek V4 FlashXHigh38.741.6153 / 135 / 18672.3$0.0100Provider-reported38%37%25%n 8
57Laguna S 2.1None38.039.5378 / 423 / 40857.9$0.0030Provider-reported50%36%14%n 14
58Gemini 3.6 FlashMinimal37.837.8670 / 861 / 91655.0$0.06Provider-reported75%25%0%n 16
59Minimax M3None37.640.4207 / 119 / 19852.0$0.03Provider-reported50%25%25%n 8
60Hy3 PreviewNone37.240.0135 / 212 / 12572.9$0.0013Provider-reported50%25%25%n 8
61MiMo-V2.5Reasoning36.339.9272 / 254 / 28452.1$0.0024Provider-reported44%25%31%n 16
62GPT-5.4 NanoNone36.138.7321 / 307 / 38656.1$0.0092List-price estimate31%44%25%n 16
63GPT-OSS 120BMedium35.636.8134 / 284 / 20346.6$0.0025Provider-reported13%75%12%n 8
64MiMo-V2.5Reasoning35.538.2172 / 116 / 22864.9$0.0038Provider-reported63%12%25%n 8
65Gemini 3.5 Flash LiteHigh35.436.6542 / 435 / 81254.4$0.06Provider-reported81%6%13%n 16
66Hy3 PreviewHigh35.037.2174 / 134 / 24459.3$0.0032Provider-reported56%22%22%n 9
67GPT-OSS 120BHigh34.839.190 / 151 / 12767.3$0.0088Provider-reported0%63%37%n 8
68Nemotron 3 Ultra 550B A55BNone34.636.9118 / 149 / 18359.7$0.03Provider-reported56%22%22%n 9
69InklingNone34.036.6315 / 338 / 46569.0$0.04Provider-reported44%31%25%n 16
70Ox AlphaHigh33.935.295 / 80 / 13342.8$0.00Provider-reported70%17%13%n 30
71LongCat 2.0None33.935.0347 / 471 / 49055.2$0.01Provider-reported50%38%12%n 16
72North Mini CodeEnabled33.636.1168 / 170 / 22850.4$0.00Provider-reported38%37%25%n 8
73MiMo-V2.5Reasoning33.535.7210 / 188 / 27847.5$0.02Provider-reported0%78%22%n 9
74DeepSeek V4 FlashMedium32.133.2219 / 328 / 38145.2$0.0032Provider-reported63%25%12%n 16
75MiMo-V2.5Medium31.835.2133 / 298 / 23155.9$0.02Provider-reported9%58%33%n 12
76Nemotron 3 Ultra 550B A55BHigh31.336.1153 / 303 / 25946.2$0.03Provider-reported12%44%44%n 16
77LongCat 2.0Enabled31.133.4489 / 425 / 87345.1$0.09Provider-reported56%19%25%n 16
78MiMo-V2.5High31.034.845 / 203 / 7458.6$0.02Provider-reported0%63%37%n 8
79Gemini 3.1 Flash LiteHigh30.634.487 / 120 / 22046.4$0.03Provider-reported0%63%37%n 8
80GPT-5.4 MiniNone30.630.6120 / 298 / 24043.8$0.02List-price estimate75%25%0%n 8
81Ox AlphaLow30.631.183 / 83 / 15545.1$0.00Provider-reported63%30%7%n 30
82Nemotron 3 Ultra 550B A55BMedium29.834.5112 / 162 / 15061.9$0.04Provider-reported0%56%44%n 9
83Gemini 3.1 Flash LiteMedium29.531.7149 / 141 / 28445.6$0.01Provider-reported25%50%25%n 8
84DeepSeek V4 FlashNone29.230.2122 / 226 / 22444.2$0.0023Provider-reported63%25%12%n 8
85Step 3.7 FlashHigh29.132.7134 / 71 / 22246.6$0.06Provider-reported0%63%37%n 8
86Laguna S 2.1Enabled28.330.1309 / 691 / 70939.4$0.0034Provider-reported50%29%21%n 14
87Minimax M3Enabled28.028.9140 / 211 / 30747.8$0.07Provider-reported0%88%12%n 8
88Nex N2 ProEnabled27.831.2116 / 51 / 25748.6$0.10Provider-reported0%63%37%n 8
89InklingMax27.431.6259 / 396 / 64760.5$0.26Provider-reported25%31%44%n 16
90Gemini 3.5 Flash LiteMedium27.028.0377 / 680 / 99540.4$0.04Provider-reported88%0%12%n 16
91Ling 3.0 FlashNone25.628.7191 / 607 / 40840.6$0.00Provider-reported38%25%37%n 16
92Qwen3.7 PlusNone25.227.9110 / 160 / 23548.5$0.0087Provider-reported45%22%33%n 9
93GPT-5.6 LunaNone20.520.588 / 187 / 41437.1$0.0046List-price estimate100%0%0%n 8
94Gemini 3.1 Flash LiteNone16.318.466 / 80 / 20041.3$0.0059Provider-reported38%25%37%n 8
95GPT-OSS 120BLow14.515.0154 / 230 / 71025.7$0.0016Provider-reported63%25%12%n 8
PGLM-5.2 ProvisionalXHigh53.353.395 / 79 / 6453.3$0.09Provider-reported100%0%0%n 2
PClaude Opus 4.5 ProvisionalMedium52.352.3562 / 324 / 38984.3$0.18List-price estimate100%0%0%n 8
PMuse Spark 1.2 ProvisionalXHigh52.052.089 / 19 / 7053.3$0.12Provider-reported100%0%0%n 2
PKimi K3 ProvisionalMax48.953.2322 / 219 / 17381.0$0.53Provider-reported14%57%29%n 7
PClaude Opus 4.5 ProvisionalHistorical45.746.5322 / 337 / 30561.7$0.26List-price estimate88%6%6%n 16
PClaude Opus 4.5 ProvisionalHigh43.043.0458 / 351 / 44765.2$0.59List-price estimate100%0%0%n 8
PMiMo-V2.5-Pro ProvisionalEnabled43.046.2168 / 110 / 18669.0$0.13Provider-reported0%75%25%n 8
PQwen3.8 Max ProvisionalXHigh42.444.864 / 36 / 2968.7$0.35Provider-reported80%0%20%n 10
PQwen3.8 Max ProvisionalMedium40.140.159 / 44 / 4155.3$0.10Provider-reported100%0%0%n 8
PGPT 4.1 Nano ProvisionalReasoning39.739.760 / 35 / 6739.7$0.0023Combined: Recorded attempt + Standard list-price estimate100%0%0%n 1
PMiMo-V2.5-Pro ProvisionalNone39.140.7199 / 122 / 18353.0$0.03Provider-reported57%29%14%n 7
PGemma 4 31B ProvisionalEnabled38.043.8155 / 66 / 17159.5$0.0066Provider-reported0%57%43%n 7
PInkling Small ProvisionalMax35.237.340 / 44 / 3355.4$0.04Provider-reported60%20%20%n 10
PQwen3.8 Max ProvisionalMinimal34.734.739 / 52 / 4565.3$0.12Provider-reported100%0%0%n 8
PStep 3.7 Flash ProvisionalMedium32.338.492 / 137 / 13951.5$0.04Provider-reported0%50%50%n 8
PInkling Small ProvisionalNone31.932.930 / 50 / 6142.2$0.0076Provider-reported56%33%11%n 9
PInkling Small ProvisionalMedium31.932.830 / 54 / 6247.5$0.0085Provider-reported45%44%11%n 9
PNex N2 Pro ProvisionalNone31.438.6141 / 261 / 17463.5$0.16Provider-reported0%44%56%n 16
PGemma 4 31B ProvisionalNone31.132.4155 / 168 / 23544.0$0.0024Provider-reported72%14%14%n 7
PStep 3.7 Flash ProvisionalLow31.036.9199 / 110 / 32152.9$0.02Provider-reported38%12%50%n 8
PMistral Medium 3.5 ProvisionalNone30.736.5117 / 104 / 15540.9$0.04Provider-reported25%25%50%n 8
PInkling ProvisionalMedium24.731.5172 / 237 / 38457.4$0.08Provider-reported31%6%63%n 16
PO3 ProvisionalLow22.827.2116 / 172 / 35154.0$0.11Combined: Recorded attempt + Standard list-price estimate50%0%50%n 8
PLing 3.0 Flash ProvisionalEnabled21.827.4227 / 254 / 49343.0$0.00Provider-reported7%33%60%n 15
PNorth Mini Code ProvisionalNone20.427.378 / 182 / 23233.9$0.00Provider-reported19%12%69%n 16
PMistral Medium 3.5 ProvisionalHigh17.821.875 / 17 / 22631.1$0.19Provider-reported22%22%56%n 9

How this is scored

1. Generate players

Each model is asked to create a program that can play every benchmark game.

2. Play matches

The generated programs compete head-to-head, with both players receiving comparable opportunities.

3. Score each game

Results and how certain they are produce a score from 0 to 100 for each game.

4. Combine games

Game scores are combined into the main leaderboard score. A known failure to create a usable player contributes zero.

5. Keep settings clear

Each model and reasoning setting keeps its own row in the standings.

6. Add context

Detailed tables provide uncertainty, match records, and other clues for careful comparison.