AI chess leaderboard

Watch a game

Complete games. Compare playing strength across the field.

32 models · 7 providersUpdated

Leaderboard

ChessBench model results. Select a model for details. Benchmark Elo measures field-relative strength. Cost per Task is in US dollars; a dash means not reported.
Rank Model / provider

Higher Elo is better. Costs in USD per task; — = not reported.

How to read these results

What the numbers mean

A field-relative measure of model performance, not a human chess rating.

Complete games
Models play full games. Planning, tactical calculation, recovery, and consistency all affect the result.
Benchmark Elo
A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.
Cost per Task
The reported cost in US dollars per task. A dash (—) means no cost has been reported, not that the task was free. Select the column heading to sort by cost; missing values remain at the end.

ChessBench TV

A complete game, one move at a time.

Archive · Match 001
Gemini 3.1 ProGoogleBlack
GPT-5.6 Sol ProOpenAIWhite to move
0 / 59

Focus the replay to use to step and Space to play.

GPT-5.6 Sol Pro
vs. Gemini 3.1 Pro

Finished · Black resigned

1–0
Starting position · White to move

Benchmark updates

New results, corrections, and product notes.

Gemini 3.8 Flash (medium) joins ChessBench

Gemini 3.8 Flash (medium) has been added with 2,300 benchmark Elo, 30.74 ACPL, and a cost of $0.08/task. It enters at rank 13, after Gemini 3.6 Flash, which also has 2,300 benchmark Elo, and ahead of GPT-4.5.

View leaderboard

Claude Fable 5.1 (Max) joins ChessBench; Cost per Task added

Claude Fable 5.1 (Max) has been added as a separate model with 2,370 benchmark Elo, 18.54 ACPL, and a cost of $44.46/task. It enters at rank 10, between Gemma 4 and GPT-6 Pro (Max).

The leaderboard now includes Cost per Task. Costs not yet reported are shown as a dash (—). ACPL and cost are also available in model details.

View leaderboard

GPT-6 Pro (Max) joins ChessBench

GPT-6 Pro (Max) has been added to ChessBench with a benchmark Elo of 2,340. It enters at rank 10, between Gemma 4 and Gemini 3.6 Flash.

The leaderboard now includes 30 models across seven providers.

View leaderboard

Gemini 3.7 Flash (High) joins ChessBench

Gemini 3.7 Flash (High) has been added to ChessBench with a benchmark ELO of 2,230 It enters at rank 17, between MiniMax M3 and ChatGPT 5.4 Pro.

The leaderboard consolidates GPT-5.6 Sol Pro (Max) and its Stealth Testing checkpoint into a single rank-one entry, with the checkpoint’s 2,650 ELO shown as secondary information. Claude Fable 5 (Max) therefore moves to rank 2.

View leaderboard

Grok 4.6 (xhigh) joins ChessBench

Grok 4.6 (xhigh) has been added to ChessBench with a benchmark ELO of 2,480 The result places it at rank 7, between Qwen 3.8 Max and Claude Opus 5 (high), and makes it the strongest xAI model in the current field.

View leaderboard

Qwen 3.8 Max joins ChessBench

Qwen 3.8 Max joins with 2,500 benchmark ELO. It places at rank 6, between Gemini 3.1 Pro (High) and Grok 4.6 (xhigh), and adds Alibaba to the provider field.

View leaderboard

Claude Opus 5 (high) joins ChessBench

Claude Opus 5 (high) enters at 2,400 benchmark ELO. Compared to the original Opus 5 configuration, this represents a substantial improvement in playing strength, moving the entry closer to the current frontier.

View leaderboard

Gemini Flash additions and a Grok 4.5 correction

Gemini 3.6 Flash joins at 2,300 benchmark ELO, alongside Gemini 3.5 Flash-Lite at 2,200 benchmark ELO. Flash-Lite’s benchmark ELO is lower because illegal moves and several quick terminations affected game outcomes.

Grok 4.5 was rebenchmarked after errors were found in the original evaluation. Its benchmark ELO changes from 1,457 to 1,680.

View leaderboard

Claude Fable 5 (Max) and Kimi K3 (Max) join ChessBench

Claude Fable 5 (Max) from Anthropic and Kimi K3 (Max) from Moonshot AI join the leading group. Claude Fable 5 (Max) holds rank 2 with 2,640 benchmark ELO, while Kimi K3 (Max) holds rank 4 with 2,590 benchmark ELO.

The additions broaden the provider mix near the top of the table and add two high-strength reference points to the leaderboard.

View leaderboard

Introducing ChessBench TV

ChessBench TV replays complete model-vs-model games on an interactive board, with the full move score and playback controls. The first replay features GPT-5.6 Sol Pro against Gemini 3.1 Pro.

Watch Match 001

Benchmark Elo
ACPL
Cost per Task

Elo is relative to the ChessBench field.