Complete games
Models are evaluated over full games, so planning, tactical calculation, recovery and consistency all affect the result.
Models play complete games. Stockfish evaluates every decision. ChessBench turns those results into a field-relative rating that is easy to inspect and compare.
Benchmark ELO determines the field rank and compares tested strength, while ACPL measures move accuracy; lower values indicate cleaner play.
| Rank | Model | Benchmark ELO | ACPL |
|---|
Change the search term or return to all providers.
The strongest models combine a high benchmark rating with a low ACPL. The chart uses a logarithmic precision scale so the full field remains readable.
ChessBench focuses on one demanding environment rather than combining unrelated tasks into a single score.
Models are evaluated over full games, so planning, tactical calculation, recovery and consistency all affect the result.
Stockfish reviews each decision. Average centipawn loss quantifies how much value a model gives away compared with stronger engine choices.
Benchmark ELO is calculated within the ChessBench field. It describes performance among these models and is not presented as a direct human chess rating.
Product notes, benchmark releases and new ways to inspect how frontier models play chess.
ChessBench now includes Gemini 3.6 Flash at 2,300 benchmark ELO and 24.81 ACPL, alongside Gemini 3.5 Flash-Lite at 2,200 benchmark ELO and 16.40 ACPL. Flash-Lite records the lower ACPL, but its benchmark ELO is lower because it produced many illegal moves and several games ended quickly.
Grok 4.5 has also been rebenchmarked after errors were found in the original evaluation. The corrected run moves its benchmark ELO from 1,457 to 1,680 and reduces ACPL from 125.26 to 53.00. The results have improved, but are still far from the frontier.
View the updated leaderboardChessBench has expanded its field with Claude Fable 5 (Max) from Anthropic and Kimi K3 (Max) from Moonshot AI. Both models debut in the leading group, combining high playing strength with exceptionally low move-error rates. Claude Fable 5 (Max) enters at rank 3 with 2,640 benchmark ELO and 4.01 ACPL, while Kimi K3 (Max) opens at rank 5 with 2,590 benchmark ELO and 5.35 ACPL. Fable’s result is especially notable: it is an unexpectedly strong ChessBench showing for an Anthropic model and places substantially ahead of the other Anthropic entries in the current field. Kimi’s debut is similarly convincing, establishing Moonshot AI as an immediate top-five contender. Together, the two additions broaden the provider mix near the top of the table and add two new high-strength, high-precision reference points to the performance map.
View the updated leaderboardChessBench now includes a dedicated TV view for complete model-vs-model games. The first broadcast replays GPT-5.6 Sol Pro against Gemini 3.1 Pro on an interactive board, with the full move score, playback controls and move-by-move context.
Watch Match 001