Sudoku-Bench
From GRPO to GPT-5: Why Sudoku Variants remain a Grand Challenge in AI Reasoning
Oct 2025
Introduction
In May 2025, Sakana AI released the Sudoku-Bench, a collection of handcrafted sudoku puzzles that were meant to test the reasoning capabilities of LLMs, presenting itself as a grand challenge to AI reasoning. At the release of the Sudoku-Bench, reasoning models such as ChatGPT-o3 were not able to solve any of the classic 9x9 problems.
Since then, a new generation of models have been released, boasting impressive results on domains such as math, coding, and agentic tasks. But what about the Sudoku-Bench?
In this post, we are proud to present:
- Our evaluations of GPT-5 have revealed impressive results on the Sudoku-Bench, claiming the top spot of our leaderboards and having a strong lead over the previous lead, GPT-o3 mini. With a weighted average solve rate of 33% on
challenge_100, GPT-5 is also the first model that is able to solve a 9x9 modern Sudoku problem, showcasing its strong capabilities in spatial and logical reasoning. - Despite the impressive results from GPT-5, the Sudoku-Bench remains a grand challenge to AI reasoning (only 33% solved!). In this post, we present our own findings from experiments with currently popular GRPO and other methods like Thought-Cloning to study the gaps between AI and human intelligence.
Sudoku, the beloved logic puzzle that was popularized in Japan and exploded in global popularity through the 1980s and 2000s, presents a deceptively simple challenge: fill a 9×9 grid so that each row, column, and 3×3 box contains all digits 1-9. Together with the popular classic puzzles, the Sudoku-Bench includes “Modern Sudokus”, complex variants with unique constraints that can involve everything from following colored pathways to understanding abstract scenarios like guiding rats through teleporter mazes. These modern variants present an extraordinary challenge for AI reasoning systems because, unlike games with fixed rules like Chess or Go, each puzzle requires meta-reasoning to first understand entirely new rulesets before attempting to solve them. While current AI models can often comprehend these novel rules and make progress through locally consistent steps, they frequently fail at maintaining global consistency over long reasoning chains, especially when encountering the creative “break-in points” that human experts use to elegantly unlock solutions. This is precisely why Sakana AI developed Sudoku-Bench, a carefully curated benchmark ranging from simple puzzles current models can solve to impossibly complex variants that push the absolute boundaries of AI reasoning capabilities, featuring hand-crafted puzzles from Nikoli and thousands of hours of expert human reasoning data from Cracking The Cryptic.
Sudoku-Bench Leaderboard
| Model | Multi-Step | Single-Shot |
|---|---|---|
| 4x4 | 6x6 | |
| --- | --- | --- |
| ASR | ACP | ASR |
| --- | --- | --- |
| 👑GPT-5 High | 80.0% | 12.8 |
| GPT-5 Med | 66.7% | 11.3 |
| GPT-5 Low | 53.3% | 8.5 |
| O3 Mini High | 60.0% | 9.7 |
| Gemini 2.5 Pro | 73.3% | 11.6 |
| DeepSeek R1 | 60.0% | 9.5 |
(Note: A ‘-’ indicates insufficient data to meet reporting thresholds due to cost limitations.)
Models are evaluated using one of two configurations:
- Single-Shot: The LLM attempts to solve the entire puzzle grid in one response.
- Multi-Step: The LLM is prompted to provide one or more cell placements in each turn. The user displays the updated board. The interaction continues until the LLM solves the puzzle or makes an incorrect move.
The evaluation measures performance based on two primary metrics:
- Average Solve Rate (ASR): The percentage of puzzles for which the model produces the complete and correct final solution grid. This is the primary metric for overall success.
- Average Correct Placements (ACP) for Multi-Step Mode: The average number of correct cell values placed before the puzzle is solved, an incorrect placement is made, or another termination condition (such as an API error or reaching a maximum number of steps) occurs.
The benchmark includes 100 puzzles of different grid sizes (15 4x4, 15 6x6, 70 9x9).
GPT-5 versus the Sudoku-Bench
As of this post, GPT-5 overtakes the previous lead on the Sudoku-Bench to be the leader, boasting impressive solve rates on both the multi-step and single-shot setting with solve rates of 33% with high reasoning effort! This presents a two-times lead over the previous leader ChatGPT-o3-mini. Most impressively, GPT-5 is the first LLM to solve a 9x9 modern sudoku problem, Theta. Despite its impressive results, there are still shortcomings of the models. In the later sections, we present examples of outputs from GPT-5 and other models we experimented with, and share why the Sudoku-Bench is still challenging for current models to solve. Note that while querying GPT-5 via the API doesn’t reveal its reasoning traces, the models can be prompted to provide a summary of its insights. The examples presented below are, therefore, are GPT-5′s summary of the reasoning traces.
Theta, the First 9x9 Sudoku Variant, Solved!
Puzzle: Theta Author: Gerhard1963 Link: Try it on SudokuPad Model: GPT-5
Puzzle Rules:
- Standard sudoku rules apply: The digits 1 through 9 appear in every row, column, and box.
- Region Sum Line: The sum of the digits on a blue line within a particular region must be the same for all of the regions the line passes through.
- XV Pairs: Digits in cells separated by a V sum to 5. Digits in cells separated by an X sum to 10. Not all possible Xs and Vs are necessarily given.
- Roman Numeral Cages: Numbers in cages connected by a C sum to 100. Numbers in cages connected by a D sum to 500. Cages are read top-to-bottom or left-to-right.
Analysis
This represents the first 9×9 Sudoku variant solved by an LLM. GPT-5 successfully coordinates four distinct constraint types: standard Sudoku rules, region sum lines, XV pairs, and Roman numeral cages. Each operates with different logic, yet GPT-5 satisfies all simultaneously.
The Roman numeral cage interpretation is particularly sophisticated. The model understands that digits form concatenated numbers (cells containing 2 and 4 form "24", not sum to 6) and that paired cages must sum to specific values (C=100, D=500). This semantic understanding goes beyond simple arithmetic. Region sum line verification requires spatial awareness to determine box boundaries and arithmetic precision to verify sums. XV pair verification shows systematic checking of all marked positions.
As we show in other examples, GPT-5's success likely stems from its mathematical reasoning capabilities and ability to identify key break-ins.
Identifying Break-Ins
Puzzle: Clearskies Author: Michael Lefkowitz Link: Try it on SudokuPad Model: GPT-5
Puzzle Rules:
- Normal 4x4 Sudoku rules apply
- Sunburn: A digit on a sun indicates how many surrounding cells (i.e orthogonally or diagonally adjacent cells) contain smaller digits
Analysis
GPT-5's solution demonstrates sophisticated reasoning that mirrors human expert problem-solving strategies, motivated by constraints and logic rather than brute force row-by-row search. The model immediately identified a crucial break-in point by recognizing the logical impossibility of certain sun values in a 4×4 Sudoku setup, significantly constraining the puzzle from the outset. Building on this insight, GPT-5 strategically focused on the most informative sun cells and systematically explored different values through smart case testing. This approach showcases the model's ability to use high-level logical deduction to dramatically reduce the search space before switching to more granular reasoning.
Mathematical Thinking
Puzzle: Unique Sum Sudoku Author: I Love Sleeping & Myxo Link: Try it on SudokuPad Model: GPT-5
Puzzle Rules:
- Normal Sudoku rules do NOT apply
- Fill the grid with digits 1-9, such that no digit repeats in any row, column, or box
- The set of digits in each row or column is unique, e.g. if a row contains the digits 1234, no other row or column many contain exactly those digits
- The digits in every row, column and box sum to x, where x has to be determined by the solver
- Digits separated by an X sum to 10. Digits separated by a V sum to 5. Not all Xs and Vs are necessarily given.
Analysis
GPT-5's approach to this algebraic Sudoku variant reveals a significant departure from previous generations of frontier models, showcasing sophisticated mathematical thinking that extends far beyond typical puzzle-solving strategies.
The model employs rigorous set notation throughout its analysis, maintaining possible values and systematically tracking constraints to create a form of working memory it can reference and build upon. This mathematical framework becomes particularly evident in GPT-5's extensive use of algebraic arithmetic, where it derives complex relationships to establish interconnected constraint systems across the entire grid. The model's ability to recognize parity relationships and systematically branch on variable values demonstrates computational reasoning that mirrors advanced mathematical problem-solving techniques. This mathematical sophistication likely stems from the puzzle's inherent algebraic nature, given the unique row, column, box rule, along with arithmetic constraints from X and V symbols. These algebraic requirements appear to have primed GPT-5 to leverage the same mathematical reasoning skills that contribute to its strong performance on other math benchmarks, suggesting that when puzzles align with a model's trained mathematical competencies, it can achieve remarkably sophisticated logical reasoning.
Spatial Reasoning is still difficult!
Puzzle: Secret Kingdoms Author: Senator Gronk Link: Try it on SudokuPad Model: GPT-5
Puzzle Rules:
- Divide the grid into 9 kingdoms of orthogonally connected cells. Fill the grid with digits 1-9 so that digits don't repeat in a row, column or kingdom
- A digit in a blue square indicates its kingdom's strength. Each kingdom has a different strength, which is determined by summing the digits in its castle(s)
- Digits connected by bridges are in different kingdoms and for each their value is the strength of the kingdom on the other side of the bridge. Two kingdom borders are provided
Analysis
GPT-5's performance on this spatial reasoning puzzle reveals the limitations of what appears to be a very strong Samurai that has trained exceptionally well with a single weapon but struggles when other tools are required.
While the model excelled at mathematical and algebraic reasoning in previous examples, it seems unable to effectively tackle problems that demand primarily spatial reasoning over mathematical computation. This challenge is admittedly extreme, departing significantly from standard Sudoku conventions by lifting the familiar 3×3 box constraints and instead requiring solvers to determine arbitrary, non-consistent kingdom boundaries where each kingdom must contain unique numbers.
The model does identify some interesting, albeit insufficient, break-in points about the relationship between blue squares and castles. However, much of its reasoning output reads more like an academic analysis or tutorial than active problem-solving, with generic statements sounding distinctly textbook-like rather than reflecting real-time puzzle engagement. Unlike previous examples where GPT-5 immediately identified the most constrained and informative regions to focus on, here the model shows no spatial intuition about where to begin or how kingdom boundaries might naturally form.
GRPO Multi-Turn: Limited Working Memory
Puzzle: Standard 9×9 Sudoku Model: Qwen-7bi + GRPO
Puzzle Rules:
- Standard Sudoku rules
- 9×9 grid with 3×3 boxes
- Each row, column, and box contains digits 1-9
Analysis
The multi-turn dialogue reveals fundamental limitations that contrast sharply with GPT-5's sophisticated constraint management.
The model is unable to handle the 3 Sudoku constraints simultaneously. In its reasoning, it analyzes only the row constraints, completely ignoring column and box constraints. This leads to catastrophic errors: the top-left 3×3 box contains three instances of the digit 2 (in r1c2, r2c1, and r3c1), violating both column and box uniqueness rules. Evidently, the model is not able to handle the simultaneous constraints that define Sudoku.
The model's spatial reasoning is critically flawed. It claims to analyze "the second row and the top-middle 3×3 grid" to make deductions about r2c7, but r2c7 is not in the top-middle box at all. This represents a basic failure to understand Sudoku's spatial structure.
Conclusion
We have discussed our evaluation of GPT-5 and ran experiments training smaller models with modern techniques on the Sudoku-Bench. Once again, we congratulate the OpenAI team for their breakthroughs on the Sudoku-Bench. Nevertheless, more research needs to be done to bridge the gap between human thinking and AI reasoning processes, as our experiments here have shown that even advanced training methodologies like GRPO and thought cloning face fundamental limitations when applied to Sudoku. While GPT-5 demonstrated impressive mathematical reasoning capabilities and human-like strategic thinking on algebraically-constrained puzzles, it struggled significantly with spatial reasoning challenges that require spatial understanding. Our smaller model experiments revealed that current fine-tuning approaches often lead to superficial pattern matching rather than genuine logical reasoning development. The Sudoku-Bench continues to expose critical gaps between computational problem-solving and authentic human-like reasoning, particularly in tasks that demand the seamless integration of mathematical logic, spatial awareness, and creative insight.