Classic sudoku has become too easy a test for artificial intelligence: a simple computer program can solve any valid 9x9 grid in a fraction of a second, and many well-known puzzles are already in the training data of large language models. Tokyo-based research company Sakana AI therefore turned to modern variant sudoku, where extra rules such as arrows, thermometers, killer cages or unusual grid shapes force the solver to find a fresh logical "break-in" for every puzzle.
What Sudoku-Bench is
Sakana AI first announced the project on 21 March 2025 and published the technical report on arXiv on 22 May 2025. The paper, "Sudoku-Bench: Evaluating creative reasoning with Sudoku variants", is by Jeffrey Seely, Yuki Imajuku, Tianyu Zhao, Edoardo Cetin and Llion Jones. The core test set, called challenge_100, has 100 puzzles: 15 on 4x4 grids, 15 on 6x6 grids and 70 on 9x9 grids, arranged in a difficulty ramp from easy to extremely hard. Puzzles are given to the models in a standardised text format, and the accompanying tools work with thousands of publicly available puzzles.
The project was built in partnership with Cracking the Cryptic, the YouTube channel run by Simon Anthony and Mark Goodliffe, both of whom have represented the UK at the World Sudoku and World Puzzle Championships. Sakana also released a large dataset taken from the channel's videos:
- more than 2,500 videos of puzzle solving;
- over 2,000 hours of spoken reasoning transcribed into text, roughly 10 million words;
- around 2 million solving actions (digit entries, pencil marks and so on).
The idea is that a model could learn not only answers, but the way expert humans think their way into a hard puzzle. As Sakana's page puts it, "The hardest puzzles in this benchmark are extremely difficult even for professional puzzle solvers."
Results so far
In the May 2025 paper, state-of-the-art language models solved fewer than 15% of the puzzles without help, and performance fell sharply on the larger 9x9 grids. Sakana noted that at release, reasoning models could not solve any of the 9x9 puzzles in the set.
In October 2025, Sakana published an update: OpenAI's GPT-5 reached a weighted average solve rate of 33% on challenge_100, about twice the previous leader, and became the first language model to solve one of the benchmark's modern 9x9 sudokus. Sakana stressed that this still leaves most of the benchmark unsolved. The same update reported that fine-tuning a small open model on ordinary sudoku did not transfer well to the variants, which suggests that practising classic grids alone is not enough.
Why it matters for puzzle fans
Sudoku-Bench confirms something experienced solvers know: the hardest part of a good puzzle is spotting the key idea, not filling in digits. Variant setters design puzzles around a single elegant insight, and that kind of reasoning is still hard to automate. It also shows that the solving videos many fans watch for entertainment have become valuable research data.
If you want to build the skills that make variants approachable, start with classic logic. Practise scanning and candidate techniques on our sudoku puzzles, and use the SE rating guide to see which techniques make a puzzle harder. Note that AI results change quickly, so the figures above describe the benchmark as reported in 2025.