O

Ozerlyn Editorial Team

Boulder, Colorado, USAOctober 11, 2026 at 02:09 AM4 min
Research

AI can solve some sudokus, but CU Boulder study finds it can't explain how

A Findings of ACL 2025 paper tested five large language models on about 2,300 six-by-six sudokus and found the explanations were the weakest link.

Researchers at the University of Colorado Boulder asked five large language models to solve and explain 6x6 sudoku puzzles. The best model solved roughly 65% of them, but none produced explanations that reflected real strategic reasoning. For puzzle fans, the study is a useful reminder that a correct grid and a clear solving path are two very different things.

On 28 July 2025, the University of Colorado Boulder publicised a study that used sudoku as a test of whether artificial intelligence can be trusted to explain its own decisions. The paper, "Explaining Puzzle Solutions in Natural Language: An Exploratory Study on 6x6 Sudoku", was written by Anirudh Maiya, Razan Alghamdi, Maria Leonor Pacheco, Ashutosh Trivedi and Fabio Somenzi. It was accepted to Findings of the Association for Computational Linguistics (ACL 2025), one of the main venues for natural language processing research.

What the researchers did

The team created nearly 2,300 original sudoku puzzles on a six-by-six grid, a smaller version of the classic 9x9 board. Using fresh puzzles matters: well-known puzzles and their solutions may already appear in a model's training data, so a new set makes it harder for a model to simply recall an answer. Five large language models (LLMs) were then asked to do two things: fill in the grid, and explain in plain language how they reached the solution.

Key findings

  • Solving was uneven. According to CU Boulder, OpenAI's o1 model (preview version) led the field and solved roughly 65% of the puzzles correctly. The paper's abstract notes that only one model showed even limited success.
  • Explaining was much worse. According to the paper, none of the models could explain the solution in a way that reflected strategic reasoning or intuitive problem-solving.
  • Some explanations were inaccurate or simply strange. Trivedi told CU Boulder that the AI sometimes made up facts, and in one case a model answered a sudoku question with a weather forecast.

Somenzi summed up why the team chose this game: "Puzzles are fun, but they're also a microcosm for studying the decision-making process in machine learning."

Why sudoku is a good test

Sudoku has a single correct answer that can be checked instantly, and an experienced human can justify every placement with a named technique: a naked single, a hidden pair, an X-Wing and so on. That makes it easy to see whether an explanation is actually true. A language model, by contrast, produces text that sounds plausible word by word. It can arrive at a valid grid while describing steps that never happened, or describe a technique that does not apply to the cells in question.

The researchers frame this as a wider trust problem. If an AI system helps with scheduling, logistics or other decisions, people need to understand why it chose a given answer, not just accept the answer. Sudoku offers a small, controlled setting in which those explanations can be checked line by line. The team said its next step is a system that both solves and explains more complicated puzzles, starting with the Japanese logic puzzle hitori.

What it means for players

For human solvers, the study points to the skill that machines still lack: being able to say why a digit must go in a cell. That is also how puzzle difficulty is usually measured. The Sudoku Explainer (SE) rating scores a puzzle by the hardest technique needed to solve it logically, so a clear chain of reasoning is the heart of the hobby. You can read how the scale works on our SE rating page, or practise explaining each move to yourself on the puzzles in our sudoku section.

One caution: the study tested models available in 2024–2025 on 6x6 grids. AI systems are improving quickly, and later results may differ. Still, the core message holds: getting the right answer is not the same as showing your work.

Sources

Share this article

X