Chain-of-Thought has long been treated as a readable, auditable "white box": write out the model's intermediate steps and let a human or monitor read along. Anthropic and OpenAI both lean on CoT as one of the most important monitoring windows for safety alignment. On September 29, a new paper put a knife into that assumption. Subbarao Kambhampati, Paarth Iyer and co-authors argue that in a reasoning chain, a correct answer does not mean a correct derivation.
A controlled stage: iGSM
The authors pick a synthetic grade-school math benchmark, iGSM, that depends on no external world knowledge and can be verified step by step automatically. Each problem exposes which quantities and dependencies a correct solution must use, which means generated chains of thought can be checked programmatically, one step at a time. In theory, that is the most direct referee for the question of whether the chain is actually pushing the next step forward.
Three findings
In-distribution agreement, out-of-distribution decoupling. The authors first train models on minimal, valid chains only. Inside the training distribution, answer correctness and chain validity almost coincide. Push the hardest out-of-distribution problems at them and the two metrics decouple: on the toughest slice, 31.6% of correct answers are accompanied by invalid derivations. Worse, more than half of those 31.6% pass syntactic and arithmetic checks but fail semantic dependency checks. The model quietly swaps a key premise at one step, keeps calculating, and lands on a correct answer.
Training supervision gets steered. A control experiment shows that training on non-minimal (verbose) chains produces verbose outputs. Asking the same problem with a different prompt reveals that the model inherits certain computations from the original query. The intuition that "minimal equals trustworthy" is refuted.
Shuffling 10% of training traces still works. Even more striking: if you shuffle the order of 10% of the trace sentences in the training data, out-of-distribution accuracy barely moves, but every resulting trace fails verification. In other words, the model does not care whether its trace is intelligible to a human reader, as long as the final answer is right. The authors go further and swap entire trace blocks; in-distribution accuracy also holds up.
What this means for safety monitoring
In the past six months there has been a major industry discussion about whether CoT monitoring can be the next moat for AI safety (Korbak et al.'s "Chain of Thought Monitorability" offered the optimistic framework). Kambhampati's result cools it. Treating CoT as a readable ECG and inferring what the model "is thinking" is easy to mislead outside the training distribution. Correct answer, wrong derivation means a model can say "I double-checked" or "I compared three approaches" while actually producing post-hoc rationalisation rather than a real computation trace.
The paper goes further: monitoring a model's "thinking" may be weaker than monitoring its behavior, external verifiable traces, or dedicated probes. Each of those alternatives has its own blind spots, which is why the authors emphasise that the CoT route should not be abandoned, but its semantic carrying capacity must be recognised as fragile and bounded.
So what
This is not a generic "AI is not to be trusted" paper; its punch is that it verifies, rather than asserts, a default assumption. Frontier labs are now pushing agentic AI, and an agent's intermediate reasoning is precisely the entry point for evaluation and audit. If even strict iGSM verification cannot stop 31.6% of "fake reasoning", how much of the long reasoning chains now deployed in code generation, financial analysis and medical assistance are quietly running a different logic under a presentable surface?
Kambhampati has not been quiet on it either. He has spent the past two years flagging the limits of CoT monitoring, and iGSM turns the suspicion into a specific percentage. CoT will not vanish, but "reading its literal words is enough" is a claim nobody is going to make with a straight face anymore.