Opinion: Current LLMs don’t perform the kind of explicit reasoning systems like AlphaGo used
The piece argues that while AlphaGo combined an intuition-like policy network with an explicit search-based reasoning process to produce creative moves, today's large language models operate as iterative next-token predictors and lack a genuine, separate, inspectable reasoning mechanism; techniques such as chain-of-thought improve performance but do not create an independent epistemic state and can be post-hoc rationalizations. The author highlights three core shortcomings—no persistent inspectable beliefs, no separation between knowledge and manipulation, and unreliable chain-of-thought—and warns this limits trustworthiness in high-stakes domains like medicine and science.