Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

TOPIC · ENTITY #5489

chain-of-thought

Related event timeline, sources and context from the news index.

EVENT TIMELINE

3

RESEARCH · 1 SOURCE · MIT Technology Review AI

Opinion: Current LLMs don’t perform the kind of explicit reasoning systems like AlphaGo used

The piece argues that while AlphaGo combined an intuition-like policy network with an explicit search-based reasoning process to produce creative moves, today's large language models operate as iterative next-token predictors and lack a genuine, separate, inspectable reasoning mechanism; techniques such as chain-of-thought improve performance but do not create an independent epistemic state and can be post-hoc rationalizations. The author highlights three core shortcomings—no persistent inspectable beliefs, no separation between knowledge and manipulation, and unreliable chain-of-thought—and warns this limits trustworthiness in high-stakes domains like medicine and science.

7.0

RESEARCH · 1 SOURCE · Apple Machine Learning Research

The Communication Bottleneck: round-trip study of tree-structured expression serialization in language models

The paper proposes a round-trip protocol to measure how well tree-structured arithmetic expressions survive serialization into natural-language word problems and back. Evaluating 16 models pairwise, the authors find the channel is lossy and asymmetric (swapping generator and extractor can change accuracy by up to 60.4 points), the best generator–extractor pair achieves 92.9% round-trip correctness, at least 73.6% of failures originate in generation, structure (operator count, depth, right-branching) drives difficulty more than model family, and ~3,600 targeted fine-tuning examples substantially improve open-weight models (raising them above untrained Gemini-3.1-Pro and helping in disjoint-domain cases).

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Round-trip study finds lossy, asymmetric serialization of tree-structured expressions in language models

The arXiv paper (arXiv:2609.21509v1) proposes a round-trip protocol where one model generates a word problem from a procedurally created arithmetic expression and another model extracts the original expression; symbolic equivalence is used as an exact oracle. Evaluating 16 models pairwise, the study finds the natural-language channel is lossy and asymmetric (generation vs extraction can shift accuracy by up to 60.4 points), most failures stem from generation and tree structure (operator count, depth, right-branching) drives difficulty, and roughly 3,600 targeted fine-tuning examples substantially improve open-weight models above an untrained Gemini-3.1-Pro baseline.

7.0