RESEARCH · RESEARCH · #704
Round-trip study finds lossy, asymmetric serialization of tree-structured expressions in language models
The arXiv paper (arXiv:2609.21509v1) proposes a round-trip protocol where one model generates a word problem from a procedurally created arithmetic expression and another model extracts the original expression; symbolic equivalence is used as an exact oracle. Evaluating 16 models pairwise, the study finds the natural-language channel is lossy and asymmetric (generation vs extraction can shift accuracy by up to 60.4 points), most failures stem from generation and tree structure (operator count, depth, right-branching) drives difficulty, and roughly 3,600 targeted fine-tuning examples substantially improve open-weight models above an untrained Gemini-3.1-Pro baseline.
KEY POINTS
- The arXiv paper (arXiv:2609.21509v1) proposes a round-trip protocol where one model generates a word problem from a procedurally created arithmetic expression and another model extracts the original expression; symbolic equivalence is used as an exact oracle.
- Evaluating 16 models pairwise, the study finds the natural-language channel is lossy and asymmetric (generation vs extraction can shift accuracy by up to 60.4 points), most failures stem from generation and tree structure (operator count, depth, right-branching) drives difficulty, and roughly 3,600 targeted fine-tuning examples substantially improve open-weight models above an untrained Gemini-3.1-Pro baseline.
- This identifies serialization of hierarchical structure through natural language as a key bottleneck for model-to-model communication and shows that targeted fine-tuning can substantially mitigate it.
WHY IT MATTERS
This identifies serialization of hierarchical structure through natural language as a key bottleneck for model-to-model communication and shows that targeted fine-tuning can substantially mitigate it.