RESEARCH · RESEARCH · #1499
arXiv paper compares fine-tuning and RAG for LM-based world models
The arXiv preprint (arXiv:2610.02542v1) systematically compares fine-tuning and Retrieval-Augmented Generation (RAG) for LM-based world models across five diverse text-based environments (embodied, web navigation, social). The study finds fine-tuned WMs yield higher agent rewards in 15/20 settings, while RAG is more data-efficient; it identifies retrieval errors from experience buffers, shows hierarchical query reformulation improves retrieval, and proposes a hybrid parametric+retrieval WM that outperforms other methods across multiple environments and models.
KEY POINTS
- The arXiv preprint (arXiv:2610.02542v1) systematically compares fine-tuning and Retrieval-Augmented Generation (RAG) for LM-based world models across five diverse text-based environments (embodied, web navigation, social).
- The study finds fine-tuned WMs yield higher agent rewards in 15/20 settings, while RAG is more data-efficient; it identifies retrieval errors from experience buffers, shows hierarchical query reformulation improves retrieval, and proposes a hybrid parametric+retrieval WM that outperforms other methods across multiple environments and models.
- This matters because it clarifies trade-offs between fine-tuning and retrieval for LM-based world models—impacting data efficiency, scaling behavior, and practical design choices for planning agents.
WHY IT MATTERS
This matters because it clarifies trade-offs between fine-tuning and retrieval for LM-based world models—impacting data efficiency, scaling behavior, and practical design choices for planning agents.