NEWS · MODELS · #1022
Pistis: 27B and 9B multimodal LLM family with IDRL post-training (arXiv:2609.28554v1)
arXiv:2609.28554v1 presents the Pistis model family—27B- and 9B-parameter multimodal LLMs built on Qwen3.6 and Qwen3.5—trained via a scalable post-training framework that starts with large-scale multimodal supervised fine-tuning and then applies Interleaved Distillation and Reinforcement Learning (IDRL). The paper describes two specialized variants (Pistis-Thinking for deep multimodal reasoning and Pistis-Agentic for long-horizon planning and tool use), plus a system method Pistis-Auto-Harnessing (PAH); both scales reportedly outperform their corresponding base models, with the Agentic variant strong in multimodal search.
KEY POINTS
- arXiv:2609.28554v1 presents the Pistis model family—27B- and 9B-parameter multimodal LLMs built on Qwen3.6 and Qwen3.5—trained via a scalable post-training framework that starts with large-scale multimodal supervised fine-tuning and then applies Interleaved Distillation and Reinforcement Learning (IDRL).
- The paper describes two specialized variants (Pistis-Thinking for deep multimodal reasoning and Pistis-Agentic for long-horizon planning and tool use), plus a system method Pistis-Auto-Harnessing (PAH); both scales reportedly outperform their corresponding base models, with the Agentic variant strong in multimodal search.
- This matters because IDRL integrates on-policy distillation and reinforcement learning in a single loop to improve knowledge transfer, optimization stability, and long-horizon agentic performance, while PAH demonstrates inference-side improvements without changing model weights.
WHY IT MATTERS
This matters because IDRL integrates on-policy distillation and reinforcement learning in a single loop to improve knowledge transfer, optimization stability, and long-horizon agentic performance, while PAH demonstrates inference-side improvements without changing model weights.