Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

NEWS · MODELS · #1022

Pistis: 27B and 9B multimodal LLM family with IDRL post-training (arXiv:2609.28554v1)

arXiv:2609.28554v1 presents the Pistis model family—27B- and 9B-parameter multimodal LLMs built on Qwen3.6 and Qwen3.5—trained via a scalable post-training framework that starts with large-scale multimodal supervised fine-tuning and then applies Interleaved Distillation and Reinforcement Learning (IDRL). The paper describes two specialized variants (Pistis-Thinking for deep multimodal reasoning and Pistis-Agentic for long-horizon planning and tool use), plus a system method Pistis-Auto-Harnessing (PAH); both scales reportedly outperform their corresponding base models, with the Agentic variant strong in multimodal search.

KEY POINTS

  1. arXiv:2609.28554v1 presents the Pistis model family—27B- and 9B-parameter multimodal LLMs built on Qwen3.6 and Qwen3.5—trained via a scalable post-training framework that starts with large-scale multimodal supervised fine-tuning and then applies Interleaved Distillation and Reinforcement Learning (IDRL).
  2. The paper describes two specialized variants (Pistis-Thinking for deep multimodal reasoning and Pistis-Agentic for long-horizon planning and tool use), plus a system method Pistis-Auto-Harnessing (PAH); both scales reportedly outperform their corresponding base models, with the Agentic variant strong in multimodal search.
  3. This matters because IDRL integrates on-policy distillation and reinforcement learning in a single loop to improve knowledge transfer, optimization stability, and long-horizon agentic performance, while PAH demonstrates inference-side improvements without changing model weights.

WHY IT MATTERS

This matters because IDRL integrates on-policy distillation and reinforcement learning in a single loop to improve knowledge transfer, optimization stability, and long-horizon agentic performance, while PAH demonstrates inference-side improvements without changing model weights.

SOURCES & TIMELINE

1