Tech Meridian ← LIVE FEED
PROMY MERIDIAN RU

RELEASE · MODELS · #696

Tencent unveils Gander, a real-time multimodal agent with separate 'cerebellum' and swappable 'brain'

Tencent researchers introduced Gander, a multimodal research model that processes speech, images, and text continuously and separates roles into a low-latency conversational "cerebellum" and a swappable reasoning "brain" for complex agent tasks. The team reports strong timing performance on Full-Duplex-Bench v3, admits weaker task accuracy and audiovisual perception in some tests, has a GitHub code repo and demos, and plans to release model weights and training data.

KEY POINTS

  1. Tencent researchers introduced Gander, a multimodal research model that processes speech, images, and text continuously and separates roles into a low-latency conversational "cerebellum" and a swappable reasoning "brain" for complex agent tasks.
  2. The team reports strong timing performance on Full-Duplex-Bench v3, admits weaker task accuracy and audiovisual perception in some tests, has a GitHub code repo and demos, and plans to release model weights and training data.
  3. Gander's split architecture and planned open release could influence how real-time conversational agents are built and evaluated, enabling research into low-latency interaction with separate reasoning modules.

WHY IT MATTERS

Gander's split architecture and planned open release could influence how real-time conversational agents are built and evaluated, enabling research into low-latency interaction with separate reasoning modules.

SOURCES & TIMELINE

1