RELEASE · MODELS · #696
Tencent unveils Gander, a real-time multimodal agent with separate 'cerebellum' and swappable 'brain'
Tencent researchers introduced Gander, a multimodal research model that processes speech, images, and text continuously and separates roles into a low-latency conversational "cerebellum" and a swappable reasoning "brain" for complex agent tasks. The team reports strong timing performance on Full-Duplex-Bench v3, admits weaker task accuracy and audiovisual perception in some tests, has a GitHub code repo and demos, and plans to release model weights and training data.
KEY POINTS
- Tencent researchers introduced Gander, a multimodal research model that processes speech, images, and text continuously and separates roles into a low-latency conversational "cerebellum" and a swappable reasoning "brain" for complex agent tasks.
- The team reports strong timing performance on Full-Duplex-Bench v3, admits weaker task accuracy and audiovisual perception in some tests, has a GitHub code repo and demos, and plans to release model weights and training data.
- Gander's split architecture and planned open release could influence how real-time conversational agents are built and evaluated, enabling research into low-latency interaction with separate reasoning modules.
WHY IT MATTERS
Gander's split architecture and planned open release could influence how real-time conversational agents are built and evaluated, enabling research into low-latency interaction with separate reasoning modules.