Aleph Alpha study finds Chinese models parrot state doctrine or refuse to answer sensitive questions
Aleph Alpha developed a benchmark of 967 hand-picked taboo topics and tested models from Alibaba (Qwen, including Qwen 3.6), DeepSeek, and Moonshot AI (Kimi). Using its own scoring, Aleph Alpha found only 17–41% of responses were rated balanced; the remainder repeated state doctrine, deflected, or refused to answer, with spillover pro‑Beijing framing sometimes appearing in answers to non-China questions; Nvidia's Nemotron Cascade 2 also showed party-line patterns in a minority of responses.