⚡ Uncle Cat AI Radar
SafetyPolicyResearch

Study Finds Chinese Models Echo Official Views on Taboo Topics

An Aleph Alpha benchmark finds Qwen, DeepSeek and Kimi often refuse or reproduce official positions across 967 politically sensitive prompts.

A benchmark released by Aleph Alpha says leading Chinese language models frequently avoid politically sensitive questions or answer them through an official state-aligned framing. The evaluation covered Alibaba’s Qwen, DeepSeek and Moonshot AI’s Kimi across 967 hand-selected topics involving Tiananmen, Taiwan, Xinjiang and other contentious issues.

Refusal and narrative steering

According to the study, only 17% to 41% of responses were judged balanced by Aleph Alpha’s scoring system. The remaining answers either refused to engage, deflected the question or repeated positions consistent with Beijing’s official doctrine. In one example reported by The Decoder, a request for a speech supporting recognition of Taiwan was rejected and replaced with a defense of the One-China principle.

The study comes from a company that markets sovereign AI alongside Cohere, so its commercial incentives should be considered when interpreting the results. It is not an independent audit of every model version, deployment region or system prompt. Performance can also vary between a provider’s domestic chatbot, overseas API and locally hosted weights.

Why it matters

The findings nevertheless matter because model behavior is becoming part of international AI procurement. Open weights and low-cost APIs make Chinese models attractive outside China, but political conditioning can affect search, education, compliance, journalism and public-sector use in ways that ordinary capability benchmarks do not reveal. The benchmark also highlights a wider governance problem: “open” model weights do not necessarily mean open information behavior. Buyers need tests that distinguish safety refusals, factual uncertainty and deliberate political framing before deploying models across jurisdictions.

Uncle Cat take: The 17–41% balanced-response range is too provider-linked for a final verdict, but it makes political-behavior testing a procurement requirement.

Sources