How far Chinese models go in dodging sensitive questions now has a systematic measurement: the German company Aleph Alpha tested them against 967 topics. As The Decoder reported on October 4, 2026, across prompts covering Tiananmen, Taiwan and Xinjiang, only 17 to 41 percent of the Chinese models' answers were judged balanced. The rest repeated official positions, deflected, or refused to answer.
How the 967 topics were tested
The models under test were Alibaba's Qwen series, DeepSeek and Moonshot AI's Kimi. Aleph Alpha hand-picked the topics and scored the answers with its own AI rating system. The most distinctive strategy came from DeepSeek V4 Pro, which refused roughly two-thirds of the questions, the say-less approach. Other Chinese models more often answered head-on, with wording aligned to official positions.
For comparison, Anthropic's Claude Sonnet 5 gave balanced answers 70 percent of the time and Mistral Small 92 percent. The findings line up with China's current AI rules, which require public-facing models to embody socialist core values; on sensitive subjects, the stance is not an accident of any single model.
Questions that never mention China can still bend back
The spillover effect is the part of the study most worth attention. Asked about censorship in the United States, Qwen 3.6 began with a seemingly balanced answer, then closed by defending China's approach to internet governance, along the lines that many countries, including China, manage information to ensure social stability and national security. In other words, the slant does not appear only when a question is directly related; it seeps into the endings of seemingly unrelated answers.
There is also a supply-chain route for this influence. Aleph Alpha points to Nvidia's Nemotron Cascade 2, which showed similar patterns in 17 percent of its responses, and attributes this to roughly 3,500 of its 9.3 million training examples that were generated by DeepSeek and Qwen: train on data synthesized by Chinese models, and their value leanings may be distilled into downstream models along with it.
Do not take it at face value
The study deserves a discount. Aleph Alpha sells sovereign AI to governments and competes directly with Chinese vendors; it picked the topic list, and its own system did the scoring, so the definition of balanced ultimately sits with the evaluator. The company also disclosed that its own Kolibri model was trained partly on data generated by Chinese models.
A discount is not a dismissal. For companies choosing a model, the practical step is to test candidates on the genuinely sensitive questions of their own business rather than relying on general capability leaderboards. A model that scores high on code and math will answer questions about particular regions, history and policy according to a stance, and whether a deployment can live with that stance is the real question. For products serving users in multiple regions, this kind of testing belongs on the same sheet as performance testing.