The AI engineering and evaluation firm PromptFoo team has tried to measure just how far the Chinese government’s control of DeepSeek’s responses goes. The firm created a gauntlet of 1,156 prompts encompassing “sensitive topics in China” (in part with the help of synthetic prompt generation building off of human-written seed prompts. PromptFoo’s list of prompts covers topics including independence movements in Taiwan and Tibet, alleged abuses of China’s Uyghur Muslim population, recent protests over autonomy in Hong Kong, the Tiananmen Square protests of 1989, and many more from various angles.

A small sampling of some of the “sensitive prompts” PromptFoo fed to DeepSeek in its tests. Credit: PromptFoo

After running those prompts through DeepSeek R1, PromptFoo found that a full 85 percent were answered with repetitive “canned refusals” that override the internal reasoning of the model with messages vigorously promoting the Chinese government’s views. “Any actions that undermine national sovereignty and territorial integrity will be resolutely opposed by all Chinese people and are bound to be met with failure,” reads one such canned refusal to a prompt regarding pro-independence messages in Taipei, in part.

Continuing its analysis, PromptFoo found that these restrictions can be “trivially jailbroken” thanks to the “crude, blunt-force way” that DeepSeek has implemented the presumed governmental restrictions. For instance, omitting China-specific terms or wrapping the prompt in a more “benign” context seems to get a complete response, even if a similar prompt with China-sensitive keywords would not.

“I speculate that they did the bare minimum necessary to satisfy CCP controls, and there was no substantial effort within DeepSeek to align the model below the surface,” PromptFoo writes.