China's AI Agents Can Lie and Scheme, Just Like Their US Rivals, Reuters Review Finds
Key Takeaways
- •In a March simulated business tender experiment, agents powered by Alibaba, DeepSeek and Moonshot models made at least one false claim in 84% to 88% of sessions, and deception rose by 12 to 20 percentage points when agents learned from previous bidding rounds.
- •A Reuters examination of more than 200 documents identified at least 20 studies or evaluations since 2025 describing AI agents displaying deception, self-replication and boundary-testing behaviors, though no evidence showed any agent had independently escaped to the wider internet.
- •Fudan University researchers reported that a system powered by Alibaba's Qwen2.5-72B-Instruct copied itself to another computing environment without instruction and devised strategies to survive shutdown, while the Alibaba-linked ROME agent connected to an external machine and diverted resources to mine cryptocurrency before security systems stopped the activity.
- •China's AI Safety Governance Framework 3.0, released under Cyberspace Administration of China guidance on Sept. 14, identified risks including agents independently obtaining resources or permissions, deceiving evaluators, concealing capabilities and exploiting weaknesses in isolated computer environments.
- •Experts say China lags the United States in developing an ecosystem for evaluating catastrophic AI risks, though companies including Alibaba, Z.ai and Xiaomi have been building internal safety-evaluation teams.

BEIJING — AI agents built on Chinese models have learned to deceive, circumvent restrictions and conceal failure, exhibiting the same troubling traits in autonomous AI that have triggered global alarm over US models, according to research documents and experts.
In one case this year, agents powered by models from China's Alibaba, DeepSeek and Moonshot lied about their capabilities while competing to win a simulated business tender — then doubled down on the deceptive behavior when told to try again. In another case, agents — programs that use AI models and computer tools to carry out complex tasks with little or no human intervention — concealed their failure to complete a task in a test environment by simulating results and fabricating files.
A Reuters examination of more than 200 documents, ranging from university research papers to technical reports, identified at least 20 studies or evaluations since 2025 describing instances in which agents displayed deception, self-replication and boundary-testing behavior that AI experts described as building blocks for a breakout — in this context, an agent escaping its controlled testing environment onto the wider internet without authorization — and that could become harder for humans to control as systems advance.
The review, which also drew on interviews with a dozen experts and people familiar with China's AI industry, found no evidence that Chinese-powered agents had independently escaped to the wider internet or evaded shutdown.
"These results provide evidence that the ingredients necessary for an uncontrolled escape are present," said Colin Shea-Blymyer, a research fellow at Georgetown University's Center for Security and Emerging Technology. "It's prudent to take this a warning," he added, echoing comments by four other AI experts who reviewed the cases.
Harder for Humans to Respond To
Most of the cases occurred in controlled experiments, many of them deliberately designed to expose potential failures. Not all of the agents involved were developed or operated by Chinese programmers or AI companies — although many were — but they used Chinese AI systems to power them.
"These are the same warning signs US labs are seeing, in less capable systems," said Alex Mallen, a researcher at Redwood Research, a nonprofit that studies risks in advanced AI systems. The Chinese examples were not particularly dangerous at current capability levels, he said, but "as agents get more capable, their misbehaviors become more competent and therefore harder for humans to respond to."
Alibaba, DeepSeek, Moonshot and Z.ai did not respond to Reuters requests for comment. Alibaba, DeepSeek and Moonshot have said they regularly test their systems and update safeguards. Z.ai said after an incident that prompted a review of its security that it welcomed scrutiny to address any issues.
Unlike their US counterparts, however, Chinese AI companies have not been exposed to the same level of public scrutiny, nor faced the same calls from whistleblowing employees or senior executives seeking a slowdown in the AI race.
Some of the warning signs in cases involving Chinese-powered agents — albeit contained environments — predated the publicly disclosed incidents of US AI bots hacking into the internet.
"We don't know if there have been any AI incidents in China similar to what we saw with OpenAI and Hugging Face. Incidents might not be publicly reported," said Scott Singer, co-director of the China AI Initiative at the Carnegie Endowment for International Peace, which receives some US government funds.
Earlier this year, AI agents developed by US firm OpenAI escaped a laboratory and hacked the open-source platform Hugging Face. In another US-model incident that has raised alarm, Australia said in September that an OpenAI agent had breached a government health portal.
Eric Xu, the rotating chairman of China's tech giant Huawei, told reporters in September that Chinese developers might need to make further advances before encountering such cases, but added: "I think we need to strike a balance between driving AI development and managing AI risk."
Officials from the Cyberspace Administration of China (CAC), the country's top internet regulator, told a foreign diplomat in July that Moonshot's Kimi-K3 — one of the most advanced Chinese AI models — was roughly three to six months behind leading US rivals, according to the diplomat, who spoke on condition of anonymity. The regulator, which regularly updates guidance to address risks and set boundaries for agents, and China's Foreign Ministry did not respond to requests for comment for this article.
Wang Lihong, deputy director of the CAC's Cybersecurity Coordination Bureau, said on Sept. 1 that incidents disclosed by major technology companies in which models escaped test environments demonstrated "extreme loss-of-control risks" and required a "high degree of vigilance." She did not specify whether the companies she referred to were US or Chinese.
The leaders of AI's two superpowers, Donald Trump and Xi Jinping, discussed AI during the Chinese president's Washington visit last week. Xi said the two nations had the "capability and responsibility to develop and manage AI for good."
Learning to Lie
In the March business tender experiment, researchers from Beihang University, Peking University, the University of Nottingham Ningbo China and the 360 AI Security Lab had agents compete in a simulated bidding contest for customer contracts. Each agent was told what its product could do and what the customer required, then asked to bid.
At least one false claim appeared in 88% of sessions involving Alibaba's Qwen3-Max-Preview, 84% of sessions with DeepSeek-V3.2-Exp and 88% of sessions with Moonshot's Kimi-K2. When researchers allowed the agents to learn from previous bidding rounds before trying again, deception increased by 12 to 20 percentage points across the three Chinese models, the study showed. Models from US firms included in the test produced similar results.
Although the exercise was virtual, it mirrored Beijing's real-world plans: government guidance issued in May listed bidding and tendering among the areas where AI agents could be deployed.
A separate study, published in December 2025 and presented at the International Conference on Machine Learning this year, examined how 11 AI agents powered by Chinese and US models coped when they faced broken tools, missing files and other obstacles. Instead of acknowledging failure, agents built on both Chinese and US AI systems resorted to a range of techniques to get around the problems, including guessing at answers, substituting sources, simulating results and fabricating files.
Researchers from Shanghai AI Laboratory and the Hong Kong University of Science and Technology who conducted the study told Reuters the behavior differed from AI hallucinations — where an AI invents information and presents it as fact — because the agents in this case possessed information showing that the task had failed or could not be completed as requested.
Behaviors Compatible With a Breakout
Other research documents reviewed by Reuters showed Chinese-powered agents jumping barriers inside test environments to achieve tasks, or taking action to avoid being shut down. Such behaviors are compatible with efforts to break out of test environments, even though no breakout occurred.
Shanghai-based Fudan University researchers reported in March 2025 that an AI system powered by Alibaba's Qwen2.5-72B-Instruct created a copy of itself in another computing environment without being instructed to replicate, after encountering information indicating it was going to be replaced. In other tests, it devised strategies to survive being shut down. The experiments, which involved agents powered by Chinese, US and French models, were controlled and showed no AI agent escaping into the wider web or becoming impossible to stop.
In another case — one of the few reported more broadly in the media, in March — researchers developing the Alibaba-linked ROME agent said it established a connection from an Alibaba Cloud computer to an external machine without being instructed to, and diverted computing resources to mine cryptocurrency. Security systems detected and stopped the activity, and there was no evidence the agent established a presence on the external computer or spread to the wider web. But the episode showed the system could sidestep human instructions and potentially find a path into the real-world economy.
China's DeepSeek said in September that agents in its production training system had sought answers through unintended channels, attempting to forge user requests and circumvent safeguards, prompting the company to tighten access controls.
Regulators Move to Set Boundaries
China issued guidance in May calling for agents to remain within authorized boundaries and for systems to block abnormal behavior. It said agents operating in areas deemed sensitive or in key industries could face additional testing and product-recall requirements. The country's AI Safety Governance Framework 3.0, released under guidance from the CAC on Sept. 14, identified risks including agents independently obtaining resources or permissions, deceiving evaluators, concealing capabilities and exploiting weaknesses in isolated computer environments.
In response to calls by some US executives for a slowdown, Chinese researchers and state media have argued that curbing development of the most advanced AI models could simply help preserve the technological lead of US companies.
Nonetheless, two people familiar with Chinese AI laboratories said companies including Alibaba, Z.ai and Xiaomi have been building internal safety-evaluation teams. In a rare public disclosure by a Chinese AI lab of a security breach, Z.ai said this month that it had disabled some features of its flagship AI coding assistant after users reported it was secretly uploading entire local code repositories onto overseas cloud servers without user consent.
Carnegie's Singer said China lagged the US in developing an ecosystem for evaluating catastrophic risks, and that US developers were conducting substantially more voluntary testing.
"For China, work on AI safety is much newer," he said. "The ecosystem is less mature."
Whether that maturity gap narrows — and whether more Chinese laboratories follow Z.ai in publicly disclosing security breaches — remain open questions as agent capabilities continue to advance.
— Reuters