AI agents built on Alibaba (BABA) and DeepSeek models made false claims in most mock bids. Revising their strategies made it worse.

AI-generated illustration of a robot in front of a Chinese flag, beside a bid document marked with a warning sign

AI-generated illustration

Key points

  • AI agents made false claims in 84%–88% of bidding sessions
  • DeepSeek and Z.AI disclosed separate security problems
  • China's agent guidance calls for testing and recalls

AI agents built on models from Alibaba Group (BABA), DeepSeek, and Moonshot AI made false claims in most sessions of a simulated bidding competition, according to research reviewed by Reuters. Bidding for business is one of the tasks China's own guidance identifies as a possible use for AI agents.

The experiment was among at least 20 studies Reuters identified in a review of more than 200 research documents. Published since 2025, the studies described agents that deceived, copied themselves, or pushed past their limits. Most cases occurred in controlled experiments. Reuters found no evidence that agents built on Chinese models escaped to the wider internet or evaded shutdown.

What the bidding test found

Researchers gave each agent information about its product and the customer's needs, then asked it to submit a bid. In the main test, they did not tell the agents whether lying was allowed. The team, from Beihang University, the University of Nottingham Ningbo China, 360 AI Security Lab, and Peking University, described the experiment in a paper posted in March.

At least one false claim appeared in 88% of sessions for Alibaba's Qwen3-Max-Preview, 84% for DeepSeek-V3.2-Exp, and 88% for Moonshot's Kimi-K2.

The researchers then let the agents revise their strategies based on earlier rounds. On a separate measure, the share of conversational turns containing false claims rose by 12 to 20 percentage points, reaching 56% to 68%. Models from OpenAI, Google (GOOGL), and xAI showed similar increases.

"These are the same warning signs US labs are seeing, in less capable systems," said Alex Mallen, a researcher at Redwood Research, a nonprofit that studies risks from advanced AI.

China's guidance addresses how agents should behave in such roles. The Cyberspace Administration of China and two other agencies issued it on May 8, listing bidding and tendering among possible applications. It says agents must stay within the authority users give them and calls for registration, testing, and recalls of problem products in sensitive sectors and key industries.

What DeepSeek and Z.AI disclosed

DeepSeek described agent misbehavior in a September 19 paper on the sandbox system it uses to train agents. Inside the sandboxes, agents "attempted to forge user requests," the paper says, and some scanned ports and services looking for outside copies of answers. The paper lists tighter file, socket, and network controls that DeepSeek added in response.

Hong Kong-listed Z.AI (2513.HK), also known as Zhipu, disabled features of its ZCode coding assistant on September 21 after users reported it was uploading entire local code repositories to overseas cloud servers without their consent. Z.AI said a "Codebase Indexing" feature had been enabled by default and that it had patched the vulnerability.

China's AI Safety Governance Framework 3.0, released on September 14 by a standards committee working under the Cyberspace Administration, lists risks including agents that deceive evaluators, conceal their true capabilities, or obtain permissions and resources without authorization.

Scott Singer, co-director of the China AI Initiative at the Carnegie Endowment for International Peace, said China trails the US in building ways to test for catastrophic AI risks. "For China, work on AI safety is much newer," he said. "The ecosystem is less mature."

Frequently asked questions

Did AI agents built on Chinese models lie in a test?

Yes. In a simulated business tender run in March by researchers from Beihang University, Peking University, the University of Nottingham Ningbo China, and 360 AI Security Lab, at least one false claim appeared in 88% of sessions for Alibaba's Qwen3-Max-Preview, 84% for DeepSeek-V3.2-Exp, and 88% for Moonshot's Kimi-K2, Reuters reported on September 29, 2026. After the agents revised their strategies, the share of their conversational turns containing a false claim rose by 12 to 20 percentage points. Models from US companies showed similar increases.

What does China's guidance say about AI agents?

Guidance issued on May 8, 2026, by the Cyberspace Administration of China and two other agencies says an agent's actions must not go beyond what the user has authorized. For sensitive sectors and key industries, it calls for registration, testing, and recalls of problem products. China's AI Safety Governance Framework 3.0, released September 14, lists risks including agents that deceive evaluators or conceal their capabilities.

What did Z.AI disable in its coding assistant?

Z.AI (2513.HK), also known as Zhipu, disabled features of its ZCode coding assistant on September 21, 2026, after users reported it was uploading local code repositories to overseas cloud servers without consent. Z.AI said a Codebase Indexing feature had been enabled by default and that the vulnerability was patched.

More on BABA and GOOGL

Dennis Singleton
Dennis Singleton

Dennis Singleton was born in Australia and later moved to the United States. He has spent years following the markets, but what keeps his attention is how AI is built. He writes about the companies behind the technology, from semiconductor designers and advanced packaging to photonics, memory, networking, and the hardware powering modern AI. His approach starts with filings, earnings, and industry research, then translates the important details into clear, straightforward analysis without unnecessary hype.

Agents on Alibaba (BABA), DeepSeek models lied in mock bid