OpenAI Says Individuals Linked to China's Moonshot AI Tried to Copy Its Models' Hidden Reasoning
Key Takeaways
- •OpenAI attributes a core cluster of the campaign to individuals associated with Moonshot AI, while noting it is unclear whether all operators originated from a single actor.
- •The activity ran from July 1 and was fully disrupted by July 28, with 16,000 extraction requests from more than 4,000 users logged on July 24-25 alone.
- •Attackers did not break encryption or access stored conversations but manipulated model interactions, including replaying encrypted reasoning across separate chats, and OpenAI has closed the exploited pathway.
- •The suspected motive is adversarial distillation, and because AI outputs are not copyrightable, the dispute rests on terms-of-service violations rather than copyright law.
- •The episode follows a series of similar accusations involving DeepSeek, Anthropic's fraud claims, White House statements, and xAI's court admission, while Moonshot has not responded amid plans for a $3 billion Hong Kong IPO at a $50 billion valuation.

OpenAI says it has disrupted a coordinated campaign to copy the hidden reasoning its AI models generate before answering, tracing a core cluster of the activity to individuals associated with Moonshot AI, the Chinese startup behind the Kimi chatbot.
According to OpenAI's blog post, the campaign began on July 1. On July 24 and 25 alone, the company logged 16,000 extraction requests from more than 4,000 users, part of a wider cluster of more than 15,000 users. OpenAI says it fully disrupted the activity by July 28.
The target was not the models' answers but the work behind them. Modern AI models "reason" before they reply, working through a problem step by step in an internal scratchpad before presenting a clean result. The technique became competitive focus after OpenAI's o1 arrived in late 2024, with DeepSeek's R1 following in January 2025. OpenAI keeps that scratchpad encrypted and says extracting it can reveal information the final answer leaves out—material that could be used to train another model without the original safeguards.
"The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations," OpenAI said. "Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service."
One method, according to OpenAI, involved copying encrypted reasoning out of one conversation and asking a model to decode it in a different conversation. The company has since closed a pathway that allowed someone who already held another user's encrypted reasoning to replay it and recover its contents.
OpenAI's post does not directly connect the campaign to K3, but it leaves space for reasonable doubt. "It is unclear whether all operators we observed during the relevant time period originated from a single actor. However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi," OpenAI said.
Why attackers want hidden reasoning
The motive is distillation, a standard machine-learning technique: training a new AI on the outputs of a stronger one, which yields better results from smaller models without heavy training. Done without authorization, OpenAI calls the practice "adversarial distillation," which it defines as "the systematic and unauthorized use of one model's outputs or reasoning to train, reproduce, or improve another model."
The episode adds to a string of controversies involving AI companies, the most prominent being the unauthorized use of copyrighted data to train models. Distillation is a different matter: AI outputs are not copyrightable, so companies rely on prohibitions and safeguards in their terms of service to prevent competitors from using those outputs. In other words, this dispute turns on contract terms rather than copyright law.
A familiar accusation
OpenAI has been here before. In January 2025, it said it was reviewing signs that DeepSeek may have distilled its models, as Washington weighed national security risks.
Anthropic followed in February, accusing Chinese labs of using about 24,000 fraudulent accounts to generate more than 16 million exchanges with Claude. Online critics shot back that Claude itself was trained on the open internet.
By April, the White House was saying foreign entities, primarily in China, were running industrial-scale distillation campaigns. A week later, Elon Musk acknowledged in court that xAI had used distillation on OpenAI models to train Grok.
In June, Anthropic took the fight to Congress, asking for penalties for large-scale model extraction.
In August, researchers showed that OpenAI, Anthropic and Google each protected reasoning with a single provider-wide encryption key, and that attackers could coax models into spitting out hidden thoughts in plain text. All three companies deployed server-side patches after disclosure, though session logs shared earlier remain decodable.
Moonshot has not publicly responded to OpenAI's post. The company is targeting a $3 billion IPO in Hong Kong at a $50 billion valuation.