OpenAI Says Moonshot-Linked Operators Tried to Unlock Its Hidden Reasoning. Anthropic Named the Same Lab Three Weeks Ago. Tokentheft Just Changed Shape.
On September 11, Anthropic said a lab called Moonshot AI ran more than 23 million exchanges against Claude through 5,380 fraudulent accounts to copy its reasoning. This week OpenAI said a core cluster of operators associated with Moonshot AI went after its models too, in the same July, with a very different technique. Two vendors, one lab, one month. That is the second dated observation of a category we named on September 1: tokentheft, theft whose objective is inference itself.
Two observations, side by side
Observation one: Anthropic's September threat intelligence report, published September 11, 2026. Moonshot, tracked as GTG-16002, ran over 23 million exchanges between May and July through 5,380 fraudulent accounts, mostly presenting from Singapore and Japan, and silently relayed its own Kimi customers' prompts to Claude so it could keep Claude's answers. Anthropic called its attribution high confidence. We covered it here: https://www.dugganusa.com/post/seven-prc-labs-ran-190-million-exchanges-against-claude-to-copy-its-reasoning-that-is-tokentheft-at
Observation two: OpenAI's report, "Disrupting a coordinated model distillation campaign," covered by The Hacker News (Ravie Lakshmanan) and CNBC on October 1, 2026. The activity began July 1, spiked on July 24 and 25 to 16,000 extraction requests from more than 4,000 users, connected to a wider cluster of more than 15,000 users, and was fully disrupted by July 28. OpenAI ties a core cluster of the activity to individuals associated with Moonshot AI, and is careful to say it does not attribute every operator to one actor or establish that every attempt worked.
What is new
The prize moved. In the Anthropic case the target was reasoning the model already shows: chain-of-thought transcripts, harvested by buying millions of exchanges and by relaying other people's prompts. In the OpenAI case the target was reasoning the model deliberately hides. OpenAI returns that reasoning to clients only in encrypted form. According to the reporting, the operators copied encrypted reasoning blocks from one conversation, inserted them into others, and prompted the model to decrypt and transcribe them. OpenAI's own line is the important one: the operators "did not break our encryption, compromise a database, or gain direct access to stored user conversations." They did not need to. They found a pathway where a protected artifact could be replayed somewhere it was never meant to go, and asked the model to read it out loud.
The traffic shrank. Anthropic's Moonshot figure is tens of millions of exchanges over three months. OpenAI's peak is 16,000 requests over two days. Those are different units over different windows, so we will not put a ratio on it, but the direction is plain: when the visible reasoning gets expensive to buy at volume, the cheaper move is a protocol trick that needs far fewer requests per stolen thought.
What each vendor caught, and how
Neither vendor caught this on volume, and the OpenAI case makes the point sharper than the first one did. Sixteen thousand requests spread across four thousand users is a handful per user. No rate limit trips on that. What OpenAI describes catching is a shape: the same extraction pattern, repeated across thousands of new accounts, aimed at an artifact that should only ever travel inside the conversation that produced it. Anthropic described the same kind of detector in September: account pools, their replacement pools, and the content they were after, not the count. That is the detector we argued for when we named the category. Tokentheft is invisible to identity, DLP and EDR because every request is a valid account making a well-formed call. The only things that give it away are the shape of the traffic and the bill.
What stays the same
The objective is still inference, not data. The access is still a legitimate account, or thousands of them. And the payer is still someone other than the thief: a lab's compute, a customer's prompt, a company's research budget. The same lab now appears in two vendors' reports in the same window, which is the part we would underline for anyone tempted to treat the first disclosure as a one-off. Attribution here is each vendor's, not ours, and OpenAI's is explicitly partial.
If you run an AI API, or a product built on one
Bind opaque artifacts to where they were born. Encrypted reasoning, session tokens, tool results and cached context should be valid only in the conversation and account that produced them. If a client can paste one somewhere else and get it processed, you have OpenAI's pathway.
Alert on shape, not rate. Look for one prompt template repeated across many new accounts, signup clusters sharing payment patterns, disposable email and residential proxies, and accounts that only ever ask for the one thing your model protects.
Watch your outputs for your own internals. If a response reproduces a format you never meant users to see, that is a detection, not a quirk.
Watch for relays. In September, Moonshot's customers did not know their prompts were going to Claude. If your traffic includes requests that look like another product's users, someone may be reselling your inference.
Meter per caller and read the meter. The invoice is still the last sensor standing; make it the first.
Receipts
We named tokentheft on September 1, 2026, before either of these disclosures, from a much smaller case: a research non-profit that lost about $600,000 of inference to someone who asked its agent for the API key. That is a naming date, not a detection lead. Neither campaign touched our feeds, and neither vendor published network indicators, so there is nothing new for the DugganUSA feed from this story. What the two reports give us is the thing a single incident cannot: demonstrated evolution, two dated observations of one actor, with a changed technique between them.
Credit for the primary research belongs to OpenAI ("Disrupting a coordinated model distillation campaign," https://openai.com/index/disrupting-a-coordinated-model-distillation-campaign/) and to Anthropic's September threat intelligence report. Reporting by Ravie Lakshmanan at The Hacker News (https://thehackernews.com/2026/10/openai-disrupts-reasoning-extraction.html), by CNBC, and by Ryan Merket at Runtime Wire. Every number above is the vendors'. The framing is ours, at our usual 95% confidence cap.
Was this useful? Rate it in the box below. It takes thirty seconds, and it is how we decide what to write next.
How do AI models see YOUR brand?
AIPM has audited 250+ domains. 15 seconds. Free while still in beta.
Was this useful? Thirty seconds, no cookies, no tracking, no third parties, your address hashed and never stored. If the box below does not load, the same question lives at https://analytics.dugganusa.com/nps.html?post=openai-says-moonshot-linked-operators-tried-to-unlock-its-hidden-reasoning-anthropic-named-the-same



Comments