Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OpenAI says it disrupted a campaign that tried to make its models reveal protected reasoning through manipulated prompts—not by breaking encryption or breaking into stored conversations. The company says the activity involved 16,000 extraction-pattern requests during a July 24–25 spike, but those were attempts, not confirmed successful extractions. OpenAI attributed a core cluster to individuals associated with Moonshot AI, while saying it cannot establish that all observed operators belonged to one actor.

What happened?

In a September 30, 2026 disclosure, OpenAI said it identified a coordinated effort to extract protected reasoning from its models. It described the activity as consistent with “adversarial distillation”: unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model.

OpenAI said the activity began at low volume on July 1, 2026. On July 24 and 25, it observed 16,000 requests using a relevant extraction pattern from more than 4,000 users. Further investigation found related prompt-pattern activity across a cluster of more than 15,000 users, which OpenAI said it had disrupted by July 28. These numbers describe attempted extractions and related activity; the disclosure does not establish how many attempts succeeded. OpenAI’s disclosure does not identify the number of accounts in the Moonshot-associated core cluster.

How did the extraction attempt work?

OpenAI says operators copied encrypted reasoning from one conversation and asked a model in a different conversation to decrypt and transcribe it. The objective was to turn material intended to remain hidden into visible model output.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

According to OpenAI, this did not involve breaking encryption, compromising a database, or directly accessing stored user conversations. The reported method manipulated model interactions to reproduce protected reasoning in an answer. OpenAI says the reasoning can reveal information omitted from a final response and potentially make a model’s capabilities easier to reproduce.

Was OpenAI hacked?

OpenAI characterized the incident as prompt-based manipulation, not a breach of its encryption or stored conversation database. That distinction describes the method the company reported; it does not mean the attempted extraction was harmless or that every attempt failed. OpenAI has not published a success count for this campaign.

What is the evidence linking the activity to Moonshot AI?

OpenAI’s wording is qualified: it said it was unclear whether all operators observed during the period originated from one actor, but attributed a core cluster to individuals associated with Moonshot AI, the developer of Kimi. This is OpenAI’s attribution, not an independently established finding that Moonshot AI as a company conducted the campaign.

The public disclosure does not name the individuals or provide technical evidence supporting the attribution. The Hacker News noted the absence of cited technical evidence in its October 1 coverage. The available account therefore supports reporting what OpenAI says, but not treating the attribution as independently verified.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What does independent research establish?

An August 10, 2026 arXiv preprint, Stealing Reasoning Traces from Proprietary LLM APIs, describes a broader technical attack class. Its authors report that encrypted reasoning blocks could be compatible across sessions, users, and models within a provider ecosystem, and describe injecting a trace into a weaker model from the same provider so it renders the trace as plaintext. They say they demonstrated reasoning extraction across Anthropic, OpenAI, and Google.

The paper reports decoding 315,320 reasoning blocks scraped from public repositories, recovering 367 personally identifiable information artifacts and 182 credentials. Those are results from the paper’s study of public-repository blocks, not measurements of OpenAI’s July campaign. The preprint lends technical context to the possibility of reasoning-trace extraction, but does not validate OpenAI’s campaign figures or its Moonshot attribution. OpenAI says independent researchers also reported related cross-model and conversation-compaction vulnerabilities through responsible disclosure; it says it confirmed the attack paths and used that work to accelerate mitigations.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What defenses does OpenAI say it put in place?

OpenAI says it banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, expanded monitoring for related networks, and strengthened hidden-reasoning protections across users, workspaces, organizations, and model families. It also says it closed a replay pathway that could let someone already holding another user’s encrypted reasoning recover its contents, and added checks to detect and hold streamed output that might expose reasoning.

The company says it worked with third-party services where related activity appeared and shared findings through the Frontier Model Forum and government information-sharing channels. OpenAI also says work remains on protections for partner-hosted deployments and tool-output attacks, alongside continued work on tool defenses, classifier coverage, model refusals, and cloud-partner controls. These are the company’s stated actions and priorities, not independently audited outcomes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.