OpenAI says a campaign detected in July 2026 manipulated model interactions to expose encrypted reasoning artifacts—not that attackers cracked its encryption. The company describes the method as an attempt to copy hidden reasoning from one conversation and get a model in another conversation to decrypt and transcribe it. OpenAI says it disrupted the activity, but its public account does not establish how many attempts succeeded or independently prove its attribution of a core cluster to people associated with Moonshot AI.
What was OpenAI’s “encryption bypass”?
In a September 30, 2026 statement, OpenAI said it identified and disrupted a coordinated effort to extract “protected reasoning” from its models. OpenAI uses that phrase for a model’s internal record of working through a task; it says the reasoning may contain information not included in the final answer.
The novel method described by OpenAI relied on model interactions: operators copied encrypted reasoning from one conversation, then asked a model in a separate conversation to decrypt and transcribe the hidden content. In that sense, “bypass” describes a way of eliciting protected material through prompts and replay—not a demonstrated cryptographic break.
OpenAI put the distinction plainly: “The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations.” The statement is the company’s account of the incident, not an independent technical audit.
#1 Best Overall
Did hackers break OpenAI’s encryption?
OpenAI says they did not. Its description is that operators manipulated model interactions to make protected reasoning visible to requesters. That is different from defeating the encryption algorithm, breaching a database, or directly reading stored conversations.
The mechanism still matters: encrypted reasoning artifacts could be copied and presented back to a model, which might then be prompted to reveal their contents. OpenAI says it closed a pathway that allowed someone already holding another user’s encrypted reasoning to replay it and recover its contents. That account does not establish that every attempt worked, or that every protected reasoning artifact was exposed.
Rank #2
What is adversarial distillation?
OpenAI defines adversarial distillation as the systematic, unauthorized use of another model’s outputs or reasoning to train, reproduce, or improve a model. The goal is not simply to get an answer to an isolated question; it is to collect material that could help reproduce a model’s capabilities.
OpenAI’s concern is that reasoning could provide useful training material beyond the model’s user-facing answers, potentially transferring advanced capabilities without the same investment in safety. The company highlights risks in dual-use domains. Those are OpenAI’s stated risk assessments, not findings that this particular campaign successfully trained another model.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →When did the campaign happen, and how large was it?
OpenAI’s reported timeline describes detection and attempted extraction activity, not a confirmed tally of stolen reasoning:
| Date | What OpenAI reported |
|---|---|
| July 1, 2026 | OpenAI says it first observed low-volume activity. |
| July 24–25, 2026 | OpenAI says it recorded 16,000 requests using a relevant extraction pattern from more than 4,000 users. These were attempts, not necessarily successful disclosures. |
| By July 28, 2026 | OpenAI says it had disrupted related prompt-pattern activity across more than 15,000 users. |
The figures count requests and users associated with patterns OpenAI considered relevant; they are not counts of successful extractions, confirmed victims, or members of one identified group. OpenAI has not provided a success rate or a definitive list of targeted models in the cited reporting.
Did Moonshot AI steal OpenAI’s reasoning?
OpenAI attributes a core cluster of activity to individuals it says were associated with Moonshot AI, the developer of Kimi. It has not attributed every observed operator to Moonshot and says it is unclear whether all the activity came from one actor.
That attribution should be read as OpenAI’s assessment, not as independently established responsibility. CyberScoop reported that OpenAI’s public post did not cite technical evidence or explain its reasoning for the attribution, and that the company declined to share more information for security reasons. The public material cited here therefore does not let readers independently verify the link.
Free tools Windows power users keep installed
One-click scans. No signup required.
How does this relate to other reasoning-extraction research?
An August 10, 2026 arXiv preprint by Alexander Panfilov and seven coauthors, “Stealing Reasoning Traces from Proprietary LLM APIs”, describes encrypted reasoning blocks returned to clients and passed back in later requests. The authors report that blocks could be compatible across sessions, users, and models within a provider ecosystem, and describe a decryption-jailbreak approach demonstrated across Anthropic, OpenAI, and Google.
This paper describes a broader attack class; it is not proof that every vector or detail in the paper was used in OpenAI’s campaign. OpenAI says independent researchers responsibly disclosed related cross-model and conversation-compaction vulnerabilities and that it confirmed their attack paths.
What protections does OpenAI say it added?
OpenAI reports several responses to the activity:
- Banning or restricting accounts it considered fraudulent.
- Strengthening signup and infrastructure controls and expanding monitoring.
- Closing a pathway for replaying another user’s encrypted reasoning to recover its contents.
- Adding checks to detect and hold streamed output that might expose reasoning.
- Coordinating with third-party providers, the Frontier Model Forum, and government information-sharing channels.
OpenAI also says partner-hosted deployments need equivalent protections and that tool-output attacks require safeguards that inspect beyond ordinary visible text. The company describes its mitigation and investigation as ongoing.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →

