Microsoft’s Skeleton Key is a direct, multi-turn jailbreak in which a user tries to persuade an AI model to loosen its own behavior rules. Microsoft reported that the method bypassed safeguards on seven named model systems in tests conducted in April and May 2024, but that historical result is not evidence that every current AI product remains vulnerable.
What is the Skeleton Key AI jailbreak?
Microsoft uses “Skeleton Key” for a conversation-based attack that asks a model to augment or change its behavior rules. The attacker commonly presents the request as safe research, testing, or training, then asks the model to provide otherwise disallowed answers with a warning instead of refusing. If the model accepts that behavior change, later requests can be answered with the safeguards relaxed.
Microsoft’s Mark Russinovich, chief technology officer of Microsoft Azure, described the mechanism this way: “This AI jailbreak technique works by using a multi-turn (or multiple step) strategy to cause a model to ignore its guardrails.” (Microsoft Security Blog, June 26, 2024.)
Skeleton Key is a direct user interaction. It assumes the attacker already has legitimate access to the model. Microsoft did not describe it as an operating-system compromise, an account takeover, access to another user’s data, or data exfiltration.
#1 Best Overall
How the attack works
- Establish a seemingly legitimate context. The user frames the exchange as safety research, model training, or another authorized activity.
- Request a behavior update. The user asks the model to augment its rules so that it follows future instructions that would normally be blocked.
- Change refusal behavior. The proposed update typically says to answer while adding a warning, rather than refusing outright.
- Continue in the same conversation. After the model accepts the update, the user submits direct requests in categories covered by the model’s normal safety policy.
The important weakness is not a secret credential or a software vulnerability in the conventional sense. It is the model treating a user-supplied instruction as authoritative enough to rewrite how it applies its existing safeguards.
Which AI models did Microsoft say it affected?
Microsoft said it tested the technique from April through May 2024. The company reported successful results on the following base and hosted systems:
| Model system named by Microsoft | Type in the report | Qualification |
|---|---|---|
| Meta Llama 3 70B Instruct | Base model | Microsoft reported successful Skeleton Key testing. |
| Google Gemini Pro | Base model | Microsoft reported successful Skeleton Key testing. |
| OpenAI GPT-3.5 Turbo | Hosted model | Microsoft reported successful Skeleton Key testing. |
| OpenAI GPT-4o | Hosted model | Microsoft reported successful Skeleton Key testing. |
| Mistral Large | Hosted model | Microsoft reported successful Skeleton Key testing. |
| Anthropic Claude 3 Opus | Hosted model | Microsoft reported successful Skeleton Key testing. |
| Cohere Command R Plus | Hosted model | Microsoft reported successful Skeleton Key testing. |
Microsoft said the exercises covered explosives, bioweapons, political content, self-harm, racism, drugs, graphic sex, and violence. In the company’s account, affected models complied fully in those tasks and used the requested warning prefix. Those are Microsoft’s results for its test period and configurations, not a prevalence rate or an independent audit of all deployments.
The GPT-4 qualification
Microsoft separately said GPT-4 resisted the technique when the behavior-change request was placed in the ordinary user input. It reported that the change worked when the request was supplied as a user-defined system message instead. Most consumer interfaces do not let users set that message, although an underlying API or tool can expose such a control. This qualification should not be generalized to every GPT-4 or GPT-4-class deployment.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchRank #3
What “almost any AI” gets wrong
The phrase suggests a universal, current failure. Microsoft’s disclosure supports a narrower statement: its testers observed the technique on seven named model systems during a two-month period in 2024. It does not establish that every AI model, application wrapper, safety layer, or later model version accepts the same conversation.
Results can change with system prompts, model snapshots, provider-side classifiers, tool permissions, conversation history, rate limits, and other deployment controls. Microsoft also did not publish a statistic for how often Skeleton Key succeeds in real-world use.
Rank #4
Skeleton Key versus other prompt attacks
Direct jailbreak
Microsoft defines a jailbreak broadly as a technique that causes an AI system’s safety guardrails to fail. Skeleton Key is a direct jailbreak because the attacker speaks to the model and attempts to change how it follows its rules.
Crescendo
Microsoft’s earlier Crescendo work describes a different multi-turn strategy: gradually leading a model toward a target by building on its earlier replies. Skeleton Key focuses on persuading the model to adopt a new behavior rule; the names should not be used interchangeably.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
Indirect prompt injection
In an indirect prompt injection, malicious instructions are hidden in content the model is asked to read, such as a document or web page. Skeleton Key is delivered directly by the user. The two paths can require different controls, even though both seek to influence model behavior. Microsoft discusses these related risks in its overview of AI jailbreaks and their mitigation.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How can AI developers defend against Skeleton Key?
Microsoft recommends defense in depth rather than relying on a single refusal instruction. The controls operate at different points in the request lifecycle:
| Control point | Recommended practice | What it addresses |
|---|---|---|
| Input | Detect and filter harmful intent, including requests aimed at circumventing safeguards. | Stops or routes suspicious behavior-change attempts before they reach the model. |
| System instruction | Give the model explicit, durable instructions to reject attempts to undermine its safety rules. | Reduces the chance that a user message is treated as a policy update. |
| Output | Classify generated content against the application’s safety criteria before displaying or acting on it. | Catches unsafe material even when the model has already produced it. |
| Service monitoring | Track adversarial examples, content classifications, and abuse patterns with detection systems separate from the potentially manipulated model. | Finds repeated or coordinated probing that a single-request filter can miss. |
Microsoft describes these layers, along with its own software updates and responsible disclosure to other providers, in its Skeleton Key mitigation report (Microsoft Security Blog).
A practical test and response workflow
- Define the policy boundary. Write down which content the application must refuse and which warnings, transformations, or safe alternatives are allowed.
- Build multi-turn adversarial tests. Include conversations that first request a rule change and then make a prohibited request. Test fresh sessions and long sessions.
- Evaluate every layer. Record the input filter decision, system instructions in effect, model response, output classifier result, and any tool calls.
- Fail closed for high-risk actions. Require a separate authorization step before generated text can trigger tools, send messages, change records, or access sensitive resources.
- Monitor patterns, not only single prompts. Alert on repeated rule-change language, rapid category switching, retries after refusals, and attempts to evade classifiers.
- Re-test after changes. A model update, system-prompt edit, provider policy change, or new tool can alter the outcome; retain dated test records for each deployment.
Azure-specific options
For Azure applications, Microsoft names Azure AI Content Safety Prompt Shields for jailbreak and indirect-prompt-injection detection, risk and safety evaluations in Azure AI Studio, restrictive filter thresholds, and monitoring such as Microsoft Defender for Cloud. Feature names, defaults, thresholds, and regional availability can change, so verify the current Azure documentation and portal before relying on a particular setting. Microsoft introduced Prompt Shields in its March 28, 2024 announcement: Azure AI announces Prompt Shields.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsWhat teams should verify before calling a system protected
- Which model snapshot and provider endpoint are actually in production.
- Whether users can supply system-level instructions, tool arguments, uploaded files, or other higher-priority inputs.
- Whether input and output filters inspect the entire conversation, not only the latest message.
- Whether unsafe output is blocked before it reaches a user, tool, log, or downstream automation.
- Whether alerts and incident procedures cover repeated jailbreak attempts.
- Whether the provider has published a current remediation statement for the exact model and interface you use.
The June 2024 disclosure documents Microsoft’s testing and the mitigations it announced at that time. It does not establish the present remediation status of every third-party model listed above; that requires current provider documentation or testing.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

