What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In tested cases, abliteration has reduced an AI model’s tendency to refuse harmful requests while leaving selected capability scores unchanged. That is a narrower result than saying the model keeps all its knowledge or behaves the same in every other way: the findings depend on the model, the edit, and the tests used.
What does abliteration do to an AI model?
Abliteration is a family of refusal-reduction techniques applied to open-weight models. These edits target representations or weights associated with refusing requests. They can change a model’s refusal behavior, but the term does not describe one standardized procedure with predictable effects.
Here, “obedience” means refusal behavior, not general instruction following. A model that refuses fewer requests has not necessarily become better at following every instruction, nor does that result show that all its other behavioral tendencies are unchanged.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Does removing refusals make a model less knowledgeable?
Not necessarily on the tasks measured. Anthropic reports that, after abliteration of GLM-5.3, refusal rates fell on JailbreakBench, HarmBench, and StrongREJECT, while the standard and edited versions received the same GPQA-Diamond score. On a tested CyberGym subset, the edited model’s score was a few percent lower. These are results from Anthropic’s evaluation of this model and setup, not proof that its entire knowledge or capabilities were preserved. Anthropic’s GLM-5.3 evaluation also reports that its abliteration process took about 2,200 GPU hours and approximately $4,400 in computation cost; those figures describe that team’s setup, not a typical cost. (Anthropic)
#1 Best Overall
Refusal and capability scores measure different outcomes. A model can change substantially on the first while appearing stable on a particular knowledge benchmark. A stable score on one benchmark cannot establish that everything the model knows, can do, or tends to do remained intact.
What do broader evaluations show?
Safety-pretraining configurations can respond differently
Agnihotri and colleagues evaluated 20 systems—10 base models and their abliterated counterparts—using 100 prompts per system, divided evenly between harmful and harmless prompts. The team used multiple judges and checked a small subset with human labels. It reported that safety-pretraining variants combining signals such as rephrasing, metatags, and refusals were more resilient to abliteration than simpler variants. The study’s prompt set is a bounded evaluation, not a measure of performance across every real-world request. Read the 2025 preprint or the Keuper Labs project page.
Other behavior may shift too
A July 2026 preprint by Aleksander Fafuła reports disposition changes after abliteration in two model families on a financial decision task. Across 60 Warsaw Stock Exchange equities over 18 weeks, the study recorded 21,600 decisions and reported greater optimism and changes in expressed uncertainty. The direction of confidence effects differed between the model families. This is preliminary, task-specific evidence: it does not establish how models behave generally, but it shows why unchanged scores on selected capability tests cannot guarantee that unrelated behavior stayed fixed. Read the preprint.
Is removing false refusals the same as removing refusals broadly?
No. A model may wrongly refuse a harmless request, and reducing those false refusals is a different objective from suppressing refusals to harmful requests. Wang and colleagues’ ICLR 2025 paper proposes removing a single false-refusal vector, with the stated aim of reducing refusals to safe requests while preserving harmful-request safety and general capability. That targeted approach addresses over-refusal; it should not be treated as equivalent to broadly reducing refusals. Read the ICLR 2025 paper.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How should you judge whether an abliterated model changed?
A refusal-rate result by itself is insufficient: it cannot show whether an edit specifically reduced unwanted refusals or degraded behavior more generally. A meaningful evaluation should report the model and edit, test setup, harmful-request refusals, harmless-request false refusals, and capability measures together.
Quick Recap
Best Value
- Identify the exact model and edit. Abliteration methods and safety-pretraining configurations differ; results from one setup should not be assumed to transfer to another.
- Test both harmful and harmless prompts. The two categories reveal whether fewer refusals reflect the intended change or a broader loss of discrimination.
- Use more than one outcome measure. Report refusal behavior separately from capability benchmarks, and consider behavioral measures relevant to the model’s intended use.
- Read results in context. The cited studies use different models, prompts, judges, and procedures, so their results are not directly comparable as if they were one standardized test.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

