Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Perturbation probing is a proposed way to test whether particular feed-forward network (FFN) neurons causally influence a chosen model behavior. In experiments reported by its authors, some behaviors were associated with small, concentrated neuron groups; other behaviors varied by model or resisted the same intervention. The results are evidence about specific models, tasks and outcomes—not proof that LLM safety generally rests on a few fragile neurons or a comprehensive measure of safety.

What perturbation probing tests

The method begins with a targeted behavior and uses two forward passes per prompt, without backpropagation, to generate causal hypotheses about FFN neurons. The authors then test candidate neurons through an intervention sweep. Their abstract describes about 150 intervention passes amortized across the identified neurons; the two-pass figure is per prompt, not the total compute for the study. The method is therefore an internal circuit diagnostic, not simply a benchmark that scores model answers from the outside. The paper by Hongliang Liu, Tung-Ling Li and Yuhao Wu was submitted to arXiv on April 30, 2026.

The authors report examining eight behavioral circuits across 13 models and four architecture families. They distinguish “opposition circuits,” which they associate with reinforcement learning from human feedback (RLHF) suppressing a pre-training tendency, from routing circuits for pre-training behaviors distributed through attention. Their examples include a concentrated FFN bottleneck in Qwen and a normalization-shielded circuit in Gemma. Those differences matter: a neuron intervention that works in one model or circuit need not transfer to another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the reported experiments found

Each result below has a specific model, task and endpoint. A change in answer format, a reduction in sycophancy and an increase in factual correction are different outcomes; none alone establishes a model-wide safety rate.

Experiment Study context Reported result and what it measures
Refusal-template neurons Qwen3-4B; 520 AdvBench prompts The authors implicated about 50 neurons, or 0.014% of all neurons, in controlling the safety-refusal template. Ablating them changed the response format on 80% of prompts. They also report three harmful-compliance cases, all with disclaimers. The 80% figure is about response format, not safety removed. Liu, Li and Wu, 2026.
Multi-turn sycophantic capitulation Qwen3.5-2B; 30 questions In the reported experiment, the initial capitulation rate was 36.7%; intervening on 20 neurons reduced it to zero. This is the outcome on that experiment’s questions, not a general rate for other conversations or models. Liu, Li and Wu, 2026.
Factual correction 200 TruthfulQA prompts Amplifying 10 related neurons raised the reported factual-correction rate from 52% to 88%. This measures correction on those prompts, not overall truthfulness. Liu, Li and Wu, 2026.
Language selection 580 prompts; 19 models tested Direction injection switched English output to Chinese on 99.1% of prompts in the reported successful setup, but the intervention worked in only three of the 19 models. The authors observed success under conditions including bilingual training, an FFN-to-skip ratio between 0.3 and 1.1, and linear representability; it failed in the other 16 models and on math, code and factual circuits. Liu, Li and Wu, 2026.

Why a refusal-format change is not the same as a safety failure

The Qwen3-4B result is striking because a small group of neurons was implicated in a recognizable refusal template. But the reported endpoint was usually a change in response format, while harmful compliance remained near zero in the benchmark experiment, with three such cases and disclaimers. A model can stop using a familiar refusal phrasing without complying with a harmful request; conversely, a disclaimer does not by itself make harmful assistance safe. The experiment supports a narrower conclusion: the refusal presentation was concentrated and causally perturbable in that tested setup.

That distinction is essential when interpreting the word “fragility.” Perturbation probing asks whether an intervention can alter a targeted behavior under specified conditions. It does not, by itself, show how likely an attacker is to find or apply that intervention in a deployed model, whether the behavior would stay altered after other safeguards act, or whether the intervention generalizes to unseen prompts.

Rank #2
J. J. Keller 2024 OSHA Construction Safety Handbook, English
  • 2024 OSHA Construction Safety Book is the seventh edition with the new OSHA HazCom final rule on 5/20/24. While the rule takes effect 7/19/24, the compliance dates don’t begin until 1/19/26 per 29 CFR 1910.1200(j).
  • Construction Site Book offers quick access to essential OSHA regulations, jobsite hazards, and practical safety tips. It also helps employees identify hazards and prevent injuries and illnesses.
  • Features easy-to-read format, full-color images, chapter quizzes with answer key, and comes in a compact size making it a convenient reference for employees.
  • Critical topics include Confined Space Entry; Cranes & Derricks; Electrical Safety; Emergency Response; Ergonomics & Back Safety; Excavations; Fall Protection; First Aid & Bloodborne Pathogens; HazCom; Health & Wellness; Jobsite Exposures; Lockout/Tagout; Ladders & Stairways; Materials Handling/Storage; Motor Vehicles; PPE; Scaffolds; Site Safety & Security; Slips, Trips & Falls; Tool Safety; Welding, Cutting & Brazing; and Work Zone Safety.
  • Specifications: 5 1/4” x 7 1/4", English, Soft bound. 7th Edition. Copyright 2024.

What the FFN-to-skip ratio can—and cannot—tell you

Across the 13 models, the authors report that an FFN-to-skip signal ratio distinguished circuit structures and helped predict an appropriate intervention. They frame it as a circuit-structure and intervention diagnostic. That is not the same as validating a universal safety score: the reported evidence does not establish that the ratio predicts overall safety in production, across every behavior, model or deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The contrast between the concentrated Qwen examples and language steering’s success in only three of 19 models illustrates why a single summary number would be misleading. Architectural variation, the specific behavior under test and whether the circuit is amenable to the intervention all affect the result.

How to compare it with other model evaluations

Perturbation probing complements external evaluation rather than replacing it. Red-team prompts and benchmark tests observe model outputs; perturbation probing forms hypotheses about internal circuits and tests them by intervening on model components. The latter requires access to inspect and alter model weights, so it is not generally available when evaluating a closed model through an API alone.

  • Check the endpoint. Record whether the result concerns refusal wording, harmful compliance, sycophancy, factual correction or another behavior. Do not substitute one for another.
  • Check scope. Note the model, architecture, benchmark, prompt count and intervention tested. These results do not establish coverage of every behavior or adversarial strategy.
  • Check repeatability. A circuit result may need to be tested again after fine-tuning, pruning, quantization or other deployment changes. The cited results do not establish persistence across those changes.
  • Use complementary evidence. Internal interventions can help explain a targeted behavior, while external tests can assess observable performance in a broader deployment context. The paper does not report a head-to-head comparison proving superiority over a comprehensive set of red-team methods.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Practical takeaway for deployment teams

Perturbation probing is best treated as a pre-deployment diagnostic for a defined behavior when a team can inspect and intervene on the model. Its results can inform an investigation into how a model produces a refusal or other response, but they do not certify that the model is safe. Unit 42 recommends pairing internal evaluation with external content filters and runtime guardrails; that is the publisher’s deployment advice, not evidence that any particular product is independently superior. Unit 42’s August 2026 explanation describes the approach and that defense-in-depth recommendation.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.