Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ShadowLogic is a way to hide a backdoor inside an AI model’s computational graph. The graph can detect a chosen trigger and route inference to attacker-defined behavior, while ordinary inputs still follow the model’s usual path. It does not require conventional injected code or a large poisoned training set, but it does require someone to tamper with the model artifact or its graph.

What ShadowLogic is—and what “codeless” means

A neural network’s computational graph describes the operations used during inference and how data flows between them. ShadowLogic adds graph operations that detect a condition and a branch that changes what the model does when that condition is met. Without the trigger, inference can continue along the normal path.

“Codeless” means the hidden behavior is expressed through model-graph operations rather than conventional executable code injected into the model. It does not mean that the attack takes no expertise, tooling, or access: the attacker must alter the model artifact or its graph. HiddenLayer introduced ShadowLogic in 2024; a peer-reviewed paper in the Proceedings of Machine Learning Research (PMLR) followed in 2025.

How the backdoor is triggered

The attacker adds a detector and a branch

The trigger detector recognizes an input condition. A conditional branch then directs execution to the attacker’s intended output or behavior when the condition matches. The added logic can be obfuscated to resemble ordinary model operations, making it less obvious from a casual inspection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall

Triggers can take different forms

HiddenLayer’s 2024 demonstrations included a red-pixel trigger for ResNet, trigger logic in a YOLO object-detection model, and controlled-token behavior in Phi-3. The described trigger possibilities also include keywords, sentences, checksums, and a separate embedded model. A trigger need not be a conspicuous image pattern; it can be a condition in text or other input data.

How ShadowLogic differs from training-time poisoning

Traditional data-poisoning backdoors are planted during training by contaminating training data. ShadowLogic instead modifies the computational graph in a model artifact, so the insertion point is after training. That distinction changes which part of the supply chain must be trusted.

Comparison ShadowLogic graph backdoor Training-time data-poisoning backdoor
Insertion point Graph in the model artifact, after training Training data or process, during training
Access implicated Access to alter the model artifact or graph Access to influence the training pipeline or its data
Trigger behavior Graph logic can test conditions such as pixels, words, sentences, or checksums Varies by attack; a directly comparable trigger range is not stated in the cited ShadowLogic sources
Persistence after fine-tuning or conversion HiddenLayer reports persistence in its experiments Not stated in the cited ShadowLogic sources
What ordinary testing can miss Tests without the trigger may exercise only the normal path Not stated in the cited ShadowLogic sources

The comparison does not mean every graph modification is malicious, or that every data-poisoning attack behaves alike. The distinctive ShadowLogic concern is conditional behavior inserted into the graph with minimal parameter changes, rather than a backdoor that depends on poisoned training examples.

What published experiments show

The reported figures are results from specific experiments, not a general estimate of how often ShadowLogic succeeds. The available summaries do not specify every model, dataset, or test condition for each metric, so those details should not be inferred.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Source and year Reported result Scope
PMLR, 2025 Greater than 60% attack success rate for further malicious queries The paper reports ShadowLogic implementations using ONNX graph manipulation in Phi-3 and Llama 3.2; the cited summary does not assign this rate to a specific model or provide further test conditions.
HiddenLayer, 2025 76.77% clean accuracy and 100% backdoor-trigger accuracy for the base model Experiment-specific; model identity and other test conditions are not stated in the cited summary.
HiddenLayer, 2025 77.43% clean accuracy and 100% trigger accuracy after fine-tuning the ShadowLogic model Experiment-specific; model identity and fine-tuning details are not stated in the cited summary.
HiddenLayer, 2025 35.68% trigger accuracy in the fine-tuning-only comparison after clean fine-tuning Experiment-specific comparison; model identity and other test conditions are not stated in the cited summary.

These results illustrate why ordinary accuracy alone is not a sufficient check: a model can retain strong clean performance while producing the attacker’s behavior on triggered inputs.

Why fine-tuning and conversion do not guarantee removal

HiddenLayer reports that its graph backdoor persisted through fine-tuning and model-format conversion, while ordinary model performance remained effectively unchanged. The 2025 measurements above are examples from those experiments, not a guarantee that every backdoor survives every fine-tuning method or conversion pipeline.

This persistence makes ShadowLogic a supply-chain concern at several handoffs: when a model is downloaded, converted, fine-tuned, or deployed. A clean validation set can miss a dormant branch if it never includes inputs that activate the trigger. Conversion or fine-tuning should therefore be treated as transformations that require renewed validation, not as proof that a model is clean.

What changes when the model controls an AI agent

In an agentic system, a language model may return a structured tool call—such as a destination and arguments—that framework software then executes. HiddenLayer’s January 2026 Agentic ShadowLogic work extends the graph-backdoor idea to this setting: a graph-level branch could alter a tool call’s destination, arguments, or action.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That creates a route from hidden model behavior to an action performed by downstream software. It is a demonstrated research risk, not evidence in the cited work of a confirmed criminal campaign or a known in-the-wild incident. The practical implication is that an agent framework should apply its own authorization and validation rules to tool calls rather than treating model output as inherently safe.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to inspect and validate an ONNX model

ONNX makes a model’s computational graph available for inspection, and the PMLR paper describes manipulating graphs in ONNX. Inspection can help identify suspicious changes, but the cited sources do not establish that any one scanner or review technique guarantees detection.

  1. Verify provenance. Obtain the model from an expected source, record its hash, and compare that hash with a trusted value when one is available. A hash can reveal that a file changed relative to a known copy; it cannot establish that the original copy was benign.
  2. Compare the graph with a trusted baseline. Inspect unexpected nodes, branches, or operations, especially when a known-good version of the same model is available. Review changes introduced during conversion or fine-tuning as well as the original artifact.
  3. Test more than ordinary inputs. Keep clean validation examples, but also test plausible trigger classes relevant to the model—such as unusual pixel patterns or targeted text conditions. A test suite without the trigger cannot establish how the model behaves when the hidden condition is met.
  4. Repeat checks after transformations. Reinspect and validate the artifact after format conversion or fine-tuning, then validate the exact file that will be deployed.
  5. Put policy checks in front of agent actions. For tool-using systems, independently validate allowed destinations, arguments, and actions before execution. The policy layer should enforce what the application permits even if the model proposes something else.

These steps reduce reliance on a model’s ordinary behavior as evidence of integrity. They are layered checks, not a claim that graph inspection or trigger testing alone can prove a model is free of hidden behavior.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.