Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
An adversarial AI attack exploits weaknesses in how an AI system is trained, queried, deployed, or connected to other software. It may manipulate an input, compromise training data, expose information, disrupt availability, or misuse an application’s access to documents and tools. The model’s weights do not have to change for an attack to succeed.
For developers and system owners, the practical question is not only whether a model can be fooled. It is what an attacker can control, what the system can reach, and what harm follows when it behaves as the attacker intends.
What is an adversarial attack on AI?
Adversarial machine learning (AML) covers attacks that exploit the behavior or surrounding infrastructure of machine-learning systems. NIST’s Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (AI 100-2e2025), published in March 2025, organizes these attacks by type, learning context, attacker goal, capabilities, and knowledge. NIST says the taxonomy is intended to evolve as the field changes.
“The model” is therefore a useful shorthand, but not always the direct target. An attacker may instead target the data used to build it, the interface that accepts queries, the application that supplies context, or the system’s confidentiality, integrity, and availability. A security review should follow the full path from data and development through deployment and use.
#1 Best Overall
How can someone attack an AI model?
Classify a possible attack along four practical dimensions: the outcome the attacker wants, what access or control they have, when they act in the system’s lifecycle, and what the AI application can affect. Those dimensions help explain why a similar technique can have very different consequences in a standalone classifier and an agent with access to private files or external tools.
| Attack pattern | Typical attacker objective | Where it acts | Key distinction |
|---|---|---|---|
| Evasion | Integrity: obtain an incorrect or attacker-favored prediction | At inference, through inputs or their presentation | Does not require changing training data or model weights |
| Poisoning or a backdoor | Integrity or other compromised behavior | Training or other model-development inputs | A backdoor can make a behavior conditional on a trigger |
| Availability attack | Make a model or service unavailable or degrade its operation | Model, service, or supporting resources | Targets access or reliable operation rather than necessarily changing an answer |
| Privacy attack or model extraction | Infer information about training data, user data, or the model | Queries, outputs, or exposed data paths | Extraction seeks to learn or reproduce information about a model; it is not the same as stealing its weights |
| Prompt injection or jailbreak | Misuse: steer a generative system around intended instructions or restrictions | User prompts or content the application processes | Prompt injection supplies malicious instructions; a jailbreak is an attempt to bypass safeguards |
The categories can overlap. An attack may combine access to a training pipeline with a privacy objective, or use malicious retrieved content to prompt an agent to disclose data. NIST also distinguishes among system contexts such as base models, retrieval-augmented generation (RAG) applications, chatbots, and agents; a technique’s relevance depends on the system’s design and access.
Rank #2
How do attacks on predictive AI differ from generative AI?
Predictive AI: manipulate a decision or compromise its development
Predictive systems classify, score, detect, or estimate. Evasion attempts to change the result by manipulating an input or how the system receives it. The exact technique depends on the model and input modality, so an example from image recognition should not be assumed to apply to every classifier.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPoisoning acts earlier: an attacker influences data or other inputs used during model development, potentially compromising model integrity. A backdoor is one possible result, in which a trigger causes behavior that would not normally occur. Availability attacks instead seek to prevent or degrade use. Privacy attacks seek information about training data or a model; they do not necessarily recover complete records.
Rank #3
Generative AI: steer outputs, extract information, or misuse connected capabilities
Generative systems introduce attacks involving instructions and generated content, alongside familiar integrity, availability, and privacy risks. NIST’s generative-AI taxonomy includes direct prompting attacks, indirect prompt injection, jailbreaks, prompt extraction, training-data extraction, data or model poisoning, and leakage of data from user interactions. These labels describe different methods or targets, not interchangeable names for a model producing a false answer.
Prompt extraction aims to elicit information about prompts or instructions. Training-data extraction seeks information that may have been memorized during training. Privacy risk can also arise from user data handled during interaction. None of these categories, by itself, establishes that a complete private record can be recovered.
Rank #4
What is prompt injection, and how is it different from a jailbreak?
Prompt injection occurs when a generative system receives malicious instructions intended to redirect its behavior. A direct attack places those instructions in a user prompt. An indirect attack hides them in content the application retrieves or otherwise processes, such as a document supplied to a RAG system. The application may treat that content as context even though it comes from an untrusted source.
A jailbreak is an attempt to bypass a model’s restrictions or elicit disallowed behavior. The terms are related but not synonymous: prompt injection describes instructions used to steer a system, while jailbreak describes an objective or attempt to defeat restrictions. Neither term means ordinary hallucination, and neither requires poisoning the model during training.
Best Value
Why agents and connected applications raise the stakes
A model that only returns text has a different impact boundary from an application that can search private repositories, send messages, modify records, or invoke other tools. In a connected system, the attack path can run from malicious content to model behavior to an application action. The relevant question is not just whether the model followed an instruction, but what permissions the application gave it and what data or actions those permissions exposed.
Assume untrusted content may contain hostile instructions whenever a system retrieves or processes outside material. A prompt that labels text “untrusted” can help communicate intended boundaries, but it is not a guarantee that the model will enforce them. Keep the trust boundary and the permissions boundary aligned: retrieved content should not gain authority to trigger actions merely because a model has read it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How do you defend an AI model against attacks?
No single filter, training technique, or prompt design guarantees protection. NIST describes mitigations for prompt injection—including task-specific training, detection, input processing, and designs that distinguish trusted from untrusted content—but warns that current mitigations do not fully protect against every technique. Treat defenses as layers that reduce likelihood or impact, then test them against the system’s actual data paths and capabilities.
- Reduce exposure during development. Protect training and fine-tuning data pipelines, review data sources, and control who can change model-development inputs. These controls address poisoning risks before a model is deployed.
- Define trust boundaries. Identify which instructions and data are trusted, which are user-provided, and which come from retrieved or third-party content. Use input processing and task-specific training where appropriate, while treating them as risk reduction rather than proof of immunity.
- Constrain application permissions. Give models and agents only the access needed for their task. Separate access to untrusted sources from authority to take consequential actions, and expose tools through well-defined interfaces rather than broad, unnecessary permissions.
- Test for the outcomes that matter. Assess integrity, privacy, availability, and misuse risks in the relevant lifecycle stage and system context. Include tests for how the model and application respond to hostile inputs, retrieved instructions, and attempted access to protected information.
- Keep conventional security controls in scope. Secure the software, infrastructure, data, and interfaces around the model. AI systems still face confidentiality, integrity, and availability risks familiar from conventional systems, while AI adds model-specific attack surfaces and failure modes.
- Reassess as the system changes. New data sources, model versions, tools, permissions, and deployment settings can change the attack surface. Repeat evaluations when those changes alter what the system can access or do.
What NIST’s guidance establishes—and what it does not
NIST’s 2025 report provides a taxonomy and terminology for reasoning about attacks and mitigations; it is not an exhaustive catalogue of every attack or a claim that every listed technique applies to every AI system. NIST notes that the literature it considered included more than 11,354 arXiv.org references since 2021, as of July 2024. That figure describes the literature, not real-world incident counts or an increase in attacks.
NIST’s AI security and resilience work also describes Dioptra as a research testbed for assessing model vulnerabilities and the effectiveness of defenses. It is a resource for evaluation research, not a consumer security product or an endorsement of a commercial mitigation. NIST notes that conventional security frameworks do not comprehensively address several AI-specific attacks or the full complexity of AI systems, so model-focused testing belongs alongside—not instead of—sound software and data security.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

