Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Red-team the AI system as it will actually be deployed—not just the model in a chat window. Define written authorization and boundaries, map the system’s data and trust boundaries, test AI-specific and conventional security risks across the model, application, integrations, infrastructure, and runtime controls, then document, remediate, and retest findings before making a release decision.

What should an AI red-team exercise cover?

In the AI context, red-teaming is a structured testing effort, often using adversarial methods, to find flaws and vulnerabilities in an AI system—including unexpected or undesirable behavior and risks from misuse. NIST’s definition, attributed to NIST AI 100-2e2025 in the NIST CSRC glossary, describes the activity; it does not make red-teaming a guarantee that a system is safe.

Set the boundary around the deployed system and its context. Depending on the design, that can include:

  • The model, its configuration, and any fine-tuning or other adaptation.
  • The application that constructs prompts, handles outputs, and enforces permissions or policy.
  • Connected tools, APIs, retrieval systems, and the data they can access.
  • Training, evaluation, and deployment pipelines, plus supporting software, hardware, and infrastructure.
  • Runtime controls such as access restrictions, output checks, logging, detection, and incident response.
  • The users, tasks, and operating environment for which the system is intended.

This broad boundary matters because an AI system has ordinary security risks as well as AI-specific ones. NIST identifies confidentiality, integrity, and availability concerns involving systems, training and output data, and underlying software and hardware. A model-level test alone can miss a flaw in the application, a connected service, or a control that is supposed to limit the model’s access.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How does red-teaming fit with other AI evaluations?

NIST’s AI Risk Management Framework Generative AI Profile (AI 600-1) treats model testing, red-teaming, and field testing as distinct evaluation levels. They answer related but different questions; a red-team exercise should be part of a wider evaluation and security program, not a substitute for either.

Evaluation What it helps examine How it differs
Model testing Model behavior against selected tests or evaluation criteria. Can examine the model without exercising the full deployed system and its controls.
Red-teaming How adversarial or misuse-oriented testing can expose flaws, vulnerabilities, and undesirable behaviors in the AI system. Organized around finding weaknesses; scope can include the model, application, integrations, infrastructure, and controls.
Field testing System behavior in its operating context. Evaluates use in the field rather than relying only on a pre-deployment adversarial exercise.

NIST says red-teaming can happen before or after a model or system is made broadly available; this article focuses on pre-deployment. Even a well-run pre-deployment exercise cannot establish that a system is risk-free, so pair its findings with ordinary security engineering and ongoing monitoring.

How should you scope and authorize the exercise?

Before testing, write down what is authorized and how the exercise will be run. OWASP’s red-teaming guidance emphasizes authorization, data logging, reporting, deconfliction, communications and operational security, and data disposition. A practical scope record should make those decisions explicit.

  • Purpose and target: Identify the system, version, intended users and tasks, deployment context, and questions the exercise is meant to answer.
  • In-scope environments and access: Name the staging or other approved environments, components and integrations in scope, tester permissions, and any out-of-scope systems.
  • Time and coordination: Set the test window, contacts, escalation route, deconfliction process, and communications rules so testing is distinguishable from an incident or another team’s work.
  • Data handling: Specify what testers may access or retain, what may appear in logs or reports, who can see the records, and how test data and evidence will be disposed of.
  • Safety and stop conditions: Define conditions that require pausing or stopping—for example, activity outside the authorized environment or an unexpected impact on service or data—and name who can make that call.
  • Evidence and reporting: Agree where to record test cases and findings, who receives the report, and how urgent findings will be escalated.

Authorization should cover the actual activities and connected resources being tested. If scope or access changes, resolve that with the system owner before proceeding rather than treating an adjacent system as implicitly in scope.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do you threat-model the deployed system?

Map how data and authority move through the system, then use that map to choose tests. Record the assets that matter, the people and components that can reach them, the trust boundaries between components, and the actions the AI system can take. Include realistic misuse as well as failures in ordinary software and infrastructure.

  1. List assets and impact: Identify sensitive data, credentials or other secrets, model and application components, service availability, and actions whose misuse could cause harm.
  2. Draw the data and control flow: Trace user input through the application to the model and onward to retrieval sources, tools, APIs, output handling, and logs. Mark where access is granted, constrained, or checked.
  3. Identify actors and trust boundaries: Consider intended users, users with different privileges, external content sources, connected services, and the people or systems that administer or deploy the service.
  4. Describe plausible failure and misuse paths: Ask how an adversarial input, compromised or untrusted content, excessive permission, or a weakness in the surrounding software could affect confidentiality, integrity, or availability.
  5. Turn risks into test objectives: For each important asset and boundary, state what a tester should try to make happen and what control should prevent, detect, or limit it.

For example, if an assistant can use a connected tool, the test objective should cover both its behavior and whether the application actually limits the tool’s permissions. A safe-sounding model response does not by itself show that the integration enforces the intended access boundary.

Which attack paths should the team test?

Choose cases that match the architecture, data, users, and likely misuse of this system. The following are test areas, not an exhaustive threat list. Test the controls around each path as well as whether an adversarial input can elicit undesirable model behavior.

  • Prompt injection and adversarial inputs: Check whether untrusted instructions or content can override the intended task, cross a trust boundary, or influence tool use in an unsafe way.
  • Unsafe cyber assistance: Assess whether the system provides malicious-code generation or materially enhances phishing when prompted in ways relevant to the intended use. Evaluate the applicable safeguards and escalation or monitoring controls.
  • Sensitive or training-data exposure: Test whether prompts or system behavior can reveal information the user should not receive, including data that should remain private.
  • Data poisoning: Where training, fine-tuning, retrieval, or other data-ingestion paths are in scope, examine whether manipulated data can undermine intended behavior or controls.
  • Membership inference and model extraction: Assess whether the system or its interfaces expose signals that could reveal whether data was used in training or enable unauthorized extraction of model behavior or information.
  • Agent and connected-tool misuse: Where tools or agents are present, test whether the system can be induced to take an unintended action, access data beyond its authority, or bypass application-level restrictions.
  • Conventional software and infrastructure security: Include confidentiality, integrity, and availability risks in the application, supporting services, pipeline, infrastructure, and handling of training or output data.
  • Guardrails and post-training changes: Check access controls, output checks, detection, and response. If the system has been fine-tuned or otherwise adapted, verify that safety and security controls still work after the change.

For every case, record the intended policy or boundary, the control expected to enforce it, the adversarial condition being tested, and the impact if that control fails. Avoid testing a component in isolation when the real risk depends on a chain of model behavior, application logic, permissions, and external services.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Who should participate, and which red-team approach fits?

Choose participants for the system and its use context. NIST describes expert, general-public, combined, and human/AI red-team approaches. These are options with different strengths, not interchangeable guarantees; results need analysis before they inform organizational governance or risk decisions.

Approach Useful when Trade-off to plan for
Expert-led The exercise needs cybersecurity expertise, deployment knowledge, or careful interpretation of complex findings. Specialist insight may not capture every way representative users will interact with the system.
General-user participation Realistic user perspectives or less specialized interaction patterns are important to the use case. Participants may need guidance, safe access, and support to make findings reproducible and interpretable.
Combined team The system needs both technical attack analysis and perspectives grounded in its intended use. Coordination is needed to keep tests within scope and relate different observations to the same controls and risks.
Human/AI-assisted AI assistance can support coverage or exploration alongside human testers. Generated test activity and results still require human review, authorization boundaries, and contextual interpretation.

Whichever approach you choose, match expertise to the deployment domain and ensure the people evaluating results understand the system’s controls and intended use. A collection of successful prompts is not yet a risk assessment.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should you measure and document results?

Use metrics that answer the exercise’s stated questions, and preserve enough evidence for another person to reproduce and assess a finding. OWASP defines attack success rate, also called jailbreak success rate, as the percentage of adversarial inputs that successfully exploit vulnerabilities or elicit undesired behavior. Treat it as a use-case-specific measure, not a universal release threshold.

For each finding, capture:

  • The test case and conditions needed to reproduce it, including relevant model and configuration versions.
  • The observed output or behavior, recorded in line with the exercise’s data-handling rules.
  • The expected behavior or control and how the observed result differed.
  • The affected asset, component, trust boundary, or control, plus the plausible impact.
  • The severity rationale and the evidence supporting it.
  • A recommended remediation, an accountable owner, and a way to verify the fix.

Choose metrics that fit the objective—for example, whether adversarial inputs elicited a defined undesired behavior—then state exactly what was measured and under what test conditions. Do not turn a rate from one test set into a claim that the system will behave the same way for all users or inputs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How should findings affect the deployment decision?

  1. Triage and assign: Route each finding to an owner who can address the affected model, application, integration, infrastructure, or operational control.
  2. Remediate and retest: Test the proposed mitigation against the reproducible case and relevant variations. Check that a fix has not weakened another safeguard or changed access behavior unexpectedly.
  3. Record residual risk: Document what remains unresolved, its impact and rationale, and who is accountable for accepting it under the organization’s risk process.
  4. Make a release decision: Use the analyzed findings alongside other evaluation and security evidence to decide whether to deploy, restrict, delay, or add controls and monitoring.

NIST advises further analysis of red-team results before incorporating them into governance and risk management. The exercise therefore informs a deployment decision; it does not make that decision automatically or certify the system as secure.

What does current guidance establish—and what does it not?

NIST’s Generative AI Profile is dated July 26, 2024. NIST’s AI security page, updated August 14, 2026, describes the security area as active and notes that existing guidance does not comprehensively address all AI attack surfaces and machine-learning attacks. OWASP’s red-teaming material offers practical scoping, documentation, and measurement guidance; its recommendations should be applied to the architecture and use case at hand.

These sources do not establish one exhaustive list of threats, one attack-success rate that is a suitable pass mark for every system, or a certification that an exercise can award. Nor does this methodology determine the legal obligations for a particular organization or jurisdiction. Tailor scope, tests, and release criteria to the model type, architecture, deployment context, risk tolerance, applicable obligations, and authorized access.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.