Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI red teaming actively probes an AI model, application, or deployment for weaknesses—for example, whether chosen adversarial inputs can elicit harmful outputs or bypass safeguards. A successful exploit shows a failure under the conditions tested. A clean test does not prove the system is safe or that other failures do not exist.

What does an AI red-team assessment actually check?

The scope depends on the target and the access granted. A report should say whether the team tested a model by itself, an application built around it, a deployed system, or some combination.

Model behavior and safeguards

Testers try prompts, jailbreaks, and other adversarial inputs to see how the model responds and whether the specific safeguards under examination withstand those inputs. In a joint U.S. and UK AI Safety Institute evaluation of an upgraded Claude 3.5 Sonnet, machine-learning experts attempted to develop jailbreaks that would induce answers to malicious requests. NIST’s account reports that most publicly available jailbreaks tested by the U.S. institute circumvented the built-in safeguards examined in that exercise. This finding applies to the tested model version, jailbreak set, and safeguards—not to every model or later version. NIST’s account of the evaluation describes its scope and caveats.

Applications, tools, and infrastructure

When scope and access allow, a red team can probe interfaces, connected tools, infrastructure, and interactions among system components, not just model responses. Microsoft’s response to NIST describes red teaming as probing for harmful capabilities and outputs as well as infrastructure threats, spanning responsible-AI and cybersecurity concerns. That is Microsoft’s practitioner perspective, not a binding NIST standard. Microsoft’s NIST submission discusses that broader scope.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What can’t red teaming prove?

It cannot certify that an AI system is safe

Red teaming is a way to discover weaknesses, including unexpected ones, rather than to create a complete inventory of a system’s capabilities or risks. The prompts and scenarios selected, the team’s expertise, the time available, and its permissions and tools all affect what it can find. If a test uncovers no issue, that means only that it did not surface one under those conditions.

NIST’s account of the joint U.S. and UK evaluation says that “the results of this evaluation cannot on their own determine the model’s risks.” The warning refers to that evaluation’s safeguard results: the work took place over a limited period with finite resources, and judgments about harmfulness can be subjective and jurisdiction-dependent.

It cannot measure every risk in ordinary or ongoing use

A scoped, point-in-time exercise does not by itself establish how prevalent a known risk is in normal use, continuously measure behavior after release, detect or prevent malicious activity in production, or provide every sector-specific impact assessment. Those questions need other evidence, such as systematic measurement, production monitoring and auditing, and domain-specific impact reviews. Microsoft’s NIST submission also discusses these complementary practices.

How do red-team tests fit into a broader evaluation?

Red teaming is one evidence stream, not a substitute for other forms of testing. NIST’s AI Risk Management Framework Assessing Risks and Impacts of AI (ARIA) materials distinguish model testing, red teaming, and field or user testing. Its 2025 ARIA 0.1 pilot report says five organizations submitted seven AI applications and describes three evaluation levels: model testing, red teaming, and field testing. The September 2026 planning manual describes a holistic approach combining Model Testing, Red Teaming, and User Testing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These levels answer different questions: model testing measures defined behavior, red teaming searches for weaknesses through adversarial probing, and field or user testing examines performance in context. NIST’s ARIA program page, pilot report, and evaluation planning manual describe the program’s evaluation approach.

What do published test figures actually tell you?

Specific results illustrate a bounded evaluation; they are not general benchmarks for red-team effectiveness or overall system security.

Reported result What it applies to
40 cybersecurity challenges; 32.5% task success The upgraded Claude 3.5 Sonnet on the U.S. AI Safety Institute’s public cybersecurity challenge suite, as reported in NIST’s account of the 2024 evaluation.
47 cybersecurity challenges; 36% success on apprentice-level tasks The UK AI Safety Institute’s suite, comprising 15 public and 32 privately developed challenges, as reported in NIST’s account of the 2024 evaluation. The percentage is specific to apprentice-level tasks.
Five organizations; seven AI applications Submissions to NIST’s ARIA 0.1 pilot, reported in its November 2025 report.

The two cybersecurity results should not be treated as a direct comparison: the challenge sets, task levels, and evaluation conditions differ. NIST’s account also notes that smaller performance differences in its evaluation might fall within test margins of error.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should you compare two red-team reports?

A headline such as “passed red teaming” or a count of discovered issues is not meaningful without the test’s boundaries. Check whether the reports cover comparable targets, attack goals, and evidence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Target and version: Identify whether the subject was a model, application, or deployed system, and record its exact version and configuration.
  2. Scope and access: Check which interfaces, tools, permissions, rate limits, and system components were in scope, and whether the team received special access.
  3. Threats and harm definitions: Find out which adversaries and attack goals were considered, and how the test defined a harmful output or successful attack.
  4. Test design: Note whether the team used public or private cases, manual exploration or a repeatable suite, which domains it covered, and what it excluded.
  5. Evidence and outcomes: Look for how findings were validated, how severity was assigned, and whether mitigations were tested.
  6. Timing and uncertainty: Record when the test occurred and what resources were available. Look for confidence estimates or margins of error; limited time and resources can leave important uncertainty.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.