Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Adversarial attacks can get tested AI image generators to create not-safe-for-work (NSFW) images despite safety measures. They do this by targeting different parts of the generation process: altering prompts, combining text and images, changing an image supplied to an image-to-image model, or exploiting interactions between several safeguards. Published results apply to the particular models and test setups studied—not necessarily to the versions of commercial services available today.

How AI image-generator safeguards can be bypassed

Image generators may use more than one safety check. A prompt filter can block or alter text before generation; a concept-erasure component can suppress certain content in the model; and an image checker can inspect the result afterward. These measures address different points in the pipeline, so testing a prompt filter alone does not show how the full system will behave.

Researchers have studied attacks against each of these points. Some attempt to make a harmful request pass as acceptable text. Others supply visual input, or probe how defenses work together. The underlying issue is that a filter or checker must interpret inputs and outputs that can be indirect or unexpected; a safety system can miss cases even when its intended rules are clear.

What kinds of attacks have researchers tested?

Iteratively changing text prompts

SneakyPrompt automates the search for prompt wording that gets past a filter. As described by IEEE Spectrum, the method repeatedly replaces filtered words with alternatives, queries the generator, and adjusts the prompt based on the output. In the study’s reported setup, it achieved an average bypass rate of about 96% on Stable Diffusion and roughly 57% on DALL·E 2. The researchers estimated that prior manual attempts against Stable Diffusion succeeded about 33% of the time. These figures are experimental results, not current success rates for those services.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

Prompt-learning attacks take a related but distinct approach. The authors of PLA: Prompt Learning Attack against Text-to-Image Generative Models, published in the ICCV 2025 proceedings, study black-box attacks intended to bypass text-to-image safety mechanisms. Black-box means the attack is designed around querying a system rather than assuming access to its internal model. IEEE Spectrum also reported on a later automated method called Jailbreaking Prompt Attack (JPA), which researchers said worked on the open and closed models they tested, including Stable Diffusion, DALL·E, and Midjourney. That report does not establish how current versions behave.

Combining text and visual input

MMA-Diffusion: MultiModal Attack on Diffusion Models, presented at CVPR 2024, studies attacks that use both text and image input to get around prompt filters and post-generation checkers. This matters because a system that scrutinizes text may still have a separate path for visual information. A defense that works on one input type is not, by itself, evidence that the combined system is robust.

Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

Altering an image given to an image-to-image model

Not every image-generation request begins with text alone. Image-to-image systems use an existing picture as input, often alongside a prompt. In AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion Models, published in ICML 2025, researchers describe optimizing the input image to induce NSFW output without changing the text prompt. The paper reports attacks against defenses including Safe Latent Diffusion. The attack’s target is therefore a different part of the workflow from methods that search for alternative prompt words.

Targeting multiple defense layers

Some attacks focus on how safeguards interact, rather than trying to defeat only one. The NeurIPS 2025 paper Transstratal Adversarial Attack: Compromising Multi-Layered Defenses in Text-to-Image Models describes prompt filters, concept erasers, and image filters as sequential layers. Across 14 text-to-image models and 17 safety modules, its authors report an 85.6% average attack success rate in their evaluation, exceeding the compared state-of-the-art methods by 73.5%. These are benchmark-specific results; they should not be read as a forecast of how often a current consumer product can be bypassed.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

Why the reported percentages cannot be compared directly

A percentage is meaningful only alongside the test conditions that produced it. The studies vary in what the attack can submit, which defenses are present, and which models are evaluated. They may also define a successful attack differently. For example, a prompt-only experiment is not directly comparable to one that can alter an input image or target several defense layers.

Study Attack input or focus What the reported evidence establishes
SneakyPrompt, reported by IEEE Spectrum Automated, iterative text-prompt changes About 96% average bypass on Stable Diffusion and roughly 57% on DALL·E 2 in the study’s tested setup; the report also estimates about 33% for earlier manual attempts on Stable Diffusion. The figures are not a current product scorecard.
PLA, ICCV 2025 Prompt-learning attack designed for black-box access Studies prompt-based bypasses of text-to-image safety mechanisms; the cited source does not establish a directly comparable rate for current commercial versions.
MMA-Diffusion, CVPR 2024 Text and image inputs together Studies multimodal attacks against prompt filters and post-generation checkers.
AdvI2I, ICML 2025 Adversarially altered image input Reports attacks on image-to-image diffusion models, including against Safe Latent Diffusion; the cited source does not establish a directly comparable rate for current commercial versions.
Transstratal, NeurIPS 2025 Attacks across sequential safety layers Reports an 85.6% average attack success rate across the study’s evaluation of 14 models and 17 safety modules.

The figures in this table come from different papers and experiments, not a shared benchmark. They cannot be used to rank the approaches or infer the odds of bypassing a particular service. Model versions, access assumptions, defense configurations, and success criteria all matter.

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How red-teaming helps identify overlooked cases

Some harmful outputs arise from prompts that are adversarial in effect but not obviously suspicious. Google Research’s 2024 Adversarial Nibbler work focuses on these “implicitly adversarial prompts.” It reports a collection of more than 10,000 prompt-image pairs with machine safety annotations, including a 1,500-sample subset with richer human annotations of harm types and attack styles.

Red-teaming is useful because it can reveal failure cases that straightforward, known-bad prompts may not expose. The Nibbler authors emphasize continual auditing and adaptation as new vulnerabilities emerge. The Transstratal findings add a related lesson: safeguards that appear effective when evaluated separately may still have weaknesses when combined.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

What these findings mean for users and developers

For people using image generators

Safety controls reduce risk, but the cited studies show they are not guarantees. They also do not provide a reliable way to predict the behavior of a particular service today. Products can change their models, policies, and safeguards, and the reviewed sources do not provide a comprehensive, independently verified test of current commercial versions as of October 5, 2026.

For teams building or evaluating generators

  • Evaluate the full pipeline, including input filters, model-level safeguards, and output checks.
  • Test text, image, and combined-input paths when a product supports them.
  • Use red-team prompts that include indirect or unexpected cases, not only obvious examples.
  • Report model versions, defense configurations, access assumptions, and success definitions so results can be interpreted accurately.
  • Repeat evaluations as models and safeguards change; passing an earlier test does not establish present-day robustness.

What remains uncertain

The studies establish that tested systems and defenses can be bypassed under their experimental conditions. They do not show that every generator is vulnerable in the same way, or how frequently current commercial systems would fail. Answering that requires fresh, version-specific testing with clearly stated models, safeguards, and success criteria.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.