Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI is used in quality engineering both to assist testing work and to help test products that contain AI. Generative AI can draft test ideas, help with automation, and summarize results—but its outputs need review against requirements and evidence. When the product itself uses AI, the team must also test risks related to its data, model, and behavior in context. In both cases, AI can support quality work; it does not prove that quality has been achieved.

Two different jobs: using AI for testing and testing AI

“AI in quality engineering” can mean either of two things. Keeping them separate helps teams choose the right tests and avoid treating an AI-generated artifact as proof that an AI-enabled product is reliable.

  • AI for testing: use AI tools to assist with activities such as analyzing requirements, drafting test cases, writing or maintaining automation, and summarizing test results.
  • Testing AI: evaluate an AI component or system as part of a product, including risks arising from its model, data, and behavior in its intended use context.

A team can use AI to help test conventional software without testing an AI system. It can also test an AI-enabled product without using generative AI to do the testing. Some teams do both, but each requires its own checks and evidence.

How generative AI can assist quality engineers

Generative AI is most useful as a drafting, analysis, or summarization aid. A person familiar with the requirements and product should check its work before that work becomes part of a test suite, defect record, or release decision.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.

Analyze requirements and acceptance criteria

A model can restate a requirement, flag ambiguous wording, suggest questions for a product owner, or propose scenarios. For example, given a requirement that an account locks after repeated failed sign-ins, it might suggest testing the threshold, a successful sign-in before the threshold, and behavior after a lockout. Those are candidate questions and cases—not an authoritative interpretation of the policy. A stakeholder must confirm the intended threshold, reset rules, and user-facing behavior.

Draft test cases and test-data ideas

AI can turn a description into candidate positive, negative, boundary, or exploratory tests and suggest data values to exercise them. Review each case for correctness, useful coverage, duplication, and traceability to a requirement or risk. Check test data for privacy and security concerns before using real or sensitive information in a prompt or test environment.

Help write and maintain test automation

A model can translate a described interaction into a candidate script, explain unfamiliar test code, suggest a refactor, or help identify tests that may be redundant. Generated code still needs normal review and execution checks. A script can run successfully while asserting the wrong expected result, using an unstable locator, or testing a behavior the requirement never promised.

Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

Summarize test runs and defects

AI can draft a summary from execution logs, screenshots, and defect notes, or help organize a failure report. Verify the summary against the underlying artifacts, including the build, environment, test data, and relevant logs. A concise summary that omits a failing configuration or misstates a failure is not reliable release evidence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Look for improvement opportunities

Teams may use AI assistance to identify recurring failure patterns or propose changes to test suites and processes. Treat suggested improvements as hypotheses. Evaluate them against a baseline and measures that matter to the team; the fact that a model produced more tests does not establish that coverage improved or risk fell.

How to use AI assistance without losing control of quality

Keep accountability with the team and preserve the connection between what was required, what risk was addressed, what was tested, and what happened. A practical review flow is:

Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
  1. Provide bounded context. Give the tool the relevant requirement, constraints, and requested output. Do not include secrets or personal data unless the tool and organizational policy explicitly permit it.
  2. Ask for proposals, not verdicts. Request candidate scenarios, questions, or code, and make clear what behavior is in scope. Do not ask an AI-generated result to certify that a release is safe.
  3. Review against an independent oracle. Check suggestions against approved requirements, business rules, known interfaces, and expected results. Where behavior is uncertain, resolve it with the responsible stakeholder instead of letting the model decide.
  4. Run and inspect generated artifacts. Review code as code, execute tests in the intended environment, and inspect logs or screenshots rather than relying on a generated narrative.
  5. Keep traceability and provenance. Record which requirement or risk a test addresses, who reviewed material that was adopted, and the relevant run or build details. This makes it possible to investigate failures and revise tests when the product changes.
  6. Measure the workflow. Compare the assisted process with the existing one using agreed outcomes, including correction time and maintenance effort—not just the quantity of generated output.

ISTQB’s updated CT-GenAI syllabus identifies prompt engineering, evaluation of generated outputs, and applying generative AI through the testing lifecycle as practical areas of focus. The important operational point is evaluation: an answer that reads plausibly can still be wrong, incomplete, or disconnected from the requirement.

How to test a product that contains AI

For an AI-enabled product, conventional functional checks may not cover the relevant risks by themselves. ISO/IEC TS 42119-2:2025 describes applying the ISO/IEC/IEEE 29119 testing series to AI systems and components using a risk-based approach. The test plan should reflect the system’s intended use, the consequences of failure, and the ways its behavior or inputs may vary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with risks and requirements

Identify what could go wrong, how likely it is, and the consequence if it happens. Prioritize risk exposure, then choose test activities and evidence appropriate to the risk. Requirements matter alongside risk: a risk-based strategy is not a reason to omit required behavior or acceptance criteria.

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

Choose test levels and techniques to match the risk

Depending on the system, a plan may combine model-level testing, functional tests, data-representativeness checks, static reviews, and other suitable test design techniques and coverage measures. For example, if model performance is a material risk, include tests that assess that performance. If input data may fail to represent the intended use context, evaluate data representativeness. These are risk-driven choices, not a universal checklist that every AI feature must apply in the same way.

Account for behavior that can change

Some AI systems can change behavior as models, data, configuration, or their operating context changes. Consider whether continuous testing or other ongoing checks are warranted, particularly when production conditions may affect behavior. Define what changes trigger reassessment and what evidence is needed before a changed system is relied on.

Assess the system in its use context

Testing a component in isolation may not reveal how users, surrounding software, input data, and operational conditions affect the complete product. Define the intended use and relevant context, then select tests that provide evidence about the risks that arise there. Do not assume a model-level result alone settles the quality of the product that uses it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to decide whether an AI-assisted workflow is helping

Establish a baseline before expanding an AI-assisted process. Choose measures that represent useful quality outcomes for the work, and account for the time and effort required to review and maintain AI-generated material.

  • Usefulness of reviewed tests: whether the cases retained after review address meaningful requirements or risks.
  • Coverage and traceability: whether requirements and prioritized risks have appropriate tests, rather than simply whether the suite is larger.
  • Defects found and escaped: interpret these alongside changes in scope, test effort, and environment; a raw count alone can be misleading.
  • Correction and review effort: how much time is spent finding and fixing incorrect, redundant, or unusable generated material.
  • Maintenance burden: whether generated scripts and cases remain understandable and workable as the product changes.

These are practical evaluation measures, not published performance guarantees. A 2025 secondary study mapping industry-context research on AI adoption in software testing reported that many use cases were proposed, while actual implementations and observed benefits in the literature it reviewed were limited. That finding does not show that organizations do not use AI; it is a reason to avoid treating proposed capability as proof of universal adoption or measured productivity.

Standards and guidance: what each one says

Reference What it covers Status and use
ISO/IEC TS 42119-2:2025, Artificial intelligence — Testing of AI — Part 2: Overview of testing AI systems Applies the ISO/IEC/IEEE 29119 testing series to AI systems and components, with risk-based test planning. Published. Its overview discusses using risk to guide test levels, test types, design techniques, static reviews, and coverage measures.
ISO/IEC TS 25058:2024, Guidance for quality evaluation of artificial intelligence systems Guidance for evaluating AI systems using an AI system quality model. Published; its stated scope includes organizations developing or using AI systems.
ISO/IEC 25059:2023 AI system quality model. Previously published edition. The second-edition ISO/IEC FDIS 25059 was identified as a draft in the approval phase on the ISO page consulted for this article, not as a published replacement. Check ISO’s current status before citing a newer edition in policy or procurement documents.
NIST AI Risk Management Framework (AI RMF) Voluntary AI risk-management guidance, with related playbook, profiles, use cases, and testing, evaluation, verification, and validation (TEVV) resources. NIST’s AI Resource Center describes the framework as voluntary and says version 1.0 is under revision. Use it as guidance, not as a mandatory standard unless a separate obligation makes it applicable.

These references can help teams structure evaluation, but naming a standard or framework does not itself demonstrate that a particular product is safe, effective, or fit for purpose. Select the requirements and evidence that apply to the system and its risks.

Capture a web-test screenshot as supporting evidence

A screenshot can help a reviewer understand a rendered page associated with a test result; it is one artifact, not a replacement for assertions, logs, or review of the test environment. For a capture endpoint rather than a browser setup, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Here is a one-call cURL example for a webpage:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo API documentation for request options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
  • ScreenshotNeo can accept cookie or consent banners before capture and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether the request was billed.
  • Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.
  • The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. All listed features are available on every plan.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Common mistakes to avoid

  • Counting generated tests as coverage. A large set can still miss a critical requirement, repeat the same scenario, or assert incorrect behavior. Map retained tests to requirements and risks.
  • Accepting generated code because it runs. Successful execution does not establish that the script exercises the right behavior or checks the right result. Review the test oracle and inspect the run.
  • Using an AI summary as the original evidence. Keep the underlying logs, screenshots, and environment information available so that a reviewer can verify the summary.
  • Testing only the model. A component result may not address risks from data, product integration, or the system’s intended use. Choose system and data checks where the risk calls for them.
  • Making productivity claims without a baseline. Compare outcomes and review costs with the existing workflow before attributing an improvement to AI.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.