Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →AI vulnerability discovery is best understood as a security investigation workflow, not a model taking a guess at a code snippet. Systems can build repository context, inspect code or commits for suspicious behavior, use tools to investigate a lead, try to reproduce it in an isolated environment, and propose a patch. That can help researchers and maintainers find and validate flaws faster, but it does not prove a codebase is secure. People still need to assess impact, review fixes, and coordinate disclosure.
How does AI find vulnerabilities in code?
A useful way to understand the process is as a loop: build context, search for a lead, investigate it, test whether it is real, and then review a proposed fix. The details vary by system. OpenAI describes Aardvark as analyzing a repository and its security objectives, scanning commits against that context, and testing candidate findings in a sandbox. Google Project Zero’s Naptime work emphasizes interactive execution and tools such as debuggers and scripting environments.
- Build repository context. The system analyzes code structure, intended behavior, and, in some workflows, security objectives or a threat model. A repository-wide view can help make a suspicious code path more meaningful than an isolated snippet.
- Search for candidates. A system may review current source, newly submitted commits, or repository history. Aardvark says it scans commits and, when first connected, repository history.
- Investigate with tools. Rather than relying only on a text explanation, an agent may write a test, run code, use scripts, or inspect runtime behavior. Naptime’s approach specifically emphasizes interactive environments and specialized tools.
- Try to reproduce the flaw. A crash or other observable failure can strengthen a finding. Aardvark says it tests potential findings in a sandbox; Naptime describes benchmark tasks whose results can be checked through observable crashes.
- Propose and review a patch. A candidate fix must remove the vulnerability without breaking intended behavior. AIxCC scoring valued both patches and preserved functionality; Aardvark presents patches for human review.
- Coordinate disclosure. A credible technical finding still needs severity assessment and communication with the affected maintainers. OpenAI’s disclosure policy says its default is to contact affected parties privately first, with timelines open-ended by default.
Reproducibility matters because a plausible explanation is not the same as a demonstrated bug. A repeatable test or proof of vulnerability gives a reviewer something concrete to inspect, though it does not by itself settle severity or prove that a proposed patch is safe.
What do AI vulnerability discovery results actually show?
Published results indicate genuine progress, but they come from different kinds of evidence: a competition, research frameworks, a reported real-world discovery, and a benchmark. Their numbers should not be combined into a single accuracy ranking.
#1 Best Overall
| Evidence | What was reported | What it does—and does not—establish |
|---|---|---|
| DARPA AI Cyber Challenge (AIxCC), final results announced August 8, 2025 | In the scored final round, systems identified 86% of the competition’s synthetic vulnerabilities and patched 68% of the vulnerabilities identified. DARPA also reported 54 unique synthetic vulnerabilities discovered and 43 patched, plus 18 real, non-synthetic vulnerabilities discovered and 11 real-issue patches provided. | These are results on defined challenge projects and rules, not a rate for arbitrary production code. DARPA said the real findings were being responsibly disclosed to open-source maintainers. |
| Google Project Zero’s Naptime framework | Project Zero reported up to 20-times-better performance on Meta’s CyberSecEval 2 vulnerability tests compared with the original paper’s results, with scores of 1.00 on Buffer Overflow tests and 0.76 on Advanced Memory Corruption tests. | These are scores under the authors’ benchmark methodology, not a probability that the system will find a bug in any codebase. Project Zero said substantial progress remained before such systems could meaningfully affect security researchers’ daily work. |
| Google DeepMind and Project Zero’s Big Sleep, reported in October 2024 | Big Sleep found an exploitable stack buffer underflow in SQLite. The team reported it to developers in early October 2024, and maintainers fixed it the same day, before the issue appeared in an official release. | This is a concrete reported discovery, not evidence of universal reliability. Project Zero described the research as early-stage and said variant analysis—looking for related bugs based on a known flaw—was a better fit for current LLMs than open-ended vulnerability research. |
| OpenAI’s EVMbench, announced February 18, 2026 | The benchmark uses 117 curated vulnerabilities from 40 audits and evaluates smart-contract agents in detection, patching, and sandboxed exploitation modes. | OpenAI reports that detection and patching remain below full coverage. In detect mode, the benchmark cannot yet reliably determine whether additional agent-reported issues are genuine vulnerabilities or false positives. |
AIxCC also reported that its systems analyzed more than 54 million lines of code, with an average cost of about $152 per competition task and an average of 45 minutes to submit patches. Those figures describe the competition setting; they are not general estimates for commercial security work or production repair times.
Can AI detect zero-day vulnerabilities?
AI systems can help discover previously unknown flaws, as Big Sleep’s reported SQLite finding illustrates. But one successful discovery does not establish that an AI can reliably find zero-days across arbitrary software, or that it will find them before attackers do. Evidence is strongest when a system is given a bounded task or a useful lead, such as a code change to inspect or a previously fixed flaw to use for variant analysis.
Open-ended research is harder because the system must decide where to look, distinguish unusual code from exploitable behavior, and establish that a suspected issue is real and consequential. Project Zero’s assessment of Big Sleep emphasizes this distinction: variant analysis narrows the search by starting from a known issue, whereas general vulnerability research begins with much more ambiguity.
Can AI fix security vulnerabilities?
AI can propose patches, and competition results show that systems can produce fixes for some challenge findings. But a patch is not successful merely because it removes a crash or silences a warning. It must address the underlying flaw while preserving intended functionality, and a reviewer needs to check the change in context.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteRank #3
OpenAI’s EVMbench highlights how difficult this can be for smart contracts: agents are evaluated on patching while retaining intended functionality, and OpenAI reports that performance remains below full coverage. DARPA’s CHESS program likewise calls for human-generated insight, proof of vulnerability, and a specific, non-disruptive patch. In practice, patch review belongs in the workflow rather than being treated as an automatic final step.
Where do AI vulnerability tools fall short?
- A scan is not an exhaustive audit. OpenAI reports that EVMbench agents sometimes stop after finding one issue in detect mode. A tool that returns a real finding has not necessarily searched every relevant path or found every flaw.
- False positives and missing ground truth complicate measurement. If a system reports an issue absent from a benchmark’s human-audited list, it can be difficult to determine whether the report is a genuine additional vulnerability or a false positive. EVMbench says its detect mode cannot yet reliably resolve that question.
- Benchmarks simplify real environments. EVMbench draws on Code4rena audits and uses a local chain environment with sequential transaction replay. Its authors note limits involving timing-dependent behavior, mainnet state, and multi-chain settings. Performance there does not settle performance under those omitted conditions.
- Correct fixes require semantic judgment. Removing a flaw while retaining intended functionality can be especially difficult when a change affects subtle program behavior. A patch that breaks legitimate use is not a complete security solution.
- Open-ended discovery is a harder problem than following a lead. A known flaw, suspicious commit, or reproducible failure gives an agent a narrower task. Finding an unknown issue without such a starting point requires more independent judgment.
- Disclosure is a human coordination problem as well as a technical one. Maintainers need actionable details, an appropriate severity assessment, and time to respond. OpenAI’s policy describes private-first contact and leaves disclosure timelines open-ended by default.
How should you evaluate an AI security tool?
Compare systems by the work they perform and the evidence they return, not by an isolated benchmark score. Useful questions include:
Rank #4
- Scope: Does it examine a whole repository, new commits, code related to a known flaw, or only a bounded benchmark task?
- Evidence: Does it provide a plausible explanation, a reproducible test, or a proof of vulnerability that a reviewer can run?
- Validation: Does it test findings in an isolated sandbox or a relevant local harness? What important production conditions are absent?
- Patch quality: Does the proposed change remove the underlying issue while preserving intended behavior, and can a maintainer review why?
- Workflow and access: Who can use the system, how are findings handled, and does it fit the team’s review and disclosure practices? For example, OpenAI describes Aardvark as a private-beta agent, not as a generally available service.
Human judgment remains central in the model DARPA’s CHESS program describes: people contribute insight, establish proof, and shape a non-disruptive patch. AI can extend a security team’s investigative capacity, but the available results support treating it as an aid to research and maintenance—not as a substitute for a security review or a guarantee that software has no vulnerabilities.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

