What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Google is using large language models to write and repair fuzz-test harnesses, then running those harnesses through its established OSS-Fuzz infrastructure. In experiments reported from 2023 to 2024, the approach expanded code coverage and helped surface vulnerabilities—but the results are project-specific, and Google has not described the system as an autonomous way to secure arbitrary software.
What AI adds to fuzz testing
Fuzz testing feeds software large volumes of varied or malformed inputs to expose crashes and other defects. A fuzz target, also called a harness, is the small piece of code that routes those inputs into a chosen function or component.
Google’s approach uses a large language model (LLM) to draft or improve that harness. The LLM does not replace the fuzzing engine: conventional fuzzing still generates and mutates inputs and explores program execution. The harness gives the engine a route into code that existing targets may not reach.
Google’s Open Source Security Team noted in 2023 that OSS-Fuzz covered around 30% of open-source project code on average at that time, leaving substantial code untouched by fuzzing. That figure was the team’s reported baseline in its 2023 announcement, not a current independent measurement.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →How Google’s AI-assisted workflow works
- Find code that may be under-tested. Fuzz Introspector identifies code with low runtime coverage that could benefit from additional fuzzing.
- Give the model project context. The evaluation framework selects a function and can provide project code, examples of existing targets, FuzzedDataProvider usage, and examples of anti-patterns.
- Generate and build a target. The LLM drafts a harness, and the framework attempts to compile it. If compilation fails, an iterative prompt can ask the model to fix the code.
- Run the target and measure outcomes. The framework checks for successful execution, crashes, and additional coverage. Google’s technical description details the workflow and evaluation measures.
Compilation is an early hurdle, not proof that a target is useful. A target can compile yet call APIs incorrectly, crash immediately, or add little coverage. Likewise, reaching more lines is a useful signal of exploration, but it does not prove that every bug in those lines has been found or that a vulnerability exists.
What Google reported, and when
The headline figures come from different experiments, time periods, and scopes. They should not be combined into a single benchmark or treated as the expected result for a new project.
| Report and scope | Google-reported result | How to read it |
|---|---|---|
| Google Open Source Security Team, 2023: TinyXML2 | Line coverage rose from 38% to 69% without intervention from Google’s team. | A result for one named project, not a guaranteed uplift for other codebases. |
| Google Open Source Security Team, 2023: sample projects | Reported code-coverage gains ranged from 1.5% to 31%. | Experimental results across selected projects, not a universal range. |
| OSS-Fuzz technical report: preliminary experiment | New targets compiled and increased coverage in 14 of 31 tested projects. | An early evaluation; the technical report describes prompt engineering and a compiler wrapper used to improve compilation outcomes. |
| OSS-Fuzz-Gen repository: sample experiment dated January 31, 2024 | More than 1,300 benchmarks from 297 open-source projects; successful targets increased coverage for 160 C/C++ projects, with a maximum 29% line-coverage increase over existing human-written targets. | A distinct sample experiment described in the OSS-Fuzz-Gen repository, not the later 2024 project-wide report. |
| Google Open Source Security Team, November 2024: 272 C/C++ projects | More than 370,000 new covered lines; the largest single-project increase cited was from 77 to 5,434 covered lines. | A later report with a broader project scope; Google does not present it as directly comparable to every earlier experiment. |
The initial 2023 results and the later 2024 figures are described in Google’s 2023 announcement and November 2024 follow-up, respectively.
Coverage gains and vulnerability discoveries are different outcomes
Google’s November 2024 account said AI-generated or enhanced targets had found 26 new vulnerabilities in OSS-Fuzz projects. One named example was OpenSSL CVE-2024-9143: Google said it reported the issue on September 16, 2024, and that a fix was published on October 16, 2024. These are findings reported by Google, not evidence that every increase in coverage produces a vulnerability.
An earlier OpenSSL example illustrates why the distinction matters. In 2023, a generated target reached code that rediscovered CVE-2022-3602, a previously known vulnerability. That showed the target could expose a missed path; it was not a new vulnerability discovery. Google’s 2023 and 2024 accounts discuss the two cases separately.
What the results do—and do not—show
The early work focused on existing OSS-Fuzz projects, initially C and C++. Google’s technical page says some blockers came from deficiencies in existing targets rather than the fuzzing engines, and frames automatically onboarding entirely new projects as a more difficult problem.
Rank #4
For teams evaluating AI-generated targets, coverage alone is an incomplete quality score. A meaningful comparison with human-written targets should consider:
- Build success: Does the target compile in the project’s environment?
- Runtime stability: Does it run reliably, or produce immediate and misleading crashes?
- Incremental coverage: Does it reach code that existing targets do not?
- Validated findings: Do crashes lead to reproducible bugs or maintainer-confirmed vulnerabilities?
- Engineering effort: How much project context and developer time are needed to get a useful target?
- Triage and review: Can developers identify which results matter and verify them?
Google’s later report describes human review and automated triage as ongoing work, rather than presenting the whole process as hands-off. It also reports that interactive tools such as debuggers can help an agent reach correct results; that observation is not a guarantee for other projects or workflows.
Best Value
Can an LLM fuzz a project on its own?
Google’s work supports a narrower claim: an LLM can help create or repair fuzz targets, while OSS-Fuzz and its evaluation framework build, run, and assess those targets. Google described more automated triage, tool-using agent workflows, and closer OSS-Fuzz integration as future work in its November 2024 account. That roadmap is not a claim that the model can independently secure arbitrary software.
For details on the service and its availability to open-source projects, see the OSS-Fuzz documentation. The OSS-Fuzz-Gen repository provides information about Google’s open framework and its sample experiment.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

