Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Harvard Census II of Free and Open Source Software report is a 2022 study of application libraries observed in production applications represented in data from software composition analysis (SCA) partners. It offers a set of evidence-based views of package use—not a complete inventory of software, a current popularity chart, or a security-risk ranking.

What is the Harvard Census II report?

Census II of Free and Open Source Software — Application Libraries is a Linux Foundation and Laboratory for Innovation Science at Harvard report published in March 2022. Its authors are Frank Nagle, James Dana, Jennifer Hoffman, Steven Randazzo, and Yanuo Zhou. It builds on Census I, which examined lower-level operating-system libraries and utilities; Census II focuses on application-level packages.

The study aggregated private usage data supplied by SCA partners Snyk, Synopsys Cybersecurity Research Center (CyRC), and FOSSA. The Linux Foundation describes the dataset as containing over half a million observations of FOSS libraries in production applications at thousands of companies. The release announcement says the report identified more than one thousand widely deployed application libraries and published eight rankings of 500 packages, each using a different view of the contributed data.

What do the rankings show?

The report does not offer one universal ranking. Its appendices separate ecosystems and distinguish views such as direct versus indirect dependencies and version-agnostic versus versioned packages. A package’s position only makes sense in the context of the particular list, its ecosystem, and how inclusion was counted.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, the Linux Foundation’s release announcement lists lodash, react, axios, debug, @babel/core, express, semver, uuid, react-dom, and jquery among the top ten version-agnostic npm packages called directly in applications represented in the study. That is an example from the 2022 data—not a current 2026 popularity ranking.

How was the data collected, and what does it represent?

Census II combined data from scans conducted by the SCA partners. The results therefore describe packages present in the applications and software layers those customers and tools scanned. They do not represent every application, organization, or software stack. A scan may omit layers: for example, an application scan running on Linux may not include the full operating system beneath it.

The authors characterize the lists as their best estimate of which FOSS packages were most widely used by the applications represented, given the broad but non-exhaustive data and the study’s time constraints. They do not claim a definitive census of all FOSS use.

Does a high rank mean a package is critical or risky?

No. Census II measures observed prevalence within its scope, not software criticality or security risk. The authors explicitly say the report does not determine which packages are used by the most widely used applications, identify packages most critical to infrastructure, or measure software risk profiles. A high rank is not a security score and does not establish that a package is safe, vulnerable, or systemically important.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What did Census II find about open-source software?

The report’s executive summary identifies five challenges that shape how organizations should interpret and manage FOSS dependencies:

  • Component names are not standardized. Different naming practices make it harder to identify and compare the same software consistently.
  • Versions complicate counts. A package counted without regard to version can appear differently when specific releases are separated.
  • Some widely used projects have few contributors. Popularity does not necessarily mean a project has a broad contributor base.
  • Developer account security matters. Individual maintainers’ accounts are part of the security picture, alongside the code itself.
  • Legacy software persists. Older packages can remain in dependency trees, so an inventory should be read with version and maintenance context.

How should organizations use the report?

Use Census II as historical evidence about the difficulty of measuring dependency prevalence and as context for building a more reliable software inventory. It can help frame questions for an organization’s own dependency analysis, but it cannot substitute for examining that organization’s actual applications and dependency trees.

  • Record the exact component identity and ecosystem; names alone may not be consistent across tools.
  • Keep version-specific findings separate from version-agnostic counts.
  • Distinguish direct dependencies—declared or included by the application—from indirect dependencies brought in through other packages.
  • Compare rankings only when they use the same ecosystem, dependency relationship, version treatment, and list construction.
  • Assess maintenance, contributor concentration, and account protections separately from usage rank.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where can you read the report?

The Linux Foundation Research page for Census II provides the report information and download. The March 2, 2022 release announcement summarizes its scale and ranking structure.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.