Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Meta’s April 29, 2025 release introduced a collection of AI safeguards, cybersecurity evaluation tools, and partner-facing services—not a single safety product. The publicly accessible tools include Llama Guard 4, LlamaFirewall, Llama Prompt Guard 2, and CyberSecEval 4; other offerings are limited to selected participants in Meta’s Llama Defenders Program.

What did Meta release?

Meta said developers could access its latest Llama Protection tools through its Llama Protections page, Hugging Face, or GitHub. The tools serve different purposes: some screen content or prompts, one coordinates safeguards across an AI system, and another set evaluates cybersecurity capabilities. Meta also announced offerings for selected organizations through its Llama Defenders Program.

Llama Guard 4: text and image safeguards

Llama Guard 4 is an update to Meta’s customizable Llama Guard tool. Meta describes it as a unified safeguard for understanding text and images. The company also said the model was available through a limited-preview Llama API.

Llama Prompt Guard 2: jailbreak and prompt-injection detection

This updated classifier is intended to detect jailbreak attempts and prompt injection. Meta introduced versions labeled 86M and 22M; these are model-size identifiers, not safety scores. Meta says the smaller version can reduce latency and compute costs with minimal performance trade-offs, but that is the company’s characterization, not an independent comparative test.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

LlamaFirewall: safeguards across an AI system

LlamaFirewall is designed as a guardrail tool for building secure AI systems. Meta says it can orchestrate across guard models and work with its protection-tool suite to detect or prevent risks such as prompt injection, insecure code, and risky interactions with LLM plug-ins. Unlike a classifier focused on a particular input, it is described as coordinating protections across system workflows.

CyberSecEval 4: cybersecurity evaluation

CyberSecEval 4 is an updated open-source benchmark suite for assessing cybersecurity capabilities in AI systems. Meta announced two additions:

  • CyberSOC Eval, developed with CrowdStrike, measures AI systems’ efficacy in security operations centers.
  • AutoPatchBench evaluates whether AI systems can automatically patch vulnerabilities in native code before exploitation.

A benchmark provides an assessment against its chosen tasks; it does not guarantee that a system will defend effectively in real-world security operations.

How do the tools differ?

Tool Main role Scope or threats described by Meta Access described in the announcement
Llama Guard 4 Content safeguard Text and image understanding Meta’s Llama Protections page, Hugging Face, GitHub; limited-preview Llama API
Llama Prompt Guard 2 Prompt classifier Jailbreaks and prompt injection Meta’s Llama Protections page, Hugging Face, GitHub
LlamaFirewall System-level guardrail and orchestration Prompt injection, insecure code, and risky LLM plug-in interactions Meta’s Llama Protections page, Hugging Face, GitHub
CyberSecEval 4 Cybersecurity benchmark suite Security-operations efficacy and automated vulnerability patching, among its evaluations Open-source suite; Meta’s announcement points developers to its Llama Protection access routes

The announcement does not provide a head-to-head independent performance comparison. It also does not fully specify whether every tool works with arbitrary models or deployment setups, so check each tool’s documentation and integration requirements before adopting it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What is the Llama Defenders Program?

Meta described the Llama Defenders Program as a program for selected partners and developers, with access to a mix of open, early-access, and closed AI solutions for security needs. Its announcement included an automated sensitive-document classification tool, intended to label internal documents or filter sensitive material from retrieval-augmented generation (RAG) systems.

Meta also described generated-audio and audio-watermark detectors intended to help organizations identify threats such as scams, fraud, and phishing. ZenDesk, Bell Canada, and AT&T were named as integration partners for the audio tools at launch. The announcement does not establish a public signup path for the program.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How does this fit Meta’s broader AI safety approach?

In its February 3, 2025 Frontier AI Framework announcement, Meta said the framework focuses on cybersecurity threats and risks involving chemical and biological weapons. Meta described a process of identifying catastrophic outcomes, threat modeling, setting risk thresholds, and applying mitigations. It also argued that open access can help the company learn from independent community assessments of model capabilities and improve risk evaluation.

That is Meta’s stated rationale and process, not independent evidence that releasing tools or models makes every system safe. Meta’s framework announcement put its openness argument this way: “Our open source approach also helps us to better anticipate and mitigate risk because it enables us to learn from the broader community’s independent assessments of our models’ capabilities.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can you use these tools with your own model?

Meta’s announcement supports describing the intended functions and access routes, but it does not establish compatibility with every third-party model, application, or deployment. Treat the release as a set of components to evaluate for a particular system, rather than a universal safety layer. Confirm the relevant tool’s documentation, supported inputs, integration requirements, and access status before relying on it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.