The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A2A-G is an early student project for publishing AI-agent safety test results, not an established certification standard. Its author describes an optional, consent-based mock-sandbox test: an agent faces a set of hijacking prompts, an LLM judge evaluates its responses, and a result is signed and added to a hash-chained history. Those mechanisms may make records easier to check for alteration, but they do not prove that an agent is safe, that testing is comprehensive, or that the signing key belongs to the operator.
What A2A-G is designed to do
The project’s author describes A2A-G as a public registry for agent access claims and test results. It starts with a “Grey Badge,” an unverified self-report of what an agent says it can access. Owners can optionally choose to have the agent tested. The author says testing takes place in a safe, consented mock sandbox rather than against a live production system, and that both passing and failing results are published. These are the author’s descriptions of the project, not independently audited implementation details. The project author’s post
How the described test works
According to the author, the test sends 18 OWASP hijacking prompts to the agent three times each, in randomized order. An LLM judges whether the agent blocked or complied with each prompt. A result that blocks at least 80% earns a “Blue Badge.” The author characterizes the method as imperfect; the figures are project-specific design claims, not independently validated measures of safety.
- What the badge indicates: a result against this described prompt set and evaluation method.
- What it does not establish: broad coverage of agent-safety risks, reliable performance in every context, or a demonstrated link between the 80% threshold and real-world risk.
- What the judge adds: an LLM’s interpretation of responses, which is itself part of the test method and may affect the result.
What the cryptographic record can prove
The author says each result is signed with Ed25519 and linked in a hash chain, with the goal of making the record’s history harder to alter or conceal. A correctly verified signature can show that a record matches a particular signing key; a hash chain can help reveal changes to linked records. Neither mechanism establishes that the test was good, the judge was reliable, or the key belongs to the operator named in a record.
#1 Best Overall
Key provenance matters. The August 2026 Internet-Draft The Agent Record: Transparent, Witness-Countersigned Event Logs for AI Agent Identity, History, and Memory explains that checking a signature against a public key supplied inside the same artifact establishes internal consistency, not external identity. Its proposed approach uses externally anchored keys, signed checkpoints over append-only Merkle logs, and independent witnesses. The draft includes an “unanchored” verification outcome, which makes no authenticity claim, and says its independent-implementation gate had not yet been met. It is a proposal, not a finalized standard. Its terminology section puts the distinction plainly: “A registry is NOT a trusted party.”
Why a Blue Badge can become stale
The author says badges do not currently revoke automatically when a dependency update changes an agent’s behavior. A Blue Badge may remain in place until its stated 90-day expiry. The author mentioned possible webhook triggers, but said continuous monitoring would require recurring LLM-judge compute costs the solo student could not currently cover. The expiry is therefore not evidence that the agent is monitored continuously or retested whenever its software changes. The author’s post and discussion
For a user assessing an agent, the practical question is not just whether a badge exists, but when and under what conditions its result was produced, and whether the agent has changed since then. The described registry does not currently provide automatic revocation on dependency updates.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsHow this differs from evidence-conformance tooling
A behavioral safety test and a tool that checks whether execution evidence conforms to a format address different problems. OECD.AI’s catalog entry describes the Agent Evidence Conformance Suite as software for machine-checkable tests of signed AI-agent execution evidence, not as a behavioral certification service or a policy-setting system. The entry reports over 270 conformance vectors and a four-stage verification pipeline; those figures apply to that toolkit, not A2A-G. OECD.AI: Agent Evidence Conformance Suite
Rank #3
When comparing registry or verification projects, look separately at the behavior tested, reproducibility, evaluator independence, sandbox and consent requirements, how signing keys are anchored, what result details are public, and how results are updated or revoked. A conformance check can help verify evidence structure without certifying agent behavior; a prompt-based behavioral test does not, by itself, prove the authenticity of the operator’s identity.
Quick Recap
Rank #4
How to interpret an A2A-G badge
- Read it as a record of an optional test under the project’s described conditions—not as a general safety certification.
- Check the result’s date and whether relevant software or dependencies have changed since testing.
- Distinguish a valid signature from a verified identity: the signing key needs a trustworthy external anchor for that stronger claim.
- Consider the test’s limited prompt set and LLM judge when deciding how much weight to give a pass or failure.
- Remember that publishing failures can improve transparency, but publication alone does not make a test comprehensive or a registry authoritative.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

