iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
To red-team an AI agent’s pull requests in GitHub Actions, add a separate adversarial-scan job that runs when agent behavior changes, measure its findings against the main branch before it can block merges, and treat a failure as a reason for human review rather than a verdict. Ayan Pahwa’s case study from Humanbound, published September 23, 2026, shows the workflow on a support agent. Its reported cost was about $0.26 per run for the attacker and judge model calls. Those figures are the author’s own example, not independently reproduced results, and the numbers below should be read that way.
What the case study describes
The scenario is a single prompt change. The support agent’s prompt was edited to tell it to issue a refund based on an order ID and an amount, without verifying that the order exists or belongs to the customer. Ordinary checks passed. An adversarial scan then found a high-severity, refund-related failure and turned the pull-request gate red. The author also reports that the main branch was already red when the change was tested, which matters for how the gate should be configured (more on that below). The source is Humanbound’s case study by Ayan Pahwa.
The example pipeline runs when any of the following change:
- agent code
- prompts
- tool definitions
- scope or configuration files that define what the agent may do
- the workflow file itself
The author’s list of paths is an example, not a universal set. Your list should follow wherever your agent’s behavior is actually defined, and a prompt stored in a database or a remote configuration service would need a different trigger.
#1 Best Overall
- POWERFUL SECURITY KEY: The Security Key C NFC is the essential physical passkey for protecting your digital life from phishing attacks. It ensures only you can access your accounts.
- WORKS WITH 1000+ ACCOUNTS: Compatible with Google, Microsoft, and Apple. A single Security Key C NFC secures 100 of your favorite accounts, including email, password managers, and more.
- FAST & CONVENIENT LOGIN: Plug in your Security Key C NFC via USB-C and tap it, or tap it against your phone (NFC) to authenticate. No batteries, no internet connection, and no extra fees required.
- TRUSTED PASSKEY TECHNOLOGY: Uses the latest passkey standards (FIDO2/WebAuthn & FIDO U2F) but does not support One-Time Passwords. For complex needs, check out the YubiKey 5 Series.
- BUILT TO LAST: Made from tough, waterproof, and crush-resistant materials. Manufactured in Sweden and programmed in the USA with the highest security standards.
Why ordinary tests can miss this
A prompt change rarely alters a function signature, a type, or a unit-level return value, so a conventional test suite can pass while the agent becomes willing to do something it should not. The refund shortcut in the case study is that kind of defect: the code compiles, the schema validates, and the behavior is wrong only when a user pushes on it.
OpenAI’s red-teaming guidance frames adversarial testing as a complement to ordinary evaluations, not a replacement. In its words: “Red teaming uses adversarial test cases to help uncover unsafe, insecure, or policy-violating behavior before deployment.” (OpenAI API documentation: Red teaming). Adversarial scans add coverage for misuse and unexpected interactions. They do not prove that an agent is secure, and a green check is not that proof either.
Build the pipeline
Choose the scan mode for each trigger
The case study uses two modes of scan, and each suits a different point in the release cycle.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →| Mode | Where it runs in the case study | What it does | Trade-off |
|---|---|---|---|
| Single-turn scan | Pull requests; fails the job on a high-severity finding | Sends one-shot adversarial prompts and judges each response | Faster and broader across attack types, but it cannot test how an agent handles a long, shifting conversation |
| Multi-turn agentic scan | Optional scheduled run | Runs longer adversarial conversations that adapt over several turns | Deeper, but slower and more expensive per run; better suited to scheduled or review use than to blocking every PR |
| Manual workflow dispatch | Before a release | Lets a maintainer run a scan on demand | Requires someone to remember to run it, so it complements rather than replaces automated triggers |
The case study does not publish cost or duration for the multi-turn mode separately, so compare the modes on depth and suitability first and measure the cost in your own environment.
Bound runs and usage
The case study’s job uses several controls that you should copy the intent of, while checking current GitHub Actions syntax for each:
Rank #2
- POWERFUL SECURITY KEY: The Security Key NFC is the essential physical passkey for protecting your digital life from phishing attacks. It ensures only you can access your accounts.
- WORKS WITH 1000+ ACCOUNTS: Compatible with Google, Microsoft, and Apple. A single Security Key NFC secures 100 of your favorite accounts, including email, password managers, and more.
- FAST & CONVENIENT LOGIN: Plug in your Security Key NFC via USB-A and tap it, or tap it against your phone (NFC) to authenticate. No batteries, no internet connection, and no extra fees required.
- TRUSTED PASSKEY TECHNOLOGY: Uses the latest passkey standards (FIDO2/WebAuthn & FIDO U2F) but does not support One-Time Passwords. For complex needs, check out the YubiKey 5 Series.
- BUILT TO LAST: Made from tough, waterproof, and crush-resistant materials. Manufactured in Sweden and programmed in the USA with the highest security standards.
- Path filters so the job runs only on relevant changes.
- Concurrency with cancellation so that a newer push cancels superseded runs of the same scan.
- Job timeouts so a stuck scan cannot consume runner time indefinitely.
- Token metering so you can see how many model tokens each run consumed.
- SARIF upload so findings can appear in the repository’s code scanning view in a structured form.
- Stored artifacts so transcripts and reports can be reviewed after the run. Read the permissions section below before enabling this.
Set a baseline before blocking merges
The most useful lesson in the case study is that the main branch was red as well. A gate that counts every existing finding will fail pull requests that did not introduce them, and teams often respond by disabling the gate. Introduce the scan in stages instead:
- Run the scan in report-only mode on the main branch and record every finding with its attack category and severity.
- Review the findings with the agent’s owners and decide which ones are accepted risks, which are known bugs, and which are blockers.
- Run the same scan on your pull request branch and compare the result with the baseline finding by finding, not only by total count.
- Set the failure threshold (for example, fail only on new high-severity findings) once the comparison works reliably.
Counts alone are not enough for step three. In the case study, the pull-request run reported 48 failures against 39 on the baseline, a difference of nine. That gap does not tell you which findings are new, so the comparison has to match findings by attack and content.
What the reported numbers show
The case study reports the following results for one demonstration agent and one configuration. They are not general performance rates for any model or product.
| Run (as reported by the author) | Attacks or conversations | Pass | Fail | Pass rate (computed from the reported counts) |
|---|---|---|---|---|
| Pull-request branch, single-turn | 432 attacks | 384 | 48 | about 88.9% |
| Baseline (main branch), single-turn | 432 attacks | 393 | 39 | about 91.0% |
| Pull-request branch, multi-turn agentic | 97 conversations | 7 | 90 | about 7.2% |
The multi-turn figure is far lower than the single-turn one. The case study does not report a multi-turn baseline, so you cannot read that gap as a regression caused by the prompt change. It does show why the two modes should be interpreted separately.
What the cost figure covers
The headline figure is the author’s approximate cost of about $0.26 for each of three reported runs. It covers the attacker and judge model calls in the example, which used gpt-4o-mini, and it is calculated with the author’s token-metering method. The author states that the agent’s own model calls were not metered. If your agent calls a larger model, its spend is outside the $0.26 figure.
Rank #3
Scan cost scales with the model you choose, the number of attacks or conversations, and how many times the scan runs. Under the conditions the author reports, the arithmetic is simple: 50 scans at about $0.26 each is roughly $13, and 200 scans is roughly $52. Those totals are illustrations of the reported unit cost, not a forecast for your repository, and they exclude the agent’s own model usage, any retries, and any change in the scan’s model or attack set.
Recommended Free Tools
Fork pull requests and secrets
The case study skips fork pull requests. Its workflow needs a secret to call the scanning service, and secrets are not provided to fork PRs in the usual setup, so the author chose not to run the scan there. That is a deliberate security decision, but it leaves a coverage gap: changes from outside contributors are not scanned.
GitHub’s guidance points in the same direction for a different product. Its documentation for Copilot CLI in Actions warns: “Workflows that run on pull request events from forks are at higher risk of prompt injection.” (GitHub Docs: About using Copilot CLI in GitHub Actions). That warning applies to Copilot CLI workflows. Do not assume the same risk profile or behavior applies to the Humanbound Action or any other Action without checking its documentation.
If you do need scans on fork PRs, decide the policy explicitly. Options include leaving them unscanned and requiring a maintainer to run a manual dispatch, or designing a workflow that never exposes secrets to untrusted code. Document the choice so reviewers know what was and was not covered.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Permissions and artifacts
The case study author warns that run artifacts on public repositories can be downloaded by signed-in GitHub users. Agent transcripts often contain attack prompts, system details, and sometimes data the agent saw, so an artifact that is harmless in a private repository can expose information in a public one. Decide before enabling artifacts whether transcripts should be stored at all, how long they are retained, and who can read them.
Rank #4
- Ultra-Compact FIDO2 Security Key - Plug-and-stay or carry on a keychain. This USB-A hardware security key offers portable, always-on protection for desktop and mobile use. (Item Size: 0.75 X 0.74 IN x 0.25 IN)
- USB-A Hardware Key for All Devices - Works with USB-A ports on PC, Mac, Android, and other laptop/notebook device. Enables secure, cross-platform login with FIDO2.0 passkey support.
- FIDO Certified Security Key - Meets FIDO and FIDO2 standards. Works with Google, Microsoft, GitHub, Dropbox, and more. Please check service compatibility before purchase.
- Passwordless Login with Passkey - Supports passkey login via WebAuthn and CTAP2. Enjoy password-free sign-ins where supported. Not all websites or services currently support passkeys.
- Advanced Multi-Factor Authentication - Offers 200 FIDO2 passkey slots and 50 OATH-TOTP slots. Strong, flexible 2FA/MFA support across various apps and authentication platforms.
For the workflow itself, grant only the permissions the job needs, keep secrets scoped to the job that uses them, and review any third-party Action before pinning it. GitHub describes a different system, GitHub Agentic Workflows, with read-only defaults, safe outputs, secret isolation, threat detection, and firewalled execution (GitHub Docs: About GitHub Agentic Workflows). Those features belong to that system. They are not evidence about how the Humanbound Action handles permissions. GitHub’s tutorial on developing agentic workflows notes a public-preview status (GitHub Docs: Develop agentic workflows in GitHub Actions), so check that status before depending on it.
Handle a red build
A failing gate is a prompt to investigate. Work through the following checks before deciding whether the change should merge:
- Open the reported finding and reproduce the attack against the pull-request branch.
- Confirm whether the same finding appears in the baseline. If it does, decide whether the PR worsens it or is unrelated.
- Check whether the finding is a false positive from the judge, and record the reason if it is.
- If the finding is real, fix the prompt, tool scope, or policy, then rerun the scan rather than only editing the threshold.
- If the scan is a blocker you cannot fix quickly, document the accepted risk, its owner, and an expiry date.
The refund case shows what a useful outcome looks like. The scan did not decide whether the change was safe; it made a specific, reproducible behavior visible to the people who could judge it, and the prompt change had to be reviewed against that behavior before merging.
Scope of the evidence
No independent study or official statistical series was found that validates the case study’s scan counts or its 26-cent estimate. Treat them as one author’s example. Your own repository’s cost, detection rate, and false-positive rate will depend on your agent, your models, and your attack set, so measure them in report-only mode before setting any threshold.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallFor the underlying attack category, OpenAI defines prompt injection as an attack in which “a third-party—not the user nor the AI—misleads the model by injecting malicious instructions into the conversation context” (OpenAI: Understanding prompt injections). The refund case is a different kind of failure, a behavior the prompt itself permitted, so scans should cover both direct policy violations and injected instructions.
Humanbound is the service named in the example. The case study does not establish any affiliate or referral arrangement, and this article does not assert one.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

