An LLM stub replaces a model call with a predetermined or request-aware response, so you can test how your agent application handles that response without contacting a model provider. Use scripted stubs to verify orchestration—tool calls, authorization, guardrails, handoffs, retries, state changes, and failures—then test the real model and provider adapter separately where model behavior or integration details matter. A passing stub test is evidence about your application’s response to a known scenario, not proof that a deployed model will resist prompt injection.
What an LLM stub can test
An agent commonly passes model output into application code that chooses whether to validate a tool request, ask for approval, execute a tool, retry, hand off to another agent, or return a final answer. A scripted model lets a test supply known outputs at those decision points and check what the surrounding application does.
For example, a stub can return a request to look up a synthetic invoice, followed by a final answer after the test’s fake lookup tool returns marker data. The test can verify the selected tool and its arguments, whether authorization allowed the request, whether the dummy state changed as expected, and what final response the application produced.
- Orchestration: Confirm the application routes requests, executes tools, performs handoffs, and advances or ends a workflow as intended.
- Policy enforcement: Check that application-side authorization, guardrails, and approval requirements are applied to proposed actions.
- Reliability paths: Exercise retries, retry limits, timeouts, errors, and circuit-breaker behavior with controlled responses.
- State and audit behavior: Inspect session or memory changes, recorded decisions, and whether the resulting trace reflects the expected workflow.
OpenAI’s current Agents SDK testing guides for Python and JavaScript describe deterministic, provider-neutral testing utilities that make no model-provider requests. They cover orchestration behavior such as tools, handoffs, guardrails, retries, streaming, and sessions. In LangChain Core’s v1.6.2 reference, fake chat-model options include FakeMessagesListChatModel, FakeListChatModel, and GenericFakeChatModel. These are framework-specific options; do not assume their APIs or behavior are identical across versions or languages.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
Build a deterministic test around the production entry point
- Choose a test double supported by your stack. Prefer an SDK-supported scripted model or a model abstraction designed for dependency injection over patching unrelated internals. Script the sequence of responses needed by the scenario, including tool-call responses and the response after tool execution where applicable.
- Use the same application entry point as production. Supply the test double through the application’s normal configuration or dependency boundary. This lets the test exercise the orchestration code that will handle model output in deployment.
- Make tools safe and observable. Replace effectful tools with instrumented fakes that record attempted calls, validated arguments, permission decisions, and dummy-state changes. Use synthetic credentials and marker data; do not let fixtures or test tools reach production systems.
- Assert actions as well as the final response. Check normalized model input, requested tool, arguments, authorization outcome, approval state, tool result, state mutation, and final output. A plausible refusal text is not sufficient evidence if a prohibited tool already ran.
- Check the script was consumed as expected. Assert that the test double used the expected responses or steps. Otherwise an unexpected change in control flow may leave a step unused while the test still appears to pass.
- Control test telemetry. Disable tracing or capture it in a test-safe destination when traces could otherwise export prompts, fixture content, or simulated actions.
Exercise both allowed and denied tool requests
For an allowed workflow, script a tool request with known arguments, let an instrumented fake return synthetic data, and script the follow-up response. Assert that the application validated the arguments, applied the intended authorization and approval rules, invoked only the expected fake, and produced the intended state transition.
For a denied workflow, script a request for a prohibited action and assert that authorization blocks it before any effectful implementation is reached. Include malformed arguments and denied approval as separate cases: they test different checks. Also exercise timeout or error responses, bounded retries, and any circuit breaker. These tests verify how application code handles those conditions; they do not establish that a real model will choose the same action or generate the same malformed request.
Design security cases around trust boundaries
OWASP’s AI Agent Security Cheat Sheet recommends structured abuse cases and testing controls around the agent’s actual capabilities. For each case, define the threat, input surface, intended policy, safe synthetic context, and observable result before writing the test.
Rank #2
- Prompt override: Test instructions that ask the agent to disregard its rules through both direct user input and indirect content, such as a retrieved document.
- Unauthorized use and privilege escalation: Attempt tools or permissions outside the user’s intended authorization, and assert that the application denies the action.
- Memory poisoning: Feed untrusted content that attempts to persist a false instruction or privilege into memory, then inspect what is stored and how later actions are authorized.
- Sensitive-data exfiltration: Use synthetic secrets or marker values and assert they are not passed to disallowed tools or exposed in an output channel that should not receive them.
- Recursive tool use and resource abuse: Exercise loops, excessive calls, retries, and token or cost limits; verify caps and circuit breakers stop further work.
- Approval bypass: Attempt to proceed with a sensitive action without the required approval, including after a denial or timeout.
- Multi-agent boundary violations: Check that one agent cannot use another agent’s permissions or treat untrusted inter-agent content as authorization.
Put indirect injections in the content channel under test
A prompt injection embedded in a retrieved page, email, file, or tool result is an indirect injection: the agent encounters it through an external-data channel. Put the test payload in that document or tool output and run the normal ingestion path. Placing the same text only in a user message tests a different trust boundary and does not show how the system handles malicious retrieved content.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsOWASP’s LLM Prompt Injection Prevention Cheat Sheet presents its hand-picked attack and benign examples as illustrative smoke tests, not a representative benchmark of real application traffic or attacks. Adapt cases to your supported tasks, permissions, tools, and content sources. Include benign controls alongside abuse cases; an agent that refuses every request should not pass as secure simply because it blocks attacks.
What passing stub tests do not establish
A stub is instructed by the test author to emit known output. It can show what the surrounding code does when that output arrives, but it cannot show whether a real probabilistic model will select a safe action, follow instructions reliably, resist an unfamiliar attack, or choose the right tool. Do not report a passing scripted suite as evidence that the deployed model is robust to prompt injection.
Likewise, a scripted model does not establish provider request serialization, authentication, HTTP headers or defaults, response parsing, or wire-protocol correctness. For those concerns, test the real provider adapter with a mocked or controlled HTTP transport. Use separate sandboxed or provider integration tests when you need evidence about actual execution or isolation behavior.
| Question to answer | Suitable test approach | What the result supports |
|---|---|---|
| Does application code authorize a proposed tool action and handle the resulting workflow correctly? | Deterministic scripted-model test with instrumented fake tools | Behavior of the tested application path for the scripted responses and fixture configuration |
| Does the provider adapter construct and parse requests correctly? | Real adapter with a mocked or controlled HTTP transport | Adapter behavior for the exercised request, transport, and response cases |
| Does the actual model behave safely on supported tasks and attacks? | Model-backed evaluations and scoped red-team tests against the supported configuration | Observed outcomes for the tested model, configuration, cases, and attempts—not a guarantee against novel attacks |
| Does an action execute safely in its intended environment? | Sandbox or provider integration tests | Behavior in the tested execution and isolation setup |
Pair deterministic tests with model-backed evaluations
Use model-backed evaluations when the question depends on real model behavior: instruction following, tool-selection quality, or resistance to direct and indirect attacks. NIST’s Center for AI Standards and Innovation (CAISI) explains that LLM outputs can vary from attempt to attempt and recommends adaptive evaluations, multiple attempts, and task-specific analysis alongside aggregate results.
In a January 17, 2025 CAISI article about agent hijacking, an evaluation on particular held-out tasks found that the strongest new attack raised measured attack success from 11% for the strongest baseline to 81%. Those are results for that evaluation’s models, attacks, tasks, and setup—not a general failure rate for agents. The article also describes indirect prompt injection arriving in content an agent ingests, including email, files, and websites.
Rank #4
When selecting or designing an evaluation, compare:
- How closely tasks and tools represent the agent’s supported work.
- Whether attacks cover the relevant channels, including retrieved and tool-returned content.
- Whether the evaluation is adaptive and how many attempts each case receives.
- Whether results show task-specific outcomes as well as an aggregate score.
- Whether runs are repeatable enough to compare changes and include useful trace or audit evidence.
AgentDojo’s authors describe an extensible environment for testing prompt-injection attacks and defenses. Their June 19, 2024 paper reports 97 realistic tasks and 629 security test cases in that research release; it also notes that state-of-the-art models fail some ordinary tasks even without an attack. A benchmark is a test environment, not a certification or guarantee of production safety.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Make the suite a release control and retain evidence
Run structured security tests before launch and again after material changes to prompts, tools, memory, retrieval, policies, or model providers. Keep known failures as regression cases. Review changes to security tests alongside changes to agent behavior; deleting or weakening a test can conceal a regression just as surely as a failing test can reveal one.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →For each release, retain a record that lets reviewers reconstruct what was tested and what happened:
- Agent version and relevant model provider, model identifier, and configuration.
- Tool policy and retrieval configuration used in the run.
- Fixture and test-case identifiers, expected outcomes, and observed outcomes.
- Approval, denial, timeout, retry, and circuit-breaker behavior, including failures and remediation.
- Accepted residual risks and any compensating controls.
OWASP’s LLMSVS v2.0, published in 2026, organizes verification requirements into eight groups, V1–V8, covering areas including secure configuration and maintenance, model lifecycle, model memory and storage, secure LLM integration, agents and plugins, dependencies, and monitoring. It is a verification framework; consulting its checklist does not certify a system.
NIST ITL’s AI Program project, Building Evaluation Probes into Agentic AI, created May 1 and updated May 5, 2026, offers a complementary traceability pattern: connect claims or decisions to source evidence and assess whether the evidence supports the claim (faithfulness), captures the source’s message (completeness), and meets the claim’s evidentiary burden (sufficiency). That project concerns evaluation probes and grounding; it is not a complete security governance standard.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

