Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Test an LLM application by defining observable success criteria, running representative cases through the complete system, grading each result with checks suited to the task, and repeating the evaluation after meaningful changes. A single score or a few hand-picked prompts cannot tell you reliably whether the app works: results depend on the model, prompt, data, tools, and evaluation method. Build a versioned test set, inspect failures, and treat each important score as evidence about one specific setup—not a universal measure of quality. This workflow applies to chat features, retrieval-augmented generation (RAG), and tool-using agents.

What an LLM application test should measure

An evaluation needs three parts: a task, inputs that exercise it, and grading logic that decides whether the system succeeded. For an individual case, that can mean a user request, relevant context, the app’s response or actions, and a judgment against explicit criteria. OpenAI’s evals guide describes the cycle as defining the task, running test inputs, and analyzing results so the system can be improved.

Write success in observable terms rather than relying on “the answer seems good.” Depending on your feature, success might mean answering the question correctly, grounding the response in supplied evidence, returning valid JSON, calling an authorized tool with the right arguments, or leaving an external system in the intended state. Define unacceptable outcomes too: for example, a fabricated citation or an action taken without required confirmation.

Keep the evaluation target clear. A model call is not always the same thing as the application a user experiences. The deployed behavior may also depend on retrieval, prompt assembly, application code, tools, retries, filters, and the surrounding environment. If you only test the model’s final text, you can miss failures elsewhere in that path.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a test set that resembles real use

Start with representative cases rather than a large pile of convenient examples. OpenAI recommends expert-authored labels and a mix of typical, edge, and adversarial inputs in its evaluation best practices. Include examples that cover the ways people actually use the feature, the cases where it is likely to fail, and behavior that must not occur.

  • Typical cases: common requests, ordinary phrasing, and the usual range of input lengths.
  • Edge cases: missing or conflicting context, ambiguous requests, unusual formatting, long inputs, and boundary conditions in your product rules.
  • Adversarial and abuse cases: attempts to override instructions, extract hidden prompts, expose private information, trigger disallowed actions, or consume disproportionate resources.
  • Observed failures: reviewed production examples, support reports, and user feedback that can be safely and appropriately turned into test cases.

For each case, record the input and only the context and setup the application is meant to receive. Add a reference answer or label when a defensible one exists; for open-ended tasks, write a rubric that describes what a good response must include and what would make it fail. Version the dataset and preserve the cases that caught regressions. When a failure reveals a new behavior, add a minimized, understandable example to the suite so it stays testable.

Do not let the test set become a proxy for the whole product. A case collection is useful only to the extent that it represents the feature’s real users and risks. Track why each important case exists and review the set as the product, user population, and failure patterns change.

Choose a grader that matches the requirement

Different requirements call for different evidence. A deterministic check is more dependable than an LLM judge for a requirement that can be stated exactly; a human or rubric-based judgment is more appropriate for qualities such as helpfulness or tone. Model graders can help scale judgments, but should be validated against human labels rather than assumed correct. OpenAI’s guidance also flags position and verbosity biases in model-based judging.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Exact or programmatic checks: use for schema validity, required fields, exact identifiers, allowed values, forbidden strings, or whether a required citation field is present. A JSON parser can establish valid JSON; it cannot establish that the answer’s facts are true.
  • Reference-based checks: compare against an expected label or answer when there is a stable ground truth. Allow for legitimate variation if multiple responses can be correct.
  • Rubric grading: assess qualities that need judgment using criteria written in advance. A human reviewer is valuable for ambiguous or high-impact cases. A model judge can increase throughput, but calibrate it against human-rated examples and inspect disagreements.
  • Pairwise comparison: compare two system versions on the same case when choosing between alternatives. Randomize which answer appears first and account for verbosity and presentation effects that can sway judges.
  • Pass/fail gates: use explicit thresholds for requirements that must not regress, such as a prohibited tool call or a malformed required output. Keep the underlying case results available; an aggregate pass rate alone can hide a serious failure.

Use more than one grader when a critical property has distinct parts. For example, a response can satisfy formatting checks yet be factually wrong. Make each check answer a specific question, and retain the evidence that supports the judgment.

Evaluate RAG retrieval and generation separately

A RAG answer can fail because the system retrieved the wrong material, because the model misused good material, or both. Measure retrieval quality separately from answer correctness and grounding where possible. For a retrieval test, check whether the expected relevant documents or passages appear in the retrieved context, and whether irrelevant material overwhelms them. For generation, check whether the response answers the request using the provided evidence, whether its claims are supported, and whether it acknowledges missing evidence when that is the correct behavior.

Keep the retrieved context in the test record. Without it, an incorrect answer may look like a generation defect even though the needed source never reached the model. Likewise, a correct answer alone does not show that the retriever is reliable: the model may know the answer from prior training or guess it. OpenAI’s evaluation guidance recommends evaluating components as well as the overall result.

Test agents by their actions and outcomes

For an agent, the answer is only one part of success. Record the tools available, tool calls and arguments, relevant intermediate steps or trace, and the resulting state of the environment. A final message that says “done” is not proof that a ticket was created, a file was changed, or a booking was correct.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Define the intended outcome and the allowed path separately. Your grader might check that the agent reached the right state, avoided forbidden actions, and used tools appropriately. Some tasks may have several valid action sequences; score the required outcome and constraints rather than demanding a single exact transcript. Anthropic’s agent evaluation guide discusses tasks, trials, graders, transcripts, outcomes, and evaluation harnesses as parts of agent evaluation.

When the environment is consequential or the model’s behavior varies, run multiple trials on selected cases. A single successful run may conceal an unreliable policy. Record how many trials were run and report variation or failure rates alongside the outcome; do not present one lucky trajectory as proof of consistency.

Add safety and abuse evaluations

Normal product examples do not adequately test misuse or deliberate attacks. Create cases for risks relevant to your application, including prompt injection, prompt extraction, privacy leakage, adversarial inputs, denial of service, and policy-violating behavior. Check not only what the model says, but whether application controls prevent a harmful action or unauthorized disclosure.

Red teaming complements ordinary quality evaluation: it probes weaknesses rather than just checking expected task performance. Google’s Responsible Generative AI evaluation guidance and OpenAI’s red-teaming guide provide safety-evaluation context and risk categories. Choose tests for your product’s actual data, capabilities, and threat model; a generic set of adversarial prompts cannot establish that an application is safe.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Automate regression checks and inspect failures

Run the suite when a change could affect behavior: a prompt edit, model change, retrieval update, tool modification, application-code change, or safety-policy update. Compare results with a known baseline, then inspect individual failures before deciding whether the change is acceptable. OpenAI recommends continuous evaluation on changes and monitoring for nondeterminism in its best-practices guide.

  1. Freeze the test setup: record the dataset version, model identifier, prompt version, application and retrieval configuration, tool definitions, grader version, and relevant run settings.
  2. Run the same cases: keep inputs and grading logic stable when comparing versions. If any part changes, record it so the comparison remains interpretable.
  3. Review changed outcomes: inspect failures and regressions case by case. Determine whether the cause is the model, prompt, retrieved context, tool behavior, grader, or test data.
  4. Decide with risk in mind: an aggregate score can improve while a critical safety case fails. Set explicit release gates for must-pass behaviors and review trade-offs rather than relying on one headline number.
  5. Extend the suite: add a compact test for each newly understood failure and rerun it on future changes.

LLM outputs can vary between runs, so repeat tests where variability matters, especially for agents and high-impact behaviors. Report the number of trials and how results were aggregated. Do not treat repeated runs as interchangeable evidence if the model, prompt, tools, or environment changed between them.

Make evaluation results reproducible and bounded

A score only describes the system and procedure that produced it. When sharing results, identify the exact model, prompt, tools, harness, safeguards, dataset, budget, and grader. State what claim the evaluation tests, and disclose checks for shortcuts, dataset contamination, refusals, and evaluation awareness where relevant. OpenAI’s playbook for trustworthy third-party evaluations emphasizes the conditions that make evaluation claims interpretable.

Keep raw case-level results alongside summaries. A mean or pass rate is easier to scan, but can obscure which user group, input type, or safety behavior failed. If two runs differ, first check whether the setup was actually held constant; then examine individual cases and grader judgments. Avoid comparing scores produced by different datasets or grading rules as though they were a controlled head-to-head result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use an evaluation framework only where it helps

You can begin with a small script and a versioned data file. A framework becomes useful when it saves work you need—such as provider integration, structured test cases, red-team workflows, or CI/CD execution. Promptfoo documents a CLI and library for LLM evaluation and red teaming, including CI/CD usage; DeepEval documents end-to-end, trajectory-based, and component-level approaches. Compare them against your test target, evidence needs, integration requirements, and maintenance capacity rather than assuming one is universally best: Promptfoo introduction; DeepEval introduction.

The following standard-library Python example shows the shape of a minimal deterministic regression harness. Replace the illustrative function with a call to your application, and replace its toy checks with criteria appropriate to your product. It requires Python 3 and does not call an LLM or claim to evaluate open-ended answer quality.

import json

CASES = [
    {"id": "capital-france", "question": "What is the capital of France?", "expected": "Paris"},
    {"id": "capital-japan", "question": "What is the capital of Japan?", "expected": "Tokyo"},
]

def app_under_test(question):
    # Replace this example with your application call.
    answers = {
        "What is the capital of France?": "Paris",
        "What is the capital of Japan?": "Tokyo",
    }
    return {"answer": answers.get(question, "I don't know")}

failures = []
for case in CASES:
    result = app_under_test(case["question"])
    passed = isinstance(result, dict) and result.get("answer") == case["expected"]
    record = {"id": case["id"], "passed": passed, "result": result}
    print(json.dumps(record, ensure_ascii=False))
    if not passed:
        failures.append(case["id"])

raise SystemExit(1 if failures else 0)

For a real application, preserve the input, relevant retrieval context, tool trace, raw output, grader result, and configuration identifiers needed to diagnose a failure. Avoid logging secrets or personal data unnecessarily; evaluation records should follow the same access and retention rules as other sensitive application data.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Capture web inputs for browser-based LLM features

If the feature you are testing reads or acts on websites, use controlled page inputs: record the URL, viewport or device settings, relevant page state, and any screenshot or extracted content the model receives. A screenshot can be useful for a vision-capable model or a browser-agent test, but it evaluates the input capture—not the LLM’s answer or behavior. Check that the captured page represents the state your product is meant to handle.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For this specific browser-input task, ScreenshotNeo is a screenshot API and MCP server, not an LLM evaluation framework. Its capture options include full-page screenshots, element capture, viewport and device settings, custom CSS or JavaScript, and waiting for a selector, delay, or network idle. Cookie-banner and popup cleanup may change the page input; the cleanup steps can be turned off when a test needs an untouched page.

Or skip the browser setup

One GET request can return a screenshot; read the ScreenshotNeo API documentation for parameters and response details. This cURL example writes a WebP file:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Or in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, and failed loads are not billed. Its MCP server offers screenshot tools for AI agents, including Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, with no card required.

Common evaluation failures and fixes

  • The score is high, but users still find obvious errors. The cases may be too easy or unrepresentative, or the grader may reward style over correctness. Add representative failures and edge cases, then review judgments against human labels.
  • A RAG answer fails, but the cause is unclear. Save the retrieved context and test retrieval separately from answer generation. The model cannot ground its answer in evidence it never received.
  • An agent passes once and fails on a rerun. The task may be nondeterministic or depend on environment state. Capture traces and outcomes, run repeated trials where needed, and make setup and reset behavior explicit.
  • The model judge favors one candidate. Position or verbosity may influence its choice. Randomize presentation order, use a defined rubric, and compare a sample of judgments with human review.
  • A code or schema test passes despite a bad answer. Structural validity is not semantic correctness. Keep format validation as its own check and add a separate factual or task-specific grader.
  • A regression appears only after deployment. The test set may omit real traffic patterns or production dependencies. Use appropriately reviewed examples to grow the dataset and test the components that differ between the harness and deployed system.
  • Scores from two reports appear directly comparable. Check whether the dataset, model, prompt, harness, grader, and run budget match. If they do not, describe the setups separately instead of treating the score difference as a controlled result.

Further reading

For a broader treatment of evaluation alongside prompt engineering, RAG, agents, and AI application development, see the publisher’s listing for Chip Huyen’s AI Engineering (ISBN 9781098166298). A book can provide context; your own representative cases are still needed to test your application.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Do I need a benchmark before I can test an LLM feature?

No. A small, carefully labeled set of representative cases can reveal regressions before you have a formal benchmark. Expand and version it as you learn where the feature fails.

Can an evaluation score guarantee that an LLM application is safe?

No single score establishes safety. Results are limited to the tested cases, setup, and grading method; safety work also needs risk-specific probes and appropriate controls.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.