Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Replay fixtures let you rerun an agent evaluation using recorded model responses instead of calling the live model API for every case. That reduces API availability and response variation as confounders—but a replay score describes the recorded fixture and the harness that ran it, not every condition or external effect of a live run.

What a replay fixture changes

In a live evaluation, each case calls the target model or agent. A replay substitutes a saved response for that target call, so repeated runs can focus on whether the evaluation harness, orchestration, and graders handle the same recorded interaction consistently. The agent-eval-kit Core Concepts documentation describes live mode, replay mode, and judge-only mode; it says replay runs from fixtures without API calls.

This is useful when you want to compare grader changes or investigate a regression without making the result depend on whether a provider is reachable or returns a different answer today. It does not establish that the agent would produce the same response in a fresh live call.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the score does—and does not—prove

It can test the recorded path

The OpenAI Agents Python discussion proposes capturing normalized model requests and responses, then replaying them through a ScriptedModel while the actual Runner and orchestration logic execute again. That boundary can help isolate model-boundary inputs and outputs while checking how the harness processes them. The discussion is a proposal, not evidence that this capability shipped: OpenAI Agents Python issue #4795.

It cannot certify uncaptured effects

The proposal leaves external tools and sandbox side effects outside its initial fixture scope. If a run sends an email, changes a record, invokes a tool, or depends on a retriever index, replaying a model response alone does not prove that the real effect occurred or that the external system behaved the same way. Stub or test those boundaries separately and state plainly what the fixture did not capture.

Build a replay workflow you can trust

  1. Capture deliberately. Record a representative live interaction when needed rather than recording everything by default. The OpenAI proposal recommends opt-in capture because fixtures may contain prompts, tool arguments, and model output; it also proposes a redaction or transformation hook before writing data.
  2. Version the fixture and harness. Identify the fixture, target, evaluation suite, orchestration, and grader versions with each result. The OpenAI proposal calls for versioned deterministic JSON fixtures. The agent-eval-kit documentation describes a configuration hash based on suite name and targetVersion.
  3. Replay the same interaction and grading setup. Run the fixture through the intended orchestration and graders, and report the fixture identity with the score. A score without those details is difficult to interpret or reproduce.
  4. Test excluded boundaries independently. Exercise external tools, sandboxes, data stores, and other side effects with dedicated tests or controlled stubs. Label which of them the replay omitted.
  5. Refresh fixtures when behavior changes. A fixture can become stale when the target, prompt, tools, or other behavior changes. Agent-eval-kit documentation describes a configurable fixture TTL with a 14-day default and warnings or errors in strict mode. Treat that as this project’s documented setting, not a universal freshness standard; check the current version’s documentation before relying on it.

Keep grading criteria visible

Replay controls the responses being graded; it does not make every grading method deterministic. Agent-eval-kit documents deterministic graders, LLM graders, weighted scores, and optional gates for pass rate, maximum cost, and p95 latency. State which graders contributed to a score, the rubric they applied, and which gates were required. If an LLM judge is involved, distinguish its result from deterministic checks rather than presenting the combined score as wholly deterministic.

The project also describes judge-only mode, which re-grades an existing run with updated graders. That is different from replaying a fixture: judge-only mode changes the grading pass without rerunning the target, while replay runs from recorded fixtures. See the agent-eval-kit Core Concepts documentation for its mode definitions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to include in a replay report

  • Fixture identifier, capture date, and version or content hash.
  • Target, suite, orchestration, and grader versions relevant to the run.
  • Whether the run was live, replayed, or judge-only.
  • Capture boundary: what was recorded and what was stubbed, omitted, or tested separately.
  • Grading rubric, deterministic and LLM grader results, and any pass-rate, cost, or latency gates applied.
  • Fixture freshness policy and any warning or strict-mode failure.

These details make a replay result interpretable: it tells readers what was held constant, what was re-evaluated, and what the score cannot speak to.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Protect fixtures as sensitive data

Recorded prompts, arguments, and outputs can include confidential or personal information. Limit capture to what the evaluation needs, apply redaction or transformation before writing fixtures, and control access and retention as you would for the underlying data. The OpenAI issue proposes these safeguards; it does not establish a universal storage or retention policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.