Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

If an LLM evaluation returns the same completed result after its input changes, first check for a cache in your application or evaluation harness. That is different from provider prompt caching: OpenAI describes prompt caching as reusing computation for a matching rendered input prefix, not as returning a previously completed evaluation result. Without your code or logs, the exact cause cannot be identified from the symptom alone.

Which cache could be involved?

Start by distinguishing what was reused. One cache may reuse intermediate computation while processing a request; another may return a stored model or grader output without running the evaluation again.

Cache layer What is reused What to inspect
Provider prompt cache Key-value computation for a matching rendered prompt prefix Rendered prefix, compatible request settings, and provider cache diagnostics or usage where available. OpenAI prompt caching documentation
Application or evaluation-harness result cache A completed output or evaluation result The cache-hit record, cache key, and provenance of the stored result. The precise behavior depends on your implementation; OpenAI’s Create eval API reference does not establish a universal result-cache key schema.

A provider prompt cache is not, by itself, evidence that an old completed evaluation result was returned. Attribute the symptom to a particular layer only when request traces, logs, or code support that conclusion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How can I tell where the old result came from?

  1. Reproduce one changed evaluation item. Record the old and new input, expected and actual output, evaluation or run identifiers, and timestamps. This creates a trace to compare; it does not assume a particular cache caused the result.
  2. Check whether a model or grader call occurred. If the old completed output is returned before such a call, investigate the harness or application result cache. If a request reached the provider, inspect that request and its prompt-cache evidence. This is a diagnostic distinction, not proof on its own.
  3. Log the exact result-cache key. Compare the key for the old and new runs, then verify that every value used to construct it reflects the changed input. Look for omitted fields, stale normalization, accidental reuse across dataset rows, mutable references, and missing version information.
  4. Trace the cached result to its inputs and configuration. Record enough provenance to identify what produced it, including the input or a stable digest and the relevant prompt, model configuration, dataset example, grader, and tool or retrieval versions. These are general engineering recommendations, not a key schema specified by the cited OpenAI documentation.

What should I check if the provider prompt cache is involved?

OpenAI describes prompt caching as reuse of key-value tensors for a matching rendered input prefix. It does not store the prompt tokens themselves. Cache compatibility depends on the rendered prefix and request settings, so compare the actual request sent—not only the source template or the input field you changed. See OpenAI’s prompt-caching guide.

OpenAI’s prompt-cache diagnostics can compare a current request with an earlier response to investigate why an expected prefix was not reused. Documented diagnostic reasons include input_changed, tools_changed, text_format_changed, reasoning_effort_changed, verbosity_changed, and context_compacted. For input_changed, earlier input may have changed or been reordered; timestamps or request IDs in instructions are examples of dynamic content that can change a prefix.

When you want stable prefix reuse, keep reusable content before dynamic content and place changing values after the reusable prefix and its cache breakpoint, as the diagnostics guidance recommends. Also compare settings that affect compatibility: model, tools and tool ordering, output format or schema, reasoning effort, verbosity, and context management. A changed request can explain why a prefix was not reused; it does not, by itself, explain an old completed evaluation output.

How should an evaluation result cache avoid stale hits?

There is no universal cache-key schema established by the cited API documentation. As an engineering practice, a result key should distinguish inputs and configurations whenever they can change the output. Depending on the evaluation, that may mean incorporating:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The complete evaluation input, or a stable digest of it.
  • The prompt or template version and relevant model configuration.
  • The dataset or example identity and version.
  • The grader configuration and version.
  • Tool, retrieval, or other dependency versions that can affect the result.

Store provenance alongside the cached result so a cache hit can be traced to the input and configuration that produced it. If the cache key is correct but a value is still served after a relevant change, inspect invalidation, normalization, and versioning logic. Choose expiration or invalidation rules for your own result cache based on its requirements; provider prompt-cache retention settings are not a guide to evaluation-result TTL.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What do OpenAI prompt-cache retention settings mean here?

OpenAI documents prompt-cache retention as model-dependent. Its current guide gives GPT-5.6-and-later a supported minimum-lifetime setting or default of 30m and describes retention options for earlier models. Those details concern provider prompt caching, not how long an application should retain a completed evaluation result. Check the guide for the applicable model before relying on a retention setting: Prompt caching.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.