What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To keep a retrieved passage meaningful after chunking, generate a short, chunk-specific explanation of where it belongs in its source document, then use the contextualized text for both embedding-based retrieval and BM25. In Spring AI, treat that as an ingestion and indexing step; use Spring AI’s RAG components to retrieve and assemble results at query time. Virtual threads are a separate choice for dispatching Anthropic HTTP requests, not a way to improve retrieval quality.

Why chunks lose useful context

Splitting a document makes passages easier to retrieve, but a passage can lose facts that were stated elsewhere. A chunk might say, “The company’s revenue grew by 3% over the previous quarter,” without naming the company or the period. A reader asking “What was the revenue growth for ACME Corp in Q2 2023?” needs those details to connect the passage to the question.

Anthropic describes Contextual Retrieval as a preprocessing technique: give a language model the full document and one chunk, ask it for concise text that situates that chunk, and prepend the result to the chunk. The added context can identify the relevant entity, time period, section, or argument. It is specific to the chunk, rather than a single generic summary repeated across every passage.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How contextual retrieval fits into an index

  1. Parse and split the source. Keep the original document and its provenance, and create chunks using boundaries and overlap appropriate to the material.
  2. Generate context for each chunk. Supply the full document and the particular chunk to a contextualizer. Ask for a short description of the chunk’s place in the document, and have it return only that context.
  3. Keep context distinguishable from source text. Store the original passage and generated context separately, or mark the contextual prefix clearly. This helps preserve provenance and lets downstream prompts distinguish generated framing from source evidence.
  4. Index the contextualized text in both retrieval channels. Use it as the embedding input and include it in the BM25 index. If the system uses only one channel, evaluate that choice against hybrid retrieval rather than assuming the same result.
  5. Retrieve, assemble, and evaluate. At query time, retrieve candidate passages and include them in the answer prompt. Test representative questions against known relevant passages, and check whether the contextual prefix improves retrieval without obscuring the original evidence.

Anthropic reports that generated context is typically 50–100 tokens. That is a starting point, not a universal setting: document style, terminology, chunk size, overlap, embedding model, and retrieval depth can all change what works. Its article also reports limited gains from generic document summaries in its evaluation, which is why the chunk-specific relationship to the full document matters.

#1 Best Overall

What Anthropic’s reported results show—and do not show

Anthropic’s 2024 engineering article reports top-20-chunk retrieval failure rates averaged across codebases, fiction, arXiv papers, and science papers, using the top-performing embedding configuration in its analysis. The reported figures are:

Approach in Anthropic’s evaluation Top-20 retrieval failure rate
Baseline 5.7%
Contextual Embeddings 3.7% (35% lower than the baseline)
Contextual Embeddings plus Contextual BM25 2.9% (49% lower than the baseline)

These are results from Anthropic’s evaluation, not a guarantee for a different production corpus. They measure whether relevant material was missed among the top 20 retrieved chunks; they do not establish answer quality for every downstream model or task.

A separate Anthropic Cookbook example evaluates nine codebases with basic character splitting and 248 queries, each with a “golden chunk.” In that setting, Anthropic reports Pass@10 improving from about 87% to about 95% with Contextual Embeddings. That result uses a different dataset and metric from the top-20 failure-rate figures, so it should not be treated as a directly interchangeable measurement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Budget for contextualization as ingestion work

Anthropic’s 2024 article gives an illustrative one-time cost of $1.02 per million document tokens under a specific set of assumptions: 800-token chunks, 8,000-token documents, 50 tokens of context instructions, and 100 generated context tokens per chunk. The figure assumes prompt caching and is a historical estimate from that article, not a current provider quote or a general per-million-token rate.

For a real pipeline, measure the cost and operational impact against the corpus and model you use. Include contextualizer input and output tokens, cache behavior, document update frequency, and the cost of regenerating context and rebuilding affected index entries when a source changes. Contextualization can improve recall while adding model calls and preprocessing latency; those are separate trade-offs to assess.

Choose the Spring AI retrieval layer that fits the application

Spring AI documents two entry points for retrieval-augmented generation. QuestionAnswerAdvisor is the simpler path: it queries a vector store and appends retrieved documents to the prompt. The more modular RetrievalAugmentationAdvisor supports composing stages such as query transformation, retrieval, document joining, post-processing, and query augmentation. The documented dependency for the first path is spring-ai-vector-store-advisor; the modular RAG API is documented with spring-ai-rag.

The Spring AI reference displayed version 2.0.1 when accessed on October 7, 2026. Check the live reference and the Spring AI BOM used by your application before relying on version-specific dependency coordinates or APIs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Neither advisor should be confused with the preprocessing step. In particular, Spring AI’s ContextualQueryAugmenter augments a user query with contextual data from documents that have already been retrieved. It does not generate Anthropic-style chunk-specific context from a full source document before indexing.

Use virtual threads for Anthropic HTTP dispatch when it fits

Spring AI’s Anthropic integration documents a dispatcherExecutor option for the synchronous and asynchronous streaming clients. Its configuration pattern is:

AnthropicChatModel chatModel = AnthropicChatModel.builder()
    .options(...)
    .dispatcherExecutor(Executors.newVirtualThreadPerTaskExecutor())
    .build();

This makes a virtual-thread-per-task executor available for HTTP dispatch. The documentation presents it as an option for high HTTP concurrency or Java 21-and-later workloads, not as a universal performance improvement. Whether it helps depends on the application’s concurrency and workload; it does not reduce contextualization calls or guarantee faster retrieval.

Own the executor lifecycle if you supply it

If the application creates and supplies the ExecutorService, the application is responsible for shutting it down. Spring AI does not call shutdown() on an executor supplied this way. Arrange cleanup as part of the application’s lifecycle so the executor is closed when its owning component or application stops. If the option is omitted, Spring AI creates and cleans up its internal executor.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the RAG advisor executor separate

A Spring AI engineering example describes a different issue in a modular RAG advisor: its per-query retrieval threads are non-daemon, so the command-line example remained alive after printing its answer. In that example, passing Spring Boot’s auto-configured TaskExecutor through .taskExecutor(...) addressed the issue, and spring.threads.virtual.enabled=true enabled virtual threads for that configuration.

That advisor setting is not the Anthropic HTTP dispatcher setting. They belong to different parts of the system and have separate configuration and lifecycle responsibilities. The same engineering example notes that its illustrated flow makes two LLM calls before retrieval and one service call per retrieved chunk; measure latency and cost for the actual flow before adopting it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Account for streaming trace context

Spring AI’s Anthropic integration reference says synchronous HTTP spans are nested beneath the model operation, but streaming HTTP spans may not be. It attributes the gap to the Anthropic Java SDK’s asynchronous implementation switching to ForkJoinPool.commonPool() before calling Spring AI’s HTTP client, which can lose the calling thread’s observation context. The reference says traceparent is still propagated and suggests correlating okhttp.requests with the model operation by trace ID or timestamp range.

Verify this behavior against the exact Spring AI and Anthropic SDK versions in use: asynchronous tracing behavior can change. If streaming requests appear as separate spans, inspect the trace ID and timing relationship rather than assuming that the HTTP request did not occur.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Evaluate the whole retrieval design on your corpus

Contextual Retrieval adds information that can help a retriever find a passage whose original wording is underspecified. It also adds preprocessing, generated text, and another part of the index to manage. Compare alternatives using the same representative queries and retrieval setup, and record at least:

  • Retrieval quality: Use a defined measure such as failure at a stated top-k, recall, or a task-specific metric, and preserve the corpus and query-set context alongside results.
  • Context strategy: Compare no prefix, a generic summary, and per-chunk context generated from the full source.
  • Retrieval channel: Evaluate semantic embeddings and BM25 separately as well as in a hybrid approach, where applicable.
  • Preprocessing burden: Track contextualizer calls and token use, caching, document-change frequency, and index update work.
  • Operational behavior: Check how original text and generated context are stored, how the Spring AI retrieval stages are composed, whether supplied executors are closed, and whether streaming traces can be correlated.

Anthropic’s results support testing contextual embeddings and contextual BM25 as candidate techniques; they do not settle the right chunking, prompt, embedding choice, retrieval depth, or architecture for another corpus. Treat the reported improvements as a reason to run a controlled evaluation, not as a production promise.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.