Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To build a RAG knowledge base with OpenSearch Serverless and Node.js 22, store document chunks and their embeddings in a vector search collection, retrieve relevant chunks for each question, then pass those passages to a language model to generate a grounded answer. OpenSearch handles retrieval; it does not, by itself, complete the generation step. “Real-time” is an architecture goal, not a guarantee of zero-delay updates or a service-wide latency SLA.

How the RAG flow works

Retrieval-augmented generation (RAG) gives a language model relevant source material at answer time. The application prepares and indexes knowledge ahead of time, then retrieves a small set of passages for each question.

  1. Prepare sources: clean documents, divide them into chunks, and attach useful metadata such as a document ID, title, or access category.
  2. Embed and index: turn each chunk into a vector with an embedding model, then store the vector alongside the text and metadata in an OpenSearch Serverless vector search collection.
  3. Retrieve: embed the user’s question using compatible embedding behavior, and search for relevant chunks. Depending on the search design, retrieval may also use lexical keywords.
  4. Generate: give the retrieved passages to a language model as context and instruct it to answer from that material. The application can call the model, or the architecture can use an OpenSearch remote-model connector.
  5. Refresh: when a source changes or is removed, update or delete its indexed chunks so future retrieval reflects the change.

The embedding model used for documents and queries must produce compatible vectors, including matching dimensions. A collection’s index mapping must also match the chosen vector representation; there is no universal mapping or chunk size that fits every model and corpus.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the collection before indexing

OpenSearch Serverless offers collection types for different search workloads, including vector search. The collection type is chosen when the collection is created and cannot later be changed. Decide on the collection generation and type before loading data, and verify the current feature support and constraints for the workload you need.

AWS documentation describes both NextGen and Classic generations. NextGen is described as offering instant auto scaling and scale-to-zero, but generation capabilities and constraints differ. Do not assume every feature or behavior is available in both generations.

  • Identity and permissions: give the application an AWS identity with only the data access it needs. Configure the collection’s data access policy for the intended principal.
  • Network access: configure network access so the application can reach the collection endpoint from its runtime. A correctly signed request cannot overcome a network policy that blocks access.
  • Encryption: select and configure encryption in line with the data and account requirements.
  • Index design: define fields for chunk text, vector data, and metadata before ingestion. The vector field’s dimensions and representation must agree with the embedding model in use.

Connect a Node.js application with SigV4

A JavaScript client must sign requests for OpenSearch Serverless using AWS Signature Version 4. AWS’s JavaScript example uses the OpenSearch client package, AwsSigv4Signer, signing service aoss, a region, and the collection endpoint. The example is not a Node.js 22 compatibility certification; check the current package and runtime support for your deployment before treating that combination as validated.

const { Client } = require('@opensearch-project/opensearch');
const { AwsSigv4Signer } = require('@opensearch-project/opensearch/aws');
const { defaultProvider } = require('@aws-sdk/credential-provider-node');

const getCredentials = defaultProvider();

const client = new Client({
  ...AwsSigv4Signer({
    region: process.env.AWS_REGION,
    service: 'aoss',
    getCredentials,
  }),
  node: process.env.OPENSEARCH_COLLECTION_ENDPOINT,
});

Set AWS_REGION to the collection’s region and OPENSEARCH_COLLECTION_ENDPOINT to its endpoint. Use an AWS credential provider appropriate to the runtime, such as the runtime’s assigned role in a deployed environment; do not hard-code access keys. The signer authenticates requests, but it does not create the collection, grant data permissions, or configure network access.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Once access and the index mapping are in place, the client can create an index and write documents. A document should keep each chunk’s text, vector, and metadata together, so retrieval can return useful context and the application can trace it to its source. The exact index mapping and vector field depend on the embedding model and search design.

Ingest content: direct writes or managed pipelines

There are two broad ingestion approaches. Pick based on who owns source-change events and how much transformation and streaming infrastructure the application should manage.

Approach Best fit Trade-off
Application writes through the OpenSearch JavaScript client The application already detects changes and needs direct control over processing and writes. More control over event handling, but the application owns embedding, retries, updates, deletes, and related operational work.
OpenSearch Ingestion A managed pipeline is useful for collecting, transforming, or streaming incoming data. Centralizes pipeline processing, while adding pipeline configuration and operations.
S3-based vector ingestion Content and vector-ingestion workflow fit an S3-based loading path. Provides a managed vector-ingestion option; confirm its supported input and processing behavior for the workload.

For direct writes, the application typically processes a source change, prepares chunks, obtains embeddings, and indexes the resulting records. Make the update path handle replacement and deletion as well as first-time indexing: otherwise, old chunks can remain searchable after the source has changed.

OpenSearch Ingestion is a managed pipeline for ingesting and transforming data. AWS also documents vector ingestion from S3. These approaches change where collection and transformation work happens; they do not remove the need to decide how source changes map to indexed chunks or how stale records are removed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retrieve passages that match a question

At question time, the application embeds the query with behavior compatible with the document embeddings, then searches the vector index. Neural search is useful when a question expresses the same idea in different words from the source. Hybrid search combines semantic retrieval with lexical search, which can help when exact terms, names, identifiers, or phrases matter alongside conceptual similarity.

Retrieval method What it emphasizes Useful when
Semantic or neural Meaning represented by embeddings The question and source use different wording for the same concept.
Hybrid lexical and semantic Both keyword matches and semantic similarity Exact terms and conceptual matches both affect relevance.

Search quality depends on choices the service does not make for every application: chunk boundaries, metadata filters, the number of passages returned, and how results are ranked or combined. Test those choices against representative questions and source material. Metadata can also restrict retrieval to the right document set or access category, but the application must enforce its own authorization rules rather than treating search relevance as access control.

AWS documents up to 15 seconds of latency for neural searches against a vector index or recently created search or ingestion pipelines. That qualification applies to those documented circumstances, not to every RAG query. Measure freshness and query latency in the actual deployment, and avoid promising that a newly changed document will be visible immediately.

Connect retrieval to generation

After search returns passages, send the relevant text and source identifiers to a language model along with the user’s question. Make the instruction clear that the answer should be based on the supplied context; if the context does not support an answer, the model should say so rather than inventing details. Preserve source identifiers so the application can display citations or let users inspect the underlying material.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Model-access approach What the application owns Trade-off
Application calls the model separately Retrieval-to-model orchestration, request handling, and context formatting Offers application control over the generation flow, while requiring the application to coordinate the model call.
OpenSearch Serverless remote-model connector Configuration and permissions for the connector and its model access Can place model integration closer to the search workflow, with additional connector configuration and coupling to that setup.

OpenSearch supplies retrieval, not the full RAG answer by itself. A separate model call or configured remote-model path supplies generation. AWS architecture guidance presents Bedrock as one possible model route, not the only required choice.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What “real-time” means in practice

In this architecture, real-time usually means that source changes trigger ingestion promptly and that the application can query the updated index without a separate batch release. It does not establish a fixed end-to-end freshness or response-time guarantee. Indexing, embedding, pipeline processing, and search visibility all affect when a change can influence an answer.

  • Track source-change time, indexing completion, and query visibility separately.
  • Update or delete all affected chunks when a source is edited or removed.
  • Monitor retrieval latency and freshness in the chosen collection generation, ingestion path, and model configuration.
  • Set user-facing expectations from observed behavior rather than assuming immediate visibility.

AWS describes OCU allocation-based charging for vector ingestion, but no price is established here. Check the current AWS pricing information when estimating operating cost; ingestion design affects not only latency and control but also the services and capacity used.

A practical build sequence

  1. Choose collection generation and type. Confirm the required vector-search features and constraints, then create the collection. Treat the collection-type choice as irreversible.
  2. Set access and connectivity. Configure encryption, network access, and a least-privilege data access policy for the application identity.
  3. Define the embedding and index design. Choose compatible document- and query-embedding behavior, vector dimensions, chunking rules, metadata, and index mapping.
  4. Implement ingestion. Use signed direct writes when the application owns source events, or evaluate OpenSearch Ingestion or S3 vector ingestion when a managed path suits the data flow.
  5. Implement retrieval. Embed user questions, run semantic or hybrid search, apply appropriate metadata filters, and return passages with source identifiers.
  6. Implement generation. Pass retrieved context and the question to a separately configured model call or remote-model connector, and preserve sources for traceability.
  7. Test updates and operations. Verify first ingestion, edits, deletions, access boundaries, retrieval relevance, and observed freshness and latency with representative data.

Node.js 22 can be the application runtime for this design, but the cited AWS JavaScript example does not itself establish compatibility with Node.js 22. Confirm support for the current client and signer packages in the exact runtime and deployment environment you plan to use.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.