Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To add an LLM feature with LangChain4j, start with a Java 17-or-newer project, add an integration for your chosen model provider, supply its API key through environment configuration, and make one direct ChatModel call. Once that works, move to an AI Service for a typed application interface, then add memory, tools, or retrieval only when the feature needs them. The examples below use the versions and names shown in LangChain4j documentation; check the current docs before copying them because artifacts and provider/model identifiers can change.

What LangChain4j does in a Java application

LangChain4j is a Java library for integrating large language models and related components. Its unified APIs cover model providers and embedding stores, with capabilities such as AI Services, prompt templates, chat memory, streaming, output parsing, tool calling, agents, and retrieval-augmented generation (RAG). The project introduction currently reports integrations with 20+ LLM providers and 30+ embedding stores; those are LangChain4j’s own current documentation claims and may change. It also describes integrations with Spring Boot, Quarkus, Helidon, and Micronaut. See the LangChain4j introduction.

You can work at two levels. Lower-level primitives such as chat models, messages, embeddings, and stores give you more direct control, but you write more orchestration code. AI Services provide a declarative interface implemented by a LangChain4j proxy, handling common input formatting and output parsing and optionally connecting memory, tools, or RAG. For new application code, start with AI Services when a typed method is a natural fit; use the lower-level APIs when you need more explicit control. LangChain4j describes Chains as legacy and says it does not currently plan to add more to them, so they are not the recommended starting point. The AI Services tutorial explains the higher-level approach.

Make the smallest working model call

1. Check the Java version and build setup

The official getting-started guide specifies JDK 17 as the minimum supported version. Confirm the project uses a compatible JDK before adding dependencies. This example uses Maven; if your project uses another build tool, use the matching dependency coordinates and configuration from its documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Add the provider integration

The getting-started example uses the OpenAI integration artifact at version 1.21.0. Add the core module as well if you plan to use AI Services. Treat this version as the version shown in the documentation, not a permanent recommendation; check the current Get Started guide and compatibility details before choosing versions.

<dependencies>
    <dependency>
        <groupId>dev.langchain4j</groupId>
        <artifactId>langchain4j-open-ai</artifactId>
        <version>1.21.0</version>
    </dependency>

    <!-- Needed for the high-level AI Services API -->
    <dependency>
        <groupId>dev.langchain4j</groupId>
        <artifactId>langchain4j</artifactId>
        <version>1.21.0</version>
    </dependency>
</dependencies>

Use compatible versions for the integration and core artifacts. The provider-specific module is not interchangeable with the provider-independent concepts described later: it supplies the implementation that connects LangChain4j’s model abstraction to a particular provider.

3. Configure credentials outside the source code

Set the provider key in the process environment as OPENAI_API_KEY, then read it in Java. Keeping the key out of source files reduces the risk of exposing it publicly. Use your deployment platform’s secret or environment-variable configuration for production rather than committing a real key.

String apiKey = System.getenv("OPENAI_API_KEY");
if (apiKey == null || apiKey.isBlank()) {
    throw new IllegalStateException("Set OPENAI_API_KEY before starting the application");
}

4. Call the chat model directly

The direct API is a useful first check: it confirms that the dependency, credentials, network path, provider configuration, and chosen model can work together before you add application orchestration. The model name below is the documentation example; confirm a currently available model name for your provider account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import dev.langchain4j.model.openai.OpenAiChatModel;

public class LlmExample {
    public static void main(String[] args) {
        String apiKey = System.getenv("OPENAI_API_KEY");
        if (apiKey == null || apiKey.isBlank()) {
            throw new IllegalStateException("Set OPENAI_API_KEY before starting the application");
        }

        OpenAiChatModel model = OpenAiChatModel.builder()
                .apiKey(apiKey)
                .modelName("gpt-4o-mini")
                .build();

        String answer = model.chat("Explain dependency injection in one sentence.");
        System.out.println(answer);
    }
}

The exact builder methods, artifact version, and model identifier are subject to change; compare this example with the current provider integration documentation if compilation or provider access fails. LangChain4j’s chat and language models guide describes the available model abstractions. For new chat features, prefer ChatModel, which works with chat messages. The simpler LanguageModel API is becoming obsolete, and the documentation says new feature support will not be expanded for it. Embedding, image, moderation, and scoring model abstractions are relevant to retrieval, image workflows, moderation, or reranking—not required for a basic text exchange.

Choose between a direct ChatModel and an AI Service

Approach Best fit Trade-off
Direct ChatModel A small feature, a connectivity check, or an application that needs explicit control over messages and model calls. Flexible and close to the model API, but your code owns more formatting, parsing, and orchestration.
AI Service interface A feature that should be exposed to the rest of the application as typed Java methods. Less boilerplate for common input/output handling, but the interface and annotations become part of the application design.

An AI Service is a good next step once a direct call works and the application benefits from a stable method boundary. Define an interface with a meaningful method signature and create its implementation through LangChain4j’s AiServices API. For example, a method that takes a support question and returns a concise answer can keep provider plumbing out of the controller or service that calls it. Consult the current AI Services documentation for the exact imports and construction syntax for the version in use.

AI Services can also be configured to use memory, tools, and retrieval, but those are distinct behaviors with their own product and operational consequences. Add them in response to a concrete requirement rather than treating them as prerequisites for a model call.

Add conversation memory only when turns need context

Conversation history and chat memory solve different problems. History is the complete exchange the application preserves and may show to the user. Chat memory is the context supplied to the model so its next response behaves as if it remembers earlier turns. A memory strategy can evict messages, summarize them, remove details, or inject additional information or instructions. Consequently, a bounded memory window is a model-context policy, not a replacement for a complete user-visible transcript if the product requires one. See the chat memory guide.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a stateless interaction—such as summarizing one submitted text—send only the current request. For a multi-turn assistant, decide how memory is scoped to a conversation, what context is retained, and how old context is handled. Keep durable transcript storage in the application when needed, and configure model memory separately to control what context is sent on each call.

Add tools when the model must trigger application actions

Tool or function calling lets an LLM request a defined operation, such as looking up an order or checking an appointment slot, through application-provided functions. It is useful when a response depends on live or private application state, but it also means the application must define and execute those operations. Treat tool definitions as an application boundary: expose only actions the feature needs, validate inputs, and apply normal authorization and business rules in the code that performs the action. LangChain4j lists tool/function calling among its capabilities in the project introduction; the exact implementation depends on the model integration and API version.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Add private or domain knowledge with RAG

RAG retrieves relevant material from application data and adds it to the model prompt before a response is generated. LangChain4j describes two broad stages: indexing the source material and retrieving relevant passages at answer time. Retrieval can use keyword/full-text search, vector/semantic search, or a hybrid of both. A vector store alone does not guarantee factual answers: results depend on the source material, how it is segmented and indexed, and whether retrieval finds the passages relevant to the question. The RAG tutorial covers the workflow.

Easy RAG for a proof of concept

LangChain4j’s Easy RAG path is intended to reduce setup for a proof of concept. Its quick flow combines document ingestion, an embedding store, and a chat model, with bounded memory as an optional addition. This is a convenient way to test whether retrieval belongs in a feature, but the documentation cautions that the easier setup has lower quality than a tailored RAG configuration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Customize retrieval when quality or control matters

A more deliberate pipeline gives you control over document loading, segmentation, embeddings, storage, retrieval, and reranking. That control matters when source formats, document structure, search relevance, or response quality require tuning. Select and evaluate those stages against the actual data and questions the feature must handle rather than assuming that enabling embeddings is sufficient.

Choose a retrieval style that fits the data

Retrieval style Useful when Documented availability caveat
Vector or semantic Relevant passages may use different words from the user’s query, so semantic similarity helps locate them. Availability depends on the embedding-store integration selected.
Full-text or keyword Exact terms, identifiers, or wording are important to finding a passage. LangChain4j documentation currently says full-text search is supported only by the Azure AI Search and Elasticsearch integrations.
Hybrid A search needs to combine lexical matches with semantic relevance. The documentation currently limits full-text and hybrid search support to Azure AI Search and Elasticsearch integrations.

These integration limits are time-sensitive; verify the current RAG documentation and the chosen store’s capabilities before designing around them.

Consider local inference with Jlama only when it fits the runtime

Jlama is an optional LangChain4j integration for local model use, not the simplest default path. Its documented example requires both a LangChain4j Jlama integration dependency and a native dependency, and Jlama uses Java 21 preview features. That runtime and build configuration make it a distinct deployment choice from calling a hosted provider. The documentation lists compatibility examples for several model architectures, but does not establish a hardware recommendation or performance benchmark. Check the current Jlama integration guide before adopting it.

Practical implementation sequence

  1. Confirm the project uses JDK 17 or later and identify its build tool.
  2. Add the LangChain4j integration module for the provider or local runtime you intend to use; include the core langchain4j module if you plan to use AI Services.
  3. Configure credentials outside source code, using environment or deployment secret configuration.
  4. Make a direct ChatModel request and verify that dependency resolution, credentials, provider access, and the selected model work.
  5. Put the feature behind an AI Service interface if typed methods and reduced formatting/parsing boilerplate suit the application.
  6. Add conversation memory, tools, or RAG only to meet a defined product need, and configure their boundaries separately from the basic model call.
  7. Before upgrading or copying an example, recheck LangChain4j’s current artifact versions, provider/model identifiers, and integration support limits.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.