Recommended Free Tools
LangChain4j gives Java developers both direct building blocks for working with large language models (LLMs) and higher-level APIs that reduce orchestration code. Start with its chat API to understand what the model receives and returns, add AI Services when you want a cleaner application interface, then introduce tools, memory, and retrieval-augmented generation (RAG) as your application needs them. The official documentation lists Java 17 as the minimum supported JDK; check its live setup guide for dependency versions before creating a project.
Set up a Java project with compatible dependencies
The official LangChain4j setup guide states, “The minimum supported JDK version is 17.” It provides framework-specific setup guidance for Quarkus, Spring Boot, and Helidon; the broader overview also lists Micronaut. Follow the guide for your framework rather than mixing integration examples, and use the matching Maven or Gradle coordinates shown there: LangChain4j Get Started documentation.
LangChain4j is modular: model-provider and vector-store integrations are separate dependencies. The main langchain4j dependency is needed for high-level AI Services. The overview describes integrations for 20+ LLM providers, 30+ embedding stores, 20+ embedding models, 5+ chat memory stores, 5+ image generation models, and 5+ scoring models. These are counts stated by the documentation, not independently validated totals, and can change; check the live overview when selecting an integration: LangChain4j overview.
The Get Started page displays version 1.20.2 for its example modules. Treat that as the version shown on that page, not as a timeless recommendation: select mutually compatible core and integration versions from the current setup instructions.
Make a direct chat call with ChatModel
For a first model interaction, use the low-level chat API. A chat model accepts chat messages and returns an AI message, leaving your Java code in charge of composing the request and handling the response. This makes it a useful starting point when you need to see the boundary between application logic and model behavior.
New code and instruction should focus on the chat API. The documentation says the older LanguageModel API will no longer be expanded. The chat API and its message model are documented in the chat and language models tutorial.
Use AI Services when you want less orchestration code
AI Services are a higher-level abstraction, not a separate model provider. They let you describe an application-facing interface while LangChain4j coordinates model calls with components such as prompts, memory, parsers, tools, and RAG. They are useful when the direct chat API’s explicit composition starts adding repetitive glue code.
Rank #2
Choose between the two approaches based on the control your application needs:
Free tools Windows power users keep installed
One-click scans. No signup required.
| Approach | What it gives you | Good fit |
|---|---|---|
ChatModel |
Direct control over chat messages, calls, and response handling. | Learning the request-response flow or implementing custom orchestration. |
| AI Services | A declarative Java interface that coordinates model calls and connected components. | Reducing boilerplate in an application that combines prompts, memory, tools, parsers, or RAG. |
See the AI Services documentation for the supported interface patterns and configuration.
Add conversational memory and application tools
Memory keeps relevant conversational context
Memory manages conversational context across interactions. Decide what context should be available to a conversation and how it fits the application; do not assume that a model call automatically knows earlier turns. LangChain4j documents memory as a component that can be combined with its model and AI Service APIs.
Tools let the model request application functions
A tool is a function your application makes available to the model. The model can request a tool call, but application code—not the model—executes the function and sends its result back into the conversation. For example, an application might expose a function to look up an order status; it must still perform the lookup and decide how to handle the result.
Tool support and the model’s ability to select the right tool vary by model. Keep authorization, input validation, and side-effect decisions in application code rather than treating a model’s request as permission to perform an action. See the tools tutorial for LangChain4j’s tool-call flow.
Build toward RAG in two stages
Retrieval-augmented generation supplies relevant pieces of domain-specific or proprietary material as context for a model response. It does not make the model’s underlying knowledge current or guarantee that a response is correct; the application retrieves material and includes it in the prompt.
Rank #4
Index documents before answering questions
Indexing prepares source documents for later retrieval. In a typical pipeline, the application loads documents, splits them into segments, creates embeddings, and stores the resulting data. The precise choices depend on the documents and the retrieval system.
Retrieve relevant material for a query
At question time, retrieval finds candidate passages to supply to the model. The LangChain4j RAG tutorial discusses keyword or full-text search, vector search, and hybrid approaches that combine retrieval methods. The documentation describes full-text and hybrid support as limited to its Azure AI Search and Elasticsearch integrations; because integration support can change, check the current RAG tutorial before basing an implementation on that limitation.
Use Easy RAG for a first proof of concept
LangChain4j’s Easy RAG path reduces setup work by applying defaults to document loading, splitting, embeddings, and storage. The tutorial positions it as a way to learn or build a proof of concept, and warns that its quality is lower than a tailored RAG setup. Begin there when speed and simplicity matter, then customize ingestion and retrieval as your needs become clearer.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
The tutorial’s described defaults are segments of up to 300 tokens with 30-token overlap and the bge-small-en-v1.5 embedding model. These are implementation details from the documentation, not universal settings; verify the live tutorial before relying on them. It also says the default embedding can run offline in the same JVM process using ONNX Runtime. That applies to embedding generation on this documented route, not necessarily to chat inference or every part of an application: assess separately where the chat model and vector store run.
Choose a path that fits your constraints
| Decision | Choose this when… | Trade-off to consider |
|---|---|---|
| Direct chat API or AI Services | You need custom request composition and response handling, or prefer a declarative interface that reduces orchestration. | Direct calls expose more control; AI Services coordinate more of the flow for you. |
| Easy RAG or tailored RAG | You want a quick learning project or proof of concept, or need greater control over ingestion and retrieval. | Easy RAG offers defaults and less setup; a tailored pipeline takes more design work. |
| Provider, store, and framework integration | You have selected a model provider, vector store, and Java framework. | Confirm the relevant LangChain4j integration exists and use compatible dependencies for the combination. |
| Local embeddings or remote services | You want the documented Easy RAG embedding route to run locally, or have other deployment requirements. | Local embedding generation does not establish where chat inference or vector storage runs; evaluate each separately. |
Keep experimental agentic functionality separate from core APIs
The langchain4j-agentic module is marked experimental in the official documentation and may change. That maturity status makes it a different kind of dependency decision from building around the documented chat, AI Services, memory, tools, and RAG abstractions. Check the agentic documentation for its current status before adopting it in an application that depends on stable behavior.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

