Class LangChain4JLLMProvider
- All Implemented Interfaces:
LLMProvider
LLMProvider.
Supports both streaming and non-streaming LangChain4j models. Tool calling is
supported through LangChain4j's Tool annotation.
Streaming vs. non-streaming: The mode is determined by the constructor
used. Pass a StreamingChatModel to
LangChain4JLLMProvider(StreamingChatModel) for streaming, or a
ChatModel to LangChain4JLLMProvider(ChatModel) for
non-streaming. Streaming mode pushes partial responses to the UI as they
arrive, which requires automatic server push or polling to deliver them.
Annotate your UI class or application shell with @Push, or enable
polling with UI.setPollInterval(), before using a streaming model. A
warning is logged at runtime when neither is active.
Blocking the request thread: a ChatModel call blocks the
thread that subscribes to the response, which is the UI thread for a prompt
triggered from the browser. Call setBackgroundExecution(true) to run the call on a background thread instead,
so the request completes and the user's message renders while the LLM works.
Tool call limits: a model that keeps requesting tool calls instead of
answering would never end its turn, and each round costs another model call.
The provider ends such a turn with a ToolCallLimitExceededException
once the model has requested more than 40 calls to any one tool, or
more than 150 tool calls in total, within the turn. Adjust or remove
them with setMaxCallsPerTool(int) and
setMaxTotalToolCalls(int).
Each provider instance maintains its own chat memory. To share conversation history across components, reuse the same provider instance. The memory holds the user messages and the assistant's final answer of each turn; tool calls and their results are sent to the model only within the turn they belong to and are not replayed on later turns. A final answer without text, which a model may give once a tool call has done what was asked, is not stored either, so that the turn leaves only its user message in the memory.
Note: LangChain4JLLMProvider is not serializable. If your application uses session persistence, you will need to create a new provider instance after session restore.
- Since:
- 25.3
- Author:
- Vaadin Ltd
-
Nested Class Summary
Nested classes/interfaces inherited from interface com.vaadin.flow.component.ai.provider.LLMProvider
LLMProvider.LLMRequest, LLMProvider.ToolSpec -
Constructor Summary
ConstructorsConstructorDescriptionLangChain4JLLMProvider(dev.langchain4j.model.chat.ChatModel chatModel) Constructor with a non-streaming chat model.LangChain4JLLMProvider(dev.langchain4j.model.chat.StreamingChatModel chatModel) Constructor with a streaming chat model. -
Method Summary
Modifier and TypeMethodDescriptionintGets the maximum number of times the model may call any one tool during a turn.intGets the maximum number of tool calls the model may make during a turn, all tools together.booleanGets whether the LLM call runs on a background thread.voidsetBackgroundExecution(boolean backgroundExecution) Sets whether to run the LLM call on a background thread.voidsetHistory(List<ChatMessage> history, Map<String, List<AIAttachment>> attachmentsByMessageId) Restores the provider's conversation memory from a list of chat messages with their associated attachments.voidsetMaxCallsPerTool(int maxCallsPerTool) Sets the maximum number of times the model may call any one tool during a turn.voidsetMaxTotalToolCalls(int maxTotalToolCalls) Sets the maximum number of tool calls the model may make during a turn, all tools together.reactor.core.publisher.Flux<String> stream(LLMProvider.LLMRequest request) Streams a response from the LLM based on the provided request.
-
Constructor Details
-
LangChain4JLLMProvider
public LangChain4JLLMProvider(dev.langchain4j.model.chat.StreamingChatModel chatModel) Constructor with a streaming chat model.- Parameters:
chatModel- the streaming chat model, notnull- Throws:
NullPointerException- if chatModel isnull
-
LangChain4JLLMProvider
public LangChain4JLLMProvider(dev.langchain4j.model.chat.ChatModel chatModel) Constructor with a non-streaming chat model.- Parameters:
chatModel- the non-streaming chat model, notnull- Throws:
NullPointerException- if chatModel isnull
-
-
Method Details
-
stream
Description copied from interface:LLMProviderStreams a response from the LLM based on the provided request. This method returns a reactive stream that emits response tokens as they become available from the LLM. The provider manages conversation history internally, so each call to this method adds to the ongoing conversation context.Threading: this method is called on the thread that triggers the prompt, and the returned stream is subscribed to on that same thread. An implementation whose LLM call blocks must therefore schedule that call itself — for example with
subscribeOn(Schedulers.boundedElastic())— otherwise it occupies the UI thread and holds the session lock for the whole turn, and nothing the turn produces reaches the browser until it ends. The built-in providers expose this as asetBackgroundExecution(boolean)setting, since running the turn on the request thread is the simpler default when a turn is short.Callers only consume the stream and do not schedule it, so whether a turn runs in the background is decided entirely by the implementation.
- Specified by:
streamin interfaceLLMProvider- Parameters:
request- the LLM request containing user message, system prompt, attachments, and tools, notnull- Returns:
- a Flux stream that emits response tokens as strings, never
null
-
isBackgroundExecution
public boolean isBackgroundExecution()Gets whether the LLM call runs on a background thread.- Returns:
trueif the call runs on a background thread,falseif it runs on the thread that asks for the response
-
setBackgroundExecution
public void setBackgroundExecution(boolean backgroundExecution) Sets whether to run the LLM call on a background thread. The default isfalse, which runs it on the thread that asks for the response — the UI thread, for a prompt triggered from the browser. The setting has no effect with aStreamingChatModel, whose response already arrives on the LLM client's own threads.A
ChatModelcall blocks for the whole turn, every tool call included. On the UI thread that means holding the session lock until the turn ends, so nothing the turn produces reaches the browser and the application appears frozen. Set this totrueto run the call on a background thread instead: the request completes immediately, the user's message and the typing indicator render, and the response is added when it arrives.This requires the following from the application:
- A way to deliver the response. Annotate the application shell
or UI class with
@Push, or enable polling withUI.setPollInterval(int). Manual push mode is not enough on its own, because nothing callsui.push()for you. A warning is logged when neither is active. - Thread-safe tools. On a background thread Vaadin thread locals
such as
UI.getCurrent()and framework contexts such as Spring Security'sSecurityContextare not bound, and UI components must not be accessed directly. Wrap component access inui.access(), or capture what you need inAIController.onRequest(com.vaadin.flow.component.ai.orchestrator.RequestListener.RequestEvent), which still runs on the UI thread. This is the same requirement aStreamingChatModelalready has.
The orchestrator processes one prompt at a time. Without background execution, a message submitted while a turn is running waits for the session lock and is processed when the turn ends; with it, the submit is rejected and dropped with a warning — and a connected input has already cleared its text.
Like the streaming mode, the setting is not preserved when the session is serialized: an application that restores sessions must re-apply it when it recreates the provider.
- Parameters:
backgroundExecution-trueto run the call on a background thread,falseto run it on the thread that asks for the response
- A way to deliver the response. Annotate the application shell
or UI class with
-
getMaxCallsPerTool
public int getMaxCallsPerTool()Gets the maximum number of times the model may call any one tool during a turn.- Returns:
- the maximum number of calls per tool, or
0if there is no limit - Since:
- 25.4
-
setMaxCallsPerTool
public void setMaxCallsPerTool(int maxCallsPerTool) Sets the maximum number of times the model may call any one tool during a turn. The default is40.A model that keeps calling the same tool instead of answering, for example to look for something the application does not have, would otherwise never end its turn. Once a call would exceed the limit, the turn fails with a
ToolCallLimitExceededExceptionthat names the tool: none of the tool calls of that round are executed, the model is not called again, and you receive the exception as the error of the turn. The conversation stays usable: the failed turn leaves only its prompt in the chat memory, and the next prompt continues from there.The limit applies to each tool separately. See
setMaxTotalToolCalls(int)for the limit on all tool calls of a turn together. The value is read when a turn starts, so a change applies from the next prompt on.- Parameters:
maxCallsPerTool- the maximum number of calls per tool during a turn, or0to remove the limit- Throws:
IllegalArgumentException- if the value is negative- Since:
- 25.4
-
getMaxTotalToolCalls
public int getMaxTotalToolCalls()Gets the maximum number of tool calls the model may make during a turn, all tools together.- Returns:
- the maximum number of tool calls per turn, or
0if there is no limit - Since:
- 25.4
-
setMaxTotalToolCalls
public void setMaxTotalToolCalls(int maxTotalToolCalls) Sets the maximum number of tool calls the model may make during a turn, all tools together. The default is150.This bounds a turn that
setMaxCallsPerTool(int)does not catch because the model spreads its calls over several tools. Once a call would exceed the limit, the turn fails with aToolCallLimitExceededException: none of the tool calls of that round are executed, the model is not called again, and you receive the exception as the error of the turn. The conversation stays usable: the failed turn leaves only its prompt in the chat memory, and the next prompt continues from there.The value is read when a turn starts, so a change applies from the next prompt on.
- Parameters:
maxTotalToolCalls- the maximum number of tool calls during a turn, or0to remove the limit- Throws:
IllegalArgumentException- if the value is negative- Since:
- 25.4
-
setHistory
public void setHistory(List<ChatMessage> history, Map<String, List<AIAttachment>> attachmentsByMessageId) Description copied from interface:LLMProviderRestores the provider's conversation memory from a list of chat messages with their associated attachments. Any existing memory is cleared before the new history is applied.Providers that support setting chat history should override this method.
This method must not be called while a streaming response is in progress.
- Specified by:
setHistoryin interfaceLLMProvider- Parameters:
history- the list of chat messages to restore, notnullattachmentsByMessageId- a map fromChatMessage.messageId()to the list of attachments for that message, notnull
-