Class SpringAILLMProvider
- All Implemented Interfaces:
LLMProvider
LLMProvider.
Supports both streaming and non-streaming Spring AI models. Tool calling is
supported through Spring AI's Tool annotation.
Streaming vs. non-streaming: Streaming is enabled by default. To
disable it, call setStreaming(false).
Streaming mode pushes partial responses to the UI as they arrive, which
requires automatic server push or polling to deliver them. Annotate your UI
class or application shell with @Push, or enable polling with
UI.setPollInterval(), before using streaming mode. A warning is
logged at runtime when neither is active.
Blocking the request thread: in non-streaming mode the LLM call blocks
the thread that subscribes to the response, which is the UI thread for a
prompt triggered from the browser. Call
setBackgroundExecution(true) to run
the call on a background thread instead, so the request completes and the
user's message renders while the LLM works.
Tool call limits: Spring AI runs the tool-calling loop itself and
bounds it: once the model has requested more than 40 calls to any one
tool, or more than 150 tool calls in total, within a turn, Spring AI
stops the loop and replies with its own message about the exceeded limit. The
provider does not pass that reply on. It fails the turn with a
ToolCallLimitExceededException instead, the same way
LangChain4JLLMProvider does, so you receive the exception as the
error of the turn whichever provider runs it. The finish reason
toolCallLimitExceeded is still published in the
response metadata, and Spring AI's reply may remain
in the chat memory with either constructor, since the provider does not
rewrite what Spring AI's advisors stored. The limits belong to the
ToolCallingAdvisor of the ChatClient: a provider created from
a ChatModel builds its own client and keeps Spring AI's defaults. To
change them, build the client yourself, passing ChatClient.builder a
ToolCallingAdvisor.Builder that carries a
DefaultToolCallingManager with your limits, and create the provider
from that client. Its maxCallsPerTool and maxTotalToolCalls
set a limit, unlimitedCallsPerTool() and
unlimitedTotalToolCalls() remove one. In a Spring Boot application
the spring.ai.tools.limits properties configure the same limits on
the auto-configured ChatClient.Builder, so passing that client to
SpringAILLMProvider(ChatClient) needs no builder code.
With the SpringAILLMProvider(ChatModel) constructor the provider
maintains its own chat memory, and setHistory(List, Map) restores a
saved conversation into it. To share conversation history across components,
reuse the same provider instance. With the
SpringAILLMProvider(ChatClient) constructor the application owns the
chat memory, so giving the LLM its context is up to the application and
setHistory(List, Map) does nothing. Restoring a conversation through
AIOrchestrator.Builder.withHistory(List, Map) still matters on that
path: the message list the user sees and the orchestrator's own conversation
history are restored by the orchestrator, not by the provider.
Note: SpringAILLMProvider is not serializable. If your application uses session persistence, you will need to create a new provider instance after session restore.
- Since:
- 25.3
- Author:
- Vaadin Ltd
-
Nested Class Summary
Nested classes/interfaces inherited from interface com.vaadin.flow.component.ai.provider.LLMProvider
LLMProvider.LLMRequest, LLMProvider.ToolSpec -
Constructor Summary
ConstructorsConstructorDescriptionSpringAILLMProvider(org.springframework.ai.chat.client.ChatClient chatClient) Constructor with a chat client.SpringAILLMProvider(org.springframework.ai.chat.model.ChatModel chatModel) Constructor with a chat model. -
Method Summary
Modifier and TypeMethodDescriptionbooleanGets whether the LLM call runs on a background thread.booleanGets whether streaming mode is used.voidsetBackgroundExecution(boolean backgroundExecution) Sets whether to run the LLM call on a background thread.voidsetHistory(List<ChatMessage> history, Map<String, List<AIAttachment>> attachmentsByMessageId) Restores the provider's conversation memory from a list of chat messages with their associated attachments.voidsetStreaming(boolean streaming) Sets whether to use streaming mode.reactor.core.publisher.Flux<String> stream(LLMProvider.LLMRequest request) Streams a response from the LLM based on the provided request.
-
Constructor Details
-
SpringAILLMProvider
public SpringAILLMProvider(org.springframework.ai.chat.model.ChatModel chatModel) Constructor with a chat model.- Parameters:
chatModel- the chat model, notnull- Throws:
NullPointerException- if chatModel isnull
-
SpringAILLMProvider
public SpringAILLMProvider(org.springframework.ai.chat.client.ChatClient chatClient) Constructor with a chat client. Conversation memory must be configured on theChatClientitself, for example with aMessageChatMemoryAdvisorand a defaultChatMemory.CONVERSATION_IDadvisor parameter.The application owns that memory, so
setHistory(List, Map)does nothing on a provider created this way. A conversation loaded from external storage must be written into theChatMemorybefore the client is passed here. Passing the same conversation toAIOrchestrator.Builder.withHistory(List, Map)still restores the message list and the orchestrator's own history snapshot.- Parameters:
chatClient- the chat client, notnull- Throws:
NullPointerException- if chatClient isnull
-
-
Method Details
-
stream
Description copied from interface:LLMProviderStreams a response from the LLM based on the provided request. This method returns a reactive stream that emits response tokens as they become available from the LLM. The provider manages conversation history internally, so each call to this method adds to the ongoing conversation context.Threading: this method is called on the thread that triggers the prompt, and the returned stream is subscribed to on that same thread. An implementation whose LLM call blocks must therefore schedule that call itself — for example with
subscribeOn(Schedulers.boundedElastic())— otherwise it occupies the UI thread and holds the session lock for the whole turn, and nothing the turn produces reaches the browser until it ends. The built-in providers expose this as asetBackgroundExecution(boolean)setting, since running the turn on the request thread is the simpler default when a turn is short.Callers only consume the stream and do not schedule it, so whether a turn runs in the background is decided entirely by the implementation.
- Specified by:
streamin interfaceLLMProvider- Parameters:
request- the LLM request containing user message, system prompt, attachments, and tools, notnull- Returns:
- a Flux stream that emits response tokens as strings, never
null
-
isStreaming
public boolean isStreaming()Gets whether streaming mode is used.- Returns:
trueif streaming mode is used,falseotherwise
-
setStreaming
public void setStreaming(boolean streaming) Sets whether to use streaming mode. The default istrue.- Parameters:
streaming-trueto use streaming mode,falsefor non-streaming.
-
isBackgroundExecution
public boolean isBackgroundExecution()Gets whether the LLM call runs on a background thread.- Returns:
trueif the call runs on a background thread,falseif it runs on the thread that asks for the response
-
setBackgroundExecution
public void setBackgroundExecution(boolean backgroundExecution) Sets whether to run the LLM call on a background thread. The default isfalse, which runs it on the thread that asks for the response — the UI thread, for a prompt triggered from the browser. The setting has no effect in streaming mode, where the response already arrives on the LLM client's own threads.A non-streaming call blocks for the whole turn, every tool call included. On the UI thread that means holding the session lock until the turn ends, so nothing the turn produces reaches the browser and the application appears frozen. Set this to
trueto run the call on a background thread instead: the request completes immediately, the user's message and the typing indicator render, and the response is added when it arrives.This requires the following from the application:
- A way to deliver the response. Annotate the application shell
or UI class with
@Push, or enable polling withUI.setPollInterval(int). Manual push mode is not enough on its own, because nothing callsui.push()for you. A warning is logged when neither is active. - Thread-safe tools. On a background thread Vaadin thread locals
such as
UI.getCurrent()and framework contexts such as Spring Security'sSecurityContextare not bound, and UI components must not be accessed directly. Wrap component access inui.access(), or capture what you need inAIController.onRequest(com.vaadin.flow.component.ai.orchestrator.RequestListener.RequestEvent), which still runs on the UI thread. This is the same requirement streaming mode already has.
The orchestrator processes one prompt at a time. Without background execution, a message submitted while a turn is running waits for the session lock and is processed when the turn ends; with it, the submit is rejected and dropped with a warning — and a connected input has already cleared its text.
Like the streaming mode, the setting is not preserved when the session is serialized: an application that restores sessions must re-apply it when it recreates the provider.
- Parameters:
backgroundExecution-trueto run the call on a background thread,falseto run it on the thread that asks for the response
- A way to deliver the response. Annotate the application shell
or UI class with
-
setHistory
public void setHistory(List<ChatMessage> history, Map<String, List<AIAttachment>> attachmentsByMessageId) Restores the provider's conversation memory from a list of chat messages with their associated attachments. Any existing memory is cleared before the new history is applied.Providers that support setting chat history should override this method.
This method must not be called while a streaming response is in progress.
Restores the conversation into the provider's own chat memory. Does nothing when the provider was created with the
ChatClientconstructor, because the application owns the chat memory in that case and is expected to have populated it before passing the client in. A warning is logged when that client is missing the chat memory configuration the conversation would need. Whether the memory actually holds the conversation is not visible to the provider, so a client configured correctly but never loaded is indistinguishable from one that was. Doing nothing here does not reduce what the caller restores: an orchestrator rebuilds the message list and its own conversation history itself.- Specified by:
setHistoryin interfaceLLMProvider- Parameters:
history- the list of chat messages to restore, notnullattachmentsByMessageId- a map fromChatMessage.messageId()to the list of attachments for that message, notnull
-