NAME
Langertha::Role::Chat - Role for APIs with normal chat functionality
VERSION
version 0.503
SYNOPSIS
# Synchronous chat
my $response = $engine->simple_chat('Hello, how are you?');
# Streaming with callback
$engine->simple_chat_stream(sub {
my ($chunk) = @_;
print $chunk->content;
}, 'Tell me a story');
# Streaming with iterator
my $stream = $engine->simple_chat_stream_iterator('Tell me a story');
while (my $chunk = $stream->next) {
print $chunk->content;
}
# Async with Future (traditional style)
my $future = $engine->simple_chat_f('Hello');
my $response = $future->get;
# Async with Future::AsyncAwait (recommended)
use Future::AsyncAwait;
async sub chat_example {
my ($engine) = @_;
my $response = await $engine->simple_chat_f('Hello');
say $response;
}
# Async streaming with real-time callback
async sub stream_example {
my ($engine) = @_;
my ($content, $chunks) = await $engine->simple_chat_stream_realtime_f(
sub { print shift->content },
'Tell me a story'
);
say "\nTotal chunks: ", scalar @$chunks;
}
DESCRIPTION
This role provides chat functionality for LLM engines. It includes both synchronous and asynchronous (Future-based) methods for chat and streaming.
The Future-based _f methods are implemented using Future::AsyncAwait. The HTTP backend is selected by Langertha::Role::AsyncHTTP: an injected _async_http client wins, else Net::Async::HTTP if it can be loaded, else a synchronous LWP::UserAgent fallback (Langertha::Request::SyncHTTP). These async modules are loaded lazily only on the async path, so synchronous-only usage — and the sync fallback — does not require them.
When the sync fallback is used the _f methods still return a Future and keep working, but they run synchronously and sequentially (blocking, no concurrency): the future is already complete when returned, so several _f calls awaited "in parallel" run one after another. Install Net::Async::HTTP + IO::Async (or inject your own _async_http client) for real concurrency.
chat_model
The model name used for chat requests. Lazily defaults to default_chat_model if the engine provides it, otherwise falls back to the general model attribute from Langertha::Role::Models.
chat
my $request = $engine->chat(@messages);
Builds and returns a chat HTTP request object. Messages may be plain strings (treated as user role) or HashRefs with role and content keys. A system prompt from Langertha::Role::SystemPrompt is prepended automatically.
inline_image_fetch_timeout
Seconds each URL image fetch may take on an engine that has to inline images (see "content_format") before the call fails with the engine-named inline-image error. Defaults to 30.
On the _f paths it is enforced on the event loop of the async backend ("async_loop" in Langertha::Role::AsyncHTTP; the error says timed out after 30s), and 0 disables it; on the synchronous LWP fallback, or with an injected client without a loop, the client's own timeout applies instead (for the fallback, the user_agent's). When the synchronous methods build the request, it is the LWP timeout of "ensure_base64" in Langertha::Content::Image (the error then carries LWP's read timeout status), and 0 leaves LWP's own default of 180 seconds, because LWP cannot run without a timeout.
inline_image_max_bytes
The most bytes a URL image fetch may download on an engine that has to inline images (see "content_format"). Defaults to 20971520 (20 MiB), about the largest inline image providers take; 0 removes the cap. Enforced on every backend: a Content-Length over the cap stops the fetch before the body, and a body that grows past it stops the download (Net::Async::HTTP closes the connection, LWP stops reading). The call then fails with the engine-named inline-image error carrying
Langertha::Content::Image image at <url> exceeds inline_image_max_bytes (<n>)
on the synchronous and the _f paths alike. See "ensure_base64" in Langertha::Content::Image.
inline_image_url_filter
inline_image_url_filter => Langertha::Content::Image->deny_private_hosts,
inline_image_url_filter => sub { my ($uri) = @_; $uri->host eq 'img.example.com' },
Optional code reference that decides which image URLs an engine that has to inline images may fetch, against server-side request forgery when image URLs come from untrusted input (a gateway forwarding its clients' messages). It gets a URI object and returns true to allow the fetch. It runs on the image URL before any request and on every redirect hop before that hop is requested; a refusal fails the call with the engine-named inline-image error. Unset by default: every http and https URL is fetched. "deny_private_hosts" in Langertha::Content::Image returns a filter that refuses loopback, link-local, private, carrier-grade NAT and cloud metadata addresses; read its DNS-rebinding caveat. Only the fetches Langertha makes are filtered: engines that pass an image URL through to the provider (OpenAI, Anthropic) leave the fetch to the provider.
chat_messages
my $messages = $engine->chat_messages(@messages);
Normalises @messages into the canonical ArrayRef-of-HashRef format expected by chat_request. Plain strings become { role => 'user', content => $string }. If the engine has a system_prompt set it is prepended as a system message.
simple_chat
my $response = $engine->simple_chat(@messages);
my $response = $engine->simple_chat('Hello, how are you?');
Sends a synchronous chat request and returns the response text. Blocks until the request completes.
chat_stream
my $request = $engine->chat_stream(@messages);
Builds and returns a streaming chat HTTP request object. Croaks if the engine does not implement chat_stream_request. Use "simple_chat_stream" or "simple_chat_stream_iterator" to execute the request.
simple_chat_stream
my $content = $engine->simple_chat_stream($callback, @messages);
$engine->simple_chat_stream(sub {
my ($chunk) = @_;
print $chunk->content;
}, 'Tell me a story');
Sends a synchronous streaming chat request. Calls $callback with each Langertha::Stream::Chunk as it arrives (each chunk may carry incremental thinking, see "thinking" in Langertha::Stream::Chunk). In scalar context returns the complete concatenated content string; in list context returns ($content, $thinking) where $thinking is the aggregated chain-of-thought (undef when the engine surfaced none), as "aggregate_thinking" assembles it. Blocks until the stream completes. total_seconds is logged; for a full breakdown read "execute_streaming_request".
simple_chat_stream_iterator
my $stream = $engine->simple_chat_stream_iterator(@messages);
while (my $chunk = $stream->next) {
print $chunk->content;
}
Returns a Langertha::Stream iterator. The full response is fetched synchronously and buffered; iteration yields each Langertha::Stream::Chunk in order.
_extract_controls
my $controls = $engine->_extract_controls(\%opts);
Removes the canonical per-request controls (karr #46) from %opts and returns them as a HashRef. The engine's chat_request receives the hash under the controls key and places each control on its wire; unknown keys stay in %opts and pass straight through as before.
model_capability_exclusions
sub model_capability_exclusions {
return (
qr/some-family/ => \&_exclude_some_combination, # a model family (regex)
'exact-model-id' => \&_exclude_some_combination, # an exact model id
);
}
The per-model capability-exclusion seam (karr #148), consulted at the chat_f / chat_stream_realtime_f layer above the boolean registry. A boolean capability flag asserts the wire accepts field X; this seam expresses a mutual exclusion between two fields in one request — combining tools with a structured-output response_format — which a flag cannot spell (ADR 0021).
Returns an ordered list of ( $matcher => $rule ) pairs, keyed on chat_model exactly as "model_capability_corrections" in Langertha::Role::Capabilities is. $matcher is an exact model-id string (matched with eq) or a qr// regex (matched against chat_model) — model ids come in families and, through aggregators, carry a provider/ prefix, so a regex catches the routed backend id too. $rule is a coderef (the concrete seam — deliberately not a constraint DSL) invoked as $self->$rule(%request) with has_tools, tool_choice_forced (true for a tool_choice of any/required or a named tool, as the caller passed it), response_format (the one that goes on the wire: the per-request value, else the engine's response_format attribute) and streaming; it croaks when the request hits the combination the model rejects.
Where the rule lives depends on what the constraint belongs to. Groq and Cerebras reject tools alongside a structured-output response_format across every model they serve — a property of the serving stack — so each declares an all-models (qr//) rule by overriding this method. SGLang does the same for a forced tool_choice combined with a response_format. A constraint that belonged to one model would instead be keyed on that model id or family regex, leaving a sibling model on the same engine unaffected. The default is an empty list, so an engine that constrains nothing pays nothing.
simple_chat_f
# Traditional Future style
my $response = $engine->simple_chat_f(@messages)->get;
# With async/await (recommended)
use Future::AsyncAwait;
async sub my_chat {
my $response = await $engine->simple_chat_f(@messages);
return $response;
}
Async version of "simple_chat". Returns a Future that resolves to the response text. The HTTP backend comes from Langertha::Role::AsyncHTTP: Net::Async::HTTP when installed (loaded lazily on first call), otherwise a synchronous LWP::UserAgent fallback under which the call blocks and several _f calls run one after another rather than concurrently.
For requests that need named arguments (tools, tool_choice, response_format, etc.) use "chat_f"; simple_chat_f delegates to it.
chat_f
my $response = await $engine->chat_f(
messages => [ ... ],
tools => [ $tool, ... ],
tool_choice => { type => 'tool', name => 'extract' },
response_format => { ... },
temperature => 0.7,
max_tokens => 512,
# any other engine-specific extras pass straight through
);
Async single-turn chat with named arguments. Returns a Future resolving to a Langertha::Response. The caller is responsible for acting on any tool_calls the engine emits — chat_f does not loop. For the multi-turn MCP tool-calling loop use "chat_with_tools_f" in Langertha::Role::Tools instead.
tools in chat_f can mix Langertha::Tool objects, Langertha::ServerTool objects and HashRefs of any provider shape (OpenAI, Anthropic, MCP, Gemini, provider built-ins). Before the request is built, each item is put into the engine's tool_wire_format in place, keeping the caller's order ("request_list" in Langertha::Tool, the same path "chat_stream_realtime_f" takes): an object goes through its to; a hash already in the wire's shape goes out verbatim, extras such as function.strict and cache_control included; a function-tool hash in another shape (for example an MCP tool with inputSchema) is converted; built-ins and unknown typed items go out verbatim for the provider to judge. On Gemini all function declarations share one functionDeclarations entry. The Responses envelope decides per item itself (Langertha::Role::ResponsesCompatible). A Langertha::ServerTool croaks on an engine that does not supports('server_tools').
On a hermes engine (Langertha::Role::HermesTools) a chat_f call is one turn of "chat_with_tools_f" in Langertha::Role::Tools: the tools go into a leading system message built from hermes_tool_prompt, in MCP shape, and the body carries no tools key. Only function tools can go into the prompt: a built-in or other non-function item croaks there instead of going out verbatim. tool_choice is never sent there: none withholds the tools (no tool prompt; a warning says so), and any value other than auto is ignored with a warning, as the prompt cannot force a tool (on Langertha::Engine::NousResearch a forced named tool takes the json_schema rewrite described below instead). On NousResearch, tools together with a json_schema response_format send both the schema prompt and the tool prompt. <tool_call> blocks in the reply land on "tool_calls" in Langertha::Response and are removed from content; a block that carries no call (no valid JSON object with a name) stays in content as the model wrote it. When calls were lifted and the reply's finish_reason was stop or absent, it reads tool_calls ("raw" in Langertha::Response keeps the provider's value); any other value, such as length, stays.
A tool_choice goes on the wire only as the engine's tool_choice_* capabilities allow. Where the engine does not supports('tool_choice_<kind>') for the choice's kind, auto is dropped silently, none withholds the request's tools instead (with a warning when there were tools to withhold), and a forced choice is dropped with a warning, the model then decides; a forced named tool that the json_schema rewrite below can take is rewritten instead. Likewise parallel_tool_use reaches the wire only where the engine supports('parallel_tool_use'); a value you set elsewhere is dropped with a warning.
These drop warnings (and the temperature drops of Langertha::Role::Temperature) name the line of your own call to chat_f, simple_chat_f, chat_request and the like, not a line inside Langertha; code running inside an event-loop callback gets whatever location Carp finds. A value that comes from an engine attribute is the same on every request, so its drop warns once per engine instance; a value passed with the request warns on every request.
A model passed to chat_f is no control: it replaces the model field of the request body, or, on engines that carry the model in the URL (Langertha::Engine::Gemini, Langertha::Engine::AKI), the model named in the URL, but every model-scoped decision is still taken for the engine's chat_model — the capability picture "supports" in Langertha::Role::Capabilities answers, the model_capability_exclusions rules, the reasoning profile, the temperature gate for reasoning models, a per-model tool_wire_format and reasoning prompt (Langertha::Engine::NousResearch) and per-model body details such as the completion-length key and the default response size. When the override differs from chat_model and one of those decisions that the request uses would come out differently for it, chat_f warns and names the decisions; the request is sent unchanged. For a different model, use an engine whose chat_model is that model.
The canonical per-request controls (karr #46) are normalized like messages/tools instead of being spread as raw target-wire kwargs: temperature, max_tokens, response_format, seed, parallel_tool_use, reasoning_effort, thinking_budget, prompt_cache, prompt_cache_ttl and prompt_cache_key. Each engine's chat_request places them on its own wire (Ollama options, Gemini generationConfig, Anthropic output_config+thinking, ...) via the same value objects the engine attributes use, so the same call is correct across engine families. A per-request control beats the configured engine attribute on a per-key basis. Any other key still passes straight through to the wire as before.
When the caller asks for a forced named tool on an engine that cannot do native named-tool-forcing but supports json_schema response_format (for example Langertha::Engine::Perplexity and Langertha::Engine::NousResearch), the request is automatically rewritten to use the JSON Schema path and the response is loose-parsed; the resulting Langertha::Response exposes the parsed arguments via "tool_call_args" in Langertha::Response with synthetic => 1 on the synthesized tool_call entry.
The rewrite takes the request's response_format. Passing a forced named tool and a response_format other than text in the same chat_f call therefore croaks: the two ask for different output, so pick one. When the response_format only comes from the engine attribute, the forced tool wins for that request and a warning says so. A text response_format is replaced silently.
On a hermes engine that takes response_format (Langertha::Engine::NousResearch), every json_schema response format (the rewritten one, one passed to chat_f or "chat_stream_realtime_f", or the engine's own) also goes into a leading system message built from "hermes_schema_prompt" in Langertha::Role::HermesTools, for a backend that ignores response_format.
simple_chat_stream_f
my ($content, $chunks) = $engine->simple_chat_stream_f(@messages)->get;
Async streaming without a real-time callback. Convenience wrapper around "simple_chat_stream_realtime_f" with undef as the callback. Returns a Future that resolves to ($content, \@chunks, \%timing, $thinking) — the same tuple as "chat_stream_realtime_f"; the trailing elements are additive.
aggregate_tool_calls
my $tool_calls = $engine->aggregate_tool_calls( $chunks );
Walks an ArrayRef of Langertha::Stream::Chunk objects and returns the flat list of Langertha::ToolCall objects collected from any chunks that carry tool_calls, in stream order. Returns an empty ArrayRef if none of the chunks emitted tool calls.
This is the streaming counterpart to "tool_calls" in Langertha::Response: for a streamed response it returns the same calls, as equal Langertha::ToolCall objects, that the non-streaming reply of that response carries. Each dialect's parse_stream_chunk assembles its fragments (Chat-Completions delta.tool_calls per index, Anthropic input_json_delta per content block) in per-stream state and puts every finished call on exactly one chunk, so this helper only collects and never sees a call twice. See "tool_calls" in Langertha::Stream::Chunk for the chunk each dialect uses.
aggregate_thinking
my $thinking = $engine->aggregate_thinking( $chunks );
Walks an ArrayRef of Langertha::Stream::Chunk objects and concatenates the thinking text of every chunk that carries one, in stream order — the way "chat_stream_realtime_f" concatenates content. Returns undef when no chunk carried thinking, so the streamed result mirrors "thinking" in Langertha::Response (also undef when the engine surfaced none).
This is the streaming counterpart to the native thinking that Langertha::Response exposes on the non-streaming path. Each dialect stream parser fills Stream::Chunk->thinking from its own delta spelling (see "thinking" in Langertha::Stream::Chunk); this helper just reassembles the fragments.
aggregate_usage
my $usage = $engine->aggregate_usage( $chunks );
my $counts = Langertha::Usage->from_hash($usage) if $usage;
Walks an ArrayRef of Langertha::Stream::Chunk objects and returns the usage HashRef of the last chunk that carries one, or undef when none did. Streamed usage is cumulative, so that is the whole stream's usage. Use it rather than reading the is_final chunk: an OpenAI-compatible stream requested with stream_options => { include_usage => 1 } reports its usage on a content-less chunk after the final one.
simple_chat_stream_realtime_f
# With async/await (recommended)
use Future::AsyncAwait;
async sub my_stream {
my ($content, $chunks) = await $engine->simple_chat_stream_realtime_f(
sub { print shift->content },
@messages
);
return $content;
}
# Traditional Future style
my $future = $engine->simple_chat_stream_realtime_f($callback, @messages);
my ($content, $chunks) = $future->get;
Async streaming with real-time callback. $callback is called with each Langertha::Stream::Chunk as it arrives from the server (each chunk may carry incremental thinking). Returns a Future that resolves to ($content, \@chunks, \%timing, $thinking), the same tuple as "chat_stream_realtime_f"; the trailing elements are additive, so callers destructuring only ($content, \@chunks) keep working.
This is the recommended method for real-time streaming in async applications. Pass undef as the callback (or use "simple_chat_stream_f") if you only need the final result.
This is a thin wrapper around "chat_stream_realtime_f"; existing callers keep working unchanged. For requests that need named arguments (tools, tool_choice, response_format, temperature, max_tokens, etc.) use "chat_stream_realtime_f" directly.
chat_stream_realtime_f
my ($content, $chunks) = await $engine->chat_stream_realtime_f(
messages => [ ... ],
chunk_callback => sub { print shift->content },
temperature => 0.7,
max_tokens => 512,
# any other engine-specific extras pass straight through
);
Async single-turn streaming chat with named arguments. messages is required (ArrayRef or a single message); chunk_callback is called with each Langertha::Stream::Chunk as it arrives from the server. The canonical per-request controls (karr #46) — temperature, max_tokens, response_format, seed, parallel_tool_use, reasoning_effort, thinking_budget, prompt_cache, prompt_cache_ttl, prompt_cache_key — are extracted and handed to "chat_stream_request" under controls, exactly as in "chat_f". tools is shaped for the engine's tool_wire_format item by item, exactly as in "chat_f" ("request_list" in Langertha::Tool): Langertha::Tool objects are serialized, hashes already in the wire's shape (built-ins and extras included) pass through verbatim, other function-tool hashes (MCP inputSchema, canonical input_schema) are converted, and on Gemini all declarations are merged into one functionDeclarations entry. tool_choice is decided against the tool_choice_* capabilities as in "chat_f" (without the json_schema rewrite); any engine-specific extras pass through. Tool calls the model streams are collected with "aggregate_tool_calls". On a hermes engine the tools ride the system prompt and tool_choice is handled as in "chat_f". The text inside the <tool_call> blocks the model writes ("hermes_call_tag" in Langertha::Role::HermesTools) is not streamed, even when a tag is split across chunks, and chunks that carried only such text are not delivered; the calls land as Langertha::ToolCall objects on the final chunk, as "chat_f" puts them on "tool_calls" in Langertha::Response, and that chunk's finish_reason reads tool_calls where "chat_f"'s would (over stop or none). A block that carries no call is streamed as text where it stood, as "chat_f" keeps it in content, and a call tag inside <think> text is no call. Markup that is unclosed when the stream ends is streamed as text and gives no call. A stream that ends without a final chunk gets a closing chunk for the text still held back and any calls. A per-request model warns as in "chat_f" when it would flip a model-scoped decision.
Returns a Future that resolves to ($content, \@chunks, \%timing, $thinking) where $content is the full concatenated text, \@chunks the collected Langertha::Stream::Chunk objects, \%timing carries ttft_seconds and total_seconds, and $thinking is the aggregated chain-of-thought (undef when the engine surfaced none), assembled from the per-chunk thinking deltas by "aggregate_thinking" so it matches the native "thinking" in Langertha::Response of the non-streaming "chat_f" on the same engine and prompt. The trailing element is additive: callers destructuring only the first three keep working.
If chunk_callback dies, or a stream line cannot be parsed, the returned future fails with that exception on every HTTP backend, and no further chunk reaches chunk_callback. How the rest of the transfer ends depends on the backend (Langertha::Role::AsyncHTTP):
Net::Async::HTTP: the exception never escapes the event loop. The request is cancelled on the next loop iteration (unless the response already ended within the same read), and the engine goes on serving requests. Like any cancelled Net::Async::HTTP request this closes its connection; the client Langertha builds does not pipeline, so requests queued behind it on the same engine wait for a connection of their own and are unaffected.
the synchronous Langertha::Request::SyncHTTP fallback: LWP stops reading at once.
an injected client whose futures have no
loop(or one withoutlater, from another event system): the rest of the body is read and discarded, and the future fails when the response ends. If the transfer then fails at the transport level, the future still fails with the original exception.
Cancelling the returned future cancels the HTTP request.
This is the streaming counterpart to "chat_f". Unlike "chat_f" it does not apply the forced-tool fallback (rewriting a named tool_choice into a response_format on engines without tool_choice_named); synthesizing a tool_calls entry from the accumulated stream text is a separate follow-up concern.
response_format is honored on the streaming path only where the engine has a native wire form (Gemini responseJsonSchema, Ollama format, OpenAI-compatible response_format). Anthropic-family engines have no native form and their synthesized-tool rewrite has no streaming lift, so they consume the key and croak — use "chat_f" for structured output there.
content_format
my $fmt = $engine->content_format;
# 'openai' | 'anthropic' | 'gemini' | 'responses' | 'ollama' | 'lmstudio'
Wire format for multimodal content blocks. Controls how Langertha::Content objects embedded in a message's content arrayref are serialized during "chat_messages". Defaults to 'openai'; overridden by Langertha::Engine::AnthropicBase, Langertha::Engine::Gemini, Langertha::Role::ResponsesCompatible (responses: input_text / input_image parts, output_text on assistant turns), Langertha::Engine::Ollama (ollama: text joined into a string content, images lifted into the message images array) and Langertha::Engine::LMStudio (lmstudio).
A message whose content is a plain string is passed through unchanged on every format; so is an arrayref without any Langertha::Content object, except on gemini (always turned into parts) and ollama (text parts always joined into a string, image_url parts lifted into images).
engine_capabilities
my $caps = $engine->engine_capabilities;
if ( $caps->{tool_choice_named} ) { ... }
Returns a HashRef of capability flags so callers can avoid passing parameters the engine cannot honour.
The base implementation reports only what Langertha::Role::Chat itself provides (chat). Every other capability-bearing role (Langertha::Role::Tools, Langertha::Role::ResponseFormat, Langertha::Role::Streaming, Langertha::Role::Embedding, Langertha::Role::Transcription, Langertha::Role::ImageGeneration, Langertha::Role::HermesTools, Langertha::Role::Temperature, Langertha::Role::Seed, Langertha::Role::ContextSize, Langertha::Role::ResponseSize, Langertha::Role::SystemPrompt, Langertha::Role::ParallelToolUse) hangs its own contribution into this method via around engine_capabilities. Engines override (also via around) when the wire reality differs from the role inventory — for example to clear tool_choice_named on providers that only accept string forms.
Common keys produced by the bundled roles:
chat—simple_chat/simple_chat_fworkstreaming—chat_stream_requestis wired uptools_native— engine accepts atoolsarray on the wiretools_hermes— tools are injected via Hermes-style XML prompt rather than (or in addition to) the native APItool_choice_auto/tool_choice_any/tool_choice_none— which string-formtool_choicevalues are acceptedtool_choice_named—{type => 'tool', name => '...'}forcing works (possibly translated internally — Gemini routes named tools throughallowed_function_names, for example)response_format_json_object—{type => 'json_object'}response_format_json_schema— JSON Schema structured outputembedding,transcription,image_generation— auxiliary capabilities matching the corresponding rolestemperature,seed,context_size,response_size,system_prompt,parallel_tool_use— generation-parameter knobs the engine will honour
Callers should treat the hash as advisory — a missing key means "unknown / unsupported", a true value means "the engine claims it will honour this".
SEE ALSO
Langertha::Role::Langfuse - Observability integration (composed by this role)
Langertha::Role::SystemPrompt - System prompt injection
Langertha::Role::Streaming - Stream parsing (SSE / NDJSON)
Langertha::Role::Tools - Tool calling on top of chat
Langertha::Role::Models - Model selection
Langertha::Stream - Stream iterator
Langertha::Stream::Chunk - Individual stream chunk
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.