NAME
Langertha::Stream::Chunk - Represents a single chunk from a streaming response
VERSION
version 0.503
SYNOPSIS
my $stream = $engine->simple_chat_stream_iterator('Tell me a story');
while (my $chunk = $stream->next) {
print $chunk->content;
if ($chunk->is_final) {
say "\nModel: ", $chunk->model if $chunk->has_model;
say "Finish: ", $chunk->finish_reason if $chunk->has_finish_reason;
}
}
DESCRIPTION
A single text chunk delivered during a streaming LLM response. Each chunk carries incremental content text and optional metadata. Chunks are collected into a Langertha::Stream iterator by "simple_chat_stream_iterator" in Langertha::Role::Chat.
content
The incremental text content delivered in this chunk. Required. For most chunks this is a word or partial word; the final chunk may be an empty string.
raw
The raw parsed API response data for this chunk as a HashRef. Use has_raw to check whether it was provided.
is_final
Boolean flag set to 1 on the last chunk of a stream. Defaults to 0.
model
The model identifier returned by the provider, if present. Use has_model to check availability.
finish_reason
The reason the stream ended: stop, length, tool_calls, etc. Provider-specific values are preserved as-is. undef on non-final chunks. Use has_finish_reason to check availability.
tool_calls
Optional ArrayRef of finished Langertha::ToolCall objects that complete on this chunk. Every call the model streams lands on exactly one chunk, as the same object the non-streaming reply of that response carries on "tool_calls" in Langertha::Response; a chunk never holds a fragment. Where it lands depends on the dialect:
Chat-Completions (Langertha::Role::OpenAICompatible): the
delta.tool_callsfragments are assembled perindex, and all calls land on the chunk that carriesfinish_reason.Anthropic Messages (Langertha::Role::AnthropicCompatible): each
tool_useblock is assembled from itsinput_json_deltafragments and lands on the chunk for itscontent_block_stop, before the final chunk.Gemini and Ollama native: calls arrive whole and land on the chunk that carries them.
Open-Responses (Langertha::Role::ResponsesCompatible): the calls land on the final chunk, read from the terminal
response.completedevent.Hermes (Langertha::Role::HermesTools, tools in the prompt): the
<tool_call>blocks are withheld fromcontentand their calls land on the final chunk (see "chat_stream_realtime_f" in Langertha::Role::Chat).
Most chunks have no tool calls — use has_tool_calls to check, or "aggregate_tool_calls" in Langertha::Role::Chat to collect them all in stream order.
citations
Optional ArrayRef of search-augmented source citations, populated on the final chunk when a search-augmented engine emits them mid-stream. The Open-Responses envelope (Langertha::Engine::Perplexity) lifts the search_results block out of the terminal response.completed output[] here, so a streamed reply surfaces the same sources the non-streaming path exposes as "citations" in Langertha::Response. Most chunks carry none — use has_citations to check. "citations" in Langertha::Stream reassembles them off the stream.
thinking
Optional incremental chain-of-thought / reasoning text delivered in this chunk, parallel to "content" and "tool_calls". Populated by the dialect stream parsers from their verified per-provider delta spelling — the OpenAI-compatible delta.reasoning_content / bare delta.reasoning, Anthropic's thinking_delta, Gemini's thought parts, and Ollama native message.thinking. Most chunks carry no thinking — use has_thinking to check. The full streamed thinking is reassembled by "aggregate_thinking" in Langertha::Role::Chat, the streaming counterpart of "thinking" in Langertha::Response on the non-streaming path.
refusal
Optional fragment of a refusal delivered in this chunk: the OpenAI-compatible delta.refusal, and the whole refusal of a Responses API stream on its final chunk. Concatenated in order, the fragments are the text "refusal" in Langertha::Response carries on the non-streaming path. Use has_refusal to check.
usage
Token usage counts as a HashRef, if provided by the engine on the final chunk. Keys vary by provider. Use has_usage to check availability. An OpenAI-compatible stream requested with include_usage reports it on a content-less chunk after the final one; "aggregate_usage" in Langertha::Role::Chat finds it either way.
cached_tokens
Number of prompt tokens served from the prefix cache, if reported by the provider on the final chunk. Populated from usage.prompt_tokens_details.cached_tokens on the OpenAI-compatible wire (SGLang with return_cached_tokens_details enabled, and other servers that emit the detail block) and from usage.input_tokens_details.cached_tokens on the Open-Responses wire (OpenAI Responses / Perplexity Agent). undef when the provider does not report it. Use has_cached_tokens to check availability.
SEE ALSO
Langertha::Stream - Iterator that holds chunks
Langertha::Response - Non-streaming response object
Langertha::Role::Chat - Chat role that produces streams
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.