NAME
Langertha::Response - LLM response with metadata
VERSION
version 0.503
SYNOPSIS
my $response = $engine->simple_chat('Hello');
# Stringifies to content (backward compatible)
print $response;
print "Response: $response\n";
# Access metadata
say $response->model;
say $response->id;
say $response->finish_reason;
# Token usage
say "Prompt tokens: ", $response->prompt_tokens;
say "Completion tokens: ", $response->completion_tokens;
say "Total tokens: ", $response->total_tokens;
# Full raw response
use Data::Dumper;
print Dumper($response->raw);
DESCRIPTION
Wraps LLM response text content together with all available metadata from the API response. Uses overload for string context so existing code treating responses as plain strings continues to work.
Boolean context — a Response is always true
Langertha::Response overloads bool to a constant true: the object exists, so it is true, whatever "content" happens to hold. This is not cosmetic. With only the "" overload and fallback => 1, Perl derives boolean context from stringification, so a Response whose content is the empty string was false — and that is exactly the shape of a tool-call-only reply (an Anthropic turn that is pure tool_use, an Ollama reply with empty message.content plus a tool_calls array), as well as of a model that legitimately answers "0". Callers writing
my $resp = $engine->simple_chat(@messages);
if ($resp) { ... } # entered now; silently skipped before
$resp or die "no response"; # no longer dies on a perfectly good Response
used to drop precisely the responses that carry "tool_calls". Decided in karr #100.
Asking whether the content is empty is a different question, and keeps its own spelling:
if ( length "$resp" ) { ... } # the model produced text
if ( $resp->has_tool_calls ) { ... } # the model emitted tool calls
The "" overload is unchanged — "$response" is still "content", and eq, ne and concatenation keep routing through it. The change also makes this class agree with the rest of the distribution: Langertha::ToolCall, Langertha::Usage and Langertha::Stream::Chunk are already true-because-they-exist. Langertha::Result got the same treatment for the same reason.
TO_JSON — the canonical, bounded representation
Langertha::Response carries a TO_JSON (delegating to "to_hash") so that every JSON::MaybeXS backend encodes a bare Response identically when convert_blessed is enabled. Decided in karr #50, after karr #43 fixed the permanent shape of the class ("usage" is now a Langertha::Usage object with its own TO_JSON).
The shape is deliberately bounded: "content" plus the metadata fields that are present ("id", "model", "finish_reason", "usage", "timing", "created", "thinking", "refusal", "rate_limit", "tool_calls"). "raw" and "probes" are not included — "raw" is the entire provider payload (duplicating every other field, including the echoed prompt) and "probes" can hold megabytes of tensor data. A TO_JSON that kept them would blow up every trace it touched; one that dropped them keeps the object's JSON form useful without silently discarding the fields a consumer is likely to want.
The string overload is unchanged: "$response" still returns "content", and the TO_JSON does not replace it.
Backend uniformity is the point. Before this method existed, what an encoder with convert_blessed did with a bare Response depended on the JSON::MaybeXS backend: Cpanel::JSON::XS fell back to the "" overload and emitted the content string, while JSON::XS and JSON::PP threw. Consumers on Cpanel silently got {"response":"hello"}, consumers on JSON::XS/JSON::PP got an exception — the same application behaved differently on two machines. With TO_JSON all three emit the canonical hash. The Cpanel behavior was never a contract (it is gone now); callers that want the content string should stringify explicitly.
content
The text content of the response. Required.
raw
The full parsed API response as a HashRef.
id
Provider-specific response ID.
model
The actual model used for the response.
finish_reason
Why the response ended: stop, end_turn, length, tool_calls, etc. Provider-specific values are preserved as-is.
usage
Token usage as a Langertha::Usage object. For backward compatibility BUILDARGS upgrades a plain HashRef (the raw provider usage block) into a Langertha::Usage automatically; new code can construct the object directly. The object overloads %{}, so existing callers that dereference $response->usage->{prompt_tokens} keep working — the overload returns the provider-verbatim hash (see "HASH OVERLOAD" in Langertha::Usage).
timing
Timing information as a HashRef. Holds client-measured ttft_seconds and total_seconds (Float, seconds) on every engine, plus provider-native stage durations that engines may populate engine-specifically — e.g. Ollama's total_seconds, load_seconds, prompt_eval_seconds, eval_seconds (all Float, seconds), with its original *_duration keys in nanoseconds preserved for backward compatibility. See "ttft_seconds" and "total_seconds" for the standard accessors.
Merge policy: first-write-wins. When both a provider-supplied duration (e.g. Ollama's total_seconds derived from total_duration) and a client-measured duration (Langertha::Role::Chat wrapper) are available for the same key, the provider value wins. See "_merge_timing_field" in Langertha::Role::Chat and ADR 0011 in docs/adr/0011-response-timing-seam.md for the rationale — server-side duration excludes network jitter and is the better signal for model-latency observability. Callers that need the client wall-clock (RTT, queue, TLS handshake, network back) should read $response->timing directly: client-measured deltas are not written under ttft_seconds / total_seconds when the provider already populated those keys, but Langertha::Role::Chat does not delete other timing entries on conflict.
ttft_seconds
my $ttft = $response->ttft_seconds; # Float, undef when unmeasured
if ($response->has_ttft) { ... }
Returns time-to-first-token in seconds (Float), measured client-side between request send and the first streamed chunk. undef for sync (non-streaming) calls and for engines that did not record the metric. Use has_ttft to test availability without warnings.
total_seconds
my $total = $response->total_seconds; # Float, undef when unmeasured
if ($response->has_total) { ... }
Returns end-to-end request time in seconds (Float). On engines that populate the field from a provider-native metric (currently Ollama, from total_duration), this is the server-reported duration — time the model actually spent generating, excluding network jitter. On engines where only the client wrapper measured, it is the wall-clock duration from before user_agent->request (sync) or do_request (async) to after the response body was fully consumed.
The two are not interchangeable: server-time answers "how fast is the model", wall-clock answers "how long did my call take end-to-end". Callers that need the wall-clock RTT should read $response->timing directly (the client wrapper writes its measurement under a sibling key when the provider has not claimed total_seconds; when both are present the provider value wins — see "timing").
undef when the engine did not record the metric. Use has_total to test availability without warnings.
created
The instant the provider says the response was created, as a Langertha::Moment — a Time::Moment subclass, so the full civil time is there: sub-seconds, UTC offset, and every Time::Moment accessor.
In numeric context it is the Unix timestamp it always was. The 0+ overload yields whole seconds since the epoch, so 0 + $response->created, $response->created > $cutoff and int($response->created) keep returning exactly what they returned when this attribute was a Maybe[Int]. In string context it is the full ISO-8601 stamp instead — which is where the sub-seconds are, and which is not what the old Int stringified to. See "Comparing and printing created" below.
Engine-agnostic by contract, and normalized here rather than per engine: BUILDARGS runs every incoming value through "from_wire" in Langertha::Moment, which takes the OpenAI-compatible wire's epoch integer and Ollama's RFC3339 created_at string alike. A value it cannot read — including the 0001-01-01T00:00:00Z zero-value sentinel Go-based servers emit — drops the field instead of failing the constructor: has_created is then false and the response is built regardless. A timestamp is metadata and must never take a whole reply down with it (GitHub issue #3, karr #92 / #117).
The provider's native form always stays available verbatim under "raw" (raw.created on the OpenAI-compatible wire, raw.created_at on Ollama).
Comparing and printing created
Numeric context, arithmetic and comparison are unchanged from the old Int:
0 + $response->created # 1700000000
$response->created == 1700000000 # true
$response->created > $cutoff # true/false, as before
sprintf '%d', $response->created # 1700000000
sort { $a->created <=> $b->created } @responses
Three idioms do not survive the change to an object:
String context is the stamp, not the digits.
"$response->created"now interpolates2023-11-14T22:13:20Z, soeqagainst the epoch digits is false and a hash keyed bycreatedre-keys itself. Write0 + $response->createdwhere the number is what is wanted. This is the point of the change, not a side effect — the string form is where the sub-seconds are.refandblessedare no longer empty. Code branching on whether the field is a reference takes the other path now.A JSON encoder without
convert_blesseddies on it — the same way it already dies on "usage", "tool_calls" and "rate_limit", which have been objects for longer. Withconvert_blessedenabled it encodes as the epoch number, because Langertha::Moment carries aTO_JSONthat says so. "to_hash" andTO_JSONon the Response itself are unaffected: they emit a plain number either way.
One smaller change: created is now true in boolean context even when the stamp is the epoch zero, where the old Int 0 was false. "has_created" is the predicate for "did the provider report a stamp", and it is unaffected.
cached_tokens
Number of prompt tokens served from the prefix cache (the cache read count), when the provider reports it. This is the engine-agnostic back-compat accessor: BUILDARGS lifts it off the Langertha::Usage object's "cached_tokens" in Langertha::Usage, which parses it from either wire spelling — usage.prompt_tokens_details.cached_tokens on the OpenAI-compatible wire (SGLang with return_cached_tokens_details enabled, and other servers that emit the detail block) or usage.cache_read_input_tokens on Anthropic. An explicit cached_tokens constructor parameter always wins. undef when the provider does not report a cache-read count.
The cache write count has no Response accessor — read it from the usage object as $response->usage->cache_write_tokens ("cache_write_tokens" in Langertha::Usage).
tool_calls
ArrayRef of Langertha::ToolCall objects extracted from the response, when the engine produced any. Single source of truth for "what tool calls did the model emit" — both native provider tool calling and synthesized fallbacks (forced-name via response_format) land here in the same shape.
For backward compatibility BUILDARGS upgrades plain HashRefs ({ name => ..., arguments => ..., id => ..., synthetic => ... }) into Langertha::ToolCall objects automatically; new code should construct the objects directly.
tool_call
my $tc = $response->tool_call; # first tool call
my $tc = $response->tool_call($name); # named lookup
Returns the Langertha::ToolCall for the first tool call (or the first matching $name), or undef when no tool calls were produced.
tool_call_args
my $args = $response->tool_call_args; # first tool call
my $args = $response->tool_call_args($name); # named lookup
Returns the arguments HashRef of the first tool call, or of the first tool call matching $name. Returns undef when no tool calls were produced.
server_tool_calls
ArrayRef of Langertha::ServerToolCall records: the tool calls the provider ran itself during this request (a web_search_call, an mcp_call, ...), in wire order, when there were any. They are deliberately not on "tool_calls", which lists only calls the client must execute, so chat_with_tools_f and chat_f callers never try to run a web search themselves (ADR 0003 Update k206, ADR 0030). Each record carries the item verbatim in data. Set by the Responses wire (Langertha::Engine::OpenAIResponses); undef elsewhere. Survives "clone_with". Not part of "to_hash" / TO_JSON (the bounded shape, like "citations"); read the attribute.
rate_limit
Optional Langertha::RateLimit object with rate limit information from the API response headers. Only present when the provider returns rate limit headers.
thinking
Chain-of-thought reasoning content. Populated automatically from native API fields (DeepSeek reasoning_content, Anthropic thinking blocks, Gemini thought parts) or from <think> tag filtering when "think_tag_filter" in Langertha::Role::ThinkTag is enabled.
refusal
The model's refusal text, when the provider reports a refusal in a field of its own instead of as content: the OpenAI-compatible message.refusal (a structured-output request the model declined; content is then empty) and a Responses API refusal content part. "content" stays what the model answered, usually '', so check has_refusal before treating an empty reply as empty. A refusal that arrives as a stop reason (Anthropic stop_reason refusal) is in "finish_reason" instead. Survives "clone_with".
citations
Search-augmented source citations, when the provider reports them. An ArrayRef of source HashRefs (each typically carrying url, title, and where available snippet/date). Populated by search-augmented engines that lift them out of the raw payload — Langertha::Engine::Perplexity extracts the Agent API's search_results block here (the classic Sonar top-level citations[] is gone), and the Responses wire adds the answer's url_citation annotations as { url, title, start_index, end_index } (the url verbatim, including OpenAI's ?utm_source=openai). One page is listed once: entries are deduplicated by url, compared without utm_* query parameters, keeping the first entry and filling in only fields it lacks. undef for every engine that does not report citations. Survives "clone_with" so it is preserved through <think> tag filtering.
probes
Provider-specific probe data returned by Langertha::Engine::VLLMHook when a vLLM-Hook server captures attention or hidden-state tensors for a request. A HashRef keyed by probe cache name (e.g. qk_cache, hs_cache) — each cache holding serialized tensors as nested JSON lists, plus an optional config block of scalar metadata. undef for every other engine. Survives "clone_with" so it is preserved through <think> tag filtering.
clone_with
my $new = $response->clone_with(content => $filtered, thinking => $thought);
Returns a new Response with the same attributes as the original, except for the overrides provided. Used by Langertha::Role::ThinkTag to produce a filtered response while preserving metadata.
Every read-only attribute with a has_* predicate is automatically carried forward — the implementation iterates the Moose metaclass rather than maintaining a hand-rolled list, so new attributes added to Langertha::Response are picked up without further changes here. Required attributes (currently just content) are handled explicitly above the loop. Attributes whose names appear in %overrides are skipped during the copy so the override value is what reaches new.
to_hash
my $hash = $response->to_hash;
Returns the canonical, bounded HashRef representation of the response: content plus every metadata field that is present ("id", "model", "finish_reason", "usage", "timing", "created", "cached_tokens", "thinking", "refusal", "rate_limit", "tool_calls"). "raw" and "probes" are deliberately excluded — see "TO_JSON — the canonical, bounded representation".
"created" is emitted as a plain epoch number (via its 0+ overload), not as the Langertha::Moment object — the hash is a bounded interop shape and that key has been a number since it existed.
prompt_tokens
Returns the number of prompt/input tokens. Reads the normalized input_tokens attribute of the Langertha::Usage object (which "from_hash" in Langertha::Usage maps from prompt_tokens, input_tokens, or prompt_eval_count).
completion_tokens
Returns the number of completion/output tokens. Reads the normalized output_tokens attribute of the Langertha::Usage object (which "from_hash" in Langertha::Usage maps from completion_tokens, output_tokens, or eval_count).
total_tokens
Returns the total token count. Reads the total_tokens attribute of the Langertha::Usage object — the provider-supplied value when present, otherwise the sum of prompt and completion tokens (lazy builder).
requests_remaining
Returns the number of requests remaining from rate limit headers, or undef.
tokens_remaining
Returns the number of tokens remaining from rate limit headers, or undef.
SEE ALSO
Langertha::RateLimit - Rate limit data from response headers
Langertha::Stream::Chunk - Single chunk from a streaming response
Langertha::Role::Chat - Chat role that produces response objects
Langertha::Role::OpenAICompatible - Parses responses into this class
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.