NAME

Langertha::Response - LLM response with metadata

VERSION

version 0.503

SYNOPSIS

my $response = $engine->simple_chat('Hello');

# Stringifies to content (backward compatible)
print $response;
print "Response: $response\n";

# Access metadata
say $response->model;
say $response->id;
say $response->finish_reason;

# Token usage
say "Prompt tokens: ", $response->prompt_tokens;
say "Completion tokens: ", $response->completion_tokens;
say "Total tokens: ", $response->total_tokens;

# Full raw response
use Data::Dumper;
print Dumper($response->raw);

DESCRIPTION

Wraps LLM response text content together with all available metadata from the API response. Uses overload for string context so existing code treating responses as plain strings continues to work.

Boolean context — a Response is always true

Langertha::Response overloads bool to a constant true: the object exists, so it is true, whatever "content" happens to hold. This is not cosmetic. With only the "" overload and fallback => 1, Perl derives boolean context from stringification, so a Response whose content is the empty string was false — and that is exactly the shape of a tool-call-only reply (an Anthropic turn that is pure tool_use, an Ollama reply with empty message.content plus a tool_calls array), as well as of a model that legitimately answers "0". Callers writing

my $resp = $engine->simple_chat(@messages);
if ($resp) { ... }             # entered now; silently skipped before
$resp or die "no response";    # no longer dies on a perfectly good Response

used to drop precisely the responses that carry "tool_calls". Decided in karr #100.

Asking whether the content is empty is a different question, and keeps its own spelling:

if ( length "$resp" )         { ... }   # the model produced text
if ( $resp->has_tool_calls )  { ... }   # the model emitted tool calls

The "" overload is unchanged — "$response" is still "content", and eq, ne and concatenation keep routing through it. The change also makes this class agree with the rest of the distribution: Langertha::ToolCall, Langertha::Usage and Langertha::Stream::Chunk are already true-because-they-exist. Langertha::Result got the same treatment for the same reason.

TO_JSON — the canonical, bounded representation

Langertha::Response carries a TO_JSON (delegating to "to_hash") so that every JSON::MaybeXS backend encodes a bare Response identically when convert_blessed is enabled. Decided in karr #50, after karr #43 fixed the permanent shape of the class ("usage" is now a Langertha::Usage object with its own TO_JSON).

The shape is deliberately bounded: "content" plus the metadata fields that are present ("id", "model", "finish_reason", "usage", "timing", "created", "thinking", "refusal", "rate_limit", "tool_calls"). "raw" and "probes" are not included — "raw" is the entire provider payload (duplicating every other field, including the echoed prompt) and "probes" can hold megabytes of tensor data. A TO_JSON that kept them would blow up every trace it touched; one that dropped them keeps the object's JSON form useful without silently discarding the fields a consumer is likely to want.

The string overload is unchanged: "$response" still returns "content", and the TO_JSON does not replace it.

Backend uniformity is the point. Before this method existed, what an encoder with convert_blessed did with a bare Response depended on the JSON::MaybeXS backend: Cpanel::JSON::XS fell back to the "" overload and emitted the content string, while JSON::XS and JSON::PP threw. Consumers on Cpanel silently got {"response":"hello"}, consumers on JSON::XS/JSON::PP got an exception — the same application behaved differently on two machines. With TO_JSON all three emit the canonical hash. The Cpanel behavior was never a contract (it is gone now); callers that want the content string should stringify explicitly.

content

The text content of the response. Required.

raw

The full parsed API response as a HashRef.

id

Provider-specific response ID.

model

The actual model used for the response.

finish_reason

Why the response ended: stop, end_turn, length, tool_calls, etc. Provider-specific values are preserved as-is.

usage

Token usage as a Langertha::Usage object. For backward compatibility BUILDARGS upgrades a plain HashRef (the raw provider usage block) into a Langertha::Usage automatically; new code can construct the object directly. The object overloads %{}, so existing callers that dereference $response->usage->{prompt_tokens} keep working — the overload returns the provider-verbatim hash (see "HASH OVERLOAD" in Langertha::Usage).

timing

Timing information as a HashRef. Holds client-measured ttft_seconds and total_seconds (Float, seconds) on every engine, plus provider-native stage durations that engines may populate engine-specifically — e.g. Ollama's total_seconds, load_seconds, prompt_eval_seconds, eval_seconds (all Float, seconds), with its original *_duration keys in nanoseconds preserved for backward compatibility. See "ttft_seconds" and "total_seconds" for the standard accessors.

Merge policy: first-write-wins. When both a provider-supplied duration (e.g. Ollama's total_seconds derived from total_duration) and a client-measured duration (Langertha::Role::Chat wrapper) are available for the same key, the provider value wins. See "_merge_timing_field" in Langertha::Role::Chat and ADR 0011 in docs/adr/0011-response-timing-seam.md for the rationale — server-side duration excludes network jitter and is the better signal for model-latency observability. Callers that need the client wall-clock (RTT, queue, TLS handshake, network back) should read $response->timing directly: client-measured deltas are not written under ttft_seconds / total_seconds when the provider already populated those keys, but Langertha::Role::Chat does not delete other timing entries on conflict.

ttft_seconds

my $ttft = $response->ttft_seconds;     # Float, undef when unmeasured
if ($response->has_ttft) { ... }

Returns time-to-first-token in seconds (Float), measured client-side between request send and the first streamed chunk. undef for sync (non-streaming) calls and for engines that did not record the metric. Use has_ttft to test availability without warnings.

total_seconds

my $total = $response->total_seconds;   # Float, undef when unmeasured
if ($response->has_total) { ... }

Returns end-to-end request time in seconds (Float). On engines that populate the field from a provider-native metric (currently Ollama, from total_duration), this is the server-reported duration — time the model actually spent generating, excluding network jitter. On engines where only the client wrapper measured, it is the wall-clock duration from before user_agent->request (sync) or do_request (async) to after the response body was fully consumed.

The two are not interchangeable: server-time answers "how fast is the model", wall-clock answers "how long did my call take end-to-end". Callers that need the wall-clock RTT should read $response->timing directly (the client wrapper writes its measurement under a sibling key when the provider has not claimed total_seconds; when both are present the provider value wins — see "timing").

undef when the engine did not record the metric. Use has_total to test availability without warnings.

created

The instant the provider says the response was created, as a Langertha::Moment — a Time::Moment subclass, so the full civil time is there: sub-seconds, UTC offset, and every Time::Moment accessor.

In numeric context it is the Unix timestamp it always was. The 0+ overload yields whole seconds since the epoch, so 0 + $response->created, $response->created > $cutoff and int($response->created) keep returning exactly what they returned when this attribute was a Maybe[Int]. In string context it is the full ISO-8601 stamp instead — which is where the sub-seconds are, and which is not what the old Int stringified to. See "Comparing and printing created" below.

Engine-agnostic by contract, and normalized here rather than per engine: BUILDARGS runs every incoming value through "from_wire" in Langertha::Moment, which takes the OpenAI-compatible wire's epoch integer and Ollama's RFC3339 created_at string alike. A value it cannot read — including the 0001-01-01T00:00:00Z zero-value sentinel Go-based servers emit — drops the field instead of failing the constructor: has_created is then false and the response is built regardless. A timestamp is metadata and must never take a whole reply down with it (GitHub issue #3, karr #92 / #117).

The provider's native form always stays available verbatim under "raw" (raw.created on the OpenAI-compatible wire, raw.created_at on Ollama).

Comparing and printing created

Numeric context, arithmetic and comparison are unchanged from the old Int:

0 + $response->created                 # 1700000000
$response->created == 1700000000       # true
$response->created > $cutoff           # true/false, as before
sprintf '%d', $response->created       # 1700000000
sort { $a->created <=> $b->created } @responses

Three idioms do not survive the change to an object:

  • String context is the stamp, not the digits. "$response->created" now interpolates 2023-11-14T22:13:20Z, so eq against the epoch digits is false and a hash keyed by created re-keys itself. Write 0 + $response->created where the number is what is wanted. This is the point of the change, not a side effect — the string form is where the sub-seconds are.

  • ref and blessed are no longer empty. Code branching on whether the field is a reference takes the other path now.

  • A JSON encoder without convert_blessed dies on it — the same way it already dies on "usage", "tool_calls" and "rate_limit", which have been objects for longer. With convert_blessed enabled it encodes as the epoch number, because Langertha::Moment carries a TO_JSON that says so. "to_hash" and TO_JSON on the Response itself are unaffected: they emit a plain number either way.

One smaller change: created is now true in boolean context even when the stamp is the epoch zero, where the old Int 0 was false. "has_created" is the predicate for "did the provider report a stamp", and it is unaffected.

cached_tokens

Number of prompt tokens served from the prefix cache (the cache read count), when the provider reports it. This is the engine-agnostic back-compat accessor: BUILDARGS lifts it off the Langertha::Usage object's "cached_tokens" in Langertha::Usage, which parses it from either wire spelling — usage.prompt_tokens_details.cached_tokens on the OpenAI-compatible wire (SGLang with return_cached_tokens_details enabled, and other servers that emit the detail block) or usage.cache_read_input_tokens on Anthropic. An explicit cached_tokens constructor parameter always wins. undef when the provider does not report a cache-read count.

The cache write count has no Response accessor — read it from the usage object as $response->usage->cache_write_tokens ("cache_write_tokens" in Langertha::Usage).

tool_calls

ArrayRef of Langertha::ToolCall objects extracted from the response, when the engine produced any. Single source of truth for "what tool calls did the model emit" — both native provider tool calling and synthesized fallbacks (forced-name via response_format) land here in the same shape.

For backward compatibility BUILDARGS upgrades plain HashRefs ({ name => ..., arguments => ..., id => ..., synthetic => ... }) into Langertha::ToolCall objects automatically; new code should construct the objects directly.

tool_call

my $tc = $response->tool_call;          # first tool call
my $tc = $response->tool_call($name);   # named lookup

Returns the Langertha::ToolCall for the first tool call (or the first matching $name), or undef when no tool calls were produced.

tool_call_args

my $args = $response->tool_call_args;            # first tool call
my $args = $response->tool_call_args($name);     # named lookup

Returns the arguments HashRef of the first tool call, or of the first tool call matching $name. Returns undef when no tool calls were produced.

server_tool_calls

ArrayRef of Langertha::ServerToolCall records: the tool calls the provider ran itself during this request (a web_search_call, an mcp_call, ...), in wire order, when there were any. They are deliberately not on "tool_calls", which lists only calls the client must execute, so chat_with_tools_f and chat_f callers never try to run a web search themselves (ADR 0003 Update k206, ADR 0030). Each record carries the item verbatim in data. Set by the Responses wire (Langertha::Engine::OpenAIResponses); undef elsewhere. Survives "clone_with". Not part of "to_hash" / TO_JSON (the bounded shape, like "citations"); read the attribute.

rate_limit

Optional Langertha::RateLimit object with rate limit information from the API response headers. Only present when the provider returns rate limit headers.

thinking

Chain-of-thought reasoning content. Populated automatically from native API fields (DeepSeek reasoning_content, Anthropic thinking blocks, Gemini thought parts) or from <think> tag filtering when "think_tag_filter" in Langertha::Role::ThinkTag is enabled.

refusal

The model's refusal text, when the provider reports a refusal in a field of its own instead of as content: the OpenAI-compatible message.refusal (a structured-output request the model declined; content is then empty) and a Responses API refusal content part. "content" stays what the model answered, usually '', so check has_refusal before treating an empty reply as empty. A refusal that arrives as a stop reason (Anthropic stop_reason refusal) is in "finish_reason" instead. Survives "clone_with".

citations

Search-augmented source citations, when the provider reports them. An ArrayRef of source HashRefs (each typically carrying url, title, and where available snippet/date). Populated by search-augmented engines that lift them out of the raw payload — Langertha::Engine::Perplexity extracts the Agent API's search_results block here (the classic Sonar top-level citations[] is gone), and the Responses wire adds the answer's url_citation annotations as { url, title, start_index, end_index } (the url verbatim, including OpenAI's ?utm_source=openai). One page is listed once: entries are deduplicated by url, compared without utm_* query parameters, keeping the first entry and filling in only fields it lacks. undef for every engine that does not report citations. Survives "clone_with" so it is preserved through <think> tag filtering.

probes

Provider-specific probe data returned by Langertha::Engine::VLLMHook when a vLLM-Hook server captures attention or hidden-state tensors for a request. A HashRef keyed by probe cache name (e.g. qk_cache, hs_cache) — each cache holding serialized tensors as nested JSON lists, plus an optional config block of scalar metadata. undef for every other engine. Survives "clone_with" so it is preserved through <think> tag filtering.

clone_with

my $new = $response->clone_with(content => $filtered, thinking => $thought);

Returns a new Response with the same attributes as the original, except for the overrides provided. Used by Langertha::Role::ThinkTag to produce a filtered response while preserving metadata.

Every read-only attribute with a has_* predicate is automatically carried forward — the implementation iterates the Moose metaclass rather than maintaining a hand-rolled list, so new attributes added to Langertha::Response are picked up without further changes here. Required attributes (currently just content) are handled explicitly above the loop. Attributes whose names appear in %overrides are skipped during the copy so the override value is what reaches new.

to_hash

my $hash = $response->to_hash;

Returns the canonical, bounded HashRef representation of the response: content plus every metadata field that is present ("id", "model", "finish_reason", "usage", "timing", "created", "cached_tokens", "thinking", "refusal", "rate_limit", "tool_calls"). "raw" and "probes" are deliberately excluded — see "TO_JSON — the canonical, bounded representation".

"created" is emitted as a plain epoch number (via its 0+ overload), not as the Langertha::Moment object — the hash is a bounded interop shape and that key has been a number since it existed.

prompt_tokens

Returns the number of prompt/input tokens. Reads the normalized input_tokens attribute of the Langertha::Usage object (which "from_hash" in Langertha::Usage maps from prompt_tokens, input_tokens, or prompt_eval_count).

completion_tokens

Returns the number of completion/output tokens. Reads the normalized output_tokens attribute of the Langertha::Usage object (which "from_hash" in Langertha::Usage maps from completion_tokens, output_tokens, or eval_count).

total_tokens

Returns the total token count. Reads the total_tokens attribute of the Langertha::Usage object — the provider-supplied value when present, otherwise the sum of prompt and completion tokens (lazy builder).

requests_remaining

Returns the number of requests remaining from rate limit headers, or undef.

tokens_remaining

Returns the number of tokens remaining from rate limit headers, or undef.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.