NAME
Langertha::Usage - Immutable value object for LLM token usage with cross-provider conversion
VERSION
version 0.503
from_raw
my $data = $engine->parse_response($http_response);
my $usage = Langertha::Usage->from_raw($data)
or return; # the body reported no usage
printf "%d in / %d out\n", $usage->input_tokens, $usage->output_tokens;
Class method. Builds a Usage from a raw decoded provider response body, the HashRef parse_response returns — for callers that send their own requests and never get a Langertha::Response. It finds the usage block wherever the provider puts it: usage (OpenAI-compatible, Anthropic, Open-Responses), usageMetadata (Gemini), response.usage (an Open-Responses event envelope), the top-level prompt_eval_count / eval_count of Ollama's native API, or the top-level prompt_length / num_generated_tokens / num_cached_tokens of AKI.IO's native API. The usage block and the Ollama counts are then read by "from_hash", so every spelling it knows applies.
Returns undef when the body reports no usage (or is not a HashRef), so a caller can tell "not reported" from "zero tokens". As in Langertha::Engine::Ollama, an Ollama count of zero counts as not reported.
from_response
my $usage = Langertha::Usage->from_response($response_or_body);
Class method. Builds a Usage from a Langertha::Response (its usage), or from a raw body HashRef via "from_raw". Always returns a Usage: all-zero when nothing is reported.
from_hash
my $usage = Langertha::Usage->from_hash($usage_block);
Class method. Builds a Usage from a provider's usage block, accepting the OpenAI (prompt_tokens / completion_tokens), Anthropic and Open-Responses (input_tokens / output_tokens), Ollama (prompt_eval_count / eval_count) and Gemini (promptTokenCount / candidatesTokenCount / totalTokenCount) spellings, in that order of preference, plus the cache counts described under "cached_tokens" and "cache_write_tokens", the "reasoning_tokens" share and the provider-reported "cost_usd".
Gemini counts thinking and the tool-use prompt beside the prompt and the answer (totalTokenCount is their sum), so from the Gemini spelling output_tokens is candidatesTokenCount plus thoughtsTokenCount and input_tokens is promptTokenCount plus toolUsePromptTokenCount. Thinking is billed at the output rate; output_tokens then means what OpenAI's completion_tokens means, reasoning included.
merge
my $sum = $usage->merge($other);
Returns a new Usage holding the sum of both: input_tokens, output_tokens, and "cached_tokens" / "cache_write_tokens" / "reasoning_tokens" (a side that did not report a count adds nothing; the sum stays undef when neither did). "input_includes_cache" comes from the sides that reported a cache count. When both count it beside input_tokens (false) the sum is false. When one counts it inside (true, or undef) and the other beside, the beside side's cache counts are added to its input_tokens before summing and the sum is true, so pricing the merged Usage costs the same as pricing both parts. When both count it inside, the sum is true, or undef if either side's flag was undef. "cost_usd" is summed only when both sides report one; otherwise the sum's cost is undef, since a partial sum would read as the whole bill. "raw" is not carried over.
HASH OVERLOAD — backward compatibility
Langertha::Usage overloads %{}, so a Usage object can keep being dereferenced as a hash: $response->usage->{prompt_tokens} keeps working. This is the deliberate back-compat seam for the coercion of "usage" in Langertha::Response from a raw HashRef to a Langertha::Usage object (karr #43).
When the object was built from a provider hash (via "from_hash", which is what Langertha::Response does in BUILDARGS), the overload returns that hash verbatim — stored in "raw". Callers therefore keep seeing exactly the engine-normalized keys they always saw: provider-specific extras (Anthropic cache tokens, OpenAI prompt_tokens_details / completion_tokens_details) survive, and keys the engine normalized away (Gemini camelCase, Ollama prompt_eval_count / eval_count) stay absent. exists checks behave exactly as they did on the raw hash.
These engine-normalized keys are legacy: the accessors above read every spelling, so no engine needs to rename for Usage's sake any more. They are kept for compatibility and are not deprecated. For example Langertha::Engine::Gemini still serves prompt_tokens, completion_tokens, total_tokens and cached_content_token_count here in place of Gemini's camelCase usageMetadata names.
For a Usage constructed directly (no raw hash), the overload returns the canonical "to_hash" shape (input_tokens / output_tokens / total_tokens). New code should prefer the accessors and the to_*_format methods over hash dereferencing.
Why the field hash
A naive %{} overload on a Moose class is impossible: the overload hijacks every $self->{attr} deref on the object, including the ones Moose's generated accessors use internally, so the accessors read through the overload and recurse (verified: infinite recursion / undef reads). The canonical Perl solution is to keep the attribute values in a Hash::Util::FieldHash keyed by object identity and route both the accessors (via around modifiers) and the overload through it. The object's own hash is then never read, so the overload only ever serves caller hash derefs.
cached_tokens
Number of prompt tokens served from the provider's prefix cache (the cache read count), when the provider reports it. "from_hash" parses it from any of its wire spellings: OpenAI's Chat wire nests it at usage.prompt_tokens_details.cached_tokens (also Mistral, and any OpenAI-compatible server such as SGLang with return_cached_tokens_details), the Open-Responses envelope (OpenAI Responses, the Perplexity Agent API) nests it at usage.input_tokens_details.cached_tokens (or the Anthropic-named cache_read_input_tokens in the same block), and Anthropic reports it flat as usage.cache_read_input_tokens. The OpenAI Chat nesting wins, then the Responses nesting, then the Anthropic flat key. Gemini reports it as usageMetadata.cachedContentTokenCount (cached_content_token_count after Langertha::Engine::Gemini renames it), read after the Anthropic key. A flat canonical cached_tokens (what Langertha::Engine::AKI puts in its usage hash, from AKI.IO's native num_cached_tokens) is read last. undef when the provider does not report a cache-read count.
Note: on OpenAI (GPT-5.6 and later) this count excludes hidden tokens and rounds down to a multiple of 128, so cost arithmetic built on it is approximate by construction.
cache_write_tokens
Number of prompt tokens written to the provider's prefix cache (the cache creation count), a distinct quantity from "cached_tokens" and deliberately not folded into it. "from_hash" parses it from any of its wire spellings: OpenAI's Chat wire nests it at usage.prompt_tokens_details.cache_write_tokens, the Open-Responses envelope (OpenAI Responses, the Perplexity Agent API) nests it at usage.input_tokens_details.cache_creation_input_tokens, and Anthropic reports it flat as usage.cache_creation_input_tokens (Anthropic further splits that count across TTL tiers under usage.cache_creation, which stays verbatim in "raw"). The OpenAI Chat nesting wins, then the Responses nesting, then the Anthropic flat key; without the flat key the ephemeral_5m_input_tokens / ephemeral_1h_input_tokens of usage.cache_creation are summed. undef when the provider does not report a cache-write count.
input_includes_cache
Whether "cached_tokens" and "cache_write_tokens" are already counted in input_tokens. input_tokens keeps the meaning the wire gives it, and that meaning differs: OpenAI's prompt_tokens, the Open-Responses input_tokens, Gemini's promptTokenCount and AKI.IO's prompt_length include the cached tokens; Anthropic's input_tokens counts only the tokens after the last cache breakpoint, with cache_read_input_tokens and cache_creation_input_tokens beside it. "from_hash" and "from_raw" set this flag from where they found the cache counts: true for a count nested in prompt_tokens_details / input_tokens_details, for Gemini's and for the flat cached_tokens; false for Anthropic's flat keys. undef when no cache count was reported, or when the object was built with new and the flag was not passed.
The Anthropic-compatible shims follow the Anthropic spelling, but not always its meaning: AKI.IO's /anthropic endpoint reports the same numbers under input_tokens as its OpenAI face does under prompt_tokens (cached reads included). A canonical input_includes_cache key in the usage hash therefore beats the inference when a cache count was found; Langertha::Role::AnthropicCompatible adds it for an engine whose _usage_input_includes_cache hook answers (Langertha::Engine::AKIAnthropic answers true), so the key also shows in $response->usage->{...}.
reasoning_tokens
How many of output_tokens the model spent reasoning (thinking), when the provider reports it. The count is already part of output_tokens — never add it again. "from_hash" reads OpenAI Chat's completion_tokens_details.reasoning_tokens, then the Open-Responses output_tokens_details.reasoning_tokens, then Gemini's thoughtsTokenCount (which Gemini reports beside candidatesTokenCount, so "from_hash" adds it into output_tokens). undef when the provider does not report one.
cost_usd
What the provider says the request was billed, in US dollars, when it reports it: the actual charge after its discounts (prompt caching) and including its server-side tool fees. xAI puts it in every usage block (chat completions, Responses, image and video generation) as the integer cost_in_usd_ticks (1 USD = 10,000,000,000 ticks), and on Responses also as cost_in_nano_usd (1 USD = 1,000,000,000 nano-USD). "from_hash" converts whichever is present, cost_in_usd_ticks first as the finer unit; the integers stay verbatim in "raw", so $usage->{cost_in_usd_ticks} still gives exact integer accounting. A stream carries it only in its usage frame, which comes when the request asks for it (stream_options => { include_usage => 1 }).
The Perplexity Agent API sends usage.cost as an object that names its currency (currency, total_cost, and the parts such as input_cost and tool_calls_cost); "from_hash" reads total_cost when currency is USD, and nothing when the currency is another one or the total is missing. OpenRouter sends usage.cost as a bare number in its credits, which are US dollars; a bare number names no unit, so "from_hash" does not read it on its own. Langertha::Engine::OpenRouter adds it to the usage block as the canonical cost_usd key, which "from_hash" reads first, so it arrives on the engine's responses and stream chunks, but not through "from_raw" on an OpenRouter body. For a BYOK request cost is still only what OpenRouter charged your credits; what the key's own provider charged is in cost_details.upstream_inference_cost, which stays in "raw" and is not added.
undef when the provider reports no cost, never 0: it is not an estimate, and Langertha::Pricing does not read it — that builds a Langertha::Cost from your own price rules.
uncached_input_tokens
my $fresh = $usage->uncached_input_tokens;
The input tokens that were neither read from nor written to the prompt cache. When "input_includes_cache" is false this is input_tokens; otherwise (true or undef) it is input_tokens minus "cached_tokens" minus "cache_write_tokens", never below zero. An undef flag is read as "included" because that is what total_tokens assumes of input_tokens. "cost_for" in Langertha::Pricing prices this count at the input rate when a rule has a cache rate.
raw
The provider-verbatim hash the object was built from, when it was built via "from_hash". undef for directly constructed objects. Read-only; the canonical "to_hash" and TO_JSON deliberately do not include it — it exists solely to back the %{} overload.
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.