NAME
Langertha::RateLimit - Rate limit information from API response headers
VERSION
version 0.503
SYNOPSIS
my $response = $engine->simple_chat('Hello');
if ($response->has_rate_limit) {
my $rl = $response->rate_limit;
say "Requests remaining: ", $rl->requests_remaining // 'unknown';
say "Tokens remaining: ", $rl->tokens_remaining // 'unknown';
say "Reset in: ", $rl->requests_reset // 'unknown', " seconds";
}
# Access raw provider-specific headers
my $raw = $response->rate_limit->raw;
# Also available on the engine (always reflects latest response)
if ($engine->has_rate_limit) {
say "Engine requests remaining: ", $engine->rate_limit->requests_remaining;
}
DESCRIPTION
Normalized rate limit data extracted from HTTP response headers. Different providers use different header naming conventions; this class provides a unified interface.
Supported providers:
OpenAI, Groq, Cerebras, OpenRouter, Replicate, HuggingFace (
x-ratelimit-*)Anthropic (
anthropic-ratelimit-*)
Engines that do not return rate limit headers (DeepSeek, Ollama, vLLM, LlamaCpp, etc.) will not have a rate_limit set.
Reset timing. Providers spell the reset of a bucket in incompatible ways: the OpenAI family sends a duration (a Go time.Duration string), Anthropic sends an instant (RFC 3339), and many send nothing. Rather than leave a single untyped field holding whichever shape arrived, each bucket exposes both a typed instant — "requests_reset_at" / "tokens_reset_at" (Langertha::Moment, when) — and a typed duration — "requests_reset_after" / "tokens_reset_after" (seconds, in how long). The parser fills whichever half the wire actually spoke; the other is derived lazily against "received". Both stay undef for a bucket no header described. The verbatim header strings remain available, untyped, as "requests_reset" / "tokens_reset".
requests_limit
Maximum number of requests allowed in the current window.
requests_remaining
Number of requests remaining in the current window.
requests_reset
The verbatim *-reset-requests / requests-reset header string, exactly as the provider sent it. Its shape is provider-dependent and untyped: OpenAI and the OpenAI-compatible family send a Go time.Duration string ("6m0s", "2m59.56s", "250ms"), Anthropic sends an RFC 3339 instant, and most providers send nothing. A consumer cannot tell which kind it holds without knowing the provider.
Kept for back-compatibility. Prefer the typed pair "requests_reset_at" (WHEN, a Langertha::Moment) and "requests_reset_after" (IN HOW LONG, seconds), which normalize this string and derive the missing half against "received".
tokens_limit
Maximum number of tokens allowed in the current window.
tokens_remaining
Number of tokens remaining in the current window.
tokens_reset
The verbatim *-reset-tokens / tokens-reset header string, exactly as the provider sent it — same untyped, provider-dependent shape as "requests_reset".
Kept for back-compatibility. Prefer the typed pair "tokens_reset_at" and "tokens_reset_after".
received
Langertha::Moment stamped when the response headers were parsed. It is the reference instant against which the two halves of each reset bucket are reconciled: a provider that sent only a duration (OpenAI family) gets its *_reset_at from received + *_reset_after, and one that sent only an instant (Anthropic) gets its *_reset_after from *_reset_at - received. Defaults to "now_utc" in Time::Moment at construction; the header readers stamp it at parse time.
requests_reset_at
Maybe[Langertha::Moment] — the instant the request limit resets (when). Set directly from an RFC 3339 reset header (Anthropic); otherwise derived lazily from "received" plus "requests_reset_after" when the provider sent only a duration (OpenAI family). undef when the provider sent no request reset header at all — absence is the expected path for the many providers that send none, not an error, and no default is invented.
requests_reset_after
Maybe[Num] — seconds until the request limit resets (in how long); may be fractional. Set directly from a Go time.Duration reset header (OpenAI family); otherwise derived lazily from "requests_reset_at" minus "received" when the provider sent only an instant (Anthropic). undef when the provider sent no request reset header at all.
tokens_reset_at
Maybe[Langertha::Moment] — the instant the token limit resets (when). The token-bucket mirror of "requests_reset_at".
tokens_reset_after
Maybe[Num] — seconds until the token limit resets (in how long); may be fractional. The token-bucket mirror of "requests_reset_after".
retry_after
Maybe[Num] — seconds the provider asks the client to wait before retrying, read from "raw". A numeric retry-after-ms (Azure OpenAI and some proxies send it next to retry-after) wins, divided by 1000. Otherwise retry-after is read, a duration whichever form the wire uses: delta-seconds (8; a fractional 1.5 is accepted too) is taken as is, an HTTP-date is measured from "received" (a date already past gives 0, never a negative wait). undef when the response sent neither header, or none in a readable form (the verbatim values stay in "raw"). Providers send it mostly on a 429 or 503, which is why the engine records the rate limit of an error response before it croaks (see "rate_limit" in Langertha::Engine::Remote).
raw
HashRef of all rate-limit-related headers as returned by the provider, keyed by lower-cased header name. The header readers collect every response header matching /^(x-ratelimit-|anthropic-ratelimit-|anthropic-priority-|anthropic-fast-|ratelimitbysize-)/i plus retry-after — a strict superset of the fields the normalized attributes cover, so provider-specific extras survive here even when they carry a window in the name (Groq's per-day request bucket, Cerebras's x-ratelimit-reset-requests-day) or a shape the normalizer does not model (Anthropic's anthropic-priority-* / anthropic-fast-*, Mistral's -minute-suffixed names). Useful for accessing provider-specific fields not covered by the normalized attributes.
_resolve_retry_after
my $seconds = Langertha::RateLimit::_resolve_retry_after(\%raw, $received);
The seconds to wait from a "raw"-shaped hash: retry-after-ms divided by 1000 when it holds a number, else retry-after via "_parse_retry_after". Backs "retry_after" and the retry note in the error messages of Langertha::Role::HTTP, so both say the same number.
_parse_retry_after
my $seconds = Langertha::RateLimit::_parse_retry_after('8'); # 8
my $seconds = Langertha::RateLimit::_parse_retry_after($http_date, $received);
Reads a Retry-After value as the seconds to wait: delta-seconds directly, an HTTP-date (via HTTP::Date) as its distance from $received (a Langertha::Moment, default now), clamped at 0. Returns undef for anything else. The retry-after half of "_resolve_retry_after".
_parse_go_duration
my $seconds = Langertha::RateLimit::_parse_go_duration('2m59.56s'); # 179.56
Parses a Go time.Duration.String() string into fractional seconds. Handles compound ("6m0s", "1h2m3s"), sub-second ("250ms", "35ms") and fractional ("7.66s") forms. Returns undef — never guesses — for anything that is not that exact shape (a bare number, an RFC 3339 instant, an epoch). Used by Langertha::Role::OpenAICompatible to populate "requests_reset_after" / "tokens_reset_after" from the x-ratelimit-reset-* headers.
_collect_headers
my %raw = Langertha::RateLimit::_collect_headers($http_response);
Collects every rate-limit-related response header into a hash keyed by lower-cased name — the single source of truth for the "raw" superset. Matches the x-ratelimit- / anthropic-ratelimit- / anthropic-priority- / anthropic-fast- / ratelimitbysize- prefixes plus retry-after and retry-after-ms. The wire-envelope roles call this, then normalize the known subset out of the result.
to_hash
my $hash = $rate_limit->to_hash;
Returns a flat HashRef of all defined rate limit fields plus the raw headers. The typed reset halves appear here whenever they were sent or can be derived: "requests_reset_at" / "tokens_reset_at" as a plain epoch number (their 0+ overload, matching "created" in Langertha::Response) and "requests_reset_after" / "tokens_reset_after" as seconds, and "retry_after" in seconds when the provider sent one. A bucket the provider never spoke is omitted rather than defaulted. "received" is not included; read it from the accessor when the derivation anchor is needed.
TO_JSON
my $json = JSON::MaybeXS->new(convert_blessed => 1)->encode({ rate_limit => $rl });
Serialization hook for JSON encoders configured with convert_blessed. Returns "to_hash" without the raw key.
TO_JSON fires implicitly, from wherever the surrounding structure happens to be encoded — a trace, a log line, a queue message. The caller did not ask for this object and cannot see what it contributed, so the implicit path carries only the normalized, provider-agnostic fields. raw holds the provider's own rate-limit response headers; shipping those into third-party sinks by accident is not something a caller should have to opt out of.
A caller who wants the raw headers asks for them explicitly — via "to_hash" or "raw". That is the only difference between the two methods, and the reason they deliberately do not return the same thing.
SEE ALSO
Langertha::Response - Response objects carry rate limit data
Langertha::Role::HTTP - Extracts rate limit headers during response parsing
Langertha::Engine::Remote - Stores the latest rate limit on the engine
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.