NAME

Langertha::RateLimit - Rate limit information from API response headers

VERSION

version 0.503

SYNOPSIS

my $response = $engine->simple_chat('Hello');

if ($response->has_rate_limit) {
    my $rl = $response->rate_limit;
    say "Requests remaining: ", $rl->requests_remaining // 'unknown';
    say "Tokens remaining: ", $rl->tokens_remaining // 'unknown';
    say "Reset in: ", $rl->requests_reset // 'unknown', " seconds";
}

# Access raw provider-specific headers
my $raw = $response->rate_limit->raw;

# Also available on the engine (always reflects latest response)
if ($engine->has_rate_limit) {
    say "Engine requests remaining: ", $engine->rate_limit->requests_remaining;
}

DESCRIPTION

Normalized rate limit data extracted from HTTP response headers. Different providers use different header naming conventions; this class provides a unified interface.

Supported providers:

  • OpenAI, Groq, Cerebras, OpenRouter, Replicate, HuggingFace (x-ratelimit-*)

  • Anthropic (anthropic-ratelimit-*)

Engines that do not return rate limit headers (DeepSeek, Ollama, vLLM, LlamaCpp, etc.) will not have a rate_limit set.

Reset timing. Providers spell the reset of a bucket in incompatible ways: the OpenAI family sends a duration (a Go time.Duration string), Anthropic sends an instant (RFC 3339), and many send nothing. Rather than leave a single untyped field holding whichever shape arrived, each bucket exposes both a typed instant — "requests_reset_at" / "tokens_reset_at" (Langertha::Moment, when) — and a typed duration — "requests_reset_after" / "tokens_reset_after" (seconds, in how long). The parser fills whichever half the wire actually spoke; the other is derived lazily against "received". Both stay undef for a bucket no header described. The verbatim header strings remain available, untyped, as "requests_reset" / "tokens_reset".

requests_limit

Maximum number of requests allowed in the current window.

requests_remaining

Number of requests remaining in the current window.

requests_reset

The verbatim *-reset-requests / requests-reset header string, exactly as the provider sent it. Its shape is provider-dependent and untyped: OpenAI and the OpenAI-compatible family send a Go time.Duration string ("6m0s", "2m59.56s", "250ms"), Anthropic sends an RFC 3339 instant, and most providers send nothing. A consumer cannot tell which kind it holds without knowing the provider.

Kept for back-compatibility. Prefer the typed pair "requests_reset_at" (WHEN, a Langertha::Moment) and "requests_reset_after" (IN HOW LONG, seconds), which normalize this string and derive the missing half against "received".

tokens_limit

Maximum number of tokens allowed in the current window.

tokens_remaining

Number of tokens remaining in the current window.

tokens_reset

The verbatim *-reset-tokens / tokens-reset header string, exactly as the provider sent it — same untyped, provider-dependent shape as "requests_reset".

Kept for back-compatibility. Prefer the typed pair "tokens_reset_at" and "tokens_reset_after".

received

Langertha::Moment stamped when the response headers were parsed. It is the reference instant against which the two halves of each reset bucket are reconciled: a provider that sent only a duration (OpenAI family) gets its *_reset_at from received + *_reset_after, and one that sent only an instant (Anthropic) gets its *_reset_after from *_reset_at - received. Defaults to "now_utc" in Time::Moment at construction; the header readers stamp it at parse time.

requests_reset_at

Maybe[Langertha::Moment] — the instant the request limit resets (when). Set directly from an RFC 3339 reset header (Anthropic); otherwise derived lazily from "received" plus "requests_reset_after" when the provider sent only a duration (OpenAI family). undef when the provider sent no request reset header at all — absence is the expected path for the many providers that send none, not an error, and no default is invented.

requests_reset_after

Maybe[Num] — seconds until the request limit resets (in how long); may be fractional. Set directly from a Go time.Duration reset header (OpenAI family); otherwise derived lazily from "requests_reset_at" minus "received" when the provider sent only an instant (Anthropic). undef when the provider sent no request reset header at all.

tokens_reset_at

Maybe[Langertha::Moment] — the instant the token limit resets (when). The token-bucket mirror of "requests_reset_at".

tokens_reset_after

Maybe[Num] — seconds until the token limit resets (in how long); may be fractional. The token-bucket mirror of "requests_reset_after".

retry_after

Maybe[Num] — seconds the provider asks the client to wait before retrying, read from "raw". A numeric retry-after-ms (Azure OpenAI and some proxies send it next to retry-after) wins, divided by 1000. Otherwise retry-after is read, a duration whichever form the wire uses: delta-seconds (8; a fractional 1.5 is accepted too) is taken as is, an HTTP-date is measured from "received" (a date already past gives 0, never a negative wait). undef when the response sent neither header, or none in a readable form (the verbatim values stay in "raw"). Providers send it mostly on a 429 or 503, which is why the engine records the rate limit of an error response before it croaks (see "rate_limit" in Langertha::Engine::Remote).

raw

HashRef of all rate-limit-related headers as returned by the provider, keyed by lower-cased header name. The header readers collect every response header matching /^(x-ratelimit-|anthropic-ratelimit-|anthropic-priority-|anthropic-fast-|ratelimitbysize-)/i plus retry-after — a strict superset of the fields the normalized attributes cover, so provider-specific extras survive here even when they carry a window in the name (Groq's per-day request bucket, Cerebras's x-ratelimit-reset-requests-day) or a shape the normalizer does not model (Anthropic's anthropic-priority-* / anthropic-fast-*, Mistral's -minute-suffixed names). Useful for accessing provider-specific fields not covered by the normalized attributes.

_resolve_retry_after

my $seconds = Langertha::RateLimit::_resolve_retry_after(\%raw, $received);

The seconds to wait from a "raw"-shaped hash: retry-after-ms divided by 1000 when it holds a number, else retry-after via "_parse_retry_after". Backs "retry_after" and the retry note in the error messages of Langertha::Role::HTTP, so both say the same number.

_parse_retry_after

my $seconds = Langertha::RateLimit::_parse_retry_after('8');   # 8
my $seconds = Langertha::RateLimit::_parse_retry_after($http_date, $received);

Reads a Retry-After value as the seconds to wait: delta-seconds directly, an HTTP-date (via HTTP::Date) as its distance from $received (a Langertha::Moment, default now), clamped at 0. Returns undef for anything else. The retry-after half of "_resolve_retry_after".

_parse_go_duration

my $seconds = Langertha::RateLimit::_parse_go_duration('2m59.56s');   # 179.56

Parses a Go time.Duration.String() string into fractional seconds. Handles compound ("6m0s", "1h2m3s"), sub-second ("250ms", "35ms") and fractional ("7.66s") forms. Returns undef — never guesses — for anything that is not that exact shape (a bare number, an RFC 3339 instant, an epoch). Used by Langertha::Role::OpenAICompatible to populate "requests_reset_after" / "tokens_reset_after" from the x-ratelimit-reset-* headers.

_collect_headers

my %raw = Langertha::RateLimit::_collect_headers($http_response);

Collects every rate-limit-related response header into a hash keyed by lower-cased name — the single source of truth for the "raw" superset. Matches the x-ratelimit- / anthropic-ratelimit- / anthropic-priority- / anthropic-fast- / ratelimitbysize- prefixes plus retry-after and retry-after-ms. The wire-envelope roles call this, then normalize the known subset out of the result.

to_hash

my $hash = $rate_limit->to_hash;

Returns a flat HashRef of all defined rate limit fields plus the raw headers. The typed reset halves appear here whenever they were sent or can be derived: "requests_reset_at" / "tokens_reset_at" as a plain epoch number (their 0+ overload, matching "created" in Langertha::Response) and "requests_reset_after" / "tokens_reset_after" as seconds, and "retry_after" in seconds when the provider sent one. A bucket the provider never spoke is omitted rather than defaulted. "received" is not included; read it from the accessor when the derivation anchor is needed.

TO_JSON

my $json = JSON::MaybeXS->new(convert_blessed => 1)->encode({ rate_limit => $rl });

Serialization hook for JSON encoders configured with convert_blessed. Returns "to_hash" without the raw key.

TO_JSON fires implicitly, from wherever the surrounding structure happens to be encoded — a trace, a log line, a queue message. The caller did not ask for this object and cannot see what it contributed, so the implicit path carries only the normalized, provider-agnostic fields. raw holds the provider's own rate-limit response headers; shipping those into third-party sinks by accident is not something a caller should have to opt out of.

A caller who wants the raw headers asks for them explicitly — via "to_hash" or "raw". That is the only difference between the two methods, and the reason they deliberately do not return the same thing.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.