NAME

Langertha::Role::HTTP - Role for HTTP APIs

VERSION

version 0.503

url

Base URL for API requests. Optional — many engines hard-code their default URL internally and only require this attribute to be set when pointing at a custom or self-hosted endpoint.

connect_address

my $engine = Langertha::Engine::OpenAI->new(
  url             => 'https://llm.example.com/v1',
  api_key         => $key,
  connect_address => '203.0.113.7',   # what llm.example.com resolved to when you checked it
);

An IPv4 or IPv6 address literal (no brackets, port or scope; undef means none). When set, every request the engine sends to the host of its "url" opens its TCP connection to this address instead of resolving the name again. The request still names the host everywhere else: the Host header, and for https the TLS SNI and the name the server certificate is verified against. Certificate verification itself is configured as it would be without the pin. This is for a caller that resolved the host and checked its addresses against a policy (no loopback, private or cloud-metadata address, say): without the pin the name is resolved again at connect time, and a DNS answer that changed in between (DNS rebinding) would send the request somewhere that was never checked (karr k375).

It covers every request core sends to that host, on every backend core builds: the synchronous methods and the synchronous fallback of the _f methods through the engine's "user_agent" (a Langertha::HTTP::UserAgent built with the pin), and the Net::Async::HTTP backend of the _f methods, streaming included ("async_request_f" in Langertha::Role::AsyncHTTP). That includes list_models, the capability probe and the metrics scrape, which go to the same host. Requests to other hosts (a Langfuse endpoint, say) are not pinned. Not pinned either: image URLs fetched for inlining ("ensure_base64" in Langertha::Content::Image, ensure_base64_f) resolve their host again, even when it is the pinned host; only on the synchronous fallback, where the fetch runs through a copy of the engine's agent, is a same-host image pinned. Vet image URLs with "inline_image_url_filter" in Langertha::Role::Chat if that matters.

Redirects: a redirect from the pinned host to another host is not followed on any backend; the 3xx is returned with a Client-Warning (redirect not followed: Langertha::HTTP::Redirect: connect_address pins ...). The address was checked for this host only, and following would resolve the new host afresh. A redirect on the same host (another port or path) stays pinned. Configure the URL the redirect points to, with its own checked address, instead.

Before a request is written the connection is checked, new or reused: its peer must be the pinned address and, over https, the session's certificate chain must have verified and the certificate must be for the host. Where verification is switched off the check follows: on the synchronous side both parts apply when LWP's verify_hostname is on (its default); on Net::Async::HTTP SSL_verify_mode => 0 (on the client or the request) skips both, SSL_verifycn_scheme => 'none' the name part. A connection that fails the check is not used (a 500 response on the synchronous side, a failed future with category connect_address on Net::Async::HTTP). This matters for connection reuse: Net::Async::HTTP pools connections by address and port, so two engines pinning different names to one address on a shared client would otherwise share a TLS session verified for only one of them. With the check, the request of the engine whose name the pooled connection was not verified for fails (it is not retried on a new connection); give such engines a client each. The refused connection is closed once idle, so later requests get a fresh one. Likewise an LWP conn_cache shared with another agent could hand over a socket that agent opened elsewhere, or without verification; such a socket is refused.

What the pin cannot do is refused rather than silently skipped:

  • With a "user_agent" passed in, it must be a Langertha::HTTP::UserAgent built with the same pin (connect_host => <host of url>, connect_address => ...); anything else croaks at construction, because the synchronous requests go straight to that agent.

  • An injected _async_http client that is neither a Net::Async::HTTP nor the synchronous shim over a correctly pinned agent fails every request to the pinned host (the future fails with ... connect_address ... cannot be applied ...). So does a Net::Async::HTTP client configured with a proxy_host or proxy_path, and on the synchronous side a request that LWP would send through a proxy (a 500 response): a proxy resolves the name itself. Through a SOCKS proxy the connection's peer is the proxy, so the peer check refuses it.

Engines derived from this one carry the pin while their URL is on the same host ("whisper" in Langertha::Engine::OpenAI, Ollama->openai, LMStudio->openai / ->anthropic, also with a url of your own on that host); on another host they get none.

response_max_bytes

A sanity ceiling, in bytes, on the decoded size of a response body. Default 268435456 (256 MiB); 0 removes the cap. Decoding a Content-Encoding (gzip, deflate, bzip2) has no size bound of its own, so a hostile or broken endpoint (a self-hosted /metrics, a gateway) could make a small compressed body expand to gigabytes in memory — a decompression bomb. A body whose decoded size passes this ceiling is refused with <engine class> response body exceeds response_max_bytes (<n>). It is a generous ceiling for genuinely large provider responses, not a tight limit, and an uncompressed body under it is never affected.

Where the bound is applied depends on the backend, because both must be covered:

This covers the non-streaming provider and metrics paths: "parse_response"'s trace and error body, transcription_result, poll_metrics_f, and the non-streaming async_request_f. A true stream (SSE / NDJSON, an on_header request) is unbounded by design and not affected. An encoding the bounded decoder cannot inflate in blocks (br, zstd) is refused rather than decoded unbounded.

generate_json_body

my $body = $engine->generate_json_body(%args);

Encodes %args as a JSON string using the engine's "json" in Langertha::Role::JSON instance. Used internally when building application/json request bodies.

generate_multipart_body

my ( $body, $boundary ) = $engine->generate_multipart_body($request, %args);

Encodes %args as a multipart/form-data body (fields sorted by name) and returns it together with the boundary it used. The boundary is $Langertha::Role::HTTP::boundary unless a part contains it, in which case HTTP::Request::Common picks another one; the Content-Type header must use the returned value ("generate_http_request" does). Used internally when the OpenAPI spec specifies multipart/form-data (e.g. audio upload endpoints).

Values are read as follows:

  • A plain scalar is a text field. It is a character string and is sent UTF-8 encoded, the same as a value in a JSON body.

  • An ArrayRef under a key ending in [] (timestamp_granularities[], include[]) is a multi-valued field: one text part per element, each under the [] key (the OpenAI and Groq form).

  • { repeated => \@values } is a multi-valued field sent as one text part per element under the key exactly as given, with no [] (the form Mistral reads: timestamp_granularities => { repeated => ['segment'] }). The marker is explicit because an ArrayRef under a plain key is a file part.

  • Any other ArrayRef is a file part in the HTTP::Request::Common form_data form: [ $path ], [ $path, $filename, @headers ], or [ undef, $filename, Content => $bytes, @headers ] for in-memory content. The filename defaults to the basename of $path. A decoded (character) filename is sent UTF-8 encoded; an undecoded one is taken as the filesystem's bytes and sent unchanged, as raw UTF-8 in filename="..." (RFC 7578).

generate_http_request

my $request = $engine->generate_http_request(
    $method, $url, $response_call, %args
);

Low-level HTTP request builder. Creates a Langertha::Request::HTTP object with the appropriate headers and body encoding (JSON or multipart). Calls the engine's update_request hook if it exists, allowing engines to inject authentication headers. If the URL contains user:password userinfo, HTTP Basic authentication is set automatically.

_bounded_decoded_content

my $body = $engine->_bounded_decoded_content($http_response);
my $text = $engine->_bounded_decoded_content($http_response, default_charset => 'UTF-8');

Like "decoded_content" in HTTP::Message, but the Content-Encoding step is bounded by "response_max_bytes" (see there), so a decompression bomb is refused instead of inflated into memory. %opt passes to the charset step. Used by "parse_response", transcription_result and poll_metrics_f.

parse_response

my $data = $engine->parse_response($http_response);

Decodes a successful HTTP::Response body as JSON and returns the data structure. On failure croaks with the HTTP status line, and appends the provider's response body (whitespace-collapsed and truncated to $error_body_max_length characters) so the real cause — e.g. a provider JSON error object — is visible in the croak message; when the response sent a Retry-After or retry-after-ms, the status line is followed by (retry after Ns) (the value of "retry_after" in Langertha::RateLimit). The async paths (chat_f, simple_chat_f, chat_stream_realtime_f, chat_with_tools_f) fail with exactly this text on every backend. If the engine supports rate limiting, it records the rate limit headers via _update_rate_limit first, for an error response too, so "rate_limit" in Langertha::Engine::Remote describes the failed response after the croak (and is cleared when the response carried none). A successful response whose body is not JSON croaks with <engine class> response is not valid JSON: <body> (the body shortened the same way).

user_agent_timeout

Optional timeout in seconds for HTTP requests. The synchronous methods get it through the LWP::UserAgent (seconds without activity on the connection); when not set, LWP's own default (180 seconds) applies there.

The _f methods (and "async_request_f" in Langertha::Role::AsyncHTTP) on the Net::Async::HTTP backend apply it as well: a plain request fails after this many seconds in total, a streaming one after this many seconds without a byte (a long stream that keeps delivering is not cut off). The Future then fails with <engine class>: request to <url> timed out after Ns (streaming request ... without data (...) for a stream), the URL without its query string, and the Net::Async::HTTP category (timeout / stall_timeout) as the second failure value. When not set, the async backend has no timeout, as before. The synchronous fallback uses the LWP::UserAgent's timeout; an injected client that is not a Net::Async::HTTP keeps its own.

user_agent_agent

The User-Agent string sent with HTTP requests. Defaults to the engine's class name.

user_agent

The LWP::UserAgent instance used for synchronous HTTP requests (and by the synchronous fallback of the _f methods). Built lazily with user_agent_agent and user_agent_timeout as a Langertha::HTTP::UserAgent, which follows redirects under Langertha::HTTP::Redirect: only GET/HEAD, never from https to http, and to another origin without any credential (every header but the representation ones is dropped, and a credential the request carried in its query is removed from the new URL). http://host to https://host is another origin too: a keyed GET behind an http-to-https redirect arrives without its key and gets a 401, so configure the https URL. A POST is never redirected, even if you add it to requests_redirectable. A refused redirect comes back as the 3xx with a Client-Warning naming the reason. The Net::Async::HTTP backend follows the same policy ("async_request_f" in Langertha::Role::AsyncHTTP).

An agent passed in is used as it is, with its own redirect behaviour: a plain LWP::UserAgent clones the request with every header but Authorization (and keeps even that before LWP 6.83), so an x-api-key or similar header would reach whatever host a server redirects to. Pass a Langertha::HTTP::UserAgent (it takes the same arguments) to keep the policy.

With a "connect_address" the built agent carries the pin (connect_host / connect_address of Langertha::HTTP::UserAgent); an agent passed in must carry the same pin, or construction croaks.

execute_streaming_request

my ($chunks, $timing) = $engine->execute_streaming_request($request, $chunk_callback);
my ($chunks, $timing) = $engine->execute_streaming_request($request);

Executes a streaming HTTP request synchronously using LWP::UserAgent and delegates stream parsing to "process_stream_data" in Langertha::Role::Streaming. Requires the engine to also compose Langertha::Role::Streaming. The response's rate limit headers replace the engine's rate limit first, as in "parse_response". On a non-success response croaks with the HTTP status line and the provider's response body appended (whitespace-collapsed and length-limited), mirroring "parse_response". Returns an ArrayRef of Langertha::Stream::Chunk objects and a timing HashRef with total_seconds (Float, seconds). ttft_seconds is omitted because this method reads the whole body before parsing — use "chat_stream_realtime_f" in Langertha::Role::Chat for true TTFT (it streams incrementally on both backends of Langertha::Role::AsyncHTTP, including the synchronous LWP fallback). If $chunk_callback is provided it is called with each chunk as it is parsed.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.