NAME

Langertha::Role::OpenAICompatible - Role for OpenAI-compatible API format

VERSION

version 0.503

SYNOPSIS

# This role is not used directly - it's composed by engines
# that implement the OpenAI-compatible API format.

package My::Engine;
use Moose;

with map { 'Langertha::Role::'.$_ } qw(
    JSON HTTP OpenAICompatible OpenAPI Models Temperature
    ResponseSize SystemPrompt Streaming Chat Tools
);

sub _build_api_key { $ENV{MY_API_KEY} || die "needs api_key" }
sub default_model { 'my-model' }

__PACKAGE__->meta->make_immutable;

DESCRIPTION

This role provides the OpenAI API format methods for chat completions, embeddings, transcription, streaming, and tool calling. Engines that use the OpenAI-compatible API format (whether OpenAI itself, Ollama's /v1 endpoint, or other compatible providers) can compose this role instead of inheriting from Langertha::Engine::OpenAI.

The role provides default implementations for all OpenAI-format operations. Engines can override individual methods to customize behavior (e.g., different operation IDs for Mistral, or disabling unsupported features).

Engines should also compose these roles:

Engines using this role:

The base classes Langertha::Engine::OpenAIBase and Langertha::Engine::OpenAI also compose this role (and so every engine that extends them inherits it without listing it explicitly).

api_key

Optional API key for Bearer token authentication. Override _build_api_key in engines that require authentication (typically from an environment variable). When undef, no Authorization header is sent.

update_request

$role->update_request($http_request);

Adds Authorization: Bearer {api_key} header to outgoing requests when an API key is configured. Skipped when api_key is undef (e.g. for local servers like vLLM or llama.cpp).

openapi_file

my ($type, $path) = $role->openapi_file;

Returns the OpenAI OpenAPI spec file path used for request generation. Override in an engine to use a provider-specific spec (e.g., Mistral).

list_models_path

my $path = $engine->list_models_path;

Returns the path appended to url for the models endpoint. Default: /models. Override in engines whose API spec uses a different path (e.g. Mistral uses /v1/models because its base URL does not include /v1).

list_models_request

my $request = $engine->list_models_request;
my $request = $engine->list_models_request(after => $last_id);

Generates an HTTP GET request for the models endpoint using list_models_path. Pass after for cursor-based pagination. Returns an HTTP request object.

list_models_response

my $data = $engine->list_models_response($http_response);

Parses the /v1/models response. Returns the full response hashref including data, has_more, and last_id for pagination.

list_models

my $model_ids = $engine->list_models;
# Returns: ['gpt-4o', 'gpt-4o-mini', ...]

my $models = $engine->list_models(full => 1);
# Returns: [{id => 'gpt-4o', created => ..., ...}, ...]

my $fresh = $engine->list_models(force_refresh => 1);

Fetches available models from the /v1/models endpoint with caching. Automatically paginates through all pages using cursor-based pagination (has_more / after). By default returns an ArrayRef of model ID strings. Pass full => 1 for full model objects. Results are cached for models_cache_ttl seconds (default: 3600). Pass force_refresh => 1 to bypass the cache.

embedding_request

my $request = $engine->embedding_request($input, %extra);

Generates an OpenAI-format embedding request for $input: a string, or an ArrayRef of strings for a batch (sent as one input array). Uses embedding_model (default: text-embedding-3-large). %extra goes into the body unchanged (dimensions, encoding_format, ...); "embedding_dimensions" in Langertha::Role::Embedding, when set, is sent under the field the private hook _embedding_dimensions_field names (dimensions by default; Langertha::Engine::Mistral overrides it with output_dimension), unless %extra carries dimensions or that field. An engine whose hook returns undef documents no such field for its embedding_model: the attribute is then not sent and carps once per engine instance, while an explicit dimensions extra still goes out untouched. The request's response parser knows the input shape, so a batch comes back as one vector per input (see "embedding_response"). Returns an HTTP request object.

embedding_response

my $vector  = $engine->embedding_response($http_response);
my $vectors = $engine->embedding_response($http_response, \@inputs);

Parses an OpenAI-format embedding response. The second argument is the request's input; the parser built by "embedding_request" passes it itself. For a string input (or none) it returns the vector of the first input (data[].index 0) as an ArrayRef of floats. For an ArrayRef input it returns an ArrayRef with one vector per input, in input order (sorted by data[].index), and croaks when the number of vectors does not match the number of inputs. A response without a vector (no data array, an empty one, or an entry without an embedding) croaks too, naming the engine and any error in the body; it never returns undef.

A response requested with encoding_format => 'base64' is decoded (little-endian float32), so the result is floats either way; there is no option to get the base64 string back. Parse the HTTP::Response yourself when you need the raw form.

chat_request

my $request = $engine->chat_request($messages, %extra);

Generates an OpenAI-format chat completion request. Includes model, messages, max_tokens, temperature, response_format (if set), and stream => false. Returns an HTTP request object.

_wire_usage

Internal hook. Returns the usage block that goes onto the Langertha::Response of "chat_response" and onto a stream chunk. The default returns the wire block unchanged. An engine that knows what the wire does not say overrides it and returns a copy with a canonical key added, which "from_hash" in Langertha::Usage reads on both paths: Langertha::Engine::OpenRouter adds cost_usd from its bare usage.cost.

chat_response

my $response = $engine->chat_response($http_response);

Parses an OpenAI-format chat completion response. Returns a Langertha::Response object with content, model, finish_reason, usage, created, and raw.

finish_reason is the wire value, with one normalization: a reply that carries tool calls but says stop (gpt-oss on vLLM-style servers, e.g. AKI.IO) reports tool_calls, so it agrees with tool_calls. Every other value, length included, passes through; the wire value stays readable in raw.

message.content may be a list of content chunks instead of a string, as Mistral's reasoning models send it: the text of text chunks becomes content, the text inside thinking chunks becomes thinking (unless reasoning_content / reasoning already filled it), and other chunk types are skipped.

message.refusal (a declined structured-output request, content then null) becomes "refusal" in Langertha::Response.

A body without a choice is not an answer and croaks, naming the engine: with an error object (gateways such as OpenRouter return one in a 200 body) "<engine> response carried an error: <message> (<code>)", otherwise "<engine> response contained no choices". A choice carrying an error object, or a top-level error beside a choice whose finish_reason is error (OpenRouter reports a provider failure this way), croaks the same response carried an error; a finish_reason of error with no error object anywhere croaks "<engine> response ended with finish_reason error". Only choices[0] is read.

transcription_request

my $request = $engine->transcription_request($audio, %extra);

Generates an OpenAI-format transcription request for the given audio (a path, \$bytes or a filehandle; filename in %extra names the upload, see "transcription_file_part" in Langertha::Role::Transcription). Uses transcription_model (default: whisper-1; gpt-transcribe on Langertha::Engine::OpenAI). Returns an HTTP request object.

languages => [ 'de', 'en' ] is sent as repeated languages[] fields. gpt-transcribe takes only that plural field, so for a gpt-transcribe* model a language you pass is sent as languages[] too (merged into languages if both are given); other models get language as given.

gpt-transcribe (and the older gpt-4o-transcribe / gpt-4o-mini-transcribe) answer only response_format => 'json', which is what the API sends when no response_format is given; Langertha never defaults one. verbose_json, srt, vtt and timestamp_granularities[] need a model that supports them, such as whisper-1 on OpenAI or a Whisper server; a response_format you pass is sent as given.

transcription_result

my $result = $engine->transcription_result($http_response);
say $result->{text};
for my $segment ( @{ $result->{segments} // [] } ) { ... }

Parses a transcription response into a HashRef. A JSON answer (response_format json or verbose_json) is returned as decoded, so segments, words, language, duration and usage stay reachable. A plain-text answer (text, srt, vtt) becomes { text => $body }, decoded as UTF-8 unless the response names another charset. Croaks like "parse_response" in Langertha::Role::HTTP on an HTTP error.

transcription_response

my $text = $engine->transcription_response($http_response);

Parses an OpenAI-format transcription response and returns the transcript as a string, for every response_format: the text field of a json or verbose_json answer, the body of a text, srt or vtt answer (the subtitle markup included). Use "transcription_result" to keep segments and word timestamps.

stream_format

my $format = $engine->stream_format;

Returns 'sse' (Server-Sent Events), indicating the streaming format used by OpenAI-compatible APIs. Used by Langertha::Role::Chat to select the correct stream parser.

chat_stream_request

my $request = $engine->chat_stream_request($messages, %extra);

Generates an OpenAI-format streaming chat request (stream => true). Returns an HTTP request object for use with streaming execution.

parse_stream_chunk

my $chunk = $engine->parse_stream_chunk($data, $event, \%state);

Parses a single SSE data payload from an OpenAI-format stream. Returns a Langertha::Stream::Chunk with content, is_final, finish_reason, model, usage, cached_tokens (lifted from usage.prompt_tokens_details.cached_tokens when present), and thinking (the streamed delta.reasoning_content / bare delta.reasoning, guarded !ref). A delta.content that is a list of content chunks is read as in "chat_response": text chunks into content, thinking chunks into thinking. The usage-only frame that stream_options => { include_usage => 1 } adds after the finish chunk (an empty choices list) becomes a content-less chunk that is not is_final and carries usage and cached_tokens; collect the stream's usage with "aggregate_usage" in Langertha::Role::Chat. Returns undef only when the payload carries neither a choice nor a usage block. A frame with a top-level error object and no choice (a gateway failing mid-stream) croaks "<engine> stream carried an error: <message> (<code>)", which fails the stream; so does a choice carrying an error object, or a top-level error beside a choice with finish_reason error (OpenRouter's mid-stream failure frame). A choice with finish_reason error and no error object anywhere croaks "<engine> stream ended with finish_reason error". A delta.refusal fragment lands on the chunk's refusal.

delta.tool_calls fragments are assembled per index (a fragment without index by its id, and by its position only when it has neither) in \%state, and the finished calls land as Langertha::ToolCall objects, in stream order, on the chunk that carries a non-empty finish_reason, read by the same "extract" in Langertha::ToolCall as "chat_response". Collect them with "aggregate_tool_calls" in Langertha::Role::Chat. finish_reason is read as on the non-streaming path: the wire value, except that stop on the chunk that delivers tool calls reports tool_calls (the wire value stays in raw). A stream that ends without one drops its pending calls with a carp (see "_finish_stream_state").

\%state is one HashRef per stream. The stream paths pass it; $event is only set by "process_stream_data" in Langertha::Role::Streaming, the chat_stream_realtime_f path passes undef. A direct caller may omit \%state and share the engine's fallback, which is closed when a _process_stream_buffer flush with $final set ends the stream; a caller feeding events to parse_stream_chunk one by one should pass its own state.

_finish_stream_state

$engine->_finish_stream_state(\%state);

Internal: called once when a stream ends (by "process_stream_data" in Langertha::Role::Streaming and "chat_stream_realtime_f" in Langertha::Role::Chat, and by a final _process_stream_buffer flush that was given no state). Tool calls still pending because no chunk carried a finish_reason -- a truncated stream -- are dropped with one carp naming them, not flushed: their arguments may be cut off, and a partial JSON string would decode to {}. Clearing them also keeps them out of the next stream that shares the same state.

image_request

my $request = $engine->image_request($prompt, %extra);

Generates an OpenAI-format image generation request for the given $prompt. Uses image_model (default: gpt-image-2). Accepts optional model, size, quality and n via %extra, passed through as given. Returns an HTTP request object.

GPT image models (gpt-image-*) always return the image as b64_json and reject response_format, so a response_format in %extra is dropped with a warning for them.

image_response

my $images = $engine->image_response($http_response);

Parses an OpenAI-format image generation response. Returns an ArrayRef of image objects, each with url or b64_json (GPT image models answer b64_json only) and optionally revised_prompt. Croaks, naming the engine and any error in the body, when the response carries no image.

simple_image

my $images = $engine->simple_image('A cat in space');

Sends an image generation request and returns the result. Blocks until the request completes. Returns an ArrayRef of image objects.

_parse_rate_limit_headers

Parses x-ratelimit-* headers from the HTTP response into a Langertha::RateLimit object. Covers OpenAI, Groq, Cerebras, OpenRouter, Replicate, and all other OpenAI-compatible engines. Collects the full raw superset via "_collect_headers" in Langertha::RateLimit, then normalizes the Go time.Duration reset strings into "requests_reset_after" in Langertha::RateLimit / "tokens_reset_after" in Langertha::RateLimit (seconds); the *_reset_at instants are derived lazily against "received" in Langertha::RateLimit.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.