NAME

Langertha::Content::Image - Canonical image content block with cross-provider conversion

VERSION

version 0.503

SYNOPSIS

use Langertha::Content::Image;

# From a remote URL
my $img = Langertha::Content::Image->from_url('https://example.com/cat.jpg');

# From a local file (media_type sniffed from extension)
my $img = Langertha::Content::Image->from_file('/tmp/cat.png');

# From raw bytes
my $img = Langertha::Content::Image->from_data($bytes, media_type => 'image/jpeg');

# From an existing base64 string
my $img = Langertha::Content::Image->from_base64($b64, media_type => 'image/png');

# Embed in a chat message — Langertha::Role::Chat converts per engine
my $response = $engine->simple_chat_f({
    role    => 'user',
    content => [ 'What is in this image?', $img ],
});

DESCRIPTION

Provider-neutral image block. Carries either a remote URL, a base64 payload, or both, plus an IANA media_type. Serializes to these vision-chat wire formats:

  • OpenAI chat completions — { type = 'image_url', image_url => { url => ... } }>

  • Anthropic messages — { type = 'image', source => { type => 'url' | 'base64', ... } }>

  • Google Gemini — { inline_data = { mime_type => ..., data => <base64> } }>

  • Open-Responses — { type = 'input_image', image_url => <url or data: URL> }>

  • Ollama native — the raw base64 string, for the message images array

  • LM Studio native — { type = 'image', data_url => <data: URL> }>

The optional "detail" hint goes out only on the two wires that have the field (OpenAI chat completions and Open-Responses) and is ignored elsewhere. "TO_JSON" gives a compact description for logs and traces, never the payload.

Gemini, Ollama native and LM Studio native require inline data, so their serializers transparently download a remote URL on first call (cached on the object). Engines whose OpenAI-compatible endpoint rejects remote image URLs get the same treatment through to_openai( inline => 1 ). On the _f methods of Langertha::Role::Chat the download happens earlier, through the engine's async HTTP backend ("ensure_base64_f"), so the serializers find the payload cached and do not block the event loop.

url

Remote HTTP(S) URL of the image. May be passed through directly (OpenAI, Anthropic) or auto-downloaded and base64-encoded (Gemini).

base64

The base64-encoded image payload (no data: URL prefix). Can be supplied at construction, or populated lazily when a provider that requires inline data (Gemini) is targeted.

media_type

IANA media type (image/jpeg, image/png, image/gif, image/webp). Required for base64 payloads on Anthropic and Gemini. Sniffed from the file extension by from_file and from the URL path by from_url.

detail

Optional image-detail hint, a non-empty string. The known values are low, high and auto; any other value is sent unchanged, and the provider decides whether it takes it. Unset by default, and then no detail field goes on any wire. When set, "to_openai" sends it as image_url.detail and "to_responses" as input_image.detail; the other serializers ignore it, because their wires have no such field. Every from_* constructor takes it:

my $img = Langertha::Content::Image->from_url($url, detail => 'low');

from_url

my $img = Langertha::Content::Image->from_url($url);
my $img = Langertha::Content::Image->from_url($url, media_type => 'image/jpeg');

Builds an image block referencing a remote URL. Media type is sniffed from the URL extension when not provided.

Only http, https and data: URLs are accepted. Any other scheme (file, ftp, gopher, ...) croaks here, before any I/O, because the engines that have to inline images would otherwise read it (a file: URL from a caller's message would send a server-local file to the provider). For a local file use "from_file". See "ensure_base64" for what the fetch does not protect against.

from_file

my $img = Langertha::Content::Image->from_file('/tmp/cat.png');

Reads a local file, base64-encodes it, and sniffs the media type from the extension (unless media_type is passed).

from_data

my $img = Langertha::Content::Image->from_data($bytes, media_type => 'image/jpeg');

Builds an image block from raw bytes. media_type is required.

from_base64

my $img = Langertha::Content::Image->from_base64($b64, media_type => 'image/png');

Builds an image block from an existing base64 string.

ensure_base64

my $b64 = $img->ensure_base64;
my $b64 = $img->ensure_base64( timeout => 5 );
my $b64 = $img->ensure_base64(
  max_bytes  => 5_000_000,
  url_filter => Langertha::Content::Image->deny_private_hosts,
);

Returns the base64 payload, fetching the URL over HTTP if necessary. Populates media_type from the response Content-Type header when the image was URL-only. Caches the result on the object.

Only http and https URLs are fetched; a data: URL is decoded locally. Any other scheme croaks before any I/O:

Langertha::Content::Image refuses to fetch image URL with scheme 'file'
(only http/https; use Content::Image->from_file for local files)

A redirect to another scheme is not followed, and a fetch that ended on one anyway is not stored.

max_bytes caps the download, 20971520 (20 MiB) by default; 0 means no cap. A Content-Length over the cap stops the fetch before the body, and a body that grows past it stops reading there. The cap holds for the decoded size too: a body sent with a Content-Encoding (gzip, deflate, bzip2) is inflated in blocks and refused once it passes the cap, and one in any other encoding is refused as undecodable while a cap is set. Nothing is stored, and the call croaks:

Langertha::Content::Image image at https://... exceeds inline_image_max_bytes (20971520)

url_filter is a code reference that decides which URLs may be fetched, against server-side request forgery (SSRF) through image URLs from untrusted input. It gets each URL as a URI object and returns true to allow it. It runs on the image URL before any I/O, and again on every redirect hop before that hop is requested. Without it, the default, every http and https host is fetched. A refused URL, or a filter that dies, croaks:

Langertha::Content::Image refuses to fetch image URL http://10.0.0.5/x.png:
rejected by inline_image_url_filter

"deny_private_hosts" builds a ready-made filter. Langertha::Role::Chat passes the engine's "inline_image_max_bytes" in Langertha::Role::Chat and "inline_image_url_filter" in Langertha::Role::Chat as these two options.

The fetch is a blocking LWP::UserAgent GET that gives up after timeout seconds of inactivity, 30 by default. When Langertha::Role::Chat builds a request it passes the engine's "inline_image_fetch_timeout" in Langertha::Role::Chat. timeout => 0 leaves LWP's own default (180 seconds) in place, because LWP cannot run without a timeout.

ensure_base64_f

my $b64 = await $img->ensure_base64_f($http);
my $b64 = await $img->ensure_base64_f( $http, max_bytes => ..., url_filter => ... );

The async "ensure_base64": returns a Future of the base64 payload and fetches the URL through $http, any client that answers the async do_request contract (Langertha::Role::AsyncHTTP), instead of a blocking LWP::UserAgent. A transport error or a non-success status fails the Future. The scheme rules of "ensure_base64" apply unchanged: a non-HTTP(S) URL fails the Future before any request, a data: URL is decoded locally, and on the synchronous LWP fallback (Langertha::Request::SyncHTTP) the fetch runs over a copy of its user agent restricted to http and https.

max_bytes and url_filter work as on "ensure_base64" and fail the Future with the same text. On Net::Async::HTTP the body is counted as it arrives and the request is abandoned (its connection closed) once it passes the cap, and redirects (up to 7) are followed one hop at a time, so the filter sees each hop before it is requested. On the LWP fallback the copy of its user agent carries the cap and the filter. Any other client follows redirects on its own: there the filter sees the hops only after the fetch, and the payload of a fetch with a refused hop is not stored; the cap is checked on the finished body. The _f methods of Langertha::Role::Chat call it with the engine's backend for every URL image the engine has to inline, before the request is built.

deny_private_hosts

my $engine = Langertha::Engine::Gemini->new(
  api_key                 => $key,
  inline_image_url_filter => Langertha::Content::Image->deny_private_hosts,
);

# With a resolver of your own (tests, a caching resolver):
my $filter = Langertha::Content::Image->deny_private_hosts(
  resolver => sub { my ($host) = @_; return @ip_addresses },
);

Returns a URL filter for "ensure_base64"'s url_filter (and the engine's "inline_image_url_filter" in Langertha::Role::Chat) that refuses hosts on internal networks. It resolves the URL's host and refuses it when any address it resolves to is one of:

  • IPv4 0.0.0.0/8 (this host), 127.0.0.0/8 (loopback), 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16 (RFC 1918), 100.64.0.0/10 (carrier-grade NAT), 169.254.0.0/16 (link-local, including the cloud metadata address 169.254.169.254), 192.0.0.0/24 (IETF protocol assignments, including the NAT64 discovery addresses), 198.18.0.0/15 (benchmarking), 224.0.0.0/3 (multicast, 240.0.0.0/4 reserved, 255.255.255.255 broadcast)

  • IPv6 ::/96 (unspecified, loopback ::1, IPv4-compatible), fe80::/10 (link-local), fec0::/10 (site-local), fc00::/7 (unique local, including the AWS metadata address fd00:ec2::254), ff00::/8 (multicast)

  • an IPv6 address that carries an IPv4 address in the list above: IPv4-mapped ::ffff:a.b.c.d, SIIT ::ffff:0:a.b.c.d, NAT64 64:ff9b::a.b.c.d and 64:ff9b:1::a.b.c.d, and 6to4 2002:AABB:CCDD::/48. Any other address in the local-use NAT64 prefix 64:ff9b:1::/48 is refused, since it does not say where its IPv4 address sits.

  • Teredo 2001:0000::/32 (RFC 4380): its client IPv4 is embedded obfuscated (bit-inverted) in the low 32 bits, so unlike 6to4 and NAT64 it cannot be re-checked as a plain address — the whole prefix is refused outright. A non-Teredo 2001::/16 address (2001:db8::, 2001:4860::, ...) is not matched and passes.

A host that does not resolve, and a URL without a host, are refused too.

The default resolver is the system's getaddrinfo (Socket); an IP literal is not looked up. The lookup blocks, also on the _f paths, where it holds the event loop for its duration. resolver replaces it: a code reference that gets the host name and returns its addresses as strings.

DNS rebinding: the host is resolved here, before the HTTP client connects, and the client resolves it again on its own. A name whose DNS answer changes between the two lookups (a short TTL pointing first at a public address, then at 127.0.0.1) passes the filter and still reaches the internal address. The filter stops URLs that name or resolve to an internal address; for a guarantee against rebinding, fetch through an egress proxy or firewall that enforces the same rule on the connection itself.

data_url

my $uri = $img->data_url;   # data:image/png;base64,...

Returns the image as a data: URL, fetching a URL-only image first (see "ensure_base64"). The media type falls back to application/octet-stream.

to_openai

my $block = $img->to_openai;
# { type => 'image_url', image_url => { url => ... } }
my $block = $img->to_openai( inline => 1 );   # always a data: URL

Serializes to the OpenAI chat-completions image block. Uses the URL when available, otherwise emits a data: URL from the base64 payload. With inline => 1 it always emits the data: URL, fetching a URL-only image first; Langertha::Role::Chat passes it for engines whose endpoint rejects remote image URLs. A set "detail" goes out as image_url.detail.

to_responses

my $block = $img->to_responses;
# { type => 'input_image', image_url => 'https://...' }   (or a data: URL)

Serializes to the Open-Responses input_image part (OpenAI /v1/responses, Perplexity /v1/agent). image_url is a plain string, not an object: the URL when available, otherwise a data: URL. Takes inline => 1 like "to_openai". A set "detail" goes out as the part's detail field.

to_ollama

my $b64 = $img->to_ollama;

Returns the raw base64 payload (no data: prefix) for one entry of the Ollama native /api/chat message images array. Fetches a URL-only image first, because that wire takes no image URLs.

to_lmstudio

my $item = $img->to_lmstudio;
# { type => 'image', data_url => 'data:image/png;base64,...' }

Serializes to an LM Studio native /api/v1/chat input image item. That wire takes only base64 data URLs, so a URL-only image is fetched first.

to_anthropic

my $block = $img->to_anthropic;
# { type => 'image', source => { type => 'url', url => ... } }
# or
# { type => 'image', source => { type => 'base64', media_type => ..., data => ... } }

Serializes to the Anthropic messages image block. Prefers a URL source when available; otherwise emits an inline base64 source (media_type required).

to_gemini

my $block = $img->to_gemini;
# { inline_data => { mime_type => ..., data => <base64> } }

Serializes to the Gemini inlineData part. Auto-downloads URL-only images because Gemini has no URL-fetching equivalent.

TO_JSON

my $json = JSON::MaybeXS->new( convert_blessed => 1 )
  ->encode([ { role => 'user', content => [ 'What is this?', $img ] } ]);
# ... {"bytes":48213,"media_type":"image/png","source":"base64","type":"image"} ...

Serialization hook for JSON encoders configured with convert_blessed (the engine's own "json" in Langertha::Role::JSON is one), so a message array holding images can be written to a log or a trace. Returns a compact description, never the image data:

{ type       => 'image',
  source     => 'url' | 'base64',
  url        => ...,     # only for a URL image
  media_type => ...,     # when known
  detail     => ...,     # when set
  bytes      => ... }    # decoded payload size, when base64 is present

source is url for an image built from a URL (even after "ensure_base64" has fetched it; bytes then gives the fetched size), and base64 otherwise. A data: URL passed as url is reported as base64 without the URL, because it is the payload. This is not a wire format: request bodies are built by the to_* serializers, never through TO_JSON.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.