NAME
Langertha::Role::ImageInput - Role for an engine whose wire can carry image input
VERSION
version 0.503
SYNOPSIS
use Langertha::Content::Image;
if ( $engine->supports('image_input') ) {
my $img = Langertha::Content::Image->from_url('https://example.com/cat.jpg');
my $r = $engine->simple_chat({
role => 'user', content => [ 'What is in this image?', $img ],
});
}
DESCRIPTION
A capability role (ADR 0016): an engine composes it when its wire can carry Langertha::Content::Image parts, i.e. its "content_format" in Langertha::Role::Chat serializes an image into a shape the endpoint accepts. Composing it contributes the image_input flag to "engine_capabilities" in Langertha::Role::Capabilities.
image_input is model-scoped and means the model sees the image, not merely that the wire accepts the part (ADR 0019, k266 Update). Most providers serve text-only and vision models side by side, so an engine that composes this role refines the flag per model:
all-vision families (OpenAI, first-party Anthropic, Gemini, Hetzner) keep the flag and clear it for the listed text-only models (
model_capability_corrections);other cloud engines clear it for every model and re-assert it only for the documented vision models -- including the
/anthropicshims of MiniMax and Moonshot, which carry the same rows as their OpenAI faces;gateways, self-hosted servers, AKIAnthropic and LMStudioAnthropic clear it for every model: the model behind them is unknown to the client, or its vision is unverified on that face, so the engine makes no static claim.
The static answer can be replaced by what the provider says about its own models: "probe_model_capabilities_f" in Langertha::Role::Capabilities reads the metadata endpoint of OpenRouter, Mistral, Ollama, OllamaOpenAI, LMStudio, LMStudioOpenAI, LMStudioAnthropic, LlamaCpp and TSystems and stores image_input per model on the engine instance (ADR 0032). Nothing probes implicitly.
The flag is advisory. Nothing blocks or strips an image when it is false; an image sent to an engine without the claim goes out on the wire as usual and the provider decides.
One reader chooses a representation by it: in the tool loop, "format_tool_results" in Langertha::Role::Tools sends an image a tool returned as an image part on the responses, Gemini 3 and anthropic wires only when the flag is true, and as a text placeholder otherwise (see "DESCRIPTION" in Langertha::ToolResult). A PDF a tool returned follows the same flag on OpenAI Responses and Gemini 3, which read PDFs through the model's vision. So a claim for a model that does not see images is no longer harmless there: the provider may reject the tool-loop turn, or (as AKI.IO's /anthropic shim does) accept it while the model never sees the image.
This role has no methods or attributes of its own.
SEE ALSO
Langertha::Content::Image - Provider-agnostic image input
Langertha::Role::Capabilities - The capability registry
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.