NAME
Langertha::Engine::Hetzner - Hetzner Inference API (OpenAI-compatible)
VERSION
version 0.503
SYNOPSIS
use Langertha::Engine::Hetzner;
my $hetzner = Langertha::Engine::Hetzner->new(
api_key => $ENV{LANGERTHA_HETZNER_API_KEY},
);
print $hetzner->simple_chat('Hello from Perl!');
# Streaming
$hetzner->simple_chat_stream(sub {
print shift->content;
}, 'Write a poem');
# Vision (Qwen/Qwen3.6-35B-A3B-FP8 accepts image_url content parts)
use Langertha::Content::Image;
my $img = Langertha::Content::Image->from_url('https://example.com/cat.jpg');
my $resp = await $hetzner->simple_chat_f({
role => 'user',
content => [ 'What is in this image?', $img ],
});
# Tool calling
my $response = await $hetzner->chat_with_tools_f('Search for Perl modules');
DESCRIPTION
Provides access to Hetzner Cloud's Inference API via their OpenAI-compatible endpoint at https://inference.hetzner.com/api/v1.
Hetzner's Inference API is currently experimental and free of charge; rate limits are 10M input / 200K output tokens per 60 seconds per API key (HTTP 429 when exceeded). Bearer-token authentication via LANGERTHA_HETZNER_API_KEY.
Supports chat, streaming, tool calling, structured output (OpenAI-compatible response_format), and image inputs (image_url content parts) on the vision-capable models. Embeddings and transcription are not available on this endpoint.
DEFAULT MODEL
Qwen/Qwen3.6-35B-A3B-FP8 — one of the two currently-listed Hetzner models, an MoE (35B total / 3B activated), Apache 2.0, text + image input, 262K context window. Picked because it is distinct from the existing Langertha::Engine::Moonshot default (Kimi K3); the catalog's only other entry (Qwen3.8-27B, a dense 27B model) is likewise text + image, so the choice is a capacity/architecture preference rather than a modality one.
MODELS
The two models currently listed at /api/v1/models (Hetzner narrowed the Inference experiment to the small Qwen models in Aug 2026, retiring the earlier DeepSeek/GLM/Kimi checkpoints):
Qwen/Qwen3.6-35B-A3B-FP8—default. Apache 2.0. MoE 35B/3B. 262K context. Text + image input.Qwen3.8-27B— dense 27B. 262K context. Text + image input.
Tool support caveat: the Hetzner Inference docs do not confirm server-side tool calling or structured output on the OpenAI-compatible endpoint, and the platform is explicitly experimental. The engine composes Langertha::Role::Tools (so chat_with_tools_f exists) and Langertha::Role::ResponseFormat, but engine_capabilities clears tools_native, every tool_choice_*, parallel_tool_use and both response_format_* flags rather than advertise capabilities that may silently no-op. Tools you pass still go out in the tools array, but tool_choice and parallel_tool_calls are no longer sent: auto is dropped silently, a forced choice with a warning, and none leaves the tools out instead (see "chat_f" in Langertha::Role::Chat). If a live test confirms the gateway honors them for the model you use, re-add them in the engine's around engine_capabilities.
No embeddings or transcription: the Hetzner Inference endpoint exposes chat completions + image processing only. "embedding" and "transcription" are not composed on this engine.
Get your API key at https://inference.hetzner.com/ and set LANGERTHA_HETZNER_API_KEY in your environment.
REFRESHING THE MODEL CATALOG
To add or remove a model: edit _build_static_models below, update the "MODELS" POD block to match, add a Changes entry, and run the offline tests (t/48_hetzner.t, t/00_load.t). The live drift check in t/88_live_hetzner.t (karr #40) compares the hardcoded catalog against https://inference.hetzner.com/api/v1/models and warns on missing models.
CAPABILITIES
Advertised flags (derived from composed roles via Langertha::Role::Capabilities):
chat— Langertha::Role::Chatstreaming— Langertha::Role::Streamingtemperature— Langertha::Role::Temperatureresponse_size,system_prompt,context_size,seed— generation-parameter knobs the engine will honour
tools_native, tool_choice_*, parallel_tool_use and response_format_* are not advertised even though Langertha::Role::Tools and Langertha::Role::ResponseFormat are composed — the engine clears them in its around engine_capabilities (see the tool-support caveat above). Tools are still sent; tool_choice and parallel_tool_calls are not. If a live test confirms them, re-add them there.
Vision input is supported on both currently-listed models (Qwen/Qwen3.6-35B-A3B-FP8 and Qwen3.8-27B) via image_url content parts; this is handled by Langertha::Content::Image and Langertha::Role::Chat's normalization — there is no engine-level vision flag.
SEE ALSO
Langertha::Engine::Moonshot - Another OpenAI-compatible cloud engine with multimodal support
Langertha::Engine::XAI - Another OpenAI-compatible cloud engine with vision + tool calling
https://inference.hetzner.com/ - Hetzner Inference API
Langertha::Engine::OpenAIBase - Base class for OpenAI-compatible engines
Langertha::Role::Tools - MCP tool calling interface
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.