NAME

Langertha::Engine::Hetzner - Hetzner Inference API (OpenAI-compatible)

VERSION

version 0.503

SYNOPSIS

use Langertha::Engine::Hetzner;

my $hetzner = Langertha::Engine::Hetzner->new(
    api_key => $ENV{LANGERTHA_HETZNER_API_KEY},
);

print $hetzner->simple_chat('Hello from Perl!');

# Streaming
$hetzner->simple_chat_stream(sub {
    print shift->content;
}, 'Write a poem');

# Vision (Qwen/Qwen3.6-35B-A3B-FP8 accepts image_url content parts)
use Langertha::Content::Image;
my $img = Langertha::Content::Image->from_url('https://example.com/cat.jpg');
my $resp = await $hetzner->simple_chat_f({
    role    => 'user',
    content => [ 'What is in this image?', $img ],
});

# Tool calling
my $response = await $hetzner->chat_with_tools_f('Search for Perl modules');

DESCRIPTION

Provides access to Hetzner Cloud's Inference API via their OpenAI-compatible endpoint at https://inference.hetzner.com/api/v1.

Hetzner's Inference API is currently experimental and free of charge; rate limits are 10M input / 200K output tokens per 60 seconds per API key (HTTP 429 when exceeded). Bearer-token authentication via LANGERTHA_HETZNER_API_KEY.

Supports chat, streaming, tool calling, structured output (OpenAI-compatible response_format), and image inputs (image_url content parts) on the vision-capable models. Embeddings and transcription are not available on this endpoint.

DEFAULT MODEL

Qwen/Qwen3.6-35B-A3B-FP8 — one of the two currently-listed Hetzner models, an MoE (35B total / 3B activated), Apache 2.0, text + image input, 262K context window. Picked because it is distinct from the existing Langertha::Engine::Moonshot default (Kimi K3); the catalog's only other entry (Qwen3.8-27B, a dense 27B model) is likewise text + image, so the choice is a capacity/architecture preference rather than a modality one.

MODELS

The two models currently listed at /api/v1/models (Hetzner narrowed the Inference experiment to the small Qwen models in Aug 2026, retiring the earlier DeepSeek/GLM/Kimi checkpoints):

  • Qwen/Qwen3.6-35B-A3B-FP8 — default. Apache 2.0. MoE 35B/3B. 262K context. Text + image input.

  • Qwen3.8-27B — dense 27B. 262K context. Text + image input.

Tool support caveat: the Hetzner Inference docs do not confirm server-side tool calling or structured output on the OpenAI-compatible endpoint, and the platform is explicitly experimental. The engine composes Langertha::Role::Tools (so chat_with_tools_f exists) and Langertha::Role::ResponseFormat, but engine_capabilities clears tools_native, every tool_choice_*, parallel_tool_use and both response_format_* flags rather than advertise capabilities that may silently no-op. Tools you pass still go out in the tools array, but tool_choice and parallel_tool_calls are no longer sent: auto is dropped silently, a forced choice with a warning, and none leaves the tools out instead (see "chat_f" in Langertha::Role::Chat). If a live test confirms the gateway honors them for the model you use, re-add them in the engine's around engine_capabilities.

No embeddings or transcription: the Hetzner Inference endpoint exposes chat completions + image processing only. "embedding" and "transcription" are not composed on this engine.

Get your API key at https://inference.hetzner.com/ and set LANGERTHA_HETZNER_API_KEY in your environment.

REFRESHING THE MODEL CATALOG

To add or remove a model: edit _build_static_models below, update the "MODELS" POD block to match, add a Changes entry, and run the offline tests (t/48_hetzner.t, t/00_load.t). The live drift check in t/88_live_hetzner.t (karr #40) compares the hardcoded catalog against https://inference.hetzner.com/api/v1/models and warns on missing models.

CAPABILITIES

Advertised flags (derived from composed roles via Langertha::Role::Capabilities):

tools_native, tool_choice_*, parallel_tool_use and response_format_* are not advertised even though Langertha::Role::Tools and Langertha::Role::ResponseFormat are composed — the engine clears them in its around engine_capabilities (see the tool-support caveat above). Tools are still sent; tool_choice and parallel_tool_calls are not. If a live test confirms them, re-add them there.

Vision input is supported on both currently-listed models (Qwen/Qwen3.6-35B-A3B-FP8 and Qwen3.8-27B) via image_url content parts; this is handled by Langertha::Content::Image and Langertha::Role::Chat's normalization — there is no engine-level vision flag.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.