NAME
Langertha::Engine::vLLM - vLLM inference server
VERSION
version 0.503
SYNOPSIS
use Langertha::Engine::vLLM;
# 1. Simple chat
my $vllm = Langertha::Engine::vLLM->new(
url => 'http://localhost:8000/v1',
system_prompt => 'You are a helpful assistant',
);
print $vllm->simple_chat('Say something nice');
# 2. Streaming
$vllm->simple_chat_stream(sub {
print shift->content;
}, 'Write a haiku about Perl');
# 3. MCP tool calling (requires server started with tool-call-parser)
use Future::AsyncAwait;
my $vllm = Langertha::Engine::vLLM->new(
url => 'http://localhost:8000/v1',
model => 'Qwen/Qwen2.5-3B-Instruct',
mcp_servers => [$mcp],
);
my $response = await $vllm->chat_with_tools_f('Add 7 and 15');
# 4. Multimodal input (vision-capable models: LLaVA, Qwen-VL, PaliGemma)
use Langertha::Content::Image;
my $img = Langertha::Content::Image->from_url('https://example.com/cat.jpg');
my $resp = await $vllm->simple_chat_f({
role => 'user',
content => [ 'What is in this image?', $img ],
});
# 5. Reasoning models (Qwen3, DeepSeek-R1, QwQ — server needs
# --reasoning-parser matching the model; reasoning_effort maps to
# enable_thinking via the chat template). simple_chat takes messages
# only — set the control on the engine, or hand it to chat_f.
my $thinker = Langertha::Engine::vLLM->new(
url => 'http://localhost:8000/v1',
reasoning_effort => 'high',
);
print $thinker->simple_chat('Solve step by step: what is 7 factorial?');
# … or as a per-request control:
my $reasoned = await $vllm->chat_f(
messages => ['Solve step by step: what is 7 factorial?'],
reasoning_effort => 'high',
);
# 6. Embeddings (for embedding models: BAAI/bge-*, intfloat/e5-*, …)
my $vector = $vllm->simple_embedding('Some text to embed');
# 7. Prometheus /metrics scraping (Runtime::MetricsPoll)
my $records = await $vllm->poll_metrics_f('vllm:');
# Returns ArrayRef of { name, type, value, labels } parsed from
# GET <url-stripped-of-/v1>/metrics.
# 8. vLLM-Hook sub-engine (IBM vLLM-Hook plugin: hidden states, qk, steer)
use Langertha::Engine::VLLMHook;
my $hooked = Langertha::Engine::VLLMHook->new(
url => 'http://localhost:8770/v1',
vllm_xargs => { output_hidden_states => JSON::MaybeXS::true() },
);
DESCRIPTION
Provides access to vLLM, a high-throughput inference engine for large language models. Extends Langertha::Engine::OpenAIBase (which composes Langertha::Role::OpenAICompatible, Langertha::Role::OpenAPI, Langertha::Role::Models, Langertha::Role::Temperature, Langertha::Role::ResponseSize, Langertha::Role::SystemPrompt, Langertha::Role::ResponseFormat, Langertha::Role::Streaming, Langertha::Role::Chat, Langertha::Role::ReasoningEffort, and Langertha::Role::PromptCache); vLLM itself additionally composes Langertha::Role::Tools (MCP tool calling), Langertha::Role::Embedding (OpenAI-compatible /v1/embeddings for embedding models), and Langertha::Role::Runtime::MetricsPoll (Prometheus /metrics scrape).
Supports chat, streaming, tool calling, embeddings, multimodal input, reasoning models, and Prometheus /metrics scraping.
Only url is required. The URL must include the /v1 path prefix (e.g., http://localhost:8000/v1). Since vLLM serves exactly one model (configured at server startup), no model name or API key is needed.
TOOL CALLING
MCP tool calling requires the vLLM server to be started with --enable-auto-tool-choice and --tool-call-parser matching the model (hermes for Qwen2.5/Hermes, llama3 for Llama, mistral for Mistral).
EMBEDDINGS
Composes Langertha::Role::Embedding. vLLM exposes an OpenAI-compatible /v1/embeddings endpoint when started with an embedding model (BAAI/bge-*, intfloat/e5-*, …). The request carries embedding_model if you set it, else model if you set it, else no model field at all: the server embeds with the model it serves. (vLLM 0.10/0.11 answer 404 to a model that is not a served name, such as the 'default' placeholder.)
REASONING MODELS
vLLM serves reasoning models through the OpenAI-compatible chat completions endpoint. reasoning_effort (Langertha::Role::ReasoningEffort, Langertha::Reasoning, wire format openai) is honored on the body — vLLM translates the value into enable_thinking via the chat template:
reasoning_effort low|medium|high -> enable_thinking = true
reasoning_effort none -> enable_thinking = false
(unset) -> enable_thinking not injected
The thought itself comes back in the reasoning response field — mainline vLLM renamed it from reasoning_content to reasoning in v0.16.0, and both spellings are read — surfaced on $response->thinking. It is only populated when the server was started with the matching --reasoning-parser.
Supported reasoning-model families (vLLM's own table — verify against https://docs.vllm.ai/en/latest/features/reasoning_outputs/ for the exact list as new models land):
Qwen3 (
Qwen3-Instruct/Qwen3-Thinking) — wirereasoning_efforthonored.DeepSeek-R1 (native) and DeepSeek-R1-Distill-* — wire
reasoning_efforthonored.QwQ-32B — wire
reasoning_efforthonored (served via thedeepseek_r1reasoning parser).Gemma 4, IBM Granite 3.2, DeepSeek-V3.1 / V4 — also honor the parameter via the same chat-template auto-injection path.
The server MUST be started with the matching --reasoning-parser flag for the loaded model (deepseek_r1 for DeepSeek-R1 / QwQ, qwen3 for Qwen3 thinking, etc.) — without it, the model emits reasoning content but Langertha cannot surface it on Langertha::Response. See vLLM's reasoning-outputs docs for the canonical flag list.
Per ADR 0009 the wire format is openai on vLLM; engines diverging (Anthropic, Gemini) carry their own reasoning_wire_format tag and serialize independently.
METRICS
Composes Langertha::Role::Runtime::MetricsPoll. vLLM serves Prometheus text-format metrics at GET <url-stripped-of-/v1/metrics>; pass a prefix to "poll_metrics_f" (e.g. 'vllm:' for the engine's own gauges) or filter downstream via "filter_prefix" in Langertha::Runtime::Metrics.
SUB-ENGINE: VLLMHook
Langertha::Engine::VLLMHook extends this engine for servers running the IBM vLLM-Hook plugin (hidden_states / qk / steer probes); it injects a top-level vllm_xargs field and lifts captured tensors onto "probes" in Langertha::Response. See ADR 0004 for the seam.
See https://docs.vllm.ai/ for installation and configuration details.
CAPABILITIES
Advertised flags (derived from composed roles via Langertha::Role::Capabilities):
chat— Langertha::Role::Chatstreaming— Langertha::Role::Streamingtools_native+tool_choice_{auto,any,none,named}— Langertha::Role::Toolsembedding— Langertha::Role::Embeddingruntime_metrics— Langertha::Role::Runtime::MetricsPollresponse_format_{json_object,json_schema}— Langertha::Role::ResponseFormattemperature— Langertha::Role::Temperaturereasoning_effort— Langertha::Role::ReasoningEffort; honored by Qwen3 / DeepSeek-R1 / QwQ / Gemma 4 / Granite 3.2 when vLLM is started with--reasoning-parser. See "REASONING MODELS".response_size,system_prompt,parallel_tool_use,context_size,seed— generation-parameter knobs the engine will honour
SEE ALSO
https://docs.vllm.ai/ - vLLM documentation
Langertha::Engine::OpenAIBase - Base class for OpenAI-compatible engines
Langertha::Role::OpenAICompatible - OpenAI API format role
Langertha::Role::Tools - MCP tool calling interface
Langertha::Role::Embedding - Embedding request/response interface
Langertha::Role::Runtime::MetricsPoll - Prometheus /metrics scraper
Langertha::Engine::VLLMHook - Sub-engine for the IBM vLLM-Hook plugin
Langertha::Engine::LlamaCpp - Sister self-hosted OpenAI-compatible engine (also Embedding + MetricsPoll)
Langertha::Engine::SGLang - Sister self-hosted OpenAI-compatible engine (also MetricsPoll)
Langertha::Engine::OllamaOpenAI - Ollama's OpenAI-compatible endpoint
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.