NAME

Langertha::Engine::vLLM - vLLM inference server

VERSION

version 0.503

SYNOPSIS

use Langertha::Engine::vLLM;

# 1. Simple chat
my $vllm = Langertha::Engine::vLLM->new(
    url           => 'http://localhost:8000/v1',
    system_prompt => 'You are a helpful assistant',
);

print $vllm->simple_chat('Say something nice');

# 2. Streaming
$vllm->simple_chat_stream(sub {
    print shift->content;
}, 'Write a haiku about Perl');

# 3. MCP tool calling (requires server started with tool-call-parser)
use Future::AsyncAwait;

my $vllm = Langertha::Engine::vLLM->new(
    url         => 'http://localhost:8000/v1',
    model       => 'Qwen/Qwen2.5-3B-Instruct',
    mcp_servers => [$mcp],
);

my $response = await $vllm->chat_with_tools_f('Add 7 and 15');

# 4. Multimodal input (vision-capable models: LLaVA, Qwen-VL, PaliGemma)
use Langertha::Content::Image;

my $img  = Langertha::Content::Image->from_url('https://example.com/cat.jpg');
my $resp = await $vllm->simple_chat_f({
    role    => 'user',
    content => [ 'What is in this image?', $img ],
});

# 5. Reasoning models (Qwen3, DeepSeek-R1, QwQ — server needs
#    --reasoning-parser matching the model; reasoning_effort maps to
#    enable_thinking via the chat template). simple_chat takes messages
#    only — set the control on the engine, or hand it to chat_f.
my $thinker = Langertha::Engine::vLLM->new(
    url              => 'http://localhost:8000/v1',
    reasoning_effort => 'high',
);

print $thinker->simple_chat('Solve step by step: what is 7 factorial?');

# … or as a per-request control:
my $reasoned = await $vllm->chat_f(
    messages         => ['Solve step by step: what is 7 factorial?'],
    reasoning_effort => 'high',
);

# 6. Embeddings (for embedding models: BAAI/bge-*, intfloat/e5-*, …)
my $vector = $vllm->simple_embedding('Some text to embed');

# 7. Prometheus /metrics scraping (Runtime::MetricsPoll)
my $records = await $vllm->poll_metrics_f('vllm:');
# Returns ArrayRef of { name, type, value, labels } parsed from
# GET <url-stripped-of-/v1>/metrics.

# 8. vLLM-Hook sub-engine (IBM vLLM-Hook plugin: hidden states, qk, steer)
use Langertha::Engine::VLLMHook;
my $hooked = Langertha::Engine::VLLMHook->new(
    url        => 'http://localhost:8770/v1',
    vllm_xargs => { output_hidden_states => JSON::MaybeXS::true() },
);

DESCRIPTION

Provides access to vLLM, a high-throughput inference engine for large language models. Extends Langertha::Engine::OpenAIBase (which composes Langertha::Role::OpenAICompatible, Langertha::Role::OpenAPI, Langertha::Role::Models, Langertha::Role::Temperature, Langertha::Role::ResponseSize, Langertha::Role::SystemPrompt, Langertha::Role::ResponseFormat, Langertha::Role::Streaming, Langertha::Role::Chat, Langertha::Role::ReasoningEffort, and Langertha::Role::PromptCache); vLLM itself additionally composes Langertha::Role::Tools (MCP tool calling), Langertha::Role::Embedding (OpenAI-compatible /v1/embeddings for embedding models), and Langertha::Role::Runtime::MetricsPoll (Prometheus /metrics scrape).

Supports chat, streaming, tool calling, embeddings, multimodal input, reasoning models, and Prometheus /metrics scraping.

Only url is required. The URL must include the /v1 path prefix (e.g., http://localhost:8000/v1). Since vLLM serves exactly one model (configured at server startup), no model name or API key is needed.

TOOL CALLING

MCP tool calling requires the vLLM server to be started with --enable-auto-tool-choice and --tool-call-parser matching the model (hermes for Qwen2.5/Hermes, llama3 for Llama, mistral for Mistral).

EMBEDDINGS

Composes Langertha::Role::Embedding. vLLM exposes an OpenAI-compatible /v1/embeddings endpoint when started with an embedding model (BAAI/bge-*, intfloat/e5-*, …). The request carries embedding_model if you set it, else model if you set it, else no model field at all: the server embeds with the model it serves. (vLLM 0.10/0.11 answer 404 to a model that is not a served name, such as the 'default' placeholder.)

REASONING MODELS

vLLM serves reasoning models through the OpenAI-compatible chat completions endpoint. reasoning_effort (Langertha::Role::ReasoningEffort, Langertha::Reasoning, wire format openai) is honored on the body — vLLM translates the value into enable_thinking via the chat template:

reasoning_effort low|medium|high   ->  enable_thinking = true
reasoning_effort none              ->  enable_thinking = false
(unset)                            ->  enable_thinking not injected

The thought itself comes back in the reasoning response field — mainline vLLM renamed it from reasoning_content to reasoning in v0.16.0, and both spellings are read — surfaced on $response->thinking. It is only populated when the server was started with the matching --reasoning-parser.

Supported reasoning-model families (vLLM's own table — verify against https://docs.vllm.ai/en/latest/features/reasoning_outputs/ for the exact list as new models land):

  • Qwen3 (Qwen3-Instruct / Qwen3-Thinking) — wire reasoning_effort honored.

  • DeepSeek-R1 (native) and DeepSeek-R1-Distill-* — wire reasoning_effort honored.

  • QwQ-32B — wire reasoning_effort honored (served via the deepseek_r1 reasoning parser).

  • Gemma 4, IBM Granite 3.2, DeepSeek-V3.1 / V4 — also honor the parameter via the same chat-template auto-injection path.

The server MUST be started with the matching --reasoning-parser flag for the loaded model (deepseek_r1 for DeepSeek-R1 / QwQ, qwen3 for Qwen3 thinking, etc.) — without it, the model emits reasoning content but Langertha cannot surface it on Langertha::Response. See vLLM's reasoning-outputs docs for the canonical flag list.

Per ADR 0009 the wire format is openai on vLLM; engines diverging (Anthropic, Gemini) carry their own reasoning_wire_format tag and serialize independently.

METRICS

Composes Langertha::Role::Runtime::MetricsPoll. vLLM serves Prometheus text-format metrics at GET <url-stripped-of-/v1/metrics>; pass a prefix to "poll_metrics_f" (e.g. 'vllm:' for the engine's own gauges) or filter downstream via "filter_prefix" in Langertha::Runtime::Metrics.

SUB-ENGINE: VLLMHook

Langertha::Engine::VLLMHook extends this engine for servers running the IBM vLLM-Hook plugin (hidden_states / qk / steer probes); it injects a top-level vllm_xargs field and lifts captured tensors onto "probes" in Langertha::Response. See ADR 0004 for the seam.

See https://docs.vllm.ai/ for installation and configuration details.

CAPABILITIES

Advertised flags (derived from composed roles via Langertha::Role::Capabilities):

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.