NAME
Langertha::Engine::VLLMHook - vLLM inference server with vLLM-Hook probe capture
VERSION
version 0.503
SYNOPSIS
use Langertha::Engine::VLLMHook;
# Capture hidden states for every request
my $engine = Langertha::Engine::VLLMHook->new(
url => 'http://localhost:8770/v1',
vllm_xargs => { output_hidden_states => JSON::MaybeXS::true() },
);
my $response = $engine->simple_chat('Hello');
print $response->content, "\n";
print "probes: ", join(',', keys %{$response->probes}), "\n"
if $response->has_probes;
# Drive xargs from a vLLM-Hook model config file
use Langertha::VLLMHook::Config;
my $cfg = Langertha::VLLMHook::Config->new(
file => 'model_configs/hidden_states/Qwen2.5-3B-Instruct.json',
);
my $engine = Langertha::Engine::VLLMHook->new(
url => 'http://localhost:8770/v1',
vllm_xargs => $cfg->xargs,
);
DESCRIPTION
Talks to a vLLM server running the vLLM-Hook plugin (https://github.com/IBM/vLLM-Hook), which observes attention patterns, extracts hidden states and performs activation steering. It speaks ordinary OpenAI-compatible HTTP, so this engine extends Langertha::Engine::vLLM and only adds two things:
On the way out, the
vllm_xargsHashRef is merged into the chat request body. vLLM maps the top-levelvllm_xargsfield ontoSamplingParams.extra_args, which the plugin reads to install its hooks. Nested values (dicts, lists) are JSON-encoded as strings becausevllm_xargsonly carries scalars — the plugin JSON-decodes them again before the worker reads them.On the way back, the serialized probe tensors that the plugin attaches as a top-level
probesfield are lifted onto "probes" in Langertha::Response.
The server must be started with the matching worker, e.g. VLLM_HOOK_WORKER=hidden_states vllm serve <model> --enforce-eager.
THIS API IS WORK IN PROGRESS
vllm_xargs
HashRef of extra arguments merged into the top-level vllm_xargs field of every chat request. Activates the vLLM-Hook plugin's probes. Defaults to an empty HashRef. Nested HashRef/ArrayRef values are JSON-encoded as strings on the wire (see "_encode_xargs"); plain scalars and JSON booleans pass through (a JSON boolean rides as a native JSON true/false). Recognised keys include output_hidden_states, output_qk, hookq_mode and steer. When non-empty this takes precedence over "worker_name".
worker_name
Optional convenience naming the vLLM-Hook worker the server was started with (VLLM_HOOK_WORKER): qk, hidden_states or steer. Used only to derive a default "vllm_xargs" when none was given explicitly. hidden_states yields { output_hidden_states => true }; qk and steer need their layer/head map or steering dict supplied via vllm_xargs and therefore contribute no default on their own.
resolved_xargs
my $xargs = $engine->resolved_xargs;
Returns the effective vllm_xargs HashRef. When "vllm_xargs" is non-empty it is returned as-is; otherwise a default is derived from "worker_name". Returns an empty HashRef when neither yields anything.
_encode_xargs
my $encoded = $engine->_encode_xargs(\%xargs);
Returns a copy of %xargs in which every nested HashRef/ArrayRef value is JSON-encoded to a string, leaving plain scalars and JSON booleans untouched (a JSON boolean rides as a native JSON true/false). vLLM's vllm_xargs only accepts scalar values, so structured values must travel as JSON strings; the vLLM-Hook plugin decodes them again.
chat_request
my $request = $engine->chat_request($messages, %extra);
Wraps "chat_request" in Langertha::Role::OpenAICompatible to merge "resolved_xargs" (encoded via "_encode_xargs") into the top-level vllm_xargs field of the request body. A per-call vllm_xargs in %extra overrides the instance values key-by-key.
chat_response
my $response = $engine->chat_response($http_response);
Wraps "chat_response" in Langertha::Role::OpenAICompatible to lift the serialized probe tensors from the response body's top-level probes field onto "probes" in Langertha::Response. Behaves exactly like the parent when the server returned no probes.
SEE ALSO
Langertha::Engine::vLLM - Parent engine (plain vLLM, no probes)
Langertha::VLLMHook::Config - Loads vLLM-Hook model config JSON into
vllm_xargs"probes" in Langertha::Response - Where captured probe tensors land
https://github.com/IBM/vLLM-Hook - The vLLM-Hook plugin
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.