NAME
Langertha::Runtime::Knobs - Immutable self-hosted runtime knobs with per-server conversion
VERSION
version 0.503
SYNOPSIS
my $k = Langertha::Runtime::Knobs->new(
prefix_cache_salt => 'tenant-a',
);
my %kwargs = $k->to('vllm');
# ( cache_salt => 'tenant-a' )
my $s = Langertha::Runtime::Knobs->new(
prefix_cache_salt => 'tenant-a',
return_cached_tokens_details => 1,
);
my %kw = $s->to('sglang');
# ( cache_salt => 'tenant-a', return_cached_tokens_details => true )
DESCRIPTION
Canonical value object for the request-side runtime knobs of the self-hosted OpenAI-compatible engines (vLLM, SGLang, llama.cpp), dispatched by an engine's knob_wire_format. Mirrors Langertha::Reasoning / Langertha::PromptCache: the per-provider placement of the fields lives in this one reviewable place rather than scattered across engines (ADR 0001, ADR 0004 — no raw extra_body side-channel; these knobs serialize as top-level request-body fields exactly like reasoning_kwargs_for / prompt_cache_kwargs_for do).
The only genuinely per-request knobs on these three servers are prefix-cache isolation/reuse controls. Each to_* serializer emits only the fields that server's wire accepts and returns the body kwargs to merge into the request.
What is deliberately NOT modeled here:
Speculative decoding — a server-launch concern on all three engines (
--speculative-configon vLLM,--speculative-*on SGLang,--spec-draft-*on llama.cpp), restart-only, never a per-request knob.Continuous batching — an internal scheduler flag, not a request knob.
prefix_cache_salt
Optional prefix-cache salt string. Emitted as cache_salt on the vLLM and SGLang wires.
This is a cache-ISOLATION feature, not a performance knob. It exists for multi-tenant privacy and timing-attack mitigation: only requests sharing a salt reuse each other's KV blocks, so a per-tenant salt keeps tenants from observing each other's cache behavior. Setting a random salt per request reduces cache reuse — the opposite of a cache-warming hint. Requests that want to share cached prefix blocks must carry the same salt.
Version gating: vLLM's cache_salt requires vLLM v1.x. Older servers silently ignore unknown fields, so setting it against an old server is a no-op, not an error.
cache_prompt
Optional llama.cpp cache_prompt boolean: whether to reuse the prompt cache for this request. Emitted as a JSON boolean on the llama.cpp wire.
n_cache_reuse
Optional llama.cpp n_cache_reuse integer. Counter-intuitive: 0 means reuse all cached tokens; higher values limit reuse (the server reuses at most that many cached tokens). The default is 0 (reuse everything).
id_slot
Optional llama.cpp id_slot integer: pin this request to a specific server slot. Useful for stateful / multi-session setups where a request must land on the slot holding the relevant KV cache.
priority
Optional SGLang priority integer: request scheduling priority. Higher values are served first.
return_cached_tokens_details
Optional SGLang return_cached_tokens_details boolean: ask the server to report how many prompt tokens were served from cache (surfaced on the response as usage.prompt_tokens_details.cached_tokens). Emitted as a JSON boolean on the SGLang wire.
extra_key
Optional SGLang extra_key string: an additional cache-key component, used to partition the prefix cache beyond the prompt text itself (e.g. per-tenant or per-request-type isolation).
to_vllm
Serializes to the vLLM wire. Emits exactly one field: cache_salt (from "prefix_cache_salt"). Every other knob is clamped away — vLLM accepts no other per-request runtime knob. Empty list when no salt is set.
to_sglang
Serializes to the SGLang wire: cache_salt (from "prefix_cache_salt"), extra_key, priority, and return_cached_tokens_details (a JSON boolean). Knobs SGLang does not accept (llama.cpp's cache_prompt / n_cache_reuse / id_slot) are clamped away. Empty list when none of the four fields is set.
to_llamacpp
Serializes to the llama.cpp wire: cache_prompt (a JSON boolean), n_cache_reuse, and id_slot. Knobs llama.cpp does not accept (vLLM / SGLang's cache_salt, SGLang's extra_key / priority / return_cached_tokens_details) are clamped away. Empty list when none of the three fields is set.
to
my %kwargs = $k->to($knob_wire_format);
Dispatch to the per-format serializer. Returns the body kwargs to merge into the request (an empty list when nothing applies to that wire). Croaks on an unknown format tag.
SEE ALSO
Langertha::Role::RuntimeKnobs - The composed role exposing the knob attributes
Langertha::Reasoning - Sibling value object for the reasoning-effort knob
Langertha::PromptCache - Sibling value object for the prompt-caching knob
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.