NAME

Langertha::Runtime::Knobs - Immutable self-hosted runtime knobs with per-server conversion

VERSION

version 0.503

SYNOPSIS

my $k = Langertha::Runtime::Knobs->new(
    prefix_cache_salt => 'tenant-a',
);
my %kwargs = $k->to('vllm');
# ( cache_salt => 'tenant-a' )

my $s = Langertha::Runtime::Knobs->new(
    prefix_cache_salt            => 'tenant-a',
    return_cached_tokens_details => 1,
);
my %kw = $s->to('sglang');
# ( cache_salt => 'tenant-a', return_cached_tokens_details => true )

DESCRIPTION

Canonical value object for the request-side runtime knobs of the self-hosted OpenAI-compatible engines (vLLM, SGLang, llama.cpp), dispatched by an engine's knob_wire_format. Mirrors Langertha::Reasoning / Langertha::PromptCache: the per-provider placement of the fields lives in this one reviewable place rather than scattered across engines (ADR 0001, ADR 0004 — no raw extra_body side-channel; these knobs serialize as top-level request-body fields exactly like reasoning_kwargs_for / prompt_cache_kwargs_for do).

The only genuinely per-request knobs on these three servers are prefix-cache isolation/reuse controls. Each to_* serializer emits only the fields that server's wire accepts and returns the body kwargs to merge into the request.

What is deliberately NOT modeled here:

  • Speculative decoding — a server-launch concern on all three engines (--speculative-config on vLLM, --speculative-* on SGLang, --spec-draft-* on llama.cpp), restart-only, never a per-request knob.

  • Continuous batching — an internal scheduler flag, not a request knob.

prefix_cache_salt

Optional prefix-cache salt string. Emitted as cache_salt on the vLLM and SGLang wires.

This is a cache-ISOLATION feature, not a performance knob. It exists for multi-tenant privacy and timing-attack mitigation: only requests sharing a salt reuse each other's KV blocks, so a per-tenant salt keeps tenants from observing each other's cache behavior. Setting a random salt per request reduces cache reuse — the opposite of a cache-warming hint. Requests that want to share cached prefix blocks must carry the same salt.

Version gating: vLLM's cache_salt requires vLLM v1.x. Older servers silently ignore unknown fields, so setting it against an old server is a no-op, not an error.

cache_prompt

Optional llama.cpp cache_prompt boolean: whether to reuse the prompt cache for this request. Emitted as a JSON boolean on the llama.cpp wire.

n_cache_reuse

Optional llama.cpp n_cache_reuse integer. Counter-intuitive: 0 means reuse all cached tokens; higher values limit reuse (the server reuses at most that many cached tokens). The default is 0 (reuse everything).

id_slot

Optional llama.cpp id_slot integer: pin this request to a specific server slot. Useful for stateful / multi-session setups where a request must land on the slot holding the relevant KV cache.

priority

Optional SGLang priority integer: request scheduling priority. Higher values are served first.

return_cached_tokens_details

Optional SGLang return_cached_tokens_details boolean: ask the server to report how many prompt tokens were served from cache (surfaced on the response as usage.prompt_tokens_details.cached_tokens). Emitted as a JSON boolean on the SGLang wire.

extra_key

Optional SGLang extra_key string: an additional cache-key component, used to partition the prefix cache beyond the prompt text itself (e.g. per-tenant or per-request-type isolation).

to_vllm

Serializes to the vLLM wire. Emits exactly one field: cache_salt (from "prefix_cache_salt"). Every other knob is clamped away — vLLM accepts no other per-request runtime knob. Empty list when no salt is set.

to_sglang

Serializes to the SGLang wire: cache_salt (from "prefix_cache_salt"), extra_key, priority, and return_cached_tokens_details (a JSON boolean). Knobs SGLang does not accept (llama.cpp's cache_prompt / n_cache_reuse / id_slot) are clamped away. Empty list when none of the four fields is set.

to_llamacpp

Serializes to the llama.cpp wire: cache_prompt (a JSON boolean), n_cache_reuse, and id_slot. Knobs llama.cpp does not accept (vLLM / SGLang's cache_salt, SGLang's extra_key / priority / return_cached_tokens_details) are clamped away. Empty list when none of the three fields is set.

to

my %kwargs = $k->to($knob_wire_format);

Dispatch to the per-format serializer. Returns the body kwargs to merge into the request (an empty list when nothing applies to that wire). Croaks on an unknown format tag.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.