NAME

Langertha::Role::PromptCache - Role for an engine with a request-side prompt-caching control

VERSION

version 0.503

prompt_cache

Enable an Anthropic cache_control breakpoint on the request (the top-level auto-place form). Defaults off. No effect on the OpenAI wire, where caching is automatic — see "prompt_cache_key" for the only OpenAI-side lever.

The top-level cache_control Langertha emits is Anthropic's documented "automatic caching" form: the system applies the cache breakpoint to the last cacheable block and it consumes one of the four available breakpoints. This is the intended shape — do not "fix" it into per-block breakpoints.

Turning prompt_cache on does not prove a cache write happened. The minimum cacheable prefix is model-dependent (roughly 512–4096 tokens depending on the model), and Anthropic silently processes a shorter prompt without caching — no error is returned. A 200 response therefore says nothing; only $response->usage->{cache_creation_input_tokens} / cache_read_input_tokens confirm the cache was actually used.

prompt_cache_ttl

Optional Anthropic cache time-to-live: 5m (the current default when unset) or 1h. Only meaningful together with "prompt_cache"; the 1h window requires this set explicitly.

prompt_cache_key

Optional OpenAI prompt_cache_key routing hint. OpenAI prompt caching is automatic; this only steers which cache shard is used. No effect on the Anthropic wire. Sent only where $engine->supports('prompt_cache_key'); the self-hosted OpenAI-compatible engines (vLLM, SGLang, llama.cpp, Ollama, LM Studio) do not advertise it and use Langertha::Role::RuntimeKnobs for their own prefix-cache controls.

cache_wire_format

cache_wire_format => 'anthropic'

The per-engine enum naming which caching dialect this engine speaks — openai | anthropic. Drives the value-object dispatch in "prompt_cache_kwargs". The default follows the engine base-class hierarchy: OpenAIBase leaves it at openai, AnthropicBase overrides to anthropic.

prompt_cache_kwargs_for

my %kwargs = $engine->prompt_cache_kwargs_for( prompt_cache => 1 );

Returns the body kwargs to merge into a chat request for the caching control, serialized for "cache_wire_format" via Langertha::PromptCache. %args may carry prompt_cache, prompt_cache_ttl and/or prompt_cache_key; keys it does not carry fall back to the engine attributes, so a per-request control (chat_f, karr #46) beats the configured attribute on a per-key basis. Empty list when nothing applies to the engine's wire (caching off / no key).

A prompt_cache_key the engine does not advertise ($engine->supports('prompt_cache_key') false, e.g. on the self-hosted OpenAI-compatible servers) is dropped silently, so the request body never carries a routing key the capability registry says the wire does not honor. prompt_cache is not gated, so an explicit cache_wire_format override keeps emitting cache_control.

prompt_cache_kwargs

my %kwargs = $engine->prompt_cache_kwargs;

Returns the body kwargs to merge into a chat request for the configured caching options, serialized for "cache_wire_format" via Langertha::PromptCache. Empty list when nothing applies to the engine's wire (caching off / no key). Delegates to "prompt_cache_kwargs_for" with no per-request overrides.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.