NAME
Langertha::Runtime::Metrics::EngineContract - Wire contract for self-hosted inference engine /metrics endpoints
VERSION
version 0.503
DESCRIPTION
This document captures, per engine, the exact HTTP shape of the runtime-metrics surface that Langertha::Role::Runtime::MetricsPoll scrapes. Engines listed here are self-hosted (no API key); the contract is what their server emits at a known URL path.
The contract drives two things:
Which URL path to GET
How to parse the response body into the Prometheus text-format record shape
{ name, type, value, labels }that Langertha::Runtime::Metrics consumes.
Each engine section lists the wire format (Prometheus text vs. custom JSON), the URL path, the namespace prefix or metric-key shape to allowlist, and a short worked example of the response body.
vLLM /metrics
URL:
{base}/metrics(base is theurlattribute without the trailing/v1)Format: Prometheus text exposition format
Allowlist prefixes:
vllm:,http:(HTTP-level request / latency / queue stats),process:Authentication: none (local server)
Representative payload:
# HELP vllm:gpu_cache_usage_perc GPU cache usage
# TYPE vllm:gpu_cache_usage_perc gauge
vllm:gpu_cache_usage_perc{model_name="Qwen/Qwen2.5-7B-Instruct"} 0.42
# HELP vllm:num_requests_running Number of running requests
# TYPE vllm:num_requests_running gauge
vllm:num_requests_running{model_name="Qwen/Qwen2.5-7B-Instruct"} 3
# HELP vllm:prompt_tokens_total Prompt tokens processed
# TYPE vllm:prompt_tokens_total counter
vllm:prompt_tokens_total{model_name="Qwen/Qwen2.5-7B-Instruct"} 18234
# HELP http_requests_total Total HTTP requests
# TYPE http_requests_total counter
http_requests_total{endpoint="v1.chat.completions",method="POST",status="200"} 412
SGLang /metrics
URL:
{base}/metricsFormat: Prometheus text exposition format
Allowlist prefixes:
sglang:(SGLang-specific runtime metrics),http:
Representative payload:
# HELP sglang:num_running_reqs Number of running requests
# TYPE sglang:num_running_reqs gauge
sglang:num_running_reqs{model="Qwen/Qwen2.5-7B-Instruct"} 2
# HELP sglang:prompt_tokens_total Prompt tokens processed
# TYPE sglang:prompt_tokens_total counter
sglang:prompt_tokens_total{model="Qwen/Qwen2.5-7B-Instruct"} 9431
# HELP sglang:gen_throughput Generation throughput (tokens/s)
# TYPE sglang:gen_throughput gauge
sglang:gen_throughput{model="Qwen/Qwen2.5-7B-Instruct"} 18.4
llama.cpp server /metrics
URL:
{base}/metricsFormat: Prometheus text exposition format
Allowlist prefixes:
llama_(note the_separator; llama.cpp emits underscore-prefixed names, not colon-prefixed)
Representative payload:
# HELP llama_prompt_tokens_total Total prompt tokens processed
# TYPE llama_prompt_tokens_total counter
llama_prompt_tokens_total 12450
# HELP llama_tokens_predicted_total Total generated tokens
# TYPE llama_tokens_predicted_total counter
llama_tokens_predicted_total 8213
# HELP llama_requests_processing Number of in-flight requests
# TYPE llama_requests_processing gauge
llama_requests_processing 1
# HELP llama_request_duration_seconds Request duration histogram
# TYPE llama_request_duration_seconds histogram
llama_request_duration_seconds_bucket{le="0.5"} 47
llama_request_duration_seconds_bucket{le="1.0"} 51
llama_request_duration_seconds_bucket{le="+Inf"} 52
llama_request_duration_seconds_sum 23.4
llama_request_duration_seconds_count 52
Ollama /api/ps (NOT Prometheus)
URL:
{base}/api/psFormat: JSON (
{"models":[{name, model, size, ...}, ...]})Authentication: none
This is not a Prometheus endpoint. There is no equivalent /metrics route exposed by Ollama at the time of writing. Wiring it into Langertha::Runtime::Metrics would require a JSON-to-Prometheus adapter (one synthetic metric per loaded model, e.g. ollama_model_size_bytes{model="..."}) that lives outside the current Prometheus-only parser. The MetricsPoll role therefore does not compose onto Langertha::Engine::Ollama for now — see the follow-up karr ticket tracked alongside this contract for the adapter work.
NAME
Langertha::Runtime::Metrics::EngineContract - Wire contract for self-hosted inference engine /metrics endpoints
SEE ALSO
Langertha::Runtime::Metrics - Prometheus text-format parser
Langertha::Role::Runtime::MetricsPoll - Async scraper that hits the endpoints documented above
https://prometheus.io/docs/instrumenting/exposition_formats/ - Prometheus text exposition format spec
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.