NAME

Langertha::Runtime::Metrics::EngineContract - Wire contract for self-hosted inference engine /metrics endpoints

VERSION

version 0.503

DESCRIPTION

This document captures, per engine, the exact HTTP shape of the runtime-metrics surface that Langertha::Role::Runtime::MetricsPoll scrapes. Engines listed here are self-hosted (no API key); the contract is what their server emits at a known URL path.

The contract drives two things:

  • Which URL path to GET

  • How to parse the response body into the Prometheus text-format record shape { name, type, value, labels } that Langertha::Runtime::Metrics consumes.

Each engine section lists the wire format (Prometheus text vs. custom JSON), the URL path, the namespace prefix or metric-key shape to allowlist, and a short worked example of the response body.

vLLM /metrics

  • URL: {base}/metrics (base is the url attribute without the trailing /v1)

  • Format: Prometheus text exposition format

  • Allowlist prefixes: vllm:, http: (HTTP-level request / latency / queue stats), process:

  • Authentication: none (local server)

Representative payload:

# HELP vllm:gpu_cache_usage_perc GPU cache usage
# TYPE vllm:gpu_cache_usage_perc gauge
vllm:gpu_cache_usage_perc{model_name="Qwen/Qwen2.5-7B-Instruct"} 0.42
# HELP vllm:num_requests_running Number of running requests
# TYPE vllm:num_requests_running gauge
vllm:num_requests_running{model_name="Qwen/Qwen2.5-7B-Instruct"} 3
# HELP vllm:prompt_tokens_total Prompt tokens processed
# TYPE vllm:prompt_tokens_total counter
vllm:prompt_tokens_total{model_name="Qwen/Qwen2.5-7B-Instruct"} 18234
# HELP http_requests_total Total HTTP requests
# TYPE http_requests_total counter
http_requests_total{endpoint="v1.chat.completions",method="POST",status="200"} 412

SGLang /metrics

  • URL: {base}/metrics

  • Format: Prometheus text exposition format

  • Allowlist prefixes: sglang: (SGLang-specific runtime metrics), http:

Representative payload:

# HELP sglang:num_running_reqs Number of running requests
# TYPE sglang:num_running_reqs gauge
sglang:num_running_reqs{model="Qwen/Qwen2.5-7B-Instruct"} 2
# HELP sglang:prompt_tokens_total Prompt tokens processed
# TYPE sglang:prompt_tokens_total counter
sglang:prompt_tokens_total{model="Qwen/Qwen2.5-7B-Instruct"} 9431
# HELP sglang:gen_throughput Generation throughput (tokens/s)
# TYPE sglang:gen_throughput gauge
sglang:gen_throughput{model="Qwen/Qwen2.5-7B-Instruct"} 18.4

llama.cpp server /metrics

  • URL: {base}/metrics

  • Format: Prometheus text exposition format

  • Allowlist prefixes: llama_ (note the _ separator; llama.cpp emits underscore-prefixed names, not colon-prefixed)

Representative payload:

# HELP llama_prompt_tokens_total Total prompt tokens processed
# TYPE llama_prompt_tokens_total counter
llama_prompt_tokens_total 12450
# HELP llama_tokens_predicted_total Total generated tokens
# TYPE llama_tokens_predicted_total counter
llama_tokens_predicted_total 8213
# HELP llama_requests_processing Number of in-flight requests
# TYPE llama_requests_processing gauge
llama_requests_processing 1
# HELP llama_request_duration_seconds Request duration histogram
# TYPE llama_request_duration_seconds histogram
llama_request_duration_seconds_bucket{le="0.5"} 47
llama_request_duration_seconds_bucket{le="1.0"} 51
llama_request_duration_seconds_bucket{le="+Inf"} 52
llama_request_duration_seconds_sum 23.4
llama_request_duration_seconds_count 52

Ollama /api/ps (NOT Prometheus)

  • URL: {base}/api/ps

  • Format: JSON ({"models":[{name, model, size, ...}, ...]})

  • Authentication: none

This is not a Prometheus endpoint. There is no equivalent /metrics route exposed by Ollama at the time of writing. Wiring it into Langertha::Runtime::Metrics would require a JSON-to-Prometheus adapter (one synthetic metric per loaded model, e.g. ollama_model_size_bytes{model="..."}) that lives outside the current Prometheus-only parser. The MetricsPoll role therefore does not compose onto Langertha::Engine::Ollama for now — see the follow-up karr ticket tracked alongside this contract for the adapter work.

NAME

Langertha::Runtime::Metrics::EngineContract - Wire contract for self-hosted inference engine /metrics endpoints

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.