NAME

Langertha::Skeid::CapacityProbe::Prometheus - Capacity probe reading vLLM/SGLang/TGI Prometheus metrics

VERSION

version 0.003

SYNOPSIS

nodes:
  - id: gpu-1
    url: http://gpu-1:8000/v1
    model: qwen3-32b
    max_conns: 32
    capacity:
      probe: prometheus
      url: http://gpu-1:8000/metrics
      interval_ms: 2000
      limit: 32                       # optional; falls back to max_conns

DESCRIPTION

vLLM, SGLang and TGI already publish how many requests they are running and how many are queued. That is the number inflight is trying to reconstruct by counting — except this one includes traffic this Skeid never sent, which is the entire problem with counting (ADR 0009).

Used is running + waiting: a queued request is occupying the node as surely as a running one, and admitting more because they are "only waiting" is how a queue becomes a timeout.

Metric names

Defaults cover the common engines: vllm:num_requests_running, sglang:num_running_reqs and tgi_batch_current_size for running, vllm:num_requests_waiting, sglang:num_queue_reqs and tgi_queue_size for waiting. running (or used) and waiting in the capacity block override them, each a name or a list of names. Names are matched ignoring labels, and several series with the same name are summed — a per-model breakdown still adds up to what the node is doing.

The capacity block

url (the metrics endpoint) or path (default /metrics, see "url"), interval_ms (default 2000), running/used, waiting, and limit (the ceiling; default the node's max_conns, and with neither the reading does not narrow admission).

url

The metrics endpoint. Taken from the capacity block; if it only gives a path, it is resolved against the node's own URL, so path: /metrics is enough for the usual case where the engine serves metrics beside its API.

source

prometheus.

poll

Fetches "url" without blocking, one request at a time (a poll still out when the next tick fires is skipped), and reports used = running + waiting against the limit. A non-2xx answer, an unreachable endpoint, or a body without any of the metric names forgets the reading instead. An answer that arrives after "stop" in Langertha::Skeid::CapacityProbe is dropped.

parse_metrics

my $values = $probe->parse_metrics($body);   # { 'vllm:num_requests_running' => 3, ... }

Prometheus text format, reduced to what a probe needs: metric name to summed value, labels ignored, # lines skipped. Not a general-purpose parser and not trying to be — histograms and label selection are not what admission asks about.

SEE ALSO

Langertha::Skeid::CapacityProbe, and ADR 0009 in the distribution repository.

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha-skeid/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <torsten@raudssus.de> https://raudssus.de/

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.