NAME

Langertha::Skeid::CapacityProbe - Background probes that report what a node's real capacity is

VERSION

version 0.003

DESCRIPTION

inflight counts what this process sent to a node. That is exact for one Skeid in front of one node, and wrong the moment anything else sends work there — a second frontend, a prefork worker, a batch job, an engineer with curl. Each counter then sees only its own share, every instance believes the node is emptier than it is, and together they over-admit.

A probe reports what the node itself says instead. See ADR 0009 for why this rather than a shared counter in Redis.

Writing one

Subclass this, implement "poll", and report through "set_capacity_reading" in Langertha::Skeid. Two rules that are not negotiable:

  • Never on the request path. Polling is a timer. A probe that resolves during a request has moved a network round-trip into the latency of a request that did not ask for it.

  • Report nothing rather than something old. If the source cannot be reached, call forget_capacity. Admission falls back to inflight, which is merely imprecise; a stale reading is confidently wrong.

skeid

The control plane to report to.

node_id

Which node this probe describes.

interval_ms

How often to poll (default 2000). The right value is a trade between staleness and load on the node's metrics endpoint, and ADR 0009 leaves it open pending measurement — 2s is a starting point, not a finding.

config

The node's capacity block, verbatim.

is_stopped

True once "stop" has run. A poll whose answer arrives after that (an HTTP request still in flight at a restart) must not report: the node may be gone, or a new probe may already describe it. A subclass that reports asynchronously checks this before it writes.

source

The source this probe writes its readings under (prometheus, registry), or undef when it does not have a fixed one -- a custom callback picks its own. When the probe fails or stops it forgets only a reading under this source, so a 429 backoff the passive rate-limit probe recorded outlives a probe that went blind (ADR 0017). Undef forgets whatever is there.

poll

$probe->poll;

What a subclass implements: read the source, then either set_capacity_reading or forget_capacity on "skeid". Must not block.

poll_interval_seconds

How often this process actually polls: "interval_ms" multiplied by the worker count.

Every worker runs its own copy of the timer, so without the multiplier four workers on a 2s interval would hit the node's metrics endpoint every 500ms — the probe becoming the load it was meant to measure. What the operator configured is the rate the *node* sees from the process group (ADR 0010).

start

$probe->start;

Begins polling on a timer, and polls once immediately so the first request does not have to wait an interval for a reading. Safe to call twice.

stop

Stops polling and drops the node's reading, so admission goes back to inflight rather than acting on whatever this probe last said.

for_node

my $probe = Langertha::Skeid::CapacityProbe->for_node($skeid, $node);

Builds the probe a node's capacity block asks for, or nothing when it asks for none.

capacity:
  probe: prometheus              # or: inflight, custom, registry
  url: http://gpu-1:8000/metrics
  interval_ms: 2000

probe may also be spelled type. inflight (the default; also none, and an absent block) means no probe object at all — that is the default admission path, not a probe that reports the same thing. ratelimit is likewise not built here: it is passive, read off responses the proxy already has, and needs nothing running.

registry is for a node that is itself a Skeid: it pulls that Skeid's signed registry snapshot (Langertha::Skeid::CapacityProbe::Registry, ADR 0017).

custom takes either a code callback (given the probe, reports through the same methods) or a class to load, because Skeid is generic and the built-ins only cover the engines we happen to know. The class must look like a Perl package name, is loaded by name and built with skeid, node_id, config and interval_ms.

interval_ms in the block becomes "interval_ms". Croaks on an unknown probe, and on a custom block with neither code nor a valid class.

start_for_skeid

my $probes = Langertha::Skeid::CapacityProbe->start_for_skeid($skeid);

Builds and starts a probe for every node that asks for one, and returns them by node id. The caller holds them: a probe that goes out of scope stops polling. A node whose block does not build is skipped with a warning, so one bad block does not stop the others.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha-skeid/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <torsten@raudssus.de> https://raudssus.de/

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.