NAME

Langertha::Skeid - Dynamic routing control-plane for multi-node LLM serving with normalized metrics and cost accounting

VERSION

version 0.003

SYNOPSIS

use Langertha::Skeid;

my $skeid = Langertha::Skeid->new(
  config_file => '/etc/skeid/config.yaml',
);

my $cost = $skeid->call_function('metrics.estimate_cost', {
  model => 'gpt-4o-mini',
  usage => { prompt_tokens => 1000, completion_tokens => 200 },
});

DESCRIPTION

Langertha::Skeid is a routing control-plane for provider-style LLM operations. It keeps a live node table, routes by model/health/capacity, and records normalized token/cost usage.

Skeid is commonly used as one API edge in front of many upstream APIs (cloud + local). With pricing and usage.record/report, you can build tenant billing from one consistent ledger.

Multi-API Billing Flow

1. Define multiple nodes in config (for example OpenAI-compatible cloud APIs and local vLLM/SGLang). 2. Set model pricing via pricing or pricing.set. 3. Let tenant identity follow from the API key the caller presents. skeid keyid prints the id a given key resolves to. A deployment that authenticates callers in front of Skeid can instead pass x-skeid-key-id and set routing.trust_key_id_header. 4. Read totals by key/model/time with usage.report.

Engine IDs

nodes[].engine uses lowercased engine class names from Langertha. Examples: OpenAI => openai, OpenAIBase => openaibase, vLLM => vllm. Legacy aliases like openai-compatible are intentionally rejected. A node without engine is openaibase. See "normalize_engine_id".

Pluggable Usage Storage

The usage storage layer is pluggable. Built-in backends are jsonlog (recommended, no DBI required), sqlite, and postgresql. You can also replace the storage layer entirely via constructor callbacks or subclass override.

jsonlog backend (recommended — no DBI dependency):

# Directory mode: one JSON file per event (no collision risk)
my $skeid = Langertha::Skeid->new(
  usage_store => { backend => 'jsonlog', path => '/var/log/skeid/events/' },
);

# File mode: JSON-lines appended to a single file
my $skeid = Langertha::Skeid->new(
  usage_store => { backend => 'jsonlog', path => '/var/log/skeid/usage.jsonl', mode => 'file' },
);

Directory mode is auto-detected when the path is an existing directory or ends with /. It writes one .json file per event, which avoids file-level locking and concurrent-write collisions entirely.

Constructor callbacks (custom backend, no subclassing):

my $skeid = Langertha::Skeid->new(
  store_usage_event => sub {
    my ($self, $event) = @_;
    # $event is a hashref with all normalized usage columns
    publish_to_nats($event);
    return { ok => 1 };
  },
  query_usage_report => sub {
    my ($self, $filters) = @_;
    # $filters has: since, api_key_id, model, limit
    return { ok => 1, enabled => 1, totals => { ... } };
  },
);

Subclass override:

package MyApp::Skeid;
use Moo;
extends 'Langertha::Skeid';

sub _store_usage_event {
  my ($self, $event) = @_;
  ...
  return { ok => 1 };
}

sub _query_usage_report {
  my ($self, $filters) = @_;
  ...
}

Optional fields. A streamed request's event also carries content_bytes: the UTF-8 byte count of the content Skeid relayed (OpenAI face) or translated (Anthropic and Ollama faces). It is set on every stream, including one whose upstream reported token counts, and is absent from non-streamed events -- a sink must not assume it is there. It is an observation, not a billing quantity: Skeid never derives token counts or cost from it. The DBI stores keep it in a nullable content_bytes column (NULL for a non-streamed or pre-existing row); jsonlog writes it as part of the event.

When a callback or override is provided, the configured store is bypassed entirely and no database connection is created. DBI, DBD::SQLite and DBD::Pg are recommends dependencies — they are not required for jsonlog or when usage is handled externally. With no sink at all, "record_usage" builds no event and answers { ok => 0, enabled => 0, error => 'usage_store not configured' }. A sink that fails answers { ok => 0, error => ... } (or dies); the proxy logs that at error level as a lost usage event, naming the request id and the store backend.

Per-Key Routing Policy

Which nodes a customer key may be served from, and which models it may ask for:

policies:
  standard:  { deny_tags: [cloud] }        # our own hardware only
  burstable: {}                            # cloud is fine when local is full
default_policy: standard
names:
  alice:   k_5f0e1a2b3c4de5f60718293a4b5c6d7e8f901a2b  # id from `skeid keyid <key>` -- never the key
  bigcorp: k_9c8b7a6f5e4de5f60718293a4b5c6d7e8f901a2b
keys:
  alice: burstable                          # a keys: entry may be written by name ...
  bigcorp:
    policy: standard
    models: [house-model]                  # sparse override of one field
  k_1122334455aae5f60718293a4b5c6d7e8f901a2b: burstable  # ... or still by the raw key id

Resolved once at config load: a request costs one hash lookup, keys on the same profile share one policy object, and a key that takes the default is not listed at all. deny_tags filters node selection, not just the plan, so a denied node cannot be reached by asking for its own model name instead of an alias. A refusal is 403, never a capacity error.

The optional names: section maps a readable name to a customer key id (karr #17). It makes keys: entries legible and confines key rotation to a single line -- change the id a name points at, and every policy line that named the customer follows. Its values are ids from skeid keyid, never customer keys, so the config still holds no secret; it is read only at config load and never on the request path.

Identity comes from the API key the caller presented — see "key_id_for_key". See also ADR 0008 in the distribution repository.

Admin API Key

admin.api_key (or admin_api_key) controls access to proxy admin routes; admin.api_key_env (or admin_api_key_env) names an environment variable to read it from instead, so the key need not be written into a mounted file. If empty, admin routes are effectively disabled by returning 404. If set, the proxy expects Authorization: Bearer .... An explicit key (skeid serve --admin-api-key, "set_admin_api_key") outranks the config, and SKEID_ADMIN_API_KEY stands in when the config names none; a reload re-applies that order. See "admin".

Provider Manifest

GET /.well-known/langertha.json serves a provider manifest (core's Langertha::Manifest, schema v1) to a customer key -- listing only the models that key's own keys: entry names. Nothing is published automatically: the route answers 404 until the config turns it on, and a key without a manifest: entry gets 403, never the catalog.

manifest:
  enabled: true
  public_url: https://llm.example.com   # where clients reach Skeid -- never a node URL
  provider_id: example-llm              # default: skeid
  faces: [openai, anthropic, ollama]    # default: all three
  capabilities:                         # optional claims per model; default chat + streaming
    house-model: { tools_native: true, tool_choice_auto: true }
names:                                  # ids from `skeid keyid <key>` -- never the key
  alice: k_5f0e1a2b3c4de5f60718293a4b5c6d7e8f901a2b
  bob:   k_9c8b7a6f5e4de5f60718293a4b5c6d7e8f901a2b
keys:
  alice:
    policy: burstable
    manifest: { models: [house-model, qwen3-32b] }
  bob:
    manifest: { models: [house-model] }

One endpoint per face, all under public_url and all with one api_key auth entry (the Authorization: Bearer / x-api-key scheme every face already takes): openai (openai-chat, public_url/v1), anthropic (anthropic-compat at public_url -- Skeid translates /v1/messages to the OpenAI upstream call and does not carry output_config.format, so structured output takes the synthetic-tool path) and ollama (public_url). Every listed model appears on every published face. Its capabilities default to chat and streaming, take the claims declared for it (only flags from "model_capabilities" in Langertha::Manifest::Builder), and are then cut per face to what that face's translator carries upstream -- a claim holds at that endpoint or is not made there. The lists are "openai_manifest_endpoint" in Langertha::Skeid::Protocol, "manifest_endpoint" in Langertha::Skeid::Protocol::Anthropic and "manifest_endpoint" in Langertha::Skeid::Protocol::Ollama.

Resolved at config load, like the routing policy: each key's manifest is built and validated once and stored under the key id, so a request costs one hash lookup and one key's manifest cannot be served for another. The load croaks on a model the key's routing policy does not let it reach (not granted, or served only by nodes it is denied), on an unknown capability, one no face carries, an unknown face, and an enabled manifest without public_url. A config that fails to load keeps the previous one in force, manifests included. The route never reloads the config itself; the request paths that already do pick up a change. Without a key the route answers 401. Needs a Langertha with Langertha::Manifest; on an older one the route answers 404. See ADR 0015 in the distribution repository.

CONFIGURATION

The config is a YAML file (config_file) or the hashref a config_loader returns; both go through the same loader. The top-level keys below are every one Skeid reads -- anything else in the file is ignored. A reload is all or nothing: a config that fails anywhere leaves the previous one in force, and is reported by "reload_status".

A reload makes the running config equal to the file (skeid k65), as a restart with that file would:

  • nodes, pricing, aliases, the policy keys (policies, default_policy, names, keys) and usage_store: present, the section replaces what is loaded -- pricing too, so a model removed from it loses its price. Removed from the file, the section goes back to empty with the next reload: no nodes, no prices, no aliases, no policies.

  • routing: each key present sets its value; a key removed from the file goes back to the value it had before any config was applied -- passed to new, else from "ENVIRONMENT", else built in.

  • The admin key: absent falls back to SKEID_ADMIN_API_KEY, else off -- see "admin".

  • registry and manifest: absent means off.

Two exceptions. A section no applied config has declared belongs to whoever set it: nodes added through the admin API or "add_node", prices from "set_model_pricing", values passed to new -- a config without that section leaves them alone. And the usage store is not removed by a reload: a usage_store taken out of the file keeps the running store in force until restart, with one warning (usage_store was removed from the config ...), because a reload that stopped recording usage events would lose billing data silently (ADR 0004). A changed usage_store is swapped at once, the new store prepared before the old one is let go.

Three settings look like config but are not read from it: "capacity_max_age_ms" and "config_reload_interval" (constructor or environment only) and "worker_count" (constructor, set by skeid serve --workers).

nodes

nodes:
  - id: gpu-1                        # required, unique
    url: http://gpu-1:8000/v1        # required; /v1 is added when the URL does not end in it
    model: qwen3-32b                 # default '': matches any requested model
    engine: vllm                     # default openaibase; see "Engine IDs"
    weight: 3                        # default 1; round-robin weight, integer, below 1 counts as 1
    max_conns: 8                     # default 0 = unlimited; admission limit, see worker_max_conns
    healthy: true                    # default true; operator state, never derived from errors
    tags: [local, gb10]              # default none; "local, gb10" works too, see normalize_tags
    metadata: { rack: b2 }           # default {}; kept on the node, never read by routing
    api_key_ref: secret/skeid/remote/groq   # upstream key, resolved through the key broker
    api_key_env: GROQ_API_KEY        # upstream key from this variable when the ref resolves none
    capacity: { probe: prometheus }  # default none: inflight admission, see below

An entry without id or url is skipped; one with an unknown engine fails the load. When a node has a key of its own it replaces the client's Authorization upstream (see Langertha::Skeid::Proxy); otherwise the client's headers go through. Key references and variable names are the only key material a config holds (ADR 0003).

The section replaces the whole inventory when it changed -- nodes added through the admin API are lost then -- and keeps the node list as it is when it did not: its inventory generation, its running probes and any health set through the admin API. Removing the section empties the inventory.

capacity selects how admission learns a node's real occupancy (ADR 0009). probe (or type) is inflight (the default, also none), ratelimit, prometheus, registry or custom; interval_ms (default 2000) sets the poll rate. Rate-limit headers are read off every upstream response whatever the block says, so ratelimit starts nothing. The keys per probe are documented in "for_node" in Langertha::Skeid::CapacityProbe, Langertha::Skeid::CapacityProbe::Prometheus and Langertha::Skeid::CapacityProbe::Registry.

pricing

pricing:
  gpt-4o-mini:
    input_per_million: 0.15          # default 0
    output_per_million: 0.60         # default 0
    cached_input_per_million: 0.075  # optional; prompt-cache reads
    cache_write_per_million: 0.1875  # optional; prompt-cache writes
  '*':                               # fallback for every model not listed
    input_per_million: 0
    output_per_million: 0

Keyed by the served model -- the name the node is asked for, not an alias. The section is the whole price list: a reload replaces it, it is not merged. A cache rate must be a number >= 0, or the load fails; on a Langertha that cannot price cache tokens the cache rates are dropped with one warning and cached tokens bill at input_per_million. See "set_model_pricing".

aliases

aliases:
  house-model:
    tiers:
      - { tags: [local], model: qwen3-32b, wait_ms: 200 }
      - { tags: [cloud], model: llama-3.3-70b-versatile }

A client-facing model name as an ordered plan of tiers (ADR 0008). Per tier: tags, model (default: the alias name), engine and wait_ms (default 0). tiers may be given as a bare list. See "set_model_alias".

policies, default_policy, names, keys

Who may reach what; see "Per-Key Routing Policy". A policy is models (or aliases: the requested names the key may use; absent or '*' means all) and deny_tags (nodes the key must never reach). default_policy names the policy of every unlisted key; unset, an unlisted key is unrestricted. names maps a readable name to a key id from skeid keyid.

A keys entry is keyed by a key id or a names name and is either a policy name or a hash:

keys:
  alice: burstable
  bigcorp:
    policy: standard          # default: default_policy
    models: [house-model]     # sparse overrides: an absent field keeps the policy's value
    deny_tags: [cloud]
    manifest: { models: [house-model] }   # see "Provider Manifest"

The override fields may instead sit under an overrides hash. The load fails on a default_policy or a keys entry naming an undefined policy, on a names value that is not a non-empty string, and on two entries that resolve to one key id -- including a short (pre-ADR 0016) id beside the full id it is the prefix of. Short ids still match, with a one-time warning.

routing

routing:
  wait_timeout_ms: 2000        # default 2000: the wait of a model that has no alias
  wait_poll_ms: 25             # default 25, at least 1: how often a waiting request retries
  trust_key_id_header: false   # default false; see key_id_for_key
  frontend_count: 1            # default 1, at least 1; see frontend_count

The defaults come from "ENVIRONMENT" when set there.

admin

admin:
  api_key_env: SKEID_ADMIN_API_KEY   # or api_key: "..."

The key for the /skeid/* admin API. It comes from the first of these that has one (skeid k64):

1. An explicit key: skeid serve --admin-api-key, build_app(admin_api_key => ...), an admin_api_key passed to new, or "set_admin_api_key".
2. The config: the first of admin_api_key, admin_api_key_env, admin.api_key and admin.api_key_env that is present decides -- set empty, or naming an unset variable, it turns the admin API off.
3. SKEID_ADMIN_API_KEY, when the config names none of the four.

With none of them the admin API is off (404). Every applied config re-resolves the key in this order, so a reload never replaces an explicit key, and removing the key from the config falls back to the variable. A variable is read when the config is applied, so a changed value takes effect with the next changed config.

registry

Publishes this Skeid's capacity to a fronting Skeid: enabled, secret_env, read_key_env, ttl_s (default 10), instance_id (default the hostname) and error_window_s (default 60). Documented under "registry_enabled".

manifest

The provider manifest: enabled, public_url, provider_id (default skeid), faces (default all three) and capabilities. Documented under "Provider Manifest".

usage_store

usage_store:
  backend: jsonlog                 # jsonlog | sqlite | postgresql; inferred when absent
  path: /var/log/skeid/events/     # jsonlog: log_path or path; mode dir|file (inferred), fsync
  # sqlite:     sqlite_path (or path, db_path), schema_file, auto_migrate (default on)
  # postgresql: dsn, or host (127.0.0.1) / port (5432) / dbname or database (skeid);
  #             user, password or password_env, schema_file, auto_migrate (default on)
  # sqlite and postgresql: flush_interval_ms (default 0) -- above 0, queue events and
  #             write them in one transaction per interval (write-behind)

Where usage events go; see "normalize_config" in Langertha::Skeid::UsageStore for the inference and the defaults. usage_db_path, a top-level key, is the older spelling of a sqlite store and is read only when usage_store is absent. A changed store is swapped on reload; a removed one stays in force until restart, with a warning -- see "CONFIGURATION".

flush_interval_ms moves a database write off the request path: the request is answered at once and the queued events are written together once per interval. Queued events are held in memory until then; a process that is killed rather than stopped loses them. See "flush_interval_ms" in Langertha::Skeid::UsageStore::DBI and ADR 0005.

ENVIRONMENT

Defaults for the attributes of the same meaning; a value in the config wins. Besides these, the config names variables of its own -- api_key_env on a node, admin.api_key_env, a usage store's password_env, the registry's secret_env and read_key_env -- so that no secret has to be written into it. "build_app" in Langertha::Skeid::Proxy reads the OPENBAO_* variables, SKEID_UPSTREAM_POOL and SKEID_UPSTREAM_TIMEOUT.

SKEID_ROUTE_WAIT_TIMEOUT_MS

Default of "route_wait_timeout_ms" (2000).

SKEID_ROUTE_WAIT_POLL_MS

Default of "route_wait_poll_ms" (25).

SKEID_TRUST_KEY_ID_HEADER

Default of "trust_key_id_header": 1, true, yes or on turn it on.

SKEID_FRONTEND_COUNT

Default of "frontend_count" (1).

SKEID_ADMIN_API_KEY

The admin API key when neither an explicit key nor the config names one; see "admin".

SKEID_USAGE_DB

Default of "usage_db_path": a SQLite usage store at this path when no usage_store is given, and the path of a sqlite store that names none.

SKEID_CAPACITY_MAX_AGE_MS

Default of "capacity_max_age_ms" (5000).

SKEID_CONFIG_RELOAD_INTERVAL

Default of "config_reload_interval" (1 second).

ATTRIBUTES

Every attribute can be passed to new. Those a config sets are overwritten by the next changed config that sets them (see "CONFIGURATION") -- except "admin_api_key", which new takes as an explicit key that outranks the config.

nodes

The node inventory: an arrayref of node hashrefs, in the shape of the config's nodes entries after "add_node" normalized them (default empty). Change it through "add_node", "remove_node" and "set_node_health", or assign a whole new list: routing caches what it derives from the inventory, and only those paths invalidate the cache. Editing a node hash in place does not.

model_pricing

Served model name to pricing rule, '*' as the fallback (default empty). Written by "set_model_pricing"; read through "pricing_for_model".

model_aliases

Alias name to { tiers => [...] }, normalized (default empty). Written by "set_model_alias".

policies

Policy name to resolved policy (see "resolve_policy"), from the config's policies or "set_policy" (default empty).

default_policy

The resolved policy an unlisted key routes under -- the policy object, not its name -- or undef, which leaves unlisted keys unrestricted (default undef).

key_policies

Customer key id to resolved policy, built at config load (default empty). Keys that take the default policy are not in it; see "policy_for_key".

key_names

Readable name to customer key id, from the config's names section (default empty). See "key_id_for_name".

manifest_enabled

Whether the config enables the provider manifest (default 0).

manifest_available

Whether the installed Langertha has Langertha::Manifest; with it missing an enabled manifest answers 404 instead of failing the config (default 0).

key_manifests

Customer key id to the canonical JSON of that key's manifest, built at config load (default empty). Read through "manifest_for_key".

trust_key_id_header

Whether a client's x-skeid-key-id (or x-api-key-id) header names the customer key id instead of the key it presented (default off, or SKEID_TRUST_KEY_ID_HEADER; config routing.trust_key_id_header). Turn it on only behind a gateway that authenticates the caller and sets that header itself: the key id selects both the routing policy and the bill.

route_wait_timeout_ms

How long, in milliseconds, a request for a model without an alias waits for a free node before it is answered 429 (default 2000, or SKEID_ROUTE_WAIT_TIMEOUT_MS; config routing.wait_timeout_ms). An alias tier waits its own wait_ms instead.

route_wait_poll_ms

How often, in milliseconds, a waiting request tries again (default 25, or SKEID_ROUTE_WAIT_POLL_MS; config routing.wait_poll_ms, at least 1). The wait is a Mojo::IOLoop timer, never a sleep.

usage_db_path

A SQLite usage database path (default SKEID_USAGE_DB, else unset). At construction it configures a SQLite store when usage_store is not given; it is also the path of a sqlite store config that names none. Kept in step with the store: set to a SQLite store's path, cleared for PostgreSQL. has_usage_db_path and clear_usage_db_path are its predicate and clearer.

usage_store

The usage store config (default empty: no store). Passed to new it is a raw usage_store config; afterwards it holds the normalized form (see "normalize_config" in Langertha::Skeid::UsageStore). Change it through "configure_usage_store".

store_usage_event

store_usage_event => sub { my ($skeid, $event) = @_; ...; return { ok => 1 } },

Optional code ref that receives every usage event instead of the configured store (has_store_usage_event is its predicate). See "Pluggable Usage Storage".

query_usage_report

query_usage_report => sub { my ($skeid, $filters) = @_; ...; return { ok => 1, ... } },

Optional code ref that answers "usage_report" instead of the configured store (has_query_usage_report is its predicate). It gets since, api_key_id, model and limit.

on_usage_lost

$skeid->on_usage_lost(sub { my ($skeid, $event, $error) = @_; ... });

Optional code ref (has_on_usage_lost is its predicate), called for every usage event a write-behind store (usage_store.flush_interval_ms) queued and then could not write. The request was answered before the write was tried, so this is the only place that failure can be said. "build_app" in Langertha::Skeid::Proxy sets it to the proxy's usage event lost log line; without it Skeid warns. A failure of a synchronous write is the caller's to report, from the answer of "record_usage", and never comes here.

admin_api_key

The bearer token of the /skeid/* admin API in force; empty disables it, which the proxy answers with 404. Passed to new it is an explicit key (see "set_admin_api_key") that wins over the config; otherwise it is resolved from the config and SKEID_ADMIN_API_KEY whenever a config is applied -- see "admin" for the order. Read it here; set it through "set_admin_api_key", since a value written through this accessor lasts only until the next changed config.

key_broker

Optional Langertha::Skeid::KeyBroker that resolves a node's api_key_ref per request (has_key_broker is its predicate). "build_app" in Langertha::Skeid::Proxy passes a Langertha::Skeid::KeyBroker::OpenBao when the OPENBAO_* variables are set.

config_file

Path of a YAML config (see "CONFIGURATION"; has_config_file is its predicate). Read at construction and again whenever its mtime moves. Give either this or "config_loader", not both. A path that is not an existing file dies at construction (config file not found: ...); a file that disappears later keeps the config in force (see "maybe_reload_config").

config_loader

A code ref that returns the config as a hashref, instead of a config_file. It is called with the Skeid object, at construction and then from call_function at most once per "config_reload_interval".

It may return ($config, $version). A defined version is the change detector: the same version as last applied means nothing changed and the config is not applied again. Without a version Skeid digests the returned structure (hash keys sorted) instead. Either way an unchanged config is a no-op, and a changed one whose nodes section is unchanged keeps the node list -- its inventory generation, its running capacity probes and any health set through the admin API. A code ref inside the config (a custom probe's code) digests by identity, so a loader that builds a fresh closure every call should return a version.

A config file follows the same rules, minus the throttle for a new file version: an mtime change whose content is unchanged applies nothing. A version that fails to parse is retried with the same bounded back-off as a failing loader, while a different mtime is read immediately. An explicit config.reload does too, so a value taken from the environment (admin.api_key_env, a usage store's password_env) is read again only when the config itself changes.

config_reload_interval

The least time, in seconds, between two runs of a config_loader (default 1; fractions allowed; 0 runs it on every dispatch, as before skeid #38).

A loader-based config is re-read from call_function, which every request passes through several times. Without a throttle every request would rerun the loader. A config file normally pays only one stat per dispatch; only retries of one failed file version are throttled.

last_reload_error

Why the last config reload failed, or undef when the last one succeeded. A failed reload keeps the previous config in force; see "maybe_reload_config". last_reload_error_at is when (epoch seconds), reload_failures how many reloads in a row have failed.

capacity_max_age_ms

How long a capacity reading is trusted, in milliseconds (default 5000, 0 disables expiry).

A stale reading is worse than none: it describes a node as it was, and admission acts on it as if it were now. Past this age a reading is ignored and inflight decides again, which is the same behaviour as having configured no probe at all (ADR 0009).

worker_count

How many worker processes share this configuration (default 1).

inflight and max_conns are per-process, so N workers would each admit up to max_conns to a node that can only serve one number — max_conns: 8 across 4 workers would permit 32, silently. Setting this makes each worker take its share instead (ADR 0010). Anything on a timer is spread the same way, so the process group's aggregate poll rate stays what was configured.

frontend_count

How many separate Skeid frontends stand in front of one node (default 1).

worker_count partitions max_conns across the prefork workers of one process; this partitions it across the distinct Skeid hosts sharing a node — a number only the operator knows, because two frontends are two processes on two machines with nothing between them to count each other's inflight (there is no probe and no shared state; that absence is the whole reason this has to be declared rather than detected). With F frontends each admits its share max_conns/F, so the group as a whole never admits more than the node was configured for (ADR 0009, ADR 0012).

It composes with worker_count: the two divisors multiply, so one worker's share is max_conns/(F*N) — partitioned across frontends first, then across workers, which integer division makes the same thing either way. Unlike the worker divisor it does not scale anything on a timer: separate frontends each run their own probes and hold their own vault token, the same reason vault renewal is not scaled per worker (ADR 0010).

Explicit and opt-in: the default 1 is today's behaviour, unchanged. An operator who runs two frontends and forgets to say so over-admits the node by a factor of two, silently — which is exactly the failure this field exists to prevent.

registry_enabled

Whether this Skeid publishes its registry snapshot on GET /skeid/registry/snapshot (default off; ADR 0017). Set by the config's registry block:

registry:
  enabled: true
  secret_env: SKEID_REGISTRY_SECRET   # required when enabled; the HMAC key
  ttl_s: 10                           # how long a snapshot may be believed
  instance_id: skeid-b                # default: the hostname
  error_window_s: 60                  # how far back errors_in_window counts
  read_key_env: SKEID_REGISTRY_READ_KEY   # optional; a bearer for the snapshot route only

The fronting tier reads the snapshot with Langertha::Skeid::CapacityProbe::Registry. The secret is taken from the environment only and kept in memory (ADR 0003); an enabled registry whose variable is empty or shorter than 32 bytes does not load, because a snapshot is never published unsigned or weakly signed. Neither does one without a credential to read it: an admin API key or a registry read key.

read_key_env names the variable holding the registry read key (skeid #49). The snapshot route accepts it as a bearer token besides the admin API key; it opens no other route, so a fronting tier holding it cannot add nodes or flip health. A named variable that is empty does not load. registry_secret, registry_read_key, registry_ttl_s, registry_instance_id and registry_error_window_s hold the other fields.

METHODS

add_node

$skeid->add_node(id => 'gpu-1', url => 'http://gpu-1:8000/v1', model => 'qwen3-32b',
  tags => ['local'], max_conns => 8);

Adds a node, replacing any node with the same id. Takes the fields of a config nodes entry and fills in their defaults (see "nodes"). Returns 1. Croaks without id or url, on an unknown engine, and on a registry capacity block that "validate_config" in Langertha::Skeid::CapacityProbe::Registry rejects. The nodes.add function and POST /skeid/nodes call it.

remove_node

my $removed = $skeid->remove_node('gpu-1');

Removes a node, and with it its capacity reading and failure history, so a node later added under the same id starts clean. Returns 1 when a node was removed, else 0.

normalize_tags

my $tags = Langertha::Skeid->normalize_tags(['Local', 'gb10']);
my $tags = Langertha::Skeid->normalize_tags('local, gb10');

Tags are lowercased, trimmed, de-duplicated and kept in the order first seen. A plain string is accepted and split on commas or whitespace, because a hand-written config says tags: local, gb10 at least as often as it says a YAML list.

list_nodes

my $nodes = $skeid->list_nodes;

The inventory as an arrayref of shallow copies of the node hashes.

list_models

my $models = $skeid->list_models(api_key_id => $key_id);
# [ { model => 'house-model', engine => '' }, { model => 'qwen3-32b', engine => 'vllm' } ]

The names a client can put in model, sorted by name, each once: the distinct model of the nodes plus the alias names. A node without a model lists nothing -- it matches any name, so none reaches it in particular. A name that is both a node model and an alias is one entry. engine is the first node's engine for a node model, empty for an alias.

The list is narrowed by the policy of api_key_id (none means the default policy), so a key is never shown a name "route_plan" would refuse it: a name its models do not grant is left out, an alias with every tier denied is left out, and a node model that only nodes carrying a denied tag serve is left out. This backs /v1/models and /api/tags.

set_node_health

$skeid->set_node_health('gpu-1', 0);   # take it out of rotation

Sets the operator health flag. An unhealthy node is not eligible for routing. Health is operator state: no error, timeout or 429 ever sets it. Returns 1 when the node exists, else 0. Setting the value it already has changes nothing, round-robin position included.

set_model_pricing

$skeid->set_model_pricing('gpt-4o-mini', {
  input_per_million        => 0.15,
  output_per_million       => 0.60,
  cached_input_per_million => 0.075,   # optional
  cache_write_per_million  => 0.1875,  # optional
});

Sets the pricing rule for one served model ('*' is the fallback) and returns it as stored. Absent input and output rates are 0. Croaks without a model or a pricing hash, and on a cache rate that is not a number >= 0. On a Langertha that cannot price prompt-cache tokens the cache rates are dropped, with one warning per process.

pricing_for_model

my $rule = $skeid->pricing_for_model('gpt-4o-mini');

The rule for a model, else the '*' rule, else a rule that prices everything at 0.

reload_config

my $config = $skeid->reload_config;

Reads the config source and applies it, returning the config hashref. A config whose fingerprint matches the one last applied changes nothing. Dies when the source cannot be read or the config does not apply; the previous config then stays in force and "reload_status" reports the failure. The config.reload function calls it; the request path uses "maybe_reload_config", which never dies.

reload_status

my $status = $skeid->reload_status;
# { ok => 0, error => '...', failed_at => '2026-09-25T16:08:30Z', failures => 3 }

Whether the last config reload succeeded, and if not, why, when, and how many times in a row. Served by the admin route GET /skeid/config; the public /health carries only ok, failed_at and failures, never the message, which can name customers and models.

set_admin_api_key

$skeid->set_admin_api_key($key);   # explicit: wins over the config, on every reload
$skeid->set_admin_api_key('');     # back to the config's key, or SKEID_ADMIN_API_KEY

Sets the explicit admin API key -- what skeid serve --admin-api-key and build_app(admin_api_key => ...) set, and what an admin_api_key passed to new is. It outranks the config: a reload re-applies the order of "admin" under it instead of replacing it. An empty or undefined key clears it, and the key the config resolved to takes over at once. Returns the admin API key now in force.

maybe_reload_config

my $applied = $skeid->maybe_reload_config;

Reloads the config if its source may have changed: the file's mtime moved, or a config_loader is due ("config_reload_interval"). Returns true when a changed config was applied. call_function runs it on every dispatch, so it never dies: a reload that fails is logged, recorded in "reload_status", and the request goes on under the config that was in force before (the reload is all or nothing). A failing source is then retried with a back-off -- the interval doubling per failure, from at least a second up to a minute. A loader that keeps returning the same broken config is not applied again; a changed config-file mtime bypasses the failed version's retry window. A config file that has disappeared is not a reload at all: the config in force stays, and the loss is warned once until the file is back. Only construction and an explicit config.reload still die on a bad or missing config.

manifest_for_key

my $json = $skeid->manifest_for_key($api_key_id);   # UTF-8 JSON bytes, or undef

The provider manifest published to a customer key id, as canonical JSON, or undef when the key has no manifest: grant, the manifest is disabled, or this Langertha has no Langertha::Manifest. Built at config load; see "Provider Manifest".

configure_usage_store

$skeid->configure_usage_store({ backend => 'jsonlog', path => '/var/log/skeid/events/' });

Normalizes a usage_store config and, when it differs from the current one, prepares the new store before disconnecting the old. Returns the normalized config. Croaks on a config "normalize_config" in Langertha::Skeid::UsageStore rejects and on a store that fails to prepare, leaving the old store in place. The usage.configure function calls it.

supported_engine_ids

my $ids = $skeid->supported_engine_ids;   # [ 'aki', 'akiopenai', 'anthropic', ... ]

The engine ids a node may name, sorted: a built-in list, plus whatever Langertha->available_engine_ids reports when the installed Langertha has it.

normalize_engine_id

my $id = $skeid->normalize_engine_id('Langertha::Engine::vLLM');   # 'vllm'

Lowercases an engine name and strips a Langertha::Engine:: or LangerthaX::Engine:: prefix. Returns the empty string for an empty value and croaks on an id that is not in "supported_engine_ids".

record_usage

my $res = $skeid->record_usage(
  api_key_id => $key_id, model => 'qwen3-32b', requested_model => 'house-model',
  node_id => 'gpu-1', status_code => 200, ok => 1, duration_ms => 840,
  metrics => $normalized,   # from normalize_metrics
);

Builds one usage event and hands it to the sink: the "store_usage_event" callback, a subclass's _store_usage_event, or the configured store. Returns the sink's answer, or { ok => 0, enabled => 0, error => 'usage_store not configured' } without building an event when there is no sink.

The event carries created_at, request_id, api_format, endpoint, api_key_id, provider, engine, model (served), requested_model (default model), node_id, route_url, status_code, ok, duration_ms, error_type, error_message, and from metrics the token counts (input_tokens, output_tokens, total_tokens, cached_tokens, cache_write_tokens), tool_calls and the costs (cost_input_usd, cost_output_usd, cost_total_usd, cost_cache_read_usd, cost_cache_write_usd). Costs arrive already priced; nothing is priced here. content_bytes is set only when given. The usage.record function calls it.

flush_usage

my $res = $skeid->flush_usage;

Writes the usage events a write-behind store still holds (see "flush" in Langertha::Skeid::UsageStore::DBI) and returns its answer, or { ok => 1, written => 0 } when the store holds nothing back. bin/skeid calls it when the server stops, and an application embedding the proxy should call it before it exits: queued events live in memory only. Replacing the store on a reload and destroying this object flush as well.

usage_report

my $report = $skeid->usage_report(since => '2026-09-01T00:00:00Z', api_key_id => $id,
  model => 'qwen3-32b', limit => 50);

Asks the "query_usage_report" callback, a subclass's _query_usage_report or the configured store for a report (its shape: Langertha::Skeid::UsageStore). Empty filters are dropped; limit defaults to 20 and is capped at 500. Without a store it returns { ok => 0, enabled => 0, error => 'usage_store not configured' }. The usage.report function, GET /skeid/usage and skeid usage call it.

estimate_cost

my $cost = $skeid->estimate_cost(model => 'gpt-4o-mini',
  usage => { prompt_tokens => 1000, completion_tokens => 200 });

Prices a usage with the model's rule ("pricing_for_model") or an explicit pricing rule, and returns the hash form of the Langertha::Cost. The usage is a hash, a Langertha::Usage, or read from a response. The metrics.estimate_cost function calls it.

normalize_metrics

my $metrics = $skeid->normalize_metrics(model => 'gpt-4o-mini', response => $upstream_body,
  tool_calls => \@calls, duration_ms => 840, engine => 'openaibase', route => '/v1/chat/completions');

What a usage event is built from: the usage (as for "estimate_cost") priced into a Langertha::UsageRecord, returned as its hash, with tool_calls counted by name. Adds cached_tokens and cache_write_tokens when the installed Langertha::Usage reads them. Optional provider, started_at, finished_at and pricing_version are carried into the record. The metrics.normalize function calls it; the proxy runs it on every upstream answer, streamed or not.

worker_max_conns

my $share = $skeid->worker_max_conns($node);

This process's share of a node's max_conns (ADR 0010, ADR 0012). With one worker and one frontend that is the configured value; otherwise it is the configured value divided by the number of processes sharing the node — frontend_count separate Skeid hosts times the worker_count prefork workers of this one — so the group as a whole never admits more than was asked for.

Never less than 1 when a limit is set: a process that may admit nothing is a process that does nothing. That means a max_conns below that combined process count cannot be honoured, and "worker_share_warnings" is what says so out loud.

worker_share_warnings

warn $_ for @{ $skeid->worker_share_warnings };

The nodes whose max_conns cannot be divided among the processes sharing them without exceeding it — frontend_count frontends times worker_count workers. Returned rather than warned so the caller decides where they go; bin/skeid prints them at startup.

Silence here would be the bad kind: the operator wrote a number, and the process group is about to ignore it.

set_model_alias

$skeid->set_model_alias('our-fast-model', {
  tiers => [
    { tags => ['local'], model => 'qwen3-32b',              wait_ms => 200 },
    { tags => ['cloud'], model => 'llama-3.3-70b-versatile' },
  ],
});

Defines a client-facing model name as an ordered list of tiers. A bare arrayref of tiers is accepted as shorthand for { tiers => [...] }.

Per tier: tags selects nodes, model is the model actually asked of them (defaulting to the alias name itself, for the case where nodes carry that name), engine optionally constrains the engine, and wait_ms is how long to wait for capacity in this tier before falling through to the next.

wait_ms defaults to 0. Writing tiers means "try here, then there"; waiting is the exception you opt into, and a tier that waits by default would send traffic to a paid cloud only after a delay nobody asked for -- or, worse, make a cheap tier look slow.

set_policy

$skeid->set_policy('standard-local-only', { deny_tags => ['cloud'] });

Defines a named policy profile. Profiles are the "standard setups" most customers take unchanged; a customer needing something precise gets the profile plus overrides, or a profile of their own.

resolve_policy

my $policy = $skeid->resolve_policy({ models => ['house-model'], deny_tags => ['cloud'] });

Turns a policy spec into the immutable form routing uses: models (or aliases) becomes a lookup hash of the requested model names the key may ask for, absent meaning all of them, and deny_tags becomes a normalized tag list.

key_id_for_key

my $id = Langertha::Skeid->key_id_for_key('sk-alice-secret');   # k_5f0e...

The customer key id derived from the presented API key. This is the name a keys: entry has to use, and skeid keyid prints it, because the config must be able to name a customer without holding that customer's key.

It is the full SHA-1 hex digest of the key (160 bits), not a secret: it identifies, it does not authenticate. What authenticates is that the caller presented the key it was derived from.

Before ADR 0016 the id was the first 12 hex digits of the same digest (k_5f0e1a2b3c4d), so an old id is the prefix of the new one. A config may still name a customer by its short id: it matches the key whose full id starts with it, with a one-time deprecation warning at config load, and a config that lists both a short id and a full id it is the prefix of fails to load. Usage events keep the id they were recorded under -- events from before the change carry the short id, later ones the full id; nothing is migrated.

policy_for_key

my $policy = $skeid->policy_for_key($key_id);   # k_5f0e...

The policy a customer key routes under. Unlisted keys take the default policy, which is what makes a deployment with ten thousand identically-configured customers a config with zero key entries. Returns undef when no policies are configured at all.

key_id_for_name

my $id = $skeid->key_id_for_name('alice');   # k_5f0e1a2b3c4de5f60718293a4b5c6d7e8f901a2b, or undef

The customer key id a readable name maps to under the config names: section, or undef when the name is not registered. The registry is a config-authoring convenience -- it lets a keys: entry be written by name and confines key rotation to one line -- and is resolved to ids at config load, so it never appears on the request path.

route_plan

my $plan = $skeid->route_plan(model => 'our-fast-model', api_key_id => $key_id);
# { tiers => [...], permitted => 1, reason => '' }

The ordered tiers to try for a requested model, under the policy of the key that asked. A model with no alias yields a single implicit tier that selects on the name itself and inherits the global route_wait_timeout_ms, which is what makes an aliasless config behave exactly as it did before aliases existed.

permitted is false when the policy does not grant this model, or when every tier of it was denied. Both mean the same thing to a caller — this key may not reach this model — and neither is a capacity problem, so they must not be reported as one.

Tiers carry the policy's deny_tags down into node selection. Dropping a denied tier is only the reporting half; without the node-level filter, a key denied cloud could still reach a cloud node by asking for its raw model name instead of the alias.

select_nodes

my $local = $skeid->select_nodes(tags => ['local']);

Nodes carrying every listed tag, as copies. No tags selects everything. Selection is by tag, never by node id, so a config can talk about local or cloud without naming machines. deny_tags drops every node carrying any of those tags. Health is not considered. The nodes.select function calls it.

pick_node

my $node = $skeid->pick_node(model => 'qwen3-32b', tags => ['local'], deny_tags => ['cloud']);

The next node for a selection by weighted round-robin (over the eligible nodes, ordered by id), skipping every node admission would refuse right now. Returns a copy of the node with inflight and route_key added, or nothing. Picking does not admit: the caller still has to "start_request", which can refuse when another request took the slot in between. The route.next function calls it.

route_state

my $state = $skeid->route_state(model => 'qwen3-32b', tags => ['local']);
# { model, engine, tags, deny_tags, eligible_count, available_count,
#   has_eligible, has_available }

How many nodes a selection matches (eligible) and how many of those admission would take now (available). No eligible node is a 503; eligible but none available is worth waiting for. The route.state function calls it.

set_capacity_reading

$skeid->set_capacity_reading('gpu-1', used => 6, limit => 8, source => 'prometheus');
$skeid->set_capacity_reading('groq-1', retry_after_ms => 2000, source => 'ratelimit');

Records what a probe found. One shape for every probe, so admission never learns where a number came from (ADR 0009):

  • used / limit — occupancy. limit 0 or absent means the probe measured something it cannot turn into a ceiling, so it does not constrain admission.

  • retry_after_ms — do not send anything here until it elapses. What a 429 means.

  • source — which probe, for reports. Never consulted by admission. quota -- which of a provider's quotas (requests, tokens) the reading is about -- likewise.

  • at — when the reading was taken (default now), and expires_at — an absolute moment after which it is dropped even inside "capacity_max_age_ms". The registry probe uses both: a snapshot is as old as its generated_at, and never outlives its own ttl.

  • interval_ms — how often this source reports, for a probe on a timer. Stored with the reading; it sets how long a tighter reading can hold a looser one from another source off (below). A passive observation (a response's rate-limit headers) passes none.

A reading only ever narrows what max_conns already allows, and expires after "capacity_max_age_ms". Probing is a background activity: calling this from a request handler is a bug unless the reading was a by-product of a response already in hand.

The tighter reading wins across sources, while it is current (ADR 0017). A source always replaces its own last reading. A reading from a different source is held off by the current one only when the current one is tighter (a pending backoff is tightest, then used/limit, then a reading without a limit) and either

  • it carries a pending backoff, or

  • it is younger than the longer of the two sources' interval_ms -- neither has had a full poll since the tighter reading was taken.

Otherwise the incoming reading replaces it. So a registry snapshot saying "empty" cannot lift a 429 backoff a response just recorded; a fresh, tight probe reading is not lifted by a roomy response or by a faster, looser probe before its own next poll; and a passive reading that is never refreshed (remaining: 0 with no reset, from the last response before traffic stopped) cannot keep a fresh probe out for longer than one of that probe's polls. Returns the reading in force afterwards.

capacity_reading

my $reading = $skeid->capacity_reading('gpu-1');   # or nothing

The node's current capacity reading, or nothing when no probe has reported or the last report has aged out. A backoff outlives the age limit: a provider that said "not for another 30 seconds" told us something that is still true.

forget_capacity

$skeid->forget_capacity('gpu-1');   # or all of them with no argument
$skeid->forget_capacity('gpu-1', source => 'registry');   # only if that source holds it

Drops probe readings, so inflight decides again. What a node removal calls, and what a probe calls when it can no longer reach its source — reporting nothing beats reporting last hour.

With source, the reading is dropped only when that source wrote it. A probe forgetting what it said must not wipe a backoff another source recorded (ADR 0017).

capacity_header_names

The response headers worth looking at, so a caller on the request path can pull those few by name instead of walking every header of every response.

observe_response_headers

$skeid->observe_response_headers('groq-1', \%headers, status => 429);

The zero-cost probe: commercial providers do not publish queue depth, but they do put their rate-limit state on every response Skeid already receives. Reading it costs no extra request (ADR 0009).

Providers meter requests and tokens separately, and for an LLM API the token budget is usually what runs out first. Both are read, and the one closest to exhausted decides — compared as a fraction, since the two are not the same unit.

Retry-After on a 429 becomes a backoff. It deliberately does not touch healthy: rate-limited is busy, not broken, and nothing would ever flip that back.

Returns the reading it recorded, or nothing when the response said nothing useful.

start_request

$skeid->start_request('gpu-1') or ...;   # refused

Admits one request to a node and counts it in inflight. Returns 1 when admitted, 0 when the node is unknown, unhealthy or full. Every 1 must be paired with a "finish_request" on every path -- errors, timeouts, client aborts -- or the node loses that slot until restart. The request.start function calls it.

finish_request

$skeid->finish_request('gpu-1', ok => 1, duration_ms => 840);
$skeid->finish_request('gpu-1', aborted => 1);   # the client hung up

Releases the slot "start_request" took and counts the outcome for "node_metrics". There are three outcomes: ok, failed (neither ok nor aborted), and aborted -- the client went away, which says nothing about the node. A failed request counts toward the registry snapshot's errors_in_window and last_failure_at; an aborted one is counted apart and toward neither. aborted wins over ok. Never touches healthy. Returns 1. The request.finish function calls it.

registry_snapshot

my $snapshot = $skeid->registry_snapshot;

What this Skeid publishes to a fronting tier (ADR 0017): schema version 1, instance, generated_at, ttl, workers, and per node id, tags, healthy, inflight, max_conns (this process's share), errors_in_window, last_failure_at and, when a current reading exists, capacity (used, limit, source, retry_after).

Built from a whitelist, so it cannot carry what it must not: no node URL, no key reference, no metadata, no customer key id, no policy, no usage. Langertha::Skeid::Registry signs it.

node_metrics

my $one = $skeid->node_metrics('gpu-1');
my $all = $skeid->node_metrics;

Volatile per-process counters for operations, never billed: node_id, inflight, started, ok, error, aborted, duration_ms_total, and capacity while a current reading exists. With no id, an arrayref with one entry per node. Served by GET /skeid/metrics/nodes.

call_function

my $result = $skeid->call_function('route.plan', { model => 'house-model', api_key_id => $id });

The command surface the proxy drives Skeid through. Every call first runs "maybe_reload_config". Croaks on an unknown name, on args that are not a hashref, and where a function names a required argument.

function            args                                  returns
metrics.estimate_cost  as estimate_cost                   estimate_cost
metrics.normalize   as normalize_metrics                  normalize_metrics
pricing.set         model, pricing                        the stored rule
nodes.add           a node                                { ok }
nodes.remove        id                                    { ok }
nodes.list                                                { nodes }
nodes.select        tags, deny_tags                       { nodes } (select_nodes)
nodes.set_health    id, healthy                           { ok }
nodes.metrics       id (optional)                         { metrics }
alias.set           name, alias (or the tiers inline)     { ok }
policy.set          name, policy (or the spec inline)     { ok }
policy.for_key      api_key_id                            { policy }
route.plan          model, engine, api_key_id             route_plan
route.next          model, engine, tags, deny_tags        { node } (pick_node)
route.state         model, engine, tags, deny_tags        route_state
capacity.set        id, and the reading                   { capacity }
capacity.get        id                                    { capacity }
capacity.observe    id, headers, status                   { capacity }
capacity.forget     id (optional)                         { ok }
request.start       id                                    { ok }
request.finish      id, ok|aborted, duration_ms                 { ok }
config.reload                                             { config } (dies on failure)
config.status                                             reload_status
usage.record        as record_usage                       record_usage
usage.report        since, api_key_id, model, limit       usage_report
usage.configure     usage_store (or the config inline)    { usage_store }
engines.list                                              { engines }

Add a function here rather than reaching into the object from the app.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha-skeid/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <torsten@raudssus.de> https://raudssus.de/

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.