NAME

Data::HashMap::Shared::Cookbook - Practical recipes for shared-memory maps

DESCRIPTION

Copy-paste patterns for common cross-process tasks with Data::HashMap::Shared: counters, caches, rate limiting, liveness, dedup, and atomic state. Runnable versions of several live under eg/ in the distribution. The /tmp paths are placeholders: keep a real map in a directory only your application can write ("SECURITY" in Data::HashMap::Shared).

Pick a variant by key/value type (see "CHOOSING A VARIANT") and replace the shm_xx_* keyword prefix accordingly (shm_ii_, shm_ss_, shm_si_, ...).

RECIPES

Shared request/error counters

Pre-forked workers tallying into one map. incr/incr_by are atomic, so no update is lost.

use Data::HashMap::Shared::SI;   # string metric name -> int64 count
my $stats = Data::HashMap::Shared::SI->new('/tmp/stats.shm', 1000);

# in each worker, per request:
shm_si_incr    $stats, 'requests';
shm_si_incr_by $stats, 'bytes_out', $n;
shm_si_incr    $stats, 'errors' if $failed;

# anytime, from any process:
my %snapshot = %{ $stats->to_hash };

For write-heavy counters spread the load with new_sharded (see eg/sharded_counter.pl).

High-water marks and leaderboards

max/min store max($current, $desired) / min(...) atomically, so concurrent updates never lose a higher score. A missing key is inserted as the given value.

use Data::HashMap::Shared::SI;   # player -> high score
my $board = Data::HashMap::Shared::SI->new('/tmp/board.shm', 10_000);

shm_si_max $board, $player, $score;   # keeps only the best, race-free

my $scores = $board->to_hash;
my @top = sort { $scores->{$b} <=> $scores->{$a} } keys %$scores;
@top = @top[0 .. 9] if @top > 10;

Full example: eg/leaderboard.pl.

Cache-aside memoization (compute once)

get_or_set stores a key's value once; racing callers all receive the same stored value, so an expensive computation is shared across the fleet.

use Data::HashMap::Shared::IS;   # int id -> string result
my $cache = Data::HashMap::Shared::IS->new('/tmp/memo.shm', 100_000);

sub lookup {
    my $id = shift;
    my $hit = shm_is_get $cache, $id;
    return $hit if defined $hit;                 # fast path
    my $fresh  = compute($id);
    my $stored = shm_is_get_or_set $cache, $id, $fresh;  # first writer wins
    return defined $stored ? $stored : $fresh;
}

get_or_set returns undef when there is no room; fall back to the computed value as above, and give a long-lived cache a $max_size.

Full example: eg/memoize.pl.

LRU cache with bounded memory

Pass $max_size to cap live entries; the least-recently-used entry is evicted on insert.

use Data::HashMap::Shared::SS;
# max_entries=1_000_000 table, evict once 100_000 entries are live:
my $cache = Data::HashMap::Shared::SS->new('/tmp/cache.shm', 1_000_000, 100_000);

shm_ss_put $cache, $key, $value;          # auto-evicts LRU when full
my $v = shm_ss_get $cache, $key;          # also refreshes recency (clock bit)
my $evicted = $cache->stats->{evictions};

For large values, size the string arena for the working set (->new($path, $max_entries, $max_size, 0, 0, $arena_cap)): an arena that only just holds it evicts on nearly every insert.

Per-key TTL and dedup / idempotency

A default TTL expires entries lazily on access; add inserts only if absent. Together they make a "process each id once within a window" guard.

use Data::HashMap::Shared::IS;   # event id -> marker, 300s TTL
my $seen = Data::HashMap::Shared::IS->new('/tmp/seen.shm', 1_000_000, 0, 300);

if ($seen->add($event_id, '1')) {   # true only the first time (then TTL'd)
    handle($event_id);
} # else: already processed recently, or no room

Event ids never repeat, so the table fills with expired ids; a timer keeps them from lengthening every probe:

# once a second: the 2097152 slots this map can grow to, over a 300s TTL
my ($flushed, $done) = $seen->flush_expired_partial(7000);

put_ttl/set_ttl set a per-key TTL; persist makes a key permanent; ttl_remaining reports seconds left.

Worker heartbeat / liveness registry

Each worker refreshes a TTL'd heartbeat; a dead worker stops refreshing and its entry expires, so a supervisor sees only live workers.

use Data::HashMap::Shared::IS;   # pid -> status, 3s TTL for a 1s refresh
my $reg = Data::HashMap::Shared::IS->new('/tmp/live.shm', 4096, 0, 3);

# worker loop:
shm_is_put $reg, $$, 'ok';        # every second; resets the TTL

# supervisor:
$reg->flush_expired;              # drop stale heartbeats
my @live = $reg->keys;

Full example: eg/heartbeat.pl.

Sliding-window rate limiting

A counter per subject with a TTL window; the first hit creates it and later hits increment it. Each incr refreshes the TTL, so the window slides: a client that keeps retrying stays limited until it backs off for a full window.

use Data::HashMap::Shared::SI;   # subject -> hits, 60s window
my $rl = Data::HashMap::Shared::SI->new('/tmp/rl.shm', 100_000, 0, 60);

sub allow {
    my $who = shift;
    # auto-creates at 1 with the default TTL; croaks when there is no room
    my $n = eval { shm_si_incr $rl, $who } // return 0;
    return $n <= 100;                # cap per window
}

incr croaks only when there is no room, as in a flood of new subjects inside one window; allow then refuses rather than dies. Flush expired subjects on a timer, with a slice that cycles the whole table inside one window, or give the map a $max_size and let eviction reclaim them:

# once a second: the 262144 slots this map can grow to, over the 60s window
my ($flushed, $done) = $rl->flush_expired_partial(5000);

When clients choose the subjects, hash them under a secret first ("SECURITY" in Data::HashMap::Shared).

Full example: eg/rate_limiter.pl.

Atomic state machine (compare-and-swap)

cas swaps only if the current value matches, so processes can drive a shared state machine without a lock.

use Data::HashMap::Shared::SI;   # job -> state code (0=idle 1=running 2=done)
my $jobs = Data::HashMap::Shared::SI->new('/tmp/jobs.shm', 100_000);
shm_si_add $jobs, $_, 0 for @job_ids;   # cas needs the key to exist first

# claim a job: idle(0) -> running(1), exactly one winner
if (shm_si_cas $jobs, $job, 0, 1) {
    run($job);
    shm_si_cas $jobs, $job, 1, 2;    # running -> done
}

cas_take atomically removes a key only if it matches (handy for one-shot claims); see also the work-queue pattern in eg/work_queue.pl.

CHOOSING A VARIANT

Variants are <key><value> where I16/I32/I are signed 16/32/64-bit integers and S is a byte string. Integer variants store values inline (no arena, fastest); string sides use the shared arena.

II    int64  -> int64     counters, ids-to-ids
I16   int16  -> int16     compact enums / small counters
I32   int32  -> int32
IS    int64  -> string    id -> blob/json (memoization)
I16S  int16  -> string
I32S  int32  -> string
SI    string -> int64     named counters, scores, rate limits
SI16  string -> int16
SI32  string -> int32
SS    string -> string    config, sessions, generic caches

Use the narrowest integer width that fits your range (values wrap at the variant's width). Counter ops exist only on integer-value variants.

SIZING

->new($path, $max_entries, $max_size, $ttl, $lru_skip, $arena_cap)

  • $max_entries -- the live-key count to run at; the table grows and shrinks up to it.

  • $max_size -- LRU cap (0 = no eviction).

  • $ttl -- default per-entry TTL in seconds (0 = none); combine with $max_size for a bounded TTL cache.

  • $lru_skip -- leave at 0 unless LRU promotions on a Zipfian write workload show up as lock contention.

  • $arena_cap -- string-storage bytes (default about 128 per entry): size it from string lengths rounded up to powers of two.

  • shards (new_sharded($prefix, $shards, ...)) -- independent maps with independent locks for write-heavy workloads; the sizing arguments above are per shard.

SEE ALSO

Data::HashMap::Shared for the full API; the eg/ directory for runnable versions of these recipes.

AUTHOR

vividsnow

LICENSE

Same terms as Perl itself.