EN | JA

Broker sizing & scaling

kanade runs on a single NATS + JetStream broker. As the fleet grows (hundreds → thousands of agents) the broker, not the agents, is the scaling bottleneck. This page covers the per-agent footprint, what the #512 work changed, the single-node limits to watch, and how to capture real numbers during a scale-up so you can decide whether to add a startup splay or grow/cluster the broker.

Per-agent consumer footprint

Each running agent holds a handful of JetStream consumers — one ordered push consumer per KV watch, plus its durable command-replay consumer:

ConsumerSource
agent_config watch (key-filtered)config_supervisor
agent_groups watch (membership → effective config)config_supervisor
agent_groups watch (membership → subscriptions)groups.rs
schedules watchlocal_scheduler
jobs watchlocal_scheduler
fleet_config freeze watch (single key)local_scheduler
EXEC durable replaycommand_replay

That is ~7 consumers per agent. The count is roughly fixed per agent — it does not shrink with key-filtering — so the broker-side total grows linearly with the fleet:

Fleet size~Consumers on the broker
15~105
500~3,500
3,000~21,000

What v0.43.96 changed (and what it did not)

#512 (shipped in v0.43.96) attacked the super-linear costs, not the consumer count:

  • #832 — agent_config key-filtered watch. The agent watches only global, pcs.<self>, and its groups.<g> keys instead of the whole bucket (watch_all). This removes two blow-ups:
    • Per-PC write fan-out: a write to one pcs.<id> no longer reaches all N agents (it used to, with N−1 classifying-and-dropping it).
    • Reconnect re-sync storm: re-sync is now 3–5 direct gets per agent instead of a keys() + per-key walk over the whole bucket — aggregate O(N²) → O(N).
  • #839 — single-key freeze watch. fleet_config holds only KEY_FREEZE, so the freeze watcher uses watch(KEY_FREEZE) instead of watch_all + a client-side filter.

What is unchanged: the ~7-consumers-per-agent footprint, and the schedules / jobs watch_all (those are a shared catalog every agent must evaluate for targeting — removing them needs server-side targeting, a separate change). So at 3,000 agents you still provision for ~21,000 consumers and the connection/consumer-create burst of a synchronised reconnect.

Single-node limits and the reconnect herd

Two things to size for on a single-node JetStream:

  1. Steady-state footprint — ~21,000 consumers at 3,000 agents live in the JetStream meta (Raft) layer and cost memory + file handles. Size the broker host's RAM and max_file/max_memory JetStream limits accordingly, and watch the meta layer's health.

  2. The reconnect herd — when many agents reconnect at the same instant (broker restart, the morning power-on wave, a network event hitting many PCs), they re-establish connections and recreate their consumers in a burst. Key-filtering already cut each agent's re-sync to a few cheap gets, so the dangerous O(N²) read-storm is gone — but the connection + consumer-create burst is still O(N) and synchronised.

    Note that random, unsynchronised reconnects (one laptop's wifi flap) are not a herd — only fleet-wide synchronised events are.

Storage: the 50 GB file store, per-resource caps, and retention

RAM and consumers are only half the sizing story; the other half is disk. All JetStream data lives under one directory (store_dir: C:/ProgramData/Kanade/nats/jetstream) bounded broker-wide by max_file_store: 50GB (configs/nats-server.conf) — a soft limit shared by every stream and object store, not a per-stream one. What happens as it fills:

  1. Resources with their own retention evict oldest-first (DiscardPolicy::Old) and stay healthy.
  2. A resource without caps cannot evict, so once the file store is full, every JetStream publish fleet-wide starts failing with "insufficient storage resources available" (error 10077) — agent result uploads, collect-bundle uploads, publishes, KV puts. Reads keep working; the write side is what dies. Uncapped resources also squeeze the capped ones out of their fair share.

So every resource must carry a cap, and it does:

  • Streams (RESULTS / INVENTORY / AUDIT / OBS_EVENTS / NOTIFICATIONS / EXEC / EVENTS) each have max_age + max_bytes (bootstrap.rs, ~5.3 GiB reserved total). They are transport + replay buffers — the durable record is the backend's SQLite, which is why the caps can be tighter than the history an operator expects to see.

  • Object stores are capped per bucket by max_bytes, tunable from the SPA (Settings → server → Object store disk caps, #1247) with no restart; blank fields fall back to these built-in defaults:

    BucketDefault capHolds
    result_output1,024 MiBoversized stdout/stderr blobs (projected into SQLite within seconds)
    agent_releases2,048 MiBagent exes (~20 versions)
    app_packages5,120 MiBoperator-curated installers
    scripts256 MiBmanifest script bodies
    collections5,120 MiBcollect-job bundles (also max_age, tunable via Collected-bundle retention)

    The backend reconciles the configured values onto the backing OBJ_* streams at every boot and on every save, which is also how caps reach buckets created before the cap existed. Total default reservation ≈ 13.5 GiB — sized, with the streams, to sit well inside 50 GB.

  • Recovery. If a stream or bucket has drifted from its expected config or is corrupted, repair it with the nats CLI and an administrative credential, not with kanade. The backend recreates anything missing the next time it starts. This works even when the backend cannot start (a drifted stream config makes its startup bootstrap fail), because the scripts talk to the broker directly. Stop kanade-backend first (its projectors hold durable consumers) and start it again afterwards:

    # delete one stream / KV bucket / object store (asks you to type the name)
    ./scripts/ops/jetstream-delete.ps1 -Kind stream -Name RESULTS -Server nats://127.0.0.1:4222 -Creds ./admin.creds
    
    # wipe everything kanade uses: dry run first, then add -Yes
    ./scripts/ops/jetstream-reset.ps1 -Server nats://127.0.0.1:4222 -Creds ./admin.creds
    ./scripts/ops/jetstream-reset.ps1 -Server nats://127.0.0.1:4222 -Creds ./admin.creds -Yes
    

    Both need the nats CLI on PATH and also read NATS_URL, NATS_CREDS, NATS_USER and NATS_PASSWORD. Deleting a resource deletes its data.

  • SQLite (the projection) is not unbounded either: the backend cleanup task prunes on a 5-minute tick in bounded batches — execution_results / executions / obs_events / inventory_history 90 d, audit_log 365 d, host_perf_samples 30 d, process_perf_samples 7 d. Every DB window is deliberately longer than the matching stream window, so anything the stream can replay is guaranteed to already be in SQLite. Remaining tables are upsert/replace-shaped (live-state sized), and dead agents are pruned by agent_prune_days (also a ServerSettings knob).

Live usage per resource (used bytes vs cap) is on the SPA's JetStream page — it reads the broker's current config, so a changed cap shows up there on the next load.

Levers, in order of preference

  1. Broker sizing first. Give the single node enough RAM / file limits for the steady-state consumer count at your target N. This is the primary lever; everything else is secondary.
  2. Startup splay — only if measured. A deterministic per-PC delay (hash(pc_id)) before the reconnect re-sync would smear the consumer-create / connection burst across a window. It is not in the product yet by design: PR1 already removed the quadratic term, and async-nats' reconnect backoff + nats_retry's ±25% jitter already spread the burst somewhat. A splay also adds latency to every single reconnect (including herd-less blips), so it is a net cost unless a herd actually stresses the broker. Decide from data (next section): if a synchronised event shows consumer-create latency, connection backlog, or JetStream API errors, add the splay (gated to the first sync after a Disconnected → Connected, capped a few seconds).
  3. Reduce consumers per agent. Fold watches where possible (the freeze watch is already a single key; the two agent_groups watches are a candidate to merge) and keep KV history shallow (agent_config / agent_groups are at history: 1).
  4. Cluster JetStream. Beyond what one node can hold, move to a JetStream cluster. This is the last resort and the biggest change.

Measuring at scale (retro-analysis)

You usually cannot watch a production broker live during a ramp. Capture the numbers instead, with the bundled collect: job, and review the bundle afterwards.

Run it on the backend / NATS host, ideally during or right after a synchronised reconnect event (the herd moment is what decides the splay question):

kanade exec collect-broker-health --pcs <backend-host-id>

The job (configs/jobs/collect-broker-health.yaml) samples the broker over ~3 minutes and uploads a bundle to OBJECT_COLLECTIONS; download it from the SPA Collect page (or hand the zip to your reviewer). It is read-only and needs zero pre-setup: it reads NATS' unauthenticated HTTP monitoring port (default 8222, the /jsz endpoint), so no nats CLI on the SYSTEM PATH and no token are required. Tune the window with KANADE_BH_SAMPLES / KANADE_BH_INTERVAL_SEC (and KANADE_BH_MON_PORT if the broker's http_port differs) on the target if needed.

The bundle contains:

  • Time series (connz-*.json, jsz-*.json) — connection count (/connz) and JetStream consumer count (/jsz) per sample. A spike here at the herd moment is the splay signal.
  • Consumer footprint (jsz-full.json) — /jsz?consumers=true&streams=true: the full consumer list; confirms the ~7N total and which streams hold them.
  • Server health/resources (varz.json, healthz.json) — /varz (memory, CPU, connections, slow consumers) and /healthz status.
  • Log tails (redacted) — backend and nats-server.

What to look for

  • Smooth consumer/connection counts across the samples, healthy /healthz, head-room on mem/cpu in /varz → the broker absorbed the event; no splay needed, just keep sizing ahead of N.
  • Spiky connection backlog or consumer-create latency at the event, JetStream API errors, or mem/cpu pegged → the herd is real → add the startup splay (lever 2) and/or grow the broker.

The #828 downgrade-flap regression is checked separately from the OBS_EVENTS agent_update timeline (via the backend API), not by this job — that data is already queryable fleet-wide.