Broker sizing & scaling
kanade runs on a single NATS + JetStream broker. As the fleet grows
(hundreds → thousands of agents) the broker, not the agents, is the
scaling bottleneck. This page covers the per-agent footprint, what the
#512 work changed, the single-node limits to watch, and how to capture
real numbers during a scale-up so you can decide whether to add a startup
splay or grow/cluster the broker.
Per-agent consumer footprint
Each running agent holds a handful of JetStream consumers — one ordered push consumer per KV watch, plus its durable command-replay consumer:
| Consumer | Source |
|---|---|
agent_config watch (key-filtered) | config_supervisor |
agent_groups watch (membership → effective config) | config_supervisor |
agent_groups watch (membership → subscriptions) | groups.rs |
schedules watch | local_scheduler |
jobs watch | local_scheduler |
fleet_config freeze watch (single key) | local_scheduler |
EXEC durable replay | command_replay |
That is ~7 consumers per agent. The count is roughly fixed per agent — it does not shrink with key-filtering — so the broker-side total grows linearly with the fleet:
| Fleet size | ~Consumers on the broker |
|---|---|
| 15 | ~105 |
| 500 | ~3,500 |
| 3,000 | ~21,000 |
What v0.43.96 changed (and what it did not)
#512 (shipped in v0.43.96) attacked the super-linear costs, not the
consumer count:
#832—agent_configkey-filtered watch. The agent watches onlyglobal,pcs.<self>, and itsgroups.<g>keys instead of the whole bucket (watch_all). This removes two blow-ups:- Per-PC write fan-out: a write to one
pcs.<id>no longer reaches all N agents (it used to, with N−1 classifying-and-dropping it). - Reconnect re-sync storm: re-sync is now 3–5 direct
gets per agent instead of akeys()+ per-key walk over the whole bucket — aggregate O(N²) → O(N).
- Per-PC write fan-out: a write to one
#839— single-key freeze watch.fleet_configholds onlyKEY_FREEZE, so the freeze watcher useswatch(KEY_FREEZE)instead ofwatch_all+ a client-side filter.
What is unchanged: the ~7-consumers-per-agent footprint, and the
schedules / jobs watch_all (those are a shared catalog every agent
must evaluate for targeting — removing them needs server-side targeting,
a separate change). So at 3,000 agents you still provision for ~21,000
consumers and the connection/consumer-create burst of a synchronised
reconnect.
Single-node limits and the reconnect herd
Two things to size for on a single-node JetStream:
-
Steady-state footprint — ~21,000 consumers at 3,000 agents live in the JetStream meta (Raft) layer and cost memory + file handles. Size the broker host's RAM and
max_file/max_memoryJetStream limits accordingly, and watch the meta layer's health. -
The reconnect herd — when many agents reconnect at the same instant (broker restart, the morning power-on wave, a network event hitting many PCs), they re-establish connections and recreate their consumers in a burst. Key-filtering already cut each agent's re-sync to a few cheap
gets, so the dangerous O(N²) read-storm is gone — but the connection + consumer-create burst is still O(N) and synchronised.Note that random, unsynchronised reconnects (one laptop's wifi flap) are not a herd — only fleet-wide synchronised events are.
Storage: the 50 GB file store, per-resource caps, and retention
RAM and consumers are only half the sizing story; the other half is
disk. All JetStream data lives under one directory
(store_dir: C:/ProgramData/Kanade/nats/jetstream) bounded broker-wide
by max_file_store: 50GB (configs/nats-server.conf) — a soft
limit shared by every stream and object store, not a per-stream one.
What happens as it fills:
- Resources with their own retention evict oldest-first
(
DiscardPolicy::Old) and stay healthy. - A resource without caps cannot evict, so once the file store is full, every JetStream publish fleet-wide starts failing with "insufficient storage resources available" (error 10077) — agent result uploads, collect-bundle uploads, publishes, KV puts. Reads keep working; the write side is what dies. Uncapped resources also squeeze the capped ones out of their fair share.
So every resource must carry a cap, and it does:
-
Streams (
RESULTS/INVENTORY/AUDIT/OBS_EVENTS/NOTIFICATIONS/EXEC/EVENTS) each havemax_age+max_bytes(bootstrap.rs, ~5.3 GiB reserved total). They are transport + replay buffers — the durable record is the backend's SQLite, which is why the caps can be tighter than the history an operator expects to see. -
Object stores are capped per bucket by
max_bytes, tunable from the SPA (Settings → server → Object store disk caps, #1247) with no restart; blank fields fall back to these built-in defaults:Bucket Default cap Holds result_output1,024 MiB oversized stdout/stderr blobs (projected into SQLite within seconds) agent_releases2,048 MiB agent exes (~20 versions) app_packages5,120 MiB operator-curated installers scripts256 MiB manifest script bodies collections5,120 MiB collect-job bundles (also max_age, tunable via Collected-bundle retention)The backend reconciles the configured values onto the backing
OBJ_*streams at every boot and on every save, which is also how caps reach buckets created before the cap existed. Total default reservation ≈ 13.5 GiB — sized, with the streams, to sit well inside 50 GB. -
Recovery. If a stream or bucket has drifted from its expected config or is corrupted, repair it with the
natsCLI and an administrative credential, not withkanade. The backend recreates anything missing the next time it starts. This works even when the backend cannot start (a drifted stream config makes its startup bootstrap fail), because the scripts talk to the broker directly. Stopkanade-backendfirst (its projectors hold durable consumers) and start it again afterwards:# delete one stream / KV bucket / object store (asks you to type the name) ./scripts/ops/jetstream-delete.ps1 -Kind stream -Name RESULTS -Server nats://127.0.0.1:4222 -Creds ./admin.creds # wipe everything kanade uses: dry run first, then add -Yes ./scripts/ops/jetstream-reset.ps1 -Server nats://127.0.0.1:4222 -Creds ./admin.creds ./scripts/ops/jetstream-reset.ps1 -Server nats://127.0.0.1:4222 -Creds ./admin.creds -YesBoth need the
natsCLI onPATHand also readNATS_URL,NATS_CREDS,NATS_USERandNATS_PASSWORD. Deleting a resource deletes its data. -
SQLite (the projection) is not unbounded either: the backend cleanup task prunes on a 5-minute tick in bounded batches —
execution_results/executions/obs_events/inventory_history90 d,audit_log365 d,host_perf_samples30 d,process_perf_samples7 d. Every DB window is deliberately longer than the matching stream window, so anything the stream can replay is guaranteed to already be in SQLite. Remaining tables are upsert/replace-shaped (live-state sized), and dead agents are pruned byagent_prune_days(also a ServerSettings knob).
Live usage per resource (used bytes vs cap) is on the SPA's JetStream page — it reads the broker's current config, so a changed cap shows up there on the next load.
Levers, in order of preference
- Broker sizing first. Give the single node enough RAM / file limits for the steady-state consumer count at your target N. This is the primary lever; everything else is secondary.
- Startup splay — only if measured. A deterministic per-PC delay
(
hash(pc_id)) before the reconnect re-sync would smear the consumer-create / connection burst across a window. It is not in the product yet by design: PR1 already removed the quadratic term, and async-nats' reconnect backoff +nats_retry's ±25% jitter already spread the burst somewhat. A splay also adds latency to every single reconnect (including herd-less blips), so it is a net cost unless a herd actually stresses the broker. Decide from data (next section): if a synchronised event shows consumer-create latency, connection backlog, or JetStream API errors, add the splay (gated to the first sync after aDisconnected → Connected, capped a few seconds). - Reduce consumers per agent. Fold watches where possible (the freeze
watch is already a single key; the two
agent_groupswatches are a candidate to merge) and keep KV history shallow (agent_config/agent_groupsare athistory: 1). - Cluster JetStream. Beyond what one node can hold, move to a JetStream cluster. This is the last resort and the biggest change.
Measuring at scale (retro-analysis)
You usually cannot watch a production broker live during a ramp. Capture
the numbers instead, with the bundled collect: job, and review the
bundle afterwards.
Run it on the backend / NATS host, ideally during or right after a synchronised reconnect event (the herd moment is what decides the splay question):
kanade exec collect-broker-health --pcs <backend-host-id>
The job (configs/jobs/collect-broker-health.yaml) samples the broker
over ~3 minutes and uploads a bundle to OBJECT_COLLECTIONS; download it
from the SPA Collect page (or hand the zip to your reviewer). It is
read-only and needs zero pre-setup: it reads NATS' unauthenticated
HTTP monitoring port (default 8222, the /jsz endpoint), so no nats CLI on the SYSTEM
PATH and no token are required. Tune the window with KANADE_BH_SAMPLES /
KANADE_BH_INTERVAL_SEC (and KANADE_BH_MON_PORT if the broker's
http_port differs) on the target if needed.
The bundle contains:
- Time series (
connz-*.json,jsz-*.json) — connection count (/connz) and JetStream consumer count (/jsz) per sample. A spike here at the herd moment is the splay signal. - Consumer footprint (
jsz-full.json) —/jsz?consumers=true&streams=true: the full consumer list; confirms the ~7N total and which streams hold them. - Server health/resources (
varz.json,healthz.json) —/varz(memory, CPU, connections, slow consumers) and/healthzstatus. - Log tails (redacted) — backend and nats-server.
What to look for
- Smooth consumer/connection counts across the samples, healthy
/healthz, head-room on mem/cpu in/varz→ the broker absorbed the event; no splay needed, just keep sizing ahead of N. - Spiky connection backlog or consumer-create latency at the event, JetStream API errors, or mem/cpu pegged → the herd is real → add the startup splay (lever 2) and/or grow the broker.
The
#828downgrade-flap regression is checked separately from theOBS_EVENTSagent_updatetimeline (via the backend API), not by this job — that data is already queryable fleet-wide.