Operations
Back up, restore, verify, upgrade, and roll back a hub safely.
A hub keeps everything in one data directory: the hub store (hub.db and its
write-ahead log), a session brain file per session, a knowledge base file per
project, and the artifact blobs. Back up, verify, and upgrade it as one unit.
Back up
There are two ways to take a backup, and both write the same set: every engine
file copied through the engine itself (VACUUM INTO), every artifact blob
verbatim, and a manifest.json holding the schema version, a timestamp, and
each file's size and SHA-256. A backup is the whole set of files the manifest
names. When the embedded tailnet is in use its key state, tailnet/keys.json,
is part of the set, so a restore keeps the node's tailnet identity instead of
re-registering. A project's files that a delete moved aside and could not
remove are left out: they are no longer part of the store, and the next start
removes them. check and restore read either kind the same way.
Offline
agent-hub backup --out /path/to/backup
A running hub holds the engine's exclusive file lock, so the offline backup
refuses while one serves the store, and says to ask the hub with --url, stop
it, or snapshot the volume. With the store free it copies everything while it
holds the lock itself, so nothing changes under the copy.
Online, taken by the serving hub
The hub is the single writer for every file it serves, so it can take the
backup itself without stopping. Online backup is off until backup_dir
(HUB_BACKUP_DIR) names a directory for it, outside the data directory:
HUB_BACKUP_DIR=/srv/agent-hub-backups
The hub refuses to start with a backup_dir inside the data directory, and
checks again with symlinks resolved before each backup. The directory must
already exist: an unmounted backup volume is reported, not filled in on the
root filesystem. Then ask the running hub for a backup:
HUB_ADMIN_TOKEN=<the admin token> agent-hub backup --url http://127.0.0.1:8080
The request is the admin's, so the token is HUB_ADMIN_TOKEN, or
admin_token under [hub] in config.toml, the one the hub serves with. The
hub writes a new directory named for the UTC time to the millisecond, such as
20261007T031500.123Z, under backup_dir, and nowhere else: the request
carries no path. The command prints where the backup landed, and exits 0 on
success, 1 when the hub refused (online backup off, one already running),
69 when the hub is unreachable, and 77 when the token was refused.
HUB_TIMEOUT bounds the wait; a backup the client stops waiting for still runs
to the end on the hub.
The same request without the binary, for a host that only has curl:
curl -fsS -X POST -H "Authorization: Bearer $HUB_ADMIN_TOKEN" \
http://127.0.0.1:8080/api/v1/backups
It answers 201 with the directory's name, its path on the node, the
manifest's schema version, timestamp, file count and bytes, and how long the
backup took. Online backup off is a 404 saying how to turn it on, a backup
already running is a 409, and a missing backup directory is a 503.
What an online backup guarantees:
- Every engine file is a consistent snapshot of itself.
hub.dbis copied through the hub's own handle, inside one read transaction, so writers carry on and the copy holds exactly what was committed when it began. Each session brain and knowledge file is copied the same way while its per-file write lock is held, so no write to that file lands mid-copy. - The store never names a blob the backup lacks.
hub.dbis copied first, and a blob is always written before the row that names it commits, so every blob the snapshot names is on disk by then. Deleting or overwriting those files is held off for the whole backup (an artifact or project delete, a prune's commit, a live artifact write and the update that seals it, which wait and then finish), so none of those blobs can go or change before it is copied, and each one holds the bytes its row describes. The hub reads the copied store back and refuses to finish a backup that names a blob it lacks. - The set is not one instant across files. The brain, knowledge and blob
files are copied after
hub.db, so they can hold writes the store snapshot does not: a page written to a brain moments after the snapshot, a blob whose row committed later, a session brain with no session row yet. The first start on a restored directory removes blobs and brain files no row names, as it does after a crash. - One backup at a time, and nothing is deleted or written live while one runs.
There is no scheduler inside the hub. Run the request from cron or a systemd timer on the node, for example nightly at 03:15:
# crontab -e
15 3 * * * HUB_ADMIN_TOKEN=<the admin token> /usr/local/bin/agent-hub backup --url http://127.0.0.1:8080
or as a timer and a oneshot service:
# /etc/systemd/system/agent-hub-backup.service
[Service]
Type=oneshot
EnvironmentFile=/etc/agent-hub/backup.env
ExecStart=/usr/local/bin/agent-hub backup --url http://127.0.0.1:8080
# /etc/systemd/system/agent-hub-backup.timer
[Timer]
OnCalendar=*-*-* 03:15:00
Persistent=true
[Install]
WantedBy=timers.target
with HUB_ADMIN_TOKEN=... in /etc/agent-hub/backup.env, readable by the
service's account only. The container image is built without the client, so on
a container host run the published binary or curl from the host against the
published port.
Old backups are never deleted by the hub. You prune them, as you do sessions:
remove the oldest directories under backup_dir when the volume needs the
room.
Snapshots
On a NAS or a virtual machine the other alternative is a filesystem snapshot of
the data directory, which is consistent without stopping the hub. Either way
the hub.db-wal is part of the data: a copy of hub.db alone silently omits
recent writes.
Verify
agent-hub check --data-dir /path/to/data
It verifies every file the manifest names (present, right size, right digest),
runs PRAGMA integrity_check on the store and every brain and knowledge file,
and cross-checks artifact rows against the blobs on disk. It exits non-zero on
any problem. A store older than this binary's artifact history is checked
without the cross-check, because there are no version rows to read.
A partially corrupt store is found here, not by /readyz. The readiness probe
runs the cheap legs only (the schema version, a content read, the store
identity, free space); a zeroed page in a table the probe does not touch leaves
it answering 200 while the routes that read that page fail with 503. The
offline check, and doctor, walk the store with PRAGMA integrity_check and
name the damage, so run them on a schedule or after an unclean shutdown.
Restore
agent-hub restore --from /path/to/backup --data-dir /path/to/data
An offline and an online backup restore the same way.
It verifies every checksum before touching anything, refuses while a hub holds
the store, and refuses a non-empty data directory unless you pass --force. It
stages beside the destination and swaps by rename.
The swap is two renames: the live directory is moved to
.agent-hub-replaced-<token> beside it, then the staged tree is moved into
place. A crash between them leaves no data directory at all while the real data
sits in the sibling, so the next start would create a fresh empty one. If a
restore is interrupted that way, move the sibling back before starting the hub:
mv /path/to/.agent-hub-replaced-<token> /path/to/data.
Upgrade and roll back
Migrations run on startup, forward only, each in one transaction that records
its version. Before applying any migration the hub copies the store to
backups/pre-migration-v<from>-<timestamp>.db and keeps the three newest, so an
upgrade that goes wrong has a recent store to restore.
A binary refuses to open a store whose schema is newer than the most it
supports, and says which versions those are, so an older binary will not quietly
run against a newer store. To roll back: stop the hub, replace hub.db with the
pre-migration .db from backups/, and start the older binary. This is a manual
file copy, not agent-hub restore: a pre-migration backup is a bare .db with
no manifest.json, and restore refuses a set it cannot verify. Copy the
sidecars with it, because hub.db alone omits whatever is still in the
write-ahead log:
docker compose -f deploy/compose.yaml down # or: systemctl stop the unit
cp backups/pre-migration-v<from>-<stamp>.db data/hub.db
cp backups/pre-migration-v<from>-<stamp>.db-wal data/hub.db-wal # only if present
cp backups/pre-migration-v<from>-<stamp>.db-shm data/hub.db-shm # only if present
# start the older binary against data/
Remove a hub.db-wal or hub.db-shm the backup did not carry, so the store is
not opened over a stale log. Read the
data model before rolling back across a release
that changed the schema.
The feed ceiling and shutdown
A project's feed is bounded by events_per_project (HUB_EVENTS_PER_PROJECT,
default one million). A write past it is refused with the cap named, so a
runaway agent cannot fill the node; promote durable work to the knowledge base
or prune the feed to make room. The ceiling bounds every agent-surface writer:
signals, questions, answers, artifact publish and update, and the knowledge
base's lifecycle signal. Session lifecycle records and the hub's audit trail
are exempt, so a full feed can never refuse session_start; a knowledge base
write itself still succeeds, because its feed signal is best-effort and is
dropped when the feed is full.
On SIGTERM or SIGINT the hub stops accepting, lets in-flight requests
finish for a bounded window, then checkpoints the store and exits 0. docker stop and a service restart therefore drain rather than cut.
agent-hub health [--url URL] asks /readyz and exits 0 when ready, non-zero
otherwise. It is what the container healthcheck runs, because the runtime image
is distroless and has no shell or curl.
Readiness and disk space
GET /readyz answers 200 only when the store is usable: the schema version is
read, is supported, and matches what the process opened; the data directory and
hub.db still have the identity they had at open, so a removed or replaced store
is caught; a content read confirms it is still a hub store; and the volume has
free space above a safety margin. Any leg failing is a 503, so a supervisor can
take the node out of rotation.
GET /healthz is liveness only: the process is up.
Practical limits
The node is sized for one operator and their agents, not a fleet of tenants, so a handful of things grow with use rather than being capped. Knowing where they are is how you decide when to prune.
- Sessions. No count cap. A session lives until the human prunes it. One session brain file warns past 256 MiB and is refused at 1 GiB, and a listing returns at most 200 sessions.
- Events. A project's feed holds
events_per_project(HUB_EVENTS_PER_PROJECT, default one million). A write past it is refused and names the cap; promote durable work to the knowledge base or prune the feed to make room. The ceiling bounds signals, questions, answers, artifact publish and update, and the knowledge base's lifecycle signal. Session lifecycle records and the hub's audit trail are exempt, so a full feed never refusessession_start; a knowledge base write still succeeds, its feed signal dropping when the feed is full. - Knowledge base. One page is capped at 1 MiB and one path at 512 bytes. A project's knowledge base file warns past 256 MiB and is refused at 1 GiB, so its page count follows that file's size rather than a count of its own. History and last-write scan the write log newest first, which is the brain engine's own call table; the hub wraps it and does not index it, so those reads cost more as the log grows. A history request carries at most 200 rows and reports the real total.
- Artifacts. One blob is capped at 50 MiB. A blob now moves on the Tokio blocking pool, so a large transfer no longer holds an async worker and the hub keeps answering other requests while it lands. Many large transfers at once are felt in memory (roughly 50 MiB each in flight) and disk bandwidth rather than in the workers; past the blocking pool's ceiling they queue.
- Search. A filtered search reads at most 5000 rows before it stops, and returns at most 100 hits a page. A scope that matches more than the fetch cap is reported as truncated rather than scanned whole.
Metrics
GET /metrics reports the hub's own counters in Prometheus text: HTTP requests
by method and status class, failed responses by hub error code, events appended
by kind, MCP tool calls by tool, and sends to the notify target by result
(agenthub_notify_sends_total, delivered or failed). It also carries storage gauges read fresh
for each scrape: free bytes on the data volume, the hub.db size, and the
write-ahead log size, so a time-series system can alert before the readiness
margin trips rather than only when the node is nearly out of room. The route is
admin-gated like the rest of the control surface, so the scraper carries the
admin token in Authorization: Bearer. There is no separate scrape credential.
Set integrity_sample_secs (HUB_INTEGRITY_SAMPLE_SECS, default 0, disabled)
to a number of seconds to have the hub sample the store's own integrity in the
background, on the sweeper's cadence. Each sample runs PRAGMA quick_check on
its own connection, never inside a request, bounded by a wall-clock cap and
skipped while the previous one has not finished, so a large or damaged store
cannot pin the loop. It raises two series: agenthub_integrity_ok, a gauge that
is 1 for a passing sample and 0 for a failing one (absent until the first
sample runs), and agenthub_integrity_failures_total, a counter of samples that
timed out, errored, or reported a problem. A failed sample is what tells you a
zeroed page has appeared between offline check runs; the offline commands still
walk the store with the full PRAGMA integrity_check.
Notify a closed app
An installed app that is fully closed raises nothing on its own. If you run a
notification service on your LAN or tailnet, the hub can nudge it instead: set
notify_url to a URL that accepts a plain-text POST, and when a question, an
approval or an enrolment request starts waiting, the hub POSTs this body to it:
Something is waiting for you in Agent Hub.
That sentence is the whole payload. It names no project, agent, title, id or count, so the target learns only that it is worth opening the app. The setting is off by default and nothing is sent without it (ADR 0026). A retried write that the hub answers from its idempotency key posts nothing new, so it sends nothing.
A self-hosted ntfy on the LAN, with an access token for a topic of its own:
[hub]
notify_url = "https://ntfy.lan/agent-hub?title=Agent+Hub&click=https://hub.lan"
notify_token = "tk_..."
notify_interval_secs = 60
ntfy reads its own message options from the query, so the title and the tap target are set there, by you; the hub sends neither. Subscribe to the topic in the ntfy app on the phone. A target that wants another shape, such as Gotify's JSON message, needs a small bridge in front of it.
Sends are coalesced: at most one per notify_interval_secs, and anything that
arrives inside the quiet period is one trailing send when it ends. A send runs
in the background with a ten-second timeout and is not retried, so a slow or
unreachable target never holds up the agent's write. A failure is logged at
warn with the target's origin and the cause, and counted in
agenthub_notify_sends_total{result="failed"}. The token is sent only in the
Authorization header and is never logged or printed; agent-hub config shows
the URL by its origin alone.
Doctor
agent-hub doctor [--data-dir DIR]
One read-only report: the data directory and its identity, the schema version and the most this binary supports, free space, the write-ahead log size, the id high-water mark, and any artifact blob the store names but the tree is missing. It refuses while a hub holds the store, like the other offline commands, and exits non-zero on any problem.
Configuration
agent-hub config reports the active settings and where each came from, and
agent-hub config --check validates them. An unknown key under [hub] or
[client] is an error: the check exits non-zero and names the key, rather than
ignoring a setting you thought took effect.
Enrolment is on by default. Set enrol = off under [hub] (or HUB_ENROL=off)
to close the unauthenticated enrolment endpoint, and set trust_proxy to the
proxy addresses whose forwarded client header the hub may trust.