Runbook for running alpi seriously — at home or inside an organisation. Covers logs, services, upgrades, backup + restore, identity rotation, and monitoring.
If you just installed alpi and want to chat, you don't need this doc yet: Quickstart covers everything. Come back here when things break, or when you need to move a profile, or when it's time to ship a new version.
Logs — the files you'll actually read
Every profile writes to {home}/logs/ with the same format so alpi logs can merge them:
~/.alpi/logs/ ← default profile
~/.alpi/profiles/<name>/logs/ ← named profile
Rotating text logs cap at 1 MB; .log.1 to .log.3 hold the previous three generations, 4 MB in all. JSONL telemetry feeds are append-only, read with jq (or alpi digest) — compaction.jsonl is unbounded; runs.jsonl is capped and rolling.
The noisy websockets logger is held at WARNING inside the daemon; if service.log ever rotates through its whole window in hours, look for another library logging once per connection before raising the cap.
| File | Scope | Format | What it answers | Who writes it |
|---|---|---|---|---|
service.log | daemon-wide; ONE file at ~/.alpi/logs/service.log, never duplicated per profile | rotated text | Did the daemon start? Which services came up for which profile? Did a peer hit an ALP listener? Did a cron job fire? | the daemon supervisor + every per-profile service that logs through the root logger |
agent.log | per profile | rotated text | What has the agent been doing? One line per engine turn on every surface (TUI, schedule, workgroup post, inbound ALP): session id, elapsed, tools called, reply length, cost, user prompt preview. Sub-agents (research, delegate) run inside their caller's turn and add no line of their own. Cross-session grep index. | the engine |
llm.log | per profile | rotated text | What the provider call was doing: request start, first delta, stream end or error. The record that explains a turn killed by the idle watchdog. | the LLM transport |
approval.log | per profile | rotated text | Security audit of every non-safe shell command the LLM tried to run: caution (pending / once / session / always / deny) or dangerous (always denied). | the approval system |
admin-audit.jsonl | daemon-wide; ONE bounded trail at ~/.alpi/logs/admin-audit.jsonl | rotated JSONL, 5 MB + 3 backups, 4 KB/row | Which connection/device attempted a sensitive host-RPC mutation, what safe target it affected, and whether it succeeded, failed, or was denied. It never stores credentials, config values, prompts, replies, or chat messages. Direct CLI/setup mutations do not enter this trail yet. | the host RPC dispatcher |
compaction.jsonl | per profile | append-only JSONL | Did auto-compact run this turn? Tokens before/after, summarized-message count, tool-truncation count, manual vs auto, fired (true when the LLM summarized; false when only oversized tool outputs were truncated). Use it as the evidence source before changing compaction/memory constants. | the engine (one line whenever compaction or tool truncation ran) |
runs/<run_id>.jsonl | per profile, under runs/ not logs/ | append-only JSONL, one file per turn | The operational timeline of one turn: start (pid, model, input), tool starts/states/ends, usage, model state, every assistant_done (the one with final=True is the deliverable; earlier ones are preamble) and the finish outcome. Streaming deltas are not journaled — the chat replay sidecar carries those. Cleanup (Old and excess run journals) offers completed journals older than 30 days plus the oldest ones beyond 200 MiB per profile, skipping anything completed in the last hour; running or unreadable journals are never offered. Separately, a profile that sets retention.runs_days has the daemon delete finished journals older than that once a day (unset = never) — see Configuration. | the engine |
runs.jsonl | per profile | capped rolling JSONL | What ran and where it stopped: one record per finished run (agent turn, schedule, workgroup, terminal command — there is no duration threshold; the digest is what highlights the slow ones) — outcome, exit code, timeout reason, pid, backend, last tool, secret-redacted output tail, raw cache counts, and request-shape diagnosis. Surfaced by alpi digest. | the engine, scheduler, and terminal tool (one line per finished run) |
ledger.json | per profile | JSON | Daily USD spend ledger; live counters for the daily cap + 30-day per-day history, including raw cache counts, provider-reported cache discount, and the source of recorded cost. Every process running the profile writes it under ledger.lock, so concurrent charges add up; a charge that cannot be written is dropped with a warning in the log, never by failing the turn. Not a log; never cleaned by Subsystem logs. | every turn that records cost |
prefix_shapes.json | per profile | bounded JSON | Hash-only request-shape history for up to 20 recently used conversation affinities. It diagnoses model, params, tools, system, or history changes without storing prompt text. Best-effort and safe to delete. | the engine before provider calls |
Tail one or all:
alpi logs # merged tail of every source under the active profile
alpi logs --source service # always reads ~/.alpi/logs/service.log
# (root-scope; -p <name> doesn't change the source)
alpi logs --source agent -n 500 # last 500 lines of the active profile's agent.log
alpi -p mira logs --source agent # mira's agent.log under ~/.alpi/profiles/mira/logs/
# NOTE: `-p` belongs to the root `alpi` command,
# not to `logs` — it must come before the subcommand
alpi logs -f # follow mode (poll every 1s)
compaction.jsonl is read with jq, not alpi logs:
jq -r '[(.ts | todate), .session_id[0:8], .trigger, .tokens_before, .tokens_after] | @tsv' \
~/.alpi/logs/compaction.jsonl # .ts is a Unix epoch; todate makes it readable
Per-record fields: ts, trigger (auto|manual), session_id, model, ctx_window, fired, tokens_before, tokens_after, summarized_messages, tool_truncated.
The agent.log + approval.log pair answers what the agent did and which shell decisions were made. alpi audit-log answers which connection/device performed an administrative mutation; Desktop exposes the same bounded trail under Connections → Audit log. compaction.jsonl answers "did the context window pressure get tight this week?" and "are my trigger ratios right for this model?".
Prompt-cache evidence
/status reports the current session's cache hit rate, while alpi digest aggregates the selected calendar-day window. Both use raw provider-reported counts: cached / measured input. A completion whose provider reports no cache fields is excluded from both sides of that ratio, so no provider cache data means unknown, not a measured miss.
The digest also shows provider-reported cache discount and a cost_source histogram. Only provider is the endpoint's own number. litellm and table are cache-blind list-price calculations; none means no price was available. Use the provider invoice or dashboard for billing reconciliation.
Low-hit runs carry cache_diag in runs.jsonl. first_contact, model or tool changes, compaction, resume, reset, and edit-and-resend are expected reasons for a cold or rewritten prefix. prefix_shapes.json stores only bounded hashes and may be removed when resetting diagnostics; deleting it does not affect chats or provider caching.
Daemon — one process per machine, every profile inside
alpi runs a single daemon process — the launchd job com.alpi.daemon on macOS, the systemd-user unit alpi-daemon.service on Linux — supervising every profile under ~/.alpi/ — default plus each profiles/<name>/. Each profile gets its own per-service supervised tasks named <profile>/<service> (e.g. doc/schedule, builder/alp); a crash in one profile's service leaves siblings untouched.
| What it does | Lifecycle | Install / config | |||
|---|---|---|---|---|---|
| Boots isolated tasks on one asyncio loop: scheduler tick, ALP socket (Unix + optional TCP/Noise_XK), workgroups poller, and the default host plane. These are fixed daemon capabilities; feature-level state lives with jobs, peers, workgroups, and connections. | `alpi daemon start\ | stop\ | restart\ | status` | auto-installed on first alpi setup; manage from alpi setup → Services → Daemon (default profile only) |
There's exactly one daemon per machine, one plist / unit. Adding a new profile just creates a directory under ~/.alpi/profiles/; the daemon picks it up on its next restart. Operational verbs that aren't lifecycle survive on their own:
File-descriptor limit. One daemon hosts every profile's services (schedule / alp / workgroups / host), so a machine with many profiles holds a lot of sockets at once. The launchd/systemd unit — and the Docker compose ulimits — raise the FD ceiling to 8192; a low platform default (256 on macOS launchd) is exhausted under load (symptom: OSError: [Errno 24] Too many open files in service.log, operations failing intermittently). The limit lives in the service definition, so alpi daemon install (re-run after upgrading) — or recreating the container — applies it; launchctl limit maxfiles shows the old default but the unit's own SoftResourceLimits overrides it for the daemon process.
alpi schedule run-once # tick the scheduler once, in-process
alpi schedule fire <job-id> # ad-hoc run of a specific job
Linux: lingering
systemctl --user services die when you log out unless lingering is enabled. alpi daemon install runs loginctl enable-linger $USER automatically; on restricted environments (WSL without systemd=true in /etc/wsl.conf, minimal containers) loginctl may not exist — the install logs a warning and you'll need to keep the daemon foregrounded under tmux / screen, or fix the linger setup manually.
When stop doesn't stop
On macOS the launchd plist declares KeepAlive=true, so alpi daemon stop is followed by a respawn within seconds. On Linux the systemd unit is Restart=on-failure and a clean stop stays stopped. To stop it permanently on macOS:
alpi setup → Services → Daemon → Uninstall
When restart is really what you want
After uv tool install --reinstall, the long-running daemon still holds the old binary's code. Use:
alpi daemon restart # macOS: stop and let launchd respawn
# Linux: start it again yourself if it stays down
alpi doctor flags "stale binary — alpi daemon restart to reload" when the binary on disk is newer than the running process.
Upgrades
alpi doesn't ship silent migrations. When the on-disk schema changes, the release notes say so and ask you to move files by hand. Today's upgrade rule of thumb:
alpi update— checks PyPI, shows what changed, and runs theuv tool/pipxupgrade on confirmation. From a source checkout instead:git pull+uv tool install --reinstall ..alpi doctor— the Daemon row flags a stale binary.alpi daemon restart— one daemon supervises every profile, so a single restart picks up the new code for all of them. (launchctl list | grep com.alpi.daemonon macOS,systemctl --user status alpi-daemonon Linux.)- If the CHANGELOG entry calls for file moves, follow them for every profile.
- Re-run
alpi doctor— should be clean.
Dependencies
Every provider (OpenAI, Anthropic, Ollama, OpenRouter, Gemini, Groq, Mistral, DeepSeek…) flows through LiteLLM, so alpi pins it to a tight range and re-audits it on a schedule; the procedure lives in RELEASE.md.
Every alpi doctor run verifies the install. The Dependencies rows check the installed LiteLLM against the pin recorded in alpi's own package metadata, and every file of the distribution against the sha256 digests its installer wrote. A version outside the pin or a file that no longer matches its digest is a fail, so a cron'd doctor exits non-zero on a local swap. A missing RECORD warns instead of pretending to verify. This is a tamper check on what is installed, not an advisory scan.
CVEs. alpi audit checks installed Python packages against OSV with exact versions; alpi audit --offline skips the lookup on machines that must not make network calls. Filter findings by surface: alpi uses the LiteLLM SDK, not the Proxy server, so Proxy-only CVEs do not apply and SDK CVEs do.
Backup + restore
alpi backup writes a single passphrase-encrypted file of the whole alpi home (~/.alpi/) — every profile in one shot; alpi restore <file> reverses it. Zero-knowledge — the passphrase derives the key locally and never leaves the machine. Lose the passphrase and the archive is unrecoverable.
alpi backup # ./alpi.YYYY-MM-DD.alpi-backup
alpi backup --out ~/vault/alpi.alpi-backup
alpi restore ~/vault/alpi.alpi-backup # into ~/.alpi/
alpi restore alpi.alpi-backup --force # overwrite a non-empty home
What's in the backup. The entire ~/.alpi/ tree: default profile (memories, sessions, skills with state/ SQLite + secrets/), every named profile under profiles/<name>/, config.yaml, .env, ALP identity (alp/secrets/alp_key.{pem,pub}), peers, and host state. Excluded recursively at every depth: cache/, logs/, .trash/, sockets (*.sock), PIDs (*.pid) and OS cruft. out/ (files the agent already delivered to you) is excluded only at the home root and at profiles/<name>/, so a skill directory that happens to be named out is still backed up. A restore brings back the profile, not the generated artifacts it already handed over.
Crypto. Scrypt KDF (n=2¹⁷, r=8, p=1) → ChaCha20-Poly1305 over a gzipped tar. Same primitives as age with a passphrase recipient. The header (KDF params, salt, nonce, scope, timestamp, file count) is bound as AAD, so any tamper flips the AEAD tag with the same error a wrong passphrase produces.
Scripting. Both commands accept --passphrase-stdin to read the passphrase from stdin without a prompt. Pair with a password manager or systemd credential — never embed it in the cron line:
pass show alpi/backup | alpi backup --passphrase-stdin --out /backup/alpi.$(date +%F).alpi-backup
After restoring on a new machine, run alpi doctor — it surfaces peers whose counterpart rotated their key since the backup, and any missing optional dependency the restored skills declare.
Credential incident and restore policy
A normal restore from a trusted, encrypted backup intentionally restores the existing host connections and device credentials. It does not require re-pairing by itself. Keep the archive and passphrase in separate controlled locations; raw copies of ~/.alpi, VM snapshots, and support bundles contain active secrets and must receive the same protection.
Use the smallest response that matches the incident:
- Lost or stolen client device: from a still-trusted admin device or the local host, revoke that device immediately. Add and verify a replacement on the same connection; its role, profile scope, sessions, and usage remain on the connection while it receives a new credential.
- Suspected device-token disclosure: revoke only the affected device. The daemon closes its active WebSockets; other devices on the connection remain valid.
- Exposed
connections.yaml, Alpi home, server snapshot, or backup: remove public reachability to the affected runtime, rebuild on a trusted host, revoke every restored device credential before reopening WSS, and re-pair each client. Device tokens are stored hashed since 0.14.39, so the file alone cannot be replayed as a credential; revoke and re-pair anyway because the custody of the rest of the material is what is in doubt. Rotate every other secret present in the exposed material: provider/API keys, email credentials or OAuth tokens, skill secrets, and the ALP identity when its private key was included. - Uncertain backup custody: treat it as exposure, not as a normal restore.
- Historical credential copies (
devices.yaml.migrated,devices.yaml.bak-*,connections.yaml.damaged-*): releases before 0.14.39 left them beside the live store and they may hold cleartext device tokens. The daemon never touches them. Runalpi doctorto list them, confirm the liveconnections.yamlcarries onlytoken_hashfields and still authenticates every paired client, then delete the copies by hand. Until they are gone, do not describe the host as free of cleartext tokens. Changing only the backup passphrase cannot revoke credentials already copied from an older archive.
After any response, run alpi doctor, verify WSS from an external network, and inspect Connections → Audit log or alpi audit-log for rejected use of the old device identities.
Backups are operational snapshots. They are not a review workflow for profile changes. For Git-based profile source, promotion, and secrets boundaries, see Profiles → Versioning.
ALP identity rotation
Rotating the Ed25519 keypair is a deliberate, disruptive act. Every peer who pinned your old pubkey must update their peers.yaml before you can reach them again.
The keypair is per profile, so rotate the one that belongs to the identity you are replacing — ~/.alpi/alp/secrets/ for the default profile, ~/.alpi/profiles/<name>/alp/secrets/ for a named one.
alpi daemon stop # or: alpi setup → Services → Daemon → Stop
rm ~/.alpi/alp/secrets/alp_key.{pem,pub}
alpi daemon start # generates a fresh pair when the ALP listener boots
alpi peers key # print the new pubkey; send OOB to every peer
Every peer on the other end:
alpi peers remove <old-id>
alpi peers add <new-id> <new-pubkey> --allow link.ping --allow link.ask
Treat rotation as planned downtime. Coordinate with your mesh.
Monitoring + alerting
alpi has no built-in metrics endpoint by design (Zero Knowledge principle — no telemetry, no phone-home). For in-house observability, the signals to watch:
- Daemon liveness.
alpi doctorin a cron; exits non-zero if any live check fails. Alert on non-zero. The Daemon row covers the supervisor's PID + install backend. - Security posture.
alpi audit --offlinein cron gives a local-only posture scan of every profile: secret file permissions, public binds, disabled hardening, and uncapped budgets. Runalpi auditmanually when network is allowed to include OSV CVEs. Its exit code is non-zero only onfailfindings; warnings are for review. - Log tail error rate.
grep ERROR ~/.alpi/logs/*.log | wc -lover a window — spike = misconfig, broken credentials, LLM API outage. - Cost ceiling. Set
budget.daily_usdin the profile'sconfig.yaml(or leave it unset for unlimited) — see CONFIG.md → Budget. The ledger at the profile's ownlogs/ledger.jsonis the in-process gate; every interactive turn, scheduled job, sub-agent spawn, and inbound ALP call admits against it before running and records its actual spend after. The same file keeps a 30-dayhistorymap of per-day totals (USD, input/output tokens, raw cache counts, provider-reported cache discount, and cost-source counts) — the authoritative spend record, including non-token costs like image generation that session files never see;host.usage.daily(admin-only) serves the last 14 days of it to clients, plus atotal30aggregate over the full 30-day retention. approval.logtriggers. Any line with acaution always-approvedentry means the allowlist grew — a new command pattern is now auto-permitted for this profile. Put a trigger onapproval.logmodifications; review before accepting a new always-allowed pattern into steady state.- Disk.
alpi profile listshows the per-profile footprint. In a managed environment, bound it — a profile quietly growing past 1 GB usually means voice-cache or session-log retention that the user didn't know was on.
For enterprise setups, ship the log dir through a forwarder (rsyslog / Vector / fluentd) to whatever SIEM you already have. The log format is standard Python logging with ISO timestamps; there's no parser to write.
What changed in this profile?
alpi diff [--since 24h] summarises profile-level activity since the cutoff. mtime-driven, side-effect free; safe to run from cron or a remote SSH session.
alpi diff # last 24h, default profile
alpi diff --since 7d # weekly digest
alpi diff --since 2026-04-25 # since an explicit date
alpi -p personal diff --since 1h
alpi diff --since 7d --json # machine-readable for scripts / dashboards
What it covers: memory edits (which file, when), local sessions (count, turns, tool calls, cost, tokens, agent time), mention threads touched, skill installs, peer-list mutations, fired schedule jobs grouped by job id, and today's budget usage.
The same primitive is exposed in the TUI as /diff [since] (default 24h). One implementation — three surfaces (CLI, TUI, host-plane verb when the desktop catches up).
Use cases:
- Came back from holiday —
alpi diff --since 7danswers "what did my service do?". - Pre-backup smoke check —
alpi diff --since <last-backup>beforealpi backupso you know what's about to be archived. - Cron snapshot —
alpi diff --since 24h --jsonpiped into whatever dashboard collects per-profile activity.
Operator evidence digest
alpi digest [--since 7d] answers a different question from alpi diff: not "what changed?", but "what parts of this profile need operator attention?" It is a read-only aggregation over state Alpi already writes:
alpi digest # last 7 days, human output
alpi digest --since 24h
alpi digest --since 30m
alpi digest --json # machine-readable report
alpi -p work digest --json
The report covers unavailable tools, skill usage telemetry, memory promotion backlog / pressure, compaction rate over the window, and a Runs section folding the run ledger (runs.jsonl) — totals by kind/outcome, recent failures and timeouts, and the slowest recent runs. It does not run an LLM, write new state, make recommendations, or send telemetry anywhere.
Use it before roadmap or ops decisions: if a proposed improvement has no evidence in the digest, it probably belongs in "listening first" until a real profile starts showing the pain.
Disaster recovery checklist
You've lost a machine. Here's the order of operations to restore.
- Reinstall alpi on the replacement machine (
uv tool install). - Restore
~/.alpi/from backup. - Run
alpi setuponce — it auto-installs the daemon if the plist / unit isn't already in place. (Or manually:alpi daemon install.) alpi doctor— the Daemon row should read "running".- If your ALP identity is intact (backup included
alp/secrets/), your peers still reach you. If you had to regenerate, see ALP identity rotation above. alpi→ test a turn; verify the reply lands.- Tail
service.logandagent.logfor 24 h to confirm every profile's scheduler and ALP listener are firing normally.
If you had no backup: you've lost the profile. Start from quickstart, re-pair your ALP peers, re-install your skills. The conversation history is gone. This is by design — alpi doesn't phone home, so there's no "recover from the cloud" path.
Common failure modes
"Listener not running" when calling @peer …. The peer's daemon is down, its address is unreachable, or its ALP listener failed. Check alpi daemon status, alpi doctor, and service.log on the peer's machine.
Two daemons running simultaneously. ps aux | grep "alpi (" shows more than one alpi (daemon, …) entry. Usually after a failed reinstall, or after running alpi daemon start foreground while the supervisor was already running. Fix: pkill -f "alpi (daemon" && alpi daemon restart.
Message didn't save to memory. Check the session file: jq '.turns[-1].tools' ~/.alpi/sessions/*.json — if no memory tool call landed, the model decided the signal wasn't worth a write. Inline-learning is LLM-driven; if you want a guaranteed capture, tell alpi explicitly ("remember that…").
The email tool fails. alpi email probe <id> exercises the live login for that account. A failure means revoked or wrong credentials in <profile>/.env, or an expired Gmail OAuth token. alpi doctor flags credential problems explicitly.
Stale binary. After uv tool install --reinstall, the daemon still runs the old code. alpi doctor warns; fix with alpi daemon restart.