Docs/Operations

Operations

Day-2 runbook. Doctor, diagnostics, log rotation, backup, recovery, upgrade.

Run it·16 min·v0.17.8
On this page
  1. Logs — the files you'll actually read
  2. Daemon — one process per machine, every profile inside
  3. Upgrades
  4. Dependencies
  5. Backup + restore
  6. ALP identity rotation
  7. Monitoring + alerting
  8. What changed in this profile?
  9. Operator evidence digest
  10. Disaster recovery checklist
  11. Common failure modes

Runbook for running alpi seriously — at home or inside an organisation. Covers logs, services, upgrades, backup + restore, identity rotation, and monitoring.

If you just installed alpi and want to chat, you don't need this doc yet: Quickstart covers everything. Come back here when things break, or when you need to move a profile, or when it's time to ship a new version.

Logs — the files you'll actually read

Every profile writes to {home}/logs/ with the same format so alpi logs can merge them:

~/.alpi/logs/                     ← default profile
~/.alpi/profiles/<name>/logs/     ← named profile

Rotating text logs cap at 1 MB; .log.1 to .log.3 hold the previous three generations, 4 MB in all. JSONL telemetry feeds are append-only, read with jq (or alpi digest) — compaction.jsonl is unbounded; runs.jsonl is capped and rolling.

The noisy websockets logger is held at WARNING inside the daemon; if service.log ever rotates through its whole window in hours, look for another library logging once per connection before raising the cap.

FileScopeFormatWhat it answersWho writes it
service.logdaemon-wide; ONE file at ~/.alpi/logs/service.log, never duplicated per profilerotated textDid the daemon start? Which services came up for which profile? Did a peer hit an ALP listener? Did a cron job fire?the daemon supervisor + every per-profile service that logs through the root logger
agent.logper profilerotated textWhat has the agent been doing? One line per engine turn on every surface (TUI, schedule, workgroup post, inbound ALP): session id, elapsed, tools called, reply length, cost, user prompt preview. Sub-agents (research, delegate) run inside their caller's turn and add no line of their own. Cross-session grep index.the engine
llm.logper profilerotated textWhat the provider call was doing: request start, first delta, stream end or error. The record that explains a turn killed by the idle watchdog.the LLM transport
approval.logper profilerotated textSecurity audit of every non-safe shell command the LLM tried to run: caution (pending / once / session / always / deny) or dangerous (always denied).the approval system
admin-audit.jsonldaemon-wide; ONE bounded trail at ~/.alpi/logs/admin-audit.jsonlrotated JSONL, 5 MB + 3 backups, 4 KB/rowWhich connection/device attempted a sensitive host-RPC mutation, what safe target it affected, and whether it succeeded, failed, or was denied. It never stores credentials, config values, prompts, replies, or chat messages. Direct CLI/setup mutations do not enter this trail yet.the host RPC dispatcher
compaction.jsonlper profileappend-only JSONLDid auto-compact run this turn? Tokens before/after, summarized-message count, tool-truncation count, manual vs auto, fired (true when the LLM summarized; false when only oversized tool outputs were truncated). Use it as the evidence source before changing compaction/memory constants.the engine (one line whenever compaction or tool truncation ran)
runs/<run_id>.jsonlper profile, under runs/ not logs/append-only JSONL, one file per turnThe operational timeline of one turn: start (pid, model, input), tool starts/states/ends, usage, model state, every assistant_done (the one with final=True is the deliverable; earlier ones are preamble) and the finish outcome. Streaming deltas are not journaled — the chat replay sidecar carries those. Cleanup (Old and excess run journals) offers completed journals older than 30 days plus the oldest ones beyond 200 MiB per profile, skipping anything completed in the last hour; running or unreadable journals are never offered. Separately, a profile that sets retention.runs_days has the daemon delete finished journals older than that once a day (unset = never) — see Configuration.the engine
runs.jsonlper profilecapped rolling JSONLWhat ran and where it stopped: one record per finished run (agent turn, schedule, workgroup, terminal command — there is no duration threshold; the digest is what highlights the slow ones) — outcome, exit code, timeout reason, pid, backend, last tool, secret-redacted output tail, raw cache counts, and request-shape diagnosis. Surfaced by alpi digest.the engine, scheduler, and terminal tool (one line per finished run)
ledger.jsonper profileJSONDaily USD spend ledger; live counters for the daily cap + 30-day per-day history, including raw cache counts, provider-reported cache discount, and the source of recorded cost. Every process running the profile writes it under ledger.lock, so concurrent charges add up; a charge that cannot be written is dropped with a warning in the log, never by failing the turn. Not a log; never cleaned by Subsystem logs.every turn that records cost
prefix_shapes.jsonper profilebounded JSONHash-only request-shape history for up to 20 recently used conversation affinities. It diagnoses model, params, tools, system, or history changes without storing prompt text. Best-effort and safe to delete.the engine before provider calls

Tail one or all:

alpi logs                          # merged tail of every source under the active profile
alpi logs --source service         # always reads ~/.alpi/logs/service.log
                                   #   (root-scope; -p <name> doesn't change the source)
alpi logs --source agent -n 500    # last 500 lines of the active profile's agent.log
alpi -p mira logs --source agent   # mira's agent.log under ~/.alpi/profiles/mira/logs/
                                   #   NOTE: `-p` belongs to the root `alpi` command,
                                   #   not to `logs` — it must come before the subcommand
alpi logs -f                       # follow mode (poll every 1s)

compaction.jsonl is read with jq, not alpi logs:

jq -r '[(.ts | todate), .session_id[0:8], .trigger, .tokens_before, .tokens_after] | @tsv' \
  ~/.alpi/logs/compaction.jsonl     # .ts is a Unix epoch; todate makes it readable

Per-record fields: ts, trigger (auto|manual), session_id, model, ctx_window, fired, tokens_before, tokens_after, summarized_messages, tool_truncated.

The agent.log + approval.log pair answers what the agent did and which shell decisions were made. alpi audit-log answers which connection/device performed an administrative mutation; Desktop exposes the same bounded trail under Connections → Audit log. compaction.jsonl answers "did the context window pressure get tight this week?" and "are my trigger ratios right for this model?".

Prompt-cache evidence

/status reports the current session's cache hit rate, while alpi digest aggregates the selected calendar-day window. Both use raw provider-reported counts: cached / measured input. A completion whose provider reports no cache fields is excluded from both sides of that ratio, so no provider cache data means unknown, not a measured miss.

The digest also shows provider-reported cache discount and a cost_source histogram. Only provider is the endpoint's own number. litellm and table are cache-blind list-price calculations; none means no price was available. Use the provider invoice or dashboard for billing reconciliation.

Low-hit runs carry cache_diag in runs.jsonl. first_contact, model or tool changes, compaction, resume, reset, and edit-and-resend are expected reasons for a cold or rewritten prefix. prefix_shapes.json stores only bounded hashes and may be removed when resetting diagnostics; deleting it does not affect chats or provider caching.

Daemon — one process per machine, every profile inside

alpi runs a single daemon process — the launchd job com.alpi.daemon on macOS, the systemd-user unit alpi-daemon.service on Linux — supervising every profile under ~/.alpi/ — default plus each profiles/<name>/. Each profile gets its own per-service supervised tasks named <profile>/<service> (e.g. doc/schedule, builder/alp); a crash in one profile's service leaves siblings untouched.

What it doesLifecycleInstall / config
Boots isolated tasks on one asyncio loop: scheduler tick, ALP socket (Unix + optional TCP/Noise_XK), workgroups poller, and the default host plane. These are fixed daemon capabilities; feature-level state lives with jobs, peers, workgroups, and connections.`alpi daemon start\stop\restart\status`auto-installed on first alpi setup; manage from alpi setup → Services → Daemon (default profile only)

There's exactly one daemon per machine, one plist / unit. Adding a new profile just creates a directory under ~/.alpi/profiles/; the daemon picks it up on its next restart. Operational verbs that aren't lifecycle survive on their own:

File-descriptor limit. One daemon hosts every profile's services (schedule / alp / workgroups / host), so a machine with many profiles holds a lot of sockets at once. The launchd/systemd unit — and the Docker compose ulimits — raise the FD ceiling to 8192; a low platform default (256 on macOS launchd) is exhausted under load (symptom: OSError: [Errno 24] Too many open files in service.log, operations failing intermittently). The limit lives in the service definition, so alpi daemon install (re-run after upgrading) — or recreating the container — applies it; launchctl limit maxfiles shows the old default but the unit's own SoftResourceLimits overrides it for the daemon process.

alpi schedule run-once          # tick the scheduler once, in-process
alpi schedule fire <job-id>     # ad-hoc run of a specific job

Linux: lingering

systemctl --user services die when you log out unless lingering is enabled. alpi daemon install runs loginctl enable-linger $USER automatically; on restricted environments (WSL without systemd=true in /etc/wsl.conf, minimal containers) loginctl may not exist — the install logs a warning and you'll need to keep the daemon foregrounded under tmux / screen, or fix the linger setup manually.

When stop doesn't stop

On macOS the launchd plist declares KeepAlive=true, so alpi daemon stop is followed by a respawn within seconds. On Linux the systemd unit is Restart=on-failure and a clean stop stays stopped. To stop it permanently on macOS:

alpi setup → Services → Daemon → Uninstall

When restart is really what you want

After uv tool install --reinstall, the long-running daemon still holds the old binary's code. Use:

alpi daemon restart      # macOS: stop and let launchd respawn
                         # Linux: start it again yourself if it stays down

alpi doctor flags "stale binary — alpi daemon restart to reload" when the binary on disk is newer than the running process.

Upgrades

alpi doesn't ship silent migrations. When the on-disk schema changes, the release notes say so and ask you to move files by hand. Today's upgrade rule of thumb:

  1. alpi update — checks PyPI, shows what changed, and runs the uv tool / pipx upgrade on confirmation. From a source checkout instead: git pull + uv tool install --reinstall ..
  2. alpi doctor — the Daemon row flags a stale binary.
  3. alpi daemon restart — one daemon supervises every profile, so a single restart picks up the new code for all of them. (launchctl list | grep com.alpi.daemon on macOS, systemctl --user status alpi-daemon on Linux.)
  4. If the CHANGELOG entry calls for file moves, follow them for every profile.
  5. Re-run alpi doctor — should be clean.

Dependencies

Every provider (OpenAI, Anthropic, Ollama, OpenRouter, Gemini, Groq, Mistral, DeepSeek…) flows through LiteLLM, so alpi pins it to a tight range and re-audits it on a schedule; the procedure lives in RELEASE.md.

Every alpi doctor run verifies the install. The Dependencies rows check the installed LiteLLM against the pin recorded in alpi's own package metadata, and every file of the distribution against the sha256 digests its installer wrote. A version outside the pin or a file that no longer matches its digest is a fail, so a cron'd doctor exits non-zero on a local swap. A missing RECORD warns instead of pretending to verify. This is a tamper check on what is installed, not an advisory scan.

CVEs. alpi audit checks installed Python packages against OSV with exact versions; alpi audit --offline skips the lookup on machines that must not make network calls. Filter findings by surface: alpi uses the LiteLLM SDK, not the Proxy server, so Proxy-only CVEs do not apply and SDK CVEs do.

Backup + restore

alpi backup writes a single passphrase-encrypted file of the whole alpi home (~/.alpi/) — every profile in one shot; alpi restore <file> reverses it. Zero-knowledge — the passphrase derives the key locally and never leaves the machine. Lose the passphrase and the archive is unrecoverable.

alpi backup                                # ./alpi.YYYY-MM-DD.alpi-backup
alpi backup --out ~/vault/alpi.alpi-backup
alpi restore ~/vault/alpi.alpi-backup      # into ~/.alpi/
alpi restore alpi.alpi-backup --force      # overwrite a non-empty home

What's in the backup. The entire ~/.alpi/ tree: default profile (memories, sessions, skills with state/ SQLite + secrets/), every named profile under profiles/<name>/, config.yaml, .env, ALP identity (alp/secrets/alp_key.{pem,pub}), peers, and host state. Excluded recursively at every depth: cache/, logs/, .trash/, sockets (*.sock), PIDs (*.pid) and OS cruft. out/ (files the agent already delivered to you) is excluded only at the home root and at profiles/<name>/, so a skill directory that happens to be named out is still backed up. A restore brings back the profile, not the generated artifacts it already handed over.

Crypto. Scrypt KDF (n=2¹⁷, r=8, p=1) → ChaCha20-Poly1305 over a gzipped tar. Same primitives as age with a passphrase recipient. The header (KDF params, salt, nonce, scope, timestamp, file count) is bound as AAD, so any tamper flips the AEAD tag with the same error a wrong passphrase produces.

Scripting. Both commands accept --passphrase-stdin to read the passphrase from stdin without a prompt. Pair with a password manager or systemd credential — never embed it in the cron line:

pass show alpi/backup | alpi backup --passphrase-stdin --out /backup/alpi.$(date +%F).alpi-backup

After restoring on a new machine, run alpi doctor — it surfaces peers whose counterpart rotated their key since the backup, and any missing optional dependency the restored skills declare.

Credential incident and restore policy

A normal restore from a trusted, encrypted backup intentionally restores the existing host connections and device credentials. It does not require re-pairing by itself. Keep the archive and passphrase in separate controlled locations; raw copies of ~/.alpi, VM snapshots, and support bundles contain active secrets and must receive the same protection.

Use the smallest response that matches the incident:

After any response, run alpi doctor, verify WSS from an external network, and inspect Connections → Audit log or alpi audit-log for rejected use of the old device identities.

Backups are operational snapshots. They are not a review workflow for profile changes. For Git-based profile source, promotion, and secrets boundaries, see Profiles → Versioning.

ALP identity rotation

Rotating the Ed25519 keypair is a deliberate, disruptive act. Every peer who pinned your old pubkey must update their peers.yaml before you can reach them again.

The keypair is per profile, so rotate the one that belongs to the identity you are replacing — ~/.alpi/alp/secrets/ for the default profile, ~/.alpi/profiles/<name>/alp/secrets/ for a named one.

alpi daemon stop                       # or: alpi setup → Services → Daemon → Stop
rm ~/.alpi/alp/secrets/alp_key.{pem,pub}
alpi daemon start                      # generates a fresh pair when the ALP listener boots
alpi peers key                         # print the new pubkey; send OOB to every peer

Every peer on the other end:

alpi peers remove <old-id>
alpi peers add <new-id> <new-pubkey> --allow link.ping --allow link.ask

Treat rotation as planned downtime. Coordinate with your mesh.

Monitoring + alerting

alpi has no built-in metrics endpoint by design (Zero Knowledge principle — no telemetry, no phone-home). For in-house observability, the signals to watch:

For enterprise setups, ship the log dir through a forwarder (rsyslog / Vector / fluentd) to whatever SIEM you already have. The log format is standard Python logging with ISO timestamps; there's no parser to write.

What changed in this profile?

alpi diff [--since 24h] summarises profile-level activity since the cutoff. mtime-driven, side-effect free; safe to run from cron or a remote SSH session.

alpi diff                       # last 24h, default profile
alpi diff --since 7d            # weekly digest
alpi diff --since 2026-04-25    # since an explicit date
alpi -p personal diff --since 1h
alpi diff --since 7d --json     # machine-readable for scripts / dashboards

What it covers: memory edits (which file, when), local sessions (count, turns, tool calls, cost, tokens, agent time), mention threads touched, skill installs, peer-list mutations, fired schedule jobs grouped by job id, and today's budget usage.

The same primitive is exposed in the TUI as /diff [since] (default 24h). One implementation — three surfaces (CLI, TUI, host-plane verb when the desktop catches up).

Use cases:

Operator evidence digest

alpi digest [--since 7d] answers a different question from alpi diff: not "what changed?", but "what parts of this profile need operator attention?" It is a read-only aggregation over state Alpi already writes:

alpi digest                 # last 7 days, human output
alpi digest --since 24h
alpi digest --since 30m
alpi digest --json          # machine-readable report
alpi -p work digest --json

The report covers unavailable tools, skill usage telemetry, memory promotion backlog / pressure, compaction rate over the window, and a Runs section folding the run ledger (runs.jsonl) — totals by kind/outcome, recent failures and timeouts, and the slowest recent runs. It does not run an LLM, write new state, make recommendations, or send telemetry anywhere.

Use it before roadmap or ops decisions: if a proposed improvement has no evidence in the digest, it probably belongs in "listening first" until a real profile starts showing the pain.

Disaster recovery checklist

You've lost a machine. Here's the order of operations to restore.

  1. Reinstall alpi on the replacement machine (uv tool install).
  2. Restore ~/.alpi/ from backup.
  3. Run alpi setup once — it auto-installs the daemon if the plist / unit isn't already in place. (Or manually: alpi daemon install.)
  4. alpi doctor — the Daemon row should read "running".
  5. If your ALP identity is intact (backup included alp/secrets/), your peers still reach you. If you had to regenerate, see ALP identity rotation above.
  6. alpi → test a turn; verify the reply lands.
  7. Tail service.log and agent.log for 24 h to confirm every profile's scheduler and ALP listener are firing normally.

If you had no backup: you've lost the profile. Start from quickstart, re-pair your ALP peers, re-install your skills. The conversation history is gone. This is by design — alpi doesn't phone home, so there's no "recover from the cloud" path.

Common failure modes

"Listener not running" when calling @peer …. The peer's daemon is down, its address is unreachable, or its ALP listener failed. Check alpi daemon status, alpi doctor, and service.log on the peer's machine.

Two daemons running simultaneously. ps aux | grep "alpi (" shows more than one alpi (daemon, …) entry. Usually after a failed reinstall, or after running alpi daemon start foreground while the supervisor was already running. Fix: pkill -f "alpi (daemon" && alpi daemon restart.

Message didn't save to memory. Check the session file: jq '.turns[-1].tools' ~/.alpi/sessions/*.json — if no memory tool call landed, the model decided the signal wasn't worth a write. Inline-learning is LLM-driven; if you want a guaranteed capture, tell alpi explicitly ("remember that…").

The email tool fails. alpi email probe <id> exercises the live login for that account. A failure means revoked or wrong credentials in <profile>/.env, or an expired Gmail OAuth token. alpi doctor flags credential problems explicitly.

Stale binary. After uv tool install --reinstall, the daemon still runs the old code. alpi doctor warns; fix with alpi daemon restart.