Living technical reference for alpi at HEAD. Describes only what currently ships — historical decisions live in commit messages, planned work lives in ROADMAP.md.
Audience: any developer (or LLM) reading this codebase from cold.
What alpi is
alpi is a local-first personal AI agent. It has a Textual TUI in the terminal, a Tauri desktop app and an Expo mobile app that talk to the daemon over the host plane (Unix socket locally, WebSocket remotely), an on-demand email tool (IMAP / Gmail) the agent calls to read and send mail, inline-learning memory, scanner-gated live skills, multi-provider LLM support via LiteLLM, read-only research, write-capable delegation, scheduling, MCP integration, and ALP for private agent-to-agent links.
The architectural constraint is sovereignty: state is local, identities are per-profile, network trust is explicit, and operational surfaces stay small enough to audit. The product is intentionally not a generic agent suite, marketplace, or hosted router.
Principles
alpi is published by Satoshi Ltd. and inherits the company's six operating principles (Privacy by Design, User Sovereignty, Security First, Open Source, Zero Knowledge, Digital Sovereignty). See Why alpi exists in README.md for the mapping between principle and code. The conventions below are the engineering expression of those principles — not separate from them.
- Focused. Every feature earns its keep. No over-engineering. Maps to Satoshi's "constraint breeds coherence" heuristic.
- Solid base. Core loop, memory, tools, paths, scanner before surface features.
- User in control. No destructive action without explicit OK. No silent migrations. Expression of User Sovereignty.
- Python stack. No Go rewrite (loses LiteLLM, tests, no upside).
- No legacy code. When a schema or layout changes, it's a clean break — no compat shims, no auto-migration. Anything from yesterday's iteration is cleaned by hand, not by
ensure_home. - Closed protocol, own transport. ALP is not A2A / MCP-over-network / gRPC. Every verb we don't ship is an attack surface we don't own. Expression of Privacy by Design + Security First.
Code conventions
Contributor rules — no human-facing comments inside alpi/, English only, every comment must carry a "why" — live in AGENTS.md at the repo root.
CLI surface
Stable verbs shared across groups so a user doesn't relearn per feature.
alpi launch the TUI
alpi -c / --continue resume the last session in the TUI
alpi -p <name> profile flag, combinable with any command
alpi chat alias for `alpi`
alpi chat --once "<text>" one-shot turn to stdout (pipe-friendly)
alpi chat --once ... -c | --session <id> continue the last / a specific session (one-shot)
alpi chat --once ... --connection-id <id> local-only delegated turn visible live to that connection
alpi chat --once ... --emit-events INTERNAL — scheduler subprocess contract
alpi chat --once ... --no-save INTERNAL — do not write a session file
alpi setup interactive menu: model / email / voice / MCPs /
peers / workgroups / sandbox / service /
health check / cleanup /
delete profile (non-default only)
alpi doctor live health check (IMAP login,
Gmail token refresh, MCP handshake, service PID);
exits 1 on any failure, 0 otherwise
alpi audit whole-install security posture scan
alpi audit-log bounded administrative activity by device
alpi logs tail every subsystem log merged by timestamp
--source {service|schedule|agent|approval} restrict to one subsystem
-n N last N lines (default 100)
-f follow mode (poll every 1s)
alpi profile list list profiles, mark the active one with its object glyph
alpi profile show [name] draw the profile's object as half-block art with its model and path
alpi profile create <name> bootstrap a new profile tree
alpi profile remove <name> delete after safety checks + confirm
alpi daemon install|uninstall register / unregister the service unit
alpi daemon start|stop|restart|status lifecycle of the single per-machine daemon
alpi schedule run-once|fire <id> manual cron tick / ad-hoc job fire
alpi peers list list pinned ALP peers for this profile
alpi peers key print this profile's ALP public key
alpi peers add <id> <pubkey> pin a peer (prefer the wizard for capability selection)
alpi peers remove <id> unpin a peer
alpi peers tools <id> [--allow …|--clear] show or set the only tools an inbound turn from that peer may run
alpi peers ping <id> live probe via link.ping
alpi workgroup list list workgroups (hub-of + member-of)
alpi workgroup show <wg_id> detail + decrypted transcript
alpi workgroup create <name> --member <id> hub-side create, granting the invited peers
alpi workgroup join <hub_peer_id> <wg_id> subscribe to a peer-hosted workgroup
alpi workgroup post <wg_id> <text> encrypt + post, declaring the turn's cost
alpi workgroup pull <wg_id> fetch new posts and decrypt; cursor advances
alpi workgroup pause|resume|leave <wg_id> membership ops
alpi workgroup kick <wg_id> <member-id|pubkey> hub-only; rotates the group key
Shape rules: containers (profile, peers, workgroups) get list/create/remove (or add/remove). The daemon gets start/stop/restart/status/install/uninstall under alpi daemon; the same lifecycle is also reachable from alpi setup → Services → Daemon (default profile only) so users have one canonical place. The first alpi setup auto-installs the daemon — no opt-in step. Scheduler, ALP, workgroups, and host are fixed daemon capabilities; their useful controls live with jobs, peers/workgroups, connections, and network settings. Interactive wizards live exclusively under alpi setup; never add a per-feature wizard command.
Command ordering in --help is frequency-first, not alphabetical: chat → setup → doctor → logs → profile → peers → workgroup → schedule → daemon. See _OrderedGroup in cli.py.
alpi/ui.py is the shared interactive layer. Raw questionary.* is forbidden outside it. Helpers: banner, menu, text, password, confirm, row, ok/fail/warn/dim/saved/cancelled. The close item is added automatically with value None (callers treat None as "out").
Menu close wording: top-level (alpi setup) → Exit. Sub-menus (Email:, MCP servers:, Manage saved keys) → ← Back. Wizard aborted mid-flow → cancelled. Mixing Exit/Back/Cancel in one context is a bug.
File layout
alpi/
├── __init__.py __version__
├── cli.py entry point, --continue, --profile resolution
├── engine.py turn runner, interrupt flag, tool loop
├── llm.py litellm stream() / complete() wrappers
├── session.py Turn / ToolLog dataclasses, save/load
├── prefix_diag.py hashed conversation affinity + request-shape diagnostics
├── memory.py MemoryStore (3 files, two-tier dedup, .bak)
├── home.py profile path resolution
├── config.py YAML load/save, defaults, deep merge
├── ui.py shared wizard/menu primitives
├── service.py the one daemon: per-profile tasks, launchd / systemd unit
├── ledger.py daily spend ledger + the profile cap gate
├── outputs.py persistent inbox JSONL store (notify + schedule failures)
├── status.py canonical /status rows (TUI + apps share this)
├── prompts/
│ ├── default_agent.md
│ └── system_prompt.md
├── providers/ metadata for the model picker
│ └── {anthropic,openai,google,groq,openrouter,ollama}.py + model catalogues
├── tools/
│ ├── base.py Tool ABC + ToolResult
│ ├── _state.py per-turn emit / interrupt / usage, isolated per thread
│ ├── _paths.py resolve_path + sensitive-path denylist
│ ├── _guards.py terminal denylist, SSRF, prompt-injection scan
│ ├── _budget.py per-result char cap for LLM context (100K default, per-tool override)
│ ├── _osv.py OSV malware query for PyPI/npm names before skill/MCP install
│ ├── _sandbox.py OS-level sandbox wrapper (opt-in)
│ ├── skill.py the skill tool: mutations, scanner, quota
│ ├── search.py content + filename search (rg + stdlib fallback)
│ ├── research.py read-only sub-agent (depth: fast/normal/deep)
│ ├── terminal.py run/background/status/output/kill
│ ├── workflow.py bounded tool DAGs routed through ToolExecutor
│ ├── notify.py native push to the owner's apps
│ └── … (read_file, write_file, edit_file, delete_file, todo, web_*, schedule,
│ memory, session_search, email, config)
├── tui/ Textual app, widgets, screens, theme
├── scheduler/ cron + once jobs, hosted by the alpi daemon
├── mail/ multi-account email: accounts, IMAP+SMTP, Gmail + OAuth
├── mcp/ MCP client (stdio JSON-RPC) + registry
├── alp/ Alpi Link Protocol (spec: docs/ALP.md)
│ ├── keys.py Ed25519 identity at {home}/alp/secrets/alp_key.{pem,pub}
│ ├── envelope.py build/sign/verify JSON-RPC envelope + replay cache
│ ├── peers.py {home}/alp/peers.yaml load/save + capability check
│ ├── server.py Unix-socket listener, fail-closed dispatch
│ ├── client.py one-shot call with typed errors (TargetOffline, RemoteError)
│ ├── handlers.py link.ask / link.cancel — engine integration
│ ├── mention.py @peer parser + executor (shared by TUI + host chat)
│ ├── pending.py pending invites store (unpinned-sender capture)
│ └── setup.py `alpi setup → Peers` wizard
├── host/ control plane for desktop / mobile clients (default profile only)
│ ├── server.py Unix-socket JSON-RPC server (no envelope, no Noise — fs perms = trust)
│ ├── handlers.py read verbs (host.workgroup.transcript, host.sessions.*)
│ ├── chat.py host.chat.send/delegate (streaming) + host.chat.cancel
│ ├── runs.py host.runs.list + host.run.{read,cancel}
│ ├── config.py config mutation verbs (providers, peers, mcp, email, …)
│ ├── connections.py connection identities, device credentials, migration
│ ├── connection_context.py request-scoped connection/device attribution
│ ├── admin_audit.py bounded audit trail for administrative mutations
│ ├── attachments_rpc.py stage uploads in, serve produced files out (scoped)
│ ├── network_rpc.py bind status and the ordered WS/WSS pairing routes
│ ├── probes.py host.email.probe, host.peers.ping, host.model.ctx_window
│ ├── schedule.py host.schedule.{list,remove,set_paused,fire}
│ ├── outputs.py host.outputs.{list,read,mark_read,mark_unread,mark_all_read,delete}
│ ├── daemon.py host.daemon.{restart,update}
│ ├── device_state.py device-facing profile state for the apps
│ ├── events.py host.events.subscribe + thread-safe emit() for daemon-pushed updates
│ ├── workgroup.py transcript decryption (hub + member shapes)
│ └── sessions.py plaintext session list / read
└── knowledge/ `alpi_knowledge` answer packs (hand-maintained)
Execution spine
Every engine turn creates one immutable RunContext, one ToolExecutor, and one ExecutionWorld. Context variables bind those objects across nested tool calls without adding parameters to every tool. The executor is the only registry dispatch seam: direct model calls and workflow steps therefore use the same denylist, member restrictions, availability checks, execution world, and durable journal.
Tools are exclusive by default. Only classes that explicitly declare parallel_safe can overlap, and the engine parallelizes a model batch only when every call is safe. A mixed batch remains serial. Results and emitted states are replayed in original call order, preserving provider transcript determinism.
A call runs only with the arguments the model actually sent. The engine, delegate, research, the workgroup wrap-up handoff and the memory reviewer decode them through alpi/tools/_args.py: empty or null is {}, raw control characters inside strings are accepted, and one JSON-encoded layer around an object is unwrapped. Anything else never reaches the tool and answers one of:
arguments for <tool> are not valid JSON: <reason> (char <pos> of <len>); the call did not run.arguments for <tool> are not valid JSON: <reason>; the call did not run.— a key given two different values, nesting deeper than 100 levels, a string holding an unpaired surrogate or a number out of range (NaN,Infinity, beyond ±1.8e308): values a tool, the logs or a strict JSON reader would mishandle. A key repeated with the same value is kept once.arguments for <tool> are not a JSON object (got array); the call did not run.arguments for <tool> are missing (required: …); the call did not run.— a tool with required parameters got empty ornullarguments; an explicit{}still reaches the tool.
Only the last call of a reply cut at the output-token limit adds that the limit cut it off. A batch holding a refused call runs serially. The assistant message sent back to the provider carries the arguments as dispatched: the raw string when it is already a strict JSON object without repeated keys equal to them, canonical ASCII JSON otherwise, {} for a refused call, so a provider that parses history never sees a broken payload.
Right after decoding, a top-level argument sent as a JSON string is decoded when its schema admits objects or arrays and not strings (type, a type list or anyOf/oneOf branches), so a stringified knowledge object or workflow steps list reaches the tool, the events and the journal as the structure the schema declares. A string whose JSON would be refused above stays a string. workflow applies the same rule to each step's own arguments before ${step.output} references expand; an expanded value stays a string. A parameter that admits strings, or one without a readable type, passes through untouched.
Each turn writes runs/<run_id>.jsonl with bounded, redacted events (an unpaired surrogate in any persisted text, journal or saved session, becomes U+FFFD): the start record (pid, model, input), tool starts / states / ends, model_state when it changes, usage, every assistant_done (only the one closing the turn carries final=True), errors and the finish outcome. Streaming deltas are never journaled — the reconnect replay is the sessions sidecar — so alpi runs show and host.run.read return an operational timeline, not the stream. Terminal command text is omitted everywhere, including workflow steps; steps or step arguments that are not structured are dropped from the record. Local operators use alpi runs list|show|cancel or /runs; paired clients use host.runs.list, host.run.read, host.run.cancel, connection-scoped like sessions.
Cleanup offers completed journals older than 30 days plus the oldest beyond 200 MiB per profile, skipping anything completed within the last hour; nothing is deleted on its own. A run sweep runs at profile start and every 30 s from the daemon's maintenance loop, off the scheduler's thread pool and wrapped so a failure in it never ends the loop. One scan of runs/ feeds two rules in order. reconcile_stale closes journals whose pid is gone as interrupted (reason dead), so a dead child is reported within about a minute. reconcile_silent judges a scheduled run that has written nothing for longer than its timeout plus SILENCE_GRACE_S: the timeout_s the scheduler recorded in run.started (from ALPI_RUN_TIMEOUT_S), else the job's current timeout, else MAX_RUN_TIMEOUT_SECONDS for a deleted job or an invalid value; silence is wall-clock, but the sweep must also have watched it hold for the grace on the monotonic clock before acting, so a wedged child is reported roughly ten minutes past its timeout and a clock step alone never fires it. A live pid is killed only when it is provably this run's process — run.started records pid_start from /proc/<pid>/stat, the sweep requires it to match, and never targets its own pid or its parent's. Anything alive it cannot vouch for (a recycled pid, a pre-0.14.50 journal, a host without /proc) is reported once and left untouched, journal included: a run.finished written under a live writer would be followed by that writer's own records and the run would read as running again. Runs without a job are never judged by silence. The scheduler closes the journal of a child it ends itself (it hands the child ALPI_RUN_ID), so its own timeout yields one alert rather than one from the scheduler and one from the sweep. Each reported row files an error output and raises schedule.failed, the path a live failure already uses.
ExecutionWorld keeps filesystem resolution and terminal shell execution under one run-scoped abstraction. local preserves the previous behavior. docker wraps processes in an ephemeral container while bind-mounting the same absolute workspace/profile paths, so file tools and subprocesses observe one namespace. Foreground containers are force-removed on timeout; background terminal jobs are refused so a detached container cannot outlive its run. Dedicated workers such as skill scripts and speech transcription remain host-side; this backend is not a whole-agent filesystem sandbox.
The workflow tool executes a bounded dependency graph of registered tools. References such as ${step.output} feed prior results into later arguments; independent safe steps can overlap. Recursion is refused, failures stop the graph unless explicitly marked continue_on_error, and every nested call re-enters ToolExecutor rather than bypassing policy. Nested parallelism uses the same tools.max_parallel_tool_calls limit as direct model batches.
Runtime state (skills, sessions, memories, logs, ALP peers, keys) does not ship with the package — it's generated per profile under ~/.alpi/. The alpi/knowledge/references/ directory holds the answer packs the alpi_knowledge tool serves; there is no bundled skill namespace. See Profile home layout immediately below. The skill tool (alpi/tools/skill.py) manages user-created skills that live at {home}/skills/<category>/<name>/.
Profile home layout (~/.alpi/ or ~/.alpi/profiles/<name>/)
~/.alpi/ default profile root
├── .env API keys, IMAP/SMTP credentials, allowlists
├── config.yaml model + tools + tui + mcp
├── memories/ USER.md, MEMORY.md, AGENT.md (+ .bak)
├── skills/<category>/<name>/ SKILL.md + scripts/ + references/ +
│ assets/ + secrets/ (0700) + state/ +
│ .gitignore
├── recipes/<id>.yaml saved workgroup recipes owned by this hub profile
├── sessions/<id>.json compact turn-based session log (TUI / desktop / `--once`)
├── knowledge.sqlite sqlite-vec derived indexes for knowledge,
│ session recall, and workgroup recall
├── mentions/<sender>@<conversation>.json @-mention threads per sender and originating conversation (cap 20 turns), receiving side; <sender>.json for callers that send no conversation
├── run/ background process registry, schedule pids
├── alp/ ALP state — keypair, peer list, socket, pid
│ ├── peers.yaml pinned peers (pubkey + allow + optional address)
│ ├── alp.sock Unix-domain socket, 0600, only while listener runs
│ ├── alp.pid listener pid
│ └── secrets/alp_key.{pem,pub} Ed25519 identity (private 0600, public 0644)
├── host/ control-plane state (default profile only)
│ └── host.sock Unix socket the local desktop connects to (mobile uses the WebSocket)
├── outputs/ persistent inbox for proactive agent messages + schedule failures
│ └── outputs.jsonl JSONL store (≤500 rows, atomic compaction)
└── logs/ service.log (daemon-wide; lives only at the root, NOT
duplicated per profile), agent.log + approval.log
(per profile — only the default profile's pair is at
this level), ledger.json, compaction.jsonl, runs.jsonl
~/.alpi/profiles/<name>/ same layout MINUS service.log; agent.log + approval.log
are emitted under each profile's own logs/
Core systems
Engine loop (alpi/engine.py)
Per turn: append user message → loop {LLM stream → emit deltas → exec tool calls → append tool results} until the LLM stops emitting tool calls OR max_steps_per_turn is hit. The configured value (default 100 model iterations) is literal for every provider and budget. Hitting it does not drop gathered work: normal chats get one tools-off best-effort wrap-up; detached workgroup turns get one workgroup_post-only handoff. interrupt_requested is polled at three checkpoints (between iterations, mid-stream, between tool calls). A turn lock serializes concurrent runs so a delayed research tool from the previous turn can't bleed into the next.
Events emitted to the UI sink: user, reasoning_delta, assistant_delta, assistant_done, tool_start, tool_state, tool_end, usage, error, done, interrupted. The TUI consumes them; the scheduler subprocess consumes a subset via JSON-lines.
A model-provider failure reaches every surface as one plain sentence, never as the provider's exception text. alpi/llm_errors.py classifies the exception chain (litellm class, HTTP status, text) into context_length, insufficient_credit, auth, forbidden, model_unavailable, content_policy, rate_limited, timeout, provider_unavailable, bad_request or unknown. The error event carries text (the sentence), code (the class) and detail (the redacted, 500-character technical text, also logged and kept in the run journal). host.chat.send frames carry text and code; detail reaches the TUI, --once ([error] text (detail), or the detail key of the JSON event) and the scheduler, whose failed-job message reads agent error: text (detail). ALP peer replies carry only the sentence. Chat-level codes such as busy share the code field.
The apps put the daemon's own protocol errors in plain words through one shared map, common/plainError.mjs (fixtures in common/plainError.fixtures.mjs, exercised by both clients): too-many-connections, too-many-requests, forbidden, method-not-found and the close reasons Authorization changed, Device authorization revoked, WebSocket capacity reached, plus the transport failures, become a sentence; anything else is shown unchanged. Both clients retry a call or a stream refused with -32029 too-many-connections once after 1.5 s before showing it.
Host chat forwards usage with context_tokens; it is non-zero only for the main completion whose input becomes Session.last_ctx_tokens. Side-model usage still carries accounting fields but cannot move a client's conversation meter.
Cross-turn resume. A chat is not a long-lived object: each turn spins up a fresh Engine and rehydrates the session from disk (_hydrate_from_path in cli.py, shared by TUI --continue and the host chat; the desktop "edit message" rewrite path mirrors it in host/chat.py). The model context is rebuilt from the prior replayable turns — those that ended in a final reply or produced a file; a turn aborted before its reply (no assistant text, no output files) is dropped, so a resumed session never re-answers a dangling request. Each replayed turn contributes its user text (plus an input-attachment marker [attached: name (mime)]) and assistant text (plus a produced-file marker [produced this turn — reuse the absolute path…: name → /abs/path]). Tool calls and tool results are deliberately not replayed — they would blow the context budget — so an agent does not remember what it searched, read, or analyzed last turn, only its final reply and the absolute paths of the files it produced. A multi-turn edit ("now relight it at sunset") reuses the produced path surfaced by the marker, not a remembered tool output; an agent that needs an earlier tool's result across turns must re-run the tool or rely on a produced file.
The system prompt for each turn is assembled in a fixed order (PART_ORDER in alpi/prompt_cache.py): AGENT.md (agent profile — voice, style, identity) → base prompt → environment block (workspace, profile home, path rule) → system time → platform hint (per-surface guidance when ALPI_PLATFORM is set: cron; empty for TUI and the apps) → turn guidance → the self-knowledge rule pointing the model at alpi_knowledge (dropped when that tool is denied) → skills index → USER.md → MEMORY.md. Providers that need an explicit marker (Anthropic, including through OpenRouter) get LiteLLM's cache_control_injection_points on messages[0] from cache_kwargs_for_model; an openai-prefixed model at OpenAI's own endpoint never does (the host is judged as LiteLLM judges it, from api_base, litellm.api_base or OPENAI_BASE_URL), because LiteLLM would turn the marker into OpenAI's explicit mode, which switches off its automatic prefix cache (and the /v1/responses bridge drops the breakpoint anyway), so an OpenAI model reads its cache with no marker at all; a gateway behind the same prefix keeps it.
The scheduler (alpi/scheduler/run.py) sets ALPI_PLATFORM=cron so scheduled jobs run knowing no user is present and they cannot ask for clarification. That overwrite would otherwise erase the deployment runtime, so it also carries ALPI_DEPLOY_RUNTIME, which alpi/runtime.py prefers: ALPI_PLATFORM is the turn origin, platform_id() is the runtime, and a job inside a container still reads as Docker. Each fire runs as a subprocess capped at job_run_timeout(job) seconds — job.timeout if set, else DEFAULT_RUN_TIMEOUT_SECONDS (900). parse_run_timeout is the one duration contract for schedule(add|update), execution, schedule list / host.schedule.list (run_timeout, or null plus timeout_error) and the silence watchdog: a whole number of seconds in [30, 86400]; booleans, fractions, non-finite and out-of-range values are refused, and a bad stored value fails the fire with invalid stored timeout instead of being clamped. Fires of one profile run serially, so a long timeout also delays that profile's other due jobs. The cap is a stuck-process backstop for unattended runs, not the cost guard (budget.daily_usd is) and not a hint that jobs must be short; heavy jobs (deep research, multi-step publishing) opt into a longer budget via schedule(add|update, timeout=…). The scheduler passes the child a soft budget via ALPI_TURN_BUDGET_S (the cap minus a ~10% reserve, floor 60s); when the engine crosses it, normal jobs get one tools-off best-effort reply and detached workgroup turns get one workgroup_post-only handoff. The hard subprocess timeout remains the last-resort kill if finalization itself stalls.
Cron jobs with no_agent: true skip the LLM entirely. The prompt is shlex-tokenized and exec'd directly (shell=False); ${ALPI_HOME} expands to the profile home and the profile's .env overrides inherited env keys so skills find their declared requires_env. A form-based allowlist enforces that the command is python[3] [flags] <script> or <script> invoked directly, where <script> resolves to <home>/skills/<category>/<name>/scripts/…; non-python executables and -c/-m inline-code flags are rejected at both schedule(add) time and inside the scheduler before exec. Use this for deterministic skills (sync, file processors) — saves both tokens and the agent boot latency per fire.
LLM transport (alpi/llm.py)
Thin wrapper over litellm.completion. stream() is a generator yielding {text_delta, reasoning_delta, tool_calls_delta, finish_reason} per chunk plus a final {final, tool_calls, input_tokens, output_tokens, cost_usd, finish_reason} whose finish_reason is the last one the provider reported (litellm reports stop when a stream closes early, so only length proves a cut); the stream end breadcrumb logs it. complete() is the non-streaming variant (research, delegate and the background passes); its Completion.finish_reason carries the same signal. _silence_litellm() runs at import time to mute LiteLLM's startup banner via FD-level redirect (Textual is sensitive to stdout pollution).
Memory (alpi/memory.py)
Three files: USER.md (facts about the user), MEMORY.md (env quirks, commands, incidents), AGENT.md (the agent's own profile — tone, style, identity, language). § entry delimiter, char limits USER_CHAR_LIMIT = 3000 / MEMORY_CHAR_LIMIT = 5000 (see alpi/memory.py; AGENT.md is free-form prose with no cap). Accent+case+punctuation-insensitive dedup, plus token-Jaccard dedup at 70% max-containment to catch paraphrases. .bak snapshot before every mutating write. Approach C: every mutating call returns the full current state of the target file so the agent sees its own write in the same turn.
v2 quality metadata. Each entry carries a trailing <!-- alpi-meta conf=... captured=... reinforced=... --> comment that is stripped before the entry reaches the system prompt. conf is low / normal / high (default normal). Near-duplicate writes reinforce the existing entry (bump reinforced, upgrade low → normal at ≥ 2) instead of appending a paraphrase. Low-confidence entries with zero reinforcements expire after LOW_CONFIDENCE_MAX_AGE_DAYS = 30 (constant in alpi/memory.py; keep it fixed unless operational traces justify tuning). The memory tool's safety scanner reuses the skill scanner patterns and adds invisible/bidi unicode detection (U+200B–200F, 202A–202E, 2060, 2066–2069, FEFF) to block Trojan-Source vectors; _operational_warning surfaces non-blocking warnings when a write looks like session state (chat_id, session_id, ISO timestamps).
Batch writes. memory(action="add", entries=[...]) accepts a list of entries for the same target in a single call. Each entry runs through cross-file and same-target dup checks independently; entries that collide are skipped with a per-line note, the rest land in one write. Replaces the pathological pattern of one add call per fact (16 calls in a single turn observed in real sessions).
Post-turn reviewer. When memory.review_interval > 0 (default 0 = off), alpi/review.py spawns a daemon thread after each turn that snapshots the user/assistant text and asks the LLM whether anything durable should be added. The reviewer is constrained to memory(action="add", ...) — never replace/remove — to prevent it from deleting unrelated entries on a bad pass.
Promotion queue (alpi/promotion.py). Auto-compaction never writes to USER.md / MEMORY.md / AGENT.md directly. After every fired compaction the engine runs a second short LLM call against the summary (system prompt CANDIDATE_PROMPT) and pushes any durable facts as candidates into <home>/memories/promotion_queue.jsonl. On enqueue, each candidate is annotated with the same preview warnings the memory tool computes at write time — operational-state heuristic, cross-file duplicate check, safety scan. The queue is bounded (MAX_PENDING = 200 per profile) and pending entries expire after MAX_AGE_DAYS = 30. Per-record fields in the JSONL: id (8-char hex), created_at (unix ts), source (compaction | reviewer | manual), session_id, model, target (USER.md | MEMORY.md | AGENT.md), text, confidence (low | normal | high), warnings (list of strings).
Two memory tool actions surface the queue, both safe for the agent to call: promotion_list (read-only) and promotion_discard(id) (drops a candidate without writing). There is no agent-callable apply. The only path that writes to durable memory from the queue is the CLI alpi memory promote — interactive review with [a]pply / [d]iscard / [s]kip / [q]uit per candidate, plus --apply-all / --discard-all for unattended sweeps. This keeps the human-in-the-loop gate genuine: the agent cannot promote facts on its own regardless of how the prompt is framed. If the underlying memory add fails (safety scan, duplicate), the candidate stays in the queue so the operator can fix and retry.
Path resolution (alpi/tools/_paths.py)
Single entry point resolve_path(path):
expanduser().- Relative paths root at the active workspace (
cfg.workspaceorcwdfallback). - Resolve symlinks.
- Reject if the resulting path matches any sensitive-path entry (denylist below) —
ValueError.
Denylist: /etc/, /boot/, /sys/, /proc/, /usr/lib/systemd/, /System/, /private/etc/, the docker sockets, ~/.ssh/id_*, ~/.ssh/authorized_keys, *_key, *_ed25519, *.pem/.p12/.pfx, ~/.aws/{credentials,config}, ~/.gnupg/, ~/.netrc, ~/.npmrc, ~/.pypirc, ~/.pgpass, ~/.config/{gh,gcloud}/, shell rc/login files (.bashrc/.zshrc/.zprofile/…), ~/Library/Launch{Agents,Daemons}/, profile .env/config.yaml, and skill secrets/ dirs. Both pre-resolve and post-resolve forms are checked (macOS /var → /private/var symlink case).
suggest_similar_paths(target) lists the parent directory and fuzzy-matches siblings by basename substring/prefix. Used by read_file, edit_file, and search to turn dead-end errors into actionable suggestions.
alpi/tools/_lint.py::lint_content(path, content) runs a parser-based syntax check before every write_file / edit_file lands on disk. Parsers by suffix: .py → ast.parse (stdlib), .json → json.loads (stdlib), .yaml/.yml → yaml.safe_load_all (PyYAML, already a dep; every --- document is parsed, so Spring and Kubernetes multi-document files pass and a broken later document is reported at its line in the file), .toml → tomllib.loads (stdlib on every supported Python version). Other suffixes pass through. Failures return a one-line error with the source line/col and the write is refused — the original file (if any) is untouched. Catches the class of bug where a malformed jobs.json, config.yaml, or skill script silently breaks a downstream consumer.
alpi/secrets_io.py::safe_write_secret(path, content, mode=0o600) is the canonical write path for any credential file. It uses tempfile.mkstemp (O_EXCL + 0o600 at creation, random unique name in the target dir), then os.replace onto the target — no TOCTOU window where the file exists at umask perms, and a stale <target>.tmp lingering at looser perms cannot compromise the write because the helper picks a fresh random name. Used by model_selector._atomic_write_env (.env writes), mail/gmail_auth._save (gmail token), alp/pending.save (pending-peers yaml), and alp/keys.create (ALP private key).
Tool registry (alpi/tools/__init__.py)
register(cls) adds a Tool subclass to the dict, schemas() emits the OpenAI function-calling shape, execute(name, args) runs by name with full error capture. The registry is assembled from the sibling tool modules in alpi/tools/__init__.py, including the Playwright-backed browser tool. knowledge registers first so durable user/workspace recall has one canonical surface.
A tool's check() is what keeps schemas() honest: an unavailable tool is never offered to the model, so the probe has to test the thing that actually fails at call time, not a proxy for it. browser is the cautionary case — probing only import playwright advertised the tool on every slim Linux image, where playwright downloads the browser on demand and then cannot launch it because the distro never installed Chromium's load-time libraries. It now dlopens one soname per Debian package family (_CHROMIUM_SONAMES, Linux-only, ~2 ms when they are absent) and reports the missing set plus CHROMIUM_DEPS_COMMAND — a uvx --from playwright … invocation, because uv tool install links only alpi-agent's own entry points and a bare playwright is not on PATH.
The image cannot be held to that list by a unit test: pip install . ignores uv.lock, so the playwright inside the image outruns the one the suite imports (observed: 1.62 vs 1.58). So docker/Dockerfile derives the packages from its own playwright (playwright install-deps chromium-headless-shell) instead of carrying a hand-written copy, and publish-docker.yml launches the real headless shell in the built image before anything is pushed. ensure_chromium() installs with --only-shell because chromium.launch(headless=True) runs chrome-headless-shell; the full Chromium build was never used, and _wanted_chromium_dirs() now tracks only the shell so an existing profile's stale full build is pruned (~640 MB reclaimed per profile).
Knowledge recall (alpi/core/ + alpi/tools/knowledge_base.py)
Per-profile semantic search over synthesized user/workspace knowledge. The source of truth is Markdown under <workspace>/knowledge/; SQLite under <home>/knowledge.sqlite is only a rebuildable derived index. Raw source files and attachments are read only as inputs for synthesis; alpi does not copy them into a durable documents store.
knowledge(action="search", query, k=5)— hybrid sqlite-vec + FTS search over knowledge pages; page-level results withpath,title,type,tags,snippet,score,links. The index holds one bundle and says which. Usealpi_knowledgefor questions about alpi itself.knowledge(action="ingest", …)— the explicit learn path: resolve a file or a current-turn attachment, extract text (PDF / DOCX / EPUB / HTML / text, OCR on request), have the LLM synthesise durable Markdown pages, updateindex.mdandlog.md, lint, refresh the index. The raw file is not copied. The synthesiser reads the first 12,000 characters of the source; ingest and maintain results carrysource_budget {available, used, truncated}, and the same object is in the prompt so a cut source is not presented as whole. A DOCX keeps its tables in document order, one line per row with cells joined by|, empty rows skipped and a merged cell read once.knowledge(action="maintain", …)— LLM-wiki maintenance: every proposed page is validated before any is written, an existing page is replaced only when the synthesiser saw its full body, and the result reportsbytes_before/bytes_afterso a legitimate shrink is visible.knowledge(action="lint", path?)— requiredindex.md/log.md, minimal frontmatter (type,title,tags,updated_at,sources), relative links, orphan pages.knowledge(action="index", path?, force?)— incremental by a sha256 fingerprint of each page's text (stored inokf_files.fingerprint, added in place to an older index): a page whose fingerprint is unchanged is skipped even when its mtime moved, a page whose text changed is re-embedded even when its size and mtime did not, and a page whose chunks are unchanged but whose metadata moved (tags, type,updated_at,sources) is refreshed without calling the embedder (refreshed_pages);force=truerebuilds theokf_*table family inside one transaction and is the only way to retarget the index at another bundle.
Supported ingest formats: markdown / text / source / configs (stdlib read), HTML (html2text), PDF (pypdf for text-layer, RapidOCR fallback when ocr=true and pypdf extracts < 50 chars), DOCX (python-docx), EPUB (ebooklib), images (PIL + RapidOCR — only with ocr=true). OCR backend is rapidocr-onnxruntime (ONNX port of PaddleOCR, no torch dependency). The PDF/image/OCR extraction primitives live in alpi/extract.py and are shared verbatim with the chat-attachment path (alpi/attachments.py); the DOCX/EPUB/ HTML readers and the chunker live in alpi/tools/workspace.py, a support library, not an agent-facing recall surface.
Shared store primitive (alpi/core/store.py). open_store(home) returns a sqlite3.Connection with the sqlite-vec extension loaded. Designed to host other shapes later (workgroup search, future entity memory) — they bring their own table schemas.
Embedder (alpi/core/embed.py). Embedder Protocol; default FastembedEmbedder wraps the ONNX export of sentence-transformers/all-MiniLM-L6-v2 (384-dim, ~90 MB, no torch). Numerically equivalent to the original sentence-transformers checkpoint but ~10× lighter at runtime. Lazy-loaded under a threading.Lock so concurrent first-touch calls serialize on a single model instance instead of racing.
The bundle uses minimal YAML frontmatter, relative Markdown links, and required index.md / log.md. It is never auto-injected into the system prompt; access happens only through knowledge tool output.
Links resolve relative to the page holding them, the way GitHub, Obsidian and VS Code resolve them, and that is the only form alpi writes. The link graph reads more than it writes: bare, angle-bracketed and percent-encoded destinations, [[wikilinks]] (by path, or by page name when it is unambiguous) and CommonMark reference links, while links inside fenced blocks or code spans are examples, not edges. maintain also repoints a proposed link that is written from the bundle root onto the page it plainly means. Two rules are alpi's own rather than Markdown's, and are the usual reasons an imported vault fails lint: every page needs an inbound link, and the type frontmatter key collides with Hugo's reserved layout key if the same tree is published with Hugo.
Session recall (alpi/tools/recall.py)
Recall over past conversations, the conversational-memory peer of knowledge recall, in three layers: lexical find (session_search, term counts over sessions/*.json), exact browse (session_read, no model call), and opt-in semantic search (index_sessions / recall_sessions) for fuzzy "when did we discuss X / what did we decide about Y".
session_search(query)— lexical first layer; returns the tail thread of matching sessions, active session excluded.session_read(session?, phrase?, start?)— browse layer, no embedding/LLM call: lists recent sessions, or opens a windowed turn slice around an exact phrase orstartindex (paged). Pairs withsession_search(find → open the window).index_sessions(force?)— opt-in (sessions are never auto-indexed): walks<home>/sessions/*.json, builds a per-turn transcript (user:/alpi:lines), chunks + embeds with the samecore/embed.py+ sqlite-vec primitives as the knowledge index, into a separate table family (session_files/session_chunks/session_vec/session_meta) in the sameknowledge.sqlite. Incremental (mtime/size skip); the active session is excluded. A pass — incremental,force, or the rebuild an embedder change triggers — runs in one write transaction, so a failure part-way leaves the previous index searchable, and readers keep seeing it until the new one commits. A session file that cannot be read stops a forced or embedder-change rebuild, keeping the previous index; an incremental pass skips it and keeps its old rows.recall_sessions(query, k=5)— cosine MATCH →[{session_id, when, snippet, score}], active session excluded. Index rows carry the session'sconnection_idanddevice_id(an index built before either column is migrated in place, device ids backfilled from the session files before the first query; a versioneddevice_backfillcount of device-less rows insession_metacommits with the backfill, so an interrupted pass, a completed legacy backfill or rows written by an older alpi rerun it without rebuilding embeddings; a session whose file cannot be read is stored as?unverified, which no device matches undersession_scope: device, and is retried on every query, without writing while it stays unreadable, until the file reads; a query that cannot migrate fails instead of answering unfiltered), and results pass the samecan_read_sessioncheck assession_search(the widening search stops at sqlite-vec's k limit of 4096), so a member sees its connection and a device undersession_scope: deviceits own and deviceless sessions.
Forgettable. Recall is a derived view, so forgetting is real: deleting a session (host.sessions.delete → host/sessions.py::delete_session) purges its rows via recall.forget_session, and index_sessions orphan-sweeps any tracked session whose file is gone. No auto per-turn injection — retrieval is explicit, like the workspace tools.
Workgroup transcript search (alpi/tools/workgroup_search.py)
The third retrieval surface on the same store: semantic search over hub-owned workgroup transcripts. Workgroups are hub-owned by design, so this is profile-local and hub-only — the hub decrypts its own transcript and indexes it; there is no cross-peer / federated search and no global "search all my peers' workgroups". Two tools:
index_workgroups(workgroup_id?, force?)— opt-in: decrypts each hub-owned transcript viahost/workgroup.py::decrypt_transcript(key-history aware, so posts written before a rekey still index), groups consecutive posts into ~2 KB chunks tagged[seq · ts · author], embeds, into a separate table family (workgroup_files/workgroup_chunks/workgroup_vec/workgroup_meta) in the sameknowledge.sqlite. Posts that don't decrypt (rotated-out key, AEAD failure) are skipped. Incremental (transcript mtime/size); emptyworkgroup_idindexes all hub-owned workgroups on the profile. Like the session index, each pass (scoped or globalforce, or an embedder-change rebuild) is one write transaction: a failure leaves the previous rows searchable. A workgroup that fails to decrypt stops a globalforceor embedder-change rebuild, keeping the previous index; incremental and scoped passes list it underfailed_workgroupsand keep its old rows.workgroup_search(workgroup_id, query, k=5)— search is scoped to one workgroup (brute-force cosine over that workgroup's chunks, so per-workgroup ranking is exact rather than a filtered global KNN). Returns[{workgroup_id, seq_start, seq_end, when, authors, snippet, score}]; never returns group keys, ciphertext, or filesystem paths.
Forgettable. Removing a workgroup purges its index in both delete paths — the host RPC (host/workgroup_admin.py::_remove) and the CLI (alpi workgroup remove) call workgroup_search.forget_workgroup; index_workgroups orphan-sweeps any tracked workgroup whose directory is gone. No auto-injection into workgroup turns. ALP encryption/transcript behaviour is untouched — this only reads through the existing decrypt path.
Removal tombstones (alp/secrets/subscriptions.removed.d/). Removing a workgroup writes an empty marker named by its id in every local home, so a stale in-memory copy or a hub-side auto-join heal cannot resurrect it. Markers never cross machines and expire after TOMBSTONES_KEEP_DAYS (2) once the id is gone from subscriptions.yaml and alp/workgroups/; setup → Cleanup offers the expired set under Workgroup tombstones.
Asset prefetch (service.py::_prefetch_assets). Scheduled at boot + 600 s, past the client-reconnection rush. Gated by runtime.prefetch on the root profile: auto (default) fetches the embedding weights only when some profile has knowledge.sqlite and Chromium only when some profile leaves browser un-denied; all forces both; off — the default in Docker — skips it. Every asset still loads lazily on first use, so off costs latency, never functionality. A successful Chromium install prunes stale builds of the headless shell.
Skills
Live under <home>/skills/<category>/<name>/. Required SKILL.md plus optional scripts/, references/, assets/, secrets/ (mode 0700, gitignored, scanner skipped), state/ (gitignored, scanner skipped, runtime persistence). .gitignore auto-written on create with secrets/\nstate/\n.
Live by default — there is no pending-approval stage; the scanner and the sandbox are the guarantees.
Frontmatter (auto-populated on create): name, description, category, version, origin: agent|user, created_at, requires_env, tools, keywords, optional output_schema. 13 fixed categories including miscellaneous as the fallback. secrets/ is filesystem state, not frontmatter: it is created lazily when a skill writes a secret file. output_schema is one-line JSON and uses a deliberately small subset (type, properties, required, items, enum) so the runtime stays dependency-light.
Security scanner (~50 patterns, _DANGER_PATTERNS in alpi/scan.py — the shared scanner library used by skills, memory writes, and the recalled-memory guard): destructive shell, credential exfiltration, prompt injection, persistence (cron/launchd/systemd/authorized_keys/sudoers/shell rc), reverse shells, tunneling, obfuscation (base64/eval/exec/compile), process exec, hardcoded credentials (API keys, OpenAI sk-, GitHub ghp_, AWS AKIA), system-password-file paths, deep traversal. Runs on every create/add_file/patch for files NOT in secrets/ or state/.
Atomic writes everywhere (tmp sibling + os.replace). .bak next to SKILL.md on every edit/patch. Quota: max 40 agent-owned skills, enforced at create.
Auto-injected into the system prompt (skills_index_block(home)): every session start, all installed skills are listed by category as name: description entries, prefixed by a directive that says "check this list before reaching for general tools". Without this nudge, mimo-class models routinely went straight to web_search/terminal even when a perfect skill existed.
TUI integration: when a terminal command's path matches .alpi/(profiles/<p>/)?skills/<cat>/<name>/..., arg_hint rewrites the ToolCard label as skill: <name> (or skill: <name> · <script> when the script is the full path). Tool name stays terminal; the rewrite is display-only.
Execution: skill(action="run", name=...). Single canonical ad-hoc path. If scripts/run.py exists the action validates the skill, then spawns the script via subprocess.run with cwd = skill dir, env += {ALPI_HOME, ALPI_SKILL_NAME, ALPI_SKILL_DIR}, 600s timeout, and the skill's requires_env checked up-front. If the skill declares output_schema, stdout must be JSON and is validated before the call succeeds. Scripts are normal Python; built-in tools and MCP methods are not importable Python APIs. No script → SKILL.md is returned with a [skill X has no scripts/run.py — follow these instructions] prefix so the agent follows the prose and calls the real tools. Scheduled prompts should call this action instead of reimplementing the skill by hand; the scheduler still enters through alpi chat --once --emit-events --no-save.
Structured composition: skill(action="invoke", name=...). Same subprocess/runtime path as run, but stricter: the callee must ship scripts/run.py, must declare output_schema, and stdout must satisfy it. This keeps skill-to-skill composition machine-readable and prevents prose-only skills from pretending to be callable subroutines.
Scripted harness: skill(action="test", name=...). Thin validation layer over the same runtime path. It exists so chat/scheduler/desktop can exercise a scripted skill and verify its declared output_schema without inventing a second testing runtime. If a CLI wrapper lands later, it should call this action instead of duplicating logic.
Research (read-only sub-agent, alpi/tools/research.py)
Spawns a sub-agent with a read-only toolset (web_search, web_fetch, web_extract, read_file, search). Returns a single synthesised report; the main agent never sees the intermediate tool trace.
Depth tiers instead of a numeric max_steps: depth="fast"|"normal"|"deep". The step ceilings are product constants (DEPTH_STEPS_DEFAULTS, 8 / 15 / 30). Locks the model to three buckets (fast = single-answer, normal = comparative, deep = exhaustive); fast and deep double as the model-tier names, so a depth also picks the matching tier when the profile configures one.
Synthesis fallback: when the budget runs out, research forces one final no-tools llm.complete() with "stop investigating, report now". Avoids the "[research gave up]" footgun where the main agent retries the whole thing.
Interrupt: polls tool_state.is_interrupted() between iterations and between tools; returns [research: interrupted] on the first hit. State label during execution: <depth> · step N/M; while an inner tool runs its own emit_state label gets auto-prefixed with step N/M · … via a wrapped _emit installed for the duration of each tool-call batch (restored in a finally).
Batch mode: tasks: [{brief, depth}] up to 3 runs concurrently — see the Delegate section below for the shared ThreadPoolExecutor design (same pattern applies here).
Attachments (alpi/attachments.py)
host.chat.send accepts attachments: [{path, mime?, name?}]. The engine validates them (att.validate — magic-byte sniff for image/PDF, NUL/control-ratio guard for binary-as-text, per-type size caps, allowlist: images png/jpeg/webp, PDF, and text/source incl. py/js/ts/tsx/go/rs/sh/sql) and turns them into OpenAI content-parts (build_content_parts): images → base64 image_url data parts, text/source → inline text parts, PDFs → text extraction (bounded by tools.attachments.max_text_tokens → chars at ~4/token; default auto = half the active model's context window, no page cap). A scanned PDF (extractable text below SCANNED_PDF_TEXT_FLOOR) falls back by model capability: vision-capable → rendered page images; text-only → RapidOCR text (capped at SCAN_MAX_PAGES), so a profile with no knowledge base and no vision can still summarize a scan. PDF text/render/OCR mechanics are shared with the knowledge tool via alpi/extract.py. Images on a text-only model are not OCR'd — they degrade to a path note telling the model it can't see them. A guidance text-part tells the model the files are inline so it doesn't reflexively call filesystem or knowledge tools to "find" them.
Per-turn only. Bytes live only in the in-memory message. session_metadata is itself bytes- and path-free ({name, mime, size}), but the engine re-adds a best-effort local path to each persisted chat-turn attachment so clients can thumbnail history — the path may be unfetchable from another client (outside host.attachments.fetch roots) or after a staged file's TTL, so this is preview replay, not durable storage. The validated turn attachments ({name, path, mime}) are also published to a runtime-only ContextVar (tools/_state.set_turn_attachments) so a tool can resolve a turn's files. Remote clients (mobile, or desktop pointed at a remote daemon) can't hand the daemon a local path, so they upload bytes via the host.attachments.stage RPC (type-aware caps, content validated 1:1 with send) which writes to a TTL-swept temp dir and returns a daemon-side path. Under session_scope: device, host.attachments.fetch serves a remote member device only what it staged (a .owner marker beside the upload; uploads with no marker stay fetchable until the TTL, but an existing unreadable or malformed marker refuses access) or a path that appears in a session it owns: a turn's attachments or output_attachments, a tool's args or result, or the assistant's text, never a path the user typed (a path the user also wrote is not offered even if the agent repeats it), and only exact paths: a file named only in the assistant's text with a space or a parenthesis in its name is not offered, and one the agent lists or reads on its own becomes offered. alpi/host/offered_paths.py keeps that set per session and rebuilds it only when the session file changes. Admins, the Unix socket and session_scope: connection are not narrowed, and the roots and the secrets denylist still apply first.
Durable. knowledge(action="ingest") is the bridge from per-turn input to permanent knowledge. It reads an attachment or source file, synthesizes Markdown pages under <workspace>/knowledge/, updates index.md / log.md, and refreshes the profile-local derived index in knowledge.sqlite. The raw source is not copied into a durable documents directory. There is no auto-learn: attachments stay one-turn unless the user explicitly asks to learn/remember/save/index/compile one.
Vision (alpi/tools/read_image.py)
read_image(path, question) runs the current (or override) model in multimodal mode on an image and returns a text answer. path can be a local file OR an http(s) URL — URLs go through check_url() for SSRF (metadata hosts + private IPs blocked, redirects re-validated via httpx event_hooks).
Magic-bytes sniff accepts PNG / JPEG / GIF / WebP / BMP plus SVG (text-sniff for <svg); rejects bytes that don't match a known header even if the extension agrees. 20 MB cap on file and on download payload.
No pre-flight vision-capability check — LiteLLM's supports_vision() is wrong for openrouter/... prefixes and would bounce real vision models. If the call fails we surface the error with a hint pointing at /model when the message mentions image / vision / multimodal.
Model override via tools.read_image.model in config (surfaced as Vision model by alpi setup, desktop and mobile). When set, read_image and browser's opt-in screenshot analysis try the override first; on failure the tool retries with the main model and prefixes the answer with [fallback: <override> unavailable, used main model]. Clearing it restores main-model fallback. This route is deliberately tool-scoped: chat image attachments are multimodal parts of the main turn and are not silently moved to the override.
Same usage / cost plumbing as research and delegate. Images are auto-resized before upload (see CONFIG.md → tools.browser.vision).
Delegate (write-capable sub-agent, alpi/tools/delegate.py)
Sibling to research, but can mutate: spawn a focused sub-agent with a chosen toolset, get back a summary. Used when a task would otherwise flood the parent context (multi-file refactors, fetch+parse+write pipelines, skills that generate several output files, iterative debug loops).
Toolsets (callable presets via the toolsets param, default ["file", "web"]):
file→read_file,write_file,edit_file,searchterminal→terminalweb→web_search,web_fetch,web_extract
Blocked for sub-agents: delegate (no recursion), memory, skill, schedule, notify, email, session_search, session_read, todo (shared global state). research is not in any preset either — if you need deep investigation inside a delegate task today, do it in the main agent first and pass findings via context.
Budget: max_steps is a per-call tool parameter, defaulting to 30 and clamped to MAX_STEPS_CAP = 100; a non-positive or unparseable value falls back to the default. It's a ceiling, not a target — the sub-agent stops when done.
System prompt is built from a single template plus the workspace root (when set): relative paths resolve under workspace, absolute paths go where the goal says, and the sub-agent is explicitly warned not to invent /workspace/... style roots.
Prompt-cache contract. messages[0] is the stable system prefix and tool schemas are sorted by name. Volatile # NOW, workgroup, skill-hint, and relay state is composed once into the user turn's persisted host_context suffix, so normal history growth is append-only across live calls and every rehydrator. OpenRouter calls carry a hashed affinity for the logical conversation; other providers receive no OpenRouter-only fields. prefix_diag.py compares bounded request-shape hashes per conversation and records causes, never prompt text. Caching and diagnostics are best-effort and cannot fail a provider call.
Batch parallel mode. Both research and delegate accept tasks: [...] (up to 3) and run them concurrently via ThreadPoolExecutor(max_workers=3). Isolation is provided by _state.py: _emit, _interrupt_getter, _usage_sink are contextvars.ContextVar, so each worker thread sees its own values without racing on module globals. Workers re-seed interrupt_getter + usage_sink from the parent context (Python's ThreadPoolExecutor doesn't propagate ContextVars automatically) and install a per-task prefixed emit so TUI progress lines read [i/N] <tag> · <msg>. Results aggregate into one markdown report with per-task sections; per-task failures are captured inline as [failed: <error>] instead of aborting the batch. Cap is hardcoded at 3 — bumping would need a config knob and would multiply LLM cost linearly; not a default worth moving.
TUI (alpi/tui/)
Textual 8.2.x. Layout: AlpiTopBar (identity: version, profile, workspace) + chat scroll (VerticalScroll.anchor() auto-follows new content) + a bottom dock holding the completion popup, the multi-line composer (ChatInput, a TextArea) and the StatusLine (model · ctx % · cost · budget · sandbox · unread inbox · prompts waiting, then context-aware key hints).
Theme (themes.py): build_theme(accent, dark) returns a Textual Theme whose background, surfaces, ink and status colours come from alpi/palette.py, a mirror of common/tokens.mjs kept honest by a parity test. Every grey of both themes, in the console and in the apps, is an equal-channel neutral, so the only colour on any surface is a profile's. palette.resolve_accent maps an unset or legacy brand accent (#c8a24e, #8a5a0a) to the mode's token (#f3efe6 dark, #14110c light, the brand ink); a chosen colour, amber included, is kept. The console (ui.py, alpi profile list) goes through the same resolver. Registered in AlpiApp.__init__ (not on_mount — child widgets read theme_variables during their own mount).
Fold in the console (fold_art.py, fold_shapes.py): the console wears the profile's object. fold_shapes.py is generated by scripts/sync_fold_shapes.py from common/folds.mjs (--check fails when stale; a parity test runs node on the shared module for every object, honeycomb and alpaca), and fold_art.fold_tones is the OKLCH three-tone rule of foldTones, tested for the same hex on every accent and random colours. fold_art.art(fold, accent, rows) rasterises the polygons into half-block cells (▀ with the top sample as foreground and the bottom as background, ▄ when only the bottom is filled, a space when neither, so the terminal background shows through), two cells wide per row. fold_art.identity(home, tui) is the pair a profile wears: the alpaca for the default profile (in palette.BRAND_INK of the theme, the brand accent #14110c light and #f3efe6 dark, mirrored from common/folds.mjs), else tui.fold (the diamond when unset or unknown) in palette.profile_accent. The art draws beside the title of the alpi setup menu (when the terminal is wide enough and tall enough to show the whole menu beside it) and in alpi profile show. fold_art.GLYPHS gives each object one narrow, distinct, BMP glyph (East Asian Width N or Na); fold_art.marker(home, tui) is that glyph in the accent and replaces the diamond that marks the active entry in alpi profile list, the TUI list rows (list_row.set_marker) and the status line. supports_fold_art() is true only on a TTY stdout with UTF-8 encoding and locale, COLORTERM truecolor or 24bit, no NO_COLOR and TERM not dumb; anywhere else every mark falls back to the single-colour ◆ (and ◇ for an inactive profile) exactly as before. Menu cursors (ui.POINTER) and the tool-hint chip are status marks and stay diamonds.
Steps (StepsGroup / ToolCard in widgets.py): each turn's tool calls collapse into one ▸ N steps · Xs row that shows the live step while running. It expands to one card per call: a family glyph (file, terminal, globe, search, link, memory, chip — tool_hints.tool_family), the argument summary, the result hint and the duration; a card opens to the pretty-printed arguments and an output excerpt. Failed calls start open and open their group. ask_user answers render as their own line, not as steps. Ctrl+O toggles the last turn's reasoning and steps.
Assistant streaming: AssistantMessage streams into a cheap Static (flushed every 150 ms) and swaps to a Markdown widget once when the reply lands. Very long user messages render as plain text instead of Markdown.
Reasoning surface:
- While the model thinks, the spinner reads
Thinking…with elapsed seconds (Preparing a step…while tool calls stream). - When the first tool call or answer token arrives, reasoning deltas and any pre-tool prose collapse into one
▸ Thought for Xsrow per turn (ReasoningBlock), expandable to the text. Live turns time it in-process from turn start; replays useTurn.reasoned_sand the consolidatedTurn.reasoning(same segment dedupe ascommon/reasoningSteps.mjs). tui.show_reasoning(defaulttrue) hides the block whenfalse; data is still persisted, the engine still emits.
Persistence contract (cross-surface). The engine consolidates the whole turn's reasoning — reasoning_delta thinking + the inter-tool prose — into Turn.reasoning (str), and records Turn.reasoned_s (float) = the reasoning span from turn start to the first tool boundary, or to the first final-answer text token when there are no tools; it excludes both tool execution and final-answer streaming so the duration isn't inflated by a long-running tool or a long reply. Desktop, mobile and the TUI render a collapsible "Thought for Ns" block from Turn.reasoning, falling back to joining ToolLog.reasoning for turns logged before the field existed. ToolLog.reasoning (first tool of each batch) remains the legacy per-tool fallback.
Slash commands come from one registry (alpi/tui/commands.py) that drives /help and the completion popup (/ or @peer, with descriptions): /help, /activity, /status, /model, /new, /clear, /compact, /sessions, /outputs, /runs, /memory, /skills, /tools, /mcps, /peers, /diff [since], /fold [object] [colour], /attach <path>, /attachments, /clear-attachments, /quit (alias /exit). Panels are FloatingPanels on the overlay layer docked above the composer, dismissed by Esc or click-outside. /activity calls host.activity.list over host.sock (needs you / running / scheduled) and answers listed approvals and questions via host.approval.respond / host.clarification.respond; without a daemon it says so. Configuration verbs (workspace, email, sandbox, …) live in alpi setup — the TUI is for chat and inspection.
Approvals and questions raised by the in-process engine are prompt panels that cannot be dismissed by a click or by opening another panel. They queue (the first shows +N queued), show a live countdown to the engine deadline (60 s approval, 300 s question), and Esc answers deny / cancel immediately. A prompt that times out, shown or still queued, leaves a line in the transcript.
Interrupt: sending a new message, Esc, or Ctrl+C stops the running turn; each also resolves any open prompt so the tool thread never blocks. engine.interrupt_requested is polled at 3 points; long-running tools (research) poll tool_state.is_interrupted(). Skipped tool calls get a [skipped — user interrupted] tool message to preserve OpenAI's pairing invariant. /quit interrupts first, then exits. /clear, /new and Ctrl+L refuse while a turn runs.
Keys: Enter sends; Ctrl+J (or a trailing \ then Enter) adds a line — Shift+Enter too on terminals that report it (CSI-u); pasted newlines are kept; ↑/↓ recall sent messages from an empty composer; Tab completes. Esc answers a prompt, closes the popup or a panel, or stops the turn. Ctrl+C stops the turn, and quits on a second press within 2 s. Ctrl+O toggles details, Ctrl+L starts a new session (same as /clear), Ctrl+Y copies the last reply (pbcopy/wl-copy/xclip/xsel/OSC-52 fallback chain).
Daemon (alpi/service.py)
One alpi daemon per machine, every profile inside. A single launchd plist (com.alpi.daemon) on macOS or systemd-user unit (alpi-daemon.service) on Linux supervises one Python process that hosts every profile under ~/.alpi/ (default plus each profiles/<name>/) on the same asyncio loop. Per-profile tasks are independently guarded — a crash in one profile's scheduler leaves siblings untouched. Tasks are named <profile>/<capability> (e.g. doc/schedule, builder/alp) so logs + asyncio.all_tasks() stay readable. These are internal capabilities, not configurable services:
- schedule — cron tick loop.
- alp — ALP listener (inbound). Serves the full protocol
on a Unix socket plus optional Noise_XK on TCP:
link.ping,link.ask,link.cancel,link.put_blob/link.get_bloband everyworkgroup.*verb. - workgroups — the poller (outbound). Holds
workgroup.pullopen for active subscriptions and uses staggered nonblocking probes for idle or paused mirrors, decrypts new posts, and dispatches an autonomous agent turn when a post mentions this profile or opens a#task. Sibling preempt watcher ticks ~6× faster to abort in-flight responses when a new#tasklands. Independent fromalpbecause direction and lifecycle are different — outbound vs inbound, periodic vs reactive — so a poller crash (timeout against a dead hub, decrypt failure on a malformed post) doesn't take the listener down. Workgroups is ALP; this task is its client half. - host — default profile only. The control-plane Unix socket
(
~/.alpi/host/host.sock) the desktop / mobile client uses to drive the daemon. Refused on non-default profiles fast — the client always targets default's socket and reaches sibling profiles via theprofileparam on each verb.
All capabilities start for every profile; the host plane is default-only. Jobs and workgroups retain their own enabled/paused state, and access control lives in peer grants and connection roles/scopes.
alpi.service.serve_all(root) is the foreground entry point called from alpi daemon start and from the supervising unit's ExecStart. It:
- Walks
~/.alpi/(default + everyprofiles/<name>/) to discover profiles. - Configures the root logger at
~/.alpi/logs/service.log(stderr only when it's a TTY, to avoid double-writes under launchd). - Sets the process title to
alpi (daemon, N profiles)viasetproctitle. - Writes
~/.alpi/service.pid. - Spawns the fixed task set for every profile and waits.
_guard_taskwraps each one so a crash leaves siblings running. - SIGTERM / SIGINT cancels every task cooperatively; PID file removed on exit.
PID 1 (alpi/pid1.py). Inside the image alpi daemon start is PID 1 (ENTRYPOINT ["alpi-docker"], no init). Before anything else it forks: PID 1 stays a minimal init — waitpid(-1) in a loop, SIGTERM / SIGINT / SIGHUP / SIGQUIT / SIGUSR1 / SIGUSR2 forwarded to its only child, exit code mirrored (128 + signal on a signal death) — and serve_all runs in the child, which owns service.pid and service.lock as before. Every orphan in the container (an exited npx wrapper chain, a docker exec session, the grandchildren of a hard-killed turn) reparents to PID 1 and is reaped there, each one logged to stderr. The reaper is deliberately not inside the daemon: a waitpid(-1) in that process would steal exit statuses from Popen.poll() / wait() and from asyncio's child watcher, which then report 0 or 255 for a child they never saw exit. Outside a container os.getpid() != 1 and the fork is skipped. pytest covers the fork, the forwarding and the exit-code mirror; the PID 1 case itself runs in publish-docker.yml, which starts the built image, leaves an orphan through docker exec, asserts no zombie remains and that docker stop returns the daemon's exit code.
Operational invariants of serve_all (each one is the root cause of a real production incident; do not regress):
~/.alpi/service.lockis held under an OS-level non-blocking lock (fcntl.flockon Unix,msvcrt.lockingon Windows) for the daemon's lifetime; this guarantees one daemon per installation. A secondalpi daemon startexits with a warning instead of racing the existing one.~/.alpi/host/host.sockis published BEFORE the TCP plane is resolved or enabled. TCP bind work (resolve_host_tcp_bind,server.enable_tcp) runs off-loop viaasyncio.to_threadand is non-fatal — a TCP failure leaves the Unix socket up, so the local desktop (which talks to the daemon overhost.sock) keeps working when network detection (Tailscale, LAN) is slow or blocked. Mobile and any remote desktop go through the WebSocket transport and do need TCP to come up.- ALP TCP is auto-bound only on
default, or on a profile with its own explicitalp.tcp_port. Other profiles stay Unix-only — otherwise every named profile would fightdefaultfor the same port.
Active home isolation. Because N profiles share one process, tools that resolve their home via home.get_home() would all see the same env vars and write to default's home. The engine wraps each run_turn in a home.set_active_home(self.home) contextvar binding (per-thread); get_home() consults this binding before the env. Without it, another profile's memory tool would write to default's USER.md. See tests/core/test_home.py for the isolation tests.
daemon_status(root) is the snapshot used by alpi daemon status and by alpi setup → Services → Daemon: PID, uptime (via ps -o etime), install backend (launchd / systemd / none), and the per-profile services map.
Host plane (alpi/host/)
Control-plane for the desktop / mobile client. Not ALP — the two share a profile but live on different sockets, with different auth models. ALP is peer-to-peer (Noise on TCP, envelope-signed, peers pinned in peers.yaml); host is client-to-daemon. JSON-RPC-shaped over ~/.alpi/host/host.sock with filesystem permissions as the trust boundary; no peer identity, no envelope, no Noise handshake. Desktop and mobile talk to this API; they do not read profile files directly.
Only the default profile hosts this plane — the client always targets default's socket and reaches sibling profiles via the profile parameter on each verb. _run_host refuses to bind on any other profile even if the toggle leaks via manual config edit.
host.device_state owns the device-facing profile state contract: profile lists/summaries, bounded profile file reads, storage stats, email status/config previews, skill lists, workgroup lists, workgroup member rosters, config field edits, and local Ollama model discovery. The desktop Tauri layer keeps its existing invoke(...) command names for UI stability, but those commands proxy to host.* verbs instead of parsing ~/.alpi themselves. Mobile should use the same verb shapes rather than inventing a separate state API.
Two transports, one dispatcher:
- Unix socket (
~/.alpi/host/host.sock, mode 0600). Local trust = filesystem perms. Used by desktop on the same machine. No token required. - WebSocket (
ws://<bind>:49200by default). Used by mobile and any remote desktop.network.hostdrives the direct bind/address; the bind is derived from it (seeconfig/security): empty → auto-detected Tailscale CGNAT (100.64.0.0/10) then private RFC1918 LAN; a private/Tailscale IP → that IP; a hostname or an opted-in public IP →0.0.0.0(all interfaces); a public IP withouthost.allow_public_bind→ refused (no TCP); Docker →0.0.0.0. Loopback is never a bind target. A0.0.0.0bind leans on the device token (and a firewall/NAT) for access control, soalpi doctorwarns whenever the listener binds0.0.0.0. Per-device token required in every authenticated request'sparams.auth_token.permessage-deflateis negotiated by default (ws_serve(compression="deflate")); JSON-RPC payloads drop 50–80% on the wire. Clients that don't negotiate fall back to raw. Mobile and desktop keep a persistent multiplexed WS pool per(URL, token)so RPCs don't pay a TCP+WS handshake every call — the dominant cost of "remote alpi feels slow" on Tailscale. Streams (host.chat.send,host.events.subscribe) open their own dedicated socket.
Bind and advertised routes are intentionally separate concerns. The daemon chooses where the host-plane server listens; Connections → Network stores an ordered host.endpoints list of complete ws:// or wss:// URLs and chooses what the paired client should dial. Plain WS requires a private IP literal; hostnames require WSS, and synthesized routes pass through the same validator. WSS terminates at a certificate-validating reverse proxy and forwards to the same daemon listener; it does not create another authorization plane. On a normal Mac or Linux install those often collapse to the same private address. In Docker they do not: the daemon binds 0.0.0.0 inside the container while the QR advertises a configured host.endpoints route or a safe private IP derived from ALPI_NETWORK_HOST.
Wire shape (both transports):
{"id": "<reqid>", "method": "host.<noun>.<verb>", "params": {…, "auth_token": "<token>"}}
Unix socket payload omits auth_token — the local transport is sovereign and bypasses token validation entirely. WS requires a valid token except for one exact bootstrap verb: host.connections.exchange_pairing may redeem a locally-created, high-entropy grant once and then the daemon closes that socket. An empty or missing connections.yaml rejects every ordinary WS request (fail-closed). The connection, role, profile scope and grant are created locally over the Unix socket; remote bootstrap cannot choose them.
The daemon writes either a single response line or, for streaming verbs (host.chat.send, host.events.subscribe), multiple frames followed by a done frame and connection close.
This is distinct from ALP peer transport. Connections / host-plane remote access configures how paired desktop and mobile clients reach their own daemon (host.*). Peer TCP listener configures the optional ALP TCP listener other alpis use for link.* and workgroup.*.
Connections and device credentials (alpi/host/connections.py)
The store lives at ~/.alpi/host/connections.yaml (mode 0600). A connection is the operational identity: {id, label, role, profile_scope, status}. Its devices[] each hold the SHA-256 digest of a separate opaque token (token_hash; the cleartext exists only on the client) plus self-reported client/name/version metadata and last_seen. Auth hashes the presented token before the constant-time compare, so a copy of the file is not a credential. Desktop and mobile may therefore share one connection, its sessions and accounting, without sharing a credential. pairings[] holds only hashed temporary grants and lifecycle metadata. A pending grant expires after ten minutes; the first exchange marks it consumed and appends exactly one device under the same file lock. Terminal metadata is kept for seven days, capped at 50 entries per connection, and omitted from host.connections.list.
The daemon resolves each token to {connection_id, device_id, role, profile_scope} and binds that identity to the request context. A hit bumps the device's last_seen at most once per minute. The engine persists connection_id on new sessions; session list, read, continue, cancel and delete reject sessions owned by another connection. The daily ledger records input/output tokens and USD under by_connection; the run ledger records both IDs. Local Unix/TUI/CLI activity uses the synthetic host connection.
Sessions also persist device_id. A connection's session_scope decides whether that matters: connection (the default) shares every session among the connection's devices; device lets a remote device list, read, continue, cancel and delete only the sessions it created, drops other devices' session_changed frames from its event stream and history, serves each device its own latest_session preview in host.profile.summaries (the summary cache is keyed by connection, device and profile), and narrows the agent's session_read / session_search / recall_sessions tools the same way. Every session records its device whatever the scope, so switching to device also hides a device's earlier sessions from its siblings; sessions without a device_id (saved by alpi before 0.15.20, or started by the daemon itself through the scheduler or host.chat.delegate) stay visible to every device of the connection. ask_user clarifications and command approvals carry the owner of the turn that raised them: host.clarification.pending / host.approval.pending, their respond verbs and the clarification.* / approval.* events follow the same rule for member devices. The local socket and admin-role reads are unaffected. A device flagged provisioner (minted from a pairing grant that carried the flag) may call add_device, pairing_status, cancel_pairing and revoke_device for its own connection without the admin role; it cannot grant provisioning, revoke itself, or reach any other admin verb.
Sensitive mutations pass through one dispatcher audit boundary after their handler returns. admin_audit.py writes only allowlisted identifiers and the stable error envelope; it never serializes request params or handler results. The bootstrap pairing exchange replaces its temporary context with the newly created connection/device identity before writing. Authenticated admin denials are recorded at most once per device/method/minute; invalid unauthenticated traffic stays in operational metrics/logging so it cannot churn durable audit history. host.audit.list is local/admin-only, cursor-paginated and filters an identity whether it acted or was the target.
This boundary covers calls through host.sock and authenticated WebSockets. Direct CLI/setup code paths still mutate their stores without crossing the dispatcher and are explicitly tracked as remaining AUDIT.2 coverage rather than being represented as synthetic host-RPC events.
Three trust tiers gate every WS call:
- Unix socket — sovereign. Used by the local CLI and the desktop running on the same machine; bypasses every role check.
- WS admin — full app-level CRUD + connection/device management
(
host.connections.*). - WS member — chat, events, read-only views, schedule listing,
workgroup post/read, voice preview, deleting its own chats
(
host.sessions.deleteonly removes sessions owned by the calling connection; any other id answersnot-found). Admin verbs reject with-32001 forbidden / "admin role required".
The admin set lives in _ADMIN_METHODS; the strictly-local set in _LOCAL_ONLY_METHODS (network admin only — no role unlocks those over WS).
Lifecycle:
- Create connection:
host.connections.create(label, role, profiles, session_scope?)creates the parent identity and a ten-minute one-time grant. The grant is embedded in the QR/link shown byalpi setup → Connections → New connection. The default role ismember; the default scope isconnection. - Add device:
host.connections.add_device(connection_id, provisioner?)creates another one-time grant under the same parent identity.provisioner: trueneeds the admin role and marks the device the grant mints; a provisioner device may call this verb,pairing_status,cancel_pairingandrevoke_devicefor its own connection. - Exchange: the client sends
host.connections.exchange_pairingas its first unauthenticated WS message. The daemon atomically consumes the grant, creates the permanent device credential with its client/name/version metadata, returns it once and closes the bootstrap socket. Reuse returns-32011 pairing-used; expiry returns-32011 pairing-expired. - Observe / cancel: local/admin callers use
host.connections.pairing_statusandhost.connections.cancel_pairing. - Update / disable:
host.connections.updatechanges label, role, profile scope or session scope.host.connections.set_statusdisables or enables every linked credential without deleting sessions or usage. - Use: every WS request carries
auth_token. Fail = JSON-RPC{code: -32000, message: "auth-failed"}and the connection closes; the mobile app's auth-failed handler wipes its endpoint and bounces back to the pair screen. - Revoke / delete:
host.connections.revoke_deviceinvalidates one device.host.connections.deletetombstones the parent and clears every linked token while retaining historical session/ledger attribution. - Policy change: every WebSocket remembers the
role,profile_scopeandsession_scopeit authenticated under, and the authorization watcher closes it with1008 "Authorization changed"on its next pass (ALPI_HOST_WS_AUTH_RECHECK) once the stored connection differs, whoever edited it:host.connections.update, the legacyhost.devices.*verbs oralpi setup connections. An open stream,host.events.subscribeincluded, never outlives the policy it was opened under; clients reconnect and get the new one. A running chat turn on the closed socket is interrupted like on a revocation. A label-only or same-value update leaves sockets open. The one exception is the socket that carries the policy-changing request itself (host.connections.update,host.devices.promote|demote|set_profiles): it stays open for at most two seconds so its own response goes out first; every other socket of the connection, streams and pending calls included, closes on the next pass. - Per-device socket limit: a device credential may hold
ALPI_HOST_WS_MAX_CONNECTIONS_PER_DEVICEsockets. A socket that would exceed it pings the device's existing ones first and evicts those that do not answer within two seconds (ping and pong share one deadline; the transport is aborted, no close handshake;stale_connections_evictedinwebsocket_status()), because a client that lost its network leaves half-open sockets registered until the keepalive fails. A device whose sockets all answer is refused with-32029 too-many-connections.
The daemon migrates the credential store at startup, before it opens the WebSocket listener: a legacy devices.yaml becomes one connection per row with hashed tokens, a pre-0.14.39 connections.yaml is rewritten once with token_hash, and the source is deleted only after the destination has been re-read and verified. A corrupt or empty store is an explicit error, never zero connections, and a failed migration keeps the WebSocket listener down. Rolling back below 0.14.39 is not supported. Leftover copies are listed by alpi doctor; Operations has the removal procedure.
host.devices.* remains as a compatibility RPC alias for older management clients; generated payloads use the new one-time grant contract. Desktop and Mobile continue to consume old QR/link payloads that contain a final token. All new management uses host.connections.*.
Verb namespaces in current shape:
host.sessions.list,host.session.read— read-only.host.audit.list— local/admin-only paginated administrative activity; never returns credentials, values, payloads or chat content.host.workgroup.transcript— read-only,{after_seq?, limit?, tail?}→{posts, next_seq, limit}. Withoutafter_seqthe default istail=trueso first-paint of a long-lived workgroup ships the recent window, not the oldest 200.decrypt_transcriptopens the hub sealed group key once outside the per-post loop (was O(N) Curve25519 unseals per fetch).host.profile.summaries— lightweight inbox/sidebar shape:name,model,accent,fold,latest_session,counts,budget_*,pubkey_b64,has_any_provider. A profile's identity is itsfoldin itsaccent, and the apps draw it wherever the diamond identified a profile (a missing or unknown fold is the diamond; status marks stay diamonds): the twelve objects and the twelve colours (all hues, no grey, so no profile reads as the ink alpaca; thenearestAccentmap from any stored colour, with the retired grey#7e8792read as sky) live incommon/folds.mjsandcommon/accents.mjs, and the design kit reads its shapes from the same file. A new profile (host.profile.create,alpi profile create, or any first bootstrap of a home underprofiles/) is seeded with the next pair of a fixed roulette (appearance.ROULETTE): the least worn colour among the existing profiles, in roulette order, with its own object, so no pair repeats until all twelve are worn. The default profile (the host-plane one, never one underprofiles/) is not one of the twelve: its summary always reportsfold: "alpaca"and noaccent(the apps use the brand ink of the theme), the apps draw the flat ink alpaca for it and show its appearance read-only, andhost.config.set_fieldrefusestui.foldandtui.accentfor it (a client that does not knowalpacafalls back to the diamond). Roster objects draw unfolded (the dashed crease pattern in grey, no fill, no ripple) for a paused profile or workgroup and, for the whole roster, while the active connection is offline, disabled, auth-failed or rate-limited; such rows stay openable, a paused name in a header reads in grey, and a stale working state is dropped. Fromhost.activity.lista row reads needs you over failed (a job whose last run failed in the last day) over working, and a failed job whosejob_idhas a running turn again reads working. The daemon lists the default profile first, and both clients keep it there:common/rosterOrder.mjssplits it (is_defaultor the namedefault) out of the list, and the desktop sidebar, the phone roster and the Fold draw it as the unlabelled first row above Pinned, dimmed when paused or without a provider, never pinnable (a storeddefaultpin is ignored), never behind Show N more, kept under a filter only when it matches like any other row (the desktop checks its name and thealpilabel, the phone also its last message), and absent when the connection does not list it; ⌘1 on the desktop opens it. The desktop has no start screen: startup, a connection switch and a deleted open profile land on the first profile row's latest session once the live roster has answered (on the cached one only while the connection is offline; with no profiles, the empty roster state), and closing Settings returns to its profile or the previous view only when the active connection still serves it; a new session starts from a profile (its row menu, the Sessions menu, or ⌘N, which outside a profile uses the last profile seen on the connection or the first row), and ⌘K opens on New session, highlighted, with the recent sessions across profiles below it. No peers/models/ mcps/provider_keys/sandbox/voice — those live inhost.profile.detail({workspace, tcp_port, advertise_host, provider_keys, provider_ollama, sandbox*, voice_*, mcps, peers, models}), fetched lazily by settings/profile screens. The summaries verb is the hot poll; the detail verb is on-demand.- Workgroup pipeline display — Phase chips on desktop and mobile show completed on an 18% success tint and blocked on a 14% danger tint over the pane. Running uses the selected ground, a semibold slug and the worker's folding object (static under reduced motion); pending uses hover and skipped is struck through. Checks and state words stay out of the chips; accessible names and hover/tap details retain the state. The desktop chat strip starts with the literal
pipeline, followed by phase chips and separators; mobile shows just the chips and separators. Neither repeats the pipeline's name or adds a pinned label or run-status suffix; the header keeps the run's progress and status. host.skills.list— one row per skill:category, name, description, path, size, status(active | inactive | invalid),reason(why, when not active) andkeywords. Passinclude_body=trueto also embed each SKILL.md body.host.skill.read({name, category?})— full structured detail: frontmatter (version, origin, created_at, platforms, tools, keywords),status/reason,requires[](env/bin/config, each resolved or not), thetreeof files (secrets/reports count + mode only, never names), and the SKILL.mdbody(capped 32K).host.skill.file({name, category?, path})— read one file under a skill (SKILL.mdor<subdir>/<file>, capped 256K, binary flagged not decoded).secrets/and symlinks are refused;nameandcategorymust match[A-Za-z0-9_-]+.host.chat.send(stream),host.chat.cancel,host.chat.events_since— run an engine turn for a profile, stream tool / reasoning / assistant events back; cancel via a separate connection that targets the in-flightrequest_id. Every emitted frame is also appended to a per-turn JSONL sidecar undersessions/_events_<session_id>.jsonl;events_since(profile, session_id, after_seq)lets a desktop client whose stream socket died mid-turn replay the missed frames and reconstruct the turn without losing the model's reply. A 5-secondheartbeatframe is woven into the same stream so a long-running tool with no deltas doesn't fool the client's stall watchdog. The daemon-side emit path catchessend_framefailures and switches to "drain + persist only" so the sidecar still capturesreply+doneafter the socket dies.host.providers.*(set_key, unset_key, add_ollama, remove_ollama, add_openrouter_model, remove_openrouter_model),host.peers.{add,remove,pending_list,pending_accept,pending_discard},host.profile.{create,delete},host.mcp.{add,remove},host.email.remove,host.sandbox.{set,network},host.voice.set_voice— config mutations. Each is a thin wrapper around the same internal helper the matching CLI subcommand calls. Thehost.peers.pending_*verbs surface unpinned-sender entries recorded by the ALP server (see ALP.md → Pending invites);pending_listenriches each row withlocal_profilewhen the pubkey resolves to a profile on this machine, so the desktop / TUI can pre-fill the peer id without prompting.host.peers.removeandhost.peers.pending_discardare idempotent: they return{ok: true, existed: <bool>}instead of raising-32004 not-foundwhen the row is already gone, so a stale UI click or a parallel retry never blocks the user's intent.host.workgroup.{create,update,add_member,kick,remove,action,post}— workgroup CRUD, hub-only for create/update/add_member/kick/remove, member-side for action (pause/resume/leave) and post. The desktop Tauri layer routes them through the host plane so mobile reuses the same contract.host.connections.{list,create,add_device,exchange_pairing,pairing_status,cancel_pairing,update,set_status,delete,revoke_device,register_device,summary,usage_daily}— connection/device management and 14-day aggregate usage for the WebSocket transport. The server requires scoped members to passparams.profileexplicitly on every profile-aware RPC (a small allowlist of profile-agnostic verbs is exempt) and returns-32001 forbiddenif missing or out of scope; admin role bypasses by design.host.devices.*aliases preserve the old client contract during migration. List-style RPCs that aggregate across profiles (host.profiles.list,host.profile.summaries,host.workgroups.list,host.approval.pending,host.clarification.pending,host.events.history) are filtered to the device's scope before delivery; the event-subscribe stream drops out-of-scope frames the same way.host.email.probe,host.peers.ping,host.model.ctx_window— diagnostic probes the desktop / TUI used to invoke viaalpi email probe,alpi peers ping, andalpi ctx. Same logic, host-plane entry point.host.peers.pingresolves intra-machine targets bypubkey(not by the peer's localid), so a co-located peer pinned under any alias still finds the rightalp.sock— the alias never has to match the remote profile's name.host.usage.daily/host.usage.workgroup.daily(admin-only) — a 14-day per-day series of token usage + cost; the profile payload also carries atotal30aggregate over the ledger's 30-day retention (the same payload feeds the profile snapshot'susagesection). Profile usage reads theledger.json30-day history (authoritative for ALL spend, including non-token costs like image generation); workgroup usage reads the hub transcript (per-post declared cost). Both bucket by UTC day, so the today figure matches the budget gate /budget_used_usd.host.outputs.{list,read,mark_read,mark_unread,mark_all_read,delete}— durable inbox for proactive agent messages and schedule results. Backed by<home>/outputs/outputs.jsonl(capped at 500 rows, atomic compaction). Rows the scheduler files (a notified reply, a failure, a child agent's own notification) carry thejob_idandrun_idbehind them, so a client can open the job or its run;mark_unreadflips a row back and emitsoutput.updatedlikemark_read, andalpi outputs unread <id>does the same from the console (alpi outputs showprints the source).notifypushes to the OWNER's own apps (native, via the sharedoutputs.create_output_and_emit_messagehelper) and carries the row's singletypeaxis (info|warning|error, defaultinfo). To reach a THIRD PARTY the agent uses theemailtool, which sends over IMAP/SMTP or Gmail directly. Producers:notifyfiles an output for every successful owner push. Attachment-only deliveries with no text body skip the row — nothing to revisit.scheduler/run.pyfiles an output onschedule.failed(always) AND onschedule.donewhen the job notified the owner (notify: true→delivered_to="alpi"). Jobs where the agent notified itself (delivered_to="external") don't get a duplicate row. Silent jobs (notify: false, the default) and stdout-only summaries write nothing — operational noise the user never saw. In schedule subprocesses the parent daemon is the single source of truth: the child'snotifyis suppressed and the parent parses thetool_endargs to file one canonical output with the fulldelivered_tolist. Each row carries{id, profile, created_at, title?, body, type: info|warning|error, status: unread|read, session_id, job_id?, run_id?, delivered_to}(titlepresent when anotifycaller set one, or on scheduler failure rows — "<job> failed", whose body opens with**Reason:**, then**Exit:**/**Timeout:**when they apply and any trace in a fencedtextblock, capped at 2000 characters). Noarchiveaction — the 500-row cap handles retention so clients only render a two-state inbox.agent.message,schedule.doneandschedule.failedevents shipoutput_iddeep_link: /outputs/<profile>/<id>whenever an output was filed so clients can deep-link straight to the row. Deleting is client-side until it lands: both clients hide the row at once and hold every delete of one run in a single 5 s window (common/undoBatch.mjs), so a run of deletes raises one toast that counts them and one Undo that brings all of them back. Each delete restarts that one window, and on desktop hovering the toast pauses it, sohost.outputs.deleteis sent only for the rows the window outlived. Desktop shows at most three toasts; a fourth retires the oldest.
Contract. host.events. is transport, not durable history. The replay window (HISTORY_MAX = 500) is sized for reconnect catch-up within a session of activity — it can drop old rows under load and must never be the source of truth for anything a user can browse. Durable user-visible state lives in the per-profile stores that host.outputs. / host.sessions.* / workgroup transcripts read from. If a UI needs history older than the replay window, it queries those stores, not host.events.history.
host.events.subscribe— long-lived push channel. Daemon emits{event, data, at, seq}frames as state changes. Sources callalpi.host.events.emit(kind, data); loop is captured at first subscription and broadcasts viacall_soon_threadsafe(safe to call from worker threads). Filter optional viaparams.kinds. On connect the daemon sends a{event: "subscribed", next_seq}handshake — clients anchor their cursor here and (if they had a previous one) backfill the gap withhost.events.historyAFTER subscribing, deduping byseq. Subscribe-then-backfill is mandatory: history-then-subscribe leaves a race window where a frame fired between the two calls is counted in the daemon'sseqbut never reaches the client.host.events.history— bounded backfill, seq-only contract:{after_seq?, limit?, kinds?}→{events, next_seq}. Recent events are kept in memory and in<server.home>/host/events.jsonl; the JSONL sidecar is periodically compacted so offline clients can catch up without unbounded growth._load_historypreserves JSONL append order rather than resorting byat— clock skew / suspend-resume would otherwise scramble the replay window. The legacy wall-clocksinceparam is silently ignored; every in-repo client (CLI/TUI, desktop, mobile) advances onseq. Wired kinds:session_changed—Engine.save_session(id + subdir).wg.post/wg.done—workgroup_client.post()(hub-only;wg.doneis detected viatasks_mod.is_done, honouring handle prefixes + line-anchored grammar). Carrywg_id,seq, and a 200-char summary.workgroup_changed(action: created|updated|removed|paused| resumed|left) — workgroup lifecycle fromhost.workgroup.{create,update,remove,action}.workgroup_members—host.workgroup.{add_member,kick}.schedule.done/schedule.failed—scheduler/run.py::tickafter each job dispatch. Carriesjob_id,title,kind,message,reply,delivered_to, andsilent; clients use the explicit fields instead of parsing the operationalmessage. Silent jobs (notify: false) are activity/history only; a job withnotify: true(or one whose agent callednotifyitself) has its reply re-emitted asagent.messagefrom the scheduler daemon so it wakes the owner's apps.schedule.failedremains an interrupt — it adds the jobtitleand a plain one-linebody("reason; exit N; timeout: …", no markdown) plusoutput_id+deep_link(/outputs/<profile>/<id>), and is itself the failure notification (clients raise it; failures are NOT re-emitted asagent.message).agent.message— emitted bynotify(the owner-push tool). In daemon turns it fires from the tool process; for scheduled jobs the parent re-emits after parsing the child subprocess events. Always carriesoutput_id+deep_link(/outputs/<profile>/<id>) so clients land on the canonical output instead of the chat window.output.created— companion event for every new outputs row ({profile, id, type}). Lets inbox surfaces refresh without pollinghost.outputs.list.
Desktop and mobile stream only the active connection. Every other daemon is polled through host.events.history (desktop every 25 s, mobile on its catch-up and background wakes) with the notifiable kinds plus output.created / output.updated: the first raise native notifications, the output kinds only refresh that connection's inbox. Each client keeps one cursor per daemon, shared by its stream and its poll (desktop shares it only between admin routes; a member route sees a filtered stream and keeps its own), so switching the active connection neither drops nor replays an event; a daemon whose next_seq falls below the cursor its request carried restarts it from zero (a late answer below a cursor that moved on since is not a reset). Its replay page then takes only frames whose at is later than the last one the client saw, since a plain restart can restore a counter below the cursor too (history=False emits take a seq that is never persisted), while a frame above the head reported at the reset is always new whatever its clock; with no at seen yet it re-anchors at the head instead. A client's first poll of a daemon anchors without banners but still refreshes that inbox when the page changed it. A polled or replayed request past ts + timeout_s raises no banner.
schedule.changed(action: removed|paused|resumed) — schedule mutators on the host plane.attention.changed({profile, counts, total}) — the set of flagged items ofhost.profile.attentionchanged; refetch it. Dropped from a member's stream likeschedule.*.config_changed(scope: providers|mcp|sandbox|voice|env|<dotted-key-head>) — every cfg.save inalpi/host/config.pyplushost.config.set_field/unset_field.email_changed(name,action: configured|cleared|authorized|removed) — IMAP/SMTP env writes, gmail OAuth success, credential removal.peers_changed(action: added|removed|accepted|discarded) — peer add/remove/pending verbs.profile_changed(action: created|deleted) — profile lifecycle.budget.threshold—ledger.record()when a USD spend crosses 80% or 100% of the daily cap (highest threshold wins when a single record vaults past both). Engine passescfg_budgetinto the record callsite.
host.profile.attention (admin) answers what needs the owner in one profile without opening each panel: memory (files at 90 % of their limit or over it: file, used, limit, pct, over), skills (a skill that fails lint, or is inactive because what it requires is missing: name, category, problem = lint|missing, message), schedules (a job whose last run failed and is not paused: id, title, message, at, last_ok_at), plus counts and total. alpi/attention.py computes it; the scheduler's tick and host.profile.memory_write reconcile it against <home>/attention.json, emit attention.changed when the set changes and file one warning output for each memory or skill item the first time it is flagged (again only after it was resolved and broke again). A failed job already files its own error row. alpi profile show and the TUI /status print the same list.
Adding a new verb: create the handler in the matching host/*.py module, register on host_server.Server.register (or register_stream for multi-frame), and call from the desktop / mobile client via the platform's host-client helper. Never expose a verb outside host.* — the namespace check in register enforces it.
Email (alpi/tools/email.py, alpi/mail/)
Email is an on-demand tool, not a listener — nothing polls the inbox and nothing auto-replies. The agent calls email (actions: list, search, read, send, reply, forward, move, delete, download_attachment) whenever a chat or a scheduled job needs to read or send mail; the tool drives the IMAP/SMTP backend (mail/imap.py::ImapClient) or the Gmail backend (mail/gmail.py:: GmailClient + OAuth). Bodies pulled by email(read) pass through the prompt-injection scanner behind an untrusted-content envelope before the model sees them.
Multi-account. A profile holds N accounts — any mix of IMAP and Gmail — modelled in alpi/mail/accounts.py; each account's identity is its address and its id is a slug of that address. The email tool's account parameter picks which one (by address or id); with one account it defaults to that account. Accounts are declared in config.yaml under email.accounts (non-secret shape only). Secrets live in <home>/.env namespaced per account — an IMAP account's password is EMAIL__<ID>__PASSWORD; Gmail OAuth client creds (GMAIL_CLIENT_ID / GMAIL_CLIENT_SECRET) are shared across all Gmail accounts, while each account's token sits at <home>/secrets/gmail_tokens/<id>.json after a one-off OAuth consent. Add and manage accounts via alpi setup → Email or the apps' Email section; probe / remove a single account by id from the CLI with alpi email probe <id> and alpi email remove <id>. There are no email.* scalar config.yaml knobs beyond the email.accounts map.
Per-profile env snapshot. alpi.home.effective_profile_env(home) overlays os.environ (process-level vars: PATH, HOME, TZ, ALPI_PLATFORM…) with <home>/.env (per-profile secrets, quotes stripped) and is the source of truth for all credentials: the per-account EMAIL__<ID>__PASSWORD keys and the shared GMAIL_CLIENT_ID / GMAIL_CLIENT_SECRET. The daemon never mutates os.environ — under multi-profile supervision a global mutation would cross-contaminate every profile. The contract holds across the agent toolchain: tools/email (IMAP's ImapClient.from_env_map), the LLM-override paths in tools/web_extract / tools/read_image, alpi/identity.py, and the model selector / TUI provider gating. Credential edits via the host plane write the file atomically; a running engine reads the current .env on its next turn.
Schedule (alpi/scheduler/)
Tick loop (default 30s) hosted inside the alpi daemon. add schedules a job (kind: cron|once, expression or after_hours). run-once ticks manually for testing. LLM time grounding: when the agent calls schedule(action='add', kind='once', after_hours=N), the engine resolves now from a single source so the agent doesn't drift.
Duplicate guard + in-place edits. add rejects a job whose (kind + cron / run_at / after_hours) matches an existing one AND whose prompt fingerprint (lowercase + whitespace-collapsed first 80 chars) collides. Pass force=true to bypass when the second job is genuinely intentional. Use update to change prompt, cron, notify, or pause state without remove/recreate churn. A job carries a single delivery axis, notify: bool (default false = silent): true pushes the reply to the owner's apps. Failure is not on that axis: a failed job always files an error output and raises schedule.failed regardless of notify, and a run the scheduler could not end does too — killed with the daemon (0.14.49), or wedged past its timeout inside a live one (0.14.50) — through the run sweep described under Runs: a dead child within about a minute, a wedged one roughly ten minutes past its timeout. Legacy jobs with a platform field are migrated to notify on load (platform set → notify: true). Reaching a THIRD PARTY is an explicit email call in the prompt — that's now allowed (the old auto-delivery guard that rejected such prompts is gone).
Scheduled jobs execute through alpi chat --once --emit-events --no-save with ALPI_PLATFORM=cron. The scheduler consumes stdout events to detect tool traces, final reply text, delivery, and failure. It does not write sessions/<id>.json: cron output belongs to schedule delivery/logging, not to local TUI / desktop chat history.
Loop isolation. serve() runs tick() in a dedicated ThreadPoolExecutor(max_workers=2), and host.schedule.fire wraps fire_by_id in run_in_executor before awaiting. Both paths ultimately call subprocess.run(timeout=job_run_timeout(job)) (default 900s, per-job up to 86400s); running them inline would block every other coroutine on the daemon's asyncio loop — ALP responders and host.chat.send streams in sibling profiles all stall for the duration of the scheduled job. The dedicated executor also means the scheduler can't starve chat's default-executor turns. A regression test in tests/core/test_schedule.py::test_serve_runs_tick_off_loop_so_chat_can_progress pins the contract.
First run. A cron job runs at its next occurrence, never on the tick it appears. The schedule tool records last_run_at when it adds a job; a job that arrives without run state (written into jobs.json by hand or by a deploy, or whose schedule/runs.json entry was lost) gets first_seen_at in runs.json from the first tick that sees it, or, if it arrives paused, from the first tick after it is resumed. A fired job is stamped (last_run_at, last_run_status; a one-shot that succeeded is removed) the moment its run returns, before its outputs and events are written and not at the end of the pass, so a daemon restart in the middle of a long pass does not fire a job that already finished. A run that raises (the agent subprocess cannot start, a file is missing) is a failed outcome stamped and reported like any other failure, and the pass goes on with the next job. Each job is re-read from jobs.json just before it fires and skipped if it was removed, paused, fired by hand or no longer due meanwhile, and fires from the fresh copy; a due check that raises skips that job with a logged reason; a stamp that cannot be written is logged and kept in memory, laid over the job on every later tick while the job's kind and run_at are the ones that ran (so it is not fired again, and a one-shot stays gone, but a job edited into something else in the meantime is left alone) and retried first thing on each tick until it lands, without overwriting a newer stamp on disk (a daemon restart while the disk keeps failing loses it); an unreadable jobs.json mid-pass ends the pass. fire_by_id stamps the outcome (a failed stamp is logged, not raised) before it emits schedule.done, so the first host.activity.list after activity.changed already reads it. schedule(action="fire") runs a job now.
Job fields. A job carries an optional one-line description (at most 160 characters, set through schedule(add|update, description=…) and listed by host.schedule.list), and its run state keeps last_run_message (the redacted failure, 300 characters, cleared by the next success) and last_ok_at.
Timezone. Cron expressions evaluate against the machine's system timezone (datetime.now().astimezone() in scheduler/run.py). Jobs are stored with UTC last_run_at but fire according to local wall-clock time. Practical consequence: if you specify 10 12 * * * because you want a 12:10 reminder in Bangkok, the Mac must be set to Asia/Bangkok. Move the machine to a different timezone and the cron fires at 12:10 there, not in Bangkok. No in-job timezone override today — add it via TZ=… in the launchd plist / systemd unit if cross-timezone stability is required.
MCP client (alpi/mcp/)
Spawns user-configured MCP servers (stdio JSON-RPC, SSE planned). Their tools are wrapped and registered as alpi tools. Servers configured in config.yaml under mcp.servers.<name> (command, args, env). Management lives in alpi setup → MCPs and in the alpi mcp command group (add, remove).
External orchestration frameworks. Alpi does not embed LangGraph, CrewAI, AutoGen, or similar graph/supervisor runtimes in core. They overlap with Alpi's own agent loop and bring a heavier dependency, state, and observability model than the local-first runtime needs. Interop belongs at the edge: expose the external workflow as an MCP server and let Alpi call it as a tool, or wrap a local workflow in a scripted skill. ALP is not the adapter layer for these frameworks; ALP is reserved for sovereign profile-to-profile collaboration across machines, while MCP is the interop layer for external runtimes and tools.
Logging (alpi/_log.py, alpi/logs.py)
Every subsystem writes to a single flat folder: ~/.alpi/logs/<subsystem>.log, rotated at 1 MB with 3 backups (MAX_BYTES / BACKUP_COUNT in _log.py). Same format everywhere (%(asctime)s %(levelname)s %(name)s %(message)s) so alpi logs can merge them by timestamp prefix. The source tag on display comes from the filename.
Three sources today (file on disk + the writer that produces it):
service— the unified orchestrator's root log: subsystem start/stop, scheduler ticks, ALP listener traffic, delivery errors. Written byalpi.serviceand every subsystem that logs through the root logger.agent— one line per TUI/schedule-triggered turn: session id, elapsed, tool names, reply size, cumulative cost, truncated user prompt. Written byengine.py::run_turnviaget_subsystem_logger(home, "agent"). This is the cross-session grep index —sessions/<id>.jsoncarry the full detail;agent.loglets you answer "what has alpi been doing this week?" without iterating JSONs.approval— one line per non-SAFE terminal command classification (ALLOW / DENY with severity, pattern, reason). Written bytools/_approval.py. Security audit trail; complements the per-turn detail insessions/.
The machine-wide structured administrative trail is separate: ~/.alpi/logs/admin-audit.jsonl, 5 MB plus three rotated generations, mode
- Each row is capped at 4 KB, bootstrap/auth failures have their own
one-row-per-minute budget, and target fields are allowlisted per method. It is JSONL because Desktop and host.audit.list filter by actor, target and result. alpi audit-log renders it for the console. It is not included in alpi logs: those commands merge human-readable .log streams, while this trail has its own bounded query contract. Chat turns are not copied into it; sessions already carry their owning connection_id.
The alpi logs --source CLI choice list also accepts schedule. Inside the unified daemon, scheduler events route through the root logger and land in service.log — the filter value is kept so that any standalone or legacy schedule.log (e.g. from an older scheduler.run.ensure_running() invocation that ran out-of-process) stays selectable.
Why logs are NOT inside sessions/: sessions/ is a structured store (one JSON per conversation, indexed by id, consumed by session_search and the resume flow). Mixing freeform logs would break the glob pattern and the cleanup semantics. Logs are the index and audit trail; sessions are the content. Peers, not nested.
Why one flat folder (logs/) instead of per-subsystem dirs: tiny <subsystem>/logs/ folders with a single file each is pure noise. The service keeps non-log state in its own places (schedule/jobs.json, alp/alp.sock, service.pid at the profile root) — only the .log files consolidate.
Adding a new source is two lines: from alpi._log import get_subsystem_logger; logger = get_subsystem_logger(home, "my-sub"). alpi logs picks it up without changes; add the tag to the --source choice list in cli.py::logs_cmd if you want it filterable.
Doctor (alpi/doctor.py)
alpi doctor — live health check. Verifies external capabilities actually respond, not just that they're configured. Same entry point from the CLI and from alpi setup → Health check; the status in the setup menu row (all green / N warning(s) / N failing) runs the full check too.
Checks:
- Model —
cfg.modelset + provider's API key present in.envor env. - Workspace — configured + exists + writable.
- Email (live) — IMAP login + SMTP handshake, Gmail OAuth token refresh.
- Service — daemon installation + PID checks distinguish "installed but dead" from "running" from "not installed".
- MCPs (live) — spawn each configured server,
list_tools, stop. Parallelised; per-server timeout 8 s. - Security — sandbox backend binary on PATH (if
tools.terminal.sandbox: true), approval allowlist count. - Dependencies — the installed LiteLLM against the pin alpi's own metadata records, and every file of the distribution against the sha256 digests its installer wrote to
RECORD. Catches a local swap or edit of the one dependency every provider call routes through. The pin row is sync; the file hashing runs in the pool with the network checks, so it never delays the first frame. An absentRECORDwarns rather than fails — the check reports what it could not verify instead of claiming it did.
Parallelism: the network-bound tasks (IMAP/Gmail/MCPs) submit to a ThreadPoolExecutor(max_workers=8). Sync checks (model, workspace, services, security) run on the main thread while the pool works. Total wall time ≈ slowest single task, not sum — ~5-10 s on a healthy profile.
Progressive rendering: run_and_render() uses rich.live.Live — every row appears immediately with a cyan spinner, each resolves to ✓/✗/! as its future completes. Animation at 10 fps via a manual frame cycler (rich's Spinner objects can't be appended to Text). Layout is stable (same rows, same column widths) so the eye doesn't jump.
Exit codes: 1 if any check returns fail, 0 for warn/info/ok. Warnings don't break cron. The wizard entry ignores the exit code — it press-enter-waits so the user can read.
Ops digest (alpi/ops_digest.py)
alpi digest [--since 7d] is the read-only evidence rollup for operator decisions. It deliberately does not own new state: each section reads the primitive owned by another subsystem.
- Tools — current availability report from
alpi.tools. - Skills — summary from
skills_usage. - Memory — promotion queue counts plus memory-file pressure.
- Compaction — event count and after/before ratios from
logs/compaction.jsonlover the requested window.
The command has two renderers: a compact Rich view for humans and --json for scripts. The JSON is a dataclass dump of the report shape. It is not an observability daemon, dashboard, recommendation engine, or telemetry channel. Tests pin the read-only contract by snapshotting the profile tree before and after a digest run.
Sessions (alpi/session.py)
Turn-based JSON: schema_version: 2, turns: [{at, user, tools[], assistant}], and cumulative metrics. ToolLog carries at, name, args, result, ok, duration_s, reasoning; large user / assistant / reasoning / tool payloads are persisted as bounded previews plus {bytes, sha256, truncated} metadata, not raw unbounded blobs. host.session.read normalizes both legacy and v2 payloads back to the client-facing shape, so desktop/mobile can render old and new sessions the same way. Empty sessions (no user message) are NOT saved.
Listings are bounded. host.sessions.list fully parses normal files but uses a cheap summary path for files above the large-session threshold; host.profile.summaries uses count_sessions() and latest_chat_summary() so profile/sidebar RPCs never parse 50 MB histories just to show a count or latest row.
Live replay is a separate sidecar (sessions/_events_<id>.jsonl). It is append-only within the active turn, sequence-numbered, and bounded: incremental assistant_delta / reasoning_delta frames are preserved exactly, while very large text fields are clipped. The canonical durable review remains sessions/<id>.json; the sidecar is for reconnect/backfill, not long-term full-fidelity storage.
sessions/ is local human chat history: TUI, desktop, and manual alpi chat --once runs that should be resumable. --continue, tui.auto_resume, host latest_session, and desktop profile opening all treat only kind == "chat" as resumable local history. Historical files whose first user message starts with [SCHEDULED:], [workgroup-poller], or another system bracket are ignored by resume/profile history.
TUI resume. Bare alpi resumes the most recent session when tui.auto_resume: true; -c / --continue is the manual override.
Scheduled jobs do not persist session files. The scheduler uses --no-save because it only needs emitted final reply/tool events for delivery and audit; keeping a resumable transcript would make background jobs appear as user chats.
@-mention threads (alpi/alp/mention_thread.py). When peer A @-mentions peer B over ALP (link.ask), the receiving side runs a fresh Engine per turn — but B persists a small thread at <B-home>/mentions/<A>@<conversation>.json, capped at 20 turns, where conversation is an opaque id A derives from its own source session (the peer tool reads it from the run context; the host chat, TUI and --once pass their session explicitly). Successive mentions from the same A conversation carry conversational memory ("what I said before" resolves) without polluting B's local --continue (which only reads sessions/); a new A conversation starts clean, and the same conversation value from peer C selects nothing. A request whose conversation cannot be established runs with no history and writes none; a caller that sends no conversation at all predates the identity and keeps the legacy <A>.json thread, which is never imported into conversation threads. The result's history field says which applied, so a sender can tell when a target ignored the identity. Wipe via setup → Cleanup → Mentions.
Security model
Two layers:
- Layer 1 — application guards (always on).
_guards._DANGEROUSdenylist on terminal (rm -rf, pipe-to-interpreter, fork bomb, ...). SSRF block on web_fetch/web_extract (RFC 1918, link-local, cloud metadata). Prompt-injection scan on email + web content. Sensitive-path denylist on file tools (_paths.py). - Layer 2 — OS sandbox (normally opt-in, mandatory for scoped bare-metal pipeline turns).
tools.terminal.sandbox: truewraps shell commands insandbox-exec(macOS) orbubblewrap(Linux). Ordinary sandboxed turns can write workspace +~/.alpi/+ temp storage; network is denied by default. A bare-metal workgroup phase carryingpathsforces the wrapper even when the profile setting is off: workspace and~/.alpi/become read-only, only exact literal paths and trailing/**subtrees are writable, ambiguous globs fail closed, and temp writes use isolated storage. A scope-only wrapper preserves the network access of a profile whose sandbox is disabled. In the official Docker runtime, scoped turns omitterminalentirely; native file tools enforcepathsand daemon gates run outside the agent tool surface. Unscoped interactive turns keep the profile's configured behavior.
Threat model: prompt injection via email/web content, LLM-issued tool calls on the user's machine, direct user input (trusted), and network adversaries for ALP links. Full discussion in Security.
Cross-cutting concerns
Profiles
alpi -p <name> resolves home to ~/.alpi/profiles/<name>/. ALPI_PROFILE env var is the same. No sticky "current profile" file — resolution is fully explicit. The single daemon (com.alpi.daemon / alpi-daemon.service) supervises every profile from one process; tasks are namespaced <profile>/<service> so they stay distinguishable in logs and asyncio.all_tasks(). Inside a turn, home.set_active_home(home) binds the per-thread contextvar consulted by home.get_home() so tools resolve to the right profile even though every concurrent turn shares the daemon's env.
Workspace
cfg.workspace (or cwd fallback if unset) is the default root for relative paths — not a wall. File tools and terminal can reach absolute paths anywhere except the sensitive denylist. Real workspace-only isolation is the opt-in OS sandbox (Layer 2). Configure it via alpi setup → Workspace; the TUI top bar read-outs the resolved path but does not edit it.
Dependencies
Hard runtime deps are kept tight — every line in pyproject.toml's dependencies is actually imported by alpi/. The audited set, with one-liner for why each earns its place:
litellm— multi-provider LLM client; the one primitive the agent is built around.rich— Text formatting primitives used across the CLI wizards, TUI rendering pipeline, and tool output.textual— TUI framework.prompt_toolkit— CLI wizard input (menus, text, password).httpx— async HTTP; Gmail API, web_fetch, OAuth dance.click— CLI command dispatch.pyyaml— config.yaml + skill frontmatter.python-dotenv—.envloader.croniter— cron expression parsing for the scheduler subsystem.setproctitle— makesps auxshowalpi (<profile>)instead of identicalalpilines for every profile's service.playwright+playwright-stealth— interactive browser tool.pillow— image pre-processing forread_image(auto-resize).html2text— strip HTML to markdown inweb_fetch/web_extract.ddgs— keyless search across public engines;web_searchasks one engine at a time in a configured order (replacedduckduckgo-searchwhen that package was deprecated).edge-tts— TTS tool (local-first, no API key).faster-whisper— STT tool (local-first, no API key).cryptography— ChaCha20-Poly1305 for encrypted backups.websockets— the host plane's WSS server.httpcore— pinned-DNS transport, so a resolved address cannot be swapped under an in-flight request.sqlite-vec— vector search inside the knowledge SQLite store.fastembed— local embeddings for that store; no embedding API key.pypdf+pypdfium2— PDF text extraction, then page rasterisation when a page carries no text layer.rapidocr-onnxruntime— OCR for those rasterised pages, local.python-docx+ebooklib— Word and EPUB readers for the workspace tools.python-gnupg— PGP sign/encrypt in the email tool.qrcode— renders the pairing link as a scannable code in the terminal.
Optional dev extra: pytest + pytest-asyncio for the test suite, ruff for lint, pip-audit for CVE scans.
Security posture: uv run --with pip-audit pip-audit must run clean against the full lockfile before each release. Known-CVE deps are not allowed to accumulate — drop or upgrade.
Testing
python3 scripts/validate.py is the one command before claiming done: it runs the release check and every suite the working tree touches (see AGENTS.md). Directly, uv run pytest -q is the fast suite, --integration adds sockets and sandbox-exec, and --llm enables real-LLM tests (a few cents on free models).
Key fixtures (tests/conftest.py):
tmp_home_no_env— isolated~/.alpi/rooted at a tmp dir, no.env(safe for unit tests).tmp_home— same with the user's.envcopied (for LLM tests).
Contracts clients and consumers rely on
Breaking one of these breaks a client, a gateway or a peer. Change the contract, its consumers and this section in the same change.
- The desktop / mobile client talks to the daemon, not the filesystem.
Verbs in the
host.*namespace (inalpi/host/) are served over~/.alpi/host/host.sock(Unix socket, 0600 + same-user trust boundary; no Noise, no pairing). When adding a desktop feature, add ahost.*verb — never read~/.alpi/directly from Rust, never spawnalpias a subprocess. ALP (alpi/alp/) is a separate plane for cross-machine peer-to-peer (link.*,workgroup.*) and is not what the client calls. The one exception is the local daemon's lifecycle, which nohost.*verb can serve while the daemon is down:local_daemon_stateclassifies this machine asrunning(host.versionanswers on the local socket),stopped(analpibinary onPATHor the usual install dirs, or a launchd / systemd unit exists;~/.alpialone does not count, the app creates it),absentorunsupported(Windows), checking existence only;local_daemon_startkickstarts the supervisor unit, or without one runsalpi daemon startdetached in its own process group (with the install dirs on itsPATH, stderr tologs/desktop-start.log, reaped when it exits, and not spawned again while a previous one lives), waits up to 45 s of wall clock forhost.versionon the local socket and on failure returns the tail of the launcher's stderr and oflogs/service.log. It never installs a service. - First run is one welcome, not an error. On a local connection that
has not answered yet in this session, the desktop replaces the main pane
by state (once it has answered, an outage keeps the open view under the
reconnecting banner, whose Retry runs this flow again):
stoppedstarts alpi once on its own, then offers Start alpi; while starting it shows the steps (start requested, waiting for alpi to answer); a failed start shows the error,alpi daemon startto run by hand and Retry;absentoffers the install commands (re-detected when the window gains focus) beside a pastedalpi://link;unsupportedoffers only the link. Once alpi answers, a one-time card says how to add a phone. Both clients pair through the same three named steps (link read, reaching the host, signing in) and name each failure with its next action fromcommon/onboarding.mjs; a used, expired or malformed link is cleared, any other failure keeps it. The phone's Scan QR opens the camera directly, Paired names the host and the role, and Open inbox clears the stack. A member with nothing shared sees whom to ask, naming the device. host.versionsays whether a daemon can update itself. It returnsinstaller(uv|pipx|docker|source, cached after a successful probe byupdater.install_kind;dockercomes from the deploy runtime) andself_update, true only foruvandpipx. A client shows its update button only whenself_updateis not false (a daemon that omits it keeps today's behaviour) and otherwise shows the manual step fromcommon/updateHint.mjs, whichalpi updateprints too: for Docker the image tag to set indocker-compose.yml, for a source installgit pulland a restart.host.daemon.updateon such a daemon answersreason: "manual"and runs nothing.- Engine
assistant_doneevents:final=Truemarks the deliverable. The engine emitsAgentEvent(kind="assistant_done", ...)for every assistant message, including preamble narration that comes before tool calls ("Let me check things first.", etc.). Only the event that closes the turn carriesfinal=True. Consumers that build the canonical reply (scheduler delivery, gateway, ALP) must filter onev.final; otherwise preamble leaks into the message users receive. The TUI is the exception — it consumes everyassistant_doneto rewrite the active bubble, which is correct for live streaming. - Two messaging intents:
notify(owner) vssend_message(third party).notify(text, title?, type?)pushes to the OWNER's own paired Alpi apps — it files an inbox row in~/.alpi/outputs/and emits theagent.messagehost event (the only native push).typeis the single presentation axis:info(default) |warning|error.send_message(text, channel, chat_id?, attachment?)reaches a THIRD PARTY through a gateway (telegram / imap / gmail / matrix / webhook) —channelis required, there is no owner channel, and it carries notype(its inbox rows are alwaysinfo). The shared native-emit helpers live inalpi/outputs.py(create_output_and_emit_message,_suppress_native_emit). Clients must surface everyagent.message— do not suppress it (e.g. for the active chat). schedule.done/schedule.failedevents carry structured output. The scheduler tick emits{profile, job_id, title, kind, message, reply, delivered_to, silent}on the host event bus.messageis the operational status for daemon logs and ops UIs.replyis the clean agent/script output, capped at 2000 chars in the event, intended for native notification bodies; the archived inbox row keeps the whole reply up to 8000 chars and, past that, ends with a line saying where it was cut. A job has one delivery axis,notify: bool(defaultfalse).delivered_tois""(silent,notify:false) |"alpi"(notify:true→ the daemon re-emits the reply asagent.message) |"external"(the agent callednotifyitself → no duplicate). Failures always file anerrorinbox row and emitschedule.failed, regardless ofnotify; the failed event carries the jobtitleand a plain one-line summary, the row is titled "<job> failed" with a markdownbody(the reason — the exception line of a traceback — then exit and timeout; a timeout also says which tool was in flight and for how long, how many tool calls ran, and the agent's last message), andschedule.failedis the single failure notification — it is NOT also re-emitted asagent.message.silentmeans a successful job produced no user-facing output. Do not parsemessagein clients when an explicit field exists. When changing the contract, update desktop/mobile consumers and bump the docs here.host.network.*is the canonical network config surface for desktop/mobile.host.network.statusreturns{scope_in_use, host_in_use, is_override, port, device_name, endpoints, is_endpoints_override, candidates: {tailscale, lan, configured, docker}, diagnosis}so clients can show the live pairing endpoint AND let the user pick a different one without dropping toalpi setup.scope_in_useis the network character of the host (tailscale | lan | custom | docker) computed vianetwork.classify_scope— NOT the resolution path.is_overridecarries the "this came fromcfg.network.host" bit separately.host.network.set_advertised({host, device_name, endpoints})persistscfg.network.host,cfg.host.device_name, and the orderedcfg.host.endpoints; empty values unset their override. Endpoint URLs accept onlyws:///wss://, reject credentials/paths and public plaintext WS, and are advertisement metadata — they do not change the daemon bind. Emptyhostunsets the override (back to auto-detect). Validation rejects public IPs (token leak), loopback, multicast/link-local/reserved, and malformed hostnames — accepts RFC1918, Tailscale CGNAT (100.64/10), and any valid hostname.host.network.restart_host_serverends the current daemon process (supervisor respawns with fresh config) and is the explicit handshake clients use after writing. Known gotcha: a stale override (e.g. Tailscale IP saved in config but Tailscale now off) still classifies astailscalebecause the IP literally is one, but the daemon won't be listening on it — clients should comparehost_in_useagainstcandidatesto detect this and warn.host.activity.listis the one "what needs me / what is running" read. It aggregatesneeds_you(pending approvals + clarifications, each withsession_id),running(turnrows from the in-process registry inalpi/host/activity.py— engine turns, workgroup dispatch inservice._dispatch_workgroup_turn, scheduler fires viascheduled_run— plusworkgrouppipeline rows from the cached fold) andscheduled(admin-only). A workgroup seen from both its hub and a member profile is one row, the hub's, chosen among the profiles the caller'sConnectionContext.profile_scopeallows, so a caller limited to the member profile keeps the member's row. It reads memory and stat-keyed caches only, because clients call it on everyactivity.changed {profile}; never add a per-call scan ofruns/or transcripts. A new long-running source registers withactivity.start_run/end_run(ortracked_run), and a new event that changes what the verb returns joinsactivity._TRIGGERS.activity.changedis live-only (emit(..., history=False)) so it never evicts replay rows; the read path never writes the phase-change baseline. Running turns use thecan_handle_promptrule, the same check asneeds_you. Chat frames:tool_start.started_at,tool_end.duration_s, and onereasoning_done {seconds}per reasoning span (consecutive deltas closed by the next tool call, text delta or step end);secondsis time spent reasoning — the first span of a step counts from the model call (equal to the storedreasoned_swhen step 0 streams no prose first), later spans in the same step from their own first delta; reasoning a retry or fallback discards never reaches a span, and all span texts of a turn share the turn's reasoning cap. The engine measures spans once (_ReasoningSpansemitsreasoning_doneAgentEvents;host.chatonly forwards them) and stores them on the turn as orderedreasoning_spans: [{seconds, before_tool, text?}],before_toolbeing the index in the turn'stoolsof the first call after the span (len(tools)when it precedes the answer) andtextthat span's own reasoning (clients place replayed text by span, never by splitting the joinedreasoning;tools[].reasoningis only the inter-tool prose); absent on turns without reasoning and on pre-field sessions, where clients fall back toreasoned_s. Each step restarts the span clock at its model call.session_changed.in_flightis true only for the in-flight stub save; a crashed chat turn still closes within_flight: false.- Session ownership is
(connection_id, device_id), gated by the connection'ssession_scope.Sessionpersists both ids; every host verb that lists, reads, continues, cancels or deletes a session goes throughconnection_context.owns_session(connection_id, device_id)(row form:owns_session_row; agent tools usecan_read_session, which keeps the admin bypass). Undersession_scope: connection(default) the device clause is a no-op; underdevicea remote device sees only sessions carrying its owndevice_id, and sessions with nodevice_id(pre-flag, scheduler,host.chat.delegate) stay visible to the whole connection. The local socket never applies the device clause. The events that carry session text,session_changed,chat.turn_done(the first 200 characters of the reply),file_mutations(a diff preview) and theapproval.*andclarification.*prompts, carry both ids, andserver._filter_session_events(_OWNED_EVENTS) drops foreign ones for members, next to the role redaction; a new event that carries session text joins that set. Never filter byowns_connectionalone in a new session verb. Those two events recorded without an owner are hidden from members altogether.tests/host/test_device_scope_matrix.pywrites as device A and reads as device B through each path it lists (the session verbs, replay, runs, activity, prompts, summaries, events and the session tools); a new path is one more row. Profile-wide data is shared by design: workgroup posts, memory and the files an admin may read. The known gaps, turns a workgroup post wakes (unfenced, SCOPE.11), staged attachments, the profile session count and the admin reading of the scope, are tasks indocs/ROADMAP.md. A device withprovisioner: truemay call the_SELF_SERVICE_METHODS(add_device,pairing_status,cancel_pairing,revoke_device) on its ownconnection_idwithout the admin role; those verbs are_SCOPE_FREE_METHODSbecause they carry no profile.
Non-obvious things to know
rich.markup.escape()any user-controlled substring before passing toText.from_markup(). Several past crashes from[exit 0]-style tokens in tool output.- Tool results are capped per-tool by
alpi/tools/_budget.py(default 100,000 chars; override viatools.<name>.max_result_chars). last_ctx_tokens(current prompt size) ≠ cumulativeinput_tokens. Header shows the former.call_from_thread+ Python built-in methods (e.g.dict.pop) crashes Textual; always wrap in a regular function.cfgmust be loaded BEFOREsuper().__init__()onAlpiApp. The theme is then registered immediately after, in__init__rather thanon_mount, because child widgets readself.app.theme_variablesduring their own mount (which fires first).self.get_css_variables()is called explicitly to rebuild the var dict synchronously — settingself.themealone schedules the refresh for the next event-loop tick.- Schedule subprocess uses
alpi chat --once --emit-events --no-save— same event stream, no resumable session file. ALPI_HOMEenv var routes daemons + tests to a specific profile root.ALPI_SKIP_UPDATE_CHECK=1short-circuits the background PyPI version check (alpi/updater.py); the autouse fixture intests/conftest.pysets it so the unit suite never reaches PyPI.ALPI_UPDATE_INDEXoverrides the JSON URL the updater hits when you need a staging or local mirror.