Docs/Architecture

Architecture

How alpi is built, for contributors: code structure, turn loop, memory, scheduler, MCP.

Project·88 min·v0.17.8
On this page
  1. What alpi is
  2. Principles
  3. Code conventions
  4. CLI surface
  5. File layout
  6. Profile home layout (~/.alpi/ or ~/.alpi/profiles/<name>/)
  7. Core systems
  8. Cross-cutting concerns
  9. Dependencies
  10. Testing
  11. Contracts clients and consumers rely on
  12. Non-obvious things to know

Living technical reference for alpi at HEAD. Describes only what currently ships — historical decisions live in commit messages, planned work lives in ROADMAP.md.

Audience: any developer (or LLM) reading this codebase from cold.

What alpi is

alpi is a local-first personal AI agent. It has a Textual TUI in the terminal, a Tauri desktop app and an Expo mobile app that talk to the daemon over the host plane (Unix socket locally, WebSocket remotely), an on-demand email tool (IMAP / Gmail) the agent calls to read and send mail, inline-learning memory, scanner-gated live skills, multi-provider LLM support via LiteLLM, read-only research, write-capable delegation, scheduling, MCP integration, and ALP for private agent-to-agent links.

The architectural constraint is sovereignty: state is local, identities are per-profile, network trust is explicit, and operational surfaces stay small enough to audit. The product is intentionally not a generic agent suite, marketplace, or hosted router.

Principles

alpi is published by Satoshi Ltd. and inherits the company's six operating principles (Privacy by Design, User Sovereignty, Security First, Open Source, Zero Knowledge, Digital Sovereignty). See Why alpi exists in README.md for the mapping between principle and code. The conventions below are the engineering expression of those principles — not separate from them.

Code conventions

Contributor rules — no human-facing comments inside alpi/, English only, every comment must carry a "why" — live in AGENTS.md at the repo root.

CLI surface

Stable verbs shared across groups so a user doesn't relearn per feature.

alpi                           launch the TUI
alpi -c / --continue           resume the last session in the TUI
alpi -p <name>                 profile flag, combinable with any command

alpi chat                      alias for `alpi`
alpi chat --once "<text>"      one-shot turn to stdout (pipe-friendly)
alpi chat --once ... -c | --session <id>   continue the last / a specific session (one-shot)
alpi chat --once ... --connection-id <id>  local-only delegated turn visible live to that connection
alpi chat --once ... --emit-events     INTERNAL — scheduler subprocess contract
alpi chat --once ... --no-save         INTERNAL — do not write a session file

alpi setup                     interactive menu: model / email / voice / MCPs /
                               peers / workgroups / sandbox / service /
                               health check / cleanup /
                               delete profile (non-default only)

alpi doctor                    live health check (IMAP login,
                               Gmail token refresh, MCP handshake, service PID);
                               exits 1 on any failure, 0 otherwise

alpi audit                     whole-install security posture scan
alpi audit-log                 bounded administrative activity by device

alpi logs                      tail every subsystem log merged by timestamp
  --source {service|schedule|agent|approval}  restrict to one subsystem
  -n N                                         last N lines (default 100)
  -f                                           follow mode (poll every 1s)

alpi profile list              list profiles, mark the active one with its object glyph
alpi profile show [name]       draw the profile's object as half-block art with its model and path
alpi profile create <name>     bootstrap a new profile tree
alpi profile remove <name>     delete after safety checks + confirm

alpi daemon install|uninstall                  register / unregister the service unit
alpi daemon start|stop|restart|status          lifecycle of the single per-machine daemon
alpi schedule run-once|fire <id>               manual cron tick / ad-hoc job fire

alpi peers list                list pinned ALP peers for this profile
alpi peers key                 print this profile's ALP public key
alpi peers add <id> <pubkey>   pin a peer (prefer the wizard for capability selection)
alpi peers remove <id>         unpin a peer
alpi peers tools <id> [--allow …|--clear]  show or set the only tools an inbound turn from that peer may run
alpi peers ping <id>           live probe via link.ping

alpi workgroup list                                list workgroups (hub-of + member-of)
alpi workgroup show <wg_id>                        detail + decrypted transcript
alpi workgroup create <name> --member <id>         hub-side create, granting the invited peers
alpi workgroup join <hub_peer_id> <wg_id>          subscribe to a peer-hosted workgroup
alpi workgroup post <wg_id> <text>                 encrypt + post, declaring the turn's cost
alpi workgroup pull <wg_id>                        fetch new posts and decrypt; cursor advances
alpi workgroup pause|resume|leave <wg_id>          membership ops
alpi workgroup kick <wg_id> <member-id|pubkey>     hub-only; rotates the group key

Shape rules: containers (profile, peers, workgroups) get list/create/remove (or add/remove). The daemon gets start/stop/restart/status/install/uninstall under alpi daemon; the same lifecycle is also reachable from alpi setup → Services → Daemon (default profile only) so users have one canonical place. The first alpi setup auto-installs the daemon — no opt-in step. Scheduler, ALP, workgroups, and host are fixed daemon capabilities; their useful controls live with jobs, peers/workgroups, connections, and network settings. Interactive wizards live exclusively under alpi setup; never add a per-feature wizard command.

Command ordering in --help is frequency-first, not alphabetical: chat → setup → doctor → logs → profile → peers → workgroup → schedule → daemon. See _OrderedGroup in cli.py.

alpi/ui.py is the shared interactive layer. Raw questionary.* is forbidden outside it. Helpers: banner, menu, text, password, confirm, row, ok/fail/warn/dim/saved/cancelled. The close item is added automatically with value None (callers treat None as "out").

Menu close wording: top-level (alpi setup) → Exit. Sub-menus (Email:, MCP servers:, Manage saved keys) → ← Back. Wizard aborted mid-flow → cancelled. Mixing Exit/Back/Cancel in one context is a bug.

File layout

alpi/
├── __init__.py             __version__
├── cli.py                  entry point, --continue, --profile resolution
├── engine.py               turn runner, interrupt flag, tool loop
├── llm.py                  litellm stream() / complete() wrappers
├── session.py              Turn / ToolLog dataclasses, save/load
├── prefix_diag.py          hashed conversation affinity + request-shape diagnostics
├── memory.py               MemoryStore (3 files, two-tier dedup, .bak)
├── home.py                 profile path resolution
├── config.py               YAML load/save, defaults, deep merge
├── ui.py                   shared wizard/menu primitives
├── service.py              the one daemon: per-profile tasks, launchd / systemd unit
├── ledger.py               daily spend ledger + the profile cap gate
├── outputs.py              persistent inbox JSONL store (notify + schedule failures)
├── status.py               canonical /status rows (TUI + apps share this)
├── prompts/
│   ├── default_agent.md
│   └── system_prompt.md
├── providers/              metadata for the model picker
│   └── {anthropic,openai,google,groq,openrouter,ollama}.py + model catalogues
├── tools/
│   ├── base.py             Tool ABC + ToolResult
│   ├── _state.py           per-turn emit / interrupt / usage, isolated per thread
│   ├── _paths.py           resolve_path + sensitive-path denylist
│   ├── _guards.py          terminal denylist, SSRF, prompt-injection scan
│   ├── _budget.py          per-result char cap for LLM context (100K default, per-tool override)
│   ├── _osv.py             OSV malware query for PyPI/npm names before skill/MCP install
│   ├── _sandbox.py         OS-level sandbox wrapper (opt-in)
│   ├── skill.py            the skill tool: mutations, scanner, quota
│   ├── search.py           content + filename search (rg + stdlib fallback)
│   ├── research.py         read-only sub-agent (depth: fast/normal/deep)
│   ├── terminal.py         run/background/status/output/kill
│   ├── workflow.py         bounded tool DAGs routed through ToolExecutor
│   ├── notify.py           native push to the owner's apps
│   └── … (read_file, write_file, edit_file, delete_file, todo, web_*, schedule,
│         memory, session_search, email, config)
├── tui/                    Textual app, widgets, screens, theme
├── scheduler/              cron + once jobs, hosted by the alpi daemon
├── mail/                   multi-account email: accounts, IMAP+SMTP, Gmail + OAuth
├── mcp/                    MCP client (stdio JSON-RPC) + registry
├── alp/                    Alpi Link Protocol (spec: docs/ALP.md)
│   ├── keys.py            Ed25519 identity at {home}/alp/secrets/alp_key.{pem,pub}
│   ├── envelope.py        build/sign/verify JSON-RPC envelope + replay cache
│   ├── peers.py           {home}/alp/peers.yaml load/save + capability check
│   ├── server.py          Unix-socket listener, fail-closed dispatch
│   ├── client.py          one-shot call with typed errors (TargetOffline, RemoteError)
│   ├── handlers.py        link.ask / link.cancel — engine integration
│   ├── mention.py         @peer parser + executor (shared by TUI + host chat)
│   ├── pending.py         pending invites store (unpinned-sender capture)
│   └── setup.py           `alpi setup → Peers` wizard
├── host/                   control plane for desktop / mobile clients (default profile only)
│   ├── server.py          Unix-socket JSON-RPC server (no envelope, no Noise — fs perms = trust)
│   ├── handlers.py        read verbs (host.workgroup.transcript, host.sessions.*)
│   ├── chat.py            host.chat.send/delegate (streaming) + host.chat.cancel
│   ├── runs.py            host.runs.list + host.run.{read,cancel}
│   ├── config.py          config mutation verbs (providers, peers, mcp, email, …)
│   ├── connections.py     connection identities, device credentials, migration
│   ├── connection_context.py request-scoped connection/device attribution
│   ├── admin_audit.py     bounded audit trail for administrative mutations
│   ├── attachments_rpc.py stage uploads in, serve produced files out (scoped)
│   ├── network_rpc.py     bind status and the ordered WS/WSS pairing routes
│   ├── probes.py          host.email.probe, host.peers.ping, host.model.ctx_window
│   ├── schedule.py        host.schedule.{list,remove,set_paused,fire}
│   ├── outputs.py         host.outputs.{list,read,mark_read,mark_unread,mark_all_read,delete}
│   ├── daemon.py          host.daemon.{restart,update}
│   ├── device_state.py    device-facing profile state for the apps
│   ├── events.py          host.events.subscribe + thread-safe emit() for daemon-pushed updates
│   ├── workgroup.py       transcript decryption (hub + member shapes)
│   └── sessions.py        plaintext session list / read
└── knowledge/              `alpi_knowledge` answer packs (hand-maintained)

Execution spine

Every engine turn creates one immutable RunContext, one ToolExecutor, and one ExecutionWorld. Context variables bind those objects across nested tool calls without adding parameters to every tool. The executor is the only registry dispatch seam: direct model calls and workflow steps therefore use the same denylist, member restrictions, availability checks, execution world, and durable journal.

Tools are exclusive by default. Only classes that explicitly declare parallel_safe can overlap, and the engine parallelizes a model batch only when every call is safe. A mixed batch remains serial. Results and emitted states are replayed in original call order, preserving provider transcript determinism.

A call runs only with the arguments the model actually sent. The engine, delegate, research, the workgroup wrap-up handoff and the memory reviewer decode them through alpi/tools/_args.py: empty or null is {}, raw control characters inside strings are accepted, and one JSON-encoded layer around an object is unwrapped. Anything else never reaches the tool and answers one of:

Only the last call of a reply cut at the output-token limit adds that the limit cut it off. A batch holding a refused call runs serially. The assistant message sent back to the provider carries the arguments as dispatched: the raw string when it is already a strict JSON object without repeated keys equal to them, canonical ASCII JSON otherwise, {} for a refused call, so a provider that parses history never sees a broken payload.

Right after decoding, a top-level argument sent as a JSON string is decoded when its schema admits objects or arrays and not strings (type, a type list or anyOf/oneOf branches), so a stringified knowledge object or workflow steps list reaches the tool, the events and the journal as the structure the schema declares. A string whose JSON would be refused above stays a string. workflow applies the same rule to each step's own arguments before ${step.output} references expand; an expanded value stays a string. A parameter that admits strings, or one without a readable type, passes through untouched.

Each turn writes runs/<run_id>.jsonl with bounded, redacted events (an unpaired surrogate in any persisted text, journal or saved session, becomes U+FFFD): the start record (pid, model, input), tool starts / states / ends, model_state when it changes, usage, every assistant_done (only the one closing the turn carries final=True), errors and the finish outcome. Streaming deltas are never journaled — the reconnect replay is the sessions sidecar — so alpi runs show and host.run.read return an operational timeline, not the stream. Terminal command text is omitted everywhere, including workflow steps; steps or step arguments that are not structured are dropped from the record. Local operators use alpi runs list|show|cancel or /runs; paired clients use host.runs.list, host.run.read, host.run.cancel, connection-scoped like sessions.

Cleanup offers completed journals older than 30 days plus the oldest beyond 200 MiB per profile, skipping anything completed within the last hour; nothing is deleted on its own. A run sweep runs at profile start and every 30 s from the daemon's maintenance loop, off the scheduler's thread pool and wrapped so a failure in it never ends the loop. One scan of runs/ feeds two rules in order. reconcile_stale closes journals whose pid is gone as interrupted (reason dead), so a dead child is reported within about a minute. reconcile_silent judges a scheduled run that has written nothing for longer than its timeout plus SILENCE_GRACE_S: the timeout_s the scheduler recorded in run.started (from ALPI_RUN_TIMEOUT_S), else the job's current timeout, else MAX_RUN_TIMEOUT_SECONDS for a deleted job or an invalid value; silence is wall-clock, but the sweep must also have watched it hold for the grace on the monotonic clock before acting, so a wedged child is reported roughly ten minutes past its timeout and a clock step alone never fires it. A live pid is killed only when it is provably this run's process — run.started records pid_start from /proc/<pid>/stat, the sweep requires it to match, and never targets its own pid or its parent's. Anything alive it cannot vouch for (a recycled pid, a pre-0.14.50 journal, a host without /proc) is reported once and left untouched, journal included: a run.finished written under a live writer would be followed by that writer's own records and the run would read as running again. Runs without a job are never judged by silence. The scheduler closes the journal of a child it ends itself (it hands the child ALPI_RUN_ID), so its own timeout yields one alert rather than one from the scheduler and one from the sweep. Each reported row files an error output and raises schedule.failed, the path a live failure already uses.

ExecutionWorld keeps filesystem resolution and terminal shell execution under one run-scoped abstraction. local preserves the previous behavior. docker wraps processes in an ephemeral container while bind-mounting the same absolute workspace/profile paths, so file tools and subprocesses observe one namespace. Foreground containers are force-removed on timeout; background terminal jobs are refused so a detached container cannot outlive its run. Dedicated workers such as skill scripts and speech transcription remain host-side; this backend is not a whole-agent filesystem sandbox.

The workflow tool executes a bounded dependency graph of registered tools. References such as ${step.output} feed prior results into later arguments; independent safe steps can overlap. Recursion is refused, failures stop the graph unless explicitly marked continue_on_error, and every nested call re-enters ToolExecutor rather than bypassing policy. Nested parallelism uses the same tools.max_parallel_tool_calls limit as direct model batches.

Runtime state (skills, sessions, memories, logs, ALP peers, keys) does not ship with the package — it's generated per profile under ~/.alpi/. The alpi/knowledge/references/ directory holds the answer packs the alpi_knowledge tool serves; there is no bundled skill namespace. See Profile home layout immediately below. The skill tool (alpi/tools/skill.py) manages user-created skills that live at {home}/skills/<category>/<name>/.

Profile home layout (~/.alpi/ or ~/.alpi/profiles/<name>/)

~/.alpi/                     default profile root
├── .env                    API keys, IMAP/SMTP credentials, allowlists
├── config.yaml             model + tools + tui + mcp
├── memories/               USER.md, MEMORY.md, AGENT.md (+ .bak)
├── skills/<category>/<name>/    SKILL.md + scripts/ + references/ +
│                                 assets/ + secrets/ (0700) + state/ +
│                                 .gitignore
├── recipes/<id>.yaml        saved workgroup recipes owned by this hub profile
├── sessions/<id>.json      compact turn-based session log (TUI / desktop / `--once`)
├── knowledge.sqlite        sqlite-vec derived indexes for knowledge,
│                            session recall, and workgroup recall
├── mentions/<sender>@<conversation>.json  @-mention threads per sender and originating conversation (cap 20 turns), receiving side; <sender>.json for callers that send no conversation
├── run/                    background process registry, schedule pids
├── alp/                    ALP state — keypair, peer list, socket, pid
│   ├── peers.yaml         pinned peers (pubkey + allow + optional address)
│   ├── alp.sock           Unix-domain socket, 0600, only while listener runs
│   ├── alp.pid            listener pid
│   └── secrets/alp_key.{pem,pub}   Ed25519 identity (private 0600, public 0644)
├── host/                   control-plane state (default profile only)
│   └── host.sock          Unix socket the local desktop connects to (mobile uses the WebSocket)
├── outputs/                persistent inbox for proactive agent messages + schedule failures
│   └── outputs.jsonl       JSONL store (≤500 rows, atomic compaction)
└── logs/                   service.log (daemon-wide; lives only at the root, NOT
                            duplicated per profile), agent.log + approval.log
                            (per profile — only the default profile's pair is at
                            this level), ledger.json, compaction.jsonl, runs.jsonl

~/.alpi/profiles/<name>/     same layout MINUS service.log; agent.log + approval.log
                             are emitted under each profile's own logs/

Core systems

Engine loop (alpi/engine.py)

Per turn: append user message → loop {LLM stream → emit deltas → exec tool calls → append tool results} until the LLM stops emitting tool calls OR max_steps_per_turn is hit. The configured value (default 100 model iterations) is literal for every provider and budget. Hitting it does not drop gathered work: normal chats get one tools-off best-effort wrap-up; detached workgroup turns get one workgroup_post-only handoff. interrupt_requested is polled at three checkpoints (between iterations, mid-stream, between tool calls). A turn lock serializes concurrent runs so a delayed research tool from the previous turn can't bleed into the next.

Events emitted to the UI sink: user, reasoning_delta, assistant_delta, assistant_done, tool_start, tool_state, tool_end, usage, error, done, interrupted. The TUI consumes them; the scheduler subprocess consumes a subset via JSON-lines.

A model-provider failure reaches every surface as one plain sentence, never as the provider's exception text. alpi/llm_errors.py classifies the exception chain (litellm class, HTTP status, text) into context_length, insufficient_credit, auth, forbidden, model_unavailable, content_policy, rate_limited, timeout, provider_unavailable, bad_request or unknown. The error event carries text (the sentence), code (the class) and detail (the redacted, 500-character technical text, also logged and kept in the run journal). host.chat.send frames carry text and code; detail reaches the TUI, --once ([error] text (detail), or the detail key of the JSON event) and the scheduler, whose failed-job message reads agent error: text (detail). ALP peer replies carry only the sentence. Chat-level codes such as busy share the code field.

The apps put the daemon's own protocol errors in plain words through one shared map, common/plainError.mjs (fixtures in common/plainError.fixtures.mjs, exercised by both clients): too-many-connections, too-many-requests, forbidden, method-not-found and the close reasons Authorization changed, Device authorization revoked, WebSocket capacity reached, plus the transport failures, become a sentence; anything else is shown unchanged. Both clients retry a call or a stream refused with -32029 too-many-connections once after 1.5 s before showing it.

Host chat forwards usage with context_tokens; it is non-zero only for the main completion whose input becomes Session.last_ctx_tokens. Side-model usage still carries accounting fields but cannot move a client's conversation meter.

Cross-turn resume. A chat is not a long-lived object: each turn spins up a fresh Engine and rehydrates the session from disk (_hydrate_from_path in cli.py, shared by TUI --continue and the host chat; the desktop "edit message" rewrite path mirrors it in host/chat.py). The model context is rebuilt from the prior replayable turns — those that ended in a final reply or produced a file; a turn aborted before its reply (no assistant text, no output files) is dropped, so a resumed session never re-answers a dangling request. Each replayed turn contributes its user text (plus an input-attachment marker [attached: name (mime)]) and assistant text (plus a produced-file marker [produced this turn — reuse the absolute path…: name → /abs/path]). Tool calls and tool results are deliberately not replayed — they would blow the context budget — so an agent does not remember what it searched, read, or analyzed last turn, only its final reply and the absolute paths of the files it produced. A multi-turn edit ("now relight it at sunset") reuses the produced path surfaced by the marker, not a remembered tool output; an agent that needs an earlier tool's result across turns must re-run the tool or rely on a produced file.

The system prompt for each turn is assembled in a fixed order (PART_ORDER in alpi/prompt_cache.py): AGENT.md (agent profile — voice, style, identity) → base prompt → environment block (workspace, profile home, path rule) → system time → platform hint (per-surface guidance when ALPI_PLATFORM is set: cron; empty for TUI and the apps) → turn guidance → the self-knowledge rule pointing the model at alpi_knowledge (dropped when that tool is denied) → skills index → USER.md → MEMORY.md. Providers that need an explicit marker (Anthropic, including through OpenRouter) get LiteLLM's cache_control_injection_points on messages[0] from cache_kwargs_for_model; an openai-prefixed model at OpenAI's own endpoint never does (the host is judged as LiteLLM judges it, from api_base, litellm.api_base or OPENAI_BASE_URL), because LiteLLM would turn the marker into OpenAI's explicit mode, which switches off its automatic prefix cache (and the /v1/responses bridge drops the breakpoint anyway), so an OpenAI model reads its cache with no marker at all; a gateway behind the same prefix keeps it.

The scheduler (alpi/scheduler/run.py) sets ALPI_PLATFORM=cron so scheduled jobs run knowing no user is present and they cannot ask for clarification. That overwrite would otherwise erase the deployment runtime, so it also carries ALPI_DEPLOY_RUNTIME, which alpi/runtime.py prefers: ALPI_PLATFORM is the turn origin, platform_id() is the runtime, and a job inside a container still reads as Docker. Each fire runs as a subprocess capped at job_run_timeout(job) seconds — job.timeout if set, else DEFAULT_RUN_TIMEOUT_SECONDS (900). parse_run_timeout is the one duration contract for schedule(add|update), execution, schedule list / host.schedule.list (run_timeout, or null plus timeout_error) and the silence watchdog: a whole number of seconds in [30, 86400]; booleans, fractions, non-finite and out-of-range values are refused, and a bad stored value fails the fire with invalid stored timeout instead of being clamped. Fires of one profile run serially, so a long timeout also delays that profile's other due jobs. The cap is a stuck-process backstop for unattended runs, not the cost guard (budget.daily_usd is) and not a hint that jobs must be short; heavy jobs (deep research, multi-step publishing) opt into a longer budget via schedule(add|update, timeout=…). The scheduler passes the child a soft budget via ALPI_TURN_BUDGET_S (the cap minus a ~10% reserve, floor 60s); when the engine crosses it, normal jobs get one tools-off best-effort reply and detached workgroup turns get one workgroup_post-only handoff. The hard subprocess timeout remains the last-resort kill if finalization itself stalls.

Cron jobs with no_agent: true skip the LLM entirely. The prompt is shlex-tokenized and exec'd directly (shell=False); ${ALPI_HOME} expands to the profile home and the profile's .env overrides inherited env keys so skills find their declared requires_env. A form-based allowlist enforces that the command is python[3] [flags] <script> or <script> invoked directly, where <script> resolves to <home>/skills/<category>/<name>/scripts/…; non-python executables and -c/-m inline-code flags are rejected at both schedule(add) time and inside the scheduler before exec. Use this for deterministic skills (sync, file processors) — saves both tokens and the agent boot latency per fire.

LLM transport (alpi/llm.py)

Thin wrapper over litellm.completion. stream() is a generator yielding {text_delta, reasoning_delta, tool_calls_delta, finish_reason} per chunk plus a final {final, tool_calls, input_tokens, output_tokens, cost_usd, finish_reason} whose finish_reason is the last one the provider reported (litellm reports stop when a stream closes early, so only length proves a cut); the stream end breadcrumb logs it. complete() is the non-streaming variant (research, delegate and the background passes); its Completion.finish_reason carries the same signal. _silence_litellm() runs at import time to mute LiteLLM's startup banner via FD-level redirect (Textual is sensitive to stdout pollution).

Memory (alpi/memory.py)

Three files: USER.md (facts about the user), MEMORY.md (env quirks, commands, incidents), AGENT.md (the agent's own profile — tone, style, identity, language). § entry delimiter, char limits USER_CHAR_LIMIT = 3000 / MEMORY_CHAR_LIMIT = 5000 (see alpi/memory.py; AGENT.md is free-form prose with no cap). Accent+case+punctuation-insensitive dedup, plus token-Jaccard dedup at 70% max-containment to catch paraphrases. .bak snapshot before every mutating write. Approach C: every mutating call returns the full current state of the target file so the agent sees its own write in the same turn.

v2 quality metadata. Each entry carries a trailing <!-- alpi-meta conf=... captured=... reinforced=... --> comment that is stripped before the entry reaches the system prompt. conf is low / normal / high (default normal). Near-duplicate writes reinforce the existing entry (bump reinforced, upgrade low → normal at ≥ 2) instead of appending a paraphrase. Low-confidence entries with zero reinforcements expire after LOW_CONFIDENCE_MAX_AGE_DAYS = 30 (constant in alpi/memory.py; keep it fixed unless operational traces justify tuning). The memory tool's safety scanner reuses the skill scanner patterns and adds invisible/bidi unicode detection (U+200B–200F, 202A–202E, 2060, 2066–2069, FEFF) to block Trojan-Source vectors; _operational_warning surfaces non-blocking warnings when a write looks like session state (chat_id, session_id, ISO timestamps).

Batch writes. memory(action="add", entries=[...]) accepts a list of entries for the same target in a single call. Each entry runs through cross-file and same-target dup checks independently; entries that collide are skipped with a per-line note, the rest land in one write. Replaces the pathological pattern of one add call per fact (16 calls in a single turn observed in real sessions).

Post-turn reviewer. When memory.review_interval > 0 (default 0 = off), alpi/review.py spawns a daemon thread after each turn that snapshots the user/assistant text and asks the LLM whether anything durable should be added. The reviewer is constrained to memory(action="add", ...) — never replace/remove — to prevent it from deleting unrelated entries on a bad pass.

Promotion queue (alpi/promotion.py). Auto-compaction never writes to USER.md / MEMORY.md / AGENT.md directly. After every fired compaction the engine runs a second short LLM call against the summary (system prompt CANDIDATE_PROMPT) and pushes any durable facts as candidates into <home>/memories/promotion_queue.jsonl. On enqueue, each candidate is annotated with the same preview warnings the memory tool computes at write time — operational-state heuristic, cross-file duplicate check, safety scan. The queue is bounded (MAX_PENDING = 200 per profile) and pending entries expire after MAX_AGE_DAYS = 30. Per-record fields in the JSONL: id (8-char hex), created_at (unix ts), source (compaction | reviewer | manual), session_id, model, target (USER.md | MEMORY.md | AGENT.md), text, confidence (low | normal | high), warnings (list of strings).

Two memory tool actions surface the queue, both safe for the agent to call: promotion_list (read-only) and promotion_discard(id) (drops a candidate without writing). There is no agent-callable apply. The only path that writes to durable memory from the queue is the CLI alpi memory promote — interactive review with [a]pply / [d]iscard / [s]kip / [q]uit per candidate, plus --apply-all / --discard-all for unattended sweeps. This keeps the human-in-the-loop gate genuine: the agent cannot promote facts on its own regardless of how the prompt is framed. If the underlying memory add fails (safety scan, duplicate), the candidate stays in the queue so the operator can fix and retry.

Path resolution (alpi/tools/_paths.py)

Single entry point resolve_path(path):

  1. expanduser().
  2. Relative paths root at the active workspace (cfg.workspace or cwd fallback).
  3. Resolve symlinks.
  4. Reject if the resulting path matches any sensitive-path entry (denylist below) — ValueError.

Denylist: /etc/, /boot/, /sys/, /proc/, /usr/lib/systemd/, /System/, /private/etc/, the docker sockets, ~/.ssh/id_*, ~/.ssh/authorized_keys, *_key, *_ed25519, *.pem/.p12/.pfx, ~/.aws/{credentials,config}, ~/.gnupg/, ~/.netrc, ~/.npmrc, ~/.pypirc, ~/.pgpass, ~/.config/{gh,gcloud}/, shell rc/login files (.bashrc/.zshrc/.zprofile/…), ~/Library/Launch{Agents,Daemons}/, profile .env/config.yaml, and skill secrets/ dirs. Both pre-resolve and post-resolve forms are checked (macOS /var → /private/var symlink case).

suggest_similar_paths(target) lists the parent directory and fuzzy-matches siblings by basename substring/prefix. Used by read_file, edit_file, and search to turn dead-end errors into actionable suggestions.

alpi/tools/_lint.py::lint_content(path, content) runs a parser-based syntax check before every write_file / edit_file lands on disk. Parsers by suffix: .py → ast.parse (stdlib), .json → json.loads (stdlib), .yaml/.yml → yaml.safe_load_all (PyYAML, already a dep; every --- document is parsed, so Spring and Kubernetes multi-document files pass and a broken later document is reported at its line in the file), .toml → tomllib.loads (stdlib on every supported Python version). Other suffixes pass through. Failures return a one-line error with the source line/col and the write is refused — the original file (if any) is untouched. Catches the class of bug where a malformed jobs.json, config.yaml, or skill script silently breaks a downstream consumer.

alpi/secrets_io.py::safe_write_secret(path, content, mode=0o600) is the canonical write path for any credential file. It uses tempfile.mkstemp (O_EXCL + 0o600 at creation, random unique name in the target dir), then os.replace onto the target — no TOCTOU window where the file exists at umask perms, and a stale <target>.tmp lingering at looser perms cannot compromise the write because the helper picks a fresh random name. Used by model_selector._atomic_write_env (.env writes), mail/gmail_auth._save (gmail token), alp/pending.save (pending-peers yaml), and alp/keys.create (ALP private key).

Tool registry (alpi/tools/__init__.py)

register(cls) adds a Tool subclass to the dict, schemas() emits the OpenAI function-calling shape, execute(name, args) runs by name with full error capture. The registry is assembled from the sibling tool modules in alpi/tools/__init__.py, including the Playwright-backed browser tool. knowledge registers first so durable user/workspace recall has one canonical surface.

A tool's check() is what keeps schemas() honest: an unavailable tool is never offered to the model, so the probe has to test the thing that actually fails at call time, not a proxy for it. browser is the cautionary case — probing only import playwright advertised the tool on every slim Linux image, where playwright downloads the browser on demand and then cannot launch it because the distro never installed Chromium's load-time libraries. It now dlopens one soname per Debian package family (_CHROMIUM_SONAMES, Linux-only, ~2 ms when they are absent) and reports the missing set plus CHROMIUM_DEPS_COMMAND — a uvx --from playwright … invocation, because uv tool install links only alpi-agent's own entry points and a bare playwright is not on PATH.

The image cannot be held to that list by a unit test: pip install . ignores uv.lock, so the playwright inside the image outruns the one the suite imports (observed: 1.62 vs 1.58). So docker/Dockerfile derives the packages from its own playwright (playwright install-deps chromium-headless-shell) instead of carrying a hand-written copy, and publish-docker.yml launches the real headless shell in the built image before anything is pushed. ensure_chromium() installs with --only-shell because chromium.launch(headless=True) runs chrome-headless-shell; the full Chromium build was never used, and _wanted_chromium_dirs() now tracks only the shell so an existing profile's stale full build is pruned (~640 MB reclaimed per profile).

Knowledge recall (alpi/core/ + alpi/tools/knowledge_base.py)

Per-profile semantic search over synthesized user/workspace knowledge. The source of truth is Markdown under <workspace>/knowledge/; SQLite under <home>/knowledge.sqlite is only a rebuildable derived index. Raw source files and attachments are read only as inputs for synthesis; alpi does not copy them into a durable documents store.

Supported ingest formats: markdown / text / source / configs (stdlib read), HTML (html2text), PDF (pypdf for text-layer, RapidOCR fallback when ocr=true and pypdf extracts < 50 chars), DOCX (python-docx), EPUB (ebooklib), images (PIL + RapidOCR — only with ocr=true). OCR backend is rapidocr-onnxruntime (ONNX port of PaddleOCR, no torch dependency). The PDF/image/OCR extraction primitives live in alpi/extract.py and are shared verbatim with the chat-attachment path (alpi/attachments.py); the DOCX/EPUB/ HTML readers and the chunker live in alpi/tools/workspace.py, a support library, not an agent-facing recall surface.

Shared store primitive (alpi/core/store.py). open_store(home) returns a sqlite3.Connection with the sqlite-vec extension loaded. Designed to host other shapes later (workgroup search, future entity memory) — they bring their own table schemas.

Embedder (alpi/core/embed.py). Embedder Protocol; default FastembedEmbedder wraps the ONNX export of sentence-transformers/all-MiniLM-L6-v2 (384-dim, ~90 MB, no torch). Numerically equivalent to the original sentence-transformers checkpoint but ~10× lighter at runtime. Lazy-loaded under a threading.Lock so concurrent first-touch calls serialize on a single model instance instead of racing.

The bundle uses minimal YAML frontmatter, relative Markdown links, and required index.md / log.md. It is never auto-injected into the system prompt; access happens only through knowledge tool output.

Links resolve relative to the page holding them, the way GitHub, Obsidian and VS Code resolve them, and that is the only form alpi writes. The link graph reads more than it writes: bare, angle-bracketed and percent-encoded destinations, [[wikilinks]] (by path, or by page name when it is unambiguous) and CommonMark reference links, while links inside fenced blocks or code spans are examples, not edges. maintain also repoints a proposed link that is written from the bundle root onto the page it plainly means. Two rules are alpi's own rather than Markdown's, and are the usual reasons an imported vault fails lint: every page needs an inbound link, and the type frontmatter key collides with Hugo's reserved layout key if the same tree is published with Hugo.

Session recall (alpi/tools/recall.py)

Recall over past conversations, the conversational-memory peer of knowledge recall, in three layers: lexical find (session_search, term counts over sessions/*.json), exact browse (session_read, no model call), and opt-in semantic search (index_sessions / recall_sessions) for fuzzy "when did we discuss X / what did we decide about Y".

Forgettable. Recall is a derived view, so forgetting is real: deleting a session (host.sessions.delete → host/sessions.py::delete_session) purges its rows via recall.forget_session, and index_sessions orphan-sweeps any tracked session whose file is gone. No auto per-turn injection — retrieval is explicit, like the workspace tools.

Workgroup transcript search (alpi/tools/workgroup_search.py)

The third retrieval surface on the same store: semantic search over hub-owned workgroup transcripts. Workgroups are hub-owned by design, so this is profile-local and hub-only — the hub decrypts its own transcript and indexes it; there is no cross-peer / federated search and no global "search all my peers' workgroups". Two tools:

Forgettable. Removing a workgroup purges its index in both delete paths — the host RPC (host/workgroup_admin.py::_remove) and the CLI (alpi workgroup remove) call workgroup_search.forget_workgroup; index_workgroups orphan-sweeps any tracked workgroup whose directory is gone. No auto-injection into workgroup turns. ALP encryption/transcript behaviour is untouched — this only reads through the existing decrypt path.

Removal tombstones (alp/secrets/subscriptions.removed.d/). Removing a workgroup writes an empty marker named by its id in every local home, so a stale in-memory copy or a hub-side auto-join heal cannot resurrect it. Markers never cross machines and expire after TOMBSTONES_KEEP_DAYS (2) once the id is gone from subscriptions.yaml and alp/workgroups/; setup → Cleanup offers the expired set under Workgroup tombstones.

Asset prefetch (service.py::_prefetch_assets). Scheduled at boot + 600 s, past the client-reconnection rush. Gated by runtime.prefetch on the root profile: auto (default) fetches the embedding weights only when some profile has knowledge.sqlite and Chromium only when some profile leaves browser un-denied; all forces both; off — the default in Docker — skips it. Every asset still loads lazily on first use, so off costs latency, never functionality. A successful Chromium install prunes stale builds of the headless shell.

Skills

Live under <home>/skills/<category>/<name>/. Required SKILL.md plus optional scripts/, references/, assets/, secrets/ (mode 0700, gitignored, scanner skipped), state/ (gitignored, scanner skipped, runtime persistence). .gitignore auto-written on create with secrets/\nstate/\n.

Live by default — there is no pending-approval stage; the scanner and the sandbox are the guarantees.

Frontmatter (auto-populated on create): name, description, category, version, origin: agent|user, created_at, requires_env, tools, keywords, optional output_schema. 13 fixed categories including miscellaneous as the fallback. secrets/ is filesystem state, not frontmatter: it is created lazily when a skill writes a secret file. output_schema is one-line JSON and uses a deliberately small subset (type, properties, required, items, enum) so the runtime stays dependency-light.

Security scanner (~50 patterns, _DANGER_PATTERNS in alpi/scan.py — the shared scanner library used by skills, memory writes, and the recalled-memory guard): destructive shell, credential exfiltration, prompt injection, persistence (cron/launchd/systemd/authorized_keys/sudoers/shell rc), reverse shells, tunneling, obfuscation (base64/eval/exec/compile), process exec, hardcoded credentials (API keys, OpenAI sk-, GitHub ghp_, AWS AKIA), system-password-file paths, deep traversal. Runs on every create/add_file/patch for files NOT in secrets/ or state/.

Atomic writes everywhere (tmp sibling + os.replace). .bak next to SKILL.md on every edit/patch. Quota: max 40 agent-owned skills, enforced at create.

Auto-injected into the system prompt (skills_index_block(home)): every session start, all installed skills are listed by category as name: description entries, prefixed by a directive that says "check this list before reaching for general tools". Without this nudge, mimo-class models routinely went straight to web_search/terminal even when a perfect skill existed.

TUI integration: when a terminal command's path matches .alpi/(profiles/<p>/)?skills/<cat>/<name>/..., arg_hint rewrites the ToolCard label as skill: <name> (or skill: <name> · <script> when the script is the full path). Tool name stays terminal; the rewrite is display-only.

Execution: skill(action="run", name=...). Single canonical ad-hoc path. If scripts/run.py exists the action validates the skill, then spawns the script via subprocess.run with cwd = skill dir, env += {ALPI_HOME, ALPI_SKILL_NAME, ALPI_SKILL_DIR}, 600s timeout, and the skill's requires_env checked up-front. If the skill declares output_schema, stdout must be JSON and is validated before the call succeeds. Scripts are normal Python; built-in tools and MCP methods are not importable Python APIs. No script → SKILL.md is returned with a [skill X has no scripts/run.py — follow these instructions] prefix so the agent follows the prose and calls the real tools. Scheduled prompts should call this action instead of reimplementing the skill by hand; the scheduler still enters through alpi chat --once --emit-events --no-save.

Structured composition: skill(action="invoke", name=...). Same subprocess/runtime path as run, but stricter: the callee must ship scripts/run.py, must declare output_schema, and stdout must satisfy it. This keeps skill-to-skill composition machine-readable and prevents prose-only skills from pretending to be callable subroutines.

Scripted harness: skill(action="test", name=...). Thin validation layer over the same runtime path. It exists so chat/scheduler/desktop can exercise a scripted skill and verify its declared output_schema without inventing a second testing runtime. If a CLI wrapper lands later, it should call this action instead of duplicating logic.

Research (read-only sub-agent, alpi/tools/research.py)

Spawns a sub-agent with a read-only toolset (web_search, web_fetch, web_extract, read_file, search). Returns a single synthesised report; the main agent never sees the intermediate tool trace.

Depth tiers instead of a numeric max_steps: depth="fast"|"normal"|"deep". The step ceilings are product constants (DEPTH_STEPS_DEFAULTS, 8 / 15 / 30). Locks the model to three buckets (fast = single-answer, normal = comparative, deep = exhaustive); fast and deep double as the model-tier names, so a depth also picks the matching tier when the profile configures one.

Synthesis fallback: when the budget runs out, research forces one final no-tools llm.complete() with "stop investigating, report now". Avoids the "[research gave up]" footgun where the main agent retries the whole thing.

Interrupt: polls tool_state.is_interrupted() between iterations and between tools; returns [research: interrupted] on the first hit. State label during execution: <depth> · step N/M; while an inner tool runs its own emit_state label gets auto-prefixed with step N/M · … via a wrapped _emit installed for the duration of each tool-call batch (restored in a finally).

Batch mode: tasks: [{brief, depth}] up to 3 runs concurrently — see the Delegate section below for the shared ThreadPoolExecutor design (same pattern applies here).

Attachments (alpi/attachments.py)

host.chat.send accepts attachments: [{path, mime?, name?}]. The engine validates them (att.validate — magic-byte sniff for image/PDF, NUL/control-ratio guard for binary-as-text, per-type size caps, allowlist: images png/jpeg/webp, PDF, and text/source incl. py/js/ts/tsx/go/rs/sh/sql) and turns them into OpenAI content-parts (build_content_parts): images → base64 image_url data parts, text/source → inline text parts, PDFs → text extraction (bounded by tools.attachments.max_text_tokens → chars at ~4/token; default auto = half the active model's context window, no page cap). A scanned PDF (extractable text below SCANNED_PDF_TEXT_FLOOR) falls back by model capability: vision-capable → rendered page images; text-only → RapidOCR text (capped at SCAN_MAX_PAGES), so a profile with no knowledge base and no vision can still summarize a scan. PDF text/render/OCR mechanics are shared with the knowledge tool via alpi/extract.py. Images on a text-only model are not OCR'd — they degrade to a path note telling the model it can't see them. A guidance text-part tells the model the files are inline so it doesn't reflexively call filesystem or knowledge tools to "find" them.

Per-turn only. Bytes live only in the in-memory message. session_metadata is itself bytes- and path-free ({name, mime, size}), but the engine re-adds a best-effort local path to each persisted chat-turn attachment so clients can thumbnail history — the path may be unfetchable from another client (outside host.attachments.fetch roots) or after a staged file's TTL, so this is preview replay, not durable storage. The validated turn attachments ({name, path, mime}) are also published to a runtime-only ContextVar (tools/_state.set_turn_attachments) so a tool can resolve a turn's files. Remote clients (mobile, or desktop pointed at a remote daemon) can't hand the daemon a local path, so they upload bytes via the host.attachments.stage RPC (type-aware caps, content validated 1:1 with send) which writes to a TTL-swept temp dir and returns a daemon-side path. Under session_scope: device, host.attachments.fetch serves a remote member device only what it staged (a .owner marker beside the upload; uploads with no marker stay fetchable until the TTL, but an existing unreadable or malformed marker refuses access) or a path that appears in a session it owns: a turn's attachments or output_attachments, a tool's args or result, or the assistant's text, never a path the user typed (a path the user also wrote is not offered even if the agent repeats it), and only exact paths: a file named only in the assistant's text with a space or a parenthesis in its name is not offered, and one the agent lists or reads on its own becomes offered. alpi/host/offered_paths.py keeps that set per session and rebuilds it only when the session file changes. Admins, the Unix socket and session_scope: connection are not narrowed, and the roots and the secrets denylist still apply first.

Durable. knowledge(action="ingest") is the bridge from per-turn input to permanent knowledge. It reads an attachment or source file, synthesizes Markdown pages under <workspace>/knowledge/, updates index.md / log.md, and refreshes the profile-local derived index in knowledge.sqlite. The raw source is not copied into a durable documents directory. There is no auto-learn: attachments stay one-turn unless the user explicitly asks to learn/remember/save/index/compile one.

Vision (alpi/tools/read_image.py)

read_image(path, question) runs the current (or override) model in multimodal mode on an image and returns a text answer. path can be a local file OR an http(s) URL — URLs go through check_url() for SSRF (metadata hosts + private IPs blocked, redirects re-validated via httpx event_hooks).

Magic-bytes sniff accepts PNG / JPEG / GIF / WebP / BMP plus SVG (text-sniff for <svg); rejects bytes that don't match a known header even if the extension agrees. 20 MB cap on file and on download payload.

No pre-flight vision-capability check — LiteLLM's supports_vision() is wrong for openrouter/... prefixes and would bounce real vision models. If the call fails we surface the error with a hint pointing at /model when the message mentions image / vision / multimodal.

Model override via tools.read_image.model in config (surfaced as Vision model by alpi setup, desktop and mobile). When set, read_image and browser's opt-in screenshot analysis try the override first; on failure the tool retries with the main model and prefixes the answer with [fallback: <override> unavailable, used main model]. Clearing it restores main-model fallback. This route is deliberately tool-scoped: chat image attachments are multimodal parts of the main turn and are not silently moved to the override.

Same usage / cost plumbing as research and delegate. Images are auto-resized before upload (see CONFIG.md → tools.browser.vision).

Delegate (write-capable sub-agent, alpi/tools/delegate.py)

Sibling to research, but can mutate: spawn a focused sub-agent with a chosen toolset, get back a summary. Used when a task would otherwise flood the parent context (multi-file refactors, fetch+parse+write pipelines, skills that generate several output files, iterative debug loops).

Toolsets (callable presets via the toolsets param, default ["file", "web"]):

Blocked for sub-agents: delegate (no recursion), memory, skill, schedule, notify, email, session_search, session_read, todo (shared global state). research is not in any preset either — if you need deep investigation inside a delegate task today, do it in the main agent first and pass findings via context.

Budget: max_steps is a per-call tool parameter, defaulting to 30 and clamped to MAX_STEPS_CAP = 100; a non-positive or unparseable value falls back to the default. It's a ceiling, not a target — the sub-agent stops when done.

System prompt is built from a single template plus the workspace root (when set): relative paths resolve under workspace, absolute paths go where the goal says, and the sub-agent is explicitly warned not to invent /workspace/... style roots.

Prompt-cache contract. messages[0] is the stable system prefix and tool schemas are sorted by name. Volatile # NOW, workgroup, skill-hint, and relay state is composed once into the user turn's persisted host_context suffix, so normal history growth is append-only across live calls and every rehydrator. OpenRouter calls carry a hashed affinity for the logical conversation; other providers receive no OpenRouter-only fields. prefix_diag.py compares bounded request-shape hashes per conversation and records causes, never prompt text. Caching and diagnostics are best-effort and cannot fail a provider call.

Batch parallel mode. Both research and delegate accept tasks: [...] (up to 3) and run them concurrently via ThreadPoolExecutor(max_workers=3). Isolation is provided by _state.py: _emit, _interrupt_getter, _usage_sink are contextvars.ContextVar, so each worker thread sees its own values without racing on module globals. Workers re-seed interrupt_getter + usage_sink from the parent context (Python's ThreadPoolExecutor doesn't propagate ContextVars automatically) and install a per-task prefixed emit so TUI progress lines read [i/N] <tag> · <msg>. Results aggregate into one markdown report with per-task sections; per-task failures are captured inline as [failed: <error>] instead of aborting the batch. Cap is hardcoded at 3 — bumping would need a config knob and would multiply LLM cost linearly; not a default worth moving.

TUI (alpi/tui/)

Textual 8.2.x. Layout: AlpiTopBar (identity: version, profile, workspace) + chat scroll (VerticalScroll.anchor() auto-follows new content) + a bottom dock holding the completion popup, the multi-line composer (ChatInput, a TextArea) and the StatusLine (model · ctx % · cost · budget · sandbox · unread inbox · prompts waiting, then context-aware key hints).

Theme (themes.py): build_theme(accent, dark) returns a Textual Theme whose background, surfaces, ink and status colours come from alpi/palette.py, a mirror of common/tokens.mjs kept honest by a parity test. Every grey of both themes, in the console and in the apps, is an equal-channel neutral, so the only colour on any surface is a profile's. palette.resolve_accent maps an unset or legacy brand accent (#c8a24e, #8a5a0a) to the mode's token (#f3efe6 dark, #14110c light, the brand ink); a chosen colour, amber included, is kept. The console (ui.py, alpi profile list) goes through the same resolver. Registered in AlpiApp.__init__ (not on_mount — child widgets read theme_variables during their own mount).

Fold in the console (fold_art.py, fold_shapes.py): the console wears the profile's object. fold_shapes.py is generated by scripts/sync_fold_shapes.py from common/folds.mjs (--check fails when stale; a parity test runs node on the shared module for every object, honeycomb and alpaca), and fold_art.fold_tones is the OKLCH three-tone rule of foldTones, tested for the same hex on every accent and random colours. fold_art.art(fold, accent, rows) rasterises the polygons into half-block cells (▀ with the top sample as foreground and the bottom as background, ▄ when only the bottom is filled, a space when neither, so the terminal background shows through), two cells wide per row. fold_art.identity(home, tui) is the pair a profile wears: the alpaca for the default profile (in palette.BRAND_INK of the theme, the brand accent #14110c light and #f3efe6 dark, mirrored from common/folds.mjs), else tui.fold (the diamond when unset or unknown) in palette.profile_accent. The art draws beside the title of the alpi setup menu (when the terminal is wide enough and tall enough to show the whole menu beside it) and in alpi profile show. fold_art.GLYPHS gives each object one narrow, distinct, BMP glyph (East Asian Width N or Na); fold_art.marker(home, tui) is that glyph in the accent and replaces the diamond that marks the active entry in alpi profile list, the TUI list rows (list_row.set_marker) and the status line. supports_fold_art() is true only on a TTY stdout with UTF-8 encoding and locale, COLORTERM truecolor or 24bit, no NO_COLOR and TERM not dumb; anywhere else every mark falls back to the single-colour ◆ (and ◇ for an inactive profile) exactly as before. Menu cursors (ui.POINTER) and the tool-hint chip are status marks and stay diamonds.

Steps (StepsGroup / ToolCard in widgets.py): each turn's tool calls collapse into one ▸ N steps · Xs row that shows the live step while running. It expands to one card per call: a family glyph (file, terminal, globe, search, link, memory, chip — tool_hints.tool_family), the argument summary, the result hint and the duration; a card opens to the pretty-printed arguments and an output excerpt. Failed calls start open and open their group. ask_user answers render as their own line, not as steps. Ctrl+O toggles the last turn's reasoning and steps.

Assistant streaming: AssistantMessage streams into a cheap Static (flushed every 150 ms) and swaps to a Markdown widget once when the reply lands. Very long user messages render as plain text instead of Markdown.

Reasoning surface:

Persistence contract (cross-surface). The engine consolidates the whole turn's reasoning — reasoning_delta thinking + the inter-tool prose — into Turn.reasoning (str), and records Turn.reasoned_s (float) = the reasoning span from turn start to the first tool boundary, or to the first final-answer text token when there are no tools; it excludes both tool execution and final-answer streaming so the duration isn't inflated by a long-running tool or a long reply. Desktop, mobile and the TUI render a collapsible "Thought for Ns" block from Turn.reasoning, falling back to joining ToolLog.reasoning for turns logged before the field existed. ToolLog.reasoning (first tool of each batch) remains the legacy per-tool fallback.

Slash commands come from one registry (alpi/tui/commands.py) that drives /help and the completion popup (/ or @peer, with descriptions): /help, /activity, /status, /model, /new, /clear, /compact, /sessions, /outputs, /runs, /memory, /skills, /tools, /mcps, /peers, /diff [since], /fold [object] [colour], /attach <path>, /attachments, /clear-attachments, /quit (alias /exit). Panels are FloatingPanels on the overlay layer docked above the composer, dismissed by Esc or click-outside. /activity calls host.activity.list over host.sock (needs you / running / scheduled) and answers listed approvals and questions via host.approval.respond / host.clarification.respond; without a daemon it says so. Configuration verbs (workspace, email, sandbox, …) live in alpi setup — the TUI is for chat and inspection.

Approvals and questions raised by the in-process engine are prompt panels that cannot be dismissed by a click or by opening another panel. They queue (the first shows +N queued), show a live countdown to the engine deadline (60 s approval, 300 s question), and Esc answers deny / cancel immediately. A prompt that times out, shown or still queued, leaves a line in the transcript.

Interrupt: sending a new message, Esc, or Ctrl+C stops the running turn; each also resolves any open prompt so the tool thread never blocks. engine.interrupt_requested is polled at 3 points; long-running tools (research) poll tool_state.is_interrupted(). Skipped tool calls get a [skipped — user interrupted] tool message to preserve OpenAI's pairing invariant. /quit interrupts first, then exits. /clear, /new and Ctrl+L refuse while a turn runs.

Keys: Enter sends; Ctrl+J (or a trailing \ then Enter) adds a line — Shift+Enter too on terminals that report it (CSI-u); pasted newlines are kept; ↑/↓ recall sent messages from an empty composer; Tab completes. Esc answers a prompt, closes the popup or a panel, or stops the turn. Ctrl+C stops the turn, and quits on a second press within 2 s. Ctrl+O toggles details, Ctrl+L starts a new session (same as /clear), Ctrl+Y copies the last reply (pbcopy/wl-copy/xclip/xsel/OSC-52 fallback chain).

Daemon (alpi/service.py)

One alpi daemon per machine, every profile inside. A single launchd plist (com.alpi.daemon) on macOS or systemd-user unit (alpi-daemon.service) on Linux supervises one Python process that hosts every profile under ~/.alpi/ (default plus each profiles/<name>/) on the same asyncio loop. Per-profile tasks are independently guarded — a crash in one profile's scheduler leaves siblings untouched. Tasks are named <profile>/<capability> (e.g. doc/schedule, builder/alp) so logs + asyncio.all_tasks() stay readable. These are internal capabilities, not configurable services:

All capabilities start for every profile; the host plane is default-only. Jobs and workgroups retain their own enabled/paused state, and access control lives in peer grants and connection roles/scopes.

alpi.service.serve_all(root) is the foreground entry point called from alpi daemon start and from the supervising unit's ExecStart. It:

  1. Walks ~/.alpi/ (default + every profiles/<name>/) to discover profiles.
  2. Configures the root logger at ~/.alpi/logs/service.log (stderr only when it's a TTY, to avoid double-writes under launchd).
  3. Sets the process title to alpi (daemon, N profiles) via setproctitle.
  4. Writes ~/.alpi/service.pid.
  5. Spawns the fixed task set for every profile and waits. _guard_task wraps each one so a crash leaves siblings running.
  6. SIGTERM / SIGINT cancels every task cooperatively; PID file removed on exit.

PID 1 (alpi/pid1.py). Inside the image alpi daemon start is PID 1 (ENTRYPOINT ["alpi-docker"], no init). Before anything else it forks: PID 1 stays a minimal init — waitpid(-1) in a loop, SIGTERM / SIGINT / SIGHUP / SIGQUIT / SIGUSR1 / SIGUSR2 forwarded to its only child, exit code mirrored (128 + signal on a signal death) — and serve_all runs in the child, which owns service.pid and service.lock as before. Every orphan in the container (an exited npx wrapper chain, a docker exec session, the grandchildren of a hard-killed turn) reparents to PID 1 and is reaped there, each one logged to stderr. The reaper is deliberately not inside the daemon: a waitpid(-1) in that process would steal exit statuses from Popen.poll() / wait() and from asyncio's child watcher, which then report 0 or 255 for a child they never saw exit. Outside a container os.getpid() != 1 and the fork is skipped. pytest covers the fork, the forwarding and the exit-code mirror; the PID 1 case itself runs in publish-docker.yml, which starts the built image, leaves an orphan through docker exec, asserts no zombie remains and that docker stop returns the daemon's exit code.

Operational invariants of serve_all (each one is the root cause of a real production incident; do not regress):

Active home isolation. Because N profiles share one process, tools that resolve their home via home.get_home() would all see the same env vars and write to default's home. The engine wraps each run_turn in a home.set_active_home(self.home) contextvar binding (per-thread); get_home() consults this binding before the env. Without it, another profile's memory tool would write to default's USER.md. See tests/core/test_home.py for the isolation tests.

daemon_status(root) is the snapshot used by alpi daemon status and by alpi setup → Services → Daemon: PID, uptime (via ps -o etime), install backend (launchd / systemd / none), and the per-profile services map.

Host plane (alpi/host/)

Control-plane for the desktop / mobile client. Not ALP — the two share a profile but live on different sockets, with different auth models. ALP is peer-to-peer (Noise on TCP, envelope-signed, peers pinned in peers.yaml); host is client-to-daemon. JSON-RPC-shaped over ~/.alpi/host/host.sock with filesystem permissions as the trust boundary; no peer identity, no envelope, no Noise handshake. Desktop and mobile talk to this API; they do not read profile files directly.

Only the default profile hosts this plane — the client always targets default's socket and reaches sibling profiles via the profile parameter on each verb. _run_host refuses to bind on any other profile even if the toggle leaks via manual config edit.

host.device_state owns the device-facing profile state contract: profile lists/summaries, bounded profile file reads, storage stats, email status/config previews, skill lists, workgroup lists, workgroup member rosters, config field edits, and local Ollama model discovery. The desktop Tauri layer keeps its existing invoke(...) command names for UI stability, but those commands proxy to host.* verbs instead of parsing ~/.alpi themselves. Mobile should use the same verb shapes rather than inventing a separate state API.

Two transports, one dispatcher:

  1. Unix socket (~/.alpi/host/host.sock, mode 0600). Local trust = filesystem perms. Used by desktop on the same machine. No token required.
  2. WebSocket (ws://<bind>:49200 by default). Used by mobile and any remote desktop. network.host drives the direct bind/address; the bind is derived from it (see config / security): empty → auto-detected Tailscale CGNAT (100.64.0.0/10) then private RFC1918 LAN; a private/Tailscale IP → that IP; a hostname or an opted-in public IP → 0.0.0.0 (all interfaces); a public IP without host.allow_public_bind → refused (no TCP); Docker → 0.0.0.0. Loopback is never a bind target. A 0.0.0.0 bind leans on the device token (and a firewall/NAT) for access control, so alpi doctor warns whenever the listener binds 0.0.0.0. Per-device token required in every authenticated request's params.auth_token. permessage-deflate is negotiated by default (ws_serve(compression="deflate")); JSON-RPC payloads drop 50–80% on the wire. Clients that don't negotiate fall back to raw. Mobile and desktop keep a persistent multiplexed WS pool per (URL, token) so RPCs don't pay a TCP+WS handshake every call — the dominant cost of "remote alpi feels slow" on Tailscale. Streams (host.chat.send, host.events.subscribe) open their own dedicated socket.

Bind and advertised routes are intentionally separate concerns. The daemon chooses where the host-plane server listens; Connections → Network stores an ordered host.endpoints list of complete ws:// or wss:// URLs and chooses what the paired client should dial. Plain WS requires a private IP literal; hostnames require WSS, and synthesized routes pass through the same validator. WSS terminates at a certificate-validating reverse proxy and forwards to the same daemon listener; it does not create another authorization plane. On a normal Mac or Linux install those often collapse to the same private address. In Docker they do not: the daemon binds 0.0.0.0 inside the container while the QR advertises a configured host.endpoints route or a safe private IP derived from ALPI_NETWORK_HOST.

Wire shape (both transports):

{"id": "<reqid>", "method": "host.<noun>.<verb>", "params": {…, "auth_token": "<token>"}}

Unix socket payload omits auth_token — the local transport is sovereign and bypasses token validation entirely. WS requires a valid token except for one exact bootstrap verb: host.connections.exchange_pairing may redeem a locally-created, high-entropy grant once and then the daemon closes that socket. An empty or missing connections.yaml rejects every ordinary WS request (fail-closed). The connection, role, profile scope and grant are created locally over the Unix socket; remote bootstrap cannot choose them.

The daemon writes either a single response line or, for streaming verbs (host.chat.send, host.events.subscribe), multiple frames followed by a done frame and connection close.

This is distinct from ALP peer transport. Connections / host-plane remote access configures how paired desktop and mobile clients reach their own daemon (host.*). Peer TCP listener configures the optional ALP TCP listener other alpis use for link.* and workgroup.*.

Connections and device credentials (alpi/host/connections.py)

The store lives at ~/.alpi/host/connections.yaml (mode 0600). A connection is the operational identity: {id, label, role, profile_scope, status}. Its devices[] each hold the SHA-256 digest of a separate opaque token (token_hash; the cleartext exists only on the client) plus self-reported client/name/version metadata and last_seen. Auth hashes the presented token before the constant-time compare, so a copy of the file is not a credential. Desktop and mobile may therefore share one connection, its sessions and accounting, without sharing a credential. pairings[] holds only hashed temporary grants and lifecycle metadata. A pending grant expires after ten minutes; the first exchange marks it consumed and appends exactly one device under the same file lock. Terminal metadata is kept for seven days, capped at 50 entries per connection, and omitted from host.connections.list.

The daemon resolves each token to {connection_id, device_id, role, profile_scope} and binds that identity to the request context. A hit bumps the device's last_seen at most once per minute. The engine persists connection_id on new sessions; session list, read, continue, cancel and delete reject sessions owned by another connection. The daily ledger records input/output tokens and USD under by_connection; the run ledger records both IDs. Local Unix/TUI/CLI activity uses the synthetic host connection.

Sessions also persist device_id. A connection's session_scope decides whether that matters: connection (the default) shares every session among the connection's devices; device lets a remote device list, read, continue, cancel and delete only the sessions it created, drops other devices' session_changed frames from its event stream and history, serves each device its own latest_session preview in host.profile.summaries (the summary cache is keyed by connection, device and profile), and narrows the agent's session_read / session_search / recall_sessions tools the same way. Every session records its device whatever the scope, so switching to device also hides a device's earlier sessions from its siblings; sessions without a device_id (saved by alpi before 0.15.20, or started by the daemon itself through the scheduler or host.chat.delegate) stay visible to every device of the connection. ask_user clarifications and command approvals carry the owner of the turn that raised them: host.clarification.pending / host.approval.pending, their respond verbs and the clarification.* / approval.* events follow the same rule for member devices. The local socket and admin-role reads are unaffected. A device flagged provisioner (minted from a pairing grant that carried the flag) may call add_device, pairing_status, cancel_pairing and revoke_device for its own connection without the admin role; it cannot grant provisioning, revoke itself, or reach any other admin verb.

Sensitive mutations pass through one dispatcher audit boundary after their handler returns. admin_audit.py writes only allowlisted identifiers and the stable error envelope; it never serializes request params or handler results. The bootstrap pairing exchange replaces its temporary context with the newly created connection/device identity before writing. Authenticated admin denials are recorded at most once per device/method/minute; invalid unauthenticated traffic stays in operational metrics/logging so it cannot churn durable audit history. host.audit.list is local/admin-only, cursor-paginated and filters an identity whether it acted or was the target.

This boundary covers calls through host.sock and authenticated WebSockets. Direct CLI/setup code paths still mutate their stores without crossing the dispatcher and are explicitly tracked as remaining AUDIT.2 coverage rather than being represented as synthetic host-RPC events.

Three trust tiers gate every WS call:

The admin set lives in _ADMIN_METHODS; the strictly-local set in _LOCAL_ONLY_METHODS (network admin only — no role unlocks those over WS).

Lifecycle:

The daemon migrates the credential store at startup, before it opens the WebSocket listener: a legacy devices.yaml becomes one connection per row with hashed tokens, a pre-0.14.39 connections.yaml is rewritten once with token_hash, and the source is deleted only after the destination has been re-read and verified. A corrupt or empty store is an explicit error, never zero connections, and a failed migration keeps the WebSocket listener down. Rolling back below 0.14.39 is not supported. Leftover copies are listed by alpi doctor; Operations has the removal procedure.

host.devices.* remains as a compatibility RPC alias for older management clients; generated payloads use the new one-time grant contract. Desktop and Mobile continue to consume old QR/link payloads that contain a final token. All new management uses host.connections.*.

Verb namespaces in current shape:

Contract. host.events. is transport, not durable history. The replay window (HISTORY_MAX = 500) is sized for reconnect catch-up within a session of activity — it can drop old rows under load and must never be the source of truth for anything a user can browse. Durable user-visible state lives in the per-profile stores that host.outputs. / host.sessions.* / workgroup transcripts read from. If a UI needs history older than the replay window, it queries those stores, not host.events.history.

Desktop and mobile stream only the active connection. Every other daemon is polled through host.events.history (desktop every 25 s, mobile on its catch-up and background wakes) with the notifiable kinds plus output.created / output.updated: the first raise native notifications, the output kinds only refresh that connection's inbox. Each client keeps one cursor per daemon, shared by its stream and its poll (desktop shares it only between admin routes; a member route sees a filtered stream and keeps its own), so switching the active connection neither drops nor replays an event; a daemon whose next_seq falls below the cursor its request carried restarts it from zero (a late answer below a cursor that moved on since is not a reset). Its replay page then takes only frames whose at is later than the last one the client saw, since a plain restart can restore a counter below the cursor too (history=False emits take a seq that is never persisted), while a frame above the head reported at the reset is always new whatever its clock; with no at seen yet it re-anchors at the head instead. A client's first poll of a daemon anchors without banners but still refreshes that inbox when the page changed it. A polled or replayed request past ts + timeout_s raises no banner.

host.profile.attention (admin) answers what needs the owner in one profile without opening each panel: memory (files at 90 % of their limit or over it: file, used, limit, pct, over), skills (a skill that fails lint, or is inactive because what it requires is missing: name, category, problem = lint|missing, message), schedules (a job whose last run failed and is not paused: id, title, message, at, last_ok_at), plus counts and total. alpi/attention.py computes it; the scheduler's tick and host.profile.memory_write reconcile it against <home>/attention.json, emit attention.changed when the set changes and file one warning output for each memory or skill item the first time it is flagged (again only after it was resolved and broke again). A failed job already files its own error row. alpi profile show and the TUI /status print the same list.

Adding a new verb: create the handler in the matching host/*.py module, register on host_server.Server.register (or register_stream for multi-frame), and call from the desktop / mobile client via the platform's host-client helper. Never expose a verb outside host.* — the namespace check in register enforces it.

Email (alpi/tools/email.py, alpi/mail/)

Email is an on-demand tool, not a listener — nothing polls the inbox and nothing auto-replies. The agent calls email (actions: list, search, read, send, reply, forward, move, delete, download_attachment) whenever a chat or a scheduled job needs to read or send mail; the tool drives the IMAP/SMTP backend (mail/imap.py::ImapClient) or the Gmail backend (mail/gmail.py:: GmailClient + OAuth). Bodies pulled by email(read) pass through the prompt-injection scanner behind an untrusted-content envelope before the model sees them.

Multi-account. A profile holds N accounts — any mix of IMAP and Gmail — modelled in alpi/mail/accounts.py; each account's identity is its address and its id is a slug of that address. The email tool's account parameter picks which one (by address or id); with one account it defaults to that account. Accounts are declared in config.yaml under email.accounts (non-secret shape only). Secrets live in <home>/.env namespaced per account — an IMAP account's password is EMAIL__<ID>__PASSWORD; Gmail OAuth client creds (GMAIL_CLIENT_ID / GMAIL_CLIENT_SECRET) are shared across all Gmail accounts, while each account's token sits at <home>/secrets/gmail_tokens/<id>.json after a one-off OAuth consent. Add and manage accounts via alpi setup → Email or the apps' Email section; probe / remove a single account by id from the CLI with alpi email probe <id> and alpi email remove <id>. There are no email.* scalar config.yaml knobs beyond the email.accounts map.

Per-profile env snapshot. alpi.home.effective_profile_env(home) overlays os.environ (process-level vars: PATH, HOME, TZ, ALPI_PLATFORM…) with <home>/.env (per-profile secrets, quotes stripped) and is the source of truth for all credentials: the per-account EMAIL__<ID>__PASSWORD keys and the shared GMAIL_CLIENT_ID / GMAIL_CLIENT_SECRET. The daemon never mutates os.environ — under multi-profile supervision a global mutation would cross-contaminate every profile. The contract holds across the agent toolchain: tools/email (IMAP's ImapClient.from_env_map), the LLM-override paths in tools/web_extract / tools/read_image, alpi/identity.py, and the model selector / TUI provider gating. Credential edits via the host plane write the file atomically; a running engine reads the current .env on its next turn.

Schedule (alpi/scheduler/)

Tick loop (default 30s) hosted inside the alpi daemon. add schedules a job (kind: cron|once, expression or after_hours). run-once ticks manually for testing. LLM time grounding: when the agent calls schedule(action='add', kind='once', after_hours=N), the engine resolves now from a single source so the agent doesn't drift.

Duplicate guard + in-place edits. add rejects a job whose (kind + cron / run_at / after_hours) matches an existing one AND whose prompt fingerprint (lowercase + whitespace-collapsed first 80 chars) collides. Pass force=true to bypass when the second job is genuinely intentional. Use update to change prompt, cron, notify, or pause state without remove/recreate churn. A job carries a single delivery axis, notify: bool (default false = silent): true pushes the reply to the owner's apps. Failure is not on that axis: a failed job always files an error output and raises schedule.failed regardless of notify, and a run the scheduler could not end does too — killed with the daemon (0.14.49), or wedged past its timeout inside a live one (0.14.50) — through the run sweep described under Runs: a dead child within about a minute, a wedged one roughly ten minutes past its timeout. Legacy jobs with a platform field are migrated to notify on load (platform set → notify: true). Reaching a THIRD PARTY is an explicit email call in the prompt — that's now allowed (the old auto-delivery guard that rejected such prompts is gone).

Scheduled jobs execute through alpi chat --once --emit-events --no-save with ALPI_PLATFORM=cron. The scheduler consumes stdout events to detect tool traces, final reply text, delivery, and failure. It does not write sessions/<id>.json: cron output belongs to schedule delivery/logging, not to local TUI / desktop chat history.

Loop isolation. serve() runs tick() in a dedicated ThreadPoolExecutor(max_workers=2), and host.schedule.fire wraps fire_by_id in run_in_executor before awaiting. Both paths ultimately call subprocess.run(timeout=job_run_timeout(job)) (default 900s, per-job up to 86400s); running them inline would block every other coroutine on the daemon's asyncio loop — ALP responders and host.chat.send streams in sibling profiles all stall for the duration of the scheduled job. The dedicated executor also means the scheduler can't starve chat's default-executor turns. A regression test in tests/core/test_schedule.py::test_serve_runs_tick_off_loop_so_chat_can_progress pins the contract.

First run. A cron job runs at its next occurrence, never on the tick it appears. The schedule tool records last_run_at when it adds a job; a job that arrives without run state (written into jobs.json by hand or by a deploy, or whose schedule/runs.json entry was lost) gets first_seen_at in runs.json from the first tick that sees it, or, if it arrives paused, from the first tick after it is resumed. A fired job is stamped (last_run_at, last_run_status; a one-shot that succeeded is removed) the moment its run returns, before its outputs and events are written and not at the end of the pass, so a daemon restart in the middle of a long pass does not fire a job that already finished. A run that raises (the agent subprocess cannot start, a file is missing) is a failed outcome stamped and reported like any other failure, and the pass goes on with the next job. Each job is re-read from jobs.json just before it fires and skipped if it was removed, paused, fired by hand or no longer due meanwhile, and fires from the fresh copy; a due check that raises skips that job with a logged reason; a stamp that cannot be written is logged and kept in memory, laid over the job on every later tick while the job's kind and run_at are the ones that ran (so it is not fired again, and a one-shot stays gone, but a job edited into something else in the meantime is left alone) and retried first thing on each tick until it lands, without overwriting a newer stamp on disk (a daemon restart while the disk keeps failing loses it); an unreadable jobs.json mid-pass ends the pass. fire_by_id stamps the outcome (a failed stamp is logged, not raised) before it emits schedule.done, so the first host.activity.list after activity.changed already reads it. schedule(action="fire") runs a job now.

Job fields. A job carries an optional one-line description (at most 160 characters, set through schedule(add|update, description=…) and listed by host.schedule.list), and its run state keeps last_run_message (the redacted failure, 300 characters, cleared by the next success) and last_ok_at.

Timezone. Cron expressions evaluate against the machine's system timezone (datetime.now().astimezone() in scheduler/run.py). Jobs are stored with UTC last_run_at but fire according to local wall-clock time. Practical consequence: if you specify 10 12 * * * because you want a 12:10 reminder in Bangkok, the Mac must be set to Asia/Bangkok. Move the machine to a different timezone and the cron fires at 12:10 there, not in Bangkok. No in-job timezone override today — add it via TZ=… in the launchd plist / systemd unit if cross-timezone stability is required.

MCP client (alpi/mcp/)

Spawns user-configured MCP servers (stdio JSON-RPC, SSE planned). Their tools are wrapped and registered as alpi tools. Servers configured in config.yaml under mcp.servers.<name> (command, args, env). Management lives in alpi setup → MCPs and in the alpi mcp command group (add, remove).

External orchestration frameworks. Alpi does not embed LangGraph, CrewAI, AutoGen, or similar graph/supervisor runtimes in core. They overlap with Alpi's own agent loop and bring a heavier dependency, state, and observability model than the local-first runtime needs. Interop belongs at the edge: expose the external workflow as an MCP server and let Alpi call it as a tool, or wrap a local workflow in a scripted skill. ALP is not the adapter layer for these frameworks; ALP is reserved for sovereign profile-to-profile collaboration across machines, while MCP is the interop layer for external runtimes and tools.

Logging (alpi/_log.py, alpi/logs.py)

Every subsystem writes to a single flat folder: ~/.alpi/logs/<subsystem>.log, rotated at 1 MB with 3 backups (MAX_BYTES / BACKUP_COUNT in _log.py). Same format everywhere (%(asctime)s %(levelname)s %(name)s %(message)s) so alpi logs can merge them by timestamp prefix. The source tag on display comes from the filename.

Three sources today (file on disk + the writer that produces it):

The machine-wide structured administrative trail is separate: ~/.alpi/logs/admin-audit.jsonl, 5 MB plus three rotated generations, mode

  1. Each row is capped at 4 KB, bootstrap/auth failures have their own

one-row-per-minute budget, and target fields are allowlisted per method. It is JSONL because Desktop and host.audit.list filter by actor, target and result. alpi audit-log renders it for the console. It is not included in alpi logs: those commands merge human-readable .log streams, while this trail has its own bounded query contract. Chat turns are not copied into it; sessions already carry their owning connection_id.

The alpi logs --source CLI choice list also accepts schedule. Inside the unified daemon, scheduler events route through the root logger and land in service.log — the filter value is kept so that any standalone or legacy schedule.log (e.g. from an older scheduler.run.ensure_running() invocation that ran out-of-process) stays selectable.

Why logs are NOT inside sessions/: sessions/ is a structured store (one JSON per conversation, indexed by id, consumed by session_search and the resume flow). Mixing freeform logs would break the glob pattern and the cleanup semantics. Logs are the index and audit trail; sessions are the content. Peers, not nested.

Why one flat folder (logs/) instead of per-subsystem dirs: tiny <subsystem>/logs/ folders with a single file each is pure noise. The service keeps non-log state in its own places (schedule/jobs.json, alp/alp.sock, service.pid at the profile root) — only the .log files consolidate.

Adding a new source is two lines: from alpi._log import get_subsystem_logger; logger = get_subsystem_logger(home, "my-sub"). alpi logs picks it up without changes; add the tag to the --source choice list in cli.py::logs_cmd if you want it filterable.

Doctor (alpi/doctor.py)

alpi doctor — live health check. Verifies external capabilities actually respond, not just that they're configured. Same entry point from the CLI and from alpi setup → Health check; the status in the setup menu row (all green / N warning(s) / N failing) runs the full check too.

Checks:

Parallelism: the network-bound tasks (IMAP/Gmail/MCPs) submit to a ThreadPoolExecutor(max_workers=8). Sync checks (model, workspace, services, security) run on the main thread while the pool works. Total wall time ≈ slowest single task, not sum — ~5-10 s on a healthy profile.

Progressive rendering: run_and_render() uses rich.live.Live — every row appears immediately with a cyan spinner, each resolves to ✓/✗/! as its future completes. Animation at 10 fps via a manual frame cycler (rich's Spinner objects can't be appended to Text). Layout is stable (same rows, same column widths) so the eye doesn't jump.

Exit codes: 1 if any check returns fail, 0 for warn/info/ok. Warnings don't break cron. The wizard entry ignores the exit code — it press-enter-waits so the user can read.

Ops digest (alpi/ops_digest.py)

alpi digest [--since 7d] is the read-only evidence rollup for operator decisions. It deliberately does not own new state: each section reads the primitive owned by another subsystem.

The command has two renderers: a compact Rich view for humans and --json for scripts. The JSON is a dataclass dump of the report shape. It is not an observability daemon, dashboard, recommendation engine, or telemetry channel. Tests pin the read-only contract by snapshotting the profile tree before and after a digest run.

Sessions (alpi/session.py)

Turn-based JSON: schema_version: 2, turns: [{at, user, tools[], assistant}], and cumulative metrics. ToolLog carries at, name, args, result, ok, duration_s, reasoning; large user / assistant / reasoning / tool payloads are persisted as bounded previews plus {bytes, sha256, truncated} metadata, not raw unbounded blobs. host.session.read normalizes both legacy and v2 payloads back to the client-facing shape, so desktop/mobile can render old and new sessions the same way. Empty sessions (no user message) are NOT saved.

Listings are bounded. host.sessions.list fully parses normal files but uses a cheap summary path for files above the large-session threshold; host.profile.summaries uses count_sessions() and latest_chat_summary() so profile/sidebar RPCs never parse 50 MB histories just to show a count or latest row.

Live replay is a separate sidecar (sessions/_events_<id>.jsonl). It is append-only within the active turn, sequence-numbered, and bounded: incremental assistant_delta / reasoning_delta frames are preserved exactly, while very large text fields are clipped. The canonical durable review remains sessions/<id>.json; the sidecar is for reconnect/backfill, not long-term full-fidelity storage.

sessions/ is local human chat history: TUI, desktop, and manual alpi chat --once runs that should be resumable. --continue, tui.auto_resume, host latest_session, and desktop profile opening all treat only kind == "chat" as resumable local history. Historical files whose first user message starts with [SCHEDULED:], [workgroup-poller], or another system bracket are ignored by resume/profile history.

TUI resume. Bare alpi resumes the most recent session when tui.auto_resume: true; -c / --continue is the manual override.

Scheduled jobs do not persist session files. The scheduler uses --no-save because it only needs emitted final reply/tool events for delivery and audit; keeping a resumable transcript would make background jobs appear as user chats.

@-mention threads (alpi/alp/mention_thread.py). When peer A @-mentions peer B over ALP (link.ask), the receiving side runs a fresh Engine per turn — but B persists a small thread at <B-home>/mentions/<A>@<conversation>.json, capped at 20 turns, where conversation is an opaque id A derives from its own source session (the peer tool reads it from the run context; the host chat, TUI and --once pass their session explicitly). Successive mentions from the same A conversation carry conversational memory ("what I said before" resolves) without polluting B's local --continue (which only reads sessions/); a new A conversation starts clean, and the same conversation value from peer C selects nothing. A request whose conversation cannot be established runs with no history and writes none; a caller that sends no conversation at all predates the identity and keeps the legacy <A>.json thread, which is never imported into conversation threads. The result's history field says which applied, so a sender can tell when a target ignored the identity. Wipe via setup → Cleanup → Mentions.

Security model

Two layers:

Threat model: prompt injection via email/web content, LLM-issued tool calls on the user's machine, direct user input (trusted), and network adversaries for ALP links. Full discussion in Security.

Cross-cutting concerns

Profiles

alpi -p <name> resolves home to ~/.alpi/profiles/<name>/. ALPI_PROFILE env var is the same. No sticky "current profile" file — resolution is fully explicit. The single daemon (com.alpi.daemon / alpi-daemon.service) supervises every profile from one process; tasks are namespaced <profile>/<service> so they stay distinguishable in logs and asyncio.all_tasks(). Inside a turn, home.set_active_home(home) binds the per-thread contextvar consulted by home.get_home() so tools resolve to the right profile even though every concurrent turn shares the daemon's env.

Workspace

cfg.workspace (or cwd fallback if unset) is the default root for relative paths — not a wall. File tools and terminal can reach absolute paths anywhere except the sensitive denylist. Real workspace-only isolation is the opt-in OS sandbox (Layer 2). Configure it via alpi setup → Workspace; the TUI top bar read-outs the resolved path but does not edit it.

Dependencies

Hard runtime deps are kept tight — every line in pyproject.toml's dependencies is actually imported by alpi/. The audited set, with one-liner for why each earns its place:

Optional dev extra: pytest + pytest-asyncio for the test suite, ruff for lint, pip-audit for CVE scans.

Security posture: uv run --with pip-audit pip-audit must run clean against the full lockfile before each release. Known-CVE deps are not allowed to accumulate — drop or upgrade.

Testing

python3 scripts/validate.py is the one command before claiming done: it runs the release check and every suite the working tree touches (see AGENTS.md). Directly, uv run pytest -q is the fast suite, --integration adds sockets and sandbox-exec, and --llm enables real-LLM tests (a few cents on free models).

Key fixtures (tests/conftest.py):

Contracts clients and consumers rely on

Breaking one of these breaks a client, a gateway or a peer. Change the contract, its consumers and this section in the same change.

Non-obvious things to know