alpi's settings live in ~/.alpi/config.yaml (or ~/.alpi/profiles/<name>/config.yaml for non-default profiles). This page lists every knob, its default, and what it controls.
What ships in the YAML
On first install alpi only writes the sections you're likely to tweak — and where defaults are platform-dependent enough to deserve visibility:
model: "" # empty on a fresh scaffold; pick it in `alpi setup`
providers:
ollama: []
mcp:
servers: {}
Everything else (tool limits, TUI flags, fallback models, workspace) falls back to the defaults below at load time. Add a key to the YAML only when you want to override it.
How to change settings
Three options:
- CLI wizards:
alpi setupcovers model selection, email credentials, MCP servers, sandbox posture, voice, peers, workgroups, disk cleanup, and the alpi daemon's lifecycle.alpi setup → Cleanupinspects the profile's heavy dirs (audio cache, old sessions, run journals older than 30 days, schedule output) and the knowledge index SQLite freelist (VACUUM-not-unlink), with one-shot confirmation per category.alpi setup → Servicesexposes daemon lifecycle (default profile), the shared accessible address, schedules, and client connections. The ALP section owns its peer TCP listener. Scheduler, ALP, workgroups, and the default host plane are daemon capabilities, not user-selectable services. The firstalpi setupauto-installs the daemon, so the lifecycle row is mostly read-only after that. - Edit the YAML: open
~/.alpi/config.yaml(or~/.alpi/profiles/<name>/config.yamlfor non-default profiles) and change values manually. Restart whatever surface was affected. Cosmetic knobs (tui.*,tools.max_steps_per_turn,tools.stt.model,fallback_models) live here. - Populate
.envdirectly (non-interactive, CI / devcontainers): alpi does not ship a.env.example— the Reference sections below (Core, Email — IMAP / Gmail) name every key with its default. Create~/.alpi/.envyourself with just the keys you use and alpi picks them up on next launch.
Reference
Core
| Key | Default | Type | Takes effect | ||
|---|---|---|---|---|---|
model | "" (empty; pick via alpi setup → Model — see docs/MODELS.md) | string | next session | ||
workspace | "" (cwd at launch) | string | next session | ||
fallback_models | [] | list of strings — availability chain: when the active model fails before producing any output (provider down, credits exhausted) the turn retries down this list, sticking with the survivor for the rest of the turn | next turn | ||
tiers.fast.model | "" (use main) | string — cheap/fast model for routine work: compaction summaries, the memory reviewer, bio drafting, research(depth=fast), and delegate/schedule runs that opt into tier: fast | next turn | ||
tiers.fast.effort | "" | low \ | medium \ | high — reasoning effort for the fast tier's own model (never inherits the profile effort) | next turn |
tiers.deep.model | "" (use main) | string — stronger model for hard reasoning: research(depth=deep), delegate(tier=deep), and the escalation target when a turn accumulates 3 consecutive tool failures or an empty reply (once per turn, skipped past 80% of budget.daily_usd) | next turn | ||
tiers.deep.effort | "" | low \ | medium \ | high — reasoning effort for the deep tier's own model | next turn |
providers.ollama | [] | list of {name, url} — one per Ollama server | next session | ||
providers.openrouter.models | [] | list of OpenRouter model ids the user has picked | next session | ||
providers.openrouter.ignore | [] | OpenRouter provider slugs (lowercase, e.g. together) excluded from routing for every openrouter/… model of this profile, tiers and fallbacks included. An exclusion, not a pin: OpenRouter still picks freely among the rest, so :nitro sorting keeps working. Sent as extra_body.provider.ignore | next turn | ||
public_bio | "" | string — one-line public tag-line broadcast to every workgroup this profile joins (source of truth for Member.bio on the hub). Empty = don't publish; peers see name only. AGENT.md stays private. | next workgroup.join | ||
paused | false | bool — profile-level pause flag. Surfaced in the desktop / mobile profile summary so paired apps can show + respect the state; the daemon itself does not gate turns on this flag. Persisted only when true. | next host-plane read |
Tools
| Key | Default | Type | Takes effect | ||||
|---|---|---|---|---|---|---|---|
tools.max_steps_per_turn | 100 | int | next turn | ||||
tools.max_parallel_tool_calls | 4 | int — maximum concurrent calls in an all-parallel-safe tool batch | next turn | ||||
tools.deny | [] | list of tool names | next turn | ||||
tools.execution.backend | "local" | local \ | docker — terminal shell backend; background jobs are refused in Docker | next turn | |||
tools.execution.docker_image | "python:3.12-slim" | Docker image used when the execution backend is docker | next turn | ||||
tools.web_extract.model | "" (use main) | string — a model id, or the literal fast / deep to reference a tier | next turn | ||||
tools.read_image.model | "" (use main) | string — a model id, or the literal fast / deep to reference a tier | next turn | ||||
tools.terminal.sandbox | false | bool | next turn | ||||
tools.terminal.allow_network | false | bool | next turn | ||||
tools.terminal.approval.allowlist | [] | list of pattern descriptions and/or command globs (see below) | next turn | ||||
tools.browser.vision | false | bool | next turn | ||||
tools.web_search.max_per_turn | 25 | int | next turn | ||||
tools.web_search.backends | [duckduckgo, yahoo, brave] | list of ddgs text engines, in the order web_search asks them; names the installed ddgs does not ship are dropped, and an empty or unusable list means the default | next search | ||||
tools.browser.allow_local | false | bool — let the browser tool navigate loopback only (127.0.0.1, ::1, and hostnames that resolve to loopback such as localhost). RFC1918 / CGNAT / Tailscale stay blocked even when this is on; the exemption is loopback-only, matching _guards._is_loopback. Off blocks every local target; on is for hitting a local dev server you trust. | next turn | ||||
tools.budget.per_result_chars | 100_000 | int (-1 = unlimited) | next turn | ||||
tools.tts.voice | "en-US-AriaNeural" | Edge TTS voice id | next turn | ||||
tools.tts.rate | "" | string ("+10%", "-20%") — speed | next turn | ||||
tools.tts.pitch | "" | string ("+5Hz", "-10Hz") — pitch | next turn | ||||
tools.tts.auto_read | false | bool — apps auto-play each agent reply aloud | next turn | ||||
tools.stt.model | "base" | tiny \ | base \ | small \ | medium \ | large-v3 | next turn |
tools.stt.language | "" (auto) | ISO code (en, es, ...) | next turn | ||||
tools.attachments.max_text_tokens | 0 (auto) | int (tokens) | next turn | ||||
tools.<name>.max_result_chars | — (unset) | int (-1 = unlimited) | next turn |
web_search needs no API key. It asks one engine at a time, in the tools.web_search.backends order, and stops at the first that answers. The default of three is provisional: it comes from the one measurement available, from a residential connection, and is not a claim that the other engines ddgs ships are broken. Which engines answer depends on the machine's IP, so a host can list others. A search gives up after 25 seconds, and at most two engine requests are ever in flight. An engine that times out, is rate-limited or errors is skipped for the next 15 minutes by the running process; an empty answer never sidelines an engine. When every engine is cooling down, a search asks only the one due back first. A search where every engine asked came back empty reports no results; one that met timeouts or errors fails and names each engine and what it did.
tools.read_image.model is a dedicated override for image inspection, not a second conversational model. Set it from alpi setup → Routing models, or the Vision model row in desktop/mobile. read_image and browser(screenshot, question=...) use it; clearing it makes both fall back to the profile's main model. It does not reroute image attachments sent directly to chat: those remain part of the main model's turn.
max_steps_per_turn is a runaway-loop backstop, not the cost guard — the cost guard is budget.daily_usd. It counts model iterations, and the configured value (default 100) is literal for every provider and budget. Hitting the cap does not discard gathered work: normal chats get one tools-off best-effort wrap-up, while detached workgroup turns get one workgroup_post-only handoff. Raise it explicitly only for profiles whose skills legitimately need longer tool chains.
tools.budget.per_result_chars caps the size of any tool output the LLM sees in-context, with a … [N chars elided by tool budget] suffix when hit. Prevents a single read_file on a 5 MB log from blowing up a turn. Per-tool overrides via tools.<name>.max_result_chars — set -1 on read_file if you want the LLM to get the whole source deliberately, or lower a chatty tool's cap.
tools.max_parallel_tool_calls only applies when every call in the model's batch declares itself parallel-safe. A mutation, terminal call, unknown tool, or mixed batch is an exclusive barrier and retains serial ordering. Results are appended to the conversation in the model's original call order.
tools.execution.backend=docker runs terminal shell processes in an ephemeral docker run --rm world. The workspace, profile home, and explicit working directory are mounted at their existing absolute paths so filesystem tools and processes see the same namespace. Network is disabled unless tools.terminal.allow_network=true. Background terminal jobs are refused in this backend so a detached container cannot outlive its run. The local backend remains the default and continues to compose with tools.terminal.sandbox as before. Dedicated workers such as skill scripts and speech transcription remain host-side; this setting is not a whole-agent filesystem sandbox.
Precedence: tools.<name>.max_result_chars (if set) → tools.budget.per_result_chars → hardcoded 100_000.
Not implemented (tracked, not planned): per-turn aggregate cap and inline preview. Comparable agents carry both, but alpi only ships them if real turns start burning through several large tool results.
tools.deny is a per-profile denylist of tool names. Denied tools are absent from the schema the LLM sees (it can't reach for what it doesn't know about) AND refused by the executor as defence in depth — if a stale context or a peer's link.ask names a denied tool, the call returns tool denied for this profile: <name> instead of running. Unknown names are no-ops, so typos are harmless. Denying alpi_knowledge also drops the self-knowledge rule from the system prompt, so the model is never told to call a tool it cannot reach. An entry may end in * to cover a family — github__* denies every tool of the github MCP server. What one peer's inbound link.ask turns may run on this profile is set the other way round, as the only tools allowed, under tools.allow in its peers.yaml record; see ALP.md → Per-peer tool policy.
Canonical names are the strings used at registration time — write_file, edit_file, terminal, email, schedule, delegate, peer, knowledge, alpi_knowledge, research, browser, workgroup_post, workgroup_file, workgroup_search, etc. See alpi/tools/__init__.py for the full registry. Note: knowledge is the user's workspace Markdown wiki; alpi_knowledge is the packaged docs tool for alpi itself.
Useful for tightening a profile that is exposed to less-trusted input — e.g. a "librarian" profile that other peers reach via link.ask and that has no business writing files, running shell, or sending mail:
# ~/.alpi/profiles/archi/config.yaml
tools:
deny:
- write_file
- edit_file
- terminal
- email
- schedule
- delegate
Today this is YAML-only. There is no alpi setup wizard for deny and no surface for it in the desktop/mobile apps — the surface is power-user enough that raw names beat any UI we'd build right now.
tools.terminal.sandbox enables OS-level isolation on shell commands (macOS sandbox-exec, Linux bubblewrap). Toggle via alpi setup → Sandbox, or directly in YAML. It has no working form in the supported Docker runtime — the container grants none of the namespaces bubblewrap needs — so there it refuses every terminal call rather than run one unsandboxed, and the container plus its volume is the isolation boundary instead (see Deployments). A turn from a member device, or from an ALP peer without tools.allow, ignores false here: its commands run only in Linux bubblewrap with the alpi home hidden (or the Docker execution backend), and are refused in the Docker runtime and on macOS (Security). The TUI top bar shows the current state (sandbox on / off). Most useful on profiles that run unattended (schedule, sub-agents) — see Security for the recommended pattern + platform requirements.
allow_network has no effect unless sandbox is on. When sandbox is on and allow_network=false, the flag blocks ALL agent-initiated network:
- The
terminalsubprocess is denied sockets (sandbox-exec / bwrap). - Python-native tools (
web_fetch,web_search,web_extract,browser,tts,email,read_imageon URLs) refuse with a clear error. - The LLM call itself (litellm) is exempt — it's the agent's brain, not an exfiltration vector.
The TUI top bar shows offline instead of sandbox when network is locked, so unattended profiles can be audited at a glance.
tools.terminal.approval controls the command approval system — a layer on top of the sandbox that gives the user a chance to approve borderline destructive commands instead of blocking them outright. Each terminal call is classified by a small pattern list into three severities:
- safe (default, no match) — runs without prompting.
- caution — matches a pattern that's often legitimate but
sometimes destructive. Examples:
rm -rf <dir>,chmod 777orchmod a+w,sudo <cmd>,git push --force,git reset --hard,DROP TABLE,kill -9. These pause for user approval in the TUI with four options:Once(this call only),Session(allowlist the pattern until restart),Always(persist the pattern description totools.terminal.approval.allowlistin config), orDeny(abort the tool call). On non-interactive surfaces (schedule) these auto-deny with a clear error telling the user to rerun from the TUI or edit the config allowlist. - dangerous — matches a pattern that's almost never legitimate.
Examples:
mkfs,dd of=/dev/…, fork bomb, pipe-to-interpreter from an unknown URL (curl … | bash),chown -Ron/,~or$HOME, reading SSH private keys, writes into/etcor/var. These are always blocked. No override — if you genuinely need to run one of these, do it directly from your shell, not through the agent.
Allowlist entries come in two shapes, sharing the same list:
Pattern descriptions — the human label attached to one of the built-in caution regexes (recursive rm, sudo, git force-push, git hard reset, chmod 777 / a+w, sql drop / truncate, process kill -9). A pattern-desc entry allows every command of that severity-category. This is what the Always button writes.
Command globs — any other string is treated as an fnmatch pattern matched against the literal command (whitespace-trimmed). Use this for per-command exceptions when the category-level bypass is too broad. Globs only override caution classification; dangerous commands stay blocked. Globs also do not apply to compound commands (containing &&, ||, ;, |, newline, backticks, or $(…)) — otherwise "sudo apt *" would also approve sudo apt update && rm -rf build. Compound commands fall back to the prompt unless a category-desc bypass covers them.
tools:
terminal:
approval:
allowlist:
- recursive rm # category: every rm -rf passes
- sudo apt * # glob: any sudo apt subcommand
- git reset --hard origin/main # glob: this exact command only
- git push --force origin my-branch # glob: only this branch's force-push
Session approvals live in memory (a module-level set) and die with the TUI process. Permanent approvals persist to config.yaml via the Always button (which writes a pattern-desc) or by hand-editing the list with globs. Dangerous commands never get an allowlist entry.
This layer composes with the sandbox (tools.terminal.sandbox): the sandbox is an OS-level boundary (network, filesystem writes outside workspace) that catches what the approval layer misses; approval is user-in-loop for the subset of commands that are legitimately destructive inside the allowed scope. Both can be on at once; the approval check runs first so the user sees the prompt before the sandbox has a chance to refuse.
tools.browser.vision lets the browser(screenshot, question=…) action auto-chain the screenshot into the vision model (tools.read_image.model or the active main model) and return the answer instead of the file path. When false (default), screenshot always returns the path and a hint pointing at read_image so the LLM can decide whether to pay for vision per call. Useful to turn on in an exploratory profile; keep off in watchdog/unattended profiles so the agent doesn't burn vision tokens silently.
Image resizing is automatic: any image whose longer edge exceeds 1568 px (Anthropic's recommended bound) is downscaled before base64-encoding to the model. Vision-model cost scales with resolution — a 4K screenshot costs ~9× more tokens than its 1568-px version for the same content. Aspect ratio is preserved, PNG-with-alpha stays PNG, everything else rounds-trips through JPEG q=85. SVG (vector) is skipped. Not a knob — it is a fixed constant (alpi.tools.read_image.MAX_EDGE).
The research sub-agent's depth tiers (fast = 8 steps, normal = 15, deep = 30) are product definition, not user config. fast and deep share their names with the model tiers — a depth also picks the matching tier when configured, while normal runs on the main model. The agent picks the depth name from intent (fast = single-answer lookups, normal = comparative research, deep = exhaustive surveys); the step ceilings live in alpi.tools.research.DEPTH_STEPS_DEFAULTS.
tools.tts.voice selects the Edge TTS voice used by the tts tool. Any Microsoft Neural voice id is valid (es-ES-AlvaroNeural, en-US-AriaNeural, fr-FR-DeniseNeural, ...). Output is an MP3 cached under ~/.alpi/cache/tts/<hash>.mp3 — same text + voice reuses the cached file. Edge TTS runs against a free Microsoft endpoint (no API key), so there's no per-call cost. To use a different voice per call the agent can pass voice=... directly without touching config. alpi setup → Voice gives you a curated shortlist (10 common-language voices) plus a "custom" entry to type any voice id.
The daemon never plays audio itself — the tts tool returns the cached file path and stops. The alpi mobile / desktop apps stream playback on demand from a per-message button, and — when tools.tts.auto_read is on — auto-play each agent reply aloud as it arrives (your own messages are never read); they synthesize through the same Edge TTS path via host.voice.preview. To deliver the MP3 to a third party the agent chains email(send, attachment=<path>) as an audio attachment. Workgroups carry an analogous hub-local auto_read flag in the workgroup meta (set from the desktop/mobile workgroup settings) that auto-reads agents' messages — never your directives; it is not replicated to members.
rate and pitch are config-only (not per-call args) — persistent prosody defaults. Leave empty for neutral. Text is capped at 1000 chars (~1 minute); longer input is rejected. Output is always MP3.
tools.attachments.max_text_tokens caps how much extracted text from a single attachment reaches the model — applied identically to text/source files, digital-PDF text, and scanned-PDF OCR. Denominated in tokens; the engine converts to characters at ~4 chars/token (attachments.CHARS_PER_TOKEN).
Default 0 = auto: the cap tracks the active model's context window — half of it per attachment (AUTO_TEXT_WINDOW_FRACTION), resolved from litellm.get_model_info. So a 200k-context model gives an attachment ~100k tokens, a 1M model ~500k, a small local model proportionally less — no config needed. When litellm can't resolve the model (some openrouter/… ids, custom Ollama names) it falls back to FALLBACK_TEXT_TOKENS (100k). A positive value overrides auto with a fixed per-attachment cap — set it to bound cost on a large-context model, or to force more text on a model litellm mis-sizes.
This is the real content ceiling; the per-file byte caps (MAX_TEXT_FILE_BYTES 2 MiB for text, MAX_FILE_BYTES 20 MiB for PDF/image) only gate acceptance. It does not change page rendering or OCR page count (that is the fixed SCAN_MAX_PAGES scan cap); and to feed a text file larger than 2 MiB you would also need a higher acceptance cap.
tools.stt.{model,language} control the stt tool backed by faster-whisper running on CPU. First call downloads the model weights (~40 MB for tiny, ~150 MB for base, ~500 MB for small, ~1.5 GB for medium, ~3 GB for large-v3) into ~/.cache/huggingface/ and keeps them forever. Pick the smallest model that meets your accuracy bar — base is the sweet spot for spoken messages/voice notes; small or above for podcasts/meetings. language defaults to "" (auto-detect); set to an ISO code (en, es, fr, ...) only when auto-detect fails on short clips.
Runtime
Provider stale-call hardening for LLM streaming turns: watchdogs that fail a slow/stuck provider instead of hanging the turn, plus jittered retries before any output reaches the consumer. A timeout of 0 disables that watchdog.
| Key | Default | Type | Takes effect | |||
|---|---|---|---|---|---|---|
runtime.first_byte_timeout_s | 300 | seconds (0 = off) | next turn | |||
runtime.stream_idle_timeout_s | 120 | seconds (0 = off) | next turn | |||
runtime.stream_max_duration_s | 600 | seconds (0 = off) | next turn | |||
runtime.max_retries | 2 | int | next turn | |||
runtime.retry_backoff_s | 1.5 | seconds (base; exponential + jitter) | next turn | |||
runtime.prefetch | "" | `"" \ | auto \ | all \ | off` | next daemon start |
first_byte_timeout_s is generous so slow reasoning models aren't killed before their first token; bump it for very slow local Ollama or long-thinking models. stream_max_duration_s bounds one provider request even while it keeps emitting deltas. The ten-minute default leaves room for long reasoning while preventing one request from consuming an entire workgroup phase; set 0 to disable it. Retries fire only for transient failures (timeouts, connection drops, 429/5xx) and only before visible text reaches an interactive client. Detached workgroup turns may replay a partial attempt because their streamed text is not exposed; an enabled request-duration limit is treated as transient and can retry the same model within the turn.
An empty runtime.prefetch selects auto outside Docker and off in Docker. all forces Chromium and embedding weights to warm after daemon startup; off keeps their existing first-use loading behavior.
Model reasoning
Optional reasoning-effort hint passed alongside cfg.model to providers that support it (Anthropic extended thinking, OpenAI o-series, DeepSeek R1, etc.). Applied only to the profile's default model — mid-chat model overrides and tool sub-models (research, delegate, web_extract, read_image) ignore it. Models that don't recognise the hint are unaffected.
| Key | Default | Type | Takes effect | |||
|---|---|---|---|---|---|---|
model_reasoning.effort | "" (no reasoning param sent) | `"" \ | "low" \ | "medium" \ | "high" — "off" written to disk normalises to ""` on load | next turn |
model_reasoning:
effort: medium
Memory
| Field | Default | Notes |
|---|---|---|
memory.review_interval | 0 (off) | Post-turn reviewer cadence. N > 0 fires a daemon-thread reviewer every N user turns that snapshots the conversation and writes durable facts via memory(action="add"). Append-only — the reviewer cannot replace/remove. Opt-in by design. |
Memory file caps, duplicate thresholds, low-confidence pruning age and the auto-compaction ratios are product constants, not user knobs. They only become configurable if real logs/compaction.jsonl or memory-review traces show repeated failures a fixed default cannot solve.
Retention
Off unless a profile asks for it: without this block nothing is ever deleted. With a window set, the daemon sweeps that profile once a day and deletes what has aged past it. alpi setup → Cleanup remains the manual path and keeps its own fixed thresholds.
| Key | Default | Type | Takes effect |
|---|---|---|---|
retention.runs_days | 0 (keep forever) | days | next daily sweep |
retention.sessions_days | 0 (keep forever) | days | next daily sweep |
Windows must be non-negative YAML integers. Invalid values (including booleans, fractions and quoted numbers) disable that window rather than enabling deletion.
The sweep removes finished run journals in runs/ and idle sessions in sessions/ whose last activity is older than the window. It never touches logs/runs.jsonl (the run ledger that carries costs), a run still running or registered as active, a session with a turn in flight or named by a run journal that is still open (that covers the CLI, the TUI and scheduled children, which never register with the daemon), or a session that served a workgroup whose directory still exists under alp/workgroups/ — a dispatch turn names its workgroup as (wg_id=…) in its first message, and that name is the only link a session keeps; a workgroup session whose id cannot be read is kept rather than guessed about. Age and open runs are checked again under a file lock that every writer of sessions/ shares (a session save, a run start), so a turn or a run that lands after selection saves its session, and a session that cannot be read at that moment is left alone rather than treated as old. If any run journal or its directory cannot be verified, session deletion stops and the daemon logs the error; an incomplete inventory never means there are no active sessions. Repairing the journal allows the next sweep to retry. A session's cost is archived to the ledger before its files go. A file that cannot be deleted is reported in the daemon log and the sweep continues. A busy profile — hundreds of scheduled runs a day, one session each — wants a short window:
retention:
runs_days: 7
sessions_days: 7
TUI
alpi's TUI is built on Textual and is the primary interactive surface. Replies stream in and settle as Markdown; each turn's tool calls collapse into one N steps · Xs row that opens to per-call cards, reasoning collapses into a Thought for Xs row, and the layout collapses labels below 60 columns. tui.accent recolours highlights and the profile name, and tui.fold names the origami model that marks the profile in the apps and the console beside its colour (the diamond until you choose one; set it with /fold in the TUI or alpi setup → Appearance; the default profile is always the alpaca in the brand accent and ignores both). The console draws the object as half-block art in alpi profile show and the alpi setup header and as a one-cell glyph on the active entry of profile and TUI lists, only on a truecolor UTF-8 terminal; elsewhere it keeps the diamond. tui.theme picks dark or light.
What the top bar shows (left to right):
alpi <version> │ profile <name> <size> │ workspace <path>
<size>is the total disk footprint of the active profile home dir (~/.alpi/for default,~/.alpi/profiles/<name>/otherwise). Cached for 30 s; refreshed when you change profile, workspace, or model. For the default profile theprofiles/subtree is excluded so it doesn't conflate with sibling profiles. Hidden in narrow mode (< 60 columns).- Workspace shows the resolved workspace path, or
not setin error colour when no workspace is configured and alpi falls back to cwd.
What the status line shows (bottom): model · context use · session cost · daily budget · sandbox (or sandbox · offline when tools.terminal.allow_network=false) · unread inbox items · prompts waiting on you (this chat plus the daemon's), then key hints for the current context. Hints drop first when the terminal is narrow.
Config knobs (tui.*):
| Key | Default | Type | Takes effect | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
tui.show_cost | true | bool | next session | |||||||||||
tui.show_tokens | true | bool | next session | |||||||||||
tui.show_reasoning | true | bool | next session | |||||||||||
tui.accent | #f0b447 (the amber profile colour; the brand accent is ink, #f3efe6 dark and #14110c light, for the default profile) | CSS color (hex / named / rgb) | next session | |||||||||||
tui.fold | diamond | diamond \ | house \ | heart \ | plane \ | shield \ | rocket \ | star \ | tree \ | box \ | crown \ | feather \ | bulb | next client refresh |
tui.theme | dark | dark \ | light | next session | ||||||||||
tui.auto_resume | false | bool | next launch |
tui.auto_resume makes bare alpi behave as if -c / --continue was passed — the last session is loaded automatically. Use /new inside the TUI to start a fresh thread without changing the config. The flag does not affect alpi chat --once (scripts and scheduled jobs always start clean) or explicit -c usage (still an override).
tui.show_reasoning controls the collapsed Thought for Xs row each turn gets: it holds the model's streamed chain-of-thought (DeepSeek-R1, OpenAI o-series, Claude extended thinking) together with any prose the model wrote between tool calls, and opens on click or Ctrl+O. While the model thinks, the spinner reads Thinking… either way.
When false, the row is hidden from the screen. The reasoning is still persisted to the session file (sessions/*.json) so that re-enabling the flag later brings it back on replay, and so that debug inspection (cat sessions/<id>.json) always has the full context. Non-interactive surfaces (scheduled jobs) never rendered reasoning, so this flag has no effect there.
Email — IMAP / Gmail
Email is an on-demand integration, not a listener: the agent reads, searches, and sends mail through the email tool when a chat or a scheduled job calls for it — nothing polls your inbox.
A profile holds as many email accounts as you want, any mix of IMAP and Gmail. Each account is keyed by its address (its id is a slug of that address), so adding a second Gmail or a third IMAP mailbox is just another entry. The email tool's account parameter selects which one by address or id.
Accounts are declared in config.yaml under email.accounts, which carries no secrets — only the non-sensitive shape of each account:
email:
accounts:
you-at-work-com:
type: imap
address: you@work.com
imap_host: imap.work.com
imap_port: 993
smtp_host: smtp.work.com
smtp_port: 465
you-at-gmail-com:
type: gmail
address: you@gmail.com
Secrets live in <profile-home>/.env, namespaced per account by its id: an IMAP account's password is EMAIL__<ID>__PASSWORD (e.g. EMAIL__YOU_AT_WORK_COM__PASSWORD). Gmail accounts use OAuth — the client credentials GMAIL_CLIENT_ID / GMAIL_CLIENT_SECRET are shared across every Gmail account on the profile, while each account's refreshable token is stored per account at <profile-home>/secrets/gmail_tokens/<id>.json after a one-off consent.
Add and manage accounts with alpi setup → Email (CLI) or the Email settings section in the desktop / mobile apps — its own section, where you add and remove accounts like MCP servers. Operate on a single account by id from the CLI with alpi email probe <id> and alpi email remove <id>.
email.accounts is intentionally absent from the key tables above: it is a per-account map, not a fixed knob.
Budget
One daily spending ceiling per profile, in dollars — or unlimited.
daily_usdcaps real spend (LiteLLM reports a cost per turn for paid APIs; local/free paths report zero and so never hit the cap).- Leave it unset for no ceiling.
A token-denominated cap is intentionally not offered: tokens are not a meaningful business constraint (their cost varies by model), so the only caps that mean anything are dollars or nothing.
The cap covers every turn this profile runs: interactive TUI replies, scheduled jobs, sub-agent spawns (research, delegate, read_image), and inbound ALP calls from pinned peers. It is re-checked before each step within a turn, so a long multi-step turn aborts as soon as it crosses the ceiling rather than running to completion. Counters reset at UTC midnight; no carry-over. The ledger lives at <profile-home>/logs/ledger.json and also records a per-peer breakdown shown by /status, though only the profile total gates new turns. The daemon, scheduled jobs and the console all record into it under one file lock, so charges made at the same moment add up instead of overwriting each other.
| Key | Default | Notes |
|---|---|---|
budget.daily_usd | unset | Hard daily USD cap. Exceeding it surfaces budget-exceeded (interactive) or JSON-RPC -32005 budget-exceeded (ALP). Unset = unlimited. |
budget:
daily_usd: 5.00
Edit interactively via alpi setup → Budget or the desktop app's profile detail.
Network (shared accessible address)
One address, shared by every network listener this profile runs — the device-pairing host plane and the ALP peer listener both bind/advertise on it, each on its own port. Configure it once.
| Key | Default | Effect |
|---|---|---|
network.host | "" | Shared bind/ALP address. Empty = auto-detect a reachable private address. A private IP literal also produces the automatic direct ws:// client route. Hostnames require an explicit certificate-validated wss:// entry in host.endpoints; they never produce a plaintext route. A public IP additionally needs host.allow_public_bind: true; that gate applies to the shared bind used by both planes (host control plane + ALP listener), so without it neither binds TCP. |
network:
host: "" # empty = auto-detect a private address
The detected address is shown read-only in alpi setup → Connections → Network and Desktop (Settings → profile → Service). Set an explicit address in config.yaml, or with ALPI_NETWORK_HOST in Docker. Ports stay per-plane (host.tcp_port, alp.tcp_port). A public IP also requires host.allow_public_bind: true; set both keys by hand.
ALP
ALP always serves the per-profile Unix socket for same-machine peers. The Noise_XK TCP listener is auto-exposed only for the default profile, whenever the machine has a reachable address — bound to the shared network.host (above), or an auto-detected overlay/LAN address, or 0.0.0.0 in Docker. Named profiles stay Unix-only unless they set their own explicit, unique alp.tcp_port (otherwise every profile would fight over the shared port). With no reachable address (no network.host, no Tailscale/LAN, not Docker) even default stays Unix-only.
| Key | Default | Notes |
|---|---|---|
alp.tcp_port | 7423 (default profile only) | The ALP peer TCP port. Auto-exposed for the default profile; a named profile binds TCP only if it sets its own unique port here. The address is network.host. |
alp.link_idle_timeout_s | 60 | Cancel link.ask after this many seconds without a signed response or progress frame. 0 disables the idle watchdog. |
alp.link_max_duration_s | 0 | Optional absolute cap for one link.ask; 0 allows an active turn to run without a fixed wall-clock limit. |
alp.max_active_workgroups | 5 on the default profile | Admission threshold, not a hard cap: it is checked when a queued pipeline is admitted, so a workgroup that re-enters the active set another way (resume after a pause, a rewind, a #task re-opening a pipeline that closed #done BLOCKED) is not re-checked against it and the active count can exceed the number. A running pipeline or a deliberation with an open task each hold a slot. Excess pipeline launches and triggers wait in a persistent FIFO queue; deliberations always open and count. A hub's own value overrides the default profile's daemon-wide one; 0 = unlimited. Set with alpi workgroup limit N (--inherit removes a hub override) or from the desktop; workgroup list shows the cap, its origin and the queue. |
alp.working_after_s | 30 | Post an automatic #working heartbeat when a member turn remains silent for this many seconds. 0 disables it. |
alp:
tcp_port: 7423
link_idle_timeout_s: 60
link_max_duration_s: 0
max_active_workgroups: 5
working_after_s: 30
link.ask uses streaming internally even when the caller only needs the final reply. The target sends a start frame and periodic progress frames, so an active review can run longer than link_idle_timeout_s; the watchdog measures silence, not total duration. Set link_max_duration_s only when the operator wants a hard cap as well. A timed-out caller requests link.cancel and reports the reason explicitly instead of returning an empty transport error.
Pipeline admission is local to each hub profile. The queue limits whole active pipelines, not agent turns: admitted workgroups retain normal parallel phase execution, while excess launches and manual triggers wait FIFO on disk and survive daemon restarts. running and between-phase pipelines occupy a slot; completed or blocked pipelines release it. Both settings take effect after a configuration reload; a daemon that predates this feature still needs the updated process before it can enforce them.
Relay
Turns a profile into a read-only front door to one designated peer. When set, the engine offers the profile only the peer and decline tools and hard-gates every turn: the agent MUST consult that pinned peer via peer before it can produce a final answer — a call to any other peer id is rejected before it runs, an empty reply does not count, and if the turn ends (or hits the step/time limit) without a valid reply it fails closed with a fixed message rather than answer from the model's own knowledge. The peer's reply is surfaced as the answer. A request the relay must not forward (a change, an action, something out of scope) can be refused with decline(reason): the reason, written in the user's language, becomes the answer and the peer is never consulted; an empty reason does not count, and everything else still has to go through the peer. So you only pin that peer in peers.yaml with link.ask — no separate tools.deny needed.
This makes the relay side read-only, structurally. It does not make the target agent immutable: an inbound link.ask runs a full turn on the target with the target's own tools (fenced like a member device unless the peer has a tools.allow), so keeping the knowledge source unwritable is the target profile's responsibility — deny its mutating tools there for everyone with tools.deny, or list the only tools that relay may use with tools.allow on the relay's record in the target's peers.yaml (ALP.md → Per-peer tool policy); knowledge:search plus alpi_knowledge keeps a knowledge relay read-only, and restrict which paired devices may address it via a member connection's profile_scope (see Host below). The relay does not police the peer; the target polices the relay.
| Key | Default | Notes |
|---|---|---|
relay.peer | unset | The pinned peer_id this profile must consult before answering. Unset = no relay gate (normal profile). |
relay:
peer: agora
Ollama
Ollama is a first-class provider. One entry per server — local, remote, different ports — each with its own user-chosen name that becomes the model prefix (home/gemma4:e4b, gpu-box/qwen3:14b). On every request against an Ollama server, num_ctx is auto-resolved from /api/show and injected so the model sees the full prompt instead of being truncated to Ollama's 2K default.
providers:
ollama:
- name: home
url: http://localhost:11434
- name: gpu-box
url: http://192.168.1.50:11434
Add via alpi setup → Model → Add Ollama. Remove via alpi setup → Model → Remove keys.
The run ledger records what answered each turn. Agent turns stream, and litellm rebuilds stream chunks without OpenRouter's provider field, so those rows carry generation_id and a null provider; GET /api/v1/generation?id=<generation_id> returns provider_name, endpoint_id and the real total_cost after the fact. Non-streaming calls record provider directly.
MCP
| Key | Default | Notes |
|---|---|---|
mcp.servers | {} | Map of <name> → {command, args, env}. Each server is a local stdio subprocess the daemon spawns — alpi has no native HTTP/SSE MCP transport. Secrets in env use the env:VAR_NAME reference (resolved from the profile .env at spawn). Add via alpi setup → MCPs; hand-editing is supported. |
An entry is command + args + env. alpi launches it as a local subprocess and speaks MCP over stdio, so a remote HTTP endpoint must be bridged (below). Each env value of the form env:VAR_NAME is resolved from the profile .env at spawn and passed as the subprocess's environment — the secret lives in .env, never in config.yaml.
Pattern A — stdio server, secret via environment (server reads its credentials from env vars, e.g. a Bitbucket MCP):
mcp:
servers:
bitbucket:
command: npx
args: [-y, bitbucket-mcp@5.0.6]
env:
BITBUCKET_URL: env:BITBUCKET_URL
BITBUCKET_WORKSPACE: env:BITBUCKET_WORKSPACE
BITBUCKET_USERNAME: env:BITBUCKET_USERNAME
BITBUCKET_PASSWORD: env:BITBUCKET_PASSWORD
.env: BITBUCKET_PASSWORD=… (etc.). The server reads them from its environment.
Pattern B — remote HTTP endpoint, secret in an auth header. Because alpi is stdio-only, bridge the HTTP MCP with mcp-remote. The catch: env:VAR injects into the subprocess environment only, not into args, so a secret that must travel as an HTTP header cannot be referenced directly in the --header arg. Use mcp-remote's own ${VAR} expansion — it substitutes ${VAR} in a --header value from its environment, which you populate via the env: map:
mcp:
servers:
lobby:
command: npx
args:
- -y
- mcp-remote@latest
- https://api.example.com/mcp
- --transport
- http-only
- --header
- 'x-mcp-secret: ${MCP_SECRET}' # mcp-remote expands ${...} from its env
env:
MCP_SECRET: env:MCP_SECRET # alpi injects it from the profile .env
.env: MCP_SECRET=…. Do not hardcode the secret in --header, and do not wrap the command in sh -c to expand it — mcp-remote expands ${VAR} in headers itself.
Takes effect: MCP servers are spawned by the profile's engine; a config change is picked up when that profile's MCP subprocess is next (re)started — restart the daemon or re-bootstrap the profile.
Daemon capabilities
The one-per-machine daemon starts the scheduler, ALP listener, and workgroup poller for every profile, plus the host control plane for default. They are fixed internal tasks rather than configuration switches. Control behavior at the owning boundary instead: enable/remove jobs, pause/leave workgroups, grant/revoke peers, and scope/revoke client connections.
Legacy service.schedule, service.alp, service.workgroups, and service.host keys are ignored. The daemon logs a warning at startup and alpi doctor reports them until the block is removed; all capabilities still start. A legacy service.prefetch value migrates to runtime.prefetch on save.
host is meaningful only on the default profile; on any other profile the toggle is honoured but the runner refuses to bind a socket (the desktop / mobile client always targets default's socket and reaches sibling profiles via the profile parameter on each verb).
Host (control plane)
The host plane serves host.* verbs over a Unix socket (always) and a WebSocket on the shared network.host (see Network above); mobile / remote desktop use this path.
| Key | Default | Effect |
|---|---|---|
host.tcp_port | 49200 | WebSocket port for device pairing (the host plane's own port). |
host.device_name | "" | Optional pairing name shown in Devices. Empty = auto, otherwise embedded in the pairing QR and device list. |
host.endpoints | [] | Ordered explicit routes advertised in new pairing codes. Each row is {url, label} for wire compatibility. Desktop and setup manage one optional public wss:// route. When no explicit ws:// row exists, Alpi appends the private route derived from network.host and host.tcp_port; unsafe or unavailable addresses produce no private route. Explicit WS rows remain supported for compatibility. |
host.token_ttl_days | absent | Expire a device token after this many days without use. Absent, 0 or an unusable value (text, negative, boolean, .inf, .nan) means no expiry, today's behaviour; so does a config file the parser cannot read or decode, which never revokes anyone. Any number the parser does read is clamped to a century at most. Measured against the device's last_seen (its creation time if it never connected), so a device in regular use never expires; an expired device stops authenticating, its live WebSockets drop, and it must pair again. The judgement is made on read, so raising or removing the policy brings the device straight back. Re-read from disk on each request, no restart needed. |
host.allow_public_bind | false | Opt-in to let the shared network bind use a public IP. Affects both the host control plane and the ALP listener — both derive their bind from network.host, so without it neither binds TCP on a public address. A private or hostname address needs no opt-in; only a public IP does. |
host:
tcp_port: 49200
device_name: ""
endpoints:
- url: wss://your.domain.com
label: Secure Internet
- url: ws://100.64.10.2:49200
label: Direct
The address itself lives in network.host (shared with the ALP listener), not here. host.endpoints is advertisement only: it does not open listeners, terminate TLS, or change authentication. A wss:// route normally points at a TLS front-end — a reverse proxy such as Caddy, or a managed edge like a cloud load balancer / CDN — which forwards to the daemon's tcp_port. The Service UI shows that derived WS endpoint as Private route and the configured WSS endpoint as Public route. Removing the public route removes only WSS advertisement; private access continues whenever a safe private address exists. The listen port is editable in Desktop and alpi setup and requires a daemon restart. ALPI_HOST_TCP_PORT takes precedence over host.tcp_port; when set by a Docker deployment, change both that environment value and the matching 1:1 port mapping, then recreate the container. Multiple containers on one host use distinct effective ports such as 49200, 49201, and 49202. host.device_name controls the visible pairing label for new devices. It is optional; when empty, alpi falls back to the platform hostname. Plaintext hostname routes are rejected because DNS may resolve them to a public address (including alternate numeric IPv4 forms). Use a private IP literal for direct WS or a certificate-validated hostname with WSS.
WebSocket safety limits are daemon-wide environment settings. Defaults should fit Desktop and Mobile; changing them requires a daemon restart.
| Environment | Default | Effect |
|---|---|---|
ALPI_HOST_WS_MAX_CONNECTIONS | 128 | Maximum simultaneous WebSockets. |
ALPI_HOST_WS_MAX_CONNECTIONS_PER_DEVICE | 8 | Maximum sockets sharing one device credential. A socket over the limit first pings the device's existing ones: those that do not answer within 2 s (half-open after a network change or a VPN switch) are dropped and the new one is admitted; a device whose sockets all answer gets too-many-connections. |
ALPI_HOST_WS_MAX_RPCS_PER_DEVICE | 8 | Maximum concurrent RPC handlers or streams for one device. |
ALPI_HOST_WS_AUTH_FAILURES_PER_MINUTE | 10 | Authentication failures one source may accumulate per minute. A rejected token that names no device (unknown token) or a rejected pairing code counts against the source address, whose new sockets then close with 1013 before any token is read. A failure that names a device (revoked, expired or disabled) counts against that device instead, with the same limit: its sockets close with 1013 auth-rate-limited and the other devices behind the same address are not affected. |
ALPI_HOST_WS_TRUSTED_PROXIES | empty | Comma-separated IPs or CIDRs of reverse proxies whose X-Forwarded-For is trusted; the client is the rightmost hop not in this list. Empty means the header is ignored and the socket peer is the source. The WSS overlay pins Caddy to 172.30.250.10 and lists it. |
ALPI_HOST_WS_AUTH_TIMEOUT | 10 | Seconds allowed for the first authenticated request. |
ALPI_HOST_WS_AUTH_RECHECK | 1 | Seconds between active-socket authorization checks. |
ALPI_HOST_WS_CLOSE_TIMEOUT | 1 | Maximum graceful WebSocket close wait in seconds. |
ALPI_HOST_WS_REVOCATION_RETRY | 5 | Minimum seconds before retrying cancellation of a revoked stream. |
These can be set in the daemon process environment or the root ~/.alpi/.env. There is intentionally no global handshake-per-minute setting: before authentication, Alpi cannot distinguish a paired client from an attacker, and a shared budget would let cheap HTTP requests lock out legitimate devices. Per-IP limits belong in the public reverse proxy, firewall or WAF where the original client address is trustworthy.
The host plane lives with the default profile and its single socket serves every sibling profile. Admin connections plus direct local socket access reach any profile. To limit which profiles a paired member connection may address, scope that connection with profile_scope; daemon task configuration is not an access-control boundary.
On regular macOS/Linux installs, leaving network.host empty keeps auto mode: Tailscale first, then LAN, used as both the advertised address and the bind. Setting it advertises that address to clients/peers; the bind is derived separately — a private/Tailscale IP binds itself, a hostname or an opted-in public IP binds 0.0.0.0, and a public IP without host.allow_public_bind refuses to bind at all. In Docker the daemon binds 0.0.0.0 inside the container. A LAN or 100.x Tailscale IP in ALPI_NETWORK_HOST can produce the direct client route. A hostname, including MagicDNS, is still valid for ALP but desktop/mobile need an explicit wss:// entry in host.endpoints (see docker/README.md).
Connection identities and per-device WS credentials live at ~/.alpi/host/connections.yaml (mode 0600). Manage them through alpi setup → Connections. A connection owns its label, role and profile scope; every linked desktop/mobile device receives a different token and can be revoked independently. A new QR/link contains a one-time pairing grant, never that permanent token. The grant is stored only as a hash, expires after ten minutes and is consumed atomically by the first client. Desktop and Mobile still accept legacy token= links generated by older daemons.
On first startup after upgrading, an existing devices.yaml is migrated automatically: every legacy row becomes one connection with one device, its token stored as a SHA-256 digest and its access preserved, and the source is deleted once the destination has been re-read and verified. No backup copy is written. The old schema has no grouping key, so migration cannot safely merge rows that may belong to the same person or workload. Copies left by earlier releases stay untouched; alpi doctor lists them.
Takes-effect cheat sheet
- next turn — change is live on the agent's next response.
- next session — restart
alpito pick it up. - next daemon restart —
alpi daemon restart(or reload through launchd / systemd if installed as an autorun). - next daily sweep — the daemon's per-profile maintenance pass, first about two minutes after it starts and then every 24 hours.