Docs/Security

Security

Two-layer security model. Approval system, SSRF, prompt-injection, sensitive paths. Sandbox.

Run it·20 min·v0.17.8
On this page
  1. Layer 1 — application guards (always on)
  2. Layer 2 — OS sandbox (per profile, opt-in — with one exception)
  3. Threat model
  4. Closed system prompt (by construction)
  5. Third-party code
  6. Security posture audit
  7. Inline image reads (host plane)
  8. Audit trail & accountability
  9. Known gaps

alpi is published by Satoshi Ltd., whose three load-bearing principles for this document are:

Security First — threat-modeled from initial development; no surveillance disguised as telemetry. Privacy by Design — privacy is the foundation, not a feature. Zero Knowledge — what we don't know can't be subpoenaed, leaked, or sold.

Those are the frame every decision below lives inside: the guard is mandatory and local, the sandbox is an opt-in second wall, the LLM is treated as an adversary with user credentials, and we keep as little state off your machine as we can.

alpi runs LLM-decided tool calls on your machine. The security posture is layered — application-level guards that always run, plus an optional OS-level sandbox for shell commands.

Layer 1 — application guards (always on)

These live inside the Python process and cannot be disabled without editing source. They cover the attack vectors an OS sandbox around terminal does not reach.

Connections carry a role. admin has full control-plane CRUD; member gets chat, events, read-only views, workgroup post/read and deletion of the chats its own connection created, and every sensitive mutation answers -32001 forbidden. The local socket is sovereign and always admin. What member does not restrict is the agent: a member device can still send chat turns, and a turn can do anything the profile's tools can do. Bound agent capability with the OS sandbox and a dedicated profile, not with the device role. Profiles in one daemon share one OS trust boundary; mutually untrusted customers need separate runtimes — see Deployments. Profile file reads over the host plane carry their own deny list (any secrets component, the host/, gateway/ and cache/ trees, any .env*, key extensions, symlinks into denied trees, path escapes); secrets surface only through dedicated methods.

Layer 2 — OS sandbox (per profile, opt-in — with one exception)

Wraps terminal subprocess calls in a native OS sandbox so the kernel refuses the syscalls, not just the detector above. Persistent writes are confined to workspace + ~/.alpi/ + the system temporary trees (/tmp, plus macOS-specific /private/tmp and /private/var/folders); a small set of character devices that well-behaved CLI tools reopen (/dev/null, /dev/{u,}random, /dev/tty, std streams) is also writable but they are not persistent storage. Read posture is platform-specific: Linux/bubblewrap only makes explicitly-mounted paths readable — workspace and profile bind-mounted writable, runtime system paths (/usr, /bin, the loader and libraries the process needs, and only the parts of /etc a CLI needs to resolve names and validate certificates) mounted read-only, /tmp as an in-sandbox tmpfs — so anything not mounted is invisible. macOS/sandbox-exec runs default-allow for reads with a small explicit deny list (~/.ssh, ~/.aws, ~/.gnupg, profile .env, skill secrets/), so anything outside those denies stays readable. Network is denied by default.

Status: stable, opt-in — with one exception. A dispatched workgroup phase that declares write scopes forces terminal through the sandbox on bare metal even when the profile has it off, so a phase owner cannot write outside its lane. In the official Docker runtime a scoped phase omits terminal altogether rather than sandboxing it (see Workgroups). Everywhere else it defaults to off, because real-world dev workflows vary too much to pick a profile that never breaks: git push over SSH relies on ~/.ssh, Apple Silicon Homebrew lives in /opt/homebrew, docker needs /var/run/docker.sock, npm wants ~/.npm. For interactive chat where you approve every command, the Layer 1 denylist is already sufficient.

Where it really earns its keep: unattended profiles. The alpi daemon (scheduler task), research / delegate sub-agents — these run without a human approving each command. A prompt-injected email or a hallucinating sub-agent can issue rm -rf ~/anything with no veto. Layer 2 is the kernel-level veto you want there.

alpi's multi-profile CLI makes this ergonomic:

Each profile has its own ~/.alpi/profiles/<name>/config.yaml, so the sandbox flag is set independently.

Enabling

Interactive: alpi setup → Sandbox → toggle on/off + network.

YAML (direct): set in ~/.alpi/profiles/<name>/config.yaml:

tools:
  terminal:
    sandbox: true
    allow_network: false   # flip to true if the profile needs git push / npm install

TUI feedback

The top bar carries a muted sandbox segment next to the workspace while the sandbox is on, reading offline instead when the network is also locked. There is no segment when the sandbox is off: absence is the off state, so read the bar for the badge, not for a word saying "off".

Platform support

macOS — uses native sandbox-exec (ships with the OS at /usr/bin/sandbox-exec). No install step.

Linux — uses bubblewrap. Install once:

Requires user namespaces enabled in the kernel (default on modern distros; some hardened configs disable them).

Windows — no native sandbox path. Two options:

  1. WSL2 (recommended): wsl --install, then run alpi inside Ubuntu as if it were Linux native. bubblewrap works there.
  2. Native Windows: leave tools.terminal.sandbox: false. Layer 1 stays active; you lose the kernel-level guarantee for shell commands.

What happens when the sandbox is on

Testing the Linux path from macOS

A minimal Docker image covers the Linux code path. See docs/sandbox-linux-test.md.

Threat model

alpi's realistic attacker:

Layer 1 covers the common-case attacks (known patterns, known sensitive paths, known SSRF targets). Layer 2 adds defense-in-depth so a creative prompt that bypasses the regex still can't touch the FS or the network.

Closed system prompt (by construction)

alpi's system prompt is assembled from three narrow, controlled sources — nothing else. There is no auto-load of workspace files like AGENTS.md, .alpi.md, CLAUDE.md, or similar "bring your own context" conventions. The build in engine.py::_build_system_prompt concatenates, in order:

  1. alpi/prompts/system_prompt.md — shipped in the package; authored by us, updated with each release.
  2. Memory (USER.md, MEMORY.md, AGENT.md) from ~/.alpi/profiles/<name>/memories/ — written by the LLM itself through the memory tool, with dedup + char limits + cross-file duplicate detection.
  3. The skills index from ~/.alpi/skills/**/SKILL.md — every mutation passes through the shared scanner (_DANGER_PATTERNS in alpi/scan.py), which scans for dangerous patterns (rm -rf, curl|sh, eval(), hardcoded keys).

Workspace files — anything the user has on disk — are data, not context. The LLM reads them through the read_file tool, which labels the result as a tool response (the model is trained to treat tool output as untrusted). The usual prompt-injection warnings in system_prompt.md cover this path.

This is a deliberate departure from agents that honour convention-over-configuration context files. Those files are raw Markdown loaded before the turn starts — a documented attack vector (an attacker who can write a .agent.md to a repo you clone can steer your next turn). alpi trades the ergonomic convention for a smaller trusted-input surface. If a project needs its conventions taught to the agent, put them in a skill or in USER.md; both paths pass through explicit user approval.

Third-party code

Every runtime dependency is an attack surface. We keep the list tight (see ARCHITECTURE.md → Dependencies for why each one earns its place) and audit it before each release. The CVE pass is a single command:

uv run --with pip-audit pip-audit

Risk profile of the runtime set:

DepRiskNotes
litellmMediumLarge surface (100+ providers). Ships with telemetry=True by default — alpi flips it off in llm.py::_silence_litellm() so no request phones home. Regression test: tests/test_llm_privacy.py.
playwrightMedium-highRuns a full Chromium (~230 MB) that loads arbitrary web content. Chromium's own sandbox is the line of defence at that layer; alpi adds nothing on top. Used only by the browser tool.
playwright-stealthLowSmall patch set on navigator.webdriver and friends. Reverse-engineered detection bypass; breaks occasionally when detection vendors tighten.
pillowMediumImage parsers have a long history of CVEs. Keep on the latest minor; pip-audit catches known issues.
faster-whisperLowBundles CTranslate2 native code. Models are downloaded from HuggingFace on first use — inspect the model hash if paranoia calls for it.
edge-ttsLowReverse-engineered unofficial Microsoft Edge TTS endpoint. Small code, but the endpoint can change; have a plan B (say on macOS, espeak on Linux) ready.
textualLowPure Python, active, stable API surface we pin to.
litellm's transitive tree (openai SDK, anthropic SDK, etc.)Low-mediumFlows through. pip-audit covers.
httpx, rich, click, pyyaml, python-dotenv, prompt_toolkit, croniter, html2text, ddgsLowSmall or stable or both. Rarely updated, rarely break.

Policy

Security posture audit

alpi audit is the read-only posture scan for an installed machine. It is different from alpi doctor: doctor asks "is the active profile healthy and reachable right now?", while audit asks "is this whole install hardened enough to leave unattended?".

The command scans the entire ~/.alpi install, not just the selected profile:

alpi audit           # includes OSV CVE lookup when network is available
alpi audit --offline # local-only: permissions, binds, hardening

Checks today:

Exit code is 1 only when a fail is present. Warnings are visible but do not break cron or release scripts. The command never changes permissions, writes config, upgrades packages, or phones home unless the user explicitly runs the online CVE check by omitting --offline.

Inline image reads (host plane)

Agent-made images render inline in chat across clients. The image bytes are read by path, scoped to a fixed root set: the active profile's workspace, its home (~/.alpi/...), and temp dirs. Same roots on every client:

Implication: a client authorised for a profile can fetch any image under those roots by path — broader than "an image that appeared in this chat". This is intentional (it's what inline rendering needs and the device is already trusted for the profile), but it is a real read surface. A future tightening would restrict reads to paths that appear in the session transcript or an output manifest; not implemented today.

Documents (every attachment that is not an image) are narrower: the daemon serves them only from the profile's out/, its workspace and the upload staging area. The engine offers a produced file as an attachment only under the roots the daemon serves for its kind (servable_roots in alpi/attachments.py), so a document in /tmp or elsewhere in the home is no longer offered, and attach_file refuses it, naming the out/ folder to use. The daemon still refuses, on top, any path with a secrets folder or a .env file name.

Inbound, a remote device attaches only uploaded files: host.chat.send from a remote connection (member or admin) accepts an attachment path only when it resolves, symlinks followed, to a file in that profile's staging area (host/attachments/tmp/, filled by host.attachments.stage and shared by every device of the profile), and refuses the whole message otherwise. The local socket and host.chat.delegate still pass local paths, which is how the desktop on the same machine attaches a file.

Audit trail & accountability

alpi records what the agent and its operators do across several local surfaces. The posture is local-grade: administrative host-plane actions are attributable to a connection and device, but the files remain locally mutable and are not a tamper-evident or external compliance trail. What exists today:

What is NOT covered today (and why it matters for a fleet, not a single user):

Closing these is an explicit roadmap item — see AUDIT.2 in ROADMAP.md. It is deliberately not built into the personal product until a real fleet deployment pulls for it.

Known gaps