A local-first terminal companion you can talk to — chat, voice, and 19 tools, on your hardware.
No cloud account required. No API bill by default. Your files, memory, and voice never leave your machine unless you hand it a key.
Slash autocomplete with fuzzy filtering, right in the prompt:
A grounded answer — real tools, real system data, streamed live:
Screenshots are real SVG captures of the app running headless (docs/shot.py), not mockups. A marketing/docs site lives in website/ (Vite + React + TS).
$ sk brief
╭─ sidekick brief Sat 2026-09-19 11:58 ─╮
│ CPU: AMD Ryzen 7 4800H (16 threads) │
│ Mem: 7.2Gi · GPU: GTX 1650 4GB │
│ /dev/nvme0n1p8 133G 117G 8.5G 94% / │
╰─────────────────────────────────────────╯
│ ! disk 94% full — clean ~/Downloads… │
$ sk run "what is the ideal llm i can run on my device"
• qwen3:4b (2.5 GB): fits comfortably in your 4096 MiB VRAM.
• llama3.2:3b (2.0 GB): another good option.
$ sk talk
[Enter] to record, [Enter] to stop. /quit exits.
heard> what files are in the sidekick repo| Sidekick | Typical cloud agent | |
|---|---|---|
| Runs fully offline (Ollama) | ✅ | ❌ |
| Voice input, transcribed on your CPU | ✅ | ❌ |
| Copy/paste that works in-terminal | ✅ drag-select, ctrl+y, /copy |
varies |
| Answers grounded in your system, not guessed | ✅ deterministic grounding | prompt-only |
Skills you can read (SKILL.md, incl. superpowers) |
✅ | varies |
| Offline test suite incl. prompt-regression evals | ✅ | rare |
Deny-by-default egress for what the model fetches (sk egress, logged) |
✅ | ❌ |
Audit ledger you can hand to an auditor (sk audit, credentials redacted) |
✅ | ❌ (their product is your code) |
| Costs $0 by default, spend caps when you bring keys | ✅ | metered |
uv tool install sidekick-agent[voice] # global `sk`, STT included
sk init # guided first-run: hardware → model → verify
sk # fullscreen chat — start here (`sk tui` works too)No clone, no build — installs straight from PyPI. Requires Python 3.12+.
Without [voice] you get everything except Talk/mic (installs on first use instead). Local path needs Ollama (ollama serve, pull qwen2.5-coder:7b for smarts or llama3.2:3b for speed).
| Channel | Command |
|---|---|
| PyPI / uv | uv tool install sidekick-agent[voice] |
| PyPI / pipx | pipx install sidekick-agent[voice] |
| PyPI / pip | pip install sidekick-agent[voice] |
| AUR (Arch) | yay -S python-sidekick-agent |
| conda-forge | conda install -c conda-forge sidekick-agent (feedstock lives in a separate repo) |
The published name is sidekick-agent (the sidekick name is taken on
PyPI); the command stays sk. Version is a single source of truth in
src/sk/__init__.py. Publishing is automatic and credential-free: when a PR
is merged to main of the canonical repo
Faisal-Fayaz/sidekick with a bumped
__version__, GitHub Actions trusted-publishes to PyPI and opens a GitHub
Release (forks can never publish) — details in packaging/README.md.
From source (dev):
git clone https://github.com/Faisal-Fayaz/sidekick && cd sidekick
uv tool install -e ".[voice]" # editable dev install; STT included
sk doctorOne input, two surfaces — fullscreen TUI and plain-text REPL share every command:
sk # fullscreen chat with streaming + themes — start here
sk tui --model fast # same, explicit form
sk chat # fallback REPL: dumb terminals, screen readers, broken TUIsType / and an autocomplete popup filters all 20+ commands — Enter completes, Tab too, Esc dismisses, ↑/↓ navigates. F1 opens a generated cheatsheet (keys + commands, built from the same tables as the dispatcher, so it can't rot).
TUI keys: Enter sends · ctrl+j/alt+enter newline · ↑/↓ history · ctrl+y copies · ctrl+g push-to-talk · pgup/pgdn scroll · F1 help · F2 dark/light theme · F3 sessions drawer · F4 plan/build toggle · F5 model picker (filter + live provider list). Answers stream live as Markdown with role colors; approvals arrive as cards with timeout; the status bar shows model · session · last-turn time/tokens.
Plan mode proposes without writing: /plan blocks file writes in chat (/build reverts), F4 toggles the same in the TUI, and sk run --plan does it single-shot. Sessions carry /compact [hint] (summarize history on demand), /diff and /review (inspect the working tree), plus /rewind [n] to undo agent file edits.
sk talk [-d SECS] [--stt-model base] [--device hw:2,0] # Enter records, Enter stops
sk mic-test # peak dB + silent/quiet/good verdictCapture via the OS-native recorder (arecord/ALSA on Linux, sox/ffmpeg on macOS), transcription via local faster-whisper int8, transcript lands editable in the prompt. In the TUI, ctrl+g (or the mic pill) does the same. Voice never leaves your machine; recordings are temp files, deleted after each take.
sk mcp [--allow-writes] # JSON-RPC 2.0 over stdio, zero new depsAll 19 tools, same safety policy (SSRF guards, write blocklists, hard-refusals). Reads auto-run; shell/writes/delete need --allow-writes, else a clean denied error. Stdout carries protocol only. Claude Desktop snippet:
{ "mcpServers": { "sidekick": { "command": "sk", "args": ["mcp"] } } }sk connect # pick provider → paste key (hidden) → pick model → ping. Done.One guided flow: numbered provider list (local ones skip keys), live validation before anything saves, curated model list (TTS/image junk filtered, recommended pre-highlighted, Enter accepts), and a 5-token ping instead of a full agent turn. Advanced paths still work: sk auth add/list/status/remove, sk model, sk setup (connect + hook), sk config --provider openai --api-key sk-..., /provider groq inside chat.
Presets: ollama|openai|groq|together|deepseek|openrouter|google|lmstudio|anthropic|opencode|custom (anthropic speaks the native Messages API; the rest are OpenAI-compatible). Any OpenAI-compatible endpoint works via --provider custom --base-url https://... --api-key .... Preferred: SIDEKICK_API_KEY env (never touches disk); file keys are chmod 600 and masked in --show. The Anthropic backend marks the static system prompt + tool definitions cacheable (repeat turns up to 10x cheaper); OpenAI-compatible providers cache matching prefixes automatically server-side. Reasoning effort: sk config --reasoning-effort low (default; off|minimal|low|medium|high|max), escalated to high in plan mode. Honored on OpenRouter (reasoning.effort), OpenAI (reasoning_effort), Anthropic (thinking budget sized to fit max_tokens), and Ollama (think toggle — off disables thinking, high/max enables it).
The opencode preset points at OpenCode Zen, opencode's gateway with a set of free, tools-capable models (config.OPENCODE_FREE_MODELS) plus paid tiers. The anonymous free tier is restricted by opencode to its own app, so add a free OpenCode account key first — OPENCODE_API_KEY env or sk auth add opencode (get it at opencode.ai/auth). Fast/smart resolve to nemotron-3.5-lightning-free / muse-spark-1.3-contributor-free; run sk models opencode for the live list.
| Command | What |
|---|---|
sk / sk tui [--continue] [--model M] [--allow LIST] [--deny LIST] |
Fullscreen chat, fresh session each launch. sk --version prints the version and exits; --cwd DIR and --profile NAME work on every command |
| `sk chat [--continue] [--session S] [-y | --yes] [--model M] [--allow LIST] [--deny LIST] [--no-stream]` |
sk talk [-d SECS] [--stt-model base] [--device hw:2,0] [--session S] [-y] [--install] |
Push-to-talk voice chat (CPU transcription, Enter to record/stop; --install skips the install prompt) |
sk mic-test [-d SECS] [--device hw:2,0] |
Mic level check: peak dB + silent/quiet/good verdict |
/sessions, /resume <n>, /sessions delete <n>, /fork [n] |
List, switch, delete, branch past sessions |
/yolo, /confirm, /readonly |
Approval modes: auto-approve writes / ask every time / block all file writes (research mode) |
/compact [hint], /diff, /review |
Summarize history on demand / inspect working-tree diff / review it from inside a session |
/plan, /build |
Plan mode: propose without writing (blocked writes) / back to build mode |
/rewind [n] |
Undo an agent file edit — snapshots write/edit/delete targets (shell mutations are not tracked, use git for those) |
sk run "task" [--yes] [--plan] [--read-only] [--deny LIST] [--model auto|fast|smart|name] [--json] [--bg] [--allow LIST] [--session S] [--no-stream] |
Single-shot agent run (auto-router picks the model; --json emits one machine-readable document + exit codes, use with --yes unattended; --bg detaches, returns a job id, notifies on completion; --allow shell:pytest,write_file skips prompts for listed tools; --plan proposes without writing) |
sk jobs [-n N] |
List background jobs from sk run --bg |
sk brief [-p PATH] [--smart] |
Morning digest: system + git + todos + memories, instant without LLM |
sk digest [--force] |
Teammate pilot: brief + overnight failures, desktop nudge or log |
sk remember/recall/memories/forget |
Long-term memory (FTS5 search, auto-injected) |
sk search QUERY [--session S] |
Full-text search across past transcripts |
sk todo add/list/done/clear |
Todos |
sk history / sk oops |
Shell log / explain last failure |
sk imagine "prompt" [--out f.png] [--size WxH] [--model M] |
Generate an image via the provider images endpoint |
sk export [SESSION] [--out f.md] [--force] |
Session transcript as Markdown (turns + tool calls). Refuses to overwrite an existing --out unless --force |
sk egress [list|allow HOST|deny HOST|test URL] |
Egress policy: what the model may fetch, and what it was blocked from (docs) |
sk audit [--session S] [--format md|json] [-n N] |
Compliance log: tool runs, approve/deny, local-vs-egress (-n clamped 1–1000) |
sk stats [--session S] [--format md|json] |
Usage + cost estimates from audit rows (turns, tools, tokens); pair with spend caps for BYO-key budgets |
sk hook-install [--write] |
Bash/zsh logging hook |
sk hooks [--check] |
List event hooks + live dry-run (see docs/hooks.md) |
sk skills / sk skills-search / sk skills-registry [QUERY] / sk skills-install NAME / sk plugins / sk daemon [--once] / sk daemon-install [--schedule TXT] / sk daemon-install-macos [--schedule TXT] / sk daemon-uninstall (removes the launchd agent; on Linux disable the systemd unit) / sk daemon-schedule [--set TXT] |
Skill packs (registry + superpowers) / user-defined tools (TOOLS.md, see docs/plugins.md) / background watcher (systemd/launchd, calendar schedules) |
sk mcp [--allow-writes] |
MCP server over stdio (19 tools, safe defaults) |
sk mcp-servers |
List configured MCP client servers + live tool check (see docs/mcp-client.md) |
sk doctor [--fix] / sk report / sk models [list | pull <id> | prune <id>] / sk config / sk version / sk upgrade [--check] |
Health (+auto-remediation) / redacted diagnostics bundle / models (list, download, remove) / settings / build / self-update |
sk init / sk setup / sk connect |
Guided first-run / full setup / provider key flow |
sk model / sk auth add/list/status/remove |
Ask the provider for its live model list and set the default / manage provider keys (masked, validated live) |
Packs use the SKILL.md frontmatter format. The prompt carries a relevance-ranked index; the agent loads full instructions on demand via the skill tool. fast/smart resolve per provider (Ollama: llama3.2:3b/qwen2.5-coder:7b, Groq: gpt-oss-20b/120b).
flowchart TB
U([you]) --> CLI[sk / sk run]
U --> TUI[sk tui: autocomplete, streaming, mic pill]
U --> VOICE[sk talk: arecord + faster-whisper]
CLI --> SLASH[slash.py: /commands, no LLM]
TUI --> SLASH
VOICE --> AGENT
CLI --> AGENT[agent.py: stream → tools → synthesize]
TUI --> AGENT
AGENT --> GROUND[deterministic grounding: ~/paths, URLs,\nsysinfo — injected before the model sees the prompt]
AGENT --> TOOLS[tools/: 19 tools, allowlists,\nhard-blocks, SSRF guard]
AGENT --> MEM[(store.py: history, memories FTS5,\ntodos, shell log)]
AGENT --> SKILLS[skills: relevance-ranked SKILL.md index]
Design bets that paid off: deterministic grounding beats prompt instructions (small models ignore rules but can't argue with injected facts), text-JSON fallback (coders emit tools as text over the OpenAI endpoint), FTS5 over vectors (zero deps, instant, no embedding server on a 4GB box), parallel reads (approval-gated tools stay serial; independent reads run concurrently with failures isolated), capability profiles over hope (known models declare their tool protocol; the router reads them instead of paying a probing 400).
Reads auto-run. Writes, deletes, and general shell need approval (inline [y/N] in TUI, prompt in CLI), HOME//tmp only, ≤100KB, never ~/.ssh, ~/.gnupg, /etc, /usr. Multi-tool turns with destructive actions get one plan review up front instead of per-tool prompts (silent in --yes//yolo; denials execute nothing). shell hard-refuses rm -rf /, mkfs, dd to devices, fork bombs even with approval. read_url/web_search block localhost/private IPs.
Egress is deny-by-default. Sidekick will not fetch a destination the model asked for until you allow that host:
sk egress # what is allowed, and what was blocked
sk egress allow example.com # allow one host
sk egress allow '*.example.com' # or a whole domain
sk egress deny example.com # revoke
sk egress test https://example.com/x # check a URL without fetching itThis covers every destination the model can cause a fetch of: read_url,
web_search, and provider-supplied image URLs. With an empty allowlist, all
three are refused. The allowlist is checked on the initial URL and on every
redirect hop, so an allowlisted host cannot redirect the agent somewhere
unlisted. web_search itself only requests html.duckduckgo.com and returns
result URLs as text — it never fetches a hit, so following one is read_url,
which is governed on its own.
Your model API endpoint and the MCP servers you configured are unaffected — you
chose those, they are not something a prompt can talk the agent into adding.
Fetch decisions are written to the ledger with the host and reason, visible via
sk egress and sk audit. (Adding or removing an allowlist entry is a config
change, not a fetch, so it is not itself a ledger row.)
The ceiling, stated plainly: this governs fetches made through the agent's
tools. A shell command running curl is not covered — no string denylist can
be a network boundary. That is why shell still needs approval, and why
issue #329 (OS-level
sandboxing) is the real answer rather than another pattern list. See
docs/egress.md. API keys chmod 600, masked in output.
Your data is private to your account. ~/.sidekick is 0700 and everything inside it is 0600 — the full transcript, shell commands the agent ran, checkpoints, and the last 100 prompts. It was created with the umask default before, which left history.db world-readable and two files world-writable; an existing install is tightened the first time you run any sk command. This is a local filesystem control: it does not protect against root, and it does not stop anything the agent itself can read.
uv run --python 3.12 --with ".[test]" pytest tests -q # unit + regression + Textual pilot, no Ollama neededThe eval harness (tests/test_eval.py) locks in every past quality bug as an offline regression test. A suite-wide fixture guarantees tests never touch your live ~/.sidekick/.
~/.sidekick/config.toml (provider, model, base_url override, api_key, …). Data stays home: history.db, skills/, nudges.log, input_history, tui-errors.log.
| Env var | Effect |
|---|---|
SIDEKICK_PROVIDER / SIDEKICK_MODEL |
override the configured provider/model |
SIDEKICK_BASE_URL / SIDEKICK_API_KEY |
override the endpoint / key for this process |
SIDEKICK_SPEND_CAP |
per-session USD cap, 0 = unlimited |
SIDEKICK_REASONING_EFFORT |
reasoning effort for backends that support it |
SIDEKICK_PROFILE |
load a named config profile (see sk config --profiles) |
OPENCODE_API_KEY |
key for the opencode provider |
SIDEKICK_TRUST_REPO=1 |
opt in to repo-supplied slash commands |
SIDEKICK_TRUST_REPO— why your repo's commands do nothing. A cloned repo ships its own.sidekick/commands/*.md, and a!cmd`` expansion inside one reachessh -cwith no approval prompt. Those templates are therefore ignored unless you set `SIDEKICK_TRUST_REPO=1` in your own shell. It is read from the environment and never from repo config, so a repository cannot self-authorize (#287). Commands in `~/.sidekick/commands/` are always loaded.
History budget: the prompt budget bounds the whole assembled prompt, not just history — system prompt, auto-context, tool schemas, and a reservation for the reply are all measured and subtracted, then history gets what is left (#308). history_budget_tokens (default 3000) caps that share; over-budget sessions compact to a rolling summary via the current model (DB history stays complete). Lower it for small-context models.
Context window: context_window (0 = take it from the model profile) sets the total window, and the same number is sent to Ollama as num_ctx, so the budget and the server cannot disagree. Local models default to 16k, frontier to 200k. The fixed cost alone — system prompt plus tool schemas — measures ~3,700 tokens, so a window below ~8k leaves no usable conversation. Raise it if you have VRAM to spare; the KV cache grows with it. Run /context to see where a turn's tokens actually go.
Per-project config: a .sidekick.toml in any repo layers over the global file (nearest one walking up from cwd). It may carry a [project] table only:
[project]
docs = ["CONTRIBUTING.md", "docs/architecture.md"] # injected into the prompt
memory_namespace = "my-repo" # scopes this repo's memoryNothing at the top level of a project file is honoured — not provider, model, max_steps, temperature, egress_allow, or anything else. That is deliberate, not a gap: a repository you just cloned must not be able to steer the agent's model, budget, approvals, or network, so those keys live in ~/.sidekick/config.toml or the environment only. The same applies to [project].approved_commands (shell approvals stay global).
Every ignored key is reported rather than dropped quietly — a reduced max_steps in a project file is a refusal, not tuning, and it should not look otherwise:
$ sk config --show
provider=ollama
project: /home/me/repo/.sidekick.toml
ignoring max_steps in .sidekick.toml (global config or env only)
ignoring max_step in .sidekick.toml (not a project-file setting; …)sk --cwd PATH runs any command as if in that directory.
Project memory auto-discovery: from the cwd upward, the first SIDEKICK.md / AGENTS.md / CLAUDE.md / GEMINI.md found (root-down) is injected into the prompt automatically — no config needed for repos that already document themselves.
Config profiles: named ~/.sidekick/profiles/<name>.toml files (e.g. work, personal, low-VRAM) replace config.toml when selected via sk --profile NAME … or SIDEKICK_PROFILE=NAME (env still wins, project files still layer). Manage with sk config --profiles (list), sk config --save-profile NAME (snapshot current non-secret settings — keys stay in keyring/env). Migrating: cp ~/.sidekick/config.toml ~/.sidekick/profiles/work.toml, trim it, and switch with the flag. sk config --show prints the active profile.
See ROADMAP.md — the shared plan (vision, next milestone, done list). It changes by pull request only.
MIT — do what you want, shout-outs appreciated.