You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The CLI module (kiro_crew/cli.py) provides the kirocrew command using stdlib argparse.
Member memory commands
kirocrew agent create --name <name> explicitly creates a stable member identity
and its empty member-scoped V2 SQLite database. A member cannot share or rebind its
store through a template, display-name change or arbitrary database path. This
development-stage V2 contract has no old V2 data migration or compatibility mode.
Global and named V1 memory retain their existing behavior independently.
kirocrew doctor reports the configured member, store and integrity failures
without provisioning, repairing or replacing databases. A missing, corrupt or
wrong-member database requires explicit recovery; ordinary opening never creates
an empty replacement or falls back to Global. SQLite reads can create transient
WAL coordination files without changing learning records. See the
memory contract.
Import Weight Contract
cli.py is the shared dispatcher for every subcommand — including the
long-lived MCP stdio servers (kirocrew mcp-core / mcp-cron /
mcp-computer), which hold its module-scope imports resident for their whole
lifetime. Its module scope therefore stays light: cli_commands,
cli_server (which pulls slack.gateway) and dashboard.state (which pulls
vector_memory → numpy) are imported inside the one main() dispatch
branch that uses each name, never at module scope. Deferring them cuts a
fresh import kiro_crew.cli from ~1.3 s / ~112 MB to ~0.5 s / ~54 MB, paid
per CLI invocation and per MCP backend process.
test/test_cli_lazy_imports.py ratchets the contract: after
import kiro_crew.cli in a fresh interpreter, none of those modules may be
present in sys.modules, and every deferred dispatch import must resolve.
Bare kirocrew --version (exactly one argv entry) is answered by pre-dispatch
fast-path guards — _bootstrap.main (which then skips importing cli on the
console-script path) and __main__ (same for the python -m path, which also
covers the desktop launchers) — printing kirocrew <version> and exiting 0.
It is the one invocation that never reaches cli.main(): no platform boot, no
sandbox hygiene, no project-dir resolution, no logging or data-home setup, and
no editable-install self-heal (a broken checkout answers --version instead
of healing or exiting 1).
Only the bare form fast-paths — artifact show --version N reuses the flag
name for an artifact version int, so every other argv shape falls through to
argparse unchanged.
The entry point itself is not negotiable: every command invocation
(kirocrew <sub>, python -m kiro_crew <sub>, and the frozen desktop
binary) lands in cli.main(), whose prelude runs boot_platform()
(fail-closed for non-standalone profiles), sandbox env hygiene, and
KIROCREW_PROJECT_DIR resolution before any dispatch.
Gateway mise environment
Before starting services, cli_server._gateway() calls env.activate_mise().
It loads the user's global mise environment even when a daemon starts without
shell startup files. Binary lookup prefers the inherited PATH, then
~/.local/bin/mise. On macOS it next checks /opt/homebrew/bin/mise and
/usr/local/bin/mise, in that order. Each fallback must resolve to an executable
file; the Homebrew paths are not probed on other platforms. No shell is started,
and locating mise does not add any directory to the gateway's PATH.
The selected binary runs mise env --json from the user's home with a ten-second
timeout. Returned string values are merged into the gateway environment for
later child processes. KIROCREW_NO_MISE skips discovery and activation. Missing
mise or a failed invocation remains non-fatal. Registered MCP search directories
remain separate from the runtime PATH; this lookup does not change that contract.
Source Checkout Launcher
The POSIX wrapper at bin/kirocrew resolves symlinks to find the real checkout,
sets KIROCREW_PROJECT_DIR to that checkout unless the caller already supplied
one, and delegates every argument to .venv/bin/kirocrew. The virtualenv entry
point comes from the editable install created by the setup scripts, so it makes
src/kiro_crew importable without adding the source tree to PYTHONPATH. Any
caller-provided PYTHONPATH is inherited unchanged.
If .venv/bin/kirocrew is unavailable, the wrapper exits with source-install
guidance instead of falling through to a different Python environment.
Standalone Wheel Installer Trust Contract
cli.sh installs channel or pinned-version wheels only from an authenticated
manifest. This distribution trust boundary is independent of the runtime CLI
and of macOS signing/notarization.
Schema: kirocrew-cli-artifact-manifest-v1.
Algorithm: RSASSA_PKCS1_V1_5_SHA_256.
Key identity: sha256: plus the lowercase SHA-256 digest of the public
SubjectPublicKeyInfo DER bytes.
Signed fields: algorithm, channel, key_id, pub_date,
python_requires, schema, sha256, version, and wheel_url.
Optional signed field: min_version — the forced-update floor for a
breaking release, declared in packaging/MIN_VERSION at publish time. Must
be a bare release (0.6.0, no prerelease suffix) and must not exceed the
manifest's own version; both are enforced by
packaging/signing/cli-manifest.py before signing. The installer verifies
its format but does not act on it (it always installs the signed version);
running gateways read it from the channel feed and mark the update
REQUIRED when their own version sits below the floor — but only after
platform/feed_trust.py verifies the manifest signature against the same
pinned key, because the floor coerces the dashboard UI and an unverified
one must degrade to the ordinary dismissible prompt. That module's
verification core is shared: the hosted feature-video manifest is signed by
the SAME offline key and verified through verify_document_signature, which
differs only in allowing a nested payload and its own size cap. One trust
root, and the two documents stay non-interchangeable because each carries a
distinct schema inside the signed payload that its consumer requires. Absent means no
floor, and the canonical payload omits the key entirely so no-floor
manifests stay byte-identical to the pre-floor format.
Signature field: base64 RSA signature over sorted, compact UTF-8 JSON of all
signed fields; signature itself is excluded.
Channel source: feed/<channel>/latest-cli.json. Pinned-version source:
cli/<channel>/<version>/cli-manifest.json; pinned installs do not resolve
through the mutable channel feed.
The installer embeds the public key and expected key id. Before any network
request, it requires OpenSSL, rejects an unconfigured pin, materializes the key,
and verifies its DER fingerprint. It then applies bounded input/object sizes,
duplicate-key and exact-field-set rejection, printable-ASCII/string checks,
canonical URL and digest validation, requested channel/version matching, and
pinned-key signature verification. Artifact fields are not consumed and wheel
bytes are not fetched until the signature succeeds. The downloaded wheel must
then match the authenticated SHA-256 digest before pipx install runs.
Any unavailable trust root, malformed or unsigned legacy feed, unknown field,
wrong schema/algorithm/key id, signature failure, metadata mismatch, network
failure, or wheel digest mismatch terminates installation. There is no unsigned,
SHA256SUMS, or trust-on-first-use fallback. Until the operational public key is
pinned, the repository's explicit UNCONFIGURED state therefore makes stock
cli.sh non-installing by design. Provisioning and rollout are specified in
packaging/signing/README.md.
The publication gate in .github/workflows/publish-installer.yml checks every
live channel feed with packaging/signing/cli-manifest.py verify before it
replaces the live cli.sh. That is a second implementation of the contract
above, so what holds between them is a direction rather than equality:
whatever the gate accepts, the installer must also accept. The gate is
deliberately stricter in three places — a 16 KiB payload cap against the
installer's 64 KiB, a 2048-character cap per field, and refusing a
min_version above the version the manifest ships — because rejecting a feed
the installer would have taken costs a publisher one loud failure, while the
opposite direction publishes a feed that bricks installs and reports success
while doing it. Both sides therefore normalize the artifact base identically,
stripping at most ONE trailing slash (${ARTIFACT_BASE%/} in cli.sh), and the
gate refuses a base with repeated trailing slashes instead of normalizing to a
URL no installer reproduces. test/test_cli_manifest_signature.py holds the
direction by driving one shared fixture set — valid, wrong-channel, wrong-host,
tampered, legacy — through the gate and through a real cli.sh run.
Project Directory Detection
At startup, main() auto-detects the project root and sets KIROCREW_PROJECT_DIR:
If KIROCREW_PROJECT_DIR env var is already set, use it
Walk up from CWD looking for a directory with both skills/ and src/kiro_crew/ (_PROJECT_MARKERS). The project-level agents/ dir was removed when agent config was consolidated into src/kiro_crew/config/ (commit bbbc1f6e), so the marker no longer references it — a stale agents/ requirement left detection (and the dashboard changelog) silently broken.
Read saved path from ~/.kiro/crew/project_dir (written by kirocrew setup); the saved path is re-validated against the same markers
This allows kirocrew to find project-level agent config and skills from any directory.
Commands
Top-level help
kirocrew --help (and a bare kirocrew, which prints the banner first) does NOT
use argparse's own subcommand block. With ~40 commands that block is one flat
list in registration order, so the three commands a new install needs — gateway,
service, doctor — land in the middle of it, and the {chat,doctor,gateway,…}
choice blob makes the usage line unreadable.
cli_help.py owns the taxonomy instead:
COMMAND_GROUPS is an ordered list of sections, each an ordered list of
(command, one-line summary). It is the single source of truth for what the
top-level help lists and in what order; Start here is first and holds exactly
gateway, service, doctor.
Its notes answer the two questions the flat list never did: how gateway
(foreground, dies with the terminal) differs from service install (systemd
unit / launchd agent, detached, restarts on crash, starts at boot, only one at
a time), and that the dashboard on loopback 5476 is the only port opened
— messaging channels connect outbound.
cli.py sets help=argparse.SUPPRESS on the subparsers action to hide
argparse's listing, passes cli_help.TOP_USAGE as the top-level usage=
(the suppressed action would otherwise drop the placeholder), and pins
prog="kirocrew" on the action so each subcommand's own usage line is
usage: kirocrew <cmd> … rather than the whole top-level usage string.
Every user-facing command is registered with cli_help.add_command(sub, name),
which raises KeyError for a name that is in no section — a new command cannot
be added without appearing in the help. The section summary becomes the
subparser's description, which is what kirocrew <cmd> --help prints, so the
sentence is not duplicated. A caller may pass its own longer description
(bench does).
Internal mcp-* servers call sub.add_parser(name) with no help, which keeps
them out of both listings. They are also kept out of argparse's
invalid choice: 'x' (choose from …) message: cli_help.hide_internal_commands
swaps the subparsers action's choices for a live Mapping view over the same
parser map that ITERATES only user-facing commands, in the help's section order.
Membership is unfiltered and _name_parser_map is untouched, so kirocrew mcp-core still dispatches — the filter changes what argparse prints, never what
it accepts. It must be installed after the last add_parser and before
parse_args.
test/test_cli_help.py pins offered-vs-listed parity by reading that same
message, and pins that a hidden command still resolves.
Command
Description
kirocrew chat -m "msg"
Send a single message, print streaming response
kirocrew chat
Interactive chat mode (readline, exit with Ctrl+D)
kirocrew chat --model X
Override model for this session
kirocrew gateway
Start the Kiro Crew server (dashboard + messaging channels)
kirocrew gateway --slack-only
Start without dashboard or SSH tunnel instructions
kirocrew gateway --no-crons
Start without cron scheduler (use when another instance handles crons)
kirocrew gateway --no-tunnel
Never publish a tunnel: refuses to start or provision one for the life of the process, whatever tunnel.enabled says. SCOPED TO TUNNELS — it does not change where the dashboard binds, so a config that widens dashboard.url off loopback still does, with token auth as the control there; do not read publish_disabled() as "no published surface of any kind". Reach the instance on the loopback port it binds (ssh -L from another host). A Dev Fleet pod boots with this whenever its own checkout declares the flag — the pod's argv is built by the control plane but executed by the target worktree's gateway, so pod.runtime.target_supports_flag probes that checkout first and DROPS the flag when it is absent (passing it would make argparse exit 2, which the unit's Restart=on-failure/RestartSec=5 turns into a 5s restart loop). Such a checkout keeps the tunnel behaviour it had before this flag existed and is not given the guarantee — see security.md for why no config-side substitute is applied.
kirocrew setup
Install agent config, save project dir, configure credentials
kirocrew setup --agent-only
Only install agent config (skip credentials)
kirocrew setup --slack
Run the guided Slack credential + slash-command setup (opt-in)
kirocrew setup --whatsapp
Run the guided WhatsApp opt-in: report the optional whatsapp extra and the pairing state, then enable the channel (opt-in)
kirocrew doctor
Verify kiro-cli is installed and config is valid. MCP counts are host-probe results, not confirmation of session tool loading; strict-identity routing is reported as configuration, not a verified live channel.
kirocrew ledger-sweep
List the session and conductor-work ledgers that look finished — kind, store directory, ledger key, phase or item counts, age and why each qualified. A DRY RUN: it deletes nothing.
Delete the ledgers the dry run listed. Irreversible, explicit, and never automatic — no gateway hook, no history-delete path and no MCP tool reaches it. Each delete is re-decided under the store's own lock and refused if the ledger came back to life. Default idle window 30 days, which must be a finite non-negative number (nan and inf are refused, because every comparison against nan is false and would admit every ledger). An unreadable record is listed but kept unless --purge-unreadable is given, and a store is never purged when its slot_key breadcrumb is absent or names a key that resolves to a different directory — a copied store is reported for a human instead of aiming the delete at the canonical ledger. The sweep infers nothing about sessions: an in-flight session ledger and a conductor holding no items are kept at any age. A ledger's age is its last write, not its terminal stamp alone, so a finished record that was written to recently is left alone too. Its own command rather than a doctor mode: doctor is read-only by contract, and a flag that belongs to this command is an argparse error under doctor, so kirocrew doctor --purge cannot look like a purge that found nothing. Rules and rationale: session-work-ledger.
kirocrew cron add/list/remove
Manage cron jobs. add registers every job kind the store accepts — an agent prompt, --script FILE.py:FUNC (a zero-token Python function that must already live under <config_dir>/crons/; the CLI registers, it does not copy) or --command SHELL — on exactly one of --every, --cron (with --timezone) or one-shot --at, plus --timeout/--timeout-secs, --model, --[no-]persistent-session, --[no-]minimal-context and --hide-in-chat. Every field is forwarded into ONE locked CronService.add_job (no post-create mutation + second _save()), a script body or command is vetted at registration by the same _vet_script_file/_vet_shell_command gates cron_add uses, and any refusal prints on stderr and exits non-zero with nothing written (2 for an argparse flag-combination error, 1 for a validation, security or store refusal) — the headless contract an installer relies on. The capabilities.cron governance gate is applied at authoring (_vet_cron_capability_governance, keyed cron:cli_add) as cron_add does, so a job authored under a disabled capability is refused rather than stored unrunnable. Session-shape defaults are CronService.add_job's own (persistent_session=True, minimal_context=False); pass --no-persistent-session --minimal-context for a script/command job.
Convert a manifest-declared plugin package into an app directory, reporting every kind that has no target. Reads and copies only; nothing in the package is executed. See plugin-import.md.
kirocrew learn add/list/remove
Manage learned corrections
kirocrew run TASK.md
Run an autonomous task from a spec file
kirocrew token
Print a dashboard access URL with auth token
kirocrew logout
Revoke all active dashboard sessions, refresh chains included
kirocrew manifest
Generate Slack manifest with user alias auto-populated
kirocrew update
Update to latest version (git fetch, pin the upstream commit, refuse a revision whose requires-python this venv fails, hard reset to the pinned commit + rebuild; a diverged checkout is refused — --force discards its local commits)
kirocrew status
Show runtime stats from running gateway
kirocrew stop
Stop a running gateway (service-aware: stops the systemd/launchd service if active, otherwise terminates the gateway found by a cross-platform port lookup — lsof on POSIX, netstat on Windows). Pass --port N to bypass the service short-circuit and target a specific gateway.
kirocrew restart
Restart a running gateway (service-aware: restarts the systemd/launchd service if active — on Linux in whichever scope runs the unit, confirming it stays up afterwards and reporting a unit that lands in activating (auto-restart) as a failed restart — otherwise terminates the foreground gateway and respawns it detached). Pass --port N to bypass the service short-circuit and target a specific gateway.
kirocrew service install
Install gateway as a system-level systemd service (Linux, requires sudo for tee + systemctl only) or launchd LaunchAgent (macOS, no sudo). Auto-restarts on crash, auto-starts on boot.
kirocrew service uninstall
Stop and remove the systemd unit / launchd plist. On Linux this covers both systemd scopes — the system unit (sudo) and a per-user unit (systemctl --user, never sudo) — and prints one line per scope saying removed (<unit file>), not installed, left in place (…) (an alias, a unit running under any load state but loaded, or a definition that is not ours is never acted on), or not reachable from this shell (…); nothing installed in either scope exits 0 with that report. In either scope the order is stop → disable → verify inactive → unlink → daemon-reload; a step the manager refuses (or a unit still running after stop returned 0) leaves the file in place and the line reads left in place (\… stop kirocrew.service` failed: …; the unit file was not removed); a system unit on a host without sudoreadsleft in place (privilege unavailable: …)while the user unit is still torn down. An alias, and a unit that is running under any load state butloaded(a runtime mask, an unparseable edit, a file removed under it), are refused whole before any verb:left in place (an active unit whose load state is masked: unmask and stop it first, then run `kirocrew service uninstall` again). Every verb rests on one ownership decision: the unit is Kiro Crew's when the file holding its definition (the reported FragmentPath, followed through a symlink to its source) is the file the installer writes for that scope or carries the installer's marker line Environment="KIROCREW_SERVICE_MANAGED=1"; anything else under our name — a distribution's unit under /usr/lib, an operator's unit systemctl --user linked under our name, a definition under another name — reads left untouched (not installed by Kiro Crew: …)with nothing stopped, disabled or unlinked, and exits 0. A unit of ours reached through a link loses only its link entries —removed (the link to ; the linked unit file itself was kept)— and a direct file of ours is unlinked. A unit file of ours at the installer's path that the manager does not have loaded and that is not running (a file dropped without adaemon-reload, an unparseable edit) is removed without those verbs, in either scope; a mask of the unit name at that path (systemctl mask's symlink to /dev/null, placed by name at exactly the path the installer writes) is ours by name and is unlinked, removed (the mask …; kirocrew.service is unmasked), while a mask elsewhere (systemctl mask --runtime, under /run) is left in place (load state masked, …); a system manager the shell cannot reach removes the file only on a host not booted with systemd (/run/systemd/systemabsent), and otherwise readsleft in place (the system manager is not reachable from this shell: …). The system scope is decided by the manager, not by the unit file: with no file at /etc/systemd/system/kirocrew.servicethe manager is still asked (an unprivilegedshow, no password prompt), so a unit of ours still loaded and running after its file was deleted under it is stopped and daemon-reloaded (stopped (its unit file … was already gone, so nothing was disabled or unlinked; daemon-reload run)— the unit is gone, so the headline reads✅ kirocrew service stopped and removed.), a loaded unit that is not ours (a vendor unit with no file of ours shadowing it) is left untouched and not stopped, one a daemon-reloadleftnot-foundbut running is refused whole like the user scope's, and an unreachable manager on a systemd host readsnot reachable from this shell (…)rather thannot installed. The headline (✅ … stopped and removed./ℹ️ No kirocrew service was removed.) follows the report's removedset, never a line's wording. Thoseleft in placelines (exit 1 whenever a unit may still be running, or the module's own file stays at/etc/systemd/system/kirocrew.service`; a system refusal with no file there and nothing running — an inactive alias provided from another directory — exits 0 like the user scope's), or a unit file that cannot be removed after a successful stop, are marked ⚠️ on their own scope line, the AppArmor profile is still removed, and the exit code is 1.
kirocrew service status
Show service status. No sudo required. On Linux it names both scopes: system scope: … then user scope: …, each not installed, not reachable from this shell (…), or the unit's ActiveState (SubState), followed by the systemctl status block for every scope that has a unit — a scope with no unit never reads inactive (dead). Exit 0 only when the unit is UP in either scope — ActiveStateactive (or reloading), what systemctl is-active exits 0 for; a crash-looping unit in activating (auto-restart) and a failed one exit 1 while the headline prints that state, so a script gating on the exit code reads a gateway that never started as down. macOS: launchctl list.
kirocrew logs
Tail gateway logs from the systemd journal (the user journal, journalctl --user, when the gateway is the per-user unit), launchd stdout file, or ~/.kiro/crew/gateway.log. Hosts without systemd/launchd, including Windows, read the UTF-8 fallback file in Python without requiring tail. Read failures exit with file-access/retry guidance instead of an exception traceback.
kirocrew logs -f
Follow logs live. The Python fallback reopens the log by name on each poll, permits Windows rename-based rotation even during reads, streams appended UTF-8 text, and stops on Ctrl+C. A replacement file or detected truncation resets the read offset; a temporary missing path during rotation is retried on the next poll. Rotated backup files are not replayed.
Provision, connect to, and manage a Kiro Crew EC2 instance in the user's AWS account. iam-boundary is the one-time admin step that pre-creates the immutable instance permissions boundary — see cloud.md.
kirocrew security events
Show recent SEL audit events (-n N for count)
kirocrew security verify
Verify SEL HMAC chain integrity
kirocrew snapshot
Create a .tar.gz snapshot of all KiroCrew state
kirocrew snapshot --keep N
Auto-prune to N most recent snapshots (default 7)
kirocrew snapshot --list
List existing snapshots
kirocrew restore <file>
Restore from a snapshot (auto-detects replace vs merge)
kirocrew restore <file> --mode replace|merge
Force restore mode; merge skips malformed incoming or local cron JSON with a file-specific warning
kirocrew restore <file> --components X,Y
Selective component restore
kirocrew restore <file> --dry-run
Preview restore without writing
kirocrew restore --list-components
Show available component names
kirocrew snapshot --allow-unpinned-staging
Stage by path name where a directory cannot be pinned by descriptor
Manage Kiro Crew agent definitions (kiro_agent, workspace, memory_store bindings). reset-model clears a spec's pinned model, the narrow way back to the shipped default — see providers.md.
Manage the AppArmor profile the agent sandbox needs (Linux) — see security.md.
kirocrew tailnet status/up/down
Publish the dashboard on your tailnet (Tailscale) and trust its origin, or stop publishing — see remote-and-mobile.
kirocrew telemetry status/disable/enable
Inspect exactly what the anonymous beacon sends, or turn it off permanently — see metrics.md and governance.md.
Staging is descriptor-pinned, and refuses rather than degrading silently
Snapshot and restore stage through kiro_crew.pinned_fs: the parent chain is resolved
once, pinned component by component with openat + O_NOFOLLOW, and everything
downstream is addressed through the descriptor already held. A validated path and the
inode later opened are otherwise not the same thing, and anything running as the user
— which in this product includes an agent — can plant the swap between the two.
os.supports_dir_fd is empty and O_NOFOLLOW does not exist on Windows, so pinning
is unavailable there. The decision, recorded here rather than only in the pull request
that made it: staging is refused on such a platform unless
--allow-unpinned-staging is passed, and when it is, the archive's MANIFEST.json
carries "staging": "unpinned" and kirocrew restore --dry-run prints that the
archive was staged by name. The refusal is the default because a by-name walk is not a
slightly weaker version of a pinned one; it is the mechanism whose failure closed two
earlier attempts at this change. The flag is a permission for a platform that cannot
pin, not a switch that turns pinning off where it works.
MANIFEST.json also carries "skipped": any file omitted during staging (a hardlink
alias, a symlink, an entry that vanished mid-walk, or an entry present but refused for
permission -- unreadable_entry) with its reason, so an incomplete archive says so in
its own record instead of only in the console output of whoever ran the command.
unreadable_entry is the one reason that is a TOLERANCE rather than a screen, and it is
narrow in three ways. It belongs to snapshot creation only -- restore and merge still
stop, because there the unreadable name is the archive's own content and skipping it
would drop data the operator asked to have put back. It covers only the permission
class, so a failing device still ends the command instead of producing a backup that
quietly omits whatever the disk refused. And it covers only reads of the operator's
own file: a refusal to WRITE the staged copy is never recorded as the source being
unreadable. A data home can hold a platform-protected path that no retry will make
readable, and refusing to produce any backup over one file is worse than a bundle whose
manifest names the gap. Files a component DECLARES are not covered: those are
product-owned state, and a bundle that silently shipped without config.json would be
worse than a refusal, so that loop still fails hard.
Retention reads those reasons as a class, not by name, through one predicate,
pinned_fs.omits_wanted_data(). symlink and not_regular are screened on every run by
design, so a bundle that screened one is complete; everything else means the archive
lacks something it was asked to carry. --keep does not prune after a run in the second
class, and prints which reasons applied. A reason code the split does not know counts as
incomplete, because the alternative is pruning the operator's last complete backup on
the strength of a code nobody has classified yet.
SQLite databases are out of scope for the pinned staging described here: they keep the
sqlite3.backup() path they already had, which reopens the live name. Capturing a live
database without reopening its name is a genuine conflict of requirements — SQLite accepts
only a path and cannot be pointed at a held descriptor — so it is tracked separately rather
than solved alongside the tree walk. The exposure is unchanged from before this staging
work, not introduced by it.
A refusal to stage is a permission decision and is written to the SEL audit log —
snapshot_rejected or state_restore_rejected, both with reason=unpinnable_staging.
The dashboard's import path (portability.apply_import_zip) is the exception, and
deliberately: it has no flag and no consent surface, so refusing there would not mean
"ask the user", it would mean deleting import on that platform. It therefore proceeds with
a by-name traversal where pinning is unavailable and records "staging": "unpinned" in its
returned summary, with a logged warning. Snapshot and restore keep refusing, because
--allow-unpinned-staging lets them ask. The per-entry screens apply on both paths — the
copy opens O_NOFOLLOW and the walk rejects links and reparse points — so what the import
path gives up is ancestor-swap resistance, not link resistance.
Each rule above has one owner. kiro_crew.snapshot is the command and API facade: it
holds snapshot_main and restore_main, the MANIFEST.json writer (_build_snapshot),
the outbound redaction seam, the merge-mode driver and the notification copy.
restore_main sequences extraction and the bundle-shape refusals, and writes their
state_restore_rejected audits. The facade re-exports the owners' names, so every existing
import keeps resolving. The owners read the helpers, drivers and limits a test replaces
(_copytree_safe, _do_replace_mutations, sqlite3, _MAX_ARCHIVE_MEMBERS and the rest
of LATE_BOUND in test/test_snapshot_refactor_seams.py) through kiro_crew.snapshot when
they use them, via snapshot_components._facade(), so a patch of one of those names on the
facade reaches the owners' call sites too, and the facade stays a plain module. Every other
name an owner uses resolves in that owner's own globals, so a test patches it on the owner
module. The same test file scans test/ and every tests package under src/kiro_crew and
fails on a patch no call site sees: a facade patch of a name an owner reads from its own
globals, an owner patch of a LATE_BOUND name, and a patch of either whose name it cannot
resolve outside the sites it lists. No owner imports the facade: _facade() reads it from
sys.modules, since the facade imports every owner and is loaded before any of them runs.
test_snapshot_refactor_ownership.py pins both halves: no owner imports the facade, and no
module outside the snapshot family imports an owner.
kiro_crew.snapshot_components owns the component table, the never-ship and host-local
rules, and the tree-root check safe_tree_root.
kiro_crew.snapshot_archive owns staging and the bundle format. That covers the pinned
tree copy with its refusal (_staging_is_pinned), the consistent SQLite capture, the
extraction filter, the archive bound and the manifest readers.
kiro_crew.snapshot_restore supplies the bundle predicates those refusals rest on
(_component_payload_absent, _components_absent_from_bundle,
_trees_absent_from_bundle), the content-soundness refusal
(_refuse_corrupt_source_databases), the destination guards, and the replace transaction
with its rollback.
kiro_crew.snapshot_merge owns the merge algorithms: memory rows, cron jobs,
notification records and no-overwrite trees.
| kirocrew config get [key] | Print full config or a dot-path value |
| kirocrew config set <key> <val> | Set a config value (auto type detection). A key whose schema declares an enum refuses a value outside it on every write (exit 1, naming the selectable values), because the load path answers such a value by degrading it with a warning rather than rejecting it — the write is the last point where the mistake is still attributable to the command. What is written is the enum's own spelling: a case variant is canonicalised (agent.log_level debug stores DEBUG), and a key with a loader-side alias table (stt.model, through stt.models.canonical_name) stores the row the alias names (turbo stores large-v3-turbo). Only declared enums on concrete registry paths are checked; a wildcard path (slack_channels.*.activation) keeps reaching the loader's own degrade rule. Type is checked only on a declared leaf's first write (_declared_type_error); the enum check has no stored value to stand in for it. |
| kirocrew config set --file <path> | Replace config from a JSON file |
| kirocrew config edit | Open config in $EDITOR |
| kirocrew memory list/search/stats/audit | Inspect vector memory (entries, semantic search, counts, suspicious-content scan) |
| kirocrew memory show [preferences\|projects\|history] | Read the markdown memory layer (all three when no target given); --format md\|json, --since YYYY-MM-DD for history |
| kirocrew memory export/import/migrate | Export one store's rows to JSON, import them back, or migrate legacy markdown memory into the vector store. Both export and import take --store <name> (default: the default store), so a named store's rows are reachable in either direction as they already are for backups, restore and carve. --include-markdown adds the markdown layer and is DEFAULT-STORE ONLY: a named store's markdown root is under the fenced memory_stores/ subtree, whose refusal markdown_snapshot cannot distinguish from an empty file, so the combination is refused rather than reported as empty |
| kirocrew memory backup/backups/restore | Take hot copies of Global, declared named V1 and actively owned V2 stores (--keep <n>), list a store's copies newest-first (--store), or stage a restore (--store, --from <file>, defaulting to that store's newest). Both V1 and V2 activate restoration at gateway restart. Archived V2 stores are excluded from routine backups. restore --store <member-store> --cancel-pending cancels a staged intent while preserving current memory and its backup; it is mutually exclusive with --from. A failed activation still requires restart after cancellation. All three dispatch BEFORE the shared vector store is opened, because opening it raises on exactly the corrupt file these verbs recover. See memory-skills-hooks |
| kirocrew memory retired | List the episodes a semantic write superseded and restore one (--restore <id>, --limit). Default store only — it has no --store, so restoring a retirement inside a silo is a dashboard action. See memory-skills-hooks |
| kirocrew memory carve --store <name> | Filter or count a crew store's rows by their carve facets: one flag per facet (--scope/--surface/--crew/--session-key/--derived-from), --kind, --count-by <axis> for grouped counts, --limit/--offset. Facets exist only on a crew memory store, so the default store answers with a named refusal rather than an empty list. See memory-skills-hooks |
| kirocrew policy show/validate/explain/profile | Inspect the effective enterprise security policy, load-check it and all profiles, explain one tool/scope decision for a surface, or print a profile. show also summarizes the built-in denied-command catalog as grouped counts (--ids lists each category's rule ids), on every install regardless of whether an enterprise policy is active — the one place an agent can learn a class of work is hard-denied before planning around it. |
| kirocrew pod up/down/ls/status/token/url/scenarios/api/logs/exec/install/provision | Isolated worktree test gateways. Three per-user, no-elevation backends: Linux systemd --user, macOS launchd, Windows Task Scheduler (schtasks.exe). On a host with none of them every service-manager-touching verb refuses with a one-line message. pod api is additionally Linux + macOS only, because its authenticated request goes over an AF_UNIX socket with no TCP fallback and CPython on Windows has none. See src/kiro_crew/pod/README.md → Platform for the per-backend capability table and the two ceilings (memory/CPU, crash restart) that only systemd enforces. |
| kirocrew pod scenarios [--json] | List packaged seed scenarios in deterministic name order. Human output shortens each description to the last complete sentence that fits, cutting between words with an ellipsis when none does; --json emits an array of {name, description} rows with each fixture manifest's complete description: scalar; literal (|) blocks preserve newlines, while folded (>) blocks normalize to one paragraph. Extraction stays dependency-free without requiring PyYAML at runtime. An empty registry returns success with [] in JSON mode or an explicit human diagnostic. |
| kirocrew pod api | api <wt> <METHOD> <path> [--data JSON] [--allow-write] makes one authenticated request and prints {name, method, path, status, ok, body}; GET/HEAD are the default surface and other methods require --allow-write. It refuses caller-supplied token query parameters without echoing them, authenticates with the dashboard's query-token contract, caps response reads, and mints only after the pod PID record agrees with systemd MainPID (listener tools are optional corroboration). The authenticated request goes over the pod's private dashboard-<port>.sock in the pod's own home with no TCP fallback, so the minted token cannot reach a process that took the pod's port; an absent socket refuses through the envelope before minting. See src/kiro_crew/pod/README.md. |
| kirocrew pod up --seed <scenario\|dir> | Pre-populate the isolated home. A bare name is a packaged fixture and populates the whole home; a path contributes only its sanitized config.json. Unknown names are refused with the available list. Named fixtures copy directly into the final home through pinned source/destination directory descriptors; config/workspace setup uses the same held home, and the manifest lands last as the completion marker. Populated homes are never overwritten: a named-seed request against one refuses before start even when its marker matches; use plain pod up to restart it unchanged. Seeded config disables channel enablement and restores the sandbox floor. A per-instance systemd drop-in runs the checkout's own venv binary, post-health marker readback detects a seed that did not land, and pod down removes the drop-in with the home. Pod homes provide operational/state isolation, not protection from arbitrary same-UID host processes; Controller v1 invokes seeding from the host control plane and does not support nested pod control. |
| kirocrew pod up --no-embeddings | Boot the pod without the embedding model. Records EMBEDDINGS='0' in the per-pod env file, which makes boot export KIROCREW_SKIP_MODEL_DOWNLOAD=1into the pod's env only — never the operator's shell, profile or real data home, which keep whatever model they already had. The pod downloads no GGUF and computes no vector; memory and knowledge search answer through the documented keyword fallback, so the instance stays usable rather than degraded. Exists for load-testing ingestion, which embeds per chunk: without the model, chunk count grows while embed compute does not, so a bundle that previously hit a 30-minute wall is drivable and any residual slowness is attributable elsewhere. The setting reaches pod exec through the same pod_context seam, so a command run against an embedding-light pod cannot start the download either; build_pod_env also drops any inherited KIROCREW_EMBED_MODEL_PATH / KIROCREW_EMBED_MODEL_URL alongside the switch, since the switch gates only the download and a custom model path would otherwise load and embed anyway. boot announces the mode in the journal from the env the pod actually runs with (an inherited KIROCREW_SKIP_MODEL_DOWNLOAD=1 is announced as such), and the pod.up audit row keys embeddings=off on the merged env file plus that inherited switch, not on the flag alone. The switch is subsystem-wide, so the pod skips its speech-to-text (whisper) model download too. Read once at boot (recording it against a live pod applies on the next one, and pod up says so) and re-validated from the hand-editable env file, where an unrecognised value leaves embeddings ON — the pre-existing behavior. Because that file is hand-editable, a pod kept around for other work can be flipped persistently by editing its EMBEDDINGS= line and bringing it up again. Deliberately NOT an unreachable KIROCREW_EMBED_MODEL_URL: that spends the downloader's whole attempt budget on requests chosen to fail, and a non-https:// value is ignored in favour of the real CDN, so a malformed sentinel downloads the model the option exists to avoid. |
| kirocrew pod up --wait-secs <SECS> | Health-wait budget in seconds before pod up gives up on the gateway answering /api/health (default 90 — a first boot does config migration, CLI staging and .local_secret minting on top of serving, so the default carries margin for a loaded host). Resolution is flag > KIROCREW_POD_HEALTH_SECS env var > default; a malformed env value warns on stderr and falls back to the default rather than making pod up unbootable, and values are clamped to the 5..3600 range so a misconfigured tiny budget cannot fail every boot instantly and an absurd one cannot wedge pod up for hours. When the budget is exhausted while the unit is still active and not crash-looping, the failure says gateway still starting after Ns and names both raise paths, instead of the never became healthy verdict reserved for a gateway that actually died. |
| kirocrew knowledge dedup [--apply] | Collapse cross-source duplicate knowledge documents. Without --apply it is a dry run that lists the collapses and writes nothing: it opens the database with SQLite mode=ro (KnowledgeStore.open_read_only), so the constructor's schema migration and orphan sweep do not run, and a library behind the schema is reported (SEL schema_behind), not migrated. --apply takes the ordinary migrating open and performs the deletes. |
| kirocrew knowledge stats [--json] | Count the knowledge library: sources, documents and items in total and per source. Read-only: it opens the database with SQLite mode=ro (KnowledgeStore.open_read_only), so it never runs the constructor's schema migration or orphan sweep, and a library behind the schema is reported, not migrated. There is deliberately no flush/rebuild/repair verb beside it. Its LLM-facing twin is knowledge_list_sources, which renders the same aggregate_stats() call (see knowledge §6 and mcp § The MCP-first rule) |
| kirocrew cron preview <script> | Run a script cron locally with real MCP tools; notifications are captured and printed instead of delivered |
| kirocrew workspace create/update --dir <name> | --dir is a directory NAME that must resolve to a strict descendant of the data home (~ is expanded first); anything landing outside — and the home root itself, in any spelling — is refused with a SEL denied audit event. Containment, not an absolute-path ban: an absolute path under the home resolves where the relative form would and is accepted. The strict-descendant test is what closes the root case for tilde paths, since the per-call-site root-equality checks compare un-expanded config_dir() / ws_dir. Deliberately stricter than the dashboard's POST /api/workspaces, which accepts an absolute dir anywhere, screened by is_sensitive_path. |
| Workspace directory materialization | The declared directory is materialized before its configuration entry is committed. Create does this with one atomic mkdir immediately before the config write, in the same locked section, made relative to the pinned parent (pinned_fs discipline: each parent component is opened O_NOFOLLOW from the path as validation resolved it, so a component swapped for a link between validation and the mkdir is refused rather than followed; where the platform cannot pin, Windows, the create is by name). An existing directory is adopted (pointing a new workspace at a folder the owner already keeps is supported, and is the normal case for an absolute dir); a path that exists and is not a directory is refused, as is a missing parent, rather than fabricating a tree. The created directory is not rolled back when the config write fails: it is reachable only through the entry written in that same section, so a failure leaves an empty directory nothing references, and deleting it would race a concurrent create that has adopted and registered the same path. The same rule holds for a --copy-from create whose config write fails after the copied tree was installed: the tree is left in place and its path is reported (CLI stderr note; handler warning log), because a concurrent create can already have adopted and registered it and deleting it would leave that workspace declared with no directory. Update does the opposite half: it REFUSES a dir that is missing or is not a directory (workspace_dir_unusable; CLI exit 1) instead of creating one, because it names a destination the owner already chose and its transaction carries no rollback. Consequence: "rebind first, create the folder after" is no longer a valid sequence — create makes its own directory, update binds one that exists. Residual, tracked in #10156 rather than chased here: writers that reach the workspaces section without going through these two commands (config edit, keyed config set workspaces.<name>.dir, a backup restore) can still commit an entry naming a missing directory; config set --file refuses a new or changed one (ConfigWriteRefused, exit 1). |
| kirocrew computer doctor [--json] | Report computer-use availability: platform support, the keystone primary-enable state, and the advisory macOS Accessibility / Screen Recording probe with a responsible_hint. See Computer Use Commands. |
| kirocrew computer apps | List on-screen applications the accessibility layer can address (human-facing twin of the computer_list_apps MCP tool). Gated by the same chokepoint as call — refused while the feature is off or the session is unattended. |
| kirocrew computer call <tool> [k=v ...] | Run ONE computer-use tool through the same gated chokepoint the agent uses, and print its reply (debug / reproduction) |
| kirocrew computer call --calls '[…]' | Run a JSON array of tool calls in a SINGLE process, so element_index values from an earlier computer_get_state are still resolvable |
| kirocrew mcp-cron | MCP server for cron tools (spawned by kiro-cli) |
| kirocrew mcp-core | MCP server for spawn, learn, task tools (spawned by kiro-cli) |
| kirocrew mcp-computer | MCP server for computer-use tools (spawned by kiro-cli; hidden — registered with no help, so it is in neither listing). A thin shim — it forwards to the gateway over loopback and does no accessibility work itself. |
| kirocrew --version | Print version |
Token Command Output Streams
kirocrew token has a machine-readable stdout contract: stdout carries only
the dashboard URL(s), and every failure reason (invalid TTL, gateway not running,
gateway unreachable, gateway refused, empty token) goes to stderr.
A gateway that ANSWERS with an HTTP error is reported as a refusal carrying the
gateway's own reason (Gateway refused the token request (HTTP 403): <error body>), never as Could not reach gateway: the 403 /api/token/local returns
when it cannot prove the caller's host provenance (a sandboxed caller, an
unresolvable peer pid, a different namespace on Linux — see
security) names its remedy in the body, and reporting it as a
network failure sent the operator to the wrong fix.
The contract exists because stdout is parsed, not just read by a human. The
remote-mint path (kiro_crew.instances.token_mint.mint_remote_token) runs
kirocrew token on a remote host over SSH and regex-extracts the JWT from its
stdout. Error prose on stdout would both break the Unix convention and hide the
reason from a caller that captures stderr.
Legacy remote handling. Older remotes predate this split and still print
their failure reasons to stdout, which made a stderr-only error message degrade
to a bare <no stderr>. mint_remote_token therefore also carries a bounded,
redacted stdout tail in TokenMintError — appended only when stdout was
non-empty, so a current remote keeps the single-stream message shape. Because
stdout is the one stream that legitimately carries a token, the tail is
token-scrubbed (URL-borne and bare forms) before the generic credential and
exfiltration redactors run.
Setup Command
kirocrew setup performs:
Saves KIROCREW_PROJECT_DIR to ~/.kiro/crew/project_dir
Installs agent config to ~/.kiro/agents/kirocrew.json
Prompts for Slack credentials and the slash-command name only when --slack
is passed; the default wizard configures no messaging channels and prints a
pointer to connect them later
Offers to set up custom domain kirocrew.localhost (macOS/Linux)
The saved project dir enables running kirocrew from any directory.
Slack credentials are checked before they are stored
kirocrew setup --slack asks Slack about each value as it is pasted, so a typo,
a revoked token, or a channel ID pasted where the member ID belongs is reported
at the prompt instead of surfacing later as a "Slack disabled" line in the
gateway log. The app-level token is checked with apps.connections.open (the
call the gateway itself makes at startup) and the bot token with auth.test,
which also names the workspace — the same two calls, and the same rules, as the
dashboard's Slack credential save. The member ID is format-checked against
validation.USER_ID_RE first (so C…/B… is refused with no network, and
Enterprise Grid's W… is admitted) and then confirmed with users.info; only
user_not_found / users_not_found (cli_setup._SLACK_OWNER_REJECTIONS) indict
the ID, because any other Slack error (a missing users:read scope, a rate
limit) indicts the check.
Three verdicts, and only one of them refuses a value:
Accepted — saved, with the workspace or member name printed.
Rejected by Slack — reported with Slack's own error code and re-asked, up
to three times; if all three are refused the step writes nothing, so a
working credential already in .env is never replaced by a broken one.
Unverifiable (Slack unreachable, transport error) — a warning, and the
value is saved as typed. Being offline never costs the operator the
credentials they just typed.
The check runs only when stdin and stdout are both a terminal. Off one, nobody
can see a verdict and re-asking would consume the next line of a piped answer
file and misassign every remaining answer, so an automated run (kirocrew update re-runs setup with its output captured and stdin on /dev/null) behaves
exactly as it did before the check existed. After a successful write the step
PRINTS Restart the gateway to pick them up: kirocrew restart, since Slack
credentials are read once at startup and tokens written here stay inert until a
running gateway restarts. It deliberately does not offer to perform the restart:
doing so would drop every in-flight session, and probing whether a gateway is up
in order to decide costs a service-manager query and a port probe that can each
fail in ways the wizard then has to degrade around — all to save one command the
line already names.
First-run Kiro CLI prerequisite onboarding
KiroCrew exposes the same two-step readiness contract on every supported
platform: an executable candidate must answer kiro-cli --version, then
kiro-cli whoami must confirm authentication. Candidate discovery includes
supported fixed locations in addition to inherited PATH; unusable candidates
are reported for repair. Setup probes the same first executable candidate ACP
will launch, so a stale earlier candidate cannot produce a false-ready result
from a different later installation.
Missing CLI: the setup page offers an explicit install action on macOS,
Linux, and Windows. macOS/Linux download the fixed
https://cli.kiro.dev/install; Windows downloads the fixed
https://cli.kiro.dev/install.ps1. Every redirect and the final response
must remain on the exact cli.kiro.dev:443 endpoint and expected path, with
no userinfo, query, or fragment. Redirect destinations are resolved and
validated before any request is sent, and the chain is limited to three
redirects. Responses are size-bounded and must match a release-pinned
SHA-256 digest plus the platform-specific official installer marker. A
changed upstream script therefore fails closed until a KiroCrew release
updates the pin; the manual official guide remains available. The exact
validated bytes stay in memory and run through the fixed system interpreter's
standard input. The installer receives a system-only PATH plus explicit
HTTP(S) proxy variables, never user-writable executable directories or
ambient application credentials. The official installer additionally
verifies its downloaded package manifest and artifact checksum.
Unusable CLI candidates: the same page identifies that Kiro CLI needs repair
instead of treating a spawn failure as a signed-out session. If the upstream
POSIX installer would require an interactive /dev/tty replacement prompt
(an existing macOS app bundle or Linux ~/.local/bin/kiro-cli), automatic
repair is disabled and the user is directed to the official guide.
A candidate that already runs is directly usable for sign-in regardless of
install source; the post-installer attestation file is now write-only
bookkeeping and does not gate credential access.
Installed but signed out: the setup page names the commands the USER runs and
runs nothing itself — kiro-cli login for a personal account (Builder ID,
Google, or GitHub), or kiro-cli login --use-device-flow --license pro for
organization SSO, which prompts for the organization's start URL and region.
Both are backend code constants rendered verbatim in a <code>, never catalog
values, because a translated command cannot be typed. Both tiers are named
because the browser portal the bare command opens presents a free Builder ID
as a peer of organization SSO; Kiro Crew does not detect which tier applies,
so the gate describes the choice and the user makes it. Sign-in completion is
observed only through the read-only kiro-cli whoami probe.
Browser dashboard: the authenticated SPA gate operates on the gateway
host, not the browser host. This covers native Windows source installs,
Linux gateways, and browsers connected to another machine.
Desktop shell: the shell starts or reuses the gateway first, then displays the
same gateway-served setup gate as a browser. Remote gateways are therefore
checked on the remote host rather than the desktop host.
Offline test harness: the explicit gateway --test-mode bundle injects a
ready prerequisite state so deterministic fake-ACP smoke and Playwright
suites do not depend on a developer machine's Kiro installation, identity, or
Linux sandbox capabilities. Ordinary gateway invocations always use the real
probe/install/login service.
The setup client cannot supply a command, URL, argument, or output path. The web
API exposes only fixed install/login mutations to the configured owner (or the
signed local-app / local-startup identities before an owner exists).
Authenticated non-owner dashboard users receive only a redacted readiness bit:
they enter the dashboard once ready but cannot see host state/output or operate
setup. App tokens remain denied. Electron has no separate installer/login IPC
or subprocess implementation. Filesystem discovery runs off the event loop.
Version probes use a minimal noninteractive environment with no proxy
credentials or desktop-session IPC. They use the strict OS sandbox and
additionally hide the configured data home, ~/.kiro/crew, ~/.kirocrew, and
every known Kiro identity store. Any candidate that runs --version is eligible
for whoami and device login — trust is "it runs, and it has a valid login",
not install source, owner, or fixed path (KiroCrew is not the authority on where
Kiro CLI is installed, and its self-updater rewrites its own bytes as the user).
Auth calls execute the user's installed binary IN PLACE, never a private copy of
its bytes — a multi-call Kiro CLI resolves its sibling subcommand executable
relative to its own path, so a copy strands it (see security.md).
Sign-in itself is delegated to Kiro CLI: login --use-device-flow runs in the
standard sandbox against the user's real home, with only the Kiro Crew data
homes hidden, and the CLI writes its own credential store exactly as it does
from a terminal. KiroCrew stages nothing and publishes nothing, so no staged
state has to be reconciled after a failure, timeout, or cancellation. The
credential-minimal temporary home populated only with Kiro identity JSON and
SQLite files survives as an opt-in read-only mode — one that also hides
unrelated AWS, SSH, GitHub, and Kubernetes state — and its temporary directory
is removed on every exit path. Any allowlisted live identity artifact that is a
symlink, non-regular, oversized, unreadable, or disappears while being captured
aborts that mode before the command runs. Every
probe emits a critical invoked SEL event before spawn
and a best-effort terminal event without argv, candidate paths, output, or
environment values. Installer and login timeouts cover process exit and
output-pipe draining. On POSIX, a private supervisor remains the process-group
leader after the real command exits and keeps the group safely addressable until
all descendants close or are terminated. Windows cleanup opens an exact
primary-process handle before yielding after spawn and completes an initial
descendant snapshot even when the launcher exits immediately. It then retains
exact child handles and continues discovery from every live child, so late
helpers remain supervised and identifier reuse cannot target an unrelated
process. Numeric parent edges are accepted only when exact-handle creation and
exit times prove that the child was created during the parent's lifetime, and
the check is repeated across both tree snapshots. The primary root and each
retained child root receive one final snapshot after becoming inactive, so a
child spawned immediately before its parent exits is not lost between polls.
Failure to anchor or validate the primary process, create a Toolhelp snapshot,
or complete any process enumeration fails the operation closed. Ordinary
pipe/task errors follow the same terminate, reap, cancel, and cleanup path
before another action can start.
An auto-created config.json alone does not mark first-run setup complete; a
successful authenticated probe writes the setup marker, while existing
session/history state preserves established-install migration behavior. Fresh
installs receive the full-screen flow. Established dashboards remain navigable
and fully usable when signed out — no controls are paused. Readiness is probed
once at gateway boot and thereafter only on explicit user action, so no path runs
a subprocess probe on the message hot path; the authoritative logout signal is
the ACP attempt's AcpAuthRequired. The SPA refreshes ready status every 30 seconds,
retains cached readiness across transient refetch errors, and invalidates
prerequisite state after access-cookie refresh.
POSIX group membership ignores zombie records, which cannot retain pipes or
perform work, so an unreaping PID 1 cannot hold the supervisor forever.
The supervisor source is captured eagerly at import for replacement resistance;
if it is missing or unreadable, gateway import still succeeds and each affected
POSIX setup operation fails cleanly before spawning a command.
Sandbox launcher/profile preparation and cleanup are worker-thread operations
and do not stall the asyncio gateway loop.
Setup and ACP launch share the side-effect-free kiro_cli resolver on every OS.
Status requests never publish a discovered path by mutating KIROCREW_KIRO_BIN.
Both setup discovery and ACP launch enumerate the same candidates — inherited
PATH, the interpreter Scripts directory, package-manager dirs (incl. the
Windows Program Files Kiro-Cli tree and winget/scoop/user installs on PATH),
and an operator override — and accept a runnable candidate wherever it lives,
since trust is "the CLI runs". ACP launch runs the resolved candidate in place on
every platform — never a copy of its bytes. Setup discovery and ACP resolution therefore agree on
Windows, so a winget/scoop install is never sent to a redundant reinstall.
When a previously completed setup is no longer ready, the dashboard remains
fully navigable and fully usable — nothing is paused and no sign-in chrome is
shown. A signed-out CLI is reported by the turn itself as an actionable
kiro-cli login error card (see modules/learn-cron-dashboard.md § "The
dashboard does not guide the user to sign in"). Only the endpoints that act
BEFORE a turn still return 503: the poll-driven kiro-cli spawn sites
(/api/models, /api/sessions/usage) and the destructive reruns (regenerate,
edit-resend, rewind), which rewrite persisted history up front.
Custom Domain
After credentials, kirocrew setup offers to add 127.0.0.1 kirocrew.localhost to the system hosts file so the dashboard is accessible at http://kirocrew.localhost:5476:
macOS/Linux: Uses sudo tee -a /etc/hosts for safe append
Skipped if kirocrew.localhost is already present or user declines.
Cloud Command
kirocrew cloud is a human installer/control-plane surface for running
KiroCrew on the user's own AWS EC2 instance. Provisioning and teardown are not
LLM-facing tools. AWS credentials are resolved by the AWS CLI; KiroCrew stores
only profile, region, and the most recent instance tag in cloud.json.
kirocrew cloud launch runs a six-step wizard: check AWS reachability, explain
permissions, choose whether to keep an existing deployment or create a new one,
choose an instance size when creating a new stack, deploy or resume the
CloudFormation stack, sign in the remote kiro-cli, and open the dashboard
through SSM port forwarding. Launch is resume-safe by default: if cloud.json
contains a last_tag whose stack still exists in the same saved profile/region,
rerunning interactive launch offers to keep/resume that stack or create a new
installation. If cloud.json is missing or stale, launch discovers existing
kirocrew-* CloudFormation stacks with cloudformation:ListStacks and offers a
choice to resume one or create a new installation. kirocrew cloud launch --new
is the explicit escape hatch for creating a separate new stack. --yes keeps a
single or saved existing stack; if multiple unsaved stacks exist it fails closed
instead of choosing one arbitrarily. For a new launch, the generated tag is
written to cloud.json before the long CloudFormation deploy starts, so an
interrupted provisioning run can be found on the next launch attempt.
Launch and connect require the local AWS Session Manager plugin for
AWS-StartPortForwardingSession. If session-manager-plugin is missing,
cloud launch prompts to install AWS's official package for the current local
platform (macOS .pkg, Debian/Ubuntu .deb, or RPM Linux .rpm) before the
wizard reaches sign-in/dashboard tunneling. --yes accepts this installer
prompt. cloud connect performs the same check and installer prompt before
opening the dashboard tunnel. If installation is declined or fails, the command
exits non-zero and tells the user to retry after fixing the local prerequisite.
The instance-size picker supports arrow keys in an interactive terminal
(↑/↓, j/k, digit shortcuts, Enter to select) and falls back to the
numbered prompt for non-TTY input. Ctrl-C must interrupt prompts and long AWS
subprocesses; unhandled cloud-command interrupts return exit code 130.
Remote Kiro sign-in prefers the device-code flow over SSM. The launcher starts
kiro-cli login --use-device-flow as a background process on the instance,
captures the URL/code from its log, and leaves that same process alive while the
wizard polls for completion. It must not kill that process after scraping the
prompt or start a second hidden device-code flow. If device-code startup does
not produce an actionable URL, launch falls back to the Google/GitHub callback
flow automatically: it starts kiro-cli login on the instance with FIFO-backed
stdin, captures the printed loopback callback port, opens an
AWS-StartPortForwardingSession from the same local port to the remote port,
sends the Enter continuation back to the remote CLI, then opens or prints the
local browser URL. The temporary callback tunnel is closed after the sign-in
poll completes. In headless local terminals, browser auto-open is skipped and
the URL is printed for manual opening.
kirocrew cloud connect mints a dashboard token over SSM, opens an
AWS-StartPortForwardingSession, waits for the local tunnel port to accept TCP
connections, and opens or prints the local dashboard URL. If the tunnel port
does not become reachable, the command reports failure, does not present the
dashboard URL as usable, and does not keep a dead tunnel process open. If final
dashboard opening fails during cloud launch, the instance remains running but
launch returns non-zero and tells the user to rerun kirocrew cloud connect
after fixing the local SSM tunnel issue.
Config Command
kirocrew config manages ~/.kiro/crew/config.json:
get — prints full effective config (with defaults resolved) or a single dot-path value
set key value — sets a value with auto type detection (bool/int/float/JSON/string). Rejects unknown leaf keys.
set --file path — replaces entire config from a JSON file. File read routed through hooks.safe_read_file() (blocks sensitive paths).
edit — opens config in $EDITOR (supports args like code --wait via shlex.split). Creates default config if missing.
All write paths emit SEL audit events (config_get, config_set, config_set_file, config_edit).
Gateway Auto-Create
kirocrew gateway creates ~/.kiro/crew/config.json with defaults if the file doesn't exist. Does nothing if it already exists.
Secrets Command
kirocrew secrets maintains the encrypted secret vault. The vault's
store/list/delete surface is the dashboard Settings → Secrets tab (the
/api/secrets routes); the CLI carries only the migration importer:
import — migrate plaintext credential lines from the data-home .env
into the vault. Dry-run by default (reports what would migrate, changes
nothing); pass --apply to store the secret(s) and rewrite each migrated
.env line to a secret://KEY reference. Only the Jira credential keys the
vault-aware consumer reads are migrated (JIRA_API_TOKEN and per-host
JIRA_TOKEN_<HEX>); every other key is left untouched. There is no --file
option — the importer reads only the data-home .env, so a caller cannot
point it at an attacker-controlled file.
Once stored, a secret is referenced from mcp.json as secret://NAME and
resolved into the bound server's environment at spawn time. See
secrets-env.md for the end-to-end flow. Secret
values are write-only — there is deliberately no command that reads a stored
value back.
Verbosity
Flag
Level
What you see
(none)
WARNING
Errors only
-v
INFO
Session lifecycle, context %, compaction
-vv
DEBUG
ACP events, message updates, full traces
Interactive Mode
Prompt: you>
Exit: exit, quit, /exit, /quit, :q, Ctrl+D
Streaming output printed as chunks arrive
Tool permission requests
A backend that routes tool decisions over ACP holds the turn open until it gets
an answer, so the stream consumer must respond — an ignored request is not a
missed prompt, it is a turn that never ends.
Answering one is an authorization decision, so the CLI is a security
surface, not just a prompt. Every request runs the same ladder, in this order:
permission_request
→ HookManager.on_tool_call (sensitive paths, denied commands, ceiling ∩ profile)
deny → SEL "denied" → reject_tool → stderr notice [not overridable]
→ may this invocation ask? (command mode AND both streams a TTY)
no → SEL "denied" → reject_tool → stderr notice
→ prompt the human
exactly "a" → SEL "allowed" → approve_tool(always=False)
anything else → SEL "denied" → reject_tool
The gate is fed the event's non-model-authored fields (tool_kind,
raw_tool_params, shell_command, is_shell, mcp_server_name, tool_name),
not just title: for a shell tool the title may be an LLM-authored
description, so a dangerous command behind a benign label is exactly what
keying on the title alone lets through. The CLI identifies itself as
session_key="cli_chat" (SEL source cli) with the resolved agent, which is
what lets the gate resolve ceiling ∩ profile rather than the ceiling alone.
No second copy of the sensitive-path or denied-command rules lives in the CLI.
The hook result is used as a deny ceiling only: TOOL_DENY rejects, and
both TOOL_ALLOW and TOOL_AUTO_APPROVE still ask the human. This consumer
answers permission requests; it does not carry the dashboard's trust and
auto-approval semantics, and honouring TOOL_AUTO_APPROVE here would add a
second execution path with no human confirmation. Asking more often than the
dashboard is the safe direction.
Mode
stdin+stdout TTY
Behaviour
interactive REPL
yes
Prompt. Exactly a allows once; anything else denies.
interactive REPL
no
Deny automatically, notice on stderr, stdin untouched.
-m single message
either
Deny automatically, notice on stderr, stdin untouched.
Both conditions are required, and neither implies the other. -m is documented
as Single message (non-interactive), so a terminal does not license a prompt
there — a script wrapped in a pty would otherwise block on a question nobody is
watching for. And a prompt nobody can see is a hang, not consent, so the REPL
still needs a real terminal on both ends. The mode is passed explicitly rather
than inferred from a TTY check, because only the caller knows which mode runs.
A shell call shows the command it is asking about:
Permission required: Run a helpful script
Command: git status --short
[a] allow once [d] deny (default):
title is LLM-authored prose, so approving on it alone is consent to a
description rather than to what runs — the same reason the gate keys on
shell_command. The command is redacted with the two security helpers,
collapsed to one line, and capped with an explicit ... [truncated]. Local
paths are deliberately not redacted here: this is the operator's own
terminal, and seeing the real path is part of the consent.
A call that names a file discloses that file. A trusted tool identity says
WHICH tool runs, not what it runs against, so fs_write under a benign title is
consent to a verb:
Permission required: Tidy up the notes
Tool: fs_write
Path: /home/tester/thesis.md
[a] allow once [d] deny (default):
The path is read from the request's own raw_tool_params, under the same
path / file_path / filePath spellings hooks._SEARCH_DENY_ARG_KEYS
accepts. Sharing the spellings is the point: the gate already denies a
sensitive path or a write-protected config path read from those keys, so what
is left for a human to judge is exactly the ordinary valuable file no rule
speaks for — and a prompt reading a different field than the gate inspects would
let the two disagree about what the target is. Rendered like the command
(sanitised, one line, capped), and shown for any call that carries a path rather
than only an edit-kind one, because the kind is agent-influenced and
disclosure can only inform the decision.
Absence of a path is not a refusal. Most builtin calls legitimately act on no
file — a memory write, a tag creation — so denying whenever a path cannot be
found would refuse them all to close a gap that exists only for tools which name
one. A non-string value is treated as absent rather than raised on: raising would
leave the request unanswered, which is the hang this path exists to end.
Beyond the command and the path, the whole tool input is still not shown —
this is the question, not a detail panel.
Terminal controls are neutralised on this surface. Every untrusted string
the permission UI prints — the title, the command, and a gate reason — goes
through _for_consent, which redacts as above and then replaces ESC, the C0 set
(U+0000–U+001F), DEL (U+007F), and the C1 set (U+0080–U+009F) with
spaces before collapsing whitespace, so the result is always a single line. Lone
UTF-16 surrogate code points (U+D800–U+DFFF) are removed by the same boundary:
they are not Unicode scalar values, so even a strict UTF-8 stream cannot encode
them. The result is then round-tripped through the destination stream's codec
with backslashreplace. UTF-8 terminals retain ordinary Unicode; a strict cp1252
or other legacy Windows stream sees inert ASCII escapes for characters it cannot
represent. This is an authorization surface: OSC 52 writes the clipboard, and
CSI can move the cursor and erase what is drawn, so a model-authored title could
otherwise repaint the question a human is answering and hide what is being approved.
Scope is the permission prompt only — ordinary streamed model output is printed
raw by _send_and_print and is unchanged, which is a surface-wide rendering
question rather than part of answering a permission request.
An authorization-boundary failure is itself a refusal. If the shared gate
raises, the CLI records gate_failed best-effort and rejects the request. If TTY
detection, prompt rendering, or prompt reading raises, it records prompt_failed
best-effort and rejects. The transport response is sent before any explanatory
notice, so a closed or unencodable output stream cannot leave the backend waiting.
This does not change cancellation: Ctrl-C still aborts the session and provider
teardown owns the pending request, because waiting on a possibly wedged rejection
transport would swallow the cancellation.
The audit write never runs on the event loop.sel() opens the audit log
and replays it to recover the running HMAC chain, so the first permission of a
fresh chat pays a filesystem cost inside the call — an unbounded one on slow or
corrupt storage. The decision coroutine shares its loop with the ACP reader and
stderr-drain tasks, exactly as the gate call does, so the audit is awaited
through asyncio.to_thread for the same reason: a blocking write here would
stop draining the backend and freeze the turn the audit is about. The one
exception is the cancellation teardown, which keeps the synchronous call because
awaiting anything there would swallow the CancelledError being delivered.
There is no "always allow". A persistent approval asks the backend to stop
sending permission requests for matching calls, and a request that is never sent
is a call this ladder never runs and never audits. The dashboard offers it
because its tool pipeline re-gates every call; this consumer does not.
The answer is matched exactly, not by first letter: a prefix match reads
abort as an allow. Blank line, EOF, and any unrecognised word all deny.
Denial is not an error: the request is answered, the turn runs to completion,
and the exit code is unchanged. Only the existing transport failures
(AcpTimeoutError, AcpError) exit non-zero.
Every decision is written to the SEL log before the matching
approve_tool/reject_tool, so a transport failure cannot erase a decision
already made:
decision
outcome
error
gate denied
denied
hook_deny
gate raised before a verdict
denied
gate_failed
execute-kind request has no verified command
denied
unverified_shell
nobody to ask
denied
noninteractive
prompt availability/render/read failed
denied
prompt_failed
user allowed
allowed
—
user allowed but the critical audit failed
denied
audit_unwritable
user denied / blank / EOF / unrecognised
denied
user_denied
session cancelled with the question open
denied
session_aborted
error is a stable machine code, never the gate's reason: a reason names the
path or command that triggered it, and an audit record must not restate the
thing it protects (log_tool_invocation does not redact for its callers). The
audited tool_name is the canonical _meta.kiro identity when the backend
supplies one, falling back to a redacted title — an audit trail keyed on prose
the model wrote can be steered by the model being audited.
Responses go through provider.approve_tool() / provider.reject_tool(). The
CLI never reads the advertised options: those ids are backend-specific and the
ACP layer owns the mapping.
stdin ownership
The prompt is read on an owned daemon thread through a private duplicate of the
stdin descriptor, not on the event loop and not through sys.stdin. The turn is
parked inside an active stream — the ACP runtime holds a reader task on the
backend's stdout and a drain task on its stderr — so a read on the loop thread
would stop draining those pipes until the human answers.
A cancelled prompt ends the session. A blocking terminal read cannot be
retracted: cancelling the await frees the coroutine, not the thread, and that
reader stays parked and takes the next line the user types. So cancellation
marks stdin poisoned and propagates; _require_usable_stdin then refuses at
every later entry point — a second permission prompt and the REPL's own you>
alike — rather than racing the abandoned reader for keystrokes. There is no
input broker and no recovery path: the abandoned reader can only outlive the
session, never compete with a live prompt.
Teardown must not await the backend. The request in flight is left
unanswered and audited as session_aborted, because answering it means awaiting
a transport: CancelledError has already been raised once and nothing
re-delivers it, so a wedged reject_tool would leave the Ctrl-C that asked for
the teardown unable to land. The provider is shut down with the session, so the
unanswered request dies with it. StdinPoisonedError carries the same rule.
That last sentence is a guarantee, not an expectation: provider.start() hands
back a live backend process, so _chat runs the whole message/REPL lifecycle
under try and calls provider.shutdown() from finally. A normal return, an
exception, and a cancellation raised through a permission prompt all tear the
backend down; without it a Ctrl-C at the prompt would exit leaving the backend
running with nothing owning it. Cleanup never swallows the cancellation — a
failing shutdown() is logged rather than raised, because replacing the
exception already propagating would discard the very CancelledError the
teardown exists to clean up after — and gc.collect() is nested inside its own
finally so a raising shutdown cannot skip it.
Context Tracking
After each message, checks provider.context_usage_pct():
>= autocompact_pct - CONTEXT_WARN_MARGIN_PCT: warning printed to stderr. Relative, not absolute: the compact arm is tested first, so an absolute warn level at or above the threshold would be unreachable
CLI compaction is blocking (single-user, acceptable).
Entry Point
console_scripts in setup.cfg maps kirocrew → kiro_crew._bootstrap:main.
Gateway asyncio child watcher
_install_child_watcher() runs once on the gateway command path only (not
chat, doctor, or any other subcommand) and must be called before
asyncio.run, on the main thread. It replaces CPython's default
thread-per-child ThreadedChildWatcher — whose os.waitpid reaper threads can
starve the event loop when many kiro-cli/MCP children die at once — with a
single-descriptor alternative:
Python 3.14+ is a deliberate no-op. CPython 3.14 removed
set_child_watcher, PidfdChildWatcher, SafeChildWatcher, and
ThreadedChildWatcher; the Unix event loop reaps children itself with a single
non-thread reaper, so the loop-starvation wedge this installer exists to prevent
cannot occur. The function short-circuits on hasattr(asyncio, "set_child_watcher") — probed by capability, not sys.version_info, so a
runtime that still ships the API keeps the mitigation. Without that guard the
Linux pidfd branch raised AttributeError and kirocrew gateway died before
binding its port, while every other subcommand kept working.
Linux gateway heap reclamation
The event-loop heartbeat offers a self-gating maintenance object a tick every
five seconds. At most once every ten minutes on Linux, it reads current RSS
directly from procfs. When RSS is at least 1.5 GiB, it asks glibc
malloc_trim(0) to return wholly-free heap pages to the OS and logs reductions
of at least 16 MiB. The probe and allocator call run in a worker thread, after
the dashboard socket has bound, so maintenance cannot delay readiness or block
the event loop. Healthy gateways remain below the threshold and do no
allocator-wide work.
The heartbeat waits at most two seconds for a pass. A timed-out worker keeps the
single in-flight slot until it exits, preventing repeated submissions or a
watchdog-triggering wait when the executor is saturated. Missing ctypes,
non-glibc libc, failed current-RSS probes, and rejected trim calls are
best-effort no-ops: reclamation must never stop the liveness heartbeat or make
the gateway unavailable. macOS and Windows are unchanged.
Live-target bootstrap
On the gateway command path only, immediately after _JAILED_COMMANDS
attestation and before the --seed handler:
maybe_reexec reads the live-target pointer (config_dir() / "live_target.json")
and, when it names a different checkout, os.execves into that checkout's own
kirocrew binary. This runs before anything is written to $KIROCREW_HOME,
before the gateway lock is acquired, and before any socket is bound — so exec'ing
away leaves nothing half-done. It is fail-safe: an absent, unreadable,
malformed, or stale pointer (missing binary, same image already running, or
KIROCREW_LIVE_EXECED marker already in env) causes the function to return, and
the currently-installed build boots normally. A bad pointer can never leave the
host with no gateway.
Gateway only — a plain CLI invocation (kirocrew doctor, kirocrew chat, etc.)
keeps running the install the user typed, not a worktree someone made live.
MCP tools: @kirocrew-cron and @kirocrew-core in tools, allowedTools, and mcpServers — auto-fixes missing entries, except from an instance that must not own the shared agent home (_decline_shared_agent_home), which writes nothing and reports the skipped repairs and any ceiling-forbidden grant as issues
Global mcp.json: kirocrew MCP servers present with valid binary paths — auto-fixes stale paths
Python environment: checks Python 3.9+ availability and dependency installation. Every install command this step prints names the RUNNING interpreter and is gated the same way as step 8's (extras.pip_install_command_for behind extras.pip_install_channel_available), because each one is for a module this process imports, so a bare pip can resolve to an interpreter the gateway never imports from. The sqlite fts5 remedy is the one place the gate hides only part of the message: the pip command is withheld where it cannot run, while "use a Python whose SQLite was built with FTS5" always prints, since that is the only remaining fix in exactly those environments. The import path row reports whether the standard library resolves from the interpreter's own tree: stdlib_shadow.find_shadowed_stdlib() locates each probed stdlib name with find_spec (never an import) and reports one that a launch-directory, PYTHONPATH or site-packages entry supplied — a ~/concurrent/ directory in the home the service unit runs from, an unpacked backport on PYTHONPATH, a stdlib-named pip package. A shadow IS an issue and names the module, the resolved file and the sys.path entry with its class, because the failure it otherwise produces is a TypeError from deep inside asyncio with nothing pointing at the directory. The same probe runs at both process entries (python -m kiro_crew and the kirocrew console script) BEFORE the CLI is imported and refuses to start with exit 2 on a shadow, so doctor normally reaches this row only clean; the row then says whether the launch entry is on sys.path at all (-P / PYTHONSAFEPATH keeps it off, and every bundled launcher passes -P), which is the note that tells an operator why the same stray directory is harmless from one cwd and fatal from another. Detection is fail-closed towards "not shadowed": an entry the classifier does not recognise as one of those three roots is never reported, so an unusual layout can only escape the check, never be refused by it
Vector memory (in-process embeddings): vendored llama-cpp-python runtime importable, embedding model file present (downloads in background on gateway start; when absent, a light HTTPS-reachability probe of the resolved model URL runs); embeddings are always-on (embeddings: ✅ always-on). On platforms with no vendored native libs (_platform_libs_dirname() returns None, e.g. darwin/x86_64 — Intel Macs or a Rosetta interpreter), the runtime line reports ⏹ unsupported platform … — memory uses keyword search and is NOT counted as an issue (designed degradation per embeddings.py); only a load failure on a supported platform flags embedding runtime. When that failure is an INCOMPLETE shipped payload, doctor additionally names the absent files (Missing native libs for <platform>: …, from embeddings.verify_vendored_libs()) and says it is a packaging defect rather than an unsupported platform — the two are indistinguishable in ctypes' own Shared library with base name 'llama' not found, which reads as an architecture problem and misdirects diagnosis. When LLAMA_CPP_LIB_PATH is set, doctor reports THAT directory as the thing to check instead (mirroring the loader's exemption): the libs load from there, so blaming the bundled tree would send the operator to reinstall a package they are deliberately not loading from. A faiss: line reports whether the optional FAISS accelerator is importable — never an issue on any platform (episodic recall falls back to the stdlib cosine scan); when absent it suggests installing faiss-cpu and prints an Install: line naming the RUNNING interpreter (extras.pip_install_command_for), because a bare pip on a packaged or minimal install resolves to an interpreter the gateway never imports from, so the wheel lands out of reach and the next run repeats the same advice. That line is printed only where the command can actually run (extras.pip_install_channel_available, shared with the dashboard's install card): on the bundled desktop interpreter, on an interpreter with no pip module, and on a PEP 668 externally-managed interpreter outside a venv, doctor names no command at all, because a pip install into the code-signed bundle breaks later launches and is discarded on the next app update
Speech-to-Text (optional): recognizer, selected model and audio-decoder presence when STT is enabled. Supported desktop releases bundle and build-gate the recognizer plus the pinned imageio-ffmpeg executable; a missing decoder there is a corrupt payload whose remedy is reinstalling Kiro Crew, never Homebrew/Winget/Apt. Source installs use kirocrew[voice] for the recognizer and a fixed system FFmpeg path for compressed audio. Windows preserves its non-fatal ⚠️ marker for an optional-extra gap so enabled-by-default STT cannot block gateway startup; on macOS/Linux a missing active runtime still flags an issue.
Slack credentials (optional)
Discord (optional): the channel's enabled flag, whether a bot token is present (never any part of its value), the three allow-lists, the privileged Message Content intent, the live connection, and the install URL. Blocking issues are enabled-without-a-token, an empty discord.allowed_user_ids (the transport fails closed, so every message is denied while it is empty), a thread or channel allow-list with Message Content OFF, and a reachable gateway whose Discord connection recorded a connect_error. The intent state comes from discord/intent_probe.py: one read-only GET /oauth2/applications/@me that decodes Discord's application-flags bitfield as a tri-state per intent PAIR (enabled / limited / disabled, since a limited grant still delivers the data) and degrades to unknown on any failure rather than aborting the report. Granted-but-unused Server Members / Presence intents are hardening notes, never issues. The install URL comes from discord/install_url.py, the OAuth-authorize analogue of Slack's app manifest: named permission bits OR'd to 309237711936 for a thread-capable install (the number discord-integration.md publishes), and none at all for the recommended DM-only install
WhatsApp (optional): printed whether or not the channel is enabled, because a channel that is invisible in the preflight is the failure this section exists to catch. When enabled it reports the optional neonize extra, checked with find_spec and never imported (importing it loads a ~19 MB ctypes CDLL plus protobuf descriptors, and a health check must not initialize the subsystem it inspects, nor construct a client), and whether the linked-device session store exists at <data home>/whatsapp/session.db, resolved from the same expression the channel opens it with so the two can never describe different files. A missing extra IS an issue: the channel is enabled, cannot start, and the fix is one offline pip install. An absent store is a ⚠️ note and never an issue, because pairing is a QR scan served BY the running gateway, so failing here would break the documented kirocrew doctor && kirocrew gateway chain at the one moment the operator has to start the gateway to make progress. Group membership is not knowable offline, so the section reports the configured count and the gateway logs the unmatched JIDs on connect
kiro-cli connectivity
Gateway running status
Loop-stall crash dump attribution: when a dump with thread stacks is less than 7 days old, the section prints the wedged main-thread frames and then an attribution: block from stall_attribution.attribute_dump(dump, config_dir()), read from disk with no gateway involved: the surface the loop was serving (named by the OUTERMOST recognised frame of the wedged stack, so a cron turn that passes through the Slack gateway module still reads as cron), the gate frame when the stall is inside the permission gate, and — for a cron surface — the job, joined by PID to the in-flight markers under <data home>/cron-running/ (cron_inflight: written when a run starts executing, cleared on every finally, so a surviving marker whose PID is dead is a run a hard exit interrupted). Exactly one marker matching the dump's # PID: names the job and yields recommended: kirocrew cron pause <id> plus the job's current pause state from crons.json (job_pause_state_from_disk); several markers list the candidates and say one cannot be named; none says so; a non-cron surface is named and implicates no job. Every line is evidence or the one action it supports — the check never guesses a job. An attributed job is also added to the issues summary, unless the dump is superseded (below).
A dump is superseded when a later LOCAL session's own pre-created, still-header-only file sits after it (crash_dump_store.dump_superseded): that session started and has not wedged, so the stall describes a past incident rather than the running gateway. A superseded dump prints its stacks and attribution unchanged and adds one line saying a later session started after it, but contributes NOTHING to the issues summary — neither the dump line nor an attributed job — so one stall stops producing a ❌ Fix these issues verdict for the following week on a gateway that has been healthy since. Ordering alone does not decide it, because the shipped Restart=always unit replaces a wedged gateway within seconds and its successor's header file would supersede every stall almost immediately: the downgrade additionally requires the stall to be at least _STALL_CURRENT_SECS (24 h) old and to be the ONLY stack-bearing dump on record (dumps_with_stacks() < 2), so a recent stall and a repeating wedge both stay current faults. A successor counts only when it is readable, header-only, and names this PID domain — a newer dump carrying stacks is a second stall, and a shared data home's foreign-domain file says nothing about a local restart.
Cron job health: names cron jobs that auto-paused after repeated failures (Fix: kirocrew cron resume <id>) and jobs whose last run errored while still scheduled (Fix: kirocrew cron trigger <id>), with an aggregate count each and the named list capped at 5 plus a +N more tail. Read-only — doctor never resumes or triggers a job, because an auto-pause after _AUTO_PAUSE_THRESHOLD consecutive failures is usually load-bearing and lifting it silently would hide the problem the run is meant to find. The scan is cron.unhealthy_jobs_from_disk(), which reads crons.json directly rather than via the gateway API: the dashboard's per-job err badge and the gateway's hourly failure re-alert both run inside the gateway, so neither can report a wedged one, and that out-of-process property is the point of this check. A job the user paused explicitly is never reported (only auto_paused is a health signal, and both flags can be set at once since pausing an auto-paused job preserves auto_paused). Silent on a healthy store and on a fresh install with no crons.json. A store that EXISTS but yields no readable job list — unparseable, non-UTF-8, wrong shape, or holding no usable record — is reported instead, because the scheduler can load nothing from it and every job has stopped; reporting that as a clean bill of health would reproduce the silence this check exists to break. No corruption aborts the run
Pod user-bus reachability (optional): on Linux, runs the same systemctl --user is-system-running probe used by runtime.require_backend() at pod verb entry. Low-level unit queries keep cheap in-process gates and do not re-run the probe. Probe and unit operations resolve systemctl through platform_compat.trusted_system_bin() and spawn only that absolute path; a PATH executable is ignored, and no trusted executable is an unclassified fail-closed error. A provably absent bus address returns no user session bus without spawning systemctl; this pre-spawn check is the only source of backend absence. Every spawned failure is an unclassified operational error except a positive permission-denied match, which reports sandboxed away. Neither may become PodBackendAbsent. The row retains any raw systemctl diagnostic. Permission denied names a generic outer sandbox and tells the operator to run pod commands from a host shell. The check is advisory and never changes doctor's exit code because pods are an optional development feature. On macOS, Windows, or a host without systemctl, the row is not applicable.
Live-target pointer (Linux only, silent when fit): reports a live_target.json that will make the launcher REFUSE every agent spawn — a symlink (dangling or resolving), a non-regular file, or a second hard link. sandbox._materialize_live_target_mask_target is fail-closed on those shapes because a bind mask covers a NAME rather than an inode, so any of them leaves a writable path to the bytes the gateway execves into inside every agent namespace (the full contract is in security, Live-target pointer). This section exists because that refusal is correct but arrives too late and in the wrong place: the operator's only notice was a failed spawn plus a logger.warning in the gateway log, and the shapes are ORDINARY operation for something else on the host — cp -al, rsnapshot and other hard-link snapshot tools raise link counts on config files, and a dotfile manager may keep the pointer as a link into its own tree — so nobody did anything wrong and the first symptom is that every agent stops starting. The read is sandbox.live_target_pointer_unfitness(), which classifies the pointer the way a spawn would WITHOUT spawning, and doctor prints that function's own sentence rather than a paraphrase, so an operator who sees this line and later hits the refusal reads one diagnosis instead of two (pinned by the drift guard in test_live_target_pointer_visibility.py). An unfit pointer IS an issue and changes doctor's exit code — but only on a host that actually confines a spawn: _materialize_live_target_mask_target runs on the launcher path, which wrap_argv reaches only when it WRAPS the child, so on a host that hands back an unwrapped command the pointer is unfit and NOTHING is currently refused. The question is asked with sandbox.credential_mask_applies, which lives beside those branches so a caller whose argument depends on the mask cannot drift from them, and which counts BOTH unwrapped outcomes — the off tier and a host with no available backend — where reading the configured mode alone sees only the first. Doctor still reports the pointer there, as a latent outage that begins the moment the host starts confining spawns, but drops the "spawns will be REFUSED" wording, says only that this pointer is not what stops a spawn here and points at the Sandbox section for whether agents start at all — the predicate answers False for a host that hands the command over unwrapped AND for one with no backend, which may be refusing every spawn for its own reason, so claiming "nothing is refused" would be a second false statement in place of the first — and does not count it as an issue; claiming a current outage there would be the same false promise as reporting this on macOS, and confinement is simply the second axis of that one rule. A predicate this process cannot evaluate fails towards reporting the refusal. A probe that itself fails reports could not check and is not an issue, because doctor must not turn its own failure into a verdict about the host — that includes a pointer that EXISTS but cannot be lstat'd, whose OSError propagates out of the probe rather than reading as fit, so an unclassifiable pointer is never rendered as silence the operator would read as a clean bill of health. Every value doctor prints in this section is terminal-escaped by the formatter that built it (see security, Live-target pointer), because a symlink target is attacker-chosen bytes. Linux only: a macOS Seatbelt profile denies by path rule and needs no mount target, so naming it there would promise a spawn outage that cannot happen. Absent is fit — the launcher publishes the absent-equivalent stub for it.
Masked credential leaves (Linux and macOS, silent when clean): reports a masked leaf whose bytes are a credential and which carries a second HARD LINK, because the launcher refuses every agent spawn on that shape — a mask binds a PATH rather than an inode, so the second name reaches the same bytes unmasked for reading and for writing (the full contract is in security, the masked-leaf alias pass). It exists for the same reason as the Live-target pointer section above and answers the same objection: cp -al, rsnapshot and other hard-link snapshot tools raise link counts on files in the home as ordinary operation, so the condition arrives without anybody doing anything wrong and the first symptom is that agents stop starting. The read is sandbox.masked_credential_leaf_aliases(), which stats the leaves without spawning and reports a data home it cannot resolve as nothing rather than as a fault. A leaf whose every other name was LOCATED is masked for each spawn and nothing is refused, so doctor prints the find -samefile remedy alone (sandbox._masked_leaf_alias_search_hint); a leaf with a name the pass could not locate is the one a spawn refuses on, and doctor prints that leaf's own refusal sentence from sandbox._masked_leaf_multilink_detail rather than a paraphrase, so an operator who reads this line and later meets the refusal reads one diagnosis instead of two. Both go through _print_wrapped so the remedy survives on one line. An aliased leaf IS an issue and changes doctor's exit code only on a host that actually confines a spawn (sandbox.credential_mask_applies, which counts both unwrapped outcomes — the off tier and a host with no backend — exactly as the pointer's section does) AND only for an unlocated name in the live data home; elsewhere it is reported as a latent outage that begins the moment the host starts confining spawns, without the "will be REFUSED" wording and without counting as an issue. An unreadable mode reports the leaf rather than hiding it. A probe that itself fails prints could not check and is not an issue, because doctor must not turn its own failure into a verdict about the host. Linux and macOS, because both launch paths run the refusal: a Seatbelt profile denies by PATH rule, which says nothing about a second hard link to the same inode, so sandbox_exec_argv calls the same pass the namespace launcher does and a macOS operator meets the same refusal. Windows has no confining launcher, so the probe is skipped there rather than promising an outage that cannot happen.
Hook auto-approve platform scope: reports whether a name-based auto-approve can be satisfied on this host, read off the same name_grant.platform_scope_notice and name_grant.windows_environment_refusal helpers name_grant.name_grant_refusal consults, so the row cannot describe a posture the check does not hold. Three answers, all printed with the source explanation because the only other trace is a decline line in gateway.log per invocation and what a user would have to read to tell the states apart is the source of name_grant: ⏹ declined on this host (windows_lookup_not_modelled) when platform_scope_notice names the one Windows host state the model cannot run in — Windows could not report where the user's Documents folder is, so whether a PowerShell profile runs before the command cannot be established, a property of the host and not the user's configuration; ⚠ declined while a PowerShell profile exists (ambiguous_env) when a per-user PowerShell profile is present at one of the paths derived from Documents, printed with the file that is doing it so the user can act on it (kiro-cli starts the shell without -NoProfile, so that profile runs before every command and a function it defines resolves ahead of any program on PATH — the same threat BASH_ENV poses on POSIX); and ✅ name grants can be satisfied on this platform otherwise, which is what macOS and Linux always print because neither can enter the Windows-only refusals above. Never an issue and never part of the exit code: each fail-closed answer is the intended posture, since the POSIX tokenizer does not preserve a backslash path and cmd.exe searches the command's own directory before PATH, so resolving names loosely there is the planted-shim attack the refusal exists to stop. The same platform-scope classification bounds the log on the three surfaces that consult it: name_grant.should_log_decline gates the human-facing warning to once per session, for a code that describes the platform rather than the command, at the agent hook webhook, the dashboard chat runner and llm_helpers. The Slack gateway and handler, the messaging driver, channel, the subagent runner and the task runner call name_grant.log_decline without consulting that gate, so those warn on every decline. name_grant.log_decline itself writes the SEL audit row for every decline on every surface, because declining is a security decision and only the human-facing line is ever deduplicated.
One flag is a MODE rather than a check and short-circuits before the list above runs, because it does not answer "is this install healthy": --bundle collects the redacted diagnostics zip and prints no health report. It ends with a prefilled Open a GitHub issue link whose version, channel dropdown answer and create-time channel: label all come from ONE resolver, release_channel.provenance() (shared with the dashboard's Report-a-problem link, which prints the same fields plus the free-form ones): the version is the PUBLIC release version — a repackager's four-part BUILD_VERSION stamp such as 0.7.0.5 is folded onto 0.7.0 (changelog.release_of_build), while the release pipeline's own spellings (0.7.0-insider.4, 0.7.0rc7, a nightly stamp) are the public identity and stay — and the channel is the lane the build can PROVE from three sources that must agree: the $KIROCREW_HOME/channel record when it names a lane, a prerelease marker in the installed distribution's metadata version when it describes the same release (a repackaged insider build keeps its rc marker there after the stamp has removed it from __version__; PLAIN metadata proves nothing, because the desktop lanes pip-install the checkout and stamp only __version__, leaving pyproject.toml's bare version in the dist-info of every lane), and __version__ itself unless it is a build stamp, which says nothing about its lane. No claim, or claims that disagree (a stable record over a promoted rc build, a lane switch not yet updated onto), prefills the form's own Not sure and attaches NO channel: label, because a create-time label outlives whatever the reporter later picks and a guessed Stable is exactly how packaged insider reports arrived mislabelled (#13168). The raw stamp, the distribution version and the record are recorded in the bundle's private versions.txt and manifest.json instead. Doctor is otherwise read-only, which is why the ledger cleanup sweep is its own command (kirocrew ledger-sweep, see the command table and session-work-ledger) rather than a second mode here.
Where each row lives.cli_doctor._doctor is the orchestrator: it prints the sections above in one fixed order, threads one issues list through every section, and turns that list into the closing ❌ Fix these issues: line with exit status 1, or ✅ Kiro Crew is ready!. Sections print as they run rather than being collected first, so on an interactive terminal a probe that hangs still leaves every earlier row on screen. Most rows live in kiro_crew.doctor_checks, one module per family: render (the escaping, indent and wrapping helpers the sections share), agents, mcp, confinement (the sandbox verdict and the shapes that refuse a spawn), access (session signing, hook auto-approve, credentials), install, services, resources, workload, channels and features (vector memory, speech-to-text). cli_doctor.py keeps the rows the orchestrator composes itself (Platform, and the Dependencies, Agent, Runtime and Connectivity rows built from the values it threads onward) and the rows a repository gate pins to that file: the process-spawning probes test/test_spawn_audit.py keys by cli_doctor.py::<function> (including the node, venv-interpreter and kiro-cli --version rows _doctor spawns itself), the agent-spec reads test/test_agent_spec_hardened_reads.py inventories for that file (Model, model pins, MCP Governance), the MCP Tools repair and probe, the KAS relay rows whose ACP imports the agent-sdk boundary baseline counts, and the embedding-model URL probe the redactor registry names. kiro_crew.cli_doctor stays the one import path and patch target: a family reads every function, class and value the facade binds through the facade at call time (a module is one shared object, so its attributes are what a test patches, and a section that imported a name inside its own body keeps doing so), every name the module held before the move still resolves on the facade (a moved one as the family's own object, with a write there forwarded to the family, so mock.patch(..., create=True) must not target one), and a section this move extracted from _doctor is patched on its own family module. The families load only when a health report runs, never on import kiro_crew.cli or for --bundle.
Update Command
kirocrew update pulls the latest source and rebuilds:
git fetch, then git reset --hard <oid> from KIROCREW_PROJECT_DIR, where
<oid> is origin/<branch> resolved ONCE right after the fetch (git rev-parse --verify origin/<branch>^{commit}). Every later judgment — "already
up to date", the divergence counts, the interpreter floor below — and the
reset itself name that pinned commit rather than the ref, so a fetch run
concurrently from another terminal can move the ref but never what this
command checked and applies. A pin that cannot be resolved (or times out)
is refused with a non-zero exit before anything else is judged.
The reset only runs for a FAST-FORWARDABLE checkout — behind its upstream
and not ahead of it (git rev-list --count --left-right HEAD...<oid> shows behind > 0, ahead = 0) — mirroring the
dashboard check's verdict, because the hard reset discards committed local
work and the uncommitted-changes prompt does not cover it. A DIVERGED
checkout (both sides non-zero) is refused with a non-zero exit and a
rebase-or-merge instruction; --force is the explicit opt-in that lets the
reset discard the local commits. An ahead-only checkout has nothing to pull
and is reported as up to date without resetting (even under --force — the
flag lets a real update discard diverged work, it does not delete commits
when there is nothing to update to). An unreadable comparison refuses
(fail closed). Uncommitted tracked changes still prompt before being
discarded — and because that prompt makes the gap to the reset unbounded,
the divergence count is re-taken immediately before the reset and refuses
commits that appeared while the update was waiting (committing the listed
edits in another terminal to rescue them is the natural response to the
prompt, and is exactly what would otherwise be reset away). Only HEAD can
move in that window, so the re-check needs no second fetch.
Interpreter floor of the pinned revision, judged before the tree moves.
After the fast-forwardable verdict and before the uncommitted-changes prompt,
dep_sync.incoming_python_floor_breach() reads requires-python out of the
pinned commit itself (git show <oid>:pyproject.toml, then setup.cfg —
never the working tree, which is still the OLD revision) and compares it to
the interpreter this command runs under. A floor the venv does not meet
refuses with a non-zero exit and a remedy naming the venv, the interpreter
the revision wants (uv venv --python <floor> --seed <venv>) and the
reinstall command, and the checkout is left where it was. Step 3's pip install -e . enforces the same floor, but by then the reset has already
moved the tree to code this interpreter cannot import — the running gateway
keeps serving the old revision from memory while every lazy import reads
the new files, and every later run repeats the reset and the refusal. A
revision with no floor file does not fire; a floor file git could not READ
(IncomingFloorUnreadable: unresolvable ref, git failing, timeout) refuses
too, because on a pinned revision "could not read" is the one way left for
that stranded state to be re-admitted. The same gate guards POST /api/update and the gateway's unattended auto-apply.
Rebuilds the dashboard via build_frontend_sync() (npm; non-fatal on failure).
Non-fatal also means non-destructive: npm ci deletes node_modules before it
installs, so the tree is moved aside and restored unless the install succeeds
(node_modules_txn.NodeModulesBackup). A failed install therefore costs the
rebuild, not the dependency tree -- which matters because the registry needed
to rebuild one is usually what was unavailable. The gateway's unattended
auto-apply takes the async sibling and is protected the same way.
Reinstalls backend via pip install -e .
Client Port Resolution
kirocrew token / status / logout / stop / restart must find the port
the gateway is actually bound to. port_resolution.resolve_client_port()
(re-exported by cli_server) resolves it
in this order, first hit wins. The MCP stdio servers (mcp_core /
mcp_computer) resolve their gateway API base through the same helper —
lazily, on the first gateway call, and cached for the process lifetime — so a
loopback callback and a client CLI command always agree on which gateway they
are talking to:
An explicit --port N flag (0 counts — the check is is not None).
KIROCREW_PORT, when it parses as an int. Deliberately above the bound
export: an explicitly-set KIROCREW_PORT is how a caller retargets a
child at a DIFFERENT gateway — pod exec builds a client env with
KIROCREW_PORT=<pod-port> while the inherited KIROCREW_BOUND_PORT
still names the spawning live gateway, and the bound value outranking it
would walk pod token/status/logout into the live plane.
(build_pod_env additionally scrubs KIROCREW_BOUND_PORT outright.)
KIROCREW_BOUND_PORT, when it parses as an int — the port the parent
gateway actually bound, exported the moment the port is reserved
(dashboard.server._reserve_dashboard_port; _export_bound_port
republishes it once the site serves). Never persisted —
service_environment() deliberately does not capture it.
A port explicitly written in dashboard.url. A portless URL
(http://my.host) is not a port choice: parse_dashboard_url()
substitutes 5476 for the server's benefit, so the client re-splits the URL
and only accepts the port when it was actually named.
The sole gateway-owned run-marker. A running gateway records
<data-home>/run/gateway-<port>.bin (see
kiro_crew.instances.run_marker, written for the SSH token-mint), so its
filename already advertises the port. A client with nothing configured reads
the marker names — never the file contents — and uses that port. Two guards
keep it from being a guess:
Ownership, not reachability — clear_marker() only runs on graceful
shutdown, so a crash leaves a stale marker behind and an unrelated process
may since have bound that port. Because _token / _logout send
X-Local-Secret to whatever answers, a bare "is something listening" probe
would walk the local secret into that process. A command-line check is not
enough either — argv is attacker-chosen, so a listener started as
/tmp/kirocrew gateway would pass it. _gateway_owns_port() therefore
requires three things, none sufficient alone:
the pid recorded in run/gateway-<port>.pid (written 0600 inside the
0700run/ dir, which is on the is_sensitive_path floor, so neither
another local user nor an agent file tool can write it);
that pid must be among the pids listening on the port
(platform_compat.find_listening_pids), which is what makes a stale
recorded pid harmless;
that pid must be owned by the calling uid
(platform_compat.process_owner_uid) and look like a KiroCrew process.
The uid check is what closes pid recycling into a foreign user's
process; argv is retained only as defense in depth.
It fails closed at every step: no sidecar, an unparseable pid, a pid
that does not hold the port, an unresolvable uid, a missing lsof /
netstat, or a throwing lookup all deny, and discovery is skipped. A
same-user attacker is out of scope by construction — they can already read
.local_secret under their own uid.
On non-POSIX platforms the step denies outright.process_owner_uid
cannot report an owner on Windows, and a KIROCREW_HOME writable by another
user would let them replace both the marker and the sidecar with a forged
listener — the file-permission argument that carries requirement 1 stops
holding there. So discovery is skipped rather than approximated: Windows
users keep --port / KIROCREW_PORT, exactly where they were before this
fallback existed, so nothing regresses.
Ambiguity — with several gateways up there is no basis to pick one, so
the step refuses, prints the candidate ports and the --port /
KIROCREW_PORT hint to stderr, and falls through.
_DEFAULT_PORT (5476).
Steps 3 and 5 are what make a single gateway started on a non-default port
(kirocrew gateway --port 6776) reachable from a bare kirocrew token with
zero configuration; before it existed, the client hit a dead 5476 while the
marker naming the live gateway sat unread. Config-load, URL-parse (including a
non-string dashboard.url, which raises TypeError rather than ValueError),
and discovery failures all degrade to the next step — a client command never
dies on a bad config or an unreadable data home.
Because restart resolves a port and then polls it for readiness, it passes the
resolved port to the detached replacement (_spawn_detached_gateway(port)). The
child re-resolves independently, so without that the replacement could bind 5476
while the parent waited on the discovered port.
The marker is written for every dashboard-serving gateway, including a
source-tree python -m kiro_crew launch with no console script beside
sys.executable: in that case the .bin file is written empty, which is inert
for the token mint (its shell clause requires a non-empty executable path) but
still advertises the port for discovery. The pid always goes to the separate
.pid sidecar — never into the marker, whose contents mint cats and execs. A
--slack-only gateway serves no dashboard, so it writes no marker — there is no
client port to discover.
Writing a marker also prunes markers naming other ports. A gateway is a
singleton per data home (gateway.lock), so any other port's marker is residue
from a run that crashed before clear_marker() could fire. Unpruned, they
accumulate one per port ever used and each costs every client command an extra
listener lookup, making discovery slower the longer a dev box churns ports. The
live gateway is the only writer and knows which port is current, so it is the
right place to reap them; pruning is best-effort, and the ownership check still
rejects anything it misses.
CLI→gateway requests are built against the literal 127.0.0.1, never the name
localhost. On a dual-stack host localhost can resolve to ::1 first, and the
listener verification is address-agnostic (lsof -ti TCP:<port> cannot tell an
IPv6 squatter from the real IPv4 gateway), so a name-based URL could deliver
X-Local-Secret to a socket other than the one that was verified. The URL
printed for the browser still uses resolve_dashboard_host() (localhost) —
that must not change, because the SPA's per-origin localStorage is keyed on it.
Stop Command
kirocrew stop [--port PORT] stops a running gateway:
If a systemd/launchd service is active and the caller did not pass
--port explicitly (see Service Management), stop it via the service
manager and return — without this branch, SIGTERM-by-port would be
racing the manager's auto-restart.
Otherwise (no service active, or --port was passed explicitly to
target a non-default dev gateway): platform_compat.find_listening_pids(port)
to find PIDs — lsof -ti TCP:{port} -sTCP:LISTEN on POSIX, netstat -ano
parsing on Windows (there is no lsof there; this previously made
kirocrew stop a no-op on Windows). Both binaries are resolved through
platform_compat.trusted_system_bin() — the fixed system directories, never
PATH, which on a gateway can lead with same-uid-writable dirs — and a name
that does not resolve there counts as absent rather than falling back.
listening_pid_tool_available() performs the same pinned resolution, so it
distinguishes "no listener" from "lookup tool missing" without disagreeing
with the lookup it describes. The pinned set is the FHS directories plus
/run/current-system/sw/bin, which is root-owned and rewritten only by a
system rebuild. A tool installed anywhere else — a Homebrew or conda prefix —
still reads as absent, and deliberately so: those prefixes are writable by the
invoking user, which is the exposure the pin exists to close.
trusted_system_bin() logs a warning once per name when the tool is on
PATH but not resolvable under the pin, and tool_outside_trusted_dirs()
lets stop name where the tool actually is rather than tell an operator who
already has it to install it. That case carries SEL
reason=<tool>_outside_trusted_dirs, distinct from <tool>_not_found, so
the two are separable in the audit log.
When the trusted lookup tool exists but returns no pid, stop makes one
fixed-loopback POST /api/shutdown request carrying the gateway generation's
local secret. That handler independently requires loopback origin and a
constant-time secret match, then triggers the same graceful shutdown event as
SIGTERM. The acknowledgement body is capped at 4 KiB before JSON parsing;
excessive nesting is treated as malformed input. A successful acknowledgement
ends the stop command; a missing secret, refusal, malformed response, or
transport failure retains the ordinary no-gateway diagnostic. This fallback never converts the pid sidecar into
authority to signal a process.
platform_compat.process_command_line(pid) to verify it's a KiroCrew process —
/proc/<pid>/cmdline (Linux), ps -o command= (macOS), Win32_Process.CommandLine
via WMI (Windows). The Windows venv kirocrew.exe re-execs python.exe, so the
match is on the command line (-m kiro_crew gateway / \Scripts\kirocrew.exe gateway),
not the image name. _args_look_like_kirocrew parses the command line structurally
and keys on a server subcommand plus a recognised module name. The module name is
matched against port_resolution._gateway_module_roots(): always kiro_crew, plus
the top-level module of every installed kirocrew.plugins entry point, so a composed
edition whose launcher execs its own module with -m classifies without the core
knowing any edition's name. When a listener exists but none of the pids classify,
stop names the pid(s) — with the cmdline basename when cheap — and exits 1 with SEL
reason=unrecognized_listener, rather than reporting "no gateway running" for a port
that is occupied.
Before that refusal, the authenticated shutdown is tried once more, with SEL
reason=argv_declined_listener (the empty-lookup attempt above carries
reason=listener_lookup_empty, so the two paths stay separable). argv is the
weakest identity a gateway has — every spawn shape has to be taught to the
patterns, and the desktop app's is not among them — while the per-generation
secret is published by the gateway at startup whatever its command line reads.
Answering the request shuts down the answerer, so no pid is guessed or
signalled on that path, and a listener that does not answer keeps the
unrecognized_listener refusal exactly as before.
The request is made only once the process that will RECEIVE the secret is
proven to be this port's gateway (_verified_loopback_gateway_pids), because
the secret mints owner tokens and reachability is not identity. The proof is
the one port_resolution._gateway_owns_port documents, minus its argv step —
which that contract itself keeps as defense in depth rather than proof —
plus the start identity and the address: the pid and .start token recorded
in run/gateway-<port>.pid (0600 inside the 0700 run/ dir, on the
is_sensitive_path floor), that token still matching the live pid's, that pid
among the ones a 127.0.0.1 connect actually reaches
(platform_compat.loopback_owner_pids, which mirrors the kernel's
most-specific-bind dispatch, so a gateway bound to another address cannot
vouch for a process squatting loopback), and the pid owned by this account.
Every step fails closed, and the path denies outright off POSIX where
process_owner_uid reports no owner — the same boundary _gateway_owns_port
draws. restart reuses that proof: when a listener was found but the argv
filter named no incumbent, it resolves the answering pid BEFORE the stop
(afterwards the identity is gone) and waits for it, since who must release the
port before a replacement can bind is answered by who answered, not by argv.
Terminate each verified PID: os.kill(SIGTERM) on POSIX; taskkill /T /F
(via platform_compat.kill_process_tree) on Windows so the gateway's detached
children are reaped too. Liveness is probed with platform_compat.pid_exists
(a raw os.kill(pid, 0) would terminate the process on Windows).
If a systemd/launchd service is active and the caller did not
pass --port explicitly, ask the platform to restart it. On Linux:
systemctl restart kirocrew.service in every scope whose unit is running
— sudo systemctl for the system unit, systemctl --user (never sudo) for
the per-user one (single
atomic operation, smaller down-window than stop+start, and the
supervisor stays in charge of the lifecycle the whole time). The exit
code of systemctl restart is not the verdict: the unit is Type=simple,
whose start job completes the moment the process is forked, so a gateway
that exits on start still gets exit 0. service.linux.restart() therefore
re-reads the unit (systemctl show) for _RESTART_SETTLE_SECS (2 s, one
read per 0.25 s) and returns a per-scope RestartReport: a scope is
restarted only if the unit is active at the end of the window; one seen
in activating (auto-restart), failed, inactive or deactivating
inside it is reported at once with that state and its Result (exit-code,
start-limit-hit, …); activating (start) is waited for until the
deadline. A non-zero systemctl restart is classified by the unit's state
right after it — still up or unreadable: a REFUSAL (the manager did not run
the job), reported with the manager's diagnostic and the restart command
for THAT scope (common.system_restart_command_hint() /
common.user_restart_command_hint(), sudo systemctl restart kirocrew /
systemctl --user restart kirocrew); otherwise the job ran and failed
(failed, activating (auto-restart)), reported as NOT UP with that scope's
journal command. The shared common.restart_command_hint() the update path,
the Slack restart-failure hint and the install-time credential warning print
picks the same way, by two stats and no spawn: only the system unit file
present → the system command; only the per-user unit file at the remedy's
location → the user command; neither, or both (a file says nothing about
which scope runs) → the service-aware kirocrew restart. The gateway's
async update-failure handler awaits it in a worker thread
(asyncio.to_thread): the per-user location is under the account's home,
which can be a network mount whose stat blocks for as long as the mount is
disconnected, and on the loop thread that wait would freeze chat and the
liveness heartbeat together.
controller.manual_restart_hint() stays the unconditional system command on
systemd (it must never answer the kirocrew restart that just failed).
When the report is not ok
and a unit is still there, the command exits 1 with one line per scope
acted on — a scope that restarted reads restarted, a refusal names its
hint, a unit that did not stay up names that scope's journal command — and
never falls through to the listener path. The headline is per scope too:
with a unit in both scopes (a stale crash-looping system unit beside the
working per-user one) it reads Restarted kirocrew service in the user scope; the restart did not take in the system scope, never "the gateway
was NOT restarted", which is false for the gateway the operator uses; the
exit code is still 1 because a scope needs a hand, the shape
service uninstall gives a teardown that finished in one scope. On
macOS: launchctl unload <plist> + launchctl load <plist> (no
-w, so persistent enable state is unchanged). The deprecated
launchctl restart is avoided because under KeepAlive it behaves
like stop (SIGTERM + immediate respawn) and never re-reads the plist.
Otherwise (foreground gateway, no service, or --port passed
explicitly to target a non-default dev gateway):
platform_compat.find_listening_pids(port) (lsof on POSIX, netstat
on Windows) to detect a running gateway. If found — OR if the lookup
tool is absent (not listening_pid_tool_available(), so a missing
tool is not mistaken for a dead gateway) — run the existing _stop
path. When the trusted lookup tool exists but returns no pid, restart
independently attempts the authenticated shutdown request. An acknowledged
request always refuses the immediate replacement and tells the operator to
retry after shutdown completes: neither an absent pid sidecar nor a live
same-user pid from that sidecar can prove singleton-lock release, because a
stale pid may be recycled and exit before the incumbent. The second restart
sees no incumbent and safely starts the replacement.
If no incumbent evidence exists (e.g. the user runs restart after a
crash), skip the stop step rather than erroring — the user expects to end
up with a running gateway either way. The _stop call is wrapped in a
try / except SystemExit so a TOCTOU race (gateway exits between the
listener check and _stop's own lookup → _stop calls sys.exit(1))
does not abort the restart before the spawn.
When a listener IS present but none of the pids classify as a Kiro Crew
gateway, _stop declines it (see step 3 above) and the incumbent list is
empty, so restart resolves the incumbent through gateway.lock instead
(_incumbent_from_lock_holder): a live holder that did not acknowledge the
shutdown is refused by _refuse_lock_holder and no replacement is spawned.
A gateway booted from a module the classifier does not recognise lands exactly
there — it holds the lock and stop could not signal it — so recognition, not
the restart path, is what makes the ordinary stop → wait → spawn path run. A
foreign listener that holds no lock still leaves the spawn to proceed and fail
on its own bind.
Spawn a detached kirocrew gateway via subprocess.Popen, stdin set
to subprocess.DEVNULL, and stdout + stderr redirected to
~/.kiro/crew/gateway.log (the same file the kirocrew logs command
tails for foreground gateways). Detach is per-platform: POSIX uses
start_new_session=True; Windows uses creationflags=DETACHED_PROCESS | CREATE_NEW_PROCESS_GROUP (there is no setsid) — both via
platform_compat. The shell returns immediately and the user can
follow logs via kirocrew logs -f.
SEL audit event logged with via=service or via=fork pid=<n> so
the audit trail distinguishes the two paths.
Service Management
kirocrew service {install,uninstall,status} registers the gateway
with the OS service manager so it survives SSH disconnects, restarts
on crash, and starts on boot. Implemented in src/kiro_crew/service/.
Linux (current_platform() == SYSTEMD):
Unit file: /etc/systemd/system/kirocrew.service (root-owned).
Install: sudo install writes the unit, then sudo systemctl daemon-reload && sudo systemctl enable --now kirocrew.service.
Privilege is resolved per call: already-root (euid 0) skips sudo
entirely — required on minimal container / root-login images that
ship no sudo binary — and a non-root caller with no sudo fails
with a clear ServiceInstallError rather than an uncaught
FileNotFoundError.
The gateway runs as User=$USER Group=$(id -gn). Every elevated
executable is a stock system program, not a kirocrew one; the module
docstring in service/linux.py sits next to the call sites, names the
current set, and records what escalating the AppArmor step's
interpreter does and does not guarantee. security
carries the reasoning behind that step's four tools.
Environment: values are captured from the installer's environment
into the unit's Environment= lines at install time
(service_environment() in service/common.py) — this is how
KIROCREW_PORT=5477 kirocrew service install binds a non-default port.
The unit also reads EnvironmentFile=-/etc/kirocrew/kirocrew.env, an
operator-editable file the installer seeds create-if-absent (a reinstall
never clobbers edits). systemd applies the file AFTER — and overriding —
the baked Environment= lines, so editing it and running sudo systemctl restart kirocrew changes a value (e.g. the port) without reinstalling.
Uninstall removes the file and its /etc/kirocrew directory.
Credentials are deliberately NOT captured. Both baked locations are
world-readable — the unit lives in root-owned /etc/systemd/system and the
override file is installed 0644 — so a model credential placed there
would be readable by every local user on the host. service_environment()
therefore carries no credential — its only installer-derived values are
PATH, KIROCREW_KIRO_BIN and KIROCREW_PORT (it also returns HOME,
LANG and LC_ALL) — and a test pins the absence so a future "just
propagate it" change fails.
Consequence: a KIRO_API_KEY exported in the installing shell does not
reach the service, the readiness probe (which forwards that variable from
the gateway's own environment) sees no credential, and unless a
kiro-cli login credential store under the baked HOME supplies one
instead, the dashboard reports a signed-out state on a host where kiro-cli
itself is authenticated. install_service() prints a warning naming the
variable and the remedy when it detects that case, and ~/.kiro/crew/.env
is the supported home — load_credentials() reads every key from that file
into the gateway environment at boot and forces 0600 on it first. The
warning is diagnostic only: it is non-fatal by construction, since the unit
is already written and started by the time it runs. kirocrew doctor reports
the same condition next to its kiro login line — the one output where the
contradiction is visible, since that line runs whoami with the inherited
environment and reports signed in. Doctor's report is gated on a service
definition existing (installed_unit_path() — the system unit file, or,
when it is absent, the per-user unit the calling account's own manager
has loaded, read from the FragmentPath it reports via
linux.user_unit_path(), which answers only for a LOADED unit whose
canonical Id is kirocrew.service — a blank or malformed show
answer, a mask, an unparseable unit (bad-setting / error) and an
alias yield nothing, so doctor never reads as the definition a file the
manager runs nothing from; the same gate feeds doctor's managed-marker
check, since render_unit(user_scope=True) bakes the same
Environment= lines): without one the gateway runs
in the foreground and inherits the invoking shell, so the credential does
reach it and a warning would be a false positive. It is advisory only —
never appended to doctor's issues, which is the exit-code channel — since
that gate establishes a definition on disk, not that the serving gateway
lacks a credential; a fall-back login store, or a stopped unit beside a
foreground kirocrew gateway, both leave the host healthy while the check
fires.
Boot survival via WantedBy=multi-user.target (no linger needed —
that's a user-service concept; this is system-level).
A home another gateway already serves is not retried.kirocrew gateway takes <home>/gateway.lock before it binds anything. One
predicate decides whether a refusal is the kind a restart cannot heal:
GatewayLock._serving_verdict in gateway_lock.py, the serving-holder
predicate, asked about the process /proc/locks positively identifies as
the lock's acquirer (of the lock file, or of the home directory when the
file has been deleted or replaced) and never about the pid the lock file
merely records. It carries four conjuncts, each measured once: that
acquirer is alive; it holds the configured dashboard port (one listener
enumeration, platform_compat.find_port_listeners, filtered to the
owner's own sockets); it holds it AT the address the probe reaches — the
address this gateway is configured to bind (KIROCREW_BIND when it parses
as an IP address; an absent override or the IPv4 wildcard 0.0.0.0 is
probed at 127.0.0.1, the IPv6 wildcard :: at ::1 because the
dashboard binds it IPV6_V6ONLY and a v4 connect never reaches it) — with
the owner's wildcard binds covering that address by family only
(0.0.0.0 covers any v4 host, :: covers v6 hosts and not v4 ones), and
an unreported address or family counting as unknowable, never as covering;
and it answers HTTP there. The address conjunct is what ties port ownership
and HTTP health to ONE process: port ownership alone is address-agnostic,
so without it a stranger answering at the probe address on the same port
would be credited to a lock owner bound elsewhere. Only all four make
GatewayLockError.live_holder True — a sibling gateway serving the home,
which this one can displace on neither front while it lives — and then the
process exits gateway_lock.LIVE_HOLDER_EXIT_CODE (78, EX_CONFIG: two
supervisors pointed at one home is a host configuration, and the remedy is
to change it); the unit's RestartPreventExitStatus= names that code, so
a Restart=always unit goes failed once with the refusal line in the
journal instead of relaunching every RestartSec against a refusal the
sibling keeps permanent (bounded by StartLimit* on the shipped unit,
unbounded on a unit without them). Every other refusal exits 1, which the
unit's Restart=always relaunches, because a later attempt can find it
cleared or because the evidence for standing down is missing: a lock file
replaced faster than it can be locked, a home that cannot be opened or
measured for directory locks, an flock whose acquirer is gone (a wedged
inheritor holds it until that process dies), a live acquirer that does not
hold the port (a sibling still starting, or one shutting down that has
closed its listener and releases the lock next), a live acquirer on the
port but silent at the probed address — a wedged gateway (a hung process
keeps its listening socket bound, and a terminal exit would leave the unit
failed with nothing left to relaunch once that process dies) or, the
residual row the predicate leaves unasserted by design, a gateway that
holds the port only at another address than this one would bind, or whose
socket addresses the platform could not report (probing the holder's own
listener address instead is the follow-up tracked in the PR's
deferred-finding issue, not a guess made here); the message names what was
measured (holds port N, not answering HTTP at 127.0.0.1, or holds port N at ::1, not at 127.0.0.1) and says the refusal is not treated as
permanent — and a holder no surface could identify — no /proc/locks
(macOS, Windows), or a Linux filesystem whose device numbers never match
the lock table (btrfs subvolumes, overlayfs) — where the pid the lock file
records may be alive and on the port and still be a reused number rather
than the process that holds the lock; the predicate is never asked about
it, and that refusal's message says the holder cannot be confirmed and that
a supervisor, if one manages the gateway, will retry it (the same message
on every platform, since macOS and Windows reach this branch on every
refusal). On the identified-acquirer paths one listener enumeration and at
most one HTTP probe (its own 1.5 s budget) feed both the message and the
exit status, so they cannot disagree; the orphaned-lock path (a dead
acquirer, candidate openers) keeps its own per-candidate facts. With no
port to weigh (--port auto, --slack-only) the verdict cannot be reached
and every refusal stays restartable. The constant is defined once, beside the
refusal in gateway_lock.py, and both cli.py (the exit) and
render_unit() (the exemption) import it; test_service.py pins the
rendered directive to the constant and the constant to 78, because the
value is baked into installed units, which service install writes once
and no upgrade re-renders — a unit written by an earlier build keeps
relaunching on this refusal until it is re-rendered. An existing install
picks the directive up by re-running kirocrew service install.
Two scopes, both visible.install writes the system unit only, but
the SELinux refusal hands the operator a per-user unit
(render_unit(user_scope=True), managed with systemctl --user), so
every OTHER verb accounts for both scopes and names the one it reports on:
status prints system scope: … then user scope: …. One guard is
shared by every verb that acts on the name (_UnitState.ours): the
canonical Idshow answers must be kirocrew.service. An operator's
Alias=kirocrew.service on their own unit makes show, stop, restart,
disable and the unlink on our name act on THAT unit, so an alias is
refused whole by uninstall, is never selected by stop() / restart(),
and is not counted by is_active() / is_up() — the headline still names
it (active (running), an alias of shared.service). Two predicates,
deliberately distinct: is_active() is the REACH predicate — true when
either scope RUNS our unit, any ActiveState but inactive / failed,
so a unit crash-looping through activating (auto-restart) counts — and
stop() / restart() act on each scope that satisfies it, selected
by the same systemctl show the status verbs read rather than an
is-active probe, which answers non-zero for activating and would let
kirocrew stop issue nothing at a flapping unit (kirocrew stop /
kirocrew restart therefore reach a user-scope gateway through its own
manager, with no sudo; a stopped unit in the other scope is never started
on the side, which selecting on "installed" would do). is_up() is the
HEALTH predicate — ActiveStateactive or reloading in either scope,
exactly what systemctl is-active exits 0 for — and the
kirocrew service status exit code follows it, not is_active(): a
crash-looping unit is one stop must reach and one status must report
as down (exit 1, headline activating (auto-restart)), because a script
gating on that exit code would otherwise read a gateway that never
started as healthy. restart() returns a RestartReport — per scope,
ok only when the unit is active at the end of a 2 s settle window
(the Type=simple start job succeeds at the fork, so systemctl restart
exiting 0 says nothing about the process), a refusal carrying the
manager's diagnostic and that scope's own restart command; see the
Restart Command section. uninstall()
tears down the system unit under sudo and the user unit through
systemctl --user, and returns an UninstallReport
naming what each scope got. Every verb in either scope rests on ONE
decision, _owned_unit(): the unit the manager reports under our name is
Kiro Crew's when the file that holds its definition — the reported
FragmentPath, followed through a symlink to its source — is the file the
installer writes for that scope (/etc/systemd/system/kirocrew.service;
~/.config/systemd/user/kirocrew.service) or carries the installer's own
marker line, Environment="KIROCREW_SERVICE_MANAGED=1" (the line every
rendered unit has, and the one kirocrew doctor reads); a definition under
another name is never ours. A unit that is not ours — a distribution's
unit under /usr/lib loaded because no file of ours shadows it, an
operator's own unit systemctl --user linked under our name, a link at the
installer's path to a file the manager reports by its resolved target —
reads left untouched (not installed by Kiro Crew: …) with nothing
stopped, disabled or unlinked, and the scope finishes (exit 0). A unit of
ours reached through a link — the fragment itself a symlink, or the
installer's path a symlink resolving to it — loses its link entries and
keeps its definition: `removed (the link to ; the linked unit file
itself was kept)`; a direct file of ours is unlinked. The order
in either scope is stop → disable → verify inactive (a second `show`) →
remove → `daemon-reload`, the system scope's removal an `rm -f` under
sudo (which removes a symlink's entry, never its target) and the user
scope's an `os.unlink` as the calling user. A step the manager refuses
(`RefuseManualStop=yes`, a bus that went away mid-run, a declined
sudo), or a unit still running after `stop` returned 0, ends that
scope's teardown before anything is unlinked, its line
reads `left in place (\`… stop kirocrew.service\` failed: …; the unit
file was not removed)`, and the controller marks that line and exits 1 —
deleting the file would leave a unit that is still loaded, possibly
still running, with no unit to find it by while the report read
`removed`. Two shapes are refused WHOLE, before any verb: an alias (in
either scope — `show ` answers for the unit the name resolves to,
so `stop` / `disable` / unlink on our name would act on that unit) and a
unit that is RUNNING under any load state but `loaded` — masked at
runtime (`systemctl mask` leaves a running unit running and reports its
fragment as the mask), edited into an unparseable state, or `not-found`
with its file removed under it — reported as `left in place (an active
unit whose load state is masked: unmask and stop it first, then run
\`kirocrew service uninstall\` again)`. A refusal is unfinished (exit 1)
when it leaves a running unit or the module's own file at
`/etc/systemd/system/kirocrew.service` behind; one that leaves neither —
an inactive alias in the user scope, or in the system scope with no file
at that path — is a report (exit 0). A system
unit file the manager answers for but does not have loaded and that is
NOT running (a file dropped without a `daemon-reload`, an inactive mask,
an unparseable unit) runs nothing and is removed outright — the mask
(`systemctl mask`'s symlink to `/dev/null` at exactly that path, a mask
of OUR name) by unlinking the entry, which unmasks the name. A system
manager the shell cannot reach is decided by `sd_booted(3)`'s test,
`/run/systemd/system`: absent, no manager runs on this host (a container
where systemd is not PID 1, where `install()` leaves the file behind when
its `daemon-reload` fails) and the stale file is removed; present, the
unit may well be running and the line reads `left in place (the system
manager is not reachable from this shell: …)`. The absence of the unit
file is not the manager's word either: `uninstall()` asks the system
manager (an unprivileged `show`) before calling that scope `not
installed`, so a unit still LOADED and running after `rm
/etc/systemd/system/kirocrew.service` is stopped (checked) and
`daemon-reload`ed — with no file there is nothing for `disable` to read
or for the unlink to remove, and the line reads `stopped (its unit file …
was already gone, so nothing was disabled or unlinked; daemon-reload
run)` — one a `daemon-reload` left `not-found` but running is refused
whole as above, a not-loaded, not-running one with no file (a runtime
mask) is reported as `left in place (load state …)`, and an unreachable
manager on a systemd host reads `not reachable from this shell (…)`
(nothing left behind, so not unfinished). The report carries its outcome
as structure — `removed`, the scopes whose unit is gone from the manager
(file unlinked, link removed, or the fileless unit stopped and
`daemon-reload`ed), and `unfinished` — and the controller's headline
reads `removed`, never a line's prefix. A host without `sudo` does
not abort the call: the system line reads `left in place (privilege
unavailable: …)` and the user scope is still torn down. A unit file that
cannot be removed after a successful stop is
reported the same way on its own scope's line. The user-scope teardown
rests on the same ownership decision: a definition that is neither the
file at `~/.config/systemd/user/kirocrew.service` nor one carrying the
marker is `left untouched (not installed by Kiro Crew: …)`, a not-loaded
file of ours at that path is removed, a mask of the name at that path
(`systemctl --user mask`) is unlinked — `removed (the mask …;
kirocrew.service is unmasked)` — and a mask or an unparseable unit
anywhere else is the operator's to inspect, reported as `left in place
(…)` with nothing in that scope
stopped, disabled or deleted. Load state comes from `systemctl show -p
Id -p LoadState -p ActiveState -p SubState -p FragmentPath -p Result` (parsed as
`Key=value`, which systemd 219 supports — `--value` does not exist there),
because `is-active` answers `inactive` both for a stopped unit and for a
scope with no unit, which is what makes a running user-scope gateway read
as `inactive (dead)` in a report that asks the system scope alone. Every
`systemctl --user` the module spawns runs with
`service.common.systemctl_user_env()`, the one resolver the pod runtime's
`systemctl --user` spawns also use: it backfills `XDG_RUNTIME_DIR` and,
when the bus socket exists, `DBUS_SESSION_BUS_ADDRESS` from the account's
runtime directory (never overriding a value the caller set), so a shell
descended from a system unit — the gateway's own — reaches the account's
manager and a `not reachable` reading is a host fact, not a
missing-variable artifact. A user
scope this process cannot
see is reported as `not reachable from this shell ()`, never as
inactive, and `uninstall` leaves it alone: a root process whose
`SUDO_USER` names a human is decided in-process without spawning
(`systemctl --user` there would answer for root's manager, not the
human's), and any other failure — no session bus, a stripped environment,
a sandbox — is systemd's own non-zero exit with no `LoadState` printed,
reported with its diagnostic and never classified by stderr text.
Logs are read from the journal: sudo journalctl -u kirocrew -f,
or unprivileged if the user is in systemd-journal / adm.
Install: launchctl load -w <plist>. RunAtLoad=true and
KeepAlive ensure auto-start and crash recovery.
Stdout and stderr are written to
~/Library/Logs/KiroCrew/gateway.{log,err}.
Other platforms: install/uninstall return exit code 2 with a
message pointing to manual setup.
kirocrew stop is service-aware: if the service is active it calls
the platform's stop instead of SIGTERM, so the manager does not
immediately restart the gateway under us.
Logs Command
kirocrew logs [-n LINES] [-f] tails the gateway log from whichever
source is most appropriate:
the USER journal (journalctl --user -u kirocrew.service) when the
calling account's own manager has the unit loaded
(service.linux.user_unit_installed()) and it is either the only unit or
the one running (user_unit_active(), asked only when a system unit file
also exists — a stopped system unit left by an earlier install must not
win over the per-user gateway that replaced it). Read without privilege;
there is no sudo rung, because the user journal never needs one, so an
empty probe falls through to the next source. The user probe runs
journalctl --user … --quiet -n 1: a journal with no matching entries
prints -- No entries -- on stdout with exit 0 (systemd 252), which is
not a row and must not exec a tail of nothing — quiet suppresses the
notice so an empty journal reads empty.
systemd journal if the system service is installed on Linux. Tries
unprivileged journalctl first; falls back to sudo journalctl
only if the unprivileged probe returns no rows. Unchanged from before
the user-journal source above was added: a probe that printed anything
— the unit's rows, or journalctl's own notices — is exec'd unprivileged,
and only a probe that printed nothing takes the sudo rung (a TTY prompt,
or the "stdin is not a TTY" refusal).
launchd stdout file if a plist exists on macOS and that file is
non-empty. Both conditions matter: the platform probe reports launchd
on any macOS host, and an install that never started the agent leaves
a 0-byte log behind, so either check alone would capture the command
and tail nothing.
~/.kiro/crew/gateway.log for foreground gateways
Uses os.execvp so signals (Ctrl+C) propagate naturally to the
underlying journalctl/tail process.
Dashboard Self-Update
On gateway startup and every 12 hours, a background task runs git fetch and
reports an update when EITHER the checkout can be FAST-FORWARDED
(git rev-list --count --left-right HEAD...@{u} shows commits behind and none
ahead) OR the pull already landed and only a restart is missing (HEAD equals
its upstream while the on-disk __version__ outranks the one this process
imported).
Commit distance is the primary signal because __version__ is bumped only at a
release: comparing version strings alone reported "you're on the latest version"
to a checkout hundreds of commits behind main, for as long as the next bump
took.
Both conditions are narrower than "is it behind", because available is also
read by an unattended apply. GatewayOrchestrator._auto_apply_update applies
git fetch + git reset --hard <oid> with no prompt — <oid> being
origin/<branch> resolved once after the fetch, the same pin kirocrew update
takes, so what it floor-checks is what it resets to — so:
"Behind" alone is true both for a checkout that is purely behind and for a
DIVERGED one carrying its own commits, and the second would have those commits
reset away. Only a fast-forwardable checkout is offered an update.
The version signal is required to come with HEAD == @{u}. A checkout that
pulled a version bump and then committed on top is ahead, so its upstream
still reads newer than the imported version; without that requirement the
same reset would drop those commits.
The unattended apply is triggered by the version, not by commit distance.
The check reports version_newer beside available, and the git branch of the
auto_update path requires both. available is what the dashboard shows, and
on commit distance alone that would mean any upstream commit — resetting a
source checkout within 12 hours of one, where the version-only verdict only did
so at a release. Requiring both keeps that path firing no more often than
before. Commit distance without a version bump lights the dashboard badge, and
POST /api/update is the non-destructive way to apply it.
POST /api/update refuses before it moves the tree, and fast-forwards to a
pinned commit rather than pulling. In order: a dirty tracked tree is 409
dirty_tree; a pre-apply git fetch that fails is 409 git_fetch_failed
(500 on timeout); a divergence count it cannot read is 409 git_read_failed;
a diverged checkout is 409 checkout_diverged; then @{u} is resolved ONCE to
a commit OID (git rev-parse --verify @{u}^{commit} — unresolvable is 409
git_read_failed, a timeout 500 git_read_failed); then the interpreter floor
of THAT commit is read with git show <oid>:pyproject.toml (then setup.cfg)
and compared to the gateway's own interpreter — a floor the venv fails is 409
python_floor with the remedy in error, and a floor git could not read is
409 git_read_failed, never waved through. Only then does the handler answer
{"ok": true, "status": "updating"} and start the worker, which runs git merge --ff-only <oid> — the same OID, so the revision that was floor-checked is the
revision that lands, and an upstream rewritten in the window fails the merge
instead of minting a merge commit — then rebuilds, pip install -e ., and
restarts. A merge that fails or times out ends the worker with an error step
whose detail names the command and the pinned OID; that detail is what the
overlay's failure card shows and what its "Ask the agent" hand-off sends.
The unattended auto-apply and kirocrew update apply the same floor gate
(see the Update Command above); all three refuse with the checkout untouched.
Topbar shows 📦 v0.1.3 badge — click to check and view changelog
If newer version found: badge turns into "📦 Update Available"
Clicking opens a dismissible changelog modal with rendered markdown
"Update Now" button: POST /api/update → fast-forward to the pinned commit →
rebuild → os.execv() restart. A 409 renders inline in the modal through
ErrorNotice; on a voluntary update the notice offers "Ask the agent" (the
hand-off closes the modal for this page session and opens chat with the
refusal), on a mandatory one it does not (the modal's enforcement is staying
up) and the installer command remains the way out
Health indicator shows "Updating…" during the process. A failed or error
progress step is terminal for both the modal and the full-screen overlay: the
modal drops its restarting latch and shows the reason in the same slot a
synchronous 409 uses; the overlay replaces the spinner with a failure card
(ErrorNotice, "Ask the agent" + Dismiss) instead of stalling until the
five-minute stuck timer
SSE auto-reconnects when the new process starts
Status Command
kirocrew status queries the running gateway's /api/status endpoint
and prints uptime, sessions, messages, tool calls, subagents, crons, lessons,
and a memory line: the gateway's own resident set (gateway_rss_mb) and the
per-session tree ceiling the cleanup watchdog recycles at
(watchdog_rss_max_mb, spelled out as disabled when 0). Both fields are
published by /api/status for this line; kirocrew doctor prints the same two
readings at the top of its Memory Pressure section, on every platform.
App Dev Mode
kirocrew app dev <name> [--off] [--confirm-out-of-install-root] toggles an
installed App Kit app into (or, with --off, out of) dev mode, which speeds
the app-UI edit loop by serving UI files uncached and live-reloading the
dashboard on file change. The command writes the flag out-of-process; the
running gateway's watcher picks it up within one poll interval, so no gateway
restart is needed. Full App Kit developer docs live in
docs/app-kit/api-reference.md; the durable contract surfaces this feature
introduces are:
Persisted schema — installed.jsondev: bool (default false): a
per-app flag in each app's ~/.kiro/crew/apps/<name>/installed.json. Tolerant
on read (absent ⇒ false), reversible, no migration. This field is the
authoritative source of truth for watching and no-store serving; enabling
also records a gateway-owned operator grant binding the ui root's resolved
path, which is what authorizes serving a ui root outside the install directory
(see api-reference.md). Builtin apps cannot enter dev mode.
Endpoint — POST /api/apps/{name}/dev, body {"enabled": <bool>},
returns {"name": <name>, "dev": <bool>}. Behind standard gateway auth;
emits an app_dev_mode SEL audit event. 400 for a non-boolean body, a
builtin app, an unsafe app name, or a refused grant (a sensitive ui root is
never grantable; an out-of-install root always answers 400 over HTTP —
only the CLI flag, run on the gateway host, supplies the confirmation,
because a request-body flag from the dashboard origin is app-controllable);
404 when the app is not installed. The flag is additionally operator-only
against agent shells through three tiers: the builtin rule
self-protection-dev-mode-out-of-root-confirm (literal text plus its argv
floor) refuses any agent shell command carrying it; the flag's consumption
point performs a runtime human-vs-agent check that refuses a process
showing evidence of agent-shell confinement
(dev_mode_operator_attestation_required) — closing runtime flag
synthesis, which no command-text scan can see; and the grant record
(~/.kiro/crew/apps/.dev-grants.json) is sealed read-only inside the agent
OS sandbox
(materialized at gateway startup so the seal always has a target), so a
sandboxed process cannot write a grant at all — any grant-touching toggle
from such a process is refused up front
(dev_mode_grant_record_readonly; use the dashboard toggle instead). Only
the operator's own terminal can supply the attestation — and both the
confirmed grant and each refusal emit a SEL event
(dev_mode_out_of_install_grant for the permission decision,
dev_mode_grant_write for the sealed-record refusal).
WebSocket event — app_reload, payload {"app": <name>, "ts": <float>},
broadcast when a dev-mode app's ui/ tree changes; the dashboard reloads that
app so edits appear immediately.
Serving behavior: while an app is in dev mode the gateway serves its UI
with Cache-Control: no-store; otherwise the standard revalidation header
applies.
An internal, unstable sentinel cache under ~/.kiro/crew/apps/ mirrors the set of
dev-mode apps so the zero-dev-apps steady state costs one stat() per second.
It is a derived cache reconciled from installed.json at watcher init (under a
cross-process lock, atomic with concurrent toggles), not part of the App Kit
contract — its path and format are internal and may change without notice.
Computer Use Commands
kirocrew computer {doctor [--json] | apps | call} — hand-rolled dispatch
rather than argparse subparsers, because the parent CLI forwards REMAINDER
(see computer-use.md).
doctor reports, in order: whether the platform is supported (macOS today;
Windows and Linux report a typed refusal), whether the keystone primary enable at
~/.kiro/crew/computer_use.json is on, and the macOS TCC probe
(AXIsProcessTrusted() + CGPreflightScreenCaptureAccess()). The probe is
advisory and never a gate: macOS attributes a grant to the responsible
parent of the process tree, so both rows can read missing while a
full-fidelity capture succeeds — observed live. doctor therefore prints a
responsible_hint naming the process a user should actually grant (the packaged
app, or the terminal that launched a dev gateway) and says outright that "not
detected" does not always mean unavailable. It never calls
CGRequestScreenCaptureAccess, which would pop a system dialog from a background
process.
--json is the machine form the gateway shells out to for the Settings
permission rows. That indirection is deliberate: a short-lived subprocess keeps
native ctypes out of the gateway, so a native fault cannot take down the gateway
and with it cron, Slack and the dashboard WebSocket.
apps lists on-screen applications resolved from
CGWindowListCopyWindowInfo (layer-0 windows only, never pgrep — a pgrep -n
lookup returns short-lived helper pids whose accessibility tree is empty). It runs
computer_list_apps through the SAME gated dispatcher as call, so it is refused
while the feature is disabled, in an unattended session, or under a policy that bans
computer use — the agent can run this command with bash, so an ungated version was
an unauthorized read of every window title.
call runs one tool — call computer_get_state app=Finder — or a whole
sequence in ONE process: call --calls '[{"tool":"computer_get_state","args": {"app":"Finder"}},{"tool":"computer_click","args":{"app":"Finder", "element_index":12}}]'. The batch form exists because element_index values only
resolve against the per-process snapshot cache that produced them, so two separate
invocations cannot share them. key=value arguments are JSON-decoded when they can
be (element_index=3 → int, screenshot=false → bool) and kept as text otherwise
(app=Finder). --json emits [{tool, text}, …]; the exit code is non-zero if any
reply carries the Error: prefix, and a batch runs to completion rather than
aborting at the first refusal.
call goes through computer_use.tools.dispatch_tool, the same chokepoint an
agent call traverses, so the primary enable, the target policy and the secure-field
floors all apply — it is a reproduction tool, not a bypass. Its session key is the
attended cli_chat surface, which is what the SEL audit records. There is no
separate diagnostics opt-in and no identity proof: the unattended-surface refusal
that made one necessary was removed along with the rest of the computer-use
governance model.
All three are human-facing. apps has an MCP twin (computer_list_apps) per
the MCP-first rule; doctor is a permission diagnostic rather than a capability,
so the rule does not bind it; and call adds no capability at all — it is a
harness over the eleven existing MCP tools, and deliberately has no MCP twin,
because a tool that runs other tools would let a model launder one per-call gate
decision into many. There is deliberately nokirocrew computer state <app> —
that would be a second, CLI-shaped spelling of an LLM-facing capability and would
have to be an MCP tool instead (it is: computer_get_state).
Gateway Test Harness
Four composable flags let an integration test or eval harness boot a gateway
deterministically, with no model and no developer-machine state:
kirocrew gateway --test-mode # bundle: ephemeral port + json-ready + reads approval
kirocrew gateway --port auto # OS-assigned port, avoiding a collision with a real gateway
kirocrew gateway --json-ready # print KIROCREW_READY:{port,token,pid,home} once listening
kirocrew gateway --approval reads # auto-approve read-only tools
kirocrew gateway --approval yolo # auto-approve ALL tools
--json-ready is what makes the harness race-free: the caller waits for the
KIROCREW_READY line instead of polling a port, and reads the token from it rather
than minting one.
--approval yolo refuses to start unless KIROCREW_HOME is explicitly set to a
non-default path. The flag disables every per-call approval, so pointing it at the
real data home would let a test drive an operator's live sessions and credentials.
The guard is a startup refusal rather than a warning because a warning in CI output
is not read.