Repository navigation
release: 0.4.6 candidate — updater repair set + 13 reviewed contributions - #328
Merged
Merged
Conversation
The first-empty-write guard set `wrote = true` when it DECLINED a write, so a renderer reload that re-sent the same empty snapshot got through on the second attempt and flattened a non-empty roster (seen live: declined backup at 20:38:47, file flattened two minutes later). The guard now only disarms when a non-empty write lands — proof the renderer actually holds a roster — so a later empty write is a genuine deletion. Adds a regression test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Auto mode dropped every sandbox because a hive worker writes to its agent folder under <harnessHome>/hive/agents/<id>/, outside the project cwd. That is a path problem, not a reason to run unsandboxed. - codex: -a never -s workspace-write instead of the full bypass; hive.ts adds the agent dir, hive root and palace via --add-dir. An explicit posture on the command line (including the old bypass flag) still wins. - claude: the per-session settings.json now enables sandbox.enabled with filesystem.allowWrite plus permissions.additionalDirectories for the same dirs. bypassPermissions only silences prompts; verified live that the sandbox still denies writes outside cwd and the listed dirs. - failIfUnavailable stays off so platforms without a sandbox spawn as before. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G5X7vQmQDwrTCn2qBzuBiJ
TELEMETRY.md and the public privacy policy both promised PostHog does not retain the IP on the event. It did: $ip was non-null on 221,393 of 221,393 events, with city on 198,809, postal code on 188,963 and lat/long on 221,343. Reality moves to match the promise, not the other way round. The app was not a passive victim of this. src/main/analytics.ts explicitly set `disableGeoip: false`, overriding the posthog-node default of OFF, and the comment beside it restated the promise it was breaking. Nothing in this codebase ever read a country, so the flag bought us nothing and cost city, postal code and coordinates on every event. Three collection surfaces, all sending to the same project: - app (posthog-node) — disableGeoip: true, and $ip: null per event - blog (posthog-js) — before_send nulls $ip, sets $geoip_disable - marketing site — same before_send PostHog fills $ip from the connection only when the event does not carry one, so sending it explicitly is the one client-side lever that exists for a server-set property. The project-level "Discard client IP data" toggle is still off and is NOT changed here — see the PR body. Both published promises are corrected to describe the stronger posture: no geolocation of any kind is derived. That is a tightening, not a softening. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0111uSG4Ge2U6nUxnea7ZfeY
A transient sidecar bind failure used to degrade a proxy-tier agent (Crush) for its entire session: one log line, no retry, nothing the user sees. - startProxyBridgeWithRetry: up to three attempts with a short backoff; each attempt kills the previous sidecar so nothing leaks. - Still failing after that is deliberate degradation and stays so (routing untouched, the CLI runs), but it is no longer silent: log.jsonl gets a proxy-degraded event, the renderer gets hive:degraded, the spawn result carries a one-line reason, and main shows it as a native toast on the same notifications gate the breaker uses. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01G5X7vQmQDwrTCn2qBzuBiJ
localStorage is partitioned by origin, so every hive the app opens shares exactly one. When a hive has no roster.json yet, the store falls back to localStorage unconditionally — and that localStorage still holds the last hive's roster. Opening a brand-new config therefore rehydrated the previous workspace's agents, archived and restorable entries, each still carrying the cwd of the workspace it was hired in, and the seed flush then wrote them into the new hive's roster.json. Stamp the hive the keys were written for (cth.rosterHome) and read the fallback only when it matches. An unstamped localStorage is still adopted once, so installs upgrading into this keep their floor, and the dev <-> packaged origin bridge the mirror exists for is untouched: it runs through roster.json, not through this fallback. Closes #236 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
When the GPU process dies — a driver reset, a TDR, a Chromium GPU crash — the floor loses its context and glRecovery rebuilds it 1500ms later. If the GPU process has not come back by then, getContext() returns null and Pixi reports it as 'This browser does not support WebGL. Try using the canvas renderer'. OfficeFloor painted that stack onto the floor and stopped, so a condition that clears itself in a second or two cost a full restart, behind a message that blames the browser for something the browser did not do. Classify that rejection and rebuild through the existing generation-bump path, budget held in a ref so it survives the rebuilds it schedules. Out of budget, show a note that names the actual cause and what to do about it. Anything not recognised as a context failure still reports its stack, so a broken theme or tileset is not hidden behind six seconds of retries. Note this is NOT the eviction path: a getContext() that pushes past Chromium's ~16-context cap evicts somebody else and succeeds, so the rebuild always wins its slot back. Verified both ways over CDP against the shipped v0.4.5 build. Policy lives in glRecovery.ts with the loss path, testable with no GPU or Pixi.
…o write it An ASK ME question was printed as raw text, so an agent writing *emphasis*, `backticks` or a bulleted list of options put those characters on screen and the card read as a wall of punctuation. It now goes through the same MarkdownPreview the file preview uses, via a new `card` variant that inherits the card's VT323 face and drops the document chrome, and keeps a single newline as a line break so a question with no markdown in it looks exactly as it did. The task detail's Q&A trail renders the same way. Rendering it is only half of it: nothing had ever asked the agents to format these. The orchestrator prompt and PROTOCOL.md now do, and PROTOCOL.md gains the ASK ME section it never had. No new dependency — react-markdown and remark-gfm were already here, and the soft-break plugin is twelve lines rather than remark-breaks. Still no rehype-raw, so agent-written HTML stays inert. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EzZ8eU93JhPtiPsKsDmVkJ
Add a shared hook event contract across the Electron IPC boundary. This introduces a typed hook event model shared between the main process, preload bridge, and renderer consumers. The runtime validator checks the existing payload shape without restricting provider event names or requiring an agent id that was previously optional. Changes: - Add a shared hook event type - Export the hook event type through preload - Validate hook event payload shape before IPC delivery Existing hook routing and provider event semantics remain unchanged.
Add regression coverage for the shared hook event contract. Tests cover: - known hook event payloads - forward-compatible provider event names - optional agent ids - malformed optional fields - main and preload usage of the shared contract This protects the IPC payload shape without constraining the existing multi-provider hook event flow.
…very not enqueue Fixes #255. God can sit 'blocked' on a human prompt for hours, which is exactly when the hourly compaction trigger fires. The trigger enqueued /compact unconditionally, so it got stuck at the head of god's queue forever (blocked never becomes deliverable on its own), and the dedupe check then silently swallowed every subsequent hourly attempt against the stuck item. Two changes: (1) fire() now checks canDeliveryToAgent before enqueueing, the same gate the drain itself uses immediately before typing, so nothing is queued in a state the drain would refuse to deliver. (2) the lastCompactUsed latch moves from enqueue-time to successful-delivery-time (via a new compactUsed field carried on the queued message), since writing it at enqueue recorded a compaction that may never happened, permanently latching out feature attempts at that token count.
renameAgent() (store.ts) -> hive.ts's renameAgent() persists a rename straight into registry.json. But useHive.ts's god-spawn effect rebuilt god's agent object from scratch on every respawn with the literal 'Michael' hardcoded in three places (the `hive:` spawn payload, the Agent object added to the store, and the /remote-control session name) instead of reading the persisted name back — so a rename reverted on every app restart even though the registry still had it right. Extracts the name-resolution into a small pure helper, src/shared/godIdentity.ts::resolveGodName(), mirroring the pattern already used elsewhere in this repo for reading god's persisted name (realtime/tools.ts's fleet-status tool: `reg.agents[godId]?.name`). The god-spawn effect now calls window.cth.hiveRegistry() (already used elsewhere in this same file, effect 0) before building the spawn payload, and passes the resolved name through all three sites instead of the hardcoded literal. Closes #279. Tested on macOS (Apple Silicon), Node 26.7. `npm run typecheck` and `npm run build` both clean. `npm run test:focused`: 551 tests, 550 pass; the 1 failure (test/update-download-asset.test.cjs) is a pre-existing artifact of running `npm install --ignore-scripts` in this sandbox (better-sqlite3's native rebuild fails against this machine's V8/Node headers, unrelated to this change) and reproduces identically on a clean checkout before this diff. Added 2 new tests for resolveGodName() via the existing load-ts.cjs harness, both green.
Same root cause as 46737e5, different call site. CommandCenterPanel.tsx's header caption had 'Michael runs the floor' hardcoded as a literal string instead of reading agent.name (already available as a prop, and already correctly used for agent.accent/character/status two lines above). This one isn't a restart-persistence bug like the first fix — it's wrong the instant you rename, no restart needed: rename via the Edit Agent panel and the sidebar/message-composer update instantly (they already read agent.name correctly), but this header stays on "Michael" regardless. Verified live with Playwright end-to-end (real onboarding, real spawn, real rename through the Edit Agent panel, real restart) — see the PR description for before/after screenshots and recordings.
Autonomy & Budgets kept its save button inline at the bottom of the panel while every other modal in the app puts the dismiss action on the left and the confirming action on the right of a shared footer (EditAgentModal, AddAgentModal, QuitWarningModal). Settings was the only modal whose footer held a lone close. The footer now renders the budget note and a save button around the untouched close button, both gated on the active section so the other six sections keep a lone close. Save uses variant="primary" and size="md" to match the footer's other action buttons rather than the sm/secondary it carried inline. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e values The renderer loads the config once at start-up and hands it down as a prop. Nothing told it when a write landed, so every view seeded from that prop kept rendering the start-up value until the app restarted. Settings was where it showed. Its fields are seeded from the prop at mount, and a re-seed effect re-reads only 14 of the 28 prop-seeded fields from disk, so the other 14 displayed the pre-write value on reopen. For the two free-text breaker limits that also lost data: a stale blank is indistinguishable from "cleared", so saveBudget wrote `undefined` and JSON.stringify dropped the stored key. The "who can add agents" toggle looked like it refused to change, because it flips relative to what is displayed — the first click after reopening wrote the value already on disk. Subscribing at persistConfig rather than writeConfig is what makes this cover every field: writeConfig is one of 23 callers in the main process, and Slack, freeflow and notifications each persist by their own route. The file write is the only point they share. resetConfig now goes through persistConfig too, so it announces like any other write; the bytes written are unchanged. The re-seed effect in SettingsModal is left in place. With the prop current it is redundant, but removing that duplication is a separate change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Code review caught the notification handing subscribers the raw object passed to persistConfig. The renderer feeds one `config` state from two sources — config:get and this notification — and config:get returns a deep-filled, home-normalised config while the notification did not. A patch touching one nested key persists a half-filled sub-object, so a component could read `contextTrigger.compact. minContextPct` as a number from a read and `undefined` from a notification a moment later. config.ts already documents that hazard on withTriggerDefaults. Subscribers now get the same shape readConfig hands out. The migration is deliberately left out of that path: it persists, and it has already been applied to the config the patch was built on. The two earlier tests asserted against writeConfig's return value, which is the raw object and disagrees with a read on exactly the partial-patch case this fixes. All of them now state one contract: subscribers see what a read returns. Setup writes a config and reads it back so the one-shot trigger migration is settled before any test subscribes, matching an app that has started up. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The notes added with this change ran long, explained mechanics the code already shows, and cited an issue number a future reader may not be able to look up. Rewritten to state the intent at a level anyone can follow, and to carry their own reason rather than pointing at a tracker. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…sed one Floor windows load the same renderer and keep their own copy of the config, so sending only to the primary left every other floor showing what it opened with — the same staleness this was meant to end. The primary is also whichever window was focused last, which made the recipient arbitrary. Verified with two windows open: a save made in one arrives in both. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Same root cause as the first two fixes in this PR, just more instances of
it: god's display name was hardcoded at ~30 more UI/prompt call sites across
CommandCenterPanel, SchedulesSection, TriggerHistoryTab, SettingsModal,
WorkersTab, TasksKanban, AgentHoldButton, the two "clocking in" boot screens,
DevicePicker, CompletionToast, and three main-process call sites (a desktop
notification title, a voice-action board post + ping subject, and the
PREP ASSISTANT system-prompt template).
Renderer fixes read the live agent from the store (`agents.find(a => a.isGod)`)
or use the already-in-scope `agent`/`a` prop, matching whichever pattern the
surrounding component already used. The two pre-spawn boot screens (rendered
before the store has god's agent object at all) share a new
useResolvedGodName() hook that reads the persisted registry name directly,
same as useHive.ts's own spawn effect.
Main-process fixes resolve the name once via hive.registry()/deps.hiveRegistry(),
through the shared resolveGodName() helper. hive.ts's injectedPrompt() needed
care: its own doc comment states a PROMPT-CACHE INVARIANT (the injected prefix
must stay volatile-free — no registry state re-read per turn). The fix respects
it: the PREP ASSISTANT's system prompt reads god's name ONCE, at that
assistant's own spawn time, exactly like its own name/id/dir/root already do —
not a live re-read within an already-running session.
Left alone, deliberately: the Office character roster entry ("Michael" as a
selectable avatar, same as "Jim"/"Pam"), ambient office flavor/easter-egg
dialogue, past-release changelog copy, onboarding-only text (shown before any
hire/rename is possible), the "Realtime Michael" feature's own branded name,
and internal protocol identifiers (`from:"michael-voice"`, the
`useRealtimeMichael` hook name) — none of these claim to reflect the live
god's current display name.
Verified: `npm run typecheck` and `npm run build` clean. `npm run test:focused`:
551 tests, 550 pass — the 1 failure (test/update-download-asset.test.cjs) is
the same pre-existing sandbox artifact noted in this PR's original commits,
reproduces on main before this diff.
A stable, always-current place to install or reinstall Munder Difflin for macOS, Windows, and Linux, independent of the in-app updater. Every button points at the newest GitHub release (releases/latest) so the page never goes stale between versions, and it links the release's SHA256SUMS.txt for optional verification. This is the one recovery path that does not depend on the app updating itself, so it works even for a client whose auto-update is stuck. Uses the site's existing design tokens, nav, footer, and theme toggle. No pricing, wall data, or index.html changed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014dMe8Mm2eas1SwvUiu3XXr
Restart-to-install could get wedged with no feedback, which is the recurring "The command is disabled and cannot be executed" seen twice in updater.log (13 Aug and 22 Aug, six restart clicks before one install landed). Two causes, both fixed here: 1. Re-entry. The handler used to abort any in-flight restart and fire a fresh quitAndInstall on every click. The second quitAndInstall hits a native command Squirrel has already disabled by the first, which is exactly the disabled-command error. A guard now refuses a duplicate while one restart is pending. 2. Silent failure. quitAndInstall reports a refusal through the autoUpdater error event, not a throw, so the handler's await never settled and the button spun forever with no way back, which is what drove the repeated clicking. A failed quit now settles the pending restart (failPendingRestart) so the renderer reports the error and restores its notice instead of hanging. The pendingRestart resolver now carries an outcome so both the cancel path and the failure path report the truth. abortPendingRestart keeps its signature (index.ts app:cancelClose still calls it). Extends restart-cancel.test.cjs with the guard and settle-on-error contracts; all 5 assertions green, typecheck:node clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014dMe8Mm2eas1SwvUiu3XXr
The founder's live 0.4.5 shows the toolbar badge stuck on "checking…". Root cause: runCheck awaits electron-updater's checkForUpdates() with no timeout. If the feed request opens but never responds (a stalled connection, a captive portal, a half-open socket after sleep), that promise never settles, so runCheck never leaves the "checking" state it emitted, the badge spins forever, and nothing is written to updater.log. electron-updater caches its in-flight check promise, so once one check hangs, every later check returns the same hung promise and the native updater is wedged until the process restarts. A spinner that never resolves is worse than an error: it looks like the app is working when it is not, and it hides the fix, because a wedged check will not detect 0.4.6 either. runCheck now races the check against CHECK_TIMEOUT_MS (30s; the feed is a few-hundred-byte YAML, so anything past this is hung, not slow). A timeout falls through to the existing catch, which logs it, shows an error state, and runs the notify-only fallback poll, so the badge always reaches a terminal state and the failure is finally visible in the log. Source-string tests in update-check-timeout.test.cjs pin the cap; typecheck:node clean. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014dMe8Mm2eas1SwvUiu3XXr
The founder's actual complaint: he clicks the version badge and nothing
happens. It is working. On the latest release a manual check succeeds,
finds no update, and the badge settles back to a quiet grey "latest"
chip with no toast, no motion, no confirmation. A check that worked
perfectly and a click that never registered are indistinguishable, so a
working button reads as broken.
The through-line across this whole bug was silence on success. This
makes success speak: after a MANUAL check that finds no update, the badge
flashes a brief positive acknowledgement ("You are on the latest
version. vX is the newest release. Checked just now.") that auto-dismisses
after a few seconds. It fires only for the manual no-update result, read
back from the settled status; an available update is already loud on its
own (the chip changes), and background 6h checks stay silent as before.
Deliberately narrow: it does not touch the available/download/restart
paths, which cannot be exercised until there is a real update to find
(the 0.4.6-rc). Make success visible first, let the rc show the rest.
Source-string tests pin the wiring (no jsdom harness in this repo);
typecheck:web clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014dMe8Mm2eas1SwvUiu3XXr
The auto-updater is only exercised by the NEXT release, and the paths that matter most (timeout to fallback, restart re-entry, the error-state link) are exactly the ones a clean successful release never touches. god asked for this to live in the repo next to the release steps, not in a chat message, because a checklist in a message is a checklist that gets skipped. Captures: the proving-hop logic and the hard release gate (0.4.7 must be a complete signed/notarized pipeline run, not a tag); the happy path a clean 0.4.7 proves on its own; the three fault-injection checks a clean release cannot reach; and the tested-vs-rc-only split as the reporting standard. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014dMe8Mm2eas1SwvUiu3XXr
…m RELEASE.md Two refinements from review: 1. The founder's plan rehearses the whole hop on prereleases (0.4.6-rc.1 -> 0.4.7-rc.1) BEFORE the real 0.4.6, so the checklist now describes that: same content, executed against the rc, which is the only packaged build that will exist, so it covers the happy path AND both fault-injection tests. Records why it is safe (verified allowPrerelease behaviour, prereleases invisible to stable clients on both paths) and the two version-maths gotchas (start on 0.4.6-rc.1; a machine left on 0.4.7-rc.1 needs a manual reinstall, no auto-downgrade). 2. A checklist a sibling file nobody links to still gets skipped, so RELEASE.md now references it as a required pre-tag step. RELEASE.md is the PUBLISHED release body (body_path in release.yml), so the reference is an HTML comment: seen by the runner editing the notes, invisible to users. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014dMe8Mm2eas1SwvUiu3XXr
… download (Option B)
The founder's actual complaint was that clicking the badge "does not
update" — because the badge only ever offered a manual download (a DMG
and a drag-to-Applications card), never the native download/restart the
Settings pane runs. When 0.4.6 lands he would click, get a DMG, and
rightly say auto-update still does not work.
Now, when the native updater has staged an update, the badge drives the
same path Settings does:
- `available` -> action `download` ("vX · update"), kick the native download
- `downloaded` -> action `restart` ("vX · restart"), quit and install
Manual download is demoted to the notify-only fallback only
(`available-manual`, where the native updater could not fetch it); the
hover card with drag-to-Applications steps now shows only in that case.
The success acknowledgement from this branch is unchanged.
This is the LAST code change before the 0.4.6 freeze, and it is the thing
the rc rehearsal is meant to exercise: rehearsing without it would test
the old badge behaviour, a build that never ships.
update-state tests updated (available -> download, downloaded -> restart);
all 17 green, badge acknowledgement tests green, typecheck:web clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014dMe8Mm2eas1SwvUiu3XXr
#238 sets `$ip: null` on every event so PostHog cannot fill it from the connection. This test pins the exact property set an update-applied event carries ("nothing from the log itself rode along") and predates that key, so it fails on origin/main the moment #238 lands — not a merge interaction. The shipped behaviour is correct and unchanged; only the assertion was stale. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eUT4Mh1EZ7rkJ2DiDens
Contributor
🚫 This PR is missing its before/after evidenceEvery pull request here has to show its work. Screenshots or a short screen recording, before the change and after it.
How to fix it: edit the description, keep the A bug fix with no visible surface still needs it: show the failing behaviour, then the same steps passing. A terminal recording is fine. Genuinely nothing to show — a CI tweak, a typo, a dependency bump? A maintainer can apply the |
The build reads its version from package.json, not the tag, so this line is what names the artifacts and the update feed's version. Rehearsal build only: a `-` suffixed tag publishes as a GitHub pre-release, which electron-updater offers only to a client already on a pre-release, so no 0.4.5 install sees it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7eUT4Mh1EZ7rkJ2DiDens
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Do not merge yet. This PR exists to run CI against the exact tree that
v0.4.6-rc.1is tagged from. It merges tomainonly after the updater rehearsal passes.What is in it
Updater repair set (5): #322 download page · #324 restart re-entry guard · #325 check timeout · #326 badge acknowledges a check and drives the real auto-update · #327 release checklist
Reviewed contributions (13): #238 #282 #240 #243 #242 #286 #271 #225 #239 #248 #156 #270 #284
All 18 merged with zero conflicts. Local
npm run typecheckclean on both projects.One extra commit, and why
test/update-applied.test.cjspins the exact property set anupdate-appliedevent carries. #238 adds$ip: nullso PostHog cannot fill it from the connection, which makes that assertion stale. Verified it is not a merge interaction: the assertion exists onorigin/main,origin/mainhas no$ip, and the commit introducing it is #238's own — so #238 fails this test against untouched main. Shipped behaviour is correct and unchanged; only the assertion moved.Known local gap
test/update-download-asset.test.cjscannot run locally — Electron's binary is not installed in any local worktree. CI is the first place it actually executes, which is the point of this PR.🤖 Generated with Claude Code
https://claude.ai/code/session_01Q7eUT4Mh1EZ7rkJ2DiDens