Skip to content

Lifecycle events in Web UI; identity-based restart detection - #257

Merged
vadv merged 3 commits into
masterfrom
feat/lifecycle-events
May 19, 2026
Merged

vadv merged 3 commits into
masterfrom
feat/lifecycle-events

Conversation

@vadv

@vadv vadv commented May 19, 2026

Copy link
Copy Markdown
Contributor

Summary

  • The Web UI used to toast "pg_doorman restarted — rate baseline reset" on every routine RELOAD: counter totals are summed across the live pool set, RELOAD and dynamic-pool GC drop pools from the map, totals fall, the rollback heuristic triggered. Identity-based detection replaces it — OverviewDto.started_at_ms next to pid, the new useProcessIdentity hook toasts once per real pid/start-time change, counter rollback just skips a rate-tick.
  • New lifecycle event targets PROCESS_START and CONFIG_VALIDATION_ERROR in src/admin/events.rs, with push sites in run_server, the SIGHUP reload error path, and the admin RELOAD / /api/admin/reload error paths. The new useLifecycleEvents hook polls /api/events incrementally and surfaces CONFIG_VALIDATION_ERROR as a red toast — the deploy step that quietly fails is the one the operator needs to see.
  • New LifecycleBanner surfaces shutdown_in_progress and migration_in_progress as a persistent strip (banners don't vanish like toasts).

Why

Operators were getting "restarted" toasts on routine reloads and dynamic-pool churn. The toast was wrong: the process was still up (sidebar pid and uptime confirmed it), only the counter sum legitimately fell because pools left the live set. Counter-based restart detection is a category error — pid + start-time is the source of truth, the same way pgbouncer_exporter and hatop do it.

What is NOT in this PR (follow-up)

  • DYNAMIC_POOL_DROP / POOL_MIGRATION / BINARY_UPGRADE_BEGIN/END events
  • Operator events drawer on Overview
  • Per-target colour mapping in Sparkline annotations

Test plan

  • cargo fmt --check clean
  • cargo clippy --lib --tests --bins -- --deny warnings clean
  • cargo test --lib: 1135 pass, 1 pre-existing flaky (web::metrics::tests::test_prometheus_server_basic — fails 3/3 on clean master locally, not our regression)
  • npm run typecheck + npm run lint + npm run build (frontend dist regenerated and committed)
  • 2 BDD scenarios added under existing @web-ui matrix tag: /api/events carries PROCESS_START, /api/overview carries started_at_ms
  • Manual: SIGHUP with a bad config → red toast in UI, cfg error chip persists on next page navigation; RELOAD via psql → no restart toast; restart pg_doorman → restart toast fires once

🤖 Generated with Claude Code

dmitrivasilyev added 3 commits May 19, 2026 13:31
…ristic

Sidebar.tsx toasted "pg_doorman restarted" every time overview totals
fell, but the totals are summed across the live pool set. RELOAD that
removes dynamic pools and the dynamic-pool GC both drop pools from
the map, so the sum falls without the process going anywhere. The
operator at the UI saw a restart toast on every routine reload.

Identity-based detection replaces it. OverviewDto carries
started_at_ms next to pid; the new useProcessIdentity hook toasts
once when either field changes between polls. Counter rollback now
just skips a rate-tick — no toast, no misleading rate spike on the
next sample either.

The events ring (src/admin/events.rs) grows two new targets,
PROCESS_START and CONFIG_VALIDATION_ERROR, and acquires three new
push sites: at the end of run_server setup, in the SIGHUP reload
error path, and in the admin RELOAD / /api/admin/reload error paths.
A new useLifecycleEvents hook polls /api/events incrementally and
toasts CONFIG_VALIDATION_ERROR as a red error toast — the deploy
step that quietly fails is the one the operator needs to see.
A persistent LifecycleBanner surfaces shutdown_in_progress and
migration_in_progress because banners do not vanish like toasts.

Two BDD scenarios pin the surface contract: /api/events emits a
PROCESS_START entry on boot, and /api/overview carries started_at_ms
alongside pid.

What is NOT in this PR (follow-up): DYNAMIC_POOL_DROP /
POOL_MIGRATION / BINARY_UPGRADE_BEGIN/END events, the operator
events drawer on Overview, and per-target colour mapping in the
Sparkline annotations.
The first commit of this PR landed identity-based restart detection
and basic lifecycle events but left four operator-visible holes that
the DevOps review flagged:

- CONFIG_VALIDATION_ERROR was a 10-second toast — easy to miss after
  the operator alt-tabs to a terminal. Now it is a persistent red
  banner that stays until a successful RELOAD clears it. The backend
  rate-limits the push to 1/s/target so a SIGHUP loop with a bad
  config no longer fills the 1024-entry ring with duplicates.
- LifecycleBanner went stale silently — TanStack Query kept the last
  successful /api/overview for 5 minutes after the pooler died, so
  "draining" stayed on screen long after pg_doorman was gone. Now the
  banner switches to an "unreachable, last contact 23s ago" state
  when the dataUpdatedAt timestamp exceeds 15s or the query errors.
- isRealRestart compared only pid and started_at_ms. On a host with
  30+ days uptime a PID recycle plus an identical lazy-read start
  time could in principle slip through. Adding uptime_seconds < prev
  closes that loophole — a restarted process always has a lower
  uptime than the cached one.
- SIGHUP that re-parsed identically used to leave no trace. Now it
  emits a RELOAD entry with "config unchanged" so the audit timeline
  has one event per signal, no more.

Tailwind tokens text-warning-strong / text-accent-strong did not
exist in tailwind.css and were silently dropped; the banner is now
on the base hues (text-warning, text-accent, text-danger).

A separate cleanup hoists STARTED_AT_MS from process.rs to
app/server.rs so both overview and process collectors share one
LazyLock instead of duplicating the conversion.

/api/events and /api/overview now ship Cache-Control: no-store so
intermediate proxies cannot collapse two consecutive polls into the
same response.

A new unit test covers the rate-limit collapse behaviour
(rate_limited_collapses_burst_per_target).
Cargo.toml + Cargo.lock move to 3.10.0 alongside the existing 3.10.0
changelog header. Two new subsections cover what landed since the
3.10.0 stub was opened:

- Eviction visibility for prepared-statement caches (PR #256).
- Web UI lifecycle events: identity-based restart detection,
  PROCESS_START / CONFIG_VALIDATION_ERROR ring entries, persistent
  banner for shutdown / migration / validation error / unreachable
  (this PR).
@vadv
vadv merged commit 473c2ad into master May 19, 2026
52 checks passed
@vadv
vadv deleted the feat/lifecycle-events branch May 19, 2026 12:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant