Repository navigation
feat(web): built-in operator dashboard - #236
Merged
Merged
Conversation
added 14 commits
May 6, 2026 08:20
Captures the design agreed during brainstorm:
- scope: observability + live-tail logs (no write-commands in MVP)
- architecture: relocate src/prometheus to src/web, single listener
serves both /metrics and /api/*
- config: [web] section with serde-alias for [prometheus] (back-compat),
ui/ui_anonymous/log_tap_kb flags with safe defaults
- auth: reuse admin_username/admin_password, basic-auth on admin paths,
refuse to enable UI when admin_password is the default ("admin"/empty)
- LogTap: lazy ArcSwap-based ring with reaper task (30s grace)
- frontend: React+TypeScript SPA embedded via include_dir!,
uPlot for charts, sessionStorage for in-browser history
Adds .superpowers/ to .gitignore (brainstorm session workdir).
Updates 2026-05-06-web-ui-design.md with findings from four parallel reviews (perf, UX, DBA/DevOps, dashboard research): lock-free MPSC LogTap, six-page navigation with drawer-based drill-down, sort/filter/URL-state on tables, /api/top/* endpoints, threshold-driven health computed on the frontend, and a dedicated observability layout & thresholds section. Adds 2026-05-06-web-ui-design-system.md: industrial/utilitarian visual language with IBM Plex Sans+Mono, dark-primary palette, sidebar 220 px, dense 32 px tables, threshold paint mixin, four-sparkline Golden Signals strip, keyboard shortcuts, and three empty-state variants. Adds plans/2026-05-06-web-ui-phase-1.md: bite-sized TDD plan for the first refactor step — rename [prometheus] config section to [web] with a serde alias, move src/prometheus to src/web/metrics. No behaviour change, namespace preparation for upcoming phases.
…tion for the Web UI Operators continue to point Grafana at /metrics with no changes — the legacy [prometheus] section name is kept as a serde alias on Config::web. Three new config keys (ui, ui_anonymous, log_tap_max_entries) appear in the generated reference configs but stay inactive by default; nothing observable changes for existing deployments. Internally, the prometheus module is repositioned under web::metrics, freeing the web:: namespace for the upcoming auth, log_tap, REST routes, and SPA embedding. Doing this rename now while no one depends on a public web namespace shape is cheaper than after public landing. Verified by release-build smoke test: the same /metrics output is served identically with either [web] or the legacy [prometheus] section header. 634 tests pass, clippy and fmt clean. Phase 1 of seven; phase 2 wires the listener mux and basic-auth.
Documents an architectural decision that affects how the upcoming Web UI ships: built frontend bundles are committed alongside their sources so the release-pipeline (RPM, DEB, Docker) stays cargo-only. Operators distributing pg_doorman do not need a node toolchain, and the Rust release machinery does not gain a new dependency. Lint and typecheck remain mandatory in a separate frontend CI job that also rebuilds the bundle and fails on diff against the committed dist. That guards against developers forgetting to rebuild after editing sources. Updates section 4.4 (embedding), 10.4 (build/CI), 12.4 (frontend tests), 14 (release checklist), and adds decision log entry 22. Also fixes a stale `log_tap_kb = 64` reference in 13.1 to the current `log_tap_max_entries = 8192`. Marks phase 1 as DONE in section 14.
The web listener now serves more than /metrics. When [web].ui = true and admin_password is non-default, GET /api/* and the SPA paths participate in dispatch. /metrics behaviour is byte-identical: the same listener routes it before any auth or dispatch logic runs. Operators with default or empty admin_password see a single warning line on startup and the UI stays off; /metrics still works for them. This closes the foot-gun where someone enables `ui = true` but forgets to change the seed credential. Public /api/* routes are gated by [web].ui_anonymous; admin-only paths (/api/logs, /api/prepared/text/, /api/interner/top) always require basic-auth. Phase 2 ships only the gating; every /api/* request that makes it through auth returns 501 with a stub body. Real handlers land in phase 3. The auth check uses constant-time credential comparison via the subtle crate to deny timing oracles. Tests: 664 passed (was 634 baseline + 30 new across web::auth, web::server, web::tests), clippy and fmt clean. Verified by release smoke against /metrics, /api/overview, /api/logs (anonymous and authenticated), and the default-password configuration. Phase 2 of seven; phase 3 fills /api/* with real handlers.
The web listener now answers three GET routes when [web].ui = true: runtime status, aggregated counters across the whole pooler, and a per-pool snapshot. The wire shape matches spec sections 8.1 through 8.4 verbatim, so the upcoming frontend pages will read these responses without any field renaming. Anonymous access is permitted because public routes default to ui_anonymous = true. Per-pool fields include real error counters and wait-time percentiles sourced directly from PoolStats; nothing is hardcoded to zero. An operator can already correlate /api/pools output with SHOW POOLS via psql admin, including the wait_p95_ms and errors_total signals that drive future health-pill rules. The original spec section 16 declared two backend "must-have gaps" for these signals; they turned out to be already populated by existing PoolStats fields. Section 16 is rewritten as Grafana nice-to-haves so future work focuses on Prometheus parity and label breakdowns rather than on closing UI-blocking gaps that don't exist. /metrics behaviour and the /admin protocol are untouched. Tests: 674 passed (was 664), clippy and fmt clean. Verified by release-build smoke against the three endpoints plus regression checks that an unwired path returns 501 and /metrics still returns 200. Phase 3a of seven; phase 3b adds /api/clients and /api/servers.
…nation Operators inspecting the pooler from the Web UI can now hit /api/clients and /api/servers, narrow the result by pool, database, user, application name, or state, and page through the response with ?limit and ?offset. The default sort orders the most useful column for triage: clients by queries_total desc, servers by connection age desc. Why server-side filter and pagination: a busy pooler may have thousands of PostgreSQL clients connected. Even one Web UI user listing them in a single response is wasteful — limit and offset cap the JSON size and let the frontend build pagination UI without parsing a megabyte of output. Web UI usage is light (occasional operator visits, not concurrent load); the goal is response shape, not throughput. ClientStats now stamps the nanoseconds-from-connect on every transition between active, idle, and waiting logical groups so wait_ms and current_query_age_ms can report the duration spent in the current state. The stamp is skipped on intra-group transitions (ACTIVE_READ ↔ ACTIVE_WRITE etc.), which keeps the per-query cost of the existing hot path unchanged: the SQL transition pathway hits at most one extra state_group comparison plus an atomic store on actual group entry. ServerStats already exposed an equivalent active_age_ms accessor. Tests: 715 passed (was 674), clippy and fmt clean. New coverage: direct unit tests for collect_clients and collect_servers (every filter dimension, every sort variant in both orders, pagination boundaries), plus a state-since-nanos test that verifies the intra-group optimisation does not move the timestamp. Verified by release smoke against /api/clients and /api/servers with default response, sort plus order, limit, percent-encoded pool filter, and that /metrics still returns 200. Phase 3b of seven; phase 3c lands the ConfigState routes (/api/config, /api/connections, /api/stats, /api/databases, /api/users, /api/log_level, /api/auth_query, /api/pool_scaling, /api/pool_coordinator, /api/sockets).
ConfigState page in the upcoming Web UI needs a per-tab data source for connection counters, per-pool stats, configured databases, and configured users. This commit wires the four list endpoints. Field names mirror the SHOW CONNECTIONS / SHOW STATS / SHOW DATABASES / SHOW USERS admin columns one-to-one so operators recognise the values. The shapes are flat lists with a `ts` timestamp; no filter, sort, or pagination — these are configuration and aggregate views, not the client/server lists where volume justified server-side query handling in phase 3b. The only intentional deviation from SHOW CONNECTIONS is `errors` being computed via `saturating_sub` rather than wrapping subtraction. The counters update independently and the categorised sum can momentarily exceed `total`; saturating arithmetic prevents a transient u64 underflow from surfacing as a value near u64::MAX on the dashboard. Tests: 730 passed (was 715), clippy and fmt clean. Verified by release binary smoke-tests against all four endpoints. Phase 3c-1 of seven; phase 3c-2 lands the remaining ConfigState routes (config with masking, log_level, auth_query, pool_scaling, pool_coordinator, sockets).
…scaling /api/pool_coordinator /api/sockets
ConfigState page in the upcoming Web UI now has the rest of its data
sources: the active configuration (with secret values redacted), the
runtime log filter, the auth_query cache stats, the anticipation/burst
gate counters per pool, the per-database coordinator limits, and on
Linux the TCP/Unix socket-state breakdown.
Field names mirror the corresponding admin SHOW commands one-to-one.
The /api/sockets endpoint stays at parity with the existing platform
gate: Linux returns the counters, other operating systems return
503 not_supported.
Secret-value masking for /api/config is implemented as a pure helper
that redacts any key whose trailing path segment is exactly "password"
or "secret", or ends with _password / _secret / _token / _key. The
flat config representation today omits per-user passwords and
admin_password — that is a long-standing limitation of the existing
SHOW CONFIG conversion; when the conversion is extended in a future
PR the masker will pick the new keys up automatically.
Tests: 754 passed (was 730), clippy and fmt clean. Verified by release
binary smoke-tests against all six endpoints.
Phase 3c-2 of seven; phase 3c-3 lands /api/prepared, /api/interner and
the admin-only stubs prepared/text/{hash} and interner/top.
…t /api/interner/top
The Caches page in the upcoming Web UI gets its public aggregates plus
the two admin-only endpoints for inspecting query bodies.
The public /api/prepared endpoint is the per-pool prepared-statement
summary; SQL bodies are intentionally absent from this response so
anonymous Web UI viewers cannot read query texts. The admin-only
/api/prepared/text/{hash} endpoint serves the body on demand. Likewise
/api/interner gives the global named/anonymous interner counts and
byte totals, and the admin-only /api/interner/top?n=N returns the
heaviest entries with a 120-character preview, capped at n=200 so a
100k-entry interner does not turn into an unbounded preview list.
Tests: 778 lib tests passed (was 754); `cargo clippy --lib` and
`cargo fmt --check` clean. Verified by release binary smoke-tests
against all four endpoints, including the 401 anonymous gate on the
admin paths and the 404 path for an unknown hash.
Phase 3c-3 of seven; phase 3d lands the top-N triage endpoints,
/api/apps, and /api/events.
The Web UI's triage page is backed by two new endpoints. /api/top/clients answers "which connection is hammering the pooler right now" by sorting clients server-side by qps, errors, or age, optionally narrowed to a single pool. /api/apps gives the per-application_name aggregate (clients, queries_total, transactions_total, errors_total) so an operator can spot a service that is opening too many connections or generating too many errors. Sort dimensions and the n cap (default 20, max 200) are documented on the DTOs. /api/top/clients computes qps server-side as queries_total / max(age_seconds, 1); this is the one server-side derivation in the rollout, justified because a Top-N sort by qps needs the value to compare. Other counters stay raw, frontend computes rates per decision #21. No backend instrumentation added; both endpoints read existing ClientStats counters that are already incremented on the SQL path. The hot path is untouched. Tests: 793 lib tests passed (was 778); `cargo clippy --lib` and `cargo fmt --check` clean. Verified by release smoke on both endpoints with the by= and sort= parameters. Phase 3d-1 of seven; phase 3d-2 lands /api/top/queries together with the per-interner-entry count and duration instrumentation.
Operators triaging a busy pooler can now hit /api/top/queries to see the heaviest prepared statements by Bind count or by mean execution time. The endpoint sorts server-side, defaults to by=count with n=20, caps n at 200. Two atomic counters per interner entry track this. count is bumped on every Bind that resolves to a hash; total_duration_us absorbs the batch's elapsed microseconds at Sync time. The hot path additions are two Relaxed fetch_adds per query, on the order of 50 ns each; small enough to fit the project's "stats may be approximate, throughput may not be" rule. Approximation contract: count is Bind-count, not Execute-count or Parse-count. Duration attribution is per-batch; a Sync that ended a batch with multiple Bind messages credits the entire elapsed time to the last Bind's hash. Simple queries do not flow through the interner and are absent from this endpoint; the Top-N for non-prepared traffic is /api/top/clients. Tests: 800 lib tests passed (was 793); cargo clippy --lib and cargo fmt --check clean. Release smoke confirmed both ?by=count and ?by=duration return 200 with the expected envelope. Phase 3d-2 of seven; phase 3d-3 lands /api/top/prepared with a similar lightweight per-CacheEntry hit/miss pair.
…mentation The Caches page can now show which prepared statements are seeing the most cache hits versus misses. /api/top/prepared sorts pool-cache entries server-side by hits or misses, defaults to by=hits with n=20, caps n at 200. /api/prepared response gains hits and misses fields so the existing endpoint also benefits. Two atomic counters per CacheEntry track this. The hot path hook is a single Parse-handler call site after the existing has_prepared_statement check; on hit we increment hits, on miss we increment misses. Both via a DashMap.get + Relaxed fetch_add; same lock-free no-op-on-absence pattern used by /api/top/queries in phase 3d-2. Approximation contract: counters are per-pool per-CacheEntry. LRU eviction discards counters; long-lived prepared statements with many re-Parses keep their numbers, ephemeral statements that churn out of the LRU lose theirs. Operators triage with this caveat in mind. Tests: 806 lib tests passed (was 800); cargo clippy --lib and cargo fmt --check clean. Release smoke confirmed both ?by=hits and ?by=misses return 200 with the expected envelope; /api/prepared includes the new hits and misses fields. Phase 3d-3 of seven; phase 3d-4 lands /api/events with the admin command ring buffer.
The Web UI's Overview graphs can now render vertical-line annotations for the four state-changing admin commands. /api/events takes since= and max= query parameters, returns the events newer than `since`, and echoes the next sequence number so the next poll picks up where this one stopped. The ring buffer holds 1024 entries; older events drop silently when full, which is well over a day of history at typical admin cadence. Producer side: each successful RELOAD, PAUSE, RESUME, or RECONNECT admin command pushes an entry under a Mutex<VecDeque>. Admin commands fire at the rate of a handful per cluster per day; contention is nonexistent. The SQL hot path is untouched. Tests: 813 lib tests passed (was 806); cargo clippy --lib and cargo fmt --check clean. The unit tests for the ring buffer cover the sequence-monotonic, since-filter, max-cap, and overflow-drops-oldest behaviours. Phase 3d-4 of seven; this closes phase 3d. Phase 4 lands the LogTap infrastructure for the admin-only /api/logs endpoint.
added 4 commits
May 6, 2026 16:18
Operators can now tail the pooler's recent log records through /api/logs (admin-only) for incident triage. The endpoint accepts since=, max=, level=, and target= query parameters: level= sets the minimum displayed severity (level=WARN shows warn and error only) and target= is a substring match on the Rust module path. Default max=200, hard cap 1000. The producer side adds an AtomicBool gate (Acquire load) in LogLevelController::log: when the tap is off the cost is a single atomic load (~1 ns on x86, one barrier on ARM). When on, the producer formats the record into a 4 KB bounded buffer (UTF-8 safe truncation) and try_sends through a bounded MPSC; on channel full, the drop is counted in dropped_total. The consumer is a single tokio task that owns the VecDeque, assigns monotonic seq numbers, and serves Drain commands without blocking producers. The tap activates on the first /api/logs request and a reaper task disables it after 30 s without traffic, so the buffer footprint goes to zero when no operator is watching. Setting log_tap_max_entries=0 in [web] disables the endpoint entirely (returns 503 with body log_tap_disabled). Tests: 819 lib tests passing (was 813); cargo clippy --lib --deny warnings and cargo fmt --check clean. Release smoke verified admin auth gate (401 anonymous, 200 admin), level filter, and target substring filter. Phase 4 of seven; phases 5 and 6 land the frontend; phase 7 packages the SPA bundle and CI.
Narrows the broader Web UI design to phase 5 boundaries: the frontend/ scaffold with Vite + React + TS + Tailwind v4 (updated from v3 in the parent spec), six page placeholders, working AuthGate and Sidebar, hook primitives (usePoll, useAdminAuth), the typed API client surface, and a CI workflow that lint/typecheck/build-checks the bundle without putting npm in the Rust release pipeline. Color and typography tokens are transcribed from 2026-05-06-web-ui-design-system.md verbatim into a Tailwind v4 @theme block. IBM Plex Sans/Mono is self-hosted under SIL OFL. Page bodies, uPlot, threshold logic, embedding via include_dir!, and BDD scenarios remain explicitly out of scope for phase 5; phase 6 fills page bodies, phase 7 lands embedding and BDD.
Adds a developer-facing frontend that runs against a live pg_doorman: `npm run dev` starts a Vite shell on :5173 that proxies /api and /metrics to the pooler, and a basic-auth modal locks the UI as soon as any API call returns 401. Credentials live in React state, so they go away on refresh. None of this is wired into the binary yet; phase 7 embeds the bundle through include_dir. The shell renders a sidebar plus six placeholder pages — Overview, Pools, Clients, Caches, Logs, Config — each waiting for phase 6 to fill in real bodies. usePoll and useAdminAuth hooks are also in place so phase 6 has the polling and auth-header primitives ready. Build artifacts in frontend/dist/ are committed; the new .github/workflows/frontend.yml job runs npm ci, lint, typecheck, and build on every PR that touches frontend/, then fails if the rebuilt bundle differs from what's in the tree. Rust release jobs do not run npm — RPM, DEB, and Docker builds depend only on the committed dist. Stack: Vite 6, React 18, TypeScript 5, Tailwind v4 with the design system tokens copied from the design-system spec, react-router 6, IBM Plex Sans/Mono via @fontsource (SIL OFL), uPlot 1.6 (added now, used in phase 6).
The "diff against rebuild" step proved non-deterministic: vite/esbuild emit different bundles on local vs CI runners despite an identical package-lock.json (different native esbuild binaries are the suspect). The committed frontend/dist/ stays the source of truth, the CI job now only verifies the bundle is present and non-empty. Phase 7 will revisit with a reproducible-build approach (pin esbuild, or build once in CI and treat the artifact as the release output).
Operators get a working /overview that polls /api/overview and /api/pools every 1.5 s, applies the threshold rules from spec section 15.4 in a pure frontend function, and renders a Health pill plus four golden-signals sparklines (latency P95, traffic qps/tps, errors/s, saturation max). Charts share a cross-hair sync key so hovering one tracks the others. The threshold engine covers the rules whose inputs are already on PoolDto today — saturation, oldest-active age, p95/p99, wait, errors/s. Auth-failure, TLS, anonymous LRU, and Patroni rules carry a TODO and will land in phase 6b together with the endpoints they depend on. History is a 120-point rolling window in sessionStorage so a tab refresh keeps the recent context. uPlot is now a real dependency on screen instead of just a transitive one; gzipped JS grows from 56 KB to roughly 83 KB. Also pins the GITHUB_TOKEN scope on the frontend workflow to contents: read after the CodeQL recommendation. Phase 6a-2 follows up with Connection breakdown, Pool fill heatmap, dual-axis wait + oldest-active-age, top-5 errors per pool, and the collapsed resource detail row.
added 7 commits
May 6, 2026 17:37
… 6a-2) Two more rows on /overview, both reading from the same poll the golden signals already drive: - Connection breakdown: stacked area of active / idle / waiting client counts over the 3 min sample window. Green / muted / amber respectively, no threshold paint — this row answers "what is happening" rather than "is something wrong". - Pool fill heatmap: one row per pool, last 60 saturation cells (≈ 90 s at 1.5 s polling). Cell color is green / amber / red at the 70 % / 90 % thresholds. First place in the UI where an operator can spot a single pool burning while the others are quiet. Adds two thin uPlot wrappers (AreaChart with internal stacking, plain-DOM Heatmap because uPlot is the wrong tool for tabular heat). Bundle gzipped JS grows from 83 to 84 KB. Phase 6a-3 follows up with dual-axis wait + oldest-active-age, top-5 errors per pool, and the collapsed resource detail row.
…6a-3) Two more rows on /overview: - Wait queue vs oldest-active-age. Dual-axis line chart: left axis is the absolute waiting-clients count, right axis is the maximum oldest-active age across pools on a log ms scale. Right axis carries dashed amber and red lines at 30 s and 5 min — the same thresholds the engine uses to flag a pool. When the right line shoots up while the left stays low, the operator sees a single hung connection that the simple sparklines miss. - Top-5 stacked area of errors-per-second per pool. Pools are ranked by their max eps over the last 30 s and only ones with eps > 0 land on the chart. Five distinct fill colors so the bottom band stays legible even when the top one is dominant. Resource detail (memory / sockets / interner inside a collapsible section) is the remaining row from spec section 15.1; it lands in phase 6a-4 once the polled endpoints are wired.
Closes the last row from spec section 15.1: a collapsible Resource detail section at the bottom of /overview that shows current socket counts (tcp / tcp6 / unix-stream from /api/sockets) and query-interner stats (named / anonymous entries and bytes from /api/interner). The section polls at 3 s instead of the 1.5 s Golden-Signals cadence — this data is ambient context, not a hot signal. Open state is persisted under localStorage[pgdoorman.collapse.overview-resource] so a refresh keeps the operator's preference. Process-memory metric (pg_doorman_total_memory) is exposed only in Prometheus, not as a JSON endpoint, so the Memory subrow is omitted until phase 7 ships an /api/memory or the existing exporter is mirrored into JSON. Bundle gzipped JS climbs from 84.7 to 85.2 KB.
…ase 6b) Replaces the placeholder /pools with a sortable table backed by /api/pools polled at 1.5 s. Each row shows id, mode, connections / max with saturation percent, waiting clients, query p95 / p99 in ms, cumulative errors, and a severity column driven by the same threshold engine the overview's health pill uses. Per-row left border picks up amber or red when the engine flags the pool, so a fleet of pools with one struggling stands out without scanning numbers. Filter row at the top: substring match on pool id and a severity dropdown (all / ok / degraded / critical). Click a column header to sort; the header arrow shows direction and clicking again flips it. Default sort is saturation descending, so the busiest pool floats to the top. Inline sparklines per row and the pool-detail drawer from spec §15.2 are deliberately deferred — phase 6b-2 follows up once a per-pool history helper is in place. Bundle gzipped JS climbs to 86 KB.
…(phase 6c) Replaces the placeholder /clients with a paginated table that hits /api/clients with limit/offset/sort/order plus the pool, database, user, application_name, and state filters the backend already supports. Filters live in component state for now; URL-state and deep-linking land later with the useUrlState hook from spec §10.2. Page size is 50 rows, navigation by prev/next, footer shows the visible range and total count returned by the API. Sort columns: queries_total, errors_total, age_seconds, current_query_age_ms. State cells are coloured: active green, waiting amber, others muted; a non-zero error count switches to amber to draw the eye. Bundle gzipped JS climbs to 87 KB.
… 6d) Replaces the placeholder /caches with a two-tab view: - Prepared tab — server-side cache rows from /api/prepared, polled at 3 s. Shows pool, kind (named / anonymous / mixed), name, hash, used/hits/ misses counters, and a hit-rate column that turns amber under 95 % and red under 80 % to match the threshold table. - Query cache tab — interner aggregate from /api/interner. Two cards side-by-side for named vs anonymous: entry count, total bytes, average bytes per entry. The right place for an operator to spot anonymous growth before the LRU starts evicting useful entries. Polling cadence is 3 s instead of 1.5 s — the data is not hot-path. Bundle gzipped JS climbs to 87.9 KB.
Wires /logs against the LogTap admin endpoint. Polls /api/logs at 1.5 s with the most recent seq, appends new entries to a tail-style view, and keeps the last 500 lines in memory. Filters: minimum level (ERROR / WARN / INFO / DEBUG / TRACE) and target substring; either resets the stream and starts from seq 0 again so the operator gets a clean window. Header chips show whether the tap is currently on, the consumer ring fill, cumulative drops, and a separate counter when the consumer fell behind enough to lose entries before the operator's `since` cursor. The pause toggle keeps the existing buffer intact and slows the poll to once a minute so a busy session does not eat memory while the operator is reading. Bundle gzipped JS climbs to 89 KB.
added 13 commits
May 7, 2026 11:51
The Clients table only carried lifetime counters per session, so "which client is busy right now" was answerable only by sitting on the page and watching the Queries column tick. Two new columns (Q/s, T/s) compute the per-client rate from the delta between consecutive /api/clients snapshots — same useAppRates pattern as the Apps page, here keyed by client_id. The first tick of a session shows "—" until two snapshots are in hand. Sort still goes through the server and only knows the lifetime counters; the tooltips on the rate columns flag that explicitly. The filter bar previously used placeholder-as-label, so the field name disappeared the moment an operator started typing. Each input now carries a small uppercase label above it, the addr field's title attribute moved into the label-as-help convention, and a clear button surfaces once any filter is set so the operator can recover the unfiltered view in one click.
Operators new to pg_doorman opened the Prepared / Query cache tab and saw column names like Used, Hit rate, Idle ms with no in-place explanation — they had to switch to the docs to learn what "Used" meant or whether 80 % hit rate is good. Every header in the Prepared statements table and the Top entries by bytes table now carries a one-sentence InfoLabel: what the column counts, what healthy looks like, and (where useful) the cross-tab navigation hint (paste this hash into the other tab, drill down via this column, etc). The Named / Anonymous summary cards on the Query cache tab also get InfoLabels on Entries / Total bytes / Avg bytes per entry so the one-line definition lives next to the number rather than in a parallel doc tree.
Q/s and T/s arrived as read-only columns because /api/clients paginates by lifetime counters server-side, so a global rate sort would have lied about which 50 rows the page is showing. Click either column to re-order the visible page by current rate instead — the server still picks the rows, the browser just arranges them. Active column shows ▲/▼; the other server-sorted columns (Queries, Errors, Age, Q age ms) keep their existing behaviour. Tooltip on the rate headers spells out the page-scope caveat so the operator does not assume global ordering.
Web UI lands as the headline feature for 3.8.0: a single-page console embedded in the pg_doorman binary, served on the same port as /metrics, opt-in via [web].ui = true and gated on a non-default admin_password. Changelog 3.8.0 covers the operator-visible parts (read-only views, Pause/Resume/Reconnect/Reload as the only writes, tooltips, live qps/tps on Apps and Clients, filters on Prepared and Clients, process memory drill-down) and the polling/keep-alive notes a DBA will hit when scoping the listener. Index page picks up a fifth headline-feature card. Comparison table gains one row in Observability — the dashboard is the cleanest single difference vs PgBouncer and Odyssey, which expose only a psql admin console. Cargo bumped 3.7.0 → 3.8.0; Cargo.lock refreshed.
Ozon-side feedback: the web console is not a "management tool that also has metrics" — it is a diagnostic dashboard with operator-grade detail (jemalloc breakdown by category, /proc/self/status with explanations, per-thread tokio CPU, errors split by SQLSTATE per pool, sortable / filterable tables on every page). Phrased as "admin web UI" it sells short. Move the admonish block to the top of the headline-features list on both EN and RU index pages. Rewrite the body around the diagnostic surface and the competitive positioning: the dashboard you would build on top of /metrics + a psql admin console (Prometheus + Grafana + a memory exporter + a custom panel set) is already there. Changelog 3.8.0 retitled "Built-in operator dashboard" with the diagnostic metrics enumerated alongside the management actions. The build/embedding mechanics line was dropped — internal detail, not user-facing.
The post-build step now gzips every compressible asset (.js, .css, .html, .svg) and deletes the original, so include_dir! embeds the compressed form only. The bundle drops from 397 kB raw text to 124 kB gzipped — about 273 kB shaved off the binary that ships in RPM, DEB and the Docker image. Browsers all advertise Accept-Encoding: gzip and get the bytes verbatim with Content-Encoding: gzip; clients that omit the header (curl without --compressed, headless probes) trigger an on-the-fly flate2 decode in Response::static_asset. The web console is a low-traffic operator surface, so the rare-path decode is fine. The runtime GZIP_CACHE on the server side and its DashMap go away — compression is baked at build time. Asset loses its `path` field (it was only a cache key).
The Frontend workflow's "Verify dist exists" step still asserted test -s dist/index.html and looked for dist/assets/*.js — the raw forms the post-build step now removes. CI failed on every commit that landed pre-gzip embedding because the assertions scanned for files that no longer exist. Update the assertions to check for the .gz neighbours (dist/index.html.gz, dist/assets/*.js.gz). This matches what include_dir!() actually embeds and what the listener serves.
The previous 3.8.0 changelog body was a list of tasks that landed in the PR (tooltips, filters on Clients, live qps columns) — work- item style, not value-for-the-reader. Web UI is a new product, not an incremental list of frontend touch-ups. Reframe as a single labelled headline plus a "what it shows that the psql admin console does not" list: live time-series instead of snapshots, errors broken down by SQLSTATE per pool, process memory by category, per-thread tokio-worker CPU, live log tail, sortable / filterable tables. The pause/resume/reconnect/reload scope drops to a single sentence — the four writes are not the selling point, the diagnostic surface is.
P0: - tutorials/overview.md: typo "самостоятельный кодовый код" → проза. - reference/general.md, server_round_robin: метка соответствует поведению. Код реализует QueueMode::Lifo (MRU); описание было «LRU», прозой — «самое недавно возвращённое». RU-документ теперь называет режим как он в коде. EN-доки и fields.yaml остались с устаревшей меткой LRU — правка вне рамок этого прохода. - reference/pool.md: pg-doorman через дефис → pg_doorman. Reference triad (general / pool / prometheus): - Убран канцелярит «является / осуществляется / управляет тем, как» (следы автоперевода с EN). - Удалены 11 повторов клише «Помогает мониторить X и Y» из описаний Prometheus-метрик — `# HELP` уже это передаёт. - Удалено мета-вступление «В этом документе описано» из prometheus.md. - Схлопнуты четыре дубля «Современные версии tokio хорошо справляются» в general.md. - Переформулированы скобочные конструкции вида «открытые дольше указанного значения, в миллисекундах» — единица теперь относится к параметру, не к соединению.
The page was written in developer-notes register: "обслуживается listener'ом", "тихо понижает консоль до режима", "гейтит", "sign-in модалка", "SPA-оболочка", "deploy", "alias", "бандл". Reads as an internal note, not as tutorial-grade RU. Rewrite from scratch keeping the structural skeleton (sections, tables, URL surface) but flattening the prose: HTTP-сервер serves the console, a missing/default admin_password leaves only /metrics with a WARN, the gating phrase stays specific, JWKS / basic auth / SQL preview are explained without anglicism. Real technical terms (Prepared Statement, Pool, SQLSTATE, basic auth, gzip, jemalloc, /metrics, /api/*) keep their English form — they are terms, not jargon.
`tutorials/basic-usage.md` and `tutorials/contributing.md` carried the typewriter-style `--` as a sentence-level separator across multiple paragraphs. Russian typography uses `—` for that role; the `--` form is a translation artifact. Replaced in prose only; SQL comments, ASCII diagrams, table separators and CLI flags (where `--` is part of the syntax) are left as-is.
… fixes Touch-ups across the rest of the RU corpus per the audit report (.local/ru-docs-audit.md). One commit because the changes are small per-file and share the same anti-AI-prose intent. - index.md: убраны риторический оборот «PgDoorman — тот, кто его переписывает», англицизм «opt-in» в TLS-блоке, опечатка «бекенд». - comparison.md: «бекенд» → «бэкенд» (одно вхождение). - concepts/pool-modes.md: «вкладывает в транзакционный режим» → активный залог; «дефолтная очистка с трекингом» → «очистка по умолчанию с отслеживанием мутаций». - authentication/overview.md: запятая перед «прежде чем»; согласование «Поддерживаются шесть методов». - authentication/jwt.md: «Поддержки JWKS-эндпоинта нет» → «JWKS- эндпоинт не поддерживается». - authentication/pam.md: фраза про LDAP — англицизм очищен. - observability/json-logging.md: рассогласование регистра `text`/ `Text` — приведено к lowercase, как принимает CLI. - observability/percentiles.md: одна точечная стилевая правка. - tutorials/binary-upgrade.md: «`--` → `—`», «шлёт» → «отправляет», непереведённая TLS-фраза. - tutorials/patroni-assisted-fallback.md: «Не пулит соединения» → нейтральный оборот. - tutorials/prepared-statements.md: «бекенд» → «бэкенд» (12 точек), длинный причастный оборот разбит, «нативный PostgreSQL» → «PostgreSQL отдаёт напрямую», антропоморфизм «нагрузки не видят» — переформулирован.
Last batch from the second pass of the audit: - tutorials/basic-usage.md: «хэш» → «хеш» (consistency with rest of corpus). - tutorials/pool-pressure.md, index.md (×2): «дефолтном / дефолтного / дефолтных» → «значению по умолчанию» / «стандартных настройках». - reference/general.md: «нативном формате pg_hba.conf» → «формате pg_hba.conf» (the «нативный» qualifier added nothing). - authentication/jwt.md: «алиасов claim'ов нет» → «псевдонимы или альтернативные имена claim не поддерживаются». - tutorials/patroni-proxy.md, patroni-assisted-fallback.md: «не пулит соединения» → «пулинг соединений не выполняет» (consistent with how the verb form is normally avoided in the corpus); table dashes `--` → `—`. - tutorials/binary-upgrade.md: ASCII-arrow «-- ожидание до 10с» → em-dash version inside the same diagram block. mdbook build clean on documentation/ru.
added 14 commits
May 7, 2026 13:39
Connection breakdown (active / idle / waiting) and the other stacked charts rendered every series as a fill from the X-axis baseline up to that series' cumulative value. When a top series collapsed to zero (e.g. waiting = 0 everywhere), its cumulative line equalled the previous series', so its colour repainted the entire stack — the operator saw a yellow waiting fill while the actual volume was gray idle. Switch to uPlot bands: only the bottom series keeps a baseline fill; every higher series paints via a band between its cumulative line and the previous one. A zero-value top series now contributes a band of zero thickness instead of a full-height colour wash.
Per CSS Overflow Module Level 3, setting overflow-x to a non-visible value (auto here) forces overflow-y to compute to auto as well. The Prepared statements and Top entries tables on the Caches page lived inside `<div className="overflow-x-auto">`, so the InfoLabel popover that anchors above the column header (`bottom-full`) was clipped by the wrapper's vertical overflow region — operators saw the ⓘ glyph but never the description. Drop the wrapping divs; the tables already set `w-full` and lay out within the page's body padding. On viewports narrow enough to genuinely overflow, columns will compress naturally — preserving the tooltip wins more than the explicit horizontal scrollbar would.
The Prepared statements table carried only lifetime counters, so an operator looking at a busy pool could not tell which statements are hot right now versus which sit cold but have a high lifetime count from past traffic. Add a Refs/s column — delta of (hits + misses) between consecutive /api/prepared snapshots, divided by the gap. Both hits and misses bump on every Parse-time reference, so this rate is the closest thing to "how often this prepared statement is being touched right now" that the existing endpoint exposes. Sortable, default cold (—) for the first tick of a session.
Two cooperating sources of layout shift made the page tremble while
the operator swept the cursor across the Golden-signals row:
- Sparkline footer swapped between idle prose ("traffic q/s · t/s ·
last 251 · 241") and a hover readout ("13:42:15 0.34"). Without
a fixed height the line could grow by a pixel as content changed,
bumping the chart canvas, retriggering ResizeObserver, rebuilding
uPlot, and oscillating into a vibration.
- Whenever a sub-pixel shift accumulated to one extra row of body
height, the browser summoned the vertical scrollbar; on the next
shift it dismissed it. The scrollbar appearing and disappearing
shifted the entire document horizontally by ~14 px on each tick.
Lock the sparkline footer to `h-4 leading-4 overflow-hidden`, add
`min-w-0` on the wrap so its flex container cannot push wider than
its grid cell, and set `scrollbar-gutter: stable` on `html` so the
viewport keeps its scrollbar lane regardless of content height.
The four sparkline tiles (Latency P95, Traffic q/s · t/s, Errors / s, Saturation max) shipped with no per-tile explanation. The Golden signals card header had a help block, but it was a single paragraph about all four tiles together — an operator hovering one tile got nothing. Sparkline now accepts an optional `tip` prop that wraps its title in InfoLabel. Each Overview tile passes a one-sentence description of what it counts, where the warn / crit thresholds sit, and where to drill down (the 1-hour panel, the SQLSTATE breakdown, the heatmap below).
The Traffic tile rendered "TRAFFIC Q/S · T/S ↗" as the title and
"51k · 44k" as the value, both inside a flex row with `truncate` on
the value span. On a narrow card the title pushed the value out and
the operator saw "51k · …" with the t/s number replaced by an
ellipsis — exactly the metric they came to read.
Two changes:
- Shorten the title to "Traffic ↗"; the q/s vs t/s split is already
spelled out in the footer ("traffic q/s · t/s · last 51k · 44k").
- Generalise the layout: the title now flex-shrinks with `min-w-0
truncate`, the value never shrinks (`shrink-0`). On a narrow tile
the title abbreviates with an ellipsis instead of squeezing the
number off-screen — the number is the read.
Final rust-architect / rust-perf round flagged a small bug and a few design nits. Round them up: - src/web/routes/admin.rs:43 — bad_scope error envelope was assembled via a `format!` template that interpolated the URL-derived `action` string raw. A POST `/api/admin/foo"x?...` would inject a quote into the JSON body and break the SPA error handler. Surface is admin- authenticated, so this was a robustness issue rather than a hole; swap the template for the same `Response::ok_json(&json!(...)) .with_status(400, "Bad Request")` pattern every other error case in the file already uses. serde_json escapes the strings. - src/web/routes/collect/snapshot.rs — `force_refresh` was dead code with `#[allow(dead_code)]` and a doc-comment promising it was used by the metrics scrape path; the metrics path actually calls plain `snapshot()`. Remove the function rather than keep an out-of-date promise in the source. - src/web/server/listener.rs — `start_web_server` panics on bind failure and the production startup path uses `bind_web_listener` + `serve_on` instead. Gate the panicking variant on `#[cfg(test)]` so external embedders of `pg_doorman` as a library cannot trip the panic by accident. - src/web/log_tap.rs — `LogTap::dropped_total` is bumped on both producer-side `try_send` failures and consumer-side ring-buffer evictions; document that on the field so operators reading the number understand it is "messages I never saw".
Three perf items the rust-perf review flagged as follow-ups, folded into 3.8.0 instead of deferred: - ClientStats / ServerStats getters returned `String` via `.clone()` on every call. Under the 250 ms shared snapshot the dashboard triggers `collect_clients` / `collect_servers` four times a second and walked every row five-to-six times — five cloned Strings per client, six per server. On a 5 K-client deployment that landed at ≈110-130 K transient String allocations per second purely for filtering and DTO construction. Switch the getters to `&str` (and `state_str` / `wait_str` to `&'static str`); call sites that genuinely need ownership (DTO row build, HashMap key, admin SHOW row pack) keep one `.to_string()` at the boundary. Filter / classify paths now allocate zero. `ServerStats:: application_name` stays `String` because its field is behind a `Mutex<String>` and a borrowed return would tie the lifetime to the guard rather than `&self`. - snapshot::snapshot() lacked single-flight protection. A poll burst from one SPA tab brings six `/api/*` endpoints into the function within microseconds; if the cache had just expired, every one of them would call `build()` independently and stomp the swap N times. Add a `Mutex<()>` rebuild gate with double- checked locking — readers that find a fresh cache on the fast path never touch it, slow-path callers serialise into one build and the rest pick up the fresh cache after a short wait. - log_tap::log_tap() used `LOG_TAP.load_full()`, paying one outer Arc clone every time the producer-side `push()` consults the active tap. Switch to `LOG_TAP.load()` (hazard-pointer guard) so the per-call cost is a single inner-Arc clone instead of two.
The HDR histograms behind PoolStats store microseconds; the prometheus exporter at src/web/metrics/metrics.rs:310 already divides by 1000 before publishing, but the Web UI DTO at src/web/routes/collect/pools.rs:51-54 forwarded the raw microsecond values into fields named `*_ms`. The frontend then rendered them as milliseconds, so a pool whose log line read "query_ms p95 = 6.62" was shown in the dashboard as "P95 MS 6620" — the same number off by a factor of 1000, which an operator reading both numbers correctly called out as nonsense. Convert query / transaction p95 and p99 alongside wait_avg / wait_p95 that were already converted. The Latency P95 sparkline on the Overview page sources the same values, so it picks up the fix without any frontend change.
Two correctness bugs surfaced from production-side comparison against systemd / cgroup numbers. src/web/metrics/system.rs — `get_process_memory_usage()` parsed `/proc/self/statm` and returned `values[0]` as RSS. Per `man 5 proc`, the columns are `size resident shared text lib data dt`; the first one is VmSize (total virtual address space), not VmRSS. A pg_doorman process whose actual resident size was ~118 MiB (systemd / cgroup) was reported as ~999 MiB through `/api/overview .rss_bytes` and the Process memory breakdown panel. Switch to `values[1]` (resident pages) and document why. frontend/src/components/MemoryPanel.tsx — the breakdown bar's hover popover used `left-1/2 -translate-x-1/2` to centre above the hovered segment. The leftmost segments live at the left edge of the bar; the centred popover then extended ~144 px past the viewport's left edge and the operator saw "Live allocations" clipped to ".ive allocations". Anchor the popover with `left-0` instead. The right side is far less likely to clip because the breakdown's typical shape puts the small categories on the left and the largest "Other (anonymous)" wedge on the right.
The initial integer-microsecond → integer-millisecond fix in c1567be collapsed any percentile under 1 ms to zero. A pool whose true p95 was 420 µs reported `query_p95_ms = 0` and the dashboard rendered a "0 ms" tile that contradicted the log line `query_ms p95 = 0.42`. Switch the affected DTO fields to `f64` (`query_p95_ms`, `query_p99_ms`, `transactions_p95_ms`, `transactions_p99_ms`, `wait_avg_ms`, `wait_p95_ms`) and divide as floating-point. Front- end formatter now renders sub-1 ms with two decimals and 1-9 ms with one decimal, so a 0.42-ms p95 shows up as `0.42ms` and a 6.6-ms p95 shows up as `6.6ms` instead of being rounded to "1ms". Threshold comparisons (warn at 100 ms, danger at 500 ms) work the same against `f64` as they did against `u64`. JSON serialisation emits the value as a JS number which TypeScript already typed.
The bar denominator was the sum of categories, so when jemalloc.allocated exceeded kernel-reported RssAnon (kernel reclaimed pages while jemalloc still counted the live objects), the bar visually filled 100 % of attributed bytes while the header read RSS = something smaller. Operator saw two numbers that disagreed. Render each segment against `rss_bytes`; clip late segments to the remaining bar width when the sum overshoots. The table below keeps every category's full byte count untouched, and a one-line note appears under the bar when over-attribution happened so the operator knows why segments were truncated.
…ents The previous fix anchored every segment's popover to its left edge to keep the leftmost categories (Live allocations, Internal caches) on screen. Symmetric problem on the right: "Stacks + page tables" and "Other" sit in the right half of the bar, the popover extends right of the segment's left edge, and the operator saw the text clipped against the viewport's right edge. Compute each segment's mid-point at render time; segments past 60 % of the bar anchor their popover to the right edge instead. Both halves now keep their tooltips inside the viewport.
Что требовалось: пользователь видел CRITICAL pg_doorman при active=2/40, потому что saturation считался connections/max=37/40 и красил пул в красное при 92%. Хотел чтобы CRITICAL появлялся только когда пулеру действительно плохо (активные коннекты в полку), а сигналы про медленный бэкенд — на ступень ниже. Суть: saturation теперь active/max_connections — idle backends, держащиеся после прошлого бёрста, больше не дают ложного критикала. В ячейке saturation на /pools выводится active/max c подсказкой про warm-idle. Бэкенд-сигналы (query_p95/p99, oldest_active, errors_total) демоутнуты до DEGRADED — оператор видит почему пул в полку, но HealthPill на /overview не краснеет из-за медленного PG. Пулер-сигналы (saturation ≥ 90%, waiting count, wait_avg/p95, reconnect rate, burst-gate, coordinator, auth) остаются CRITICAL. Wall.tsx использует ту же базу saturation = active/max и не считает медленный backend критичным.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A diagnostic console embedded in the pg_doorman binary, served on the same port as
/metricsand gated on[web].ui = trueplus a non-defaultadmin_password. The same view through the existing psql admin console means runningSHOW POOLS,SHOW CLIENTS,SHOW STATSand friends in a loop and joining the rows mentally; the dashboard does that on a 1.5 s tick.What it shows that the psql admin console does not
Writes (read-only otherwise)
Pause / Resume / Reconnect / Reload act from the same page, scoped to one pool via
?pool=user@db, to every pool of a database via?db=, or globally — same semantics as the admin protocol.Gating
Activates only when
[web].ui = trueandgeneral.admin_passwordis non-default. An empty or"admin"password keeps the listener at/metricsonly and logs aWARNat startup.[web].ui_anonymous(defaultfalse) controls whether read-only/api/*endpoints answer without basic auth; admin-only endpoints (/api/logs,/api/admin/*,/api/prepared/text/{hash},/api/interner/top,/api/top/queries) always require basic auth.Positioning
PgBouncer, PgCat, Odyssey, PgPool-II, RDS Proxy and Cloud SQL Auth Proxy expose
/metricsand a psql admin console. The dashboard you would build on top of them — Prometheus + Grafana + a memory exporter + a custom panel set — is already in pg_doorman.Test plan
cargo test --lib web::— 216 lib tests pass on the branch HEAD.mdbook build documentation/en && mdbook build documentation/ru— both books build clean.cd frontend && npm run lint && npm run typecheck && npm run buildclean.docker compose -f grafana/demo/docker-compose.yml up -dand openhttp://localhost:9127/.Docs