Skip to content

feat(web): built-in operator dashboard - #236

Merged
vadv merged 128 commits into
masterfrom
feat/web-ui
May 7, 2026
Merged

vadv merged 128 commits into
masterfrom
feat/web-ui

Conversation

@vadv

@vadv vadv commented May 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A diagnostic console embedded in the pg_doorman binary, served on the same port as /metrics and gated on [web].ui = true plus a non-default admin_password. The same view through the existing psql admin console means running SHOW POOLS, SHOW CLIENTS, SHOW STATS and friends in a loop and joining the rows mentally; the dashboard does that on a 1.5 s tick.

What it shows that the psql admin console does not

  • Live time-series, not snapshots. Latency p95/p99, qps, errors/s and connection saturation as sparklines.
  • Errors broken down by SQLSTATE per pool. Plus top-N stuck queries, top-N noisy clients, top-N hottest prepared statements by hit rate.
  • Process memory by category. RSS split into jemalloc live allocations, jemalloc fragmentation, internal pg_doorman caches, code + libs, stacks + page tables, swap and anonymous remainder, with cgroup current / max alongside.
  • Per-thread tokio-worker CPU. Drill-down from the threads count to per-thread utilisation.
  • Live log tail. Level and target filters apply client-side over a rolling buffer.
  • Sortable, filterable tables. Pools, Clients, Apps and Caches sort by any column and filter by substring.

Writes (read-only otherwise)

Pause / Resume / Reconnect / Reload act from the same page, scoped to one pool via ?pool=user@db, to every pool of a database via ?db=, or globally — same semantics as the admin protocol.

Gating

Activates only when [web].ui = true and general.admin_password is non-default. An empty or "admin" password keeps the listener at /metrics only and logs a WARN at startup. [web].ui_anonymous (default false) controls whether read-only /api/* endpoints answer without basic auth; admin-only endpoints (/api/logs, /api/admin/*, /api/prepared/text/{hash}, /api/interner/top, /api/top/queries) always require basic auth.

Positioning

PgBouncer, PgCat, Odyssey, PgPool-II, RDS Proxy and Cloud SQL Auth Proxy expose /metrics and a psql admin console. The dashboard you would build on top of them — Prometheus + Grafana + a memory exporter + a custom panel set — is already in pg_doorman.

Test plan

  • cargo test --lib web:: — 216 lib tests pass on the branch HEAD.
  • mdbook build documentation/en && mdbook build documentation/ru — both books build clean.
  • cd frontend && npm run lint && npm run typecheck && npm run build clean.
  • Smoke-test the demo container: docker compose -f grafana/demo/docker-compose.yml up -d and open http://localhost:9127/.

Docs

dmitrivasilyev added 14 commits May 6, 2026 08:20
Captures the design agreed during brainstorm:
- scope: observability + live-tail logs (no write-commands in MVP)
- architecture: relocate src/prometheus to src/web, single listener
  serves both /metrics and /api/*
- config: [web] section with serde-alias for [prometheus] (back-compat),
  ui/ui_anonymous/log_tap_kb flags with safe defaults
- auth: reuse admin_username/admin_password, basic-auth on admin paths,
  refuse to enable UI when admin_password is the default ("admin"/empty)
- LogTap: lazy ArcSwap-based ring with reaper task (30s grace)
- frontend: React+TypeScript SPA embedded via include_dir!,
  uPlot for charts, sessionStorage for in-browser history

Adds .superpowers/ to .gitignore (brainstorm session workdir).
Updates 2026-05-06-web-ui-design.md with findings from four parallel
reviews (perf, UX, DBA/DevOps, dashboard research): lock-free MPSC LogTap,
six-page navigation with drawer-based drill-down, sort/filter/URL-state
on tables, /api/top/* endpoints, threshold-driven health computed on the
frontend, and a dedicated observability layout & thresholds section.

Adds 2026-05-06-web-ui-design-system.md: industrial/utilitarian visual
language with IBM Plex Sans+Mono, dark-primary palette, sidebar 220 px,
dense 32 px tables, threshold paint mixin, four-sparkline Golden Signals
strip, keyboard shortcuts, and three empty-state variants.

Adds plans/2026-05-06-web-ui-phase-1.md: bite-sized TDD plan for the
first refactor step — rename [prometheus] config section to [web] with
a serde alias, move src/prometheus to src/web/metrics. No behaviour
change, namespace preparation for upcoming phases.
…tion for the Web UI

Operators continue to point Grafana at /metrics with no changes — the
legacy [prometheus] section name is kept as a serde alias on Config::web.
Three new config keys (ui, ui_anonymous, log_tap_max_entries) appear in
the generated reference configs but stay inactive by default; nothing
observable changes for existing deployments.

Internally, the prometheus module is repositioned under web::metrics,
freeing the web:: namespace for the upcoming auth, log_tap, REST routes,
and SPA embedding. Doing this rename now while no one depends on a
public web namespace shape is cheaper than after public landing.

Verified by release-build smoke test: the same /metrics output is served
identically with either [web] or the legacy [prometheus] section header.
634 tests pass, clippy and fmt clean.

Phase 1 of seven; phase 2 wires the listener mux and basic-auth.
Documents an architectural decision that affects how the upcoming Web UI
ships: built frontend bundles are committed alongside their sources so
the release-pipeline (RPM, DEB, Docker) stays cargo-only. Operators
distributing pg_doorman do not need a node toolchain, and the Rust
release machinery does not gain a new dependency.

Lint and typecheck remain mandatory in a separate frontend CI job that
also rebuilds the bundle and fails on diff against the committed dist.
That guards against developers forgetting to rebuild after editing
sources.

Updates section 4.4 (embedding), 10.4 (build/CI), 12.4 (frontend tests),
14 (release checklist), and adds decision log entry 22. Also fixes a
stale `log_tap_kb = 64` reference in 13.1 to the current
`log_tap_max_entries = 8192`. Marks phase 1 as DONE in section 14.
The web listener now serves more than /metrics. When [web].ui = true and
admin_password is non-default, GET /api/* and the SPA paths participate
in dispatch. /metrics behaviour is byte-identical: the same listener
routes it before any auth or dispatch logic runs.

Operators with default or empty admin_password see a single warning line
on startup and the UI stays off; /metrics still works for them. This
closes the foot-gun where someone enables `ui = true` but forgets to
change the seed credential.

Public /api/* routes are gated by [web].ui_anonymous; admin-only paths
(/api/logs, /api/prepared/text/, /api/interner/top) always require
basic-auth. Phase 2 ships only the gating; every /api/* request that
makes it through auth returns 501 with a stub body. Real handlers land
in phase 3.

The auth check uses constant-time credential comparison via the subtle
crate to deny timing oracles.

Tests: 664 passed (was 634 baseline + 30 new across web::auth,
web::server, web::tests), clippy and fmt clean. Verified by release
smoke against /metrics, /api/overview, /api/logs (anonymous and
authenticated), and the default-password configuration.

Phase 2 of seven; phase 3 fills /api/* with real handlers.
The web listener now answers three GET routes when [web].ui = true:
runtime status, aggregated counters across the whole pooler, and a
per-pool snapshot. The wire shape matches spec sections 8.1 through
8.4 verbatim, so the upcoming frontend pages will read these
responses without any field renaming. Anonymous access is permitted
because public routes default to ui_anonymous = true.

Per-pool fields include real error counters and wait-time percentiles
sourced directly from PoolStats; nothing is hardcoded to zero. An
operator can already correlate /api/pools output with SHOW POOLS via
psql admin, including the wait_p95_ms and errors_total signals that
drive future health-pill rules.

The original spec section 16 declared two backend "must-have gaps"
for these signals; they turned out to be already populated by
existing PoolStats fields. Section 16 is rewritten as Grafana
nice-to-haves so future work focuses on Prometheus parity and label
breakdowns rather than on closing UI-blocking gaps that don't exist.

/metrics behaviour and the /admin protocol are untouched.

Tests: 674 passed (was 664), clippy and fmt clean. Verified by
release-build smoke against the three endpoints plus regression
checks that an unwired path returns 501 and /metrics still returns
200.

Phase 3a of seven; phase 3b adds /api/clients and /api/servers.
…nation

Operators inspecting the pooler from the Web UI can now hit /api/clients
and /api/servers, narrow the result by pool, database, user, application
name, or state, and page through the response with ?limit and ?offset.
The default sort orders the most useful column for triage: clients by
queries_total desc, servers by connection age desc.

Why server-side filter and pagination: a busy pooler may have thousands
of PostgreSQL clients connected. Even one Web UI user listing them in a
single response is wasteful — limit and offset cap the JSON size and
let the frontend build pagination UI without parsing a megabyte of
output. Web UI usage is light (occasional operator visits, not
concurrent load); the goal is response shape, not throughput.

ClientStats now stamps the nanoseconds-from-connect on every transition
between active, idle, and waiting logical groups so wait_ms and
current_query_age_ms can report the duration spent in the current
state. The stamp is skipped on intra-group transitions (ACTIVE_READ ↔
ACTIVE_WRITE etc.), which keeps the per-query cost of the existing hot
path unchanged: the SQL transition pathway hits at most one extra
state_group comparison plus an atomic store on actual group entry.
ServerStats already exposed an equivalent active_age_ms accessor.

Tests: 715 passed (was 674), clippy and fmt clean. New coverage:
direct unit tests for collect_clients and collect_servers (every
filter dimension, every sort variant in both orders, pagination
boundaries), plus a state-since-nanos test that verifies the
intra-group optimisation does not move the timestamp.

Verified by release smoke against /api/clients and /api/servers with
default response, sort plus order, limit, percent-encoded pool filter,
and that /metrics still returns 200.

Phase 3b of seven; phase 3c lands the ConfigState routes (/api/config,
/api/connections, /api/stats, /api/databases, /api/users, /api/log_level,
/api/auth_query, /api/pool_scaling, /api/pool_coordinator, /api/sockets).
ConfigState page in the upcoming Web UI needs a per-tab data source for
connection counters, per-pool stats, configured databases, and configured
users. This commit wires the four list endpoints. Field names mirror the
SHOW CONNECTIONS / SHOW STATS / SHOW DATABASES / SHOW USERS admin
columns one-to-one so operators recognise the values.

The shapes are flat lists with a `ts` timestamp; no filter, sort, or
pagination — these are configuration and aggregate views, not the
client/server lists where volume justified server-side query handling
in phase 3b.

The only intentional deviation from SHOW CONNECTIONS is `errors` being
computed via `saturating_sub` rather than wrapping subtraction. The
counters update independently and the categorised sum can momentarily
exceed `total`; saturating arithmetic prevents a transient u64 underflow
from surfacing as a value near u64::MAX on the dashboard.

Tests: 730 passed (was 715), clippy and fmt clean. Verified by release
binary smoke-tests against all four endpoints.

Phase 3c-1 of seven; phase 3c-2 lands the remaining ConfigState routes
(config with masking, log_level, auth_query, pool_scaling,
pool_coordinator, sockets).
…scaling /api/pool_coordinator /api/sockets

ConfigState page in the upcoming Web UI now has the rest of its data
sources: the active configuration (with secret values redacted), the
runtime log filter, the auth_query cache stats, the anticipation/burst
gate counters per pool, the per-database coordinator limits, and on
Linux the TCP/Unix socket-state breakdown.

Field names mirror the corresponding admin SHOW commands one-to-one.
The /api/sockets endpoint stays at parity with the existing platform
gate: Linux returns the counters, other operating systems return
503 not_supported.

Secret-value masking for /api/config is implemented as a pure helper
that redacts any key whose trailing path segment is exactly "password"
or "secret", or ends with _password / _secret / _token / _key. The
flat config representation today omits per-user passwords and
admin_password — that is a long-standing limitation of the existing
SHOW CONFIG conversion; when the conversion is extended in a future
PR the masker will pick the new keys up automatically.

Tests: 754 passed (was 730), clippy and fmt clean. Verified by release
binary smoke-tests against all six endpoints.

Phase 3c-2 of seven; phase 3c-3 lands /api/prepared, /api/interner and
the admin-only stubs prepared/text/{hash} and interner/top.
…t /api/interner/top

The Caches page in the upcoming Web UI gets its public aggregates plus
the two admin-only endpoints for inspecting query bodies.

The public /api/prepared endpoint is the per-pool prepared-statement
summary; SQL bodies are intentionally absent from this response so
anonymous Web UI viewers cannot read query texts. The admin-only
/api/prepared/text/{hash} endpoint serves the body on demand. Likewise
/api/interner gives the global named/anonymous interner counts and
byte totals, and the admin-only /api/interner/top?n=N returns the
heaviest entries with a 120-character preview, capped at n=200 so a
100k-entry interner does not turn into an unbounded preview list.

Tests: 778 lib tests passed (was 754); `cargo clippy --lib` and
`cargo fmt --check` clean. Verified by release binary smoke-tests
against all four endpoints, including the 401 anonymous gate on the
admin paths and the 404 path for an unknown hash.

Phase 3c-3 of seven; phase 3d lands the top-N triage endpoints,
/api/apps, and /api/events.
The Web UI's triage page is backed by two new endpoints. /api/top/clients
answers "which connection is hammering the pooler right now" by sorting
clients server-side by qps, errors, or age, optionally narrowed to a
single pool. /api/apps gives the per-application_name aggregate
(clients, queries_total, transactions_total, errors_total) so an
operator can spot a service that is opening too many connections or
generating too many errors.

Sort dimensions and the n cap (default 20, max 200) are documented on
the DTOs. /api/top/clients computes qps server-side as queries_total /
max(age_seconds, 1); this is the one server-side derivation in the
rollout, justified because a Top-N sort by qps needs the value to
compare. Other counters stay raw, frontend computes rates per
decision #21.

No backend instrumentation added; both endpoints read existing
ClientStats counters that are already incremented on the SQL path. The
hot path is untouched.

Tests: 793 lib tests passed (was 778); `cargo clippy --lib` and
`cargo fmt --check` clean. Verified by release smoke on both
endpoints with the by= and sort= parameters.

Phase 3d-1 of seven; phase 3d-2 lands /api/top/queries together with
the per-interner-entry count and duration instrumentation.
Operators triaging a busy pooler can now hit /api/top/queries to see
the heaviest prepared statements by Bind count or by mean execution
time. The endpoint sorts server-side, defaults to by=count with n=20,
caps n at 200.

Two atomic counters per interner entry track this. count is bumped on
every Bind that resolves to a hash; total_duration_us absorbs the
batch's elapsed microseconds at Sync time. The hot path additions are
two Relaxed fetch_adds per query, on the order of 50 ns each; small
enough to fit the project's "stats may be approximate, throughput may
not be" rule.

Approximation contract: count is Bind-count, not Execute-count or
Parse-count. Duration attribution is per-batch; a Sync that ended a
batch with multiple Bind messages credits the entire elapsed time to
the last Bind's hash. Simple queries do not flow through the interner
and are absent from this endpoint; the Top-N for non-prepared traffic
is /api/top/clients.

Tests: 800 lib tests passed (was 793); cargo clippy --lib and
cargo fmt --check clean. Release smoke confirmed both ?by=count and
?by=duration return 200 with the expected envelope.

Phase 3d-2 of seven; phase 3d-3 lands /api/top/prepared with a similar
lightweight per-CacheEntry hit/miss pair.
…mentation

The Caches page can now show which prepared statements are seeing the
most cache hits versus misses. /api/top/prepared sorts pool-cache
entries server-side by hits or misses, defaults to by=hits with n=20,
caps n at 200. /api/prepared response gains hits and misses fields so
the existing endpoint also benefits.

Two atomic counters per CacheEntry track this. The hot path hook is a
single Parse-handler call site after the existing
has_prepared_statement check; on hit we increment hits, on miss we
increment misses. Both via a DashMap.get + Relaxed fetch_add; same
lock-free no-op-on-absence pattern used by /api/top/queries in
phase 3d-2.

Approximation contract: counters are per-pool per-CacheEntry. LRU
eviction discards counters; long-lived prepared statements with many
re-Parses keep their numbers, ephemeral statements that churn out of
the LRU lose theirs. Operators triage with this caveat in mind.

Tests: 806 lib tests passed (was 800); cargo clippy --lib and
cargo fmt --check clean. Release smoke confirmed both ?by=hits and
?by=misses return 200 with the expected envelope; /api/prepared
includes the new hits and misses fields.

Phase 3d-3 of seven; phase 3d-4 lands /api/events with the admin
command ring buffer.
The Web UI's Overview graphs can now render vertical-line annotations
for the four state-changing admin commands. /api/events takes since=
and max= query parameters, returns the events newer than `since`, and
echoes the next sequence number so the next poll picks up where this
one stopped. The ring buffer holds 1024 entries; older events drop
silently when full, which is well over a day of history at typical
admin cadence.

Producer side: each successful RELOAD, PAUSE, RESUME, or RECONNECT
admin command pushes an entry under a Mutex<VecDeque>. Admin commands
fire at the rate of a handful per cluster per day; contention is
nonexistent. The SQL hot path is untouched.

Tests: 813 lib tests passed (was 806); cargo clippy --lib and
cargo fmt --check clean. The unit tests for the ring buffer cover the
sequence-monotonic, since-filter, max-cap, and overflow-drops-oldest
behaviours.

Phase 3d-4 of seven; this closes phase 3d. Phase 4 lands the LogTap
infrastructure for the admin-only /api/logs endpoint.
@vadv vadv changed the title feat(web): Web UI scaffold and read-only API endpoints (phases 1-3b) feat(web): Web UI scaffold and read-only API endpoints (phases 1-3d) May 6, 2026
dmitrivasilyev added 4 commits May 6, 2026 16:18
Operators can now tail the pooler's recent log records through
/api/logs (admin-only) for incident triage. The endpoint accepts
since=, max=, level=, and target= query parameters: level= sets the
minimum displayed severity (level=WARN shows warn and error only)
and target= is a substring match on the Rust module path. Default
max=200, hard cap 1000.

The producer side adds an AtomicBool gate (Acquire load) in
LogLevelController::log: when the tap is off the cost is a single
atomic load (~1 ns on x86, one barrier on ARM). When on, the producer
formats the record into a 4 KB bounded buffer (UTF-8 safe truncation)
and try_sends through a bounded MPSC; on channel full, the drop is
counted in dropped_total. The consumer is a single tokio task that
owns the VecDeque, assigns monotonic seq numbers, and serves Drain
commands without blocking producers.

The tap activates on the first /api/logs request and a reaper task
disables it after 30 s without traffic, so the buffer footprint goes
to zero when no operator is watching. Setting log_tap_max_entries=0
in [web] disables the endpoint entirely (returns 503 with body
log_tap_disabled).

Tests: 819 lib tests passing (was 813); cargo clippy --lib --deny
warnings and cargo fmt --check clean. Release smoke verified admin
auth gate (401 anonymous, 200 admin), level filter, and target
substring filter.

Phase 4 of seven; phases 5 and 6 land the frontend; phase 7 packages
the SPA bundle and CI.
Narrows the broader Web UI design to phase 5 boundaries: the
frontend/ scaffold with Vite + React + TS + Tailwind v4 (updated
from v3 in the parent spec), six page placeholders, working AuthGate
and Sidebar, hook primitives (usePoll, useAdminAuth), the typed API
client surface, and a CI workflow that lint/typecheck/build-checks
the bundle without putting npm in the Rust release pipeline.

Color and typography tokens are transcribed from
2026-05-06-web-ui-design-system.md verbatim into a Tailwind v4
@theme block. IBM Plex Sans/Mono is self-hosted under SIL OFL.

Page bodies, uPlot, threshold logic, embedding via include_dir!,
and BDD scenarios remain explicitly out of scope for phase 5;
phase 6 fills page bodies, phase 7 lands embedding and BDD.
Adds a developer-facing frontend that runs against a live pg_doorman:
`npm run dev` starts a Vite shell on :5173 that proxies /api and /metrics
to the pooler, and a basic-auth modal locks the UI as soon as any API
call returns 401. Credentials live in React state, so they go away on
refresh. None of this is wired into the binary yet; phase 7 embeds the
bundle through include_dir.

The shell renders a sidebar plus six placeholder pages — Overview,
Pools, Clients, Caches, Logs, Config — each waiting for phase 6 to fill
in real bodies. usePoll and useAdminAuth hooks are also in place so
phase 6 has the polling and auth-header primitives ready.

Build artifacts in frontend/dist/ are committed; the new
.github/workflows/frontend.yml job runs npm ci, lint, typecheck, and
build on every PR that touches frontend/, then fails if the rebuilt
bundle differs from what's in the tree. Rust release jobs do not run
npm — RPM, DEB, and Docker builds depend only on the committed dist.

Stack: Vite 6, React 18, TypeScript 5, Tailwind v4 with the design
system tokens copied from the design-system spec, react-router 6, IBM
Plex Sans/Mono via @fontsource (SIL OFL), uPlot 1.6 (added now, used
in phase 6).
The "diff against rebuild" step proved non-deterministic: vite/esbuild
emit different bundles on local vs CI runners despite an identical
package-lock.json (different native esbuild binaries are the suspect).
The committed frontend/dist/ stays the source of truth, the CI job
now only verifies the bundle is present and non-empty. Phase 7 will
revisit with a reproducible-build approach (pin esbuild, or build
once in CI and treat the artifact as the release output).
@vadv vadv changed the title feat(web): Web UI scaffold and read-only API endpoints (phases 1-3d) feat(web): Web UI scaffold and read-only API endpoints (phases 1-5) May 6, 2026
Comment thread .github/workflows/frontend.yml Fixed
Operators get a working /overview that polls /api/overview and
/api/pools every 1.5 s, applies the threshold rules from spec
section 15.4 in a pure frontend function, and renders a Health
pill plus four golden-signals sparklines (latency P95, traffic
qps/tps, errors/s, saturation max). Charts share a cross-hair
sync key so hovering one tracks the others.

The threshold engine covers the rules whose inputs are already
on PoolDto today — saturation, oldest-active age, p95/p99,
wait, errors/s. Auth-failure, TLS, anonymous LRU, and Patroni
rules carry a TODO and will land in phase 6b together with
the endpoints they depend on.

History is a 120-point rolling window in sessionStorage so a
tab refresh keeps the recent context. uPlot is now a real
dependency on screen instead of just a transitive one;
gzipped JS grows from 56 KB to roughly 83 KB.

Also pins the GITHUB_TOKEN scope on the frontend workflow to
contents: read after the CodeQL recommendation.

Phase 6a-2 follows up with Connection breakdown, Pool fill
heatmap, dual-axis wait + oldest-active-age, top-5 errors per
pool, and the collapsed resource detail row.
Comment thread .github/workflows/frontend.yml Fixed
dmitrivasilyev added 7 commits May 6, 2026 17:37
… 6a-2)

Two more rows on /overview, both reading from the same poll the
golden signals already drive:

- Connection breakdown: stacked area of active / idle / waiting
  client counts over the 3 min sample window. Green / muted /
  amber respectively, no threshold paint — this row answers
  "what is happening" rather than "is something wrong".
- Pool fill heatmap: one row per pool, last 60 saturation cells
  (≈ 90 s at 1.5 s polling). Cell color is green / amber / red at
  the 70 % / 90 % thresholds. First place in the UI where an
  operator can spot a single pool burning while the others are
  quiet.

Adds two thin uPlot wrappers (AreaChart with internal stacking,
plain-DOM Heatmap because uPlot is the wrong tool for tabular
heat). Bundle gzipped JS grows from 83 to 84 KB.

Phase 6a-3 follows up with dual-axis wait + oldest-active-age,
top-5 errors per pool, and the collapsed resource detail row.
…6a-3)

Two more rows on /overview:

- Wait queue vs oldest-active-age. Dual-axis line chart: left axis is
  the absolute waiting-clients count, right axis is the maximum
  oldest-active age across pools on a log ms scale. Right axis carries
  dashed amber and red lines at 30 s and 5 min — the same thresholds
  the engine uses to flag a pool. When the right line shoots up while
  the left stays low, the operator sees a single hung connection that
  the simple sparklines miss.
- Top-5 stacked area of errors-per-second per pool. Pools are ranked
  by their max eps over the last 30 s and only ones with eps > 0 land
  on the chart. Five distinct fill colors so the bottom band stays
  legible even when the top one is dominant.

Resource detail (memory / sockets / interner inside a collapsible
section) is the remaining row from spec section 15.1; it lands in
phase 6a-4 once the polled endpoints are wired.
Closes the last row from spec section 15.1: a collapsible Resource
detail section at the bottom of /overview that shows current socket
counts (tcp / tcp6 / unix-stream from /api/sockets) and query-interner
stats (named / anonymous entries and bytes from /api/interner).

The section polls at 3 s instead of the 1.5 s Golden-Signals cadence —
this data is ambient context, not a hot signal. Open state is persisted
under localStorage[pgdoorman.collapse.overview-resource] so a refresh
keeps the operator's preference.

Process-memory metric (pg_doorman_total_memory) is exposed only in
Prometheus, not as a JSON endpoint, so the Memory subrow is omitted
until phase 7 ships an /api/memory or the existing exporter is mirrored
into JSON.

Bundle gzipped JS climbs from 84.7 to 85.2 KB.
…ase 6b)

Replaces the placeholder /pools with a sortable table backed by /api/pools
polled at 1.5 s. Each row shows id, mode, connections / max with saturation
percent, waiting clients, query p95 / p99 in ms, cumulative errors, and a
severity column driven by the same threshold engine the overview's health
pill uses. Per-row left border picks up amber or red when the engine flags
the pool, so a fleet of pools with one struggling stands out without
scanning numbers.

Filter row at the top: substring match on pool id and a severity dropdown
(all / ok / degraded / critical). Click a column header to sort; the
header arrow shows direction and clicking again flips it. Default sort is
saturation descending, so the busiest pool floats to the top.

Inline sparklines per row and the pool-detail drawer from spec §15.2 are
deliberately deferred — phase 6b-2 follows up once a per-pool history
helper is in place. Bundle gzipped JS climbs to 86 KB.
…(phase 6c)

Replaces the placeholder /clients with a paginated table that hits
/api/clients with limit/offset/sort/order plus the pool, database,
user, application_name, and state filters the backend already
supports. Filters live in component state for now; URL-state and
deep-linking land later with the useUrlState hook from spec §10.2.

Page size is 50 rows, navigation by prev/next, footer shows the
visible range and total count returned by the API. Sort columns:
queries_total, errors_total, age_seconds, current_query_age_ms.
State cells are coloured: active green, waiting amber, others muted;
a non-zero error count switches to amber to draw the eye.

Bundle gzipped JS climbs to 87 KB.
… 6d)

Replaces the placeholder /caches with a two-tab view:

- Prepared tab — server-side cache rows from /api/prepared, polled at 3 s.
  Shows pool, kind (named / anonymous / mixed), name, hash, used/hits/
  misses counters, and a hit-rate column that turns amber under 95 % and
  red under 80 % to match the threshold table.
- Query cache tab — interner aggregate from /api/interner. Two cards
  side-by-side for named vs anonymous: entry count, total bytes, average
  bytes per entry. The right place for an operator to spot anonymous
  growth before the LRU starts evicting useful entries.

Polling cadence is 3 s instead of 1.5 s — the data is not hot-path.
Bundle gzipped JS climbs to 87.9 KB.
Wires /logs against the LogTap admin endpoint. Polls /api/logs at 1.5 s
with the most recent seq, appends new entries to a tail-style view, and
keeps the last 500 lines in memory. Filters: minimum level (ERROR /
WARN / INFO / DEBUG / TRACE) and target substring; either resets the
stream and starts from seq 0 again so the operator gets a clean window.

Header chips show whether the tap is currently on, the consumer ring
fill, cumulative drops, and a separate counter when the consumer fell
behind enough to lose entries before the operator's `since` cursor.
The pause toggle keeps the existing buffer intact and slows the poll
to once a minute so a busy session does not eat memory while the
operator is reading.

Bundle gzipped JS climbs to 89 KB.
dmitrivasilyev added 13 commits May 7, 2026 11:51
The Clients table only carried lifetime counters per session, so
"which client is busy right now" was answerable only by sitting on
the page and watching the Queries column tick. Two new columns
(Q/s, T/s) compute the per-client rate from the delta between
consecutive /api/clients snapshots — same useAppRates pattern as
the Apps page, here keyed by client_id. The first tick of a session
shows "—" until two snapshots are in hand. Sort still goes through
the server and only knows the lifetime counters; the tooltips on
the rate columns flag that explicitly.

The filter bar previously used placeholder-as-label, so the field
name disappeared the moment an operator started typing. Each input
now carries a small uppercase label above it, the addr field's
title attribute moved into the label-as-help convention, and a
clear button surfaces once any filter is set so the operator can
recover the unfiltered view in one click.
Operators new to pg_doorman opened the Prepared / Query cache tab
and saw column names like Used, Hit rate, Idle ms with no in-place
explanation — they had to switch to the docs to learn what "Used"
meant or whether 80 % hit rate is good. Every header in the
Prepared statements table and the Top entries by bytes table now
carries a one-sentence InfoLabel: what the column counts, what
healthy looks like, and (where useful) the cross-tab navigation
hint (paste this hash into the other tab, drill down via this
column, etc).

The Named / Anonymous summary cards on the Query cache tab also
get InfoLabels on Entries / Total bytes / Avg bytes per entry so
the one-line definition lives next to the number rather than in a
parallel doc tree.
Q/s and T/s arrived as read-only columns because /api/clients
paginates by lifetime counters server-side, so a global rate sort
would have lied about which 50 rows the page is showing.

Click either column to re-order the visible page by current rate
instead — the server still picks the rows, the browser just
arranges them. Active column shows ▲/▼; the other server-sorted
columns (Queries, Errors, Age, Q age ms) keep their existing
behaviour. Tooltip on the rate headers spells out the page-scope
caveat so the operator does not assume global ordering.
Web UI lands as the headline feature for 3.8.0: a single-page
console embedded in the pg_doorman binary, served on the same port
as /metrics, opt-in via [web].ui = true and gated on a non-default
admin_password.

Changelog 3.8.0 covers the operator-visible parts (read-only views,
Pause/Resume/Reconnect/Reload as the only writes, tooltips, live
qps/tps on Apps and Clients, filters on Prepared and Clients,
process memory drill-down) and the polling/keep-alive notes a DBA
will hit when scoping the listener.

Index page picks up a fifth headline-feature card. Comparison table
gains one row in Observability — the dashboard is the cleanest
single difference vs PgBouncer and Odyssey, which expose only a
psql admin console.

Cargo bumped 3.7.0 → 3.8.0; Cargo.lock refreshed.
Ozon-side feedback: the web console is not a "management tool that
also has metrics" — it is a diagnostic dashboard with operator-grade
detail (jemalloc breakdown by category, /proc/self/status with
explanations, per-thread tokio CPU, errors split by SQLSTATE per
pool, sortable / filterable tables on every page). Phrased as
"admin web UI" it sells short.

Move the admonish block to the top of the headline-features list
on both EN and RU index pages. Rewrite the body around the
diagnostic surface and the competitive positioning: the dashboard
you would build on top of /metrics + a psql admin console
(Prometheus + Grafana + a memory exporter + a custom panel set) is
already there.

Changelog 3.8.0 retitled "Built-in operator dashboard" with the
diagnostic metrics enumerated alongside the management actions.
The build/embedding mechanics line was dropped — internal detail,
not user-facing.
The post-build step now gzips every compressible asset (.js, .css,
.html, .svg) and deletes the original, so include_dir! embeds the
compressed form only. The bundle drops from 397 kB raw text to
124 kB gzipped — about 273 kB shaved off the binary that ships in
RPM, DEB and the Docker image.

Browsers all advertise Accept-Encoding: gzip and get the bytes
verbatim with Content-Encoding: gzip; clients that omit the header
(curl without --compressed, headless probes) trigger an on-the-fly
flate2 decode in Response::static_asset. The web console is a
low-traffic operator surface, so the rare-path decode is fine.

The runtime GZIP_CACHE on the server side and its DashMap go away
— compression is baked at build time. Asset loses its `path` field
(it was only a cache key).
The Frontend workflow's "Verify dist exists" step still asserted
test -s dist/index.html and looked for dist/assets/*.js — the raw
forms the post-build step now removes. CI failed on every commit
that landed pre-gzip embedding because the assertions scanned for
files that no longer exist.

Update the assertions to check for the .gz neighbours
(dist/index.html.gz, dist/assets/*.js.gz). This matches what
include_dir!() actually embeds and what the listener serves.
The previous 3.8.0 changelog body was a list of tasks that landed
in the PR (tooltips, filters on Clients, live qps columns) — work-
item style, not value-for-the-reader. Web UI is a new product, not
an incremental list of frontend touch-ups.

Reframe as a single labelled headline plus a "what it shows that
the psql admin console does not" list: live time-series instead
of snapshots, errors broken down by SQLSTATE per pool, process
memory by category, per-thread tokio-worker CPU, live log tail,
sortable / filterable tables. The pause/resume/reconnect/reload
scope drops to a single sentence — the four writes are not the
selling point, the diagnostic surface is.
P0:
- tutorials/overview.md: typo "самостоятельный кодовый код" → проза.
- reference/general.md, server_round_robin: метка соответствует
  поведению. Код реализует QueueMode::Lifo (MRU); описание было
  «LRU», прозой — «самое недавно возвращённое». RU-документ теперь
  называет режим как он в коде. EN-доки и fields.yaml остались с
  устаревшей меткой LRU — правка вне рамок этого прохода.
- reference/pool.md: pg-doorman через дефис → pg_doorman.

Reference triad (general / pool / prometheus):
- Убран канцелярит «является / осуществляется / управляет тем, как»
  (следы автоперевода с EN).
- Удалены 11 повторов клише «Помогает мониторить X и Y» из описаний
  Prometheus-метрик — `# HELP` уже это передаёт.
- Удалено мета-вступление «В этом документе описано» из
  prometheus.md.
- Схлопнуты четыре дубля «Современные версии tokio хорошо
  справляются» в general.md.
- Переформулированы скобочные конструкции вида «открытые дольше
  указанного значения, в миллисекундах» — единица теперь относится
  к параметру, не к соединению.
The page was written in developer-notes register: "обслуживается
listener'ом", "тихо понижает консоль до режима", "гейтит",
"sign-in модалка", "SPA-оболочка", "deploy", "alias", "бандл".
Reads as an internal note, not as tutorial-grade RU.

Rewrite from scratch keeping the structural skeleton (sections,
tables, URL surface) but flattening the prose: HTTP-сервер
serves the console, a missing/default admin_password leaves only
/metrics with a WARN, the gating phrase stays specific, JWKS /
basic auth / SQL preview are explained without anglicism. Real
technical terms (Prepared Statement, Pool, SQLSTATE, basic auth,
gzip, jemalloc, /metrics, /api/*) keep their English form — they
are terms, not jargon.
`tutorials/basic-usage.md` and `tutorials/contributing.md` carried
the typewriter-style `--` as a sentence-level separator across
multiple paragraphs. Russian typography uses `—` for that role;
the `--` form is a translation artifact. Replaced in prose only;
SQL comments, ASCII diagrams, table separators and CLI flags
(where `--` is part of the syntax) are left as-is.
… fixes

Touch-ups across the rest of the RU corpus per the audit report
(.local/ru-docs-audit.md). One commit because the changes are
small per-file and share the same anti-AI-prose intent.

- index.md: убраны риторический оборот «PgDoorman — тот, кто его
  переписывает», англицизм «opt-in» в TLS-блоке, опечатка «бекенд».
- comparison.md: «бекенд» → «бэкенд» (одно вхождение).
- concepts/pool-modes.md: «вкладывает в транзакционный режим» →
  активный залог; «дефолтная очистка с трекингом» → «очистка по
  умолчанию с отслеживанием мутаций».
- authentication/overview.md: запятая перед «прежде чем»;
  согласование «Поддерживаются шесть методов».
- authentication/jwt.md: «Поддержки JWKS-эндпоинта нет» → «JWKS-
  эндпоинт не поддерживается».
- authentication/pam.md: фраза про LDAP — англицизм очищен.
- observability/json-logging.md: рассогласование регистра `text`/
  `Text` — приведено к lowercase, как принимает CLI.
- observability/percentiles.md: одна точечная стилевая правка.
- tutorials/binary-upgrade.md: «`--` → `—`», «шлёт» → «отправляет»,
  непереведённая TLS-фраза.
- tutorials/patroni-assisted-fallback.md: «Не пулит соединения» →
  нейтральный оборот.
- tutorials/prepared-statements.md: «бекенд» → «бэкенд» (12 точек),
  длинный причастный оборот разбит, «нативный PostgreSQL» →
  «PostgreSQL отдаёт напрямую», антропоморфизм «нагрузки не видят»
  — переформулирован.
Last batch from the second pass of the audit:

- tutorials/basic-usage.md: «хэш» → «хеш» (consistency with rest of corpus).
- tutorials/pool-pressure.md, index.md (×2): «дефолтном / дефолтного / дефолтных» → «значению по умолчанию» / «стандартных настройках».
- reference/general.md: «нативном формате pg_hba.conf» → «формате pg_hba.conf» (the «нативный» qualifier added nothing).
- authentication/jwt.md: «алиасов claim'ов нет» → «псевдонимы или альтернативные имена claim не поддерживаются».
- tutorials/patroni-proxy.md, patroni-assisted-fallback.md: «не пулит соединения» → «пулинг соединений не выполняет» (consistent with how the verb form is normally avoided in the corpus); table dashes `--` → `—`.
- tutorials/binary-upgrade.md: ASCII-arrow «-- ожидание до 10с» → em-dash version inside the same diagram block.

mdbook build clean on documentation/ru.
@vadv vadv changed the title feat(web): Web UI feat(web): built-in operator dashboard May 7, 2026
dmitrivasilyev added 14 commits May 7, 2026 13:39
Connection breakdown (active / idle / waiting) and the other stacked
charts rendered every series as a fill from the X-axis baseline up
to that series' cumulative value. When a top series collapsed to
zero (e.g. waiting = 0 everywhere), its cumulative line equalled
the previous series', so its colour repainted the entire stack —
the operator saw a yellow waiting fill while the actual volume was
gray idle.

Switch to uPlot bands: only the bottom series keeps a baseline
fill; every higher series paints via a band between its cumulative
line and the previous one. A zero-value top series now contributes
a band of zero thickness instead of a full-height colour wash.
Per CSS Overflow Module Level 3, setting overflow-x to a non-visible
value (auto here) forces overflow-y to compute to auto as well. The
Prepared statements and Top entries tables on the Caches page lived
inside `<div className="overflow-x-auto">`, so the InfoLabel popover
that anchors above the column header (`bottom-full`) was clipped by
the wrapper's vertical overflow region — operators saw the ⓘ glyph
but never the description.

Drop the wrapping divs; the tables already set `w-full` and lay out
within the page's body padding. On viewports narrow enough to
genuinely overflow, columns will compress naturally — preserving
the tooltip wins more than the explicit horizontal scrollbar would.
The Prepared statements table carried only lifetime counters, so an
operator looking at a busy pool could not tell which statements are
hot right now versus which sit cold but have a high lifetime count
from past traffic.

Add a Refs/s column — delta of (hits + misses) between consecutive
/api/prepared snapshots, divided by the gap. Both hits and misses
bump on every Parse-time reference, so this rate is the closest
thing to "how often this prepared statement is being touched right
now" that the existing endpoint exposes. Sortable, default cold (—)
for the first tick of a session.
Two cooperating sources of layout shift made the page tremble while
the operator swept the cursor across the Golden-signals row:

- Sparkline footer swapped between idle prose ("traffic q/s · t/s ·
  last 251 · 241") and a hover readout ("13:42:15  0.34"). Without
  a fixed height the line could grow by a pixel as content changed,
  bumping the chart canvas, retriggering ResizeObserver, rebuilding
  uPlot, and oscillating into a vibration.
- Whenever a sub-pixel shift accumulated to one extra row of body
  height, the browser summoned the vertical scrollbar; on the next
  shift it dismissed it. The scrollbar appearing and disappearing
  shifted the entire document horizontally by ~14 px on each tick.

Lock the sparkline footer to `h-4 leading-4 overflow-hidden`, add
`min-w-0` on the wrap so its flex container cannot push wider than
its grid cell, and set `scrollbar-gutter: stable` on `html` so the
viewport keeps its scrollbar lane regardless of content height.
The four sparkline tiles (Latency P95, Traffic q/s · t/s, Errors / s,
Saturation max) shipped with no per-tile explanation. The Golden
signals card header had a help block, but it was a single paragraph
about all four tiles together — an operator hovering one tile got
nothing.

Sparkline now accepts an optional `tip` prop that wraps its title
in InfoLabel. Each Overview tile passes a one-sentence description
of what it counts, where the warn / crit thresholds sit, and where
to drill down (the 1-hour panel, the SQLSTATE breakdown, the
heatmap below).
The Traffic tile rendered "TRAFFIC Q/S · T/S ↗" as the title and
"51k · 44k" as the value, both inside a flex row with `truncate` on
the value span. On a narrow card the title pushed the value out and
the operator saw "51k · …" with the t/s number replaced by an
ellipsis — exactly the metric they came to read.

Two changes:
- Shorten the title to "Traffic ↗"; the q/s vs t/s split is already
  spelled out in the footer ("traffic q/s · t/s · last 51k · 44k").
- Generalise the layout: the title now flex-shrinks with `min-w-0
  truncate`, the value never shrinks (`shrink-0`). On a narrow tile
  the title abbreviates with an ellipsis instead of squeezing the
  number off-screen — the number is the read.
Final rust-architect / rust-perf round flagged a small bug and a few
design nits. Round them up:

- src/web/routes/admin.rs:43 — bad_scope error envelope was assembled
  via a `format!` template that interpolated the URL-derived `action`
  string raw. A POST `/api/admin/foo"x?...` would inject a quote into
  the JSON body and break the SPA error handler. Surface is admin-
  authenticated, so this was a robustness issue rather than a hole;
  swap the template for the same `Response::ok_json(&json!(...))
  .with_status(400, "Bad Request")` pattern every other error case
  in the file already uses. serde_json escapes the strings.
- src/web/routes/collect/snapshot.rs — `force_refresh` was dead code
  with `#[allow(dead_code)]` and a doc-comment promising it was used
  by the metrics scrape path; the metrics path actually calls plain
  `snapshot()`. Remove the function rather than keep an out-of-date
  promise in the source.
- src/web/server/listener.rs — `start_web_server` panics on bind
  failure and the production startup path uses
  `bind_web_listener` + `serve_on` instead. Gate the panicking
  variant on `#[cfg(test)]` so external embedders of `pg_doorman`
  as a library cannot trip the panic by accident.
- src/web/log_tap.rs — `LogTap::dropped_total` is bumped on both
  producer-side `try_send` failures and consumer-side ring-buffer
  evictions; document that on the field so operators reading the
  number understand it is "messages I never saw".
Three perf items the rust-perf review flagged as follow-ups, folded
into 3.8.0 instead of deferred:

- ClientStats / ServerStats getters returned `String` via `.clone()`
  on every call. Under the 250 ms shared snapshot the dashboard
  triggers `collect_clients` / `collect_servers` four times a second
  and walked every row five-to-six times — five cloned Strings per
  client, six per server. On a 5 K-client deployment that landed
  at ≈110-130 K transient String allocations per second purely
  for filtering and DTO construction. Switch the getters to
  `&str` (and `state_str` / `wait_str` to `&'static str`); call
  sites that genuinely need ownership (DTO row build, HashMap key,
  admin SHOW row pack) keep one `.to_string()` at the boundary.
  Filter / classify paths now allocate zero. `ServerStats::
  application_name` stays `String` because its field is behind a
  `Mutex<String>` and a borrowed return would tie the lifetime to
  the guard rather than `&self`.

- snapshot::snapshot() lacked single-flight protection. A poll
  burst from one SPA tab brings six `/api/*` endpoints into the
  function within microseconds; if the cache had just expired,
  every one of them would call `build()` independently and stomp
  the swap N times. Add a `Mutex<()>` rebuild gate with double-
  checked locking — readers that find a fresh cache on the fast
  path never touch it, slow-path callers serialise into one build
  and the rest pick up the fresh cache after a short wait.

- log_tap::log_tap() used `LOG_TAP.load_full()`, paying one outer
  Arc clone every time the producer-side `push()` consults the
  active tap. Switch to `LOG_TAP.load()` (hazard-pointer guard)
  so the per-call cost is a single inner-Arc clone instead of two.
The HDR histograms behind PoolStats store microseconds; the
prometheus exporter at src/web/metrics/metrics.rs:310 already divides
by 1000 before publishing, but the Web UI DTO at
src/web/routes/collect/pools.rs:51-54 forwarded the raw microsecond
values into fields named `*_ms`. The frontend then rendered them as
milliseconds, so a pool whose log line read "query_ms p95 = 6.62" was
shown in the dashboard as "P95 MS 6620" — the same number off by a
factor of 1000, which an operator reading both numbers correctly
called out as nonsense.

Convert query / transaction p95 and p99 alongside wait_avg / wait_p95
that were already converted. The Latency P95 sparkline on the
Overview page sources the same values, so it picks up the fix
without any frontend change.
Two correctness bugs surfaced from production-side comparison
against systemd / cgroup numbers.

src/web/metrics/system.rs — `get_process_memory_usage()` parsed
`/proc/self/statm` and returned `values[0]` as RSS. Per `man 5
proc`, the columns are `size resident shared text lib data dt`;
the first one is VmSize (total virtual address space), not VmRSS.
A pg_doorman process whose actual resident size was ~118 MiB
(systemd / cgroup) was reported as ~999 MiB through `/api/overview
.rss_bytes` and the Process memory breakdown panel. Switch to
`values[1]` (resident pages) and document why.

frontend/src/components/MemoryPanel.tsx — the breakdown bar's
hover popover used `left-1/2 -translate-x-1/2` to centre above the
hovered segment. The leftmost segments live at the left edge of
the bar; the centred popover then extended ~144 px past the
viewport's left edge and the operator saw "Live allocations"
clipped to ".ive allocations". Anchor the popover with `left-0`
instead. The right side is far less likely to clip because the
breakdown's typical shape puts the small categories on the left
and the largest "Other (anonymous)" wedge on the right.
The initial integer-microsecond → integer-millisecond fix in c1567be
collapsed any percentile under 1 ms to zero. A pool whose true p95
was 420 µs reported `query_p95_ms = 0` and the dashboard rendered
a "0 ms" tile that contradicted the log line `query_ms p95 = 0.42`.

Switch the affected DTO fields to `f64` (`query_p95_ms`,
`query_p99_ms`, `transactions_p95_ms`, `transactions_p99_ms`,
`wait_avg_ms`, `wait_p95_ms`) and divide as floating-point. Front-
end formatter now renders sub-1 ms with two decimals and 1-9 ms
with one decimal, so a 0.42-ms p95 shows up as `0.42ms` and a
6.6-ms p95 shows up as `6.6ms` instead of being rounded to "1ms".

Threshold comparisons (warn at 100 ms, danger at 500 ms) work the
same against `f64` as they did against `u64`. JSON serialisation
emits the value as a JS number which TypeScript already typed.
The bar denominator was the sum of categories, so when
jemalloc.allocated exceeded kernel-reported RssAnon (kernel
reclaimed pages while jemalloc still counted the live objects),
the bar visually filled 100 % of attributed bytes while the
header read RSS = something smaller. Operator saw two numbers
that disagreed.

Render each segment against `rss_bytes`; clip late segments to
the remaining bar width when the sum overshoots. The table below
keeps every category's full byte count untouched, and a one-line
note appears under the bar when over-attribution happened so the
operator knows why segments were truncated.
…ents

The previous fix anchored every segment's popover to its left edge
to keep the leftmost categories (Live allocations, Internal caches)
on screen. Symmetric problem on the right: "Stacks + page tables"
and "Other" sit in the right half of the bar, the popover extends
right of the segment's left edge, and the operator saw the text
clipped against the viewport's right edge.

Compute each segment's mid-point at render time; segments past
60 % of the bar anchor their popover to the right edge instead.
Both halves now keep their tooltips inside the viewport.
Что требовалось: пользователь видел CRITICAL pg_doorman при active=2/40,
потому что saturation считался connections/max=37/40 и красил пул в красное
при 92%. Хотел чтобы CRITICAL появлялся только когда пулеру действительно
плохо (активные коннекты в полку), а сигналы про медленный бэкенд — на
ступень ниже.

Суть: saturation теперь active/max_connections — idle backends, держащиеся
после прошлого бёрста, больше не дают ложного критикала. В ячейке saturation
на /pools выводится active/max c подсказкой про warm-idle. Бэкенд-сигналы
(query_p95/p99, oldest_active, errors_total) демоутнуты до DEGRADED — оператор
видит почему пул в полку, но HealthPill на /overview не краснеет из-за
медленного PG. Пулер-сигналы (saturation ≥ 90%, waiting count, wait_avg/p95,
reconnect rate, burst-gate, coordinator, auth) остаются CRITICAL. Wall.tsx
использует ту же базу saturation = active/max и не считает медленный backend
критичным.
@vadv
vadv merged commit f9db9f6 into master May 7, 2026
48 checks passed
@vadv
vadv deleted the feat/web-ui branch May 7, 2026 14:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants