A full-stack infrastructure monitoring system built with Java 21, Spring Boot, React, PostgreSQL, Redis, STOMP over WebSockets, Prometheus, Grafana, and Docker Compose. Agents collect host CPU, memory, and disk utilization; the server persists heartbeats, detects failures, evaluates alerts, and pushes live updates to an authenticated dashboard.
flowchart LR
A["Java monitoring agents<br/>OSHI metrics"] -->|"HTTPS/REST<br/>X-Agent-Key"| S["Spring Boot API"]
U["Operator browser<br/>React dashboard"] -->|"REST + JWT"| S
S -->|"STOMP events"| U
S -->|"JPA/Flyway<br/>source of truth"| P[("PostgreSQL")]
S -->|"latest metrics<br/>TTL cache"| R[("Redis")]
PR["Prometheus"] -->|"/actuator/prometheus"| S
G["Grafana"] -->|"PromQL"| PR
The Maven reactor contains monitoring-common, monitoring-server, and
monitoring-agent. The Vite application lives in monitoring-dashboard.
PostgreSQL is the durable source of truth. Redis accelerates latest-metric reads
but is deliberately non-authoritative. The server emits transaction-aware
realtime events only after database commits. In the container deployment, Nginx
serves the dashboard and proxies /api and /ws to the server.
- Git
- Docker Desktop or Docker Engine with Compose v2 (recommended startup)
- For source development: JDK 21, Maven 3.9+, Node.js 22, and npm
These commands are the shortest reproducible path on a machine with Git and Docker Compose. Replace the repository URL with the actual clone URL.
git clone <repository-url> monitoring-platform
cd monitoring-platform
cp .env.example .env
docker compose pull
docker compose build
docker compose up -d
docker compose psPowerShell equivalent for the copy step:
Copy-Item .env.example .envBefore exposing the stack beyond localhost, edit .env and replace every
development password, JWT_SECRET, and agent key. Wait until docker compose ps
shows the services healthy, then open:
- Dashboard: http://localhost:5173 (
adminand theADMIN_PASSWORDvalue) - API health: http://localhost:8080/health
- Prometheus: http://localhost:9090
- Grafana: http://localhost:3000
Grafana uses GRAFANA_ADMIN_USER and GRAFANA_ADMIN_PASSWORD. The Compose
startup creates three deterministic development agents through the one-shot
agent-credentials-init service; their raw keys come from .env. View startup
progress with docker compose logs -f. Stop without deleting data using
docker compose down. docker compose down -v permanently removes all named
volumes and should only be used for an intentional reset.
Start PostgreSQL and Redis, then run the server and dashboard in separate terminals:
docker compose up -d postgres redis
mvn -pl monitoring-server spring-boot:runcd monitoring-dashboard
npm ci
npm run devFor an agent:
mvn -pl monitoring-agent -am package
java -jar monitoring-agent/target/monitoring-agent-0.0.1-SNAPSHOT.jarLocal server defaults expect database monitoring, user monitoring, and the
password configured by DATABASE_PASSWORD. Copy .env.example and export the
variables in your shell, or use the Compose server to avoid configuration drift.
Flyway applies schema migrations automatically; Hibernate validates rather than
creates the schema.
| Variable | Used by | Default/example | Purpose |
|---|---|---|---|
POSTGRES_DB |
Compose/PostgreSQL | monitoring |
Database name |
POSTGRES_USER |
Compose/PostgreSQL | monitoring |
Database user |
POSTGRES_PASSWORD |
Compose/PostgreSQL | development value | Database password |
POSTGRES_PORT |
Compose | 5432 |
Host PostgreSQL port |
DATABASE_URL |
server | jdbc:postgresql://localhost:5432/monitoring... |
JDBC URL |
DATABASE_USERNAME |
server | monitoring |
JDBC user |
DATABASE_PASSWORD |
server | local development value | JDBC password |
REDIS_HOST |
server | localhost |
Redis hostname |
REDIS_PORT |
server/Compose | 6379 |
Redis/container host port |
SERVER_PORT |
Compose | 8080 |
Host API port |
DASHBOARD_PORT |
Compose | 5173 |
Host dashboard port |
JWT_SECRET |
server | development value | HMAC signing secret; rotate in production |
JWT_EXPIRATION |
server | PT1H |
ISO-8601 token lifetime |
ADMIN_USERNAME |
server | admin |
In-memory operator username |
ADMIN_PASSWORD |
server | development value | In-memory operator password |
MONITORING_CORS_ALLOWED_ORIGIN |
server | http://localhost:5173 |
Allowed REST browser origin |
MONITORING_SERVER_URL |
agent | required | Complete heartbeat endpoint including server UUID |
AGENT_HOSTNAME |
agent | required | Must match the registered hostname |
AGENT_API_KEY |
agent | required | Raw per-agent key returned once at registration |
HEARTBEAT_INTERVAL_SECONDS |
agent/Compose | 5 |
Positive heartbeat delay |
AGENT_1_API_KEY…AGENT_3_API_KEY |
Compose seed/agents | development values | Seeded demo credentials |
VITE_API_BASE_URL |
dashboard build/dev | current origin | API origin |
VITE_WS_URL |
dashboard build/dev | derived /ws URL |
WebSocket endpoint |
PROMETHEUS_PORT |
Compose | 9090 |
Host Prometheus port |
GRAFANA_PORT |
Compose | 3000 |
Host Grafana port |
GRAFANA_ADMIN_USER |
Grafana | admin |
Grafana administrator |
GRAFANA_ADMIN_PASSWORD |
Grafana | development value | Grafana administrator password |
Server-only tuning keys are Spring properties in application.yml:
monitoring.offline-check-interval (30 seconds),
monitoring.offline-threshold (60 seconds), and
monitoring.latest-metrics-cache-ttl (60 seconds). Override them with standard
Spring relaxed-binding environment names if needed.
The health, login, registration, heartbeat, and observability endpoints are
public at the HTTP security layer. Heartbeats authenticate independently with
X-Agent-Key; management endpoints require a bearer JWT.
Register an agent (the raw key is returned only in this response):
curl -X POST http://localhost:8080/api/v1/agents/register \
-H "Content-Type: application/json" \
-d '{"hostname":"server-1","ipAddress":"10.0.0.10","osName":"Linux"}'{"serverId":"<uuid>","hostname":"server-1","agentApiKey":"<raw-key>"}Send a heartbeat:
curl -X POST http://localhost:8080/api/v1/agents/<uuid>/heartbeats \
-H "Content-Type: application/json" \
-H "X-Agent-Key: <raw-key>" \
-d '{"hostname":"server-1","osName":"Linux","cpuUsage":42.5,"memoryUsage":71.2,"diskUsage":55.0,"recordedAt":"2026-07-30T10:00:00Z"}'Obtain a JWT and call management APIs:
curl -X POST http://localhost:8080/api/v1/auth/login \
-H "Content-Type: application/json" \
-d '{"username":"admin","password":"<ADMIN_PASSWORD>"}'
curl http://localhost:8080/api/v1/servers?page=0\&size=20 \
-H "Authorization: Bearer <token>"
curl "http://localhost:8080/api/v1/alerts?status=ACTIVE&page=0&size=20" \
-H "Authorization: Bearer <token>"Other authenticated endpoints include:
GET /api/v1/servers/{id}andDELETE /api/v1/servers/{id}(soft deregistration)GET /api/v1/servers/{id}/metrics/latestGET /api/v1/servers/{id}/metrics/history?from=...&to=...&page=0&size=20GET /api/v1/alerts/{id}andPATCH /api/v1/alerts/{id}/resolveGET /api/v1/servers/{id}/alertsGET /api/v1/dashboard/summary
Page sizes must be 1–100. Timestamps are UTC ISO-8601 instants and range bounds
are inclusive. OpenAPI UI is available at /swagger-ui/index.html but is
JWT-protected; API clients should use the endpoint list above.
Registration stores only a BCrypt hash of the generated key. Save the raw key
immediately, then set MONITORING_SERVER_URL, AGENT_HOSTNAME, and
AGENT_API_KEY. The URL is the complete
/api/v1/agents/{serverId}/heartbeats address. The agent uses OSHI, sends once
immediately, then uses a fixed delay. Connection and request timeouts prevent a
stalled server from blocking collection indefinitely. Failures are logged and
retried on the next interval; they do not terminate the agent.
The hostname in a heartbeat must match the registered identity, and a key for
one server cannot submit for another. Deregistered servers reject new
heartbeats. The agent-reported recordedAt represents measurement time, while
server receipt time controls liveness.
All three project images use multi-stage builds and unprivileged runtime users. The server and agent run on Eclipse Temurin 21 JRE Alpine; the dashboard is served by unprivileged Nginx. Build individually with:
docker build -f monitoring-server/Dockerfile -t monitoring-server .
docker build -f monitoring-agent/Dockerfile -t monitoring-agent .
docker build -f monitoring-dashboard/Dockerfile -t monitoring-dashboard .Compose adds health checks, startup dependencies, a private bridge network, and persistent volumes for PostgreSQL, Redis, Prometheus, and Grafana. It publishes ports for local development; production deployments should restrict exposure, use a TLS reverse proxy, and inject secrets from a secret manager.
Run all backend and agent tests:
mvn clean testIntegration tests use Testcontainers and therefore require a running Docker daemon. Run frontend tests and the production build:
cd monitoring-dashboard
npm ci
npm test
npm run buildBuild Java artifacts without rerunning tests:
mvn package -DskipTestsGitHub Actions runs independent jobs for Java tests, frontend tests, Java and frontend builds, and a matrix build of all three Docker images. Images are validated but not published. Maven, npm, and BuildKit layers are cached; job permissions are read-only and concurrent superseded runs are cancelled.
Latest metrics use keys latest:server:<uuid> with a 60-second default TTL.
Every accepted heartbeat is still persisted to PostgreSQL. A cache hit returns
quickly; a miss queries PostgreSQL and repopulates Redis. Invalid cached JSON or
Redis connection failures are logged and treated as misses, so Redis downtime
degrades performance rather than availability or durability. Historical metrics
always come from PostgreSQL.
Every 30 seconds the server checks active registrations. A server is stale when
its last server-received heartbeat—or registration time before its first
heartbeat—is older than the 60-second threshold. It is marked DOWN, one
offline alert is created, and a status event is emitted. A later authenticated
heartbeat marks it UP and resolves the offline alert. Partial unique indexes
and service logic prevent duplicate active alerts.
The raw WebSocket/STOMP endpoint is /ws. Send the login token as a STOMP
CONNECT native header:
Authorization: Bearer <token>
Authenticated clients subscribe to /topic/heartbeats, /topic/alerts, and
/topic/server-status. The browser client reconnects after five seconds and
uses STOMP heartbeats. REST state remains authoritative: after a reconnect,
refresh the relevant API data because the simple in-memory broker does not
replay events. Events are published after successful transaction commit.
POST /api/v1/auth/login authenticates the configured in-memory admin and
returns a signed bearer token plus expiry. REST requests send
Authorization: Bearer <token>; WebSockets use the same value in STOMP
CONNECT. The server is stateless, validates signature/expiry, and returns JSON
401 responses. Agent keys are separate credentials and must never be used as
JWTs. Current authentication is suitable for one configured administrator;
production multi-user deployments need an external identity provider, secret
rotation, TLS, and authorization policy.
MetricAlertRule is the extension point. Spring injects all rule components and
evaluates each accepted heartbeat:
| Rule | Opens when | Severity | Resolves when |
|---|---|---|---|
| High CPU | CPU > 90% | WARNING | CPU <= 90% |
| High memory | memory > 90% | WARNING | memory <= 90% |
| Disk almost full | disk > 95% | CRITICAL | disk <= 95% |
| Server offline | receipt/registration age > 60s | CRITICAL | next valid heartbeat |
Threshold comparison is intentionally strict: exactly 90% CPU or memory and
exactly 95% disk are healthy. Repeated unhealthy samples reuse the existing
active alert rather than creating duplicates. Add a new metric strategy by
implementing MetricAlertRule, defining its domain/migration values, and adding
boundary and lifecycle tests.
Prometheus scrapes monitoring-server:8080/actuator/prometheus using
infrastructure/prometheus/prometheus.yml. Metrics include processed
heartbeats, validation failures, processing duration, alerts tagged by type and
severity, and gauges for active/up and down servers. Grafana is automatically
provisioned with the Prometheus datasource and the repository dashboard at
infrastructure/grafana/dashboards/monitoring-platform.json.
For direct inspection:
curl http://localhost:8080/actuator/health
curl http://localhost:8080/actuator/prometheusDo not expose actuator metrics publicly on an internet-facing deployment; protect them at the network or reverse-proxy layer.
- A service is unhealthy: run
docker compose psanddocker compose logs <service>. PostgreSQL and Redis must become healthy before the server. - Server cannot connect to PostgreSQL: check the database name/user/password
match and remember that containers use hostname
postgres, while host-run processes uselocalhost. - Testcontainers tests fail: start Docker and verify
docker infosucceeds. - 401 from management API: obtain a fresh JWT, preserve the
Bearerprefix, and confirmJWT_SECRETdid not change after token issuance. - 401 from heartbeat API: use the raw key for that UUID and confirm the
hostname matches registration. Seeded raw keys are from
.env, not the database hash. - Dashboard loads but API calls fail: verify the dashboard URL, API health, configured CORS origin, and Vite URL variables. Compose uses Nginx same-origin proxying.
- No realtime updates: inspect the browser WebSocket connection, token
expiry,
/wsproxy upgrade headers, and STOMPCONNECTAuthorization header. - Redis is down: latest reads should fall back to PostgreSQL. Restore Redis for performance; do not treat cache loss as data loss.
- No Prometheus target: check
http://localhost:9090/targets, server health, and the shared Compose network. - Grafana has no data: confirm Prometheus is healthy and selected time range includes traffic; generate heartbeats first.
- Port already in use: change the corresponding host port in
.env. - Schema validation/migration error: inspect Flyway logs and use a fresh development volume only if data deletion is acceptable.
- Why PostgreSQL is authoritative while Redis is a disposable cache, including consistency, fallback, TTL, and failure modes.
- Per-agent credentials versus operator JWTs, BCrypt hashing, stateless REST, WebSocket authentication, expiry, and production identity improvements.
- Fixed-delay agents, server receipt time for liveness, idempotent alert lifecycles, strict threshold boundaries, and avoiding duplicate incidents.
- Transaction-after-commit events, reconnect gaps, and when to replace the simple STOMP broker with a durable external broker.
- Flyway migrations, soft deregistration, UTC timestamps, indexing, partial unique constraints, pagination, and query growth.
- Test pyramid: pure rule/agent tests, MockMvc tests, Testcontainers integration tests, frontend component tests, production builds, and Docker build gates.
- Operational signals, cardinality-safe metric tags, scrape security, dashboards, health checks, and alerting on the monitoring system itself.
- Scaling constraints: scheduled scans, in-memory WebSocket broker, a single configured admin, cache stampedes, retention/partitioning, and horizontal server coordination.
- Security hardening: TLS, secret management/rotation, restricted actuator and database ports, image scanning/signing, rate limits, audit logs, and RBAC.
The complete evidence-based Milestone 14 audit is in
docs/definition-of-done.md. It records
PASS, FAIL, or NOT TESTED for every reviewed requirement and intentionally
does not turn missing runtime validation into a pass.