Skip to content

Repository files navigation

Monitoring Platform

A full-stack infrastructure monitoring system built with Java 21, Spring Boot, React, PostgreSQL, Redis, STOMP over WebSockets, Prometheus, Grafana, and Docker Compose. Agents collect host CPU, memory, and disk utilization; the server persists heartbeats, detects failures, evaluates alerts, and pushes live updates to an authenticated dashboard.

Architecture

flowchart LR
    A["Java monitoring agents<br/>OSHI metrics"] -->|"HTTPS/REST<br/>X-Agent-Key"| S["Spring Boot API"]
    U["Operator browser<br/>React dashboard"] -->|"REST + JWT"| S
    S -->|"STOMP events"| U
    S -->|"JPA/Flyway<br/>source of truth"| P[("PostgreSQL")]
    S -->|"latest metrics<br/>TTL cache"| R[("Redis")]
    PR["Prometheus"] -->|"/actuator/prometheus"| S
    G["Grafana"] -->|"PromQL"| PR
Loading

The Maven reactor contains monitoring-common, monitoring-server, and monitoring-agent. The Vite application lives in monitoring-dashboard. PostgreSQL is the durable source of truth. Redis accelerates latest-metric reads but is deliberately non-authoritative. The server emits transaction-aware realtime events only after database commits. In the container deployment, Nginx serves the dashboard and proxies /api and /ws to the server.

Prerequisites

  • Git
  • Docker Desktop or Docker Engine with Compose v2 (recommended startup)
  • For source development: JDK 21, Maven 3.9+, Node.js 22, and npm

Exact clean-machine startup

These commands are the shortest reproducible path on a machine with Git and Docker Compose. Replace the repository URL with the actual clone URL.

git clone <repository-url> monitoring-platform
cd monitoring-platform
cp .env.example .env
docker compose pull
docker compose build
docker compose up -d
docker compose ps

PowerShell equivalent for the copy step:

Copy-Item .env.example .env

Before exposing the stack beyond localhost, edit .env and replace every development password, JWT_SECRET, and agent key. Wait until docker compose ps shows the services healthy, then open:

Grafana uses GRAFANA_ADMIN_USER and GRAFANA_ADMIN_PASSWORD. The Compose startup creates three deterministic development agents through the one-shot agent-credentials-init service; their raw keys come from .env. View startup progress with docker compose logs -f. Stop without deleting data using docker compose down. docker compose down -v permanently removes all named volumes and should only be used for an intentional reset.

Source setup

Start PostgreSQL and Redis, then run the server and dashboard in separate terminals:

docker compose up -d postgres redis
mvn -pl monitoring-server spring-boot:run
cd monitoring-dashboard
npm ci
npm run dev

For an agent:

mvn -pl monitoring-agent -am package
java -jar monitoring-agent/target/monitoring-agent-0.0.1-SNAPSHOT.jar

Local server defaults expect database monitoring, user monitoring, and the password configured by DATABASE_PASSWORD. Copy .env.example and export the variables in your shell, or use the Compose server to avoid configuration drift. Flyway applies schema migrations automatically; Hibernate validates rather than creates the schema.

Environment variables

Variable Used by Default/example Purpose
POSTGRES_DB Compose/PostgreSQL monitoring Database name
POSTGRES_USER Compose/PostgreSQL monitoring Database user
POSTGRES_PASSWORD Compose/PostgreSQL development value Database password
POSTGRES_PORT Compose 5432 Host PostgreSQL port
DATABASE_URL server jdbc:postgresql://localhost:5432/monitoring... JDBC URL
DATABASE_USERNAME server monitoring JDBC user
DATABASE_PASSWORD server local development value JDBC password
REDIS_HOST server localhost Redis hostname
REDIS_PORT server/Compose 6379 Redis/container host port
SERVER_PORT Compose 8080 Host API port
DASHBOARD_PORT Compose 5173 Host dashboard port
JWT_SECRET server development value HMAC signing secret; rotate in production
JWT_EXPIRATION server PT1H ISO-8601 token lifetime
ADMIN_USERNAME server admin In-memory operator username
ADMIN_PASSWORD server development value In-memory operator password
MONITORING_CORS_ALLOWED_ORIGIN server http://localhost:5173 Allowed REST browser origin
MONITORING_SERVER_URL agent required Complete heartbeat endpoint including server UUID
AGENT_HOSTNAME agent required Must match the registered hostname
AGENT_API_KEY agent required Raw per-agent key returned once at registration
HEARTBEAT_INTERVAL_SECONDS agent/Compose 5 Positive heartbeat delay
AGENT_1_API_KEY…AGENT_3_API_KEY Compose seed/agents development values Seeded demo credentials
VITE_API_BASE_URL dashboard build/dev current origin API origin
VITE_WS_URL dashboard build/dev derived /ws URL WebSocket endpoint
PROMETHEUS_PORT Compose 9090 Host Prometheus port
GRAFANA_PORT Compose 3000 Host Grafana port
GRAFANA_ADMIN_USER Grafana admin Grafana administrator
GRAFANA_ADMIN_PASSWORD Grafana development value Grafana administrator password

Server-only tuning keys are Spring properties in application.yml: monitoring.offline-check-interval (30 seconds), monitoring.offline-threshold (60 seconds), and monitoring.latest-metrics-cache-ttl (60 seconds). Override them with standard Spring relaxed-binding environment names if needed.

API examples

The health, login, registration, heartbeat, and observability endpoints are public at the HTTP security layer. Heartbeats authenticate independently with X-Agent-Key; management endpoints require a bearer JWT.

Register an agent (the raw key is returned only in this response):

curl -X POST http://localhost:8080/api/v1/agents/register \
  -H "Content-Type: application/json" \
  -d '{"hostname":"server-1","ipAddress":"10.0.0.10","osName":"Linux"}'
{"serverId":"<uuid>","hostname":"server-1","agentApiKey":"<raw-key>"}

Send a heartbeat:

curl -X POST http://localhost:8080/api/v1/agents/<uuid>/heartbeats \
  -H "Content-Type: application/json" \
  -H "X-Agent-Key: <raw-key>" \
  -d '{"hostname":"server-1","osName":"Linux","cpuUsage":42.5,"memoryUsage":71.2,"diskUsage":55.0,"recordedAt":"2026-07-30T10:00:00Z"}'

Obtain a JWT and call management APIs:

curl -X POST http://localhost:8080/api/v1/auth/login \
  -H "Content-Type: application/json" \
  -d '{"username":"admin","password":"<ADMIN_PASSWORD>"}'

curl http://localhost:8080/api/v1/servers?page=0\&size=20 \
  -H "Authorization: Bearer <token>"

curl "http://localhost:8080/api/v1/alerts?status=ACTIVE&page=0&size=20" \
  -H "Authorization: Bearer <token>"

Other authenticated endpoints include:

  • GET /api/v1/servers/{id} and DELETE /api/v1/servers/{id} (soft deregistration)
  • GET /api/v1/servers/{id}/metrics/latest
  • GET /api/v1/servers/{id}/metrics/history?from=...&to=...&page=0&size=20
  • GET /api/v1/alerts/{id} and PATCH /api/v1/alerts/{id}/resolve
  • GET /api/v1/servers/{id}/alerts
  • GET /api/v1/dashboard/summary

Page sizes must be 1–100. Timestamps are UTC ISO-8601 instants and range bounds are inclusive. OpenAPI UI is available at /swagger-ui/index.html but is JWT-protected; API clients should use the endpoint list above.

Agent configuration and behavior

Registration stores only a BCrypt hash of the generated key. Save the raw key immediately, then set MONITORING_SERVER_URL, AGENT_HOSTNAME, and AGENT_API_KEY. The URL is the complete /api/v1/agents/{serverId}/heartbeats address. The agent uses OSHI, sends once immediately, then uses a fixed delay. Connection and request timeouts prevent a stalled server from blocking collection indefinitely. Failures are logged and retried on the next interval; they do not terminate the agent.

The hostname in a heartbeat must match the registered identity, and a key for one server cannot submit for another. Deregistered servers reject new heartbeats. The agent-reported recordedAt represents measurement time, while server receipt time controls liveness.

Docker

All three project images use multi-stage builds and unprivileged runtime users. The server and agent run on Eclipse Temurin 21 JRE Alpine; the dashboard is served by unprivileged Nginx. Build individually with:

docker build -f monitoring-server/Dockerfile -t monitoring-server .
docker build -f monitoring-agent/Dockerfile -t monitoring-agent .
docker build -f monitoring-dashboard/Dockerfile -t monitoring-dashboard .

Compose adds health checks, startup dependencies, a private bridge network, and persistent volumes for PostgreSQL, Redis, Prometheus, and Grafana. It publishes ports for local development; production deployments should restrict exposure, use a TLS reverse proxy, and inject secrets from a secret manager.

Testing and CI

Run all backend and agent tests:

mvn clean test

Integration tests use Testcontainers and therefore require a running Docker daemon. Run frontend tests and the production build:

cd monitoring-dashboard
npm ci
npm test
npm run build

Build Java artifacts without rerunning tests:

mvn package -DskipTests

GitHub Actions runs independent jobs for Java tests, frontend tests, Java and frontend builds, and a matrix build of all three Docker images. Images are validated but not published. Maven, npm, and BuildKit layers are cached; job permissions are read-only and concurrent superseded runs are cancelled.

Redis behavior

Latest metrics use keys latest:server:<uuid> with a 60-second default TTL. Every accepted heartbeat is still persisted to PostgreSQL. A cache hit returns quickly; a miss queries PostgreSQL and repopulates Redis. Invalid cached JSON or Redis connection failures are logged and treated as misses, so Redis downtime degrades performance rather than availability or durability. Historical metrics always come from PostgreSQL.

Failure detection

Every 30 seconds the server checks active registrations. A server is stale when its last server-received heartbeat—or registration time before its first heartbeat—is older than the 60-second threshold. It is marked DOWN, one offline alert is created, and a status event is emitted. A later authenticated heartbeat marks it UP and resolves the offline alert. Partial unique indexes and service logic prevent duplicate active alerts.

WebSockets

The raw WebSocket/STOMP endpoint is /ws. Send the login token as a STOMP CONNECT native header:

Authorization: Bearer <token>

Authenticated clients subscribe to /topic/heartbeats, /topic/alerts, and /topic/server-status. The browser client reconnects after five seconds and uses STOMP heartbeats. REST state remains authoritative: after a reconnect, refresh the relevant API data because the simple in-memory broker does not replay events. Events are published after successful transaction commit.

JWT security

POST /api/v1/auth/login authenticates the configured in-memory admin and returns a signed bearer token plus expiry. REST requests send Authorization: Bearer <token>; WebSockets use the same value in STOMP CONNECT. The server is stateless, validates signature/expiry, and returns JSON 401 responses. Agent keys are separate credentials and must never be used as JWTs. Current authentication is suitable for one configured administrator; production multi-user deployments need an external identity provider, secret rotation, TLS, and authorization policy.

Alert strategies

MetricAlertRule is the extension point. Spring injects all rule components and evaluates each accepted heartbeat:

Rule Opens when Severity Resolves when
High CPU CPU > 90% WARNING CPU <= 90%
High memory memory > 90% WARNING memory <= 90%
Disk almost full disk > 95% CRITICAL disk <= 95%
Server offline receipt/registration age > 60s CRITICAL next valid heartbeat

Threshold comparison is intentionally strict: exactly 90% CPU or memory and exactly 95% disk are healthy. Repeated unhealthy samples reuse the existing active alert rather than creating duplicates. Add a new metric strategy by implementing MetricAlertRule, defining its domain/migration values, and adding boundary and lifecycle tests.

Prometheus and Grafana

Prometheus scrapes monitoring-server:8080/actuator/prometheus using infrastructure/prometheus/prometheus.yml. Metrics include processed heartbeats, validation failures, processing duration, alerts tagged by type and severity, and gauges for active/up and down servers. Grafana is automatically provisioned with the Prometheus datasource and the repository dashboard at infrastructure/grafana/dashboards/monitoring-platform.json.

For direct inspection:

curl http://localhost:8080/actuator/health
curl http://localhost:8080/actuator/prometheus

Do not expose actuator metrics publicly on an internet-facing deployment; protect them at the network or reverse-proxy layer.

Troubleshooting

  • A service is unhealthy: run docker compose ps and docker compose logs <service>. PostgreSQL and Redis must become healthy before the server.
  • Server cannot connect to PostgreSQL: check the database name/user/password match and remember that containers use hostname postgres, while host-run processes use localhost.
  • Testcontainers tests fail: start Docker and verify docker info succeeds.
  • 401 from management API: obtain a fresh JWT, preserve the Bearer prefix, and confirm JWT_SECRET did not change after token issuance.
  • 401 from heartbeat API: use the raw key for that UUID and confirm the hostname matches registration. Seeded raw keys are from .env, not the database hash.
  • Dashboard loads but API calls fail: verify the dashboard URL, API health, configured CORS origin, and Vite URL variables. Compose uses Nginx same-origin proxying.
  • No realtime updates: inspect the browser WebSocket connection, token expiry, /ws proxy upgrade headers, and STOMP CONNECT Authorization header.
  • Redis is down: latest reads should fall back to PostgreSQL. Restore Redis for performance; do not treat cache loss as data loss.
  • No Prometheus target: check http://localhost:9090/targets, server health, and the shared Compose network.
  • Grafana has no data: confirm Prometheus is healthy and selected time range includes traffic; generate heartbeats first.
  • Port already in use: change the corresponding host port in .env.
  • Schema validation/migration error: inspect Flyway logs and use a fresh development volume only if data deletion is acceptable.

Interview discussion points

  • Why PostgreSQL is authoritative while Redis is a disposable cache, including consistency, fallback, TTL, and failure modes.
  • Per-agent credentials versus operator JWTs, BCrypt hashing, stateless REST, WebSocket authentication, expiry, and production identity improvements.
  • Fixed-delay agents, server receipt time for liveness, idempotent alert lifecycles, strict threshold boundaries, and avoiding duplicate incidents.
  • Transaction-after-commit events, reconnect gaps, and when to replace the simple STOMP broker with a durable external broker.
  • Flyway migrations, soft deregistration, UTC timestamps, indexing, partial unique constraints, pagination, and query growth.
  • Test pyramid: pure rule/agent tests, MockMvc tests, Testcontainers integration tests, frontend component tests, production builds, and Docker build gates.
  • Operational signals, cardinality-safe metric tags, scrape security, dashboards, health checks, and alerting on the monitoring system itself.
  • Scaling constraints: scheduled scans, in-memory WebSocket broker, a single configured admin, cache stampedes, retention/partitioning, and horizontal server coordination.
  • Security hardening: TLS, secret management/rotation, restricted actuator and database ports, image scanning/signing, rate limits, audit logs, and RBAC.

Definition-of-done review

The complete evidence-based Milestone 14 audit is in docs/definition-of-done.md. It records PASS, FAIL, or NOT TESTED for every reviewed requirement and intentionally does not turn missing runtime validation into a pass.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages