Repository navigation
Conversation
Adds backend foundations for queue monitoring: new monitoring thresholds and DTO schemas, Bull throughput helpers, and worker heartbeat publishing to Redis so admin views can report live worker state across pods. Queue creation now enables Bull per-minute metrics, tracks added jobs with TTL-based Redis counters, and exposes a configurable history window (`QUEUE_METRICS_MAX_DATA_POINTS`, default 1440). QueueModule now starts/stops heartbeat publishing and registers tracked queues with their concurrency. Includes unit coverage to ensure queue metrics are enabled and default retention is one day.
barrfalk
marked this pull request as draft
October 1, 2026 22:19
Adds admin queue monitoring for recipient-level backlog and service health, plus live SSE change signals when queue state moves. The backend now aggregates per-channel message throughput, tracks stalled active jobs, exposes active merge-batch progress from Bull jobs, and includes a V72 index to support the hourly stats query. The frontend dashboard shows message backlog, active batches, and drill-down metrics, with matching specs covering the new monitoring views and progress reporting. No public API contract changes; this is admin-only monitoring.
Adds an overview layer to queue monitoring with last-hour totals, request counts, and delivery-time percentiles from backend data, and updates the admin UI to use tabbed Overview/Failures/System views with summary cards and URL-persisted tab state. It also fixes load-test API key autobind by introducing a dedicated guard so the route works with the global JwtGuard while still enforcing the feature flag and gateway-header checks. Finally, Coraza WAF path rules were adjusted to allow SPA /admin routes, with a route-vs-WAF regression test to prevent refresh/direct-link 403s.
Harden `autoBindApiKeyForLoadTest` so it no longer rebinds an existing credential from another tenant to the load-test tenant. The service now logs and throws a `ConflictException` when a key is already owned by a different tenant, while still allowing idempotent reuse of keys already bound to the load-test tenant. Added focused unit tests covering new bind, idempotent behavior, cross-tenant rejection, and environment guardrails in test/prod namespaces.
Update worker heartbeat tracking to use per-job IDs instead of a simple counter so local stalled-job `failed` events for jobs from other pods do not corrupt active/failed metrics. Added backend tests for unknown finished jobs and cross-pod stalled-job scenarios. In Queue Monitoring, hide the manual Refresh button while the live stream is connected (it already auto-refetches) and show it again when the stream drops. Added a frontend test to cover this visibility toggle.
Replaces the monitoring overview’s `sentPerMinute` with `sendingRate` (`perMinute` + `activeMinutes`) on both backend and frontend DTOs/interfaces. The new backend calculation averages only minutes with sends and ignores the current partial minute unless it is the only active minute, preventing idle time from diluting throughput; UI copy and specs were updated to show the new metric and cover null/aggregation behavior.
Improve the Redis stats UI to avoid misleading fragmentation ratios on small instances. Fragmentation is now hidden with a “Not meaningful at low memory use” hint below 100MB used memory, and high fragmentation (>1.5) is explicitly flagged with guidance. Added/updated QueueMonitoring tests to cover normal, low-memory, and high-fragmentation scenarios.
Add graceful queue shutdown so workers mark themselves as draining, finish in-flight jobs within the pod grace period, and then remove their heartbeat instead of appearing stalled. Expose the new draining state through backend monitoring and the admin UI. Also make retried email merge batches skip recipients already sent by an earlier attempt to avoid duplicate delivery during pod replacement or shutdown.
Remove unnecessary `forwardRef()` wrappers from `GcNotifyModule.forRoot()` when importing `NotifyModule` and `QueueModule`, preventing duplicate static module instantiation. Add an e2e guard test that scans the Nest `ModulesContainer` (excluding `TypeOrmModule` dynamic `forFeature` entries) to ensure non-TypeORM modules are only created once, protecting against double worker/heartbeat startup and related false stalled-job monitoring.
Bull's default maxStalledCount of 1 fails a job the second time its pod dies mid-run, e.g. a pod deleted and then restarted during one bulk send. The two merge batches in PR-253 ended in the failed set that way, leaving 80 recipients pending. Retries are safe now that merge batches skip recipients already sent. Overridable with QUEUE_MAX_STALLED_COUNT. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Refactors mail-merge queue payload handling so jobs carry shared content/params only, while per-recipient params are persisted on `notification_request_detail` (with migration `V73`). Adds merge reconstruction helpers and updates pending-notification retry to rebuild proper merge ingestion payloads and enqueue unnamed jobs so workers can process them. Introduces `BatchReconcilerService` to detect stale merge batches, retry/requeue missing or failed batch jobs from stored request data, and fail remaining recipients after a capped number of recovery attempts to prevent batches from staying in-flight indefinitely. Includes new unit coverage for merge batch builders, batch reconciliation behavior, and pending retry merge rebuilding.
Introduce a shutdown diagnostics helper that logs synchronous stderr markers when SIGTERM is received and when the process exits with a non-zero code. It now records the last app error from `StructuredLoggerService` so failed shutdowns include likely cause/context, wires diagnostics into app bootstrap, and adds unit tests covering SIGTERM, non-zero exit reporting, and clean exits.
Moves the Node heap limit out of the backend image and into Helm so environments can tune it without rebuilding. The deployment now sets NODE_OPTIONS from `backend.heapMb` (default 512), and backend memory requests are templated via `backend.memoryRequest` (default 200Mi). TEST and PROD deploy params in `merge.yml` now override memory request to 384Mi while keeping autoscaling settings unchanged.
Replace the separate pending-notification and batch recovery services with a single DeliveryReconciler that restores stuck pending, ingestion, batch, and plain-delivery work from Postgres. It rebuilds scheduled sends correctly, avoids duplicating detail rows or re-sending finished channels on ingestion reruns, and adds CHES request deadlines so hung token or email calls fail fast with clear 504 errors.
Adds Redis-backed reconcile activity tracking (last pass, action counts, recent actions) and exposes it through the admin monitoring API with new reconciler DTOs and status evaluation. Updates the admin Monitoring UI to show a Delivery recovery panel, include reconciler warnings in the top banner, and display recent recoveries. Also extends delivery reconciliation to handle overdue scheduled sends and adds comprehensive backend/frontend test coverage for the new behavior.
Add a Redis-backed concurrency limiter and apply it to CHES email POST requests so in-flight CHES calls are capped across pods. The limiter uses expiring leases, releases slots on success/failure, and fails open if Redis is unavailable. Update CHES config defaults by increasing `CHES_TIMEOUT_MS` to 120000 and adding `CHES_MAX_CONCURRENT_REQUESTS` (default 5, `0` disables limiting). Extend tests to cover limiter behavior and verify CHES adapter uses the shared limit only when enabled.
Introduces shared transient delivery error handling and a Redis-backed circuit breaker, then applies it to CHES and SMS so provider outages pause sends instead of failing recipients immediately. SMS adapter wiring now always runs through a breaker, provider/HTTP/network failures are classified consistently, and delivery workers keep outage-affected recipients owed (including merge batches) while only failing unsent rows when retries are exhausted for non-transient errors. Monitoring was expanded to show provider circuit state and sends-in-progress rollups (replacing active batch rows), with backend/frontend DTO, service, UI, and test updates to match.
barrfalk
marked this pull request as ready for review
October 6, 2026 16:47
This change renames the shutdown timeout env var to QUEUE_SHUTDOWN_DRAIN_MS and tightens the graceful shutdown sequence. Queue workers are paused and closed in stages (ingestion → delivery → webhooks) with a bounded drain window before app shutdown, while timed-out jobs are logged as re-run elsewhere. This avoids partial merge batches being left mid-flight during deploys or scale-downs.
Removes long TEST and PROD `params` overrides from `merge.yml` so rollout strategy, autoscaling, PDB behavior, and memory sizing come from `values-test.yaml` and `values-prod.yaml` directly. Both env value files now explicitly set `backend.memoryRequest` to `384Mi` and update quota/headroom comments to match, keeping deployment shape and resource planning reviewable in git.
Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>
revanth-banala
self-requested a review
October 6, 2026 19:33
revanth-banala
approved these changes
Oct 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
CCP-6026. This PR adds an admin Monitoring page for queues, workers and Redis, and makes bulk delivery recover from crashes, restarts and lost jobs without losing or duplicating messages.
Monitoring page (
/admin/monitoring, NOTIFY_ADMIN SSO role only)GET /api/v1/frontend/admin/monitoring/queues/events), driven by Bull's Redis pub/sub and throttled to one refresh every 2s. Falls back to polling every 15s if the stream drops. Warnings from every tab are shown in a banner at the top.GET /api/v1/frontend/admin/monitoring/queues(NotifyAdminGuard), plus two new Kong routes inroutes.yaml. Every read goes to Redis or Postgres, so any pod gives the same answer.QUEUE_METRICS_MAX_DATA_POINTS), and a per-minute "added" counter supplies the fill rate.Delivery robustness
notification_request_detail.paramscolumn (V73).BATCH_SIZE).QUEUE_SHUTDOWN_DRAIN_MS(105s, from the chart).terminationGracePeriodSecondsis 120.draining.maxStalledCountis 3 (QUEUE_MAX_STALLED_COUNT), so a pod restart no longer fails a long batch outright.DeliveryReconcilerServicereplacesPendingNotificationRetryService.One single-flight pass every 60s, behind a Redis lock. It compares what Postgres says is owed with what the queues hold.
A live job is left alone; a failed job is retried; a missing job is rebuilt. After 5 attempts, whatever is still owed is marked failed and the request is settled.
Five kinds of stuck work:
pendingingestionscheduledbatchdeliveryEach pass is recorded in Redis for the Monitoring page.
CHES_TIMEOUT_MS(120s) instead of hanging a worker slot.TransientDeliveryError. The email worker then pauses the batch, leaving unsent recipients owed for Bull's retry and then the reconciler, rather than marking each one failed. SMS works the same way: ACS and Twilio SDK errors are classified, and recipients a provider could not take are flaggedtransientand left owed. Previously an 11-second CHES restart failed about 1,500 recipients in a 2,000-recipient test. An email timeout still fails just that recipient, because CHES may have queued it and a resend could duplicate it.CHES_MAX_CONCURRENT_REQUESTS, 5) byRedisConcurrencyLimiter. CHES accepts about one email a second however many requests arrive together. Without a cap, a large merge piles requests up inside CHES until some pass the timeout: with a 30s timeout, a 400-recipient test failed one recipient this way. The cap moves that wait to our side, where nothing times out, and sends just as many emails a minute. If Redis is unreachable, sends go ahead without the cap.Fixes found along the way
GcNotifyModule.forRootimportedQueueModule/NotifyModulethroughforwardRef, which created a second instance. The e2e spec now asserts each module is created once./admin/*page returned 403. WAF rule 1004 infrontend/coraza.confno longer blocks/admin. That URL is the SPA's admin area; access is enforced by the API.routes/waf-routes.spec.tschecks every route against rules 1004 and 1007.LoadtestAutobindGuard, so it works only in load-test mode.'process'jobs that no handler picked up. It also sent scheduled sends immediately. Both are fixed by the reconciler.--max-old-space-size=150. The chart setsNODE_OPTIONSfrombackend.heapMb(512). The memory request is 200Mi by default and 384Mi in TEST/PROD (merge.yml).installShutdownDiagnosticslogs the signal, the exit code and the last error on any nonzero exit.Type of change
Rollout notes
notification_request_detail(last_attempt_at)for sent/failed rows, used by the monitoring counts;params jsonbcolumn onnotification_request_detail.app-api-secretsin f6bc3f-dev, -test and -prod. Nothing to do before merge.BATCH_SIZE=25,QUEUE_MAX_STALLED_COUNT=3,QUEUE_METRICS_MAX_DATA_POINTS=1440,CHES_TIMEOUT_MS=120000,CHES_MAX_CONCURRENT_REQUESTS=5DELIVERY_RECONCILE_INTERVAL_MS=60000,DELIVERY_RECONCILE_STALE_MS=600000,DELIVERY_RECONCILE_MAX_ATTEMPTS=5,DELIVERY_RECONCILE_PENDING_STALE_MS=30000MONITORING_*thresholdsBATCH_RECONCILE_*keys have been removed.NODE_OPTIONS,QUEUE_SHUTDOWN_DRAIN_MS,terminationGracePeriodSeconds: 120,backend.memoryRequest.api-gateway/templates/routes.yamlgo out with the usual gateway publish.How Has This Been Tested?
Automated checks (all green locally):
npm run test:unit:cov: 1946 passed. This includestest/app.e2e-spec.tsagainst local Postgres and Redis.npm run lintandnpm run buildpass.npx vitest run: 665 passed. Frontendnpm run lintandnpm run buildpass.Against real Postgres and Redis:
In the PR environment (OpenShift):
Checklist
Further comments
Why recipients moved out of job payloads. A merge job used to carry its full recipient list in Redis, so a large send was held twice: in Redis and in Postgres. Retries also had no record of who had already been sent to. With Postgres as the record of what is owed, every recovery path (Bull retry, stalled job, reconciler) can be repeated safely. Delivery is at-least-once, and workers are idempotent.
Known limits, not addressed here:
BATCH_SIZEor lengthening the grace period in TEST/PROD would close the gap.cancelOrRescheduleNotification) updatesdelayed_send_timebut doesn't move the delayed Bull job. The send still goes out at the original time. This predates this PR; the new scheduled finder neither causes nor fixes it.AGENTS.mdandCLAUDE.mdare git-ignored, so their updates (the env-var rule and the reconciler notes) aren't in this diff.Thanks for the PR!
Deployments, as required, will be available below:
Please create PRs in draft mode. Mark as ready to enable:
After merge, new images are deployed in:
🤖 Generated with Claude Code
Thanks for the PR!
Deployments, as required, will be available below:
Please create PRs in draft mode. Mark as ready to enable:
After merge, new images are deployed in: