Skip to content

Unconditional panic() on storage-init failure at startup causes crash-loop instead of graceful degrade/retry #1793

Description

@jagatsingh

Describe the bug

main.initializeStorage() (main.go:105) calls panic(err) unconditionally when
store.Initialize() fails at startup, for any underlying error — including a transient
network/DNS failure to the configured storage.type: postgres backend. Combined with a
restart: unless-stopped/always-style supervisor (Docker, systemd, k8s), this turns a
temporary storage-connectivity blip into a tight crash-loop that never lets Gatus reach the
point where it would actually run and report on the endpoints it's configured to monitor.

This is the same crash signature as #1633 (panic: EOF at main.initializeStorage/
main.go:105, fixed for that specific case by bumping lib/pq in #1644) — but the underlying
issue is broader than any one driver bug: any transient failure to reach the storage backend at
startup is fatal, by design (confirmed in #1395: "Gatus is built with fail fast in mind"). This
report is about a different trigger that isn't fixed by a driver bump.

What do you see?

Setup: storage.type: postgres, storage.path pointing at a Postgres/TimescaleDB container by
Docker Compose service name. When that container is removed from the Docker network (not
merely stopped/restarted — a genuinely different condition than a normal "connection refused"),
Docker's embedded DNS resolver (127.0.0.11:53) returns a malformed/unusual response rather than
NXDOMAIN, which the Go resolver surfaces as server misbehaving:

panic: dial tcp: lookup timescaledb-v2 on 127.0.0.11:53: server misbehaving

goroutine 1 [running]:
main.initializeStorage(0x1c611bc80000)
	/app/main.go:105 +0x1766
main.main()
	/app/main.go:34 +0xc5

lib/pq propagates this raw dial error out of store.Initialize(), and main.go panics on it
unconditionally. Under restart: unless-stopped, this repeated 206 times over ~19 hours
(224 occurrences of the panic line in the container logs) until the storage container was
restored — Gatus never once reached the point of running its configured checks (including, in
our case, a TCP check on that same host, which would have alerted on exactly this outage if
Gatus had ever gotten past initializeStorage).

Full write-up (self-inflicted circular dependency on our end, but the panic-vs-graceful-degrade
question is the part we think is worth raising upstream):
https://github.com/jagatsingh/trading_infrastructure/blob/main/docs/timescaledb-v2-removal-incident.md

What do you expect to see?

Ideally, store.Initialize() failing at startup should not be an unconditional, un-retried
panic(). A few options that would all help:

  • Retry storage-connect with backoff for some bounded window before giving up, rather than
    panicking on the first attempt.
  • Log the fatal error clearly and os.Exit(1) instead of panic(), so it's visually obvious
    this is a startup-config problem rather than an internal Gatus bug (the current panic/stack
    trace reads like an application crash, not "your storage backend is unreachable").
  • (Bigger ask, understand if out of scope) an option to fall back to an in-memory store with a
    loud warning if the configured persistent store can't be reached at startup, rather than
    refusing to start at all — since a monitoring tool arguably has more value running
    degraded than not running.

Any of the above would have meant Gatus kept running and alerting throughout the incident above
instead of crash-looping silently for ~19 hours.

Steps to reproduce

  1. docker compose with two services: a Postgres/TimescaleDB container (db) and Gatus
    configured with storage.type: postgres, storage.path: postgres://...@db:5432/gatus....
  2. Start the stack; confirm Gatus is running normally.
  3. docker rm -f db (a full remove, not docker stop) while Gatus keeps running.
  4. Wait for Gatus's own restart: unless-stopped policy (or manually restart the Gatus
    container).
  5. Observe panic: dial tcp: lookup db on 127.0.0.11:53: server misbehaving at
    main.initializeStorage/main.go:105, repeating on every restart until db exists again.

Version

v5.36.0

Additional information

  • Confirmed via source read at the v5.36.0 tag that this is not a bug in the per-check TCP
    code path (client.CanCreateNetworkConnection already returns (false, nil) gracefully on any
    dial error, DNS failures included) — the panic is specifically in the storage-initialization
    path at startup.
  • Our fix on our end is to stop pointing Gatus's own storage at a container one of its own
    checks monitors (removes the circular dependency entirely) — filing this because the
    unconditional-panic-on-storage-init-failure behavior itself seems like it could bite anyone
    whose storage backend has any transient startup-time connectivity hiccup, not just our specific
    circular-dependency case.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions