Purpose: Items worth doing eventually but not needed for Phase 1. Kept here so they don't clutter the spec but aren't lost. The design should not preclude any of these — if a current decision would make one of these harder later, reconsider.
Source: Gap analysis from Suggested-Improvements.md, filtered through practical priorities.
Expose hardcoded values as env vars (TA_PORT, TA_DATA_DIR, TA_MAX_CONCURRENT_TRACES, TA_LOG_LEVEL, etc.). Ship a .env.example. Low effort once the code exists — just replace constants with os.getenv() calls.
Configurable via TA_RETENTION_DAYS (default: 0 = keep forever). Daily cleanup job deletes stopped traces older than the retention period. Cascading deletes via foreign keys. Important once history browsing is implemented.
Client reconnects with ?since=<timestamp>. Server resumes from that point. Prevents duplicate rendering. Already noted in spec section 6.2 as a Phase 2 item.
Document sqlite3 .backup procedure in README. Not a code feature — just operational guidance for users.
Semantic HTML (<table> for hop table), ARIA labels for icon-only buttons, keyboard-navigable actions, color-blind safe palette. Good practice, not MVP-critical.
Optional TA_TARGET_ALLOWLIST env var (glob for hostnames, CIDR for IPs). Default: allow all. Only relevant for shared/public deployments, which the tool isn't designed for today — but may matter if someone runs a public demo instance.
In-memory sliding window, 10 trace starts per minute per client IP. Returns 429. Only matters if the tool is exposed to untrusted users.
GET /metrics endpoint exposing: active traces, total traces, probe results, probe failures, WebSocket connections, HTTP request counts. Design the internals so counters are easy to add later — don't bury state in closures.
Additional columns for future features:
Trace.tags(JSON array),Trace.notes(freeform text)TraceResult.destination_reached(boolean)HopResult.asn(integer, for ASN lookup),HopResult.geo_country(text, for GeoIP)
Add via schema migrations when the features that use them are built.
Switch from standard Python logging to structlog with JSON output. Include trace_id correlation. Makes log aggregation easier in multi-instance setups. For now, standard logging to stdout is fine — Docker captures it.
Document in README why CAP_NET_RAW is needed, that it's scoped to the mtr binary via setcap, and that the Python process runs as non-root. Reassures security-conscious users.
When a breaking change is actually needed: prefix routes with /api/v2/. Keep v1 for 1-2 major versions. Don't version preemptively.
TA_ENABLE_CORS and TA_CORS_ORIGINS env vars for development or split frontend/backend deployments. Not needed while backend serves the frontend.
GitHub Actions: lint (ruff), format check (black), unit tests, Docker build, integration tests, smoke tests. Auto-push to Docker Hub on main. Tag releases.
Define and measure: API P95 <200ms, probe interval accuracy ±10%, DB write <50ms, UI update <500ms. Useful for catching regressions once there's a test suite to measure against.
Define alert-worthy conditions (mtr missing, DB write failures, disk >90%, trace saturation). Expose via /health detail levels. Integrate with Prometheus Alertmanager or similar.
CSS variables for colors, detect prefers-color-scheme: dark, optional manual toggle in localStorage.
Ctrl+N (new trace), Space (pause/resume), Ctrl+E (export), arrow keys for hop navigation. Add <kbd> hints in UI.
Screen reader support, focus management, live region announcements. Lighthouse audit.
black, ruff, mypy (gradual, not strict), pre-commit hooks. Set up when the codebase is large enough to benefit. For Phase 1, just pick a formatter and be consistent.
Trunk-based: main always deployable, feature branches, squash merge, conventional commit messages. Formalize when there are multiple contributors.
These aren't features to build — they're constraints on current design decisions:
-
Configuration as constants, not magic numbers. Use named constants (e.g.,
MAX_CONCURRENT_TRACES = 10) so converting to env vars later is a one-line change per value. -
Counters where you'd want metrics. When tracking active traces or probe counts internally, use clean counter variables — not derived from querying the DB each time. Makes adding a
/metricsendpoint trivial later. -
Logging with context. Use Python's standard
loggingmodule but includetrace_idin log messages where relevant. Switching to structlog later is a logger-config change, not a code rewrite. -
Health endpoint is extensible. Return a dict from
/health. Adding new checks (disk space, probe failure rate) means adding keys, not changing the contract. -
Error response format is the contract. Every error already uses
{error, message, category}. Future features (rate limiting, allowlists) just add new error codes to the same format.