Repository navigation
feat: structured logs, Prometheus /metrics, tracing, rich /health (#21) - #49
Conversation
Expose Prometheus metrics, request/trace ids, and LLM health so operators can see when raglogs itself degrades. Closes #21. Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
Review — PR #49 (G12 self-observability)Checked Acceptance that holds: Must-fix(None) Should-fix
Nice-to-have
VerdictNeeds changes (should-fix remain) |
Classify GET ingestions as query, mark failed ingest spans as errors, emit only W3C-hex traceparent, and leave queue depth stale when the DB scrape fails. Closes #21. Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
Review — PR #49 (G12 self-observability, round 2)Round-1 should-fix items are addressed. Touched-file ruff is clean. Verified
Must-fix(None) Should-fix(None) Nice-to-have
VerdictReady to merge (0 must-fix, 0 should-fix) |
Closes #21
Self-observability so consumers can see when raglogs itself is slow or degraded.
request_id+scope(contextvars);X-Request-Idechoed (401s included)GET /metrics(Prometheus text, auth-exempt) for ingest/query/LLM/breaker/queuetraceparent+X-Trace-Idheaders (G7 body unchanged)/healthaddsllm {provider, status}; existing ok/degraded rules keptOTEL_EXPORTER_OTLP_ENDPOINT;OTEL_SDK_DISABLED=trueskips SDKMetrics
raglogs_ingest_duration_secondsraglogs_ingest_lines_total{result}raglogs_ingest_request_duration_secondsraglogs_cluster_countraglogs_query_request_duration_seconds{endpoint}raglogs_llm_request_duration_secondsraglogs_llm_fallback_totalraglogs_llm_estimated_tokens_totalraglogs_llm_breaker_stateraglogs_worker_queue_depth