Repository navigation
feat: per-scope data retention and scheduled purge (#22) - #50
Conversation
Keep cluster summaries after raw logs expire so similar-incident search still works. Closes #22. Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
Review — PR #50 (G13 data retention)Unit tests: 774 passed. Ruff on PR-touched files: clean. New migration Must-fix
Should-fix
Nice-to-have
VerdictNeeds changes (must-fix and should-fix remain) |
Dry-run now COUNTs expired rows instead of looping the same LIMIT ids. Partial chunks skip last_purge_at and re-enqueue; commits are per scope chunk; invalid intervals skip that scope; raw expiry uses created_at. Closes remaining #22 review items. Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
Review — PR #50 (G13 data retention, round 2)Unit tests: 778 passed. Ruff on PR-touched files: clean. Round-1 must-fix and all five should-fixes are in Verified:
Must-fix(None) Should-fix(None) Nice-to-have
VerdictReady to merge (0 must-fix, 0 should-fix) |
Closes #22
Logs and vectors grew unbounded. This adds configurable TTLs for raw rows vs cluster summaries/embeddings, a worker purge job, and reclaim metrics.
RETENTION_RAW=30d,RETENTION_SUMMARY=180d(0/ empty /off= never purge that tier)scope_retention; missing row → env defaultlog_entries(log_embeddings+cluster_membersCASCADE) and leavescluster_embeddingsso/v1/query/similarstill workscluster_embeddings,cluster_runs, andexplanationsfor that scopePURGE_INTERVAL_SECONDS(SKIP LOCKED; one chunk per scope so ingest is not starved)raglogs purge [--scope] [--dry-run]raglogs_purge_rows_total{kind="raw"|"summary"|"embedding"}