Skip to content

feat: per-scope data retention and scheduled purge (#22) - #50

Merged
cursor[bot] merged 2 commits into
mainfrom
cursor/data-retention-1cc6
Aug 17, 2026
Merged

cursor[bot] merged 2 commits into
mainfrom
cursor/data-retention-1cc6

Conversation

@leo-aa88

Copy link
Copy Markdown
Member

Closes #22

Logs and vectors grew unbounded. This adds configurable TTLs for raw rows vs cluster summaries/embeddings, a worker purge job, and reclaim metrics.

  • Env defaults: RETENTION_RAW=30d, RETENTION_SUMMARY=180d (0 / empty / off = never purge that tier)
  • Per-scope overrides in scope_retention; missing row → env default
  • Raw purge deletes log_entries (log_embeddings + cluster_members CASCADE) and leaves cluster_embeddings so /v1/query/similar still works
  • Summary purge then removes cluster_embeddings, cluster_runs, and explanations for that scope
  • Worker enqueues purge on idle poll every PURGE_INTERVAL_SECONDS (SKIP LOCKED; one chunk per scope so ingest is not starved)
  • CLI: raglogs purge [--scope] [--dry-run]
  • Metric: raglogs_purge_rows_total{kind="raw"|"summary"|"embedding"}
Open in Web Open in Cursor 

Keep cluster summaries after raw logs expire so similar-incident
search still works. Closes #22.

Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Review — PR #50 (G13 data retention)

Unit tests: 774 passed. Ruff on PR-touched files: clean. New migration 0010 only. Raw delete SQL is scope-filtered and does not touch cluster_embeddings; CASCADE FKs are on log_embeddings and cluster_members. 0 / empty / off skip the tier.

Must-fix

  1. src/cli/commands/purge.py:37-41 + src/core/retention/purge.py:417-428 — raglogs purge --dry-run hangs if any rows are expired. CLI passes max_chunks=None (drain until a chunk is empty). purge_raw_chunk / purge_summary_chunk still SELECT the same oldest LIMIT ids when dry_run=True and never delete them, so the batch is never empty and the while True loop never exits. Counts also inflate without bound. Worker is unaffected (max_chunks=1). Fix: on dry-run, use COUNT (or break after one pass / page with a keyset). Add a unit test that run_purge(..., dry_run=True, max_chunks=None) returns once when expired ids exist.

Should-fix

  1. src/core/retention/purge.py:429-430 + src/worker/runner.py:231-234 — a partial purge stamps last_purge_at, so remaining expired rows wait a full PURGE_INTERVAL_SECONDS. If a chunk is full, re-enqueue (or skip writing last_purge_at) so idle workers keep draining while ingest still wins FIFO.

  2. src/core/retention/purge.py:397-431 — one worker job / one transaction covers every scope. Commit per scope (or per chunk) so ingest writes are not blocked for the whole run.

  3. migrations/versions/0010_retention.py:65-73 — cluster_runs.scope and explanations.scope backfill as 'default' with no UPDATE from member log_entries. Pre-migration rows from other G8 scopes become default. Backfill from cluster_members → log_entries.scope.

  4. src/core/retention/policy.py:36-50 + src/core/retention/purge.py:414-415 — invalid interval raises and aborts the rest of the job. Catch per scope (or validate at settings load) and skip/log the bad policy.

  5. src/core/retention/purge.py:43-44 — raw age is COALESCE(timestamp, created_at). Ingesting a 45-day-old incident dump with RETENTION_RAW=30d makes those rows eligible on the next purge. Prefer created_at (or GREATEST(created_at, timestamp)) so TTL is time-in-store.

Nice-to-have

  1. Raw purge metrics omit CASCADE’d cluster_members.
  2. No compiled-SQL assertion that raw DELETE omits cluster_runs / clusters.
  3. No CLI/API to upsert scope_retention.

Verdict

Needs changes (must-fix and should-fix remain)

Dry-run now COUNTs expired rows instead of looping the same LIMIT
ids. Partial chunks skip last_purge_at and re-enqueue; commits are
per scope chunk; invalid intervals skip that scope; raw expiry uses
created_at. Closes remaining #22 review items.

Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Review — PR #50 (G13 data retention, round 2)

Unit tests: 778 passed. Ruff on PR-touched files: clean. Round-1 must-fix and all five should-fixes are in 55b2e2e.

Verified:

  • Dry-run hang: --dry-run COUNTs once per scope; no LIMIT loop.
  • last_purge_at on partial: a full chunk skips the stamp and enqueues a follow-up purge.
  • Commit per scope: db.commit() after each chunk.
  • 0010 backfill: cluster_runs from cluster_members → log_entries.scope; explanations best-effort from the window.
  • Invalid interval: skip that scope and continue.
  • Raw age: log_entries.created_at (time-in-store).

Must-fix

(None)

Should-fix

(None)

Nice-to-have

  1. log_entries has no (scope, created_at) index; large scopes may seq-scan each raw chunk.
  2. raglogs purge --scope X still writes the global last_purge_at, delaying other scopes by one interval.

Verdict

Ready to merge (0 must-fix, 0 should-fix)

@cursor
cursor Bot merged commit 6157116 into main Aug 17, 2026
2 checks passed
@leo-aa88
leo-aa88 deleted the cursor/data-retention-1cc6 branch August 17, 2026 19:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(G13): data retention & lifecycle (per-scope TTL, tiering, purge job)

2 participants