Skip to content

feat: API rate limiting, ingest backpressure, LLM concurrency (#18) - #46

Merged
cursor[bot] merged 2 commits into
mainfrom
cursor/rate-limit-backpressure-1cc6
Aug 17, 2026
Merged

cursor[bot] merged 2 commits into
mainfrom
cursor/rate-limit-backpressure-1cc6

Conversation

@leo-aa88

Copy link
Copy Markdown
Member

Closes #18

Summary

Token-bucket limits on ingest/query, the existing ingest queue ceiling, and a process-wide LLM semaphore keep a large incident dump from unbounded API and provider fan-out.

Behavior

  • API token bucket (src/api/ratelimit.py) on POST /v1/ingestions* (writes) and /v1/query* plus unversioned aliases. Identity is request.state.auth_principal.key_id, or "anonymous" when auth is off. 429 + Retry-After with {"error_code":"RATE_LIMITED",...}. /health, /docs, static UI, and /config are not limited. In-memory per process.
  • Ingest queue unchanged: INGEST_QUEUE_MAX still yields 429 INGEST_QUEUE_FULL; tail ticks still skip when full.
  • LLM cap in src/core/llm/provider.py: semaphore around generate_summary (and ask HTTP). Noop skips the wait but still uses the same entrypoint.

Config

Via get_settings() (no RAGLOGS_ prefix): RATELIMIT_*, INGEST_QUEUE_MAX, LLM_MAX_CONCURRENCY. 0 rps / concurrency = unlimited. Defaults are high (100 rps/burst, LLM concurrency 4) so existing TestClient tests do not 429.

Tests

Allow-then-429, Retry-After, per-key isolation, anonymous defaults, queue-full still INGEST_QUEUE_FULL, LLM max_concurrency=1 serialization, noop not blocking, synthetic burst that 429s. python -m pytest tests/unit/ passed locally (668).

Open in Web Open in Cursor 

Unbounded ingest dumps and explain fan-out can overload the process and
the LLM provider. Token-bucket limits on ingest/query, the existing
ingest queue ceiling, and a process-wide LLM semaphore keep G9
backpressure in place without changing queue semantics.

Closes #18

Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Review — PR #46 (G9 rate limiting)

Checked origin/main...origin/cursor/rate-limit-backpressure-1cc6 against issue #18. python3 -m pytest tests/unit/ → 668 passed. Ruff on the touched files is clean.

Token-bucket middleware sits inside auth, covers versioned and unversioned ingest writes and /query*, returns 429 RATE_LIMITED + Retry-After, and leaves INGEST_QUEUE_FULL alone. LLM calls go through CappedLLMProvider.generate_summary; ask HTTP uses the same llm_concurrency_slot. Noop skips the wait. Config is unprefixed RATELIMIT_* / INGEST_QUEUE_MAX / LLM_MAX_CONCURRENCY. Defaults (100 rps/burst) did not 429 the existing TestClient suite.

Must-fix

(None)

Should-fix

  1. tests/unit/test_ratelimit.py:187-196 — test_synthetic_load_exhausts_tiny_bucket posts {"since": "1h"} (enters explain_endpoint / get_db) with ratelimit_query_rps=1.0 and asserts statuses.count(429) >= 8. Denied requests still refill; if those eight calls take >~1s (Postgres up, slow connect), extra tokens land and the assertion fails. Use a fast 400 body (json={}) or mock get_db.

  2. src/core/retrieval/question_router.py:421 / tests/unit/test_llm_concurrency.py — _call_llm_ask is the bypass around generate_summary (OpenAI/Ollama HTTP). Concurrency tests only exercise CappedLLMProvider.generate_summary. Dropping the with llm_concurrency_slot in ask would not fail CI and would let POST /query/ask fan out unbounded provider calls. Add a test that _call_llm_ask / the ask HTTP path actually takes the semaphore.

Nice-to-have

  1. src/api/ratelimit.py:30-31 — _buckets never evicts. Fine for a small key set; a long-lived process with many key_id / oidc: identities will grow without bound.

  2. src/api/ratelimit.py:9 — docstring mentions “key-admin routes”; there are no /admin routes.

  3. src/api/ratelimit.py:120-128 — Retry-After is RATELIMIT_RETRY_AFTER_SECONDS, not time-to-next-token.

  4. src/api/ratelimit.py:84-87 — POST /ingestions/{id}:pause|resume|stop share the ingest write bucket. Lifecycle is not a dump path; exempting it would avoid operator 429s during an ingest burst.

  5. No suite-level reset_rate_limiter() in conftest.py. Files other than test_ratelimit.py share the anonymous bucket.

Verdict

Needs changes (should-fix remain)

The synthetic load case posted a valid explain body that hit get_db;
slow connects could refill the 1 rps bucket and flake the 429 count.
Ask HTTP concurrency was untested, so dropping the semaphore would
not fail CI.

Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Review — PR #46 (G9 rate limiting, round 2)

Checked origin/main...origin/cursor/rate-limit-backpressure-1cc6 (commits 6705a7c + f7320ff). python3 -m pytest tests/unit/ → 669 passed. Ruff on the touched files is clean.

Round 1 should-fix

  1. test_synthetic_load_exhausts_tiny_bucket — addressed. It now posts json={} to /v1/query/explain. That body is valid Pydantic (ExplainRequest fields are optional) and 400s in resolve_window before get_db. Burst 2 → two 400s, then eight 429s. The 1 rps refill cannot sneak tokens in from a slow DB connect.

  2. test_call_llm_ask_takes_concurrency_semaphore — addressed. Two threads call _call_llm_ask with LLM_MAX_CONCURRENCY=1 and a blocking fake httpx.Client.post. Overlapping posts would make max_seen > 1 if the with llm_concurrency_slot around the ask HTTP path were dropped.

Must-fix

(None)

Should-fix

(None)

Nice-to-have

(None)

Prior nits (unbounded _buckets, static Retry-After, pause/resume sharing the ingest write bucket, docstring “key-admin”, no suite-level reset_rate_limiter) are unchanged and are not re-raised.

Verdict

Ready to merge (0 must-fix, 0 should-fix)

@cursor
cursor Bot merged commit e28d640 into main Aug 17, 2026
2 checks passed
@leo-aa88
leo-aa88 deleted the cursor/rate-limit-backpressure-1cc6 branch August 17, 2026 19:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(G9): rate limiting & backpressure (API, ingest queue, LLM concurrency)

2 participants