Repository navigation
feat: API rate limiting, ingest backpressure, LLM concurrency (#18) - #46
Conversation
Unbounded ingest dumps and explain fan-out can overload the process and the LLM provider. Token-bucket limits on ingest/query, the existing ingest queue ceiling, and a process-wide LLM semaphore keep G9 backpressure in place without changing queue semantics. Closes #18 Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
Review — PR #46 (G9 rate limiting)Checked Token-bucket middleware sits inside auth, covers versioned and unversioned ingest writes and Must-fix(None) Should-fix
Nice-to-have
VerdictNeeds changes (should-fix remain) |
The synthetic load case posted a valid explain body that hit get_db; slow connects could refill the 1 rps bucket and flake the 429 count. Ask HTTP concurrency was untested, so dropping the semaphore would not fail CI. Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
Review — PR #46 (G9 rate limiting, round 2)Checked Round 1 should-fix
Must-fix(None) Should-fix(None) Nice-to-have(None) Prior nits (unbounded VerdictReady to merge (0 must-fix, 0 should-fix) |
Closes #18
Summary
Token-bucket limits on ingest/query, the existing ingest queue ceiling, and a process-wide LLM semaphore keep a large incident dump from unbounded API and provider fan-out.
Behavior
src/api/ratelimit.py) onPOST /v1/ingestions*(writes) and/v1/query*plus unversioned aliases. Identity isrequest.state.auth_principal.key_id, or"anonymous"when auth is off.429+Retry-Afterwith{"error_code":"RATE_LIMITED",...}./health,/docs, static UI, and/configare not limited. In-memory per process.INGEST_QUEUE_MAXstill yields429 INGEST_QUEUE_FULL; tail ticks still skip when full.src/core/llm/provider.py: semaphore aroundgenerate_summary(and ask HTTP). Noop skips the wait but still uses the same entrypoint.Config
Via
get_settings()(noRAGLOGS_prefix):RATELIMIT_*,INGEST_QUEUE_MAX,LLM_MAX_CONCURRENCY.0rps / concurrency = unlimited. Defaults are high (100rps/burst, LLM concurrency4) so existing TestClient tests do not 429.Tests
Allow-then-429,
Retry-After, per-key isolation, anonymous defaults, queue-full stillINGEST_QUEUE_FULL, LLM max_concurrency=1 serialization, noop not blocking, synthetic burst that 429s.python -m pytest tests/unit/passed locally (668).