Skip to content

feat: LLM timeouts, retries, auto-fallback, and circuit breaker (#19) - #47

Merged
cursor[bot] merged 2 commits into
mainfrom
cursor/llm-resilience-1cc6
Aug 17, 2026
Merged

cursor[bot] merged 2 commits into
mainfrom
cursor/llm-resilience-1cc6

Conversation

@leo-aa88

Copy link
Copy Markdown
Member

Closes #19

Every LLM call (explain generate_summary and ask HTTP) now has a configured timeout, bounded jittered retries, a per-request token ceiling, and a process-local circuit breaker.

On timeout, error, exhausted retries, over-budget evidence, or an open breaker, explain/ask fall back to deterministic templates and set llm.fell_back=true. The request still succeeds. Fallback output is the evidence packet rendered by existing templates — polish drops, grounding does not.

GET /health gains llm_breaker {state, consecutive_failures, cooldown_remaining_seconds}. An open breaker marks status=degraded but still returns 200. Defaults keep LLM_PROVIDER=disabled / noop unchanged. G9 CappedLLMProvider remains the outer wrapper.

Config

Unprefixed: LLM_TIMEOUT (30s), LLM_MAX_RETRIES (2 extra / 3 total), LLM_MAX_TOKENS (600), LLM_MAX_INPUT_TOKENS, LLM_BREAKER_THRESHOLD (5), LLM_BREAKER_COOLDOWN_SECONDS (60).

Tests

Injected timeout/failure → template fallback equals render_text_summary / _rules_answer. G9 concurrency tests still pass. Unit suite passed locally (688).

Open in Web Open in Cursor 

cursoragent and others added 2 commits August 17, 2026 09:51
Wire LLM_TIMEOUT, LLM_MAX_RETRIES, LLM_MAX_TOKENS, and breaker
settings through get_settings so every provider call has a bounded
timeout and jittered retries. On failure, explain/ask fall back to
deterministic templates with llm.fell_back; an open breaker is
surfaced on GET /health without failing probes.

Closes #19

Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
Ruff F401 failed on the touched health test module.

Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

@cursor

cursor Bot commented Aug 17, 2026

Copy link
Copy Markdown
Contributor

Review — PR #47 (G10 LLM resilience)

Checked origin/main...origin/cursor/llm-resilience-1cc6 (commits 99a0ad0, 87ff814). Unit suite: 688 passed. Ruff on touched files: clean.

Acceptance holds: timeout is on httpx.Client(timeout=LLM_TIMEOUT); retries use wait_exponential_jitter with LLM_MAX_RETRIES+1 attempts; explain and ask catch provider errors and keep mode="rules" so llm.fell_back is true; token guard trims only list keys (evidence, clusters, secondary_clusters, trigger_candidates) and falls back if window / primary_cluster still overflow; breaker skips the inner provider while open and is exposed on GET /health as degraded + HTTP 200; settings are unprefixed via get_settings(); noop skips ResilientLLMProvider; CappedLLMProvider remains the outer wrapper. Ask HTTP uses invoke_llm + prepare_llm_packet + complete() after unwrap, so it is not a second unguarded client.

Must-fix

(None)

Should-fix

(None)

Nice-to-have

(None)

Verdict

Ready to merge (0 must-fix, 0 should-fix)

@cursor
cursor Bot merged commit 0185bc6 into main Aug 17, 2026
2 checks passed
@leo-aa88
leo-aa88 deleted the cursor/llm-resilience-1cc6 branch August 17, 2026 19:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(G10): LLM resilience (timeout, retries, auto-fallback, circuit breaker)

2 participants