Repository navigation
feat: LLM timeouts, retries, auto-fallback, and circuit breaker (#19) - #47
Conversation
Wire LLM_TIMEOUT, LLM_MAX_RETRIES, LLM_MAX_TOKENS, and breaker settings through get_settings so every provider call has a bounded timeout and jittered retries. On failure, explain/ask fall back to deterministic templates with llm.fell_back; an open breaker is surfaced on GET /health without failing probes. Closes #19 Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
Ruff F401 failed on the touched health test module. Co-authored-by: Leonardo <leo-aa88@users.noreply.github.com>
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
Review — PR #47 (G10 LLM resilience)Checked Acceptance holds: timeout is on Must-fix(None) Should-fix(None) Nice-to-have(None) VerdictReady to merge (0 must-fix, 0 should-fix) |
Closes #19
Every LLM call (explain
generate_summaryand ask HTTP) now has a configured timeout, bounded jittered retries, a per-request token ceiling, and a process-local circuit breaker.On timeout, error, exhausted retries, over-budget evidence, or an open breaker, explain/ask fall back to deterministic templates and set
llm.fell_back=true. The request still succeeds. Fallback output is the evidence packet rendered by existing templates — polish drops, grounding does not.GET /healthgainsllm_breaker{state, consecutive_failures, cooldown_remaining_seconds}. An open breaker marksstatus=degradedbut still returns 200. Defaults keepLLM_PROVIDER=disabled/ noop unchanged. G9CappedLLMProviderremains the outer wrapper.Config
Unprefixed:
LLM_TIMEOUT(30s),LLM_MAX_RETRIES(2 extra / 3 total),LLM_MAX_TOKENS(600),LLM_MAX_INPUT_TOKENS,LLM_BREAKER_THRESHOLD(5),LLM_BREAKER_COOLDOWN_SECONDS(60).Tests
Injected timeout/failure → template fallback equals
render_text_summary/_rules_answer. G9 concurrency tests still pass. Unit suite passed locally (688).