Skip to content

fix: forward generation timeout to model HTTP requests - #462

Merged
Sun-sunshine06 merged 1 commit into
OpenCoworkAI:mainfrom
Lxr-max:fix/forward-generation-timeout-to-requests
Oct 4, 2026
Merged

Sun-sunshine06 merged 1 commit into
OpenCoworkAI:mainfrom
Lxr-max:fix/forward-generation-timeout-to-requests

Conversation

@Lxr-max

@Lxr-max Lxr-max commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Summary

generationTimeoutSec (Settings → Advanced → Generation timeout) only arms the run-level AbortController. The model HTTP requests go through pi-ai without timeoutMs, so the OpenAI and Anthropic SDK clients use their own 600 000 ms default. Any single request longer than 10 minutes gets cut by the client. This is common with large local models on LM Studio, Ollama, or vLLM, and it happens even when the user has set 3600 s or 7200 s.

@mariozechner/pi-ai 0.72.1 (already our pinned version) exposes StreamOptions.timeoutMs and maps it to the SDK timeout in openai-completions, openai-responses, azure-openai-responses, and anthropic. This PR passes the configured value through:

  • providers: GenerateOptions.timeoutMs is forwarded to completeSimple.
  • core: GenerateInput.requestTimeoutMs. pi-agent-core's Agent does not forward timeoutMs from its own options, so when a timeout is set, the agent gets a streamFn that calls streamSimple with { ...options, timeoutMs }. It applies to every turn and retry agent. With no timeout set, the agent keeps pi-agent-core's default stream.
  • desktop: the generate run derives requestTimeoutMs from generationTimeoutSec (generationRequestTimeoutMs) and passes it to the agent and the visual-parity judge complete() call. The request timeout is equal to the run timeout, so the run-level abort (with its clear "Generation aborted after Ns" message) is still what ends a run.

Relation to #411 (unlimited timeout): armGenerationTimeout already treats 0 as "disabled". generationRequestTimeoutMs(0) maps that to the largest delay Node timers accept (2 147 483 647 ms ≈ 24.8 days) instead of falling back to the SDK's 10 minutes. Huge values are clamped the same way, because larger delays overflow and fire immediately. That means #411 only needs the UI/preferences side once this lands. This PR does not change the Settings options.

Out of scope (follow-up): short auxiliary calls (run-preference router, title, memory/brief summaries) still use the SDK default. They have small output budgets, and the router call doesn't take a signal yet either.

Type of change

  • Bug fix
  • New feature
  • Refactor (no behavior change)
  • Documentation
  • Build / CI / tooling
  • Breaking change

Linked issue

Closes #199
Refs #411

Checklist

  • I checked the linked issue / relevant context before starting
  • pnpm lint && pnpm typecheck && pnpm test passes locally
  • Added/updated tests for the change
  • Added a changeset (pnpm changeset) if user-visible
  • Updated docs if behavior changed (no docs describe the SDK timeout)

Tests

  • packages/providers/src/request-timeout.test.ts: real pi-ai and OpenAI SDK against a local stalled OpenAI-compatible server. complete(..., { timeoutMs: 200 }) rejects with "Request timed out" in about 2 s. Without the fix it hangs on the SDK's 10-minute default (the test times out at 20 s).
  • packages/providers/src/index.test.ts: timeoutMs reaches completeSimple.
  • packages/core/src/agent.test.ts: with requestTimeoutMs, the agent's streamFn calls streamSimple with timeoutMs. Without it, no custom streamFn is installed.
  • apps/desktop/src/main/generation-ipc.test.ts: generationRequestTimeoutMs (1200 s → 1 200 000 ms, 7200 s → 7 200 000 ms, 0/huge → int32 max).

PRINCIPLES §5b

  • Compatibility ✅ All new fields are optional, and there are no IPC, config, or schema changes. Providers that don't support timeoutMs (for example Google and Bedrock) ignore it in pi-ai.
  • Upgradeability ✅ It uses pi-ai's public timeoutMs option (no patching of SDK internals) and nothing is persisted.
  • No bloat ✅ No new dependencies, about 40 lines of production code.
  • Elegance ✅ The existing user setting now also drives the per-request limit, so there's no new setting. Core still talks only to pi-ai, with no direct SDK imports.

generationTimeoutSec only armed the run-level AbortController. The HTTP
requests themselves went through pi-ai without timeoutMs, so the OpenAI /
Anthropic SDK clients applied their 600s default and cut long local-model
turns (LM Studio, Ollama, vLLM) at 10 minutes regardless of the setting.

Thread the configured timeout through:
- providers: complete() accepts timeoutMs and passes it to completeSimple
- core: GenerateInput.requestTimeoutMs wraps the agent stream function so
  every streamSimple call carries timeoutMs
- desktop: the generate run derives it from generationTimeoutSec for the
  agent and the visual-parity judge; a disabled timeout (0) maps to the
  largest delay Node timers accept

Refs OpenCoworkAI#199
@github-actions github-actions Bot added docs Documentation area:desktop apps/desktop (Electron shell, renderer) area:core packages/core (generation orchestration) area:providers packages/providers (pi-ai adapter, model calls) labels Oct 3, 2026

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings

No Blocker, Major, or Minor findings. The change is additive (timeoutMs on GenerateOptions, requestTimeoutMs on GenerateInput), keeps every model call inside @mariozechner/pi-ai, adds no dependencies, and includes a changeset (.changeset/forward-generation-timeout-to-requests.md:1). The per-request timeout math mirrors the existing run timer (apps/desktop/src/main/generation-ipc.ts:110) and is covered by tests.

Questions

  • apps/desktop/src/main/ipc/generate.ts:710 now reads readPreferences() unconditionally at handler entry, before the run's timeout is armed. armGenerationTimeout deliberately rethrows prefs failures as PREFERENCES_READ_FAIL / PREFERENCES_INVALID_TIMEOUT (apps/desktop/src/main/generation-ipc.ts, armGenerationTimeout). Confirm the early read sits inside that same error-mapping path, or derive requestTimeoutMs from the value armGenerationTimeout already reads, so a corrupt/invalid prefs file still yields the documented error code instead of a raw rejection.
  • packages/core/src/agent.ts:30 imports streamSimple directly instead of routing the agent stream through packages/providers. AGENTS.md prefers extending packages/providers when pi-ai lacks a capability. Was a providers-level streamWithTimeout helper considered, and is the direct streamSimple wrapper in core intentional?

Summary

  • Review mode: initial
  • Direction is sound. The provider SDKs default to a 10-minute per-request timeout that ignored generationTimeoutSec; forwarding the setting through pi-ai's timeoutMs is the right, non-invasive fix, and mapping a disabled (0) or oversized value to the int32 timer max preserves the intent of #411 without adding a new setting. Since the run-level AbortController and the per-request timeout use the same value and the run timer starts no later than the first request, run-level abort still ends a run as the PR describes.
  • Residual risk: desktop always computes a concrete requestTimeoutMs, so the agent always receives a custom streamFn; the undefined fallback back to pi-agent-core's default stream is only reachable from non-desktop callers. packages/core/src/agent.test.ts mocks pi-agent-core and only checks argument passthrough — the only end-to-end evidence that timeoutMs reaches the SDK is packages/providers/src/request-timeout.test.ts.
  • The Closes #199 / Refs #411 claims could not be validated against the issue bodies in this run (issue text not present in the provided public context).

Testing

  • Coverage is solid: real stalled-server integration test (packages/providers/src/request-timeout.test.ts), providers passthrough (packages/providers/src/index.test.ts), agent streamFn present/absent (packages/core/src/agent.test.ts), and desktop mapping (apps/desktop/src/main/generation-ipc.test.ts).
  • Gap: no test asserts the apps/desktop/src/main/ipc/generate.ts wiring actually passes requestTimeoutMs to both generateViaAgent and the visual-parity judge complete() opts. A small unit test (or an explicit comment if the preflight harness makes that impractical) would close this.

Open-CoDesign Bot

@Lxr-max

Lxr-max commented Oct 4, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for the review. I traced the preference read against armGenerationTimeout. No code change.

Preference errors. The new readPreferences() call is inside runGenerate (apps/desktop/src/main/ipc/generate.ts), and both callers await armTimeout before runGenerate (the generate handler and the comment-apply handler). armGenerationTimeout therefore still runs first. A thrown prefs read is rethrown as PREFERENCES_READ_FAIL. A non-finite or negative generationTimeoutSec is PREFERENCES_INVALID_TIMEOUT. readPersisted() already throws CodesignError with PREFERENCES_READ_FAIL for a missing, unreadable, or invalid preferences.json (including a non-positive generationTimeoutSec); it does not reject with a bare Error. A value that gets past that arming step is a finite timeout. generationRequestTimeoutMs then uses the same int32 cap as the run timer, and maps 0 to that cap so a disabled run timer is not replaced by the SDK’s 10-minute default. I am leaving the second read where it is: the request timeout is computed only after that mapped read has succeeded, from the same readPreferences() helper.

streamSimple in core. This is intentional. packages/providers forwards timeoutMs on complete() to pi-ai completeSimple (the visual-parity judge and other one-shot calls). pi-agent-core’s Agent does not take timeoutMs on its own options, so packages/core installs a streamFn at the agent constructor that calls pi-ai streamSimple with { ...options, timeoutMs }. That call stays on pi-ai. Core already imports @mariozechner/pi-ai for model and message types. A providers-level streamWithTimeout would only wrap that one streamSimple call.

Wiring test. generationRequestTimeoutMs is covered in generation-ipc.test.ts. GenerateInput.requestTimeoutMs reaching streamSimple is covered in agent.test.ts. GenerateOptions.timeoutMs reaching completeSimple, including a stalled local server, is covered in packages/providers. The desktop handler passes that same number into generateViaAgent and into the judge’s complete() options. The IPC tests mock generateViaAgent and never call the judge, so an extra assertion there would not show the SDK timeout. I am leaving the IPC suite as it is.

@Sun-sunshine06
Sun-sunshine06 merged commit f35b854 into OpenCoworkAI:main Oct 4, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core packages/core (generation orchestration) area:desktop apps/desktop (Electron shell, renderer) area:providers packages/providers (pi-ai adapter, model calls) docs Documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Generations silently cancelled at 10 min against local LM Studio / Ollama endpoints, regardless of generationTimeoutSec

2 participants