Skip to content

fix(mcp): allow slow bot deployment and clarify timeout outcomes - #261

Draft
mlguys wants to merge 1 commit into
hummingbot:mainfrom
mlguys:fix/bot-deploy-timeout
Draft

mlguys wants to merge 1 commit into
hummingbot:mainfrom
mlguys:fix/bot-deploy-timeout

Conversation

@mlguys

@mlguys mlguys commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

Bot deployment can finish after the Hummingbot API client's default 30-second timeout. In the observed case, the tool reported Failed to manage bots: while the backend created the bot roughly 42 seconds after the request, leaving the agent uncertain whether deployment succeeded.

This change gives manage_bots(action="deploy") a temporary SDK client with a timeout of at least 120 seconds, preserving any larger configured timeout. Other actions continue using the shared client. The temporary client closes on success, initialization failure, and cancellation.

If deployment still times out, the tool explicitly reports an unknown outcome, names the bot, controller configs, and account, and tells the agent to reconcile before retrying. It sends one deployment POST and performs no automatic deployment retry.

Validation:

  • 109 focused tests passed, including real SDK calls to a local mock HTTP API for delayed success and timeout outcomes.
  • Regression coverage checks timeout isolation, larger configured budgets, one POST, client cleanup, and unchanged status behavior.
  • Black, isort, and Git whitespace checks passed.
  • No live bot deployment was used for validation.

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[Medium risk] Changes bot deployment timeout handling and error messaging.

The PR is not ready to merge until deployment retains a retryable connection check before its single POST.

Findings

  1. P1 Deployment loses connection retries ▶

Summary

The PR gives bot deployment a separate SDK client with a timeout of at least 120 seconds, reports an unknown outcome when the deployment POST times out, and adds focused timeout and cleanup tests.

  • Other bot actions retain the shared client.
  • The deployment path loses the shared client's retryable API readiness check.

Reviews (1) · Last reviewed commit: "fix(mcp): allow slow bot deployment and ..."

max_controller_drawdown_quote=max_controller_drawdown_quote,
image=image,
)
async with hummingbot_client.deployment_client() as client:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Deployment loses connection retries

If the API briefly refuses connections before deployment, the new client makes no retryable readiness check before its first deployment POST. The previous shared-client path retried its accounts check according to max_retries, so a deployment that could have succeeded after the API recovered now fails. Restore retries for a pre-deployment check without retrying the POST when its outcome is unknown.

@mlguys
mlguys marked this pull request as draft September 30, 2026 12:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant