fix(shaka-lab-node): Time out the driver update at startup - #94
Merged
Conversation
start-nodes.js ran update-drivers with a blocking spawnSync and no timeout. The script talks to package registries, driver CDNs, attached devices, and OS services, and when one of those never answers, the service never finishes starting. It just sits wedged, with nothing past the last line of driver installer output to say why. This was observed on a Windows node, where the update stalled and a stray webdriver-installer process kept the service hostage until it was killed by hand. Wait for the script asynchronously with a 10 minute ceiling instead. On timeout, kill it and everything below it, then reject so the process exits non-zero and the service manager restarts us. Killing the tree matters as much as the timeout. On Windows the script runs under a shell, so killing the direct child would orphan npm and the driver installer, recreating the exact stray process seen in the incident; taskkill /T /F covers the tree, as it already does for the node processes. On other platforms the script is now spawned in its own process group so a negative PID reaches the same processes. Co-Authored-By: Claude Code (Claude Opus 5) <noreply@anthropic.com>
|
Thanks for your pull request! It looks like this may be your first contribution to a Google open source project. Before we can look at your pull request, you'll need to sign a Contributor License Agreement (CLA). View this failed invocation of the CLA check for more information. For the most up to date status, view the checks section at the bottom of the pull request. |
Member
Author
|
This should prevent the Windows service from hanging on any misbehaving update tooling in the future. The companion change in shaka-project/webdriver-installer#75 should fix the root cause of the hung update. |
avelad
approved these changes
Jul 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
start-nodes.js ran update-drivers with a blocking spawnSync and no timeout. The script talks to package registries, driver CDNs, attached devices, and OS services, and when one of those never answers, the service never finishes starting. It just sits wedged, with nothing past the last line of driver installer output to say why. This was observed on a Windows node, where the update stalled and a stray webdriver-installer process kept the service hostage until it was killed by hand.
Wait for the script asynchronously with a 10 minute ceiling instead. On timeout, kill it and everything below it, then reject so the process exits non-zero and the service manager restarts us.
Killing the tree matters as much as the timeout. On Windows the script runs under a shell, so killing the direct child would orphan npm and the driver installer, recreating the exact stray process seen in the incident; taskkill /T /F covers the tree, as it already does for the node processes. On other platforms the script is now spawned in its own process group so a negative PID reaches the same processes.