Repository navigation
Bound Flink sink retries by time, not attempt count - #206
Merged
Merged
Conversation
Central requires sources and javadoc on every component, and both coordinates upload as a single deployment bundle -- so the incomplete chalk-java-shaded component failed validation and took chalk-java 1.3.3 down with it. Both publish runs (8/10, 8/14) uploaded successfully and were rejected server-side afterward, which is why CI stayed green with nothing on Central. Reuse the unshaded jars: relocation is a bytecode rewrite, so the sources are identical and the public ai.chalk.* API is unrelocated. Deferred to afterEvaluate because the vanniktech plugin registers both tasks after this file evaluates.
Production Flink jobs restarted on transient UNAVAILABLE during query-server rollouts: "upload_features failed after 4 attempt(s)". The retries themselves worked, but a fixed maxRetries=3 with a 500ms base backoff only covers 0.5+1+2 = 3.5s -- shorter than a routine deploy, so the budget was spent while the backend was still coming back and the task failed. An attempt count can't express "survive a rollout": the real budget shifts whenever the backoff is tuned. Replace it with an explicit wall-clock budget, retryTimeout, defaulting to 120s. maxRetries stays as an opt-in hard cap and is now unlimited by default, so the time budget is the effective bound. Timing uses nanoTime so an NTP step can't move the deadline, and the loop stops when the next sleep would overrun it rather than sleeping past it. Retries still run inline on the task thread, so the budget is also the worst-case subtask stall; documented alongside the checkpoint-timeout interaction.
sjmignot
approved these changes
Aug 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Production Flink jobs restart on transient
UNAVAILABLEduring query-server rollouts:The retries worked —
after 4 attempt(s)ismaxRetries=3+ 1, andUNAVAILABLEis retryable. The budget was just too short. With the default 500ms base backoff:Connection refusedfails fast rather than hanging to the 30s upload timeout, so the real coverage is ~3.5s — shorter than a routine deploy. Channel-level gRPC retry doesn't help:maxAttempts=3with backoff capped at 0.1s adds ~0.16s.Change
An attempt count can't express "survive a rollout" — the real budget silently shifts whenever the backoff is tuned. Replaced it with an explicit wall-clock budget:
retryTimeout(new, default 120s) — total budget for retrying one batch's transient failures. Sized to outlast a normal query-server rollout.maxRetries— now an opt-in hard cap, unlimited by default, soretryTimeoutis the effective bound. Setting it explicitly preserves the old exact-attempt semantics.nanoTimeso an NTP step can't move the deadline, and the loop breaks when the next sleep would overrun the budget instead of sleeping past it.Retries still run inline on the task thread, so the budget is also the worst-case subtask stall — documented next to the checkpoint-timeout interaction.
Also included
1cd1803— the sources/javadoc fix for thechalk-java-shadedpublication, currently only onrelease/1.3.3(#205). Without it onmain, cutting 1.3.4 fails Central validation exactly as 1.3.3 did: the incomplete shaded component sinks the whole deployment bundle, takingchalk-javawith it, while CI still reports green. Harmless duplicate if #205 merges first.Tests
New:
stopsRetryingWhenTimeBudgetSpent,timeBudgetAllowsMoreAttemptsThanTheOldFixedCap,explicitMaxRetriesStillCapsBeforeBudget,retryBudgetDefaultsToTimeNotAttempts. All pre-existing retry tests pass unchanged.The 4
initializationErrorfailures inTestChalkClient*/TestGrpcClient/TestAllClientsare pre-existing credential-dependent integration tests — they fail identically on a cleanmainlocally.Behaviour note
An outage longer than the budget still surfaces as a task failure and a checkpoint restart. That is at-least-once working as intended, not data loss — this change makes the common case (a rollout) stop causing restarts at all.