Repository navigation
[Fix #1749] Publish onWorkflowCancelled before cancelling pending futures - #1751
Conversation
…fore cancelling pending futures Cancelling the futures registered through addCancelable (e.g. by a listen task) completes the execution pipeline synchronously, which runs cleanUp and clears the instance metadata. cancel()/cancelFuture() used to do that before publishing the status change and onWorkflowCancelled, so listeners could not find their per-instance metadata when the cancelled event arrived. internalCancel() now returns the futures to cancel, and they are cancelled only after the cancelled status change and onWorkflowCancelled have been published. Signed-off-by: Eric Deandrea <eric@ericdeandrea.dev>
There was a problem hiding this comment.
🟡 Changes recommended
Incoming listen events during asynchronous publication can still overwrite cancellation status or clear metadata before cancellation listeners run.
2 open findings
What changed in this PR
Addresses #1749 by publishing workflow cancellation events before cancelling registered futures, preserving metadata for cancellation listeners.
Changes:
- Separates cancellation state changes from future cancellation.
- Adds metadata lifecycle tests for both cancellation APIs with listen and wait workflows.
| File | Description |
|---|---|
| impl/test/src/test/java/io/serverlessworkflow/impl/test/CancelMetadataTest.java | Tests metadata availability during cancellation callbacks. |
| impl/core/src/main/java/io/serverlessworkflow/impl/WorkflowMutableInstance.java | Defers future cancellation until lifecycle publication finishes. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…ean up after it is published Address review feedback on the cancellation ordering: - status(WorkflowStatus) now checks and sets the status under statusLock and ignores changes once the instance is CANCELLED, so a listen task receiving an event while the cancellation is being published can no longer move the instance back to WAITING and let it complete normally. - The execution pipeline waits for the cancellation to be published before running cleanUp, so instance metadata is still available to onWorkflowCancelled listeners even when the pipeline ends on its own (an event completes the listen task, or a wait elapses) while an asynchronous status change listener is still pending. Adds regressions that hold the CANCELLED status change pending and deliver a late event or let the pipeline finish in the meantime. Signed-off-by: Eric Deandrea <eric@ericdeandrea.dev>
… cancelled instance status(WorkflowStatus) ignored every change once the instance was CANCELLED by returning a normally completed future, but publishEvents() and handleException() ignore that value. If the cancellation happened while the last task's onTaskCompleted/onTaskFailed listener was still pending, releasing it published onWorkflowCompleted (and completed start() normally) or onWorkflowFailed for a cancelled instance. A rejected COMPLETED or FAULTED transition now fails with a CancellationException, so the pipeline ends as cancelled without publishing completion or failure events. Late non-terminal changes (e.g. WAITING from a listen task) are still ignored. Signed-off-by: Eric Deandrea <eric@ericdeandrea.dev>
|
@edeandrea |
|
The intent isn’t to keep metadata around after the instance is logically gone; it’s to keep it available until Listeners often use In this bug, |
|
Ok, so rather than chanigng the order of hte listeners, why not postponing the metadata clearance till the workflow is actually cancelled?
I prefer the first one and I think is easily achievable with the previous strcuture. |
|
@fjtirado I agree |
|
@fjtirado Yes — that is the first option I’m aiming for. The goal is not to reorder listeners, but to keep metadata alive until |
|
This comment summarizes the thread and addresses the review points raised so far.
Agreed that
That is the approach this PR takes. The bug was that
Fixed by preventing late listen callbacks from reviving a cancelled instance or moving it back out of
Fixed by deferring cleanup until cancellation publication finishes, so metadata is still available when
Fixed by making terminal transitions on a cancelled instance fail with So the current behavior is: keep the existing listener model, preserve metadata through the terminal cancellation callback, and close the cancellation races around late listen callbacks and terminal transitions. |
|
Ok, the PR description makes me thing you were changing the order of the listeners. |
Fair point — the PR description is misleading if it sounds like this is only reordering listeners. The actual fix keeps metadata alive through |
|
I think we are addressing several issues with the same PR (which is not necessarily bad)
Im not sure 2) and 3) were really there before the change to 1). (Im pretty sure workflow failed was not invoked once canclled, but not so sure about workflow completed) so, what Im going to do is take your unit test, ran it with previous version (which should fail, a least for some of them) and try to make all test work starting from current state. Im doing that because the changes are not trivial, they affect core behaviour and, to be honest, the one in status looks too specific and the one for cleanup too verbose (a handle and a compoese where a compose should be enough) |
|
Btw, thanks a lot for detecting the issue and proposing solution. Ill probably send a PR over your PR tomorrow (too late for me now) |
I agree and think you are right. The other issues are actually issues Copilot pointed out in its reviews (see #1751 (review) & #1751 (review))
The problem I really care about (yes, I'm selfish :) ) is that if I call Thats what I documented in #1749 and also what I worked around in quarkiverse/quarkus-flow#1059 The other option could be to just publish from the pipeline rather than publishing before cancelling the futures: have |
|
@edeandrea |
0982b96 to
6e056c1
Compare
Signed-off-by: Francisco Javier Tirado Sarti <ftirados@ibm.com>
|
@edeandrea |
|
Thanks @fjtirado for the work! |


Many thanks for submitting your Pull Request ❤️!
What this PR does / why we need it:
Fixes #1749.
This bug shows up when
cancel()/cancelFuture()is called while an instance is WAITING on alistentask: metadata can be cleared beforeonWorkflowCancelledruns. There are also two related races: late listen callbacks can still interfere with cancellation, and terminal transitions can still publishonWorkflowCompleted/onWorkflowFailedafter cancellation.This PR fixes those cancellation races by:
onWorkflowCancelled, then clearing it after that callback completes.CANCELLED.CancellationException, so cancelled workflows do not publish completion/failure events.Special notes for reviewers:
onWorkflowCancelledremains the terminal callback; there is no newbeforeCancelledevent.cancel()andcancelFuture()), both workflow shapes (listenandwait), and the terminal-transition races.Additional information (if needed):
mvn -pl impl -amd installpasses locally.SchedulerTest.testAfter(200 ms await) is flaky onmainindependently of this change: it failed 1 of 5 runs on unmodifiedmainlocally.