What happened
When WorkflowInstance.cancel() (or cancelFuture()) is called on an instance that is WAITING on a listen task, WorkflowMutableInstance closes and clears the instance metadata (additionalObjects) before it publishes onWorkflowCancelled. A WorkflowExecutionListener that keeps per-instance state in metadata (addMetadataIfAbsent / findMetadata) therefore can't find that state when the cancelled event arrives.
This is how it surfaced in quarkus-flow's OpenTelemetry listener: the workflow.execute span of the cancelled instance was never ended or exported (quarkiverse/quarkus-flow#1058).
Why (7.35.2.Final, unchanged on main)
cancel() -> internalCancel() sets status to CANCELLED and, outside the lock, calls toCancel.forEach(t -> t.cancel(true)). For a WAITING instance that is the future ListenExecutor registered through addCancelable.
- Cancelling that future completes the start pipeline right away on the calling thread:
AbstractTaskExecutor.handleException publishes onTaskCancelled, WorkflowMutableInstance.handleException sees the CancellationException and publishes nothing, then .whenComplete(this::cleanUp) closes every AutoCloseable in additionalObjects and clears the map.
- Only after
internalCancel() returns does cancel() call publishStatusChange(..., CANCELLED) and then onWorkflowCancelled. By then findMetadata returns Optional.empty().
For an instance cancelled while RUNNING (e.g. during a wait) the order is the opposite: cancelCheck fails the next task later, so onWorkflowCancelled is published before cleanUp. So whether listeners can see metadata on cancel depends on what the instance was doing when it was cancelled.
Reproducer
https://github.com/edeandrea/quarkus-flow-reproducers (issue1058). In short:
var instance = definition.instance(input); // workflow: a single listen task for an event that never comes
instance.start();
await().until(() -> instance.status() == WorkflowStatus.WAITING);
instance.cancel();
// a listener's onWorkflowCancelled(ev): ev.workflowContext().instanceData().findMetadata(KEY, ...) is empty
Expected
Instance metadata stays available until the terminal lifecycle event (onWorkflowCompleted / onWorkflowFailed / onWorkflowCancelled) has been published, for every cancel path, the same way it already is for completion and failure.
Possible fix
Either of these, both inside WorkflowMutableInstance:
- Publish before cancelling the futures: in
cancel() / cancelFuture(), publish the status change and onWorkflowCancelled first, then cancel the collected cancelables (e.g. internalCancel() returns the futures to cancel and the caller cancels them after publishEvent(...) completes).
- Publish from the pipeline: have
handleException publish onWorkflowCancelled when it sees a CancellationException, so it runs before whenComplete(this::cleanUp) on every path, and stop publishing it from cancel() / cancelFuture(). This also covers the RUNNING path consistently, but cancelFuture() would then need to complete only once the pipeline has published.
Happy to send a PR for whichever you prefer.
Downstream workaround
quarkus-flow is working around it by ending the span from its metadata object's close() when the instance status is already CANCELLED (quarkiverse/quarkus-flow#1058). That workaround could be removed once this is fixed upstream.
What happened
When
WorkflowInstance.cancel()(orcancelFuture()) is called on an instance that is WAITING on alistentask,WorkflowMutableInstancecloses and clears the instance metadata (additionalObjects) before it publishesonWorkflowCancelled. AWorkflowExecutionListenerthat keeps per-instance state in metadata (addMetadataIfAbsent/findMetadata) therefore can't find that state when the cancelled event arrives.This is how it surfaced in quarkus-flow's OpenTelemetry listener: the
workflow.executespan of the cancelled instance was never ended or exported (quarkiverse/quarkus-flow#1058).Why (7.35.2.Final, unchanged on
main)cancel()->internalCancel()sets status toCANCELLEDand, outside the lock, callstoCancel.forEach(t -> t.cancel(true)). For a WAITING instance that is the futureListenExecutorregistered throughaddCancelable.AbstractTaskExecutor.handleExceptionpublishesonTaskCancelled,WorkflowMutableInstance.handleExceptionsees theCancellationExceptionand publishes nothing, then.whenComplete(this::cleanUp)closes everyAutoCloseableinadditionalObjectsand clears the map.internalCancel()returns doescancel()callpublishStatusChange(..., CANCELLED)and thenonWorkflowCancelled. By thenfindMetadatareturnsOptional.empty().For an instance cancelled while RUNNING (e.g. during a
wait) the order is the opposite:cancelCheckfails the next task later, soonWorkflowCancelledis published beforecleanUp. So whether listeners can see metadata on cancel depends on what the instance was doing when it was cancelled.Reproducer
https://github.com/edeandrea/quarkus-flow-reproducers (
issue1058). In short:Expected
Instance metadata stays available until the terminal lifecycle event (
onWorkflowCompleted/onWorkflowFailed/onWorkflowCancelled) has been published, for every cancel path, the same way it already is for completion and failure.Possible fix
Either of these, both inside
WorkflowMutableInstance:cancel()/cancelFuture(), publish the status change andonWorkflowCancelledfirst, then cancel the collectedcancelables(e.g.internalCancel()returns the futures to cancel and the caller cancels them afterpublishEvent(...)completes).handleExceptionpublishonWorkflowCancelledwhen it sees aCancellationException, so it runs beforewhenComplete(this::cleanUp)on every path, and stop publishing it fromcancel()/cancelFuture(). This also covers the RUNNING path consistently, butcancelFuture()would then need to complete only once the pipeline has published.Happy to send a PR for whichever you prefer.
Downstream workaround
quarkus-flow is working around it by ending the span from its metadata object's
close()when the instance status is alreadyCANCELLED(quarkiverse/quarkus-flow#1058). That workaround could be removed once this is fixed upstream.