Repository navigation
Standard error handler interface for workflow and activity failures #12152
Replies: 2 comments
|
@robzienert Thanks! I think this generally makes sense. My main concern is not having too many very-specific interceptors. Folks on the team and I have been talking about the possibility of a generic Would love to have you guys make some PRs, I think how exactly that looks depends on if we introduce a I think we will also need to enumerate what sites, exactly, constitute an unhandled error/panic as it relates to a specific workflow or activity. I think there may be (without looking too closely at this exact moment) some cases where the error may be maybe tangentially but not directly related to some specific workflow/activity run. Maybe most challenging, is for the Core based SDKs, certainly a lot of those sites are in Core and not in the language layer. This will require the introduction of a callback into the lang bridges to invoke the user's handler if the error happened inside Core. Same thing here where it's maybe not immediately clear where the line should be drawn for various sites. |
|
If a generic interceptor interface is possible, I agree that's definitely the path to take. Any new api surface is a long-term cost, so I'm aligned on being purposeful about what is introduced. Looking forward to hear what the team thinks. The nuances of error call sites and context you outlined puts the fear in me that such a change could be quite high-risk. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Is your feature request related to a problem? Please describe.
Each SDK currently makes a hardcoded choice about how that failure is logged. The application, however, has the context that would make the log actionable, but must log separately and rely on correlation IDs to provide the context.
Adding an SDK exception handler interface that is called when an exception leaks would allow the application to provide actionable log messages. While I'm talking specifically about logs here, I would see our use also hooking into telemetry, etc.
The handler should be written in such a way that the contract is consistent across SDKs. If the handler is undefined, it would default to the current implementation.
The four SDKs we use (Go, Java, Python, TypeScript) already disagree on what to do here, but the benefit of introducing an interface would benefit all SDKs:
WorkflowExecutionHandler#throwAndFailWorkflowExecution—WARN"Workflow execution failure WorkflowId='{}', RunId={}, WorkflowType='{}'"internal_task_handlers.go applyWorkflowPanicPolicy—Error"Workflow panic"+ workflow/run/attempt/error/stack tags_workflow_instance.py—WARNING"Failed activation on workflow ...",extra={"__temporal_error_identifier": "WorkflowTaskFailure"}worker.ts—error"Failed to process Workflow Activation"— but only fires for unhandled-rejection / SDK-internal paths; normal user-thrown exceptions silently become a failed completion with no TS-side logExisting extension points don't compose:
ClientOptions.Logger, TSRuntime.install({logger})) are client/process-wide, not failure-aware.runConstructor) and the SDK's own log still fires after them.logging.Filteragainst__temporal_error_identifier) work but lean on internal class names, message strings, and extras the SDK doesn't promise to keep stable.Describe the solution you'd like
Add a per-SDK exception handler invoked when a workflow or activity leaks an uncaught exception, so the application decides how that failure is logged.
Additional context
Once we agree on a contract, we would be willing to implement for some of the SDKs: starting first with Java, then our other paved road (Go, Python, TypeScript) as our capacity allows (:fingerscrossed: EOY26).
All reactions