fix(cli): keep after_run callbacks when /run client disconnects - #7408
Open
CodeAlex52 wants to merge 1 commit into
Open
CodeAlex52 wants to merge 1 commit into
CodeAlex52 wants to merge 1 commit into
Conversation
Signed-off-by: CodeAlex52 <59381946+CodeAlex52@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #7394.
Problem
When a client disconnects from
/run, the disconnect monitor cancels the worker task with a bareTask.cancel(). The runner treats a bare cancellation as an external cancel, which deliberately skipsafter_run_callback(seetest_run_async_cancellation_does_not_execute_after_run_plugin), so a plugin that finalizes a run — writing a terminal event to the session, telemetry, cleanup — gets no callback at all for a disconnected run.That path already exists: cancelling with the caller-closed-early marker makes the runner treat the run as "the caller stopped consuming events", which does run the after-run callbacks.
/run_ssegets this for free because it closes the generator; the non-streaming/runendpoint is the only path that cancels without the marker.Fix
One line in
cli/api_server.py, plus the test that pins it:I deliberately did not change the bare-cancellation semantics in
runners.py/workflow/_node_runner_utils.py. Flipping those would contradict the existing intentional behavior and its test, and it would change what happens for genuine external cancellations (shutdown, a caller cancelling its own task). The disconnect case is a caller going away, which is exactly what the marker means, so the fix belongs at the call site.The error path is unaffected: the generator still re-raises
CancelledError, the worker task still ends cancelled, and/runstill answers 499.Tests
Added
test_agent_run_disconnect_marks_cancellation_as_caller_closed_earlynext to the existing disconnect test intests/unittests/cli/test_fast_api.py, reusing the same harness and asserting theCancelledErrorthat reachesrun_asynccarries the marker.Validation:
worker_task.cancel()reachesrun_asyncwithe.args == ()) and passes with the change.tests/unittests/test_runners.py,tests/unittests/plugins,tests/unittests/workflow: 1589 passed, 1 skipped, 5 xfailed.tests/unittests/cli/test_fast_api.py: 176 passed with the change vs 175 on the unmodified tree; the 6 failures and 5 errors in that file are pre-existing in my environment (metrics/agent-identity tests and optional-dependency collection errors) and are identical before and after.pyink --checkclean on both files.Note for discussion
@Li-john1021 commented on #7394 proposing to restore
after_run_callbackfor any external cancel ofRunner.run_async. That is a different, larger behavior change and conflicts with the deliberate semantics noted above — happy to discuss, but I think the disconnect path is the actual gap worth fixing first.