Skip to content

after_run_callback is not called when the task driving Runner.run_async is cancelled (regression in 2.11.0) #7394

Description

@roanny

Summary

Since 2.11.0, when the task driving Runner.run_async(...) is cancelled from outside (which is what adk api_server's /run does on a client disconnect: worker_task.cancel()), no after_run_callback of any plugin is called. on_run_error_callback is not called either, because asyncio.CancelledError is not an Exception. A plugin that finalizes per run therefore gets no callback at all for a cancelled run: not for telemetry, not for clean-up, not for writing a terminal event to the session.

On 2.10.0 and 2.9.2 the same cancellation still reaches after_run_callback.

Reproduction

A scripted model, one slow tool, an in-memory runner, and a plugin that records its callbacks. No Google Cloud access is needed.

"""after_run_callback is not called when the task driving Runner.run_async is cancelled.

This is what `adk api_server`'s /run does on a client disconnect: worker_task.cancel().
No Google Cloud access is needed: a scripted model and an in-memory runner.
"""
import asyncio

import google.adk
from google.adk.agents.llm_agent import LlmAgent
from google.adk.apps import App
from google.adk.models.base_llm import BaseLlm
from google.adk.models.llm_response import LlmResponse
from google.adk.plugins.base_plugin import BasePlugin
from google.adk.runners import InMemoryRunner
from google.adk.tools import FunctionTool
from google.genai import types

calls = []
started = asyncio.Event()


class Recorder(BasePlugin):
    async def before_run_callback(self, *, invocation_context):
        calls.append("before_run")

    async def after_run_callback(self, *, invocation_context):
        calls.append("after_run")

    async def on_run_error_callback(self, *, invocation_context, error):
        calls.append(f"on_run_error({type(error).__name__})")


class ScriptedModel(BaseLlm):
    async def generate_content_async(self, llm_request, stream=False):
        call = types.FunctionCall(name="slow_tool", args={}, id="c1")
        yield LlmResponse(content=types.Content(role="model", parts=[types.Part(function_call=call)]))


async def slow_tool() -> dict:
    started.set()
    await asyncio.sleep(30)
    return {}


async def main():
    agent = LlmAgent(name="a", model=ScriptedModel(model="scripted"), instruction="x", tools=[FunctionTool(slow_tool)])
    runner = InMemoryRunner(app=App(name="app", root_agent=agent, plugins=[Recorder(name="recorder")]))
    await runner.session_service.create_session(app_name="app", user_id="u", session_id="s")
    message = types.Content(role="user", parts=[types.Part(text="go")])

    async def worker():
        return [e async for e in runner.run_async(user_id="u", session_id="s", new_message=message)]

    task = asyncio.create_task(worker())
    await started.wait()
    task.cancel()  # as /run does on http.disconnect
    try:
        await task
    except asyncio.CancelledError:
        pass
    print("google-adk", google.adk.__version__, "->", calls)


asyncio.run(main())

Output:

google-adk 2.9.2 -> ['before_run', 'after_run']
google-adk 2.10.0 -> ['before_run', 'after_run']
google-adk 2.11.0 -> ['before_run']

main at 63aed551 behaves like 2.11.0.

Where it comes from

Bisected between v2.10.0 and v2.11.0 with the script above: ef5bbcfe ("feat: add optional abort_signal support to Runner, Workflow, and Nodes") is the first commit where after_run is no longer called.

In 2.11.0, workflow/_node_runner_utils.py handles the cancellation like this:

except asyncio.CancelledError as e:
  if e.args and e.args[0] == _CALLER_CLOSED_EARLY_MSG:
    closing_early = True
  else:
    run_error = e
  raise

The finally that follows runs the after-run callbacks only when run_error is None. A cancellation that does not carry _CALLER_CLOSED_EARLY_MSG is treated as a run error and skips them. /run's disconnect monitor cancels the worker with no message (cli/api_server.py: worker_task.cancel()), and so does any caller that cancels its own task.

Impact

In our deployment, a plugin writes a terminal event to the session when an invocation ends without answering, so that the client stops waiting. On 2.11.0 a /run cut by the platform's request timeout leaves the session ending on unanswered function calls, and the client keeps waiting. We worked around it by watching the task in before_run_callback. The same gap affects any plugin that counts, traces or cleans up per run.

Possible directions

  • Treat an external cancellation like an early close for the purpose of after_run_callback.
  • Notify plugins of a cancelled run through on_run_error_callback, or through a dedicated callback.
  • Have /run cancel its worker with a message the runner recognizes.

Environment: Python 3.12, google-adk 2.11.0 (and main at 63aed551).

Activity

  1. Li-john1021 commented on Oct 3, 2026

    @Li-john1021

    I'd like to work on this. Could a maintainer please assign this issue to me?

    I can restore after_run_callback for an external cancel of Runner.run_async (including /run disconnecting with worker_task.cancel() and no message), and add a regression test. I won't open a PR until it's assigned.

  2. added theissue type on Oct 5, 2026
  3. added
    core[Component] This issue is related to the core interface and implementation
    on Oct 5, 2026
  4. sattyamjjain commented on Oct 5, 2026

    @sattyamjjain

    +1, this matters a lot for audit plugins. A run with a start record and no end record is ambiguous later: crashed, timed out, or still going?

    Of the three directions, I'd prefer the second (a dedicated cancelled callback, or on_run_error_callback with the CancelledError) over the first. If after_run_callback fires on a cancel the same way it does on a normal finish, any plugin that writes "completed" there will mark a cut-off run as completed. Plugins need to know it was cancelled to write the right terminal event.

  5. surajksharma07 commented on Oct 5, 2026

    @surajksharma07
    Collaborator

    Reproduced this with your script @roanny. 2.9.2 and 2.10.0 call after_run but 2.11.0 and current main only call before_run so upgrading won't help yet.

    As a workaround where your own code cancels the run: pass abort_signal=asyncio.Event() to run_async, call abort_signal.set(), then task.cancel(_CALLER_CLOSED_EARLY_MSG) (from google.adk.runners). On my side after_run fires with invocation_context.is_aborted=True and the dangling function call gets an "aborted" response so the client shouldn't hang. Could you check whether that works in your deployment? The constant is private and it doesn't cover the /run disconnect path which needs a change in ADK itself.

    @abhayjoshi201 @CodeAlex52 thanks for the PRs. Before review please test them against the plain task.cancel() case from the repro above. Right now #7408 only covers /run and #7403 only helps when abort_signal has been set.

  6. roanny commented on Oct 6, 2026

    @roanny
    Author

    Thanks for reproducing it. The workaround doesn't apply to our case: we never cancel the run ourselves. Our cut is /run's client disconnect, where api_server.py (2.11.0) runs the run in a worker task and calls worker_task.cancel() with no abort_signal — only /run_sse wires one — which is the path you note needs a change in ADK. Until then we keep watching the run's task from a plugin.

    On the direction, we'd prefer what @sattyamjjain describes: a cancelled run reported as cancelled (a dedicated callback, or on_run_error_callback with the CancelledError) rather than after_run_callback firing as if it had completed. #7408 targets the /run path; we're happy to test it against our deployment once it's ready for review.

  7. surajksharma07 commented on Oct 7, 2026

    @surajksharma07
    Collaborator

    #7403 has been merged to main in d806a36 @roanny so this should cover your case now. /run creates an abort_signal and sets it before worker_task.cancel() so after_run_callback runs when the client disconnects. Not in a release yet (2.11.0 is still the latest) so you'd need to install from main to try it before the next release.

    On cancelled vs. completed (@sattyamjjain): after_run_callback now gets invocation_context.is_aborted=True for a disconnect so your plugin can write a "cancelled" terminal event instead of "completed". The dangling function call is also closed with {"error": "Invocation was aborted by client."} so the client stops waiting. Checked this on main with your repro.

    A plain task.cancel() without abort_signal still skips after_run on main. That's intentional so if you need a dedicated cancel callback for that path a separate feature request would be the right place for it.

    Could you test main against your deployment and let us know? If /run behaves for you we can close this and you can drop the task-watching workaround once the release is out.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

core[Component] This issue is related to the core interface and implementation

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions