Repository navigation
Free-threaded CPython can starve a thread reattaching during repeated stop-the-world GC #151518
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Jun 15, 2026 - addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Jun 15, 2026 Hey, this seems like an interesting example of a thread starvation problem.
I take some time to dive into this problem, and based on my understanding, the root cause is a race in
start_the_world/park_detached_threads: a thread woken fromSUSPENDED→DETACHEDcan be re-parked before it gets a chance to re-attach, because the nextgc.collect()call fires immediately, andpark_detached_threadssees it as plainDETACHED.I have some personal suggestions on this problem fix, maybe introduce
DETACHED_WAKING (4)as a transient state:- start_the_world transitions
SUSPENDED→DETACHED_WAKINGinstead ofDETACHED. park_detached_threadsseesDETACHED_WAKING→ sets_PY_EVAL_PLEASE_STOP_BITon the eval breaker without parking. The thread gets a chance to attach first, then suspends cooperatively via_Py_HandlePending().tstate_try_attachnow acceptsDETACHED_WAKING→ATTACHEDin addition toDETACHED→ATTACHED, soPy_END_ALLOW_THREADSsucceeds immediately after a wakeup.
Side: Based on the description, this issue seems to occur only on HPC systems. I haven’t had a chance to compile it on an HPC system yet, but I haven’t encountered this problem on the M chip. I’ll try to reproduce the issue on a machine with a 128-core CPU and an RTX-6000, and I’ll report back to you later.
- start_the_world transitions
Thanks — yes, that is the release/re-park race I am seeing. I considered
representing the wakeup entirely in the thread-state value.The complication with mapping every resumed
SUSPENDEDtstate to
DETACHED_WAKINGis that some tstates remain detached because their OS thread
is still blocked in sleep or I/O and is not yet trying to attach. If the next
STW refuses to park one of those tstates, it can wait until that operation
returns; an eval-breaker bit cannot run while the thread remains detached.PR #152826 instead distinguishes tstates parked from
DETACHED, then marks a
waiter only whentstate_wait_attach()actually observes that state. A later
STW skips only active attach waiters. Passive detached tstates remain
parkable, and the normal uncontended attach path remains the existing CAS. A
state-only representation of an active waiter could also work, but it would
need another state and transition rather than marking every woken tstate.Reacted by Jiucheng(Oliver)- added 13 commits that reference this issue
on Sep 20, 2026 - added a commit that references this issue
on Sep 28, 2026
Bug report
While hardening free-threaded Python support in
numba-cuda, we found a workload that could wedge when a stress worker repeatedly calledgc.collect()in a tight loop while other threads were doing CUDA dispatch work. Removing only the tight manual-GC worker made the workload pass.I reduced this to a no-third-party CPython reproducer. The reduced case is strongest on a debug free-threaded CPython
mainbuild, where the extra free-threaded GC validation work makes the stop-the-world window long enough to reproduce reliably.Build used for the reproducer:
Observed on CPython
maindebug/free-threaded:I reproduced this on a B200 Linux system with 224 CPUs at the host level. I have not yet reproduced it on a smaller local workstation. The test process was running under a Slurm allocation that reported 28 CPUs available to the process and 224 CPUs on the system.
The script regularly times out via the faulthandler watchdog. The Python-level traceback shows the GC worker in
gc.collect()and the main thread apparently still attime.sleep(). A native stack shows the more precise state:At the time of the stall, GDB showed the interpreter stop-the-world state as:
The suspected issue is that after
start_the_world()moves suspended threads back to detached and unparks them, a thread running a tightgc.collect()loop can immediately request another stop-the-world pause and re-suspend a just-unparked thread before that thread gets to attach. The result is starvation of a thread trying to return from a detached operation such astime.sleep().This was not reproduced in my minimized matrix on:
mainbuildHowever, the original larger stress workload was first encountered while testing Python 3.14t free-threaded package support, and the reduced CPython
maindebug build gives a clean way to expose the progress bug.I have a local CPython patch that tracks thread states waiting to attach after a stop-the-world suspension and prevents a subsequent stop-the-world requester from immediately re-suspending those attach waiters. With that patch:
test_free_threading.test_gcpasses under debug and non-debug free-threaded buildstest_free_threading.test_gc test_threading -vpasses under the debug free-threaded buildI will prepare a PR with the fix and regression test once this issue exists.
CPython versions tested on:
Operating systems tested on:
Linked PRs