Problem
The prior-branch artifact lookup in ci/tools/lookup-run-id (hardened by #2643) has a known-latent gap that surfaced repeatedly today (2026-09-30): gh run list --branch 12.9.x --workflow CI --status success --limit 100 intermittently returns a bounded candidate set that omits the newest runs. When every returned candidate is older than the 90-day GitHub Actions artifact retention, all lookups fail with:
Error: No successful 'CI' run on branch '12.9.x' has all required artifacts
Ralf explicitly deferred this class in #2643:
This intentionally does not eliminate the possibility that GitHub's server-side branch filter returns a stale or incomplete bounded candidate set; that broader lookup redesign remains outside #2642.
Evidence
The 12.9.x branch has four fresh successful CI runs (all carrying the required cuda-bindings-python*-cuda12.9.1-* artifacts, unexpired):
$ gh api "repos/NVIDIA/cuda-python/actions/workflows/ci.yml/runs?branch=12.9.x&status=success&per_page=5" \
--jq '.workflow_runs[] | [.id, .created_at]'
[35797318627, "2026-09-22T23:25:12Z"]
[35362127153, "2026-09-18T15:23:32Z"]
[35166820782, "2026-09-17T00:30:14Z"]
[34881968601, "2026-09-14T18:37:49Z"]
[31739833756, "2026-08-13T20:14:11Z"]
Three CI jobs today independently invoked lookup-run-id and each received a different stale slice, none containing any 2026-09 run:
Sibling matrix rows in the same runs made the same lookup within seconds and mostly got fresher slices, so this manifests as a per-row flake. #2965's py3.10 build passed on run_attempt: 2 with no code change, confirming the flake is API-side.
Root cause
gh run list --branch --workflow --status success --limit 100 occasionally returns a bounded set that omits the true newest runs on the branch. The tool trusts this list as authoritative; when the entire returned slice is older than the 90-day artifact retention, every candidate fails the artifact-freshness check and the step errors out. This is the residual API-caching risk Ralf called out in #2643.
Possible mitigations
- Cross-check
gh run list against a direct REST call to /repos/{r}/actions/workflows/ci.yml/runs?branch=X&status=success&per_page=100, union the results, then sort. A local replay of this endpoint today consistently returned all 2026-09 runs even while the CI's gh run list did not.
- Look up the branch HEAD SHA (
/repos/{r}/branches/12.9.x) first and try the run for that SHA before falling back to the list.
- Retry
gh run list once with a short delay after exhausting candidates; the stale slice usually clears on a second call.
- On final failure, emit a diagnostic that dumps the returned slice and the branch HEAD, so a reviewer doesn't have to reproduce the API state to see what happened.
-- Leo's bot
Problem
The prior-branch artifact lookup in
ci/tools/lookup-run-id(hardened by #2643) has a known-latent gap that surfaced repeatedly today (2026-09-30):gh run list --branch 12.9.x --workflow CI --status success --limit 100intermittently returns a bounded candidate set that omits the newest runs. When every returned candidate is older than the 90-day GitHub Actions artifact retention, all lookups fail with:Ralf explicitly deferred this class in #2643:
Evidence
The
12.9.xbranch has four fresh successfulCIruns (all carrying the requiredcuda-bindings-python*-cuda12.9.1-*artifacts, unexpired):Three CI jobs today independently invoked
lookup-run-idand each received a different stale slice, none containing any 2026-09 run:aarch64 / py3.11buildlinux-64 / py3.10buildlinux-64 / py3.14ttestSibling matrix rows in the same runs made the same lookup within seconds and mostly got fresher slices, so this manifests as a per-row flake. #2965's
py3.10build passed onrun_attempt: 2with no code change, confirming the flake is API-side.Root cause
gh run list --branch --workflow --status success --limit 100occasionally returns a bounded set that omits the true newest runs on the branch. The tool trusts this list as authoritative; when the entire returned slice is older than the 90-day artifact retention, every candidate fails the artifact-freshness check and the step errors out. This is the residual API-caching risk Ralf called out in #2643.Possible mitigations
gh run listagainst a direct REST call to/repos/{r}/actions/workflows/ci.yml/runs?branch=X&status=success&per_page=100, union the results, then sort. A local replay of this endpoint today consistently returned all 2026-09 runs even while the CI'sgh run listdid not./repos/{r}/branches/12.9.x) first and try the run for that SHA before falling back to the list.gh run listonce with a short delay after exhausting candidates; the stale slice usually clears on a second call.-- Leo's bot