Skip to content

EPIC: VirtualMemoryResource ownership redesign #2906

Description

@Andy-Jost

Status: the fix, #2917, has merged. The hold on VMM pull requests is lifted; see the update comment below.

This epic collects the open VirtualMemoryResource bugs, states the fix, and asks contributors to hold new VMM pull requests until the fix lands.

What is wrong

The twelve issues and seven pull requests listed below report about twenty bugs in VirtualMemoryResource. All of them live in one file of about 640 lines. Most of the serious ones have one cause.

When a buffer is freed, the resource knows only the pointer and size it recorded at allocation. From those two values, the code must undo every driver call that built the buffer, in the right order. It does this by hand, with eight calls spread over three code paths. That works for a plain allocation. It fails after a grow.

A grow extends an existing buffer. If the address range right after the buffer is free, the resource reserves it and maps new memory there (the fast path). If that range is taken, the resource reserves a larger range elsewhere and moves the buffer (the slow path). Either way, a grown buffer owns two address reservations, two physical allocations, and two mappings. The driver frees a reservation only when the pointer and size match one reservation exactly (cuMemAddressFree). So no single deallocate(pointer, size) call can free a grown buffer, whatever size it passes.

The pool-backed resources do not have this problem. Each of their buffers carries a C++ handle that knows how to free itself and holds the handles it depends on.

The fix

The implementation plan is posted in this comment; it will land in the PR as cuda_core/cuda/core/_cpp/rt/VMM_DESIGN.md.

Move VirtualMemoryResource onto the same handle layer (_rt). Each physical allocation, address reservation, and mapping gets its own std::shared_ptr handle, and a Buffer owns its mappings. Teardown order then follows from ownership instead of from hand-paired driver calls. The redesign also decides how the free is ordered on the stream (#2886), whether a subclass's deallocate() still runs (#2615), and whether a grown buffer keeps its pointer and identity.

This fixes #2887, the unaligned size recorded on grown buffers, the second grow that fails, the slow-path failure that cannot be undone, the stream argument that allocate() drops, and the fast path that never runs (#2388 item 2). Owner: @Andy-Jost. Milestone: cuda.core 1.3.0.

Two smaller groups of work go with it:

  1. Fixes that were in flight when the epic opened: cuda.core: close the old buffer when a VMM grow moves the mapping #2880 (spurious warning when a slow-path grow moves the buffer), fix(cuda.core): order VMM unmaps on the stream #2889 (stream sync before unmap), and the failed-reservation part of Avoid freeing failed VMM grow reservations #2237. The redesign PR fix(cuda.core): move VirtualMemoryResource onto the _rt handle layer #2917 replaces the module they edit and carries their fixes, so they are superseded and close when fix(cuda.core): move VirtualMemoryResource onto the _rt handle layer #2917 merges. Fix VMM handle leaks in virtual memory paths #2235, which fixes the physical-allocation leak, has merged.
  2. Option validation, as small follow-up PRs after the redesign: addr_align, shrink requests, MANAGED, host location defaults ([BUG]: VirtualMemoryResource host_numa allocations always fail #2694), the Windows default handle type, an assert that guards user input, size 0, device_id and is_device_accessible for host-located resources, and the config= argument that persists (with Review test_memory.py::test_vmm_allocator_policy_configuration xfail #1300).

Contributors

Thanks to @fallintoplace and @aryanputta for the reports and fixes. #2344 and #2235 found and closed the largest leak, and #2886 found the missing stream sync.

Please hold new VMM pull requests until the redesign lands. For the open PRs, the table below states what happens to each. Changes that try to free a grown buffer through one deallocate(pointer, size) call cannot be merged, for the driver reason above (#2887, #2890). Held PRs stay open. Discussion of the plan belongs on this epic.

On tests: we do not merge tests that monkeypatch driver.* entry points or hand fake Buffer objects to the resource (see the review on #2235). A VMM test allocates against the real driver and checks what it can observe: the change in free memory that cuMemGetInfo reports, the buffer contents after a grow, or the absence of a CUDAWarning. To force the slow path, reserve a decoy range right after the buffer, as the #2917 tests do.

Issues

# Issue Fixed by
#2344 leak after allocate and grow #2235 (merged)
#2877 spurious warning on the slow grow path #2917 (merged)
#2886 unmap without a stream sync #2917 (merged)
#2345 a failed adjacent reservation gets freed #2917 (merged)
#2887 a fast-path grow leaves its extension unfreeable #2917 (merged)
#2388 four defects items 1, 2 and 3 by #2917 (merged); item 4 by #2418 (merged)
#2907 grown buffers record an unaligned size #2917 (merged)
#2908 a second grow of a grown buffer fails #2917 (merged)
#2694 host_numa allocations fail option validation
#2909 modify_allocation(config=) persists on the resource #2917 (merged)
#2910 options and location validation items 1, 2, 4, 5 and the host part of 6 by #2917 (merged); the rest open
#1300 xfail review of the policy test option validation, same question as the config= item

#2882 and #2884 were duplicates of #2344 and #2235 and are closed. The feature requests #2057 (multicast objects) and #2358 (logical endpoints) are not part of this epic; they depend on the redesign, because it changes how VMM buffers own their mappings.

Pull requests

PR Status
#2235 Merged. Every open VMM PR needs a rebase on main to pick up its Transaction.on_exit changes
#2880 Closed: superseded by #2917
#2889 Closed: superseded by #2917
#2237 Closed: superseded by #2917
#2407 Closed: a subset of #2237
#2440 Closed: #2880 removed the code it edits
#2890 Closed: one call cannot free two reservations
#2917 Merged: the redesign

The tables are updated as PRs merge or close.

Activity

  1. added this to the cuda.core 1.3.0 milestone on Sep 17, 2026
  2. self-assigned this
    on Sep 17, 2026
  3. added
    P1Medium priority - Should do
    cuda.coreEverything related to the cuda.core module
    EPICSoul of a release
    on Sep 17, 2026
  4. added theissue type on Sep 17, 2026
  5. Andy-Jost commented on Sep 18, 2026

    @Andy-Jost
    ContributorAuthor

    VirtualMemoryResource on the handle layer: implementation plan

    This is the design for step 2 of #2906. It moves VirtualMemoryResource onto the _rt handle
    layer that the pool-backed resources already use, in one PR targeted at cuda.core 1.3.0. The plan
    will land in the PR as cuda_core/cuda/core/_cpp/rt/VMM_DESIGN.md. Comments are welcome on this
    thread.

    Summary

    The public surface stays: the class names, VirtualMemoryResourceOptions, the method names and
    their docstrings. The implementation changes underneath. Each physical allocation, address
    reservation, and mapping gets its own std::shared_ptr handle with a deleter that knows the exact
    driver call to undo it. A buffer owns a list of mappings. Teardown order then follows from what
    holds what, and a failed multi-step operation unwinds by letting its local handles die. The module
    moves from Python to Cython, because the handles are cdef types.

    Driver behavior the plan relies on

    • Reservation. cuMemAddressReserve(size, align, hint) returns (ptr, size).
      cuMemAddressFree
      succeeds only for the exact (ptr, size) pair of one reservation. Reservations never overlap;
      growing a buffer in place yields two adjacent reservations, each of which is freed separately.
      A hint must be a multiple of max(align, 2 MiB); align == 0 means the default 2 MiB.
    • Physical allocation. cuMemCreate returns a handle with one reference; cuMemRelease drops
      one; cuMemRetainAllocationHandle adds one. Mappings are counted separately. The memory is
      freed when references are zero and no mapping remains. Releasing while mapped is legal.
    • Mapping. cuMemMap(ptr, size, 0, handle) maps the whole allocation at ptr; offset must be
      0 and size must equal the allocation's size. cuMemSetAccess(ptr, size, descs, count) grants
      access per mapped range and rejects count == 0. cuMemUnmap(ptr, size) may cover several
      whole mappings, never part of one. One allocation may be mapped at several addresses at once;
      each mapping has its own access state.
    • Coherence of aliases. Two addresses that map one allocation reach the same physical pages.
      Writes through one are visible through the other in stream order and at kernel boundaries; the
      driver adds no synchronization of its own.
    • Context and synchronization. No VMM entry point needs a current context. cuMemUnmap does
      not synchronize. cuStreamSynchronize is rejected on a capturing stream.

    So a mapping depends on exactly one reservation and one allocation, mappings never depend on
    other mappings, and reservations and allocations are independent of each other. A buffer is a
    list of mappings.

    Handles

    Three new std::shared_ptr aliases into boxes, following the conventions in types.hpp.
    CUmemGenericAllocationHandle and CUdeviceptr are both unsigned long long, so the values are
    wrapped in TaggedHandle<T, N> to keep the overload sets distinct.

    Handle Box Deleter Depends on
    MemAllocationHandle {handle, size, access descriptors} pw_cuMemRelease(handle) nothing
    VaReservationHandle {ptr, size} pw_cuMemAddressFree(ptr, size) nothing
    VaMappingHandle {ptr, size, h_alloc, h_reservation} pw_cuMemUnmap(ptr, size), then the members release allocation, reservation

    The allocation box carries the access descriptors it was created with. A mapping applies its
    allocation's descriptors (and skips the call when there are none), so a chunk keeps its access
    wherever it is mapped.

    MemAllocationHandle create_mem_allocation_handle(size_t size, const CUmemAllocationProp& prop,
                                                     const CUmemAccessDesc* descs, size_t count);
    VaReservationHandle create_va_reservation_handle(size_t size, size_t align, CUdeviceptr hint);
    VaMappingHandle     create_va_mapping_handle(CUdeviceptr ptr, const MemAllocationHandle& h_alloc,
                                                 const VaReservationHandle& h_res);
    size_t mem_allocation_size(...) noexcept;  size_t va_reservation_size(...) noexcept;
    

    Factories return the handle and put the status in thread-local err, as the other factories do;
    an empty input handle sets err too, so an empty result always carries a status.
    create_va_mapping_handle checks that the range lies inside the reservation, maps, and applies
    access; if access fails it unmaps and returns empty. The eight VMM entry points, plus
    cuStreamSynchronize and cuStreamGetCaptureInfo, join the driver_api pointer table.

    The range and the device pointer

    struct VmmRange {                             // one per buffer base address
        std::vector<VaMappingHandle> mappings;    // ascending, contiguous; sum of sizes = range total
        std::mutex mu;                            // guards `streams`; nothing under it takes the GIL
        std::vector<DeallocationStream> streams;  // every stream a dying owner forwarded, deduplicated
    };
    using VmmRangeHandle = std::shared_ptr<VmmRange>;      // deleter: sync each stream, then destroy
    DevicePtrHandle deviceptr_create_vmm(CUdeviceptr base, VmmRangeHandle range);
    VmmRangeHandle  vmm_range(const DevicePtrHandle& h);   // empty for a non-VMM or closed handle
    

    There is exactly one ownership chain: Buffer._h_ptr -> DevicePtrBox (holds the range) ->
    VmmRange -> mappings -> reservations and allocations. The Cython buffer keeps no other
    reference; grow operations call Buffer_check_open and then vmm_range(buf._h_ptr).
    Buffer.close() stays _h_ptr.reset(). A buffer's size is always a prefix of its range.

    The DevicePtrHandle must own the memory because graph memcpy nodes retain buf._h_ptr as an
    opaque owner; a non-owning handle would let a launched graph outlive its buffer. Any number of
    owners may therefore exist at once: aliases from a grow, and graph attachments.

    Deleters:

    • DevicePtrBox, VMM flavor: release the GIL; append this box's recorded DeallocationStream to
      the range under mu (skipping an empty stream and duplicates); release mu; drop the range
      reference. It never blocks.
    • VmmRange: release the GIL; unless the interpreter is finalizing, for each forwarded stream
      check the capture status and skip a capturing stream with one report, otherwise synchronize it
      under its bound context with cleanup_in_context. Then, whether or not the syncs succeeded,
      destroy the mappings. Each mapping unmaps; the reservations free and the allocations release as
      their last references go. Every forwarded stream is synchronized because two aliases may have
      recorded different streams; synchronizing only the last one to die would unmap under work
      queued on the other. This is the first blocking deleter in the layer, and it may run inside
      the deferred-cleanup drain on the main thread, with the GIL released.

    Allocations are shared by two ranges after a grow that moves the buffer. Shared ownership is
    what makes that safe: the allocation is released exactly once, when its last mapping goes.

    The resource

    • VirtualMemoryResourceOptions is unchanged. __init__ rejects location_type="host" with a
      handle type other than None, which the driver rejects, and keeps the RDMA and VMM-support
      checks. The resource gains is_ipc_enabled = False, which Buffer.ipc_descriptor reads.
    • cdef class VirtualMemoryBuffer(Buffer) carries no extra state. It is created with
      Buffer_from_deviceptr_handle(h_ptr, size, self, cls=VirtualMemoryBuffer) and documented in
      api.rst like ManagedBuffer. It overrides close(stream=None) to reject a capturing stream,
      since VMM deallocation is synchronous and cannot be captured.
    • allocate(size, *, stream=None):
      1. size == 0 returns an empty buffer without a driver call, like the other resources.
      2. Build CUmemAllocationProp and the access descriptors from the options; query the
        granularity; align the size.
      3. Create the allocation, the reservation, and the mapping as locals; on any empty handle,
        HANDLE_RETURN(get_last_error()). The locals unwind everything.
      4. Build the range and the device pointer handle.
      5. Record the deallocation stream: the caller's stream if it is a real stream, otherwise the
        legacy default token bound to the device's primary context, as _SynchronousMemoryResource
        does. allocate() therefore never needs a current context. Host-located resources record no
        stream and close without a sync.
      6. Return a VirtualMemoryBuffer whose size is the aligned size.
    • modify_allocation(buf, new_size, config=None):
      • Buffer_check_open(buf); range = vmm_range(buf._h_ptr); an empty range means the buffer
        did not come from this resource: TypeError. cfg = config or self.config governs the new
        chunk only and is not stored on the resource. Let req = align_up(new_size) and total be
        the range total.
      • req <= buf.size: return buf. The buffer already covers the request.
      • buf.size < req <= total: return a new VirtualMemoryBuffer over the same range with size
        req; no driver call. This serves a shorter alias asking for what the range already maps.
      • req > total, in place: probe cuMemAddressReserve(req - total, align=0, hint=base+total).
        If the driver grants the hint, create the new allocation with cfg's descriptors and its
        mapping as locals; mappings.reserve(n+1); create a second DevicePtrHandle on the same
        range, copying the input's recorded deallocation stream; build the new buffer; push_back
        the mapping as the last, non-throwing step. If the driver grants another address, drop the
        reservation and move the buffer instead. If the probe fails, drain the status with
        get_last_error() and move the buffer.
      • req > total, moved: reserve req with addr_align; map every mapping in the range (shared
        allocation handles, their own descriptors) at base_new + offset; create and map the new
        allocation; build the new range (copying the input's recorded stream), handle, and buffer.
      • Both paths return a new buffer and leave the input open. See "Why modify_allocation
        returns a new buffer" below.
      • modify_allocation is not thread-safe with respect to two buffers that share a range; that
        synchronization is the caller's responsibility, as elsewhere in cuda.core.
    • deallocate(ptr, size, *, stream=None) stays for the MemoryResource contract. It serves
      pointers wrapped with Buffer.from_handle(ptr, size, mr=self): synchronize stream if given,
      cuMemUnmap, cuMemAddressFree. It handles one reservation, and the caller must have released
      its own cuMemCreate reference. It is not called for buffers from allocate(), whose ranges
      free themselves. The docstring says so.
    • Removed: the Transaction helper and its undo lists, raise_if_driver_error, the
      driver.cu* calls in the resource, and the cdef public size_t _size declaration in
      _buffer.pxd, whose only user was this module.

    Why modify_allocation returns a new buffer

    The input buffer stays open and aliases the result. The chunks the input already mapped are
    shared by both buffers and are freed when the last of the two closes; the chunk the grow added
    belongs to the result. When the grow happens in place, the two buffers share one range at one
    base, and the shorter one pins the whole range until it closes. Callers who are done with the
    input close it.

    The alternative, growing the input object in place so every holder sees the new size and
    pointer, was set aside:

    • In-place update reaches only holders of the Python object. Holders of the handle, such as
      graph memcpy nodes, DLPack capsules, and IPC descriptors, would keep the old mapping alive but
      see a different address than the buffer reports.
    • An address change should be visible. When the buffer moves, an object that quietly changes
      address turns cached int(buf.handle) values into dangling pointers.
    • Buffer.__hash__ and __eq__ include the size, so in-place growth changes the hash of a live
      object.
    • The compatibility cost is small. The docstring described in-place growth with the pointer
      preserved, but the in-place path has not run in practice, so every real grow returned a new
      buffer and closed the input. No code outside this repository calls modify_allocation; the
      known external users of the resource call allocate() only. The change costs a docstring and
      a release note.
    • Leaving the input open costs the caller one line and gives them a valid, shorter alias, which
      no in-place scheme can offer.

    Behavior changes

    1. modify_allocation returns a new buffer and leaves the one passed in open; the two alias the
      same physical memory. When the driver grants the adjacent range, the pointer is preserved.
      DLPack views into the old buffer stay valid on both paths. The grow-in-place path is live and
      correct ([BUG]: VirtualMemoryResource: four pre-existing defects (grow-rollback access loss, dead fast path, finalizer warnings, handle_type docstring) #2388 item 2, cuda.core: a fast-path VMM grow leaves its extension unfreeable #2887).
    2. Buffer.size after a grow is the aligned total (cuda.core: grown VMM buffers record the requested size, not the aligned size #2907).
    3. modify_allocation(config=) applies to the new chunk only; existing chunks keep the access
      they were created with (cuda.core: VirtualMemoryResource.modify_allocation(config=...) changes the resource's policy #2909, Review test_memory.py::test_vmm_allocator_policy_configuration xfail #1300).
    4. A VMM buffer records its allocation stream, and the last close of an aliased range
      synchronizes every stream its buffers recorded before the unmap (cuda.core: VirtualMemoryResource.deallocate() unmaps without ordering on the stream #2886). An explicit close()
      on a capturing stream raises; a close from garbage collection or from a graph's cleanup skips
      the sync for a capturing stream and reports it.
    5. VirtualMemoryResource.deallocate() is not called when a buffer from allocate() closes, so a
      subclass override does not run, as for the pool-backed resources (Pool-backed MemoryResource buffers bypass overridden deallocate() methods #2615 tracks the general
      question).
    6. modify_allocation accepts only buffers returned by this resource.
    7. location_type="host" requires handle_type=None.
    8. Buffers are freed during interpreter shutdown, without the stream sync.

    Issues this PR closes

    #2887, #2907, #2908, #2909 (with the #1300 test update), #2886 (or preserves #2889 if it lands
    first), #2345 (no reservation is ever freed by hand), #2388 (item 2 here; items 1 and 3 through
    #2880; item 4 done), #2877 (through #2880). From #2910: the misaligned probe, the assert,
    size 0, and the HOST default handle type.

    Left for small follow-ups: #2694 (needs a NUMA id option), MANAGED, the Windows default handle
    type, device_id and is_device_accessible for host locations, HOST_NUMA_CURRENT.

    Failure handling

    • Rollback is RAII: locals die in reverse order, deleters run the pw_* wrappers, and failures
      become CUDAWarning.
    • Every deleter releases the GIL first; the range deleter holds no C++ lock while it synchronizes
      or reports. A failed sync (lost context, capturing stream) is reported and the unmap proceeds;
      the driver needs no context for it, so nothing leaks.
    • Empty handles always carry a status; the in-place probe drains its status before falling back.
    • Factories that allocate are declared except+ in _rt.pxd; deleters only destroy vectors.

    Tests

    Fifteen tests, all against the real driver

    All against the real driver, with provenance markers, and thread_unsafe wherever
    cuMemGetInfo, warning capture, or the context stack is involved.

    1. Leak: allocate and close eight times; the free-memory delta stays below one aligned allocation.
    2. Grow in place, forced: reserve 3s through the bindings and free it, allocate(s) with
      addr_hint at that address, grow by s. Same pointer; result size 2s; input still open with
      size s; contents preserved; no warning. Closing either first changes free memory by zero;
      the last close returns 2s. Skip with a reason if the driver placed the allocation elsewhere.
      Repeat with addr_align = 64 MiB.
    3. Grow that moves, forced with a decoy reservation: new pointer; contents preserved; no warning.
      Closing the input first frees nothing; closing the result first frees exactly the added chunk;
      the last close returns the rest.
    4. Unaligned grow (2 MiB to 3 MiB) closes cleanly. A second grow of the result works. A second
      grow of the shorter alias: to within the range (no driver call, same base) and beyond it
      (remaps every mapping of the range). A request within the aligned size returns the same
      buffer.
    5. modify_allocation on a foreign Buffer raises TypeError.
    6. Close after a memset on the recorded stream does not fault.
    7. Host-located resource with handle_type=None: allocate and close with no current context; the
      default handle type raises ValueError at construction.
    8. config= governs the new chunk only; the old chunk stays writable when the new one is
      read-only; a later allocate() uses the resource's config.
    9. self_access=None with no peers allocates.
    10. Close frees memory while the Python wrapper is still referenced.
    11. Explicit close on a capturing stream raises and leaves the buffer open.
    12. A graph memcpy node retaining the buffer keeps the memory alive across close() and across a
      grow; the exec's destruction frees it.
    13. Two aliases with distinct recorded streams, work enqueued on each, closed in both orders: no
      fault, no warning.
    14. Cross-alias visibility: a write through one alias is visible through the other after a
      stream sync, on both paths.
    15. Shutdown smoke test in a subprocess: return code 0 and empty stderr.

    Files touched

    Approximate line counts per area
    Area Lines (approx.)
    _cpp/rt/virtual_memory.cpp (three boxes, the range, four factories, accessors) 300
    memory.cpp (deviceptr_create_vmm, vmm_range) 50
    driver_api.hpp/.cpp and the _rt.pyx pointer table 35
    types.hpp, api.hpp, internal.hpp, _rt.pxd 140
    _memory/_virtual_memory_resource.pyx (from .py) and .pyi 450 + 80
    tests: replace the mocked tests with the fifteen above -250 / +380
    release note, DESIGN.md handle list, api.rst, VMM_DESIGN.md 60

    The build needs no wiring: build_hooks already globs cuda/core/**/*.pyx, _cpp/rt/*.cpp,
    and every header under _cpp/rt/.

  6. aryanputta commented on Sep 18, 2026

    @aryanputta
    Contributor

    @Andy-Jost Thanks for writing this up. Looking through #2235, #2388, #2880, and #2886, the common issue seems to be that VMM reconstructs ownership from pointer and size after a buffer can contain multiple reservations, mappings, and physical handles.

    If I understand the redesign correctly, the ownership could look like this:

      mapping = Mapping(
          allocation = shared_allocation,
          reservation = exact_reservation,
      )
    
      range.mappings.push_back(mapping)
      buffer = DevicePtr(range)
    

    Cleanup would then follow the ownership graph:

      for stream in range.streams:
          synchronize(stream)
    
      for mapping in range.mappings:
          unmap(mapping.ptr, mapping.size)
          release(mapping.reservation)
          release(mapping.allocation)
    

    For growth, the new mappings can be created and validated first, then added to the range only after mapping and access setup succeeds. That keeps rollback local and preserves the old alias if the grow fails.

    The main invariants seem to be:

    • each reservation is freed with its original pointer and size
    • failed growth leaves the old mapping and access unchanged
    • moved growth preserves existing aliases
    • all recorded streams finish before the last mapping is unmapped
    • graph and alias owners keep the range alive

    Making these invariants explicit in VMM_DESIGN.md could make the implementation and tests easier to review aswell.

  7. Andy-Jost commented on Sep 18, 2026

    @Andy-Jost
    ContributorAuthor

    The redesign is up as #2917. It implements the plan above and closes #2887, #2907, #2908, #2909, #2886 and #2345. It also addresses item 2 of #2388 and the size-0, misaligned-probe and host handle-type parts of #2910. Review there.

  8. Andy-Jost commented on Sep 21, 2026

    @Andy-Jost
    ContributorAuthor

    Follow-up: synchronous frees wait on a stream from a destructor. The VirtualMemoryResource range deleter (_cpp/rt/virtual_memory.cpp) and _SynchronousMemoryResource.deallocate call cuStreamSynchronize on the recorded stream before they free. The wait runs on whichever thread drops the last reference, which could be during GC. Beyond latency it has two failure modes: it deadlocks when the host is part of the stream's dependency chain (a persistent kernel waiting on a host flag, or a buffer released inside a host callback), and it cannot run during a graph capture, so the release skips the wait, unmaps anyway, and warns.

    A stream-ordered release would be sounder: enqueue a host function on each recorded stream that ties into the existing deferred cleanup mechanism.

    An alternative is what PyTorch and RMM do: mark cleanup points with events and reap finished releases on the next call into the resource. This avoids callbacks but makes frees lazy and sometimes requires an explicit flush.

  9. Andy-Jost commented on Sep 21, 2026

    @Andy-Jost
    ContributorAuthor

    Hi @aryanputta thanks for your comments. Your diagnosis is right: every listed bug comes from bookkeeping errors. The common theme is that tracking the destruction recipe as data members in the VMM memory manager is difficult and prone to errors. It also gave rise to complex rollback logic implemented in a Transaction class.

    With the proposed fix, each reservation, allocation, and mapping is a shared_ptr whose deleter holds the exact undo call. For example, a successful cuMemAddressReserve that returns a for size b - a captures its matching cuMemAddressFree(a, b - a) in the same scope, and nothing more needs to be tracked. It guarantees your first invariant.

    Relationships are structural, so teardown and rollback come automatically. A mapping that owns its allocation and reservation guarantees unmap precedes release and free, and a failed CUDA API call rolls back pending work by letting local variables die. This matches your cleanup pseudocode while also handling the additional complication that allocations can be mapped multiple times (so not every unmap() is immediately followed by release()).

    Making these invariants explicit in VMM_DESIGN.md could make the implementation and tests easier to review aswell.

    Good suggestion. I will add this to #2917.

  10. leofang commented on Sep 30, 2026

    @leofang
    Member

    Bumping this epic to P0.

  11. added
    P0High priority - Must do!
    and removed
    P1Medium priority - Should do
    on Sep 30, 2026
  12. Andy-Jost commented on Oct 2, 2026

    @Andy-Jost
    ContributorAuthor

    #2917 has merged, so the hold on VMM pull requests is lifted. Pull requests for the open sub-issues (#2910, #2694, #1300) are welcome; build on the merged module and read cuda_core/cuda/core/_cpp/rt/VMM_DESIGN.md first.

    What landed differs from the plan above in one way. A range is immutable: a grow copies the input's mapping list and builds a new range for its result, so two buffers never share mutable state, and each buffer owns one range and one recorded stream. That removed the stream union in the range deleter, the range mutex and the base-address registry, and it is what makes concurrent grows of aliased buffers safe. modify_allocation always returns a new buffer, so closing the result never closes the input.

    Two notes for follow-up work. The release of a VMM buffer waits on its deallocation stream, because cuMemUnmap is not stream-ordered; #2989 tracks replacing that wait with a host callback into the deferred-cleanup queue, for every synchronous release in cuda.core. The VMM tests no longer compare device-wide free-memory counts, which other processes move; they ask the driver that no mapping and no reservation remain at the addresses a buffer used.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

EPICSoul of a releaseP0High priority - Must do!cuda.coreEverything related to the cuda.core module

Type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions