Skip to content

VRAM leak: orphaned DMA-BUFs sized to physical monitor accumulate while streaming a virtual headless output (wlr capture, NVENC) #5810

Description

@dWacing

Describe the Bug

Streaming from a virtual headless output leaks VRAM as orphaned DMA-BUFs that
are visible in neither nvidia-smi's process list nor any process's fd table.
The leak accumulates continuously (~1.3 GB/h) while a stream is active, and
survives client disconnect — only restarting Sunshine frees it.

Measured on an RTX 5090 32GB over two sessions:

  • ~30 GB VRAM unaccounted for after ~2 days of intermittent streaming
  • 24.9 GB leaked after one overnight session; freed instantly by
    systemctl --user restart sunshine (24893 → 2723 MiB)
  • 1-hour instrumented run: unattributed VRAM grows 1819 → 2570 MiB while
    streaming, stays flat after client disconnect, grows again on reconnect

Key observation: leaked buffers match the physical monitor, not the streamed one

/sys/kernel/debug/dma_buf/bufinfo (as root) shows thousands of leaked
buffers with refcount 2 and no attached devices:

size       flags      mode       count    exp_name  ino       name
39321600   00000002   02080007   00000002 drm       00974892  <none>
  Attached Devices:
  Total 0 devices attached
... (repeated ~800x, plus a second size class)
20971520   00000002   02080007   00000002 drm       00974746  <none>
  • 39321600 bytes = exactly 3840×2560×4 → the physical monitor (BenQ RD280U)
  • 20971520 bytes = 2048×2560×4 (aligned variant)
  • The streamed output is a virtual headless monitor "SUNSHINE" at 1512×982
    (output_name = SUNSHINE), i.e. the leaked buffers are the wrong size for
    the capture target — they appear to be allocated per capture cycle against
    the physical display and never released.

Config

capture = wlr
output_name = SUNSHINE

Hyprland (Wayland). Sunshine uses the zwlr_screencopy_manager_v1 path
("Screencasting with Wayland's protocol" in the log), NVENC encoder.

Environment

Reproduction

  1. Hyprland with one physical 4K monitor + one headless output (e.g.
    hyprctl output create headless SUNSHINE), Sunshine output_name = the
    headless output, capture = wlr
  2. Stream for ~1 hour; periodically measure unattributed VRAM =
    nvidia-smi --query-gpu=memory.used minus the process sum
  3. Disconnect the client — leak does not drop
  4. systemctl --user restart sunshine — leak drops to baseline instantly

Expected behavior

DMA-BUFs used for capture should be released when the frame is consumed; VRAM
usage should stay flat (minus normal encoder buffers) for the session
lifetime and drop on client disconnect.

Activity

  1. added
    aiPR has signs of heavy ai usage (either indicated by user or assumed)
    on Sep 28, 2026
  2. jeffscottward commented on Sep 29, 2026

    @jeffscottward

    Independent reproduction on a physical output, with a likely mechanism reproduced outside Sunshine.

    Environment: Arch Linux, Hyprland 0.56.2, RTX 3080 Ti 12 GB, NVIDIA 610.57.04, Sunshine 2026.830.165455 (f53f0f9), capture = wlr (auto), NVENC H.264, physical output DP-1 2560×1440.

    Symptom: an idle Moonlight session was left open for about 1.5 h. VRAM that nvidia-smi does not attribute to any process grew to 11.8 of 12 GB at 0 % GPU utilization, until NVENC failed with CUDA_ERROR_OUT_OF_MEMORY / InitializeEncoder failed: out of memory. systemctl --user restart of Sunshine freed about 10.5 GB immediately. The idle session had let the output power off (hyprctl monitors -j → "dpmsStatus": false). The log also showed bursts of Couldn't import RGB Image: 0000300C, each followed by five encoder re-creations within about 250 ms.

    Likely mechanism (reproduced with a standalone client; I have not instrumented the running Sunshine process):

    • wlr_t::snapshot() calls dmabuf.listen(), which starts a new capture_output request, on every call.
    • While the output is off, Hyprland never completes the copy (no ready/failed). It does still answer every new request immediately with buffer/linux_dmabuf/buffer_done.
    • So after each 1000 ms timeout, create_and_copy_dmabuf() allocates another full-size GBM buffer object and wl_buffer, overwriting current_bo/current_wl_buffer from the request still in flight.
    • ~dmabuf_t() also never closes the render-node fd passed to gbm_create_device().

    Reproduction without Sunshine: a minimal client that mirrors snapshot(): GBM buffer, linux-dmabuf, screencopy, a 1000 ms dispatch timeout, a new request on every call, and overwritten pointers. Against the powered-off output:

    • Current behaviour: 30 requests → 30 buffer_done → 30 × 14,745,600-byte allocations, 0 ready, 0 failed. Unattributed VRAM rose about 480 MiB in 30 s and dropped back only when the process exited.
    • With no new request while one is in flight: 30 timeouts → 1 allocation, and VRAM stayed flat.

    Both parts already have open PRs:

    Together they should fix this. The guard also matters with plain copy whenever the compositor defers a frame, so it could be landed on its own if the damage change needs more review. I haven't reproduced your headless-output setup. Your leaked buffers are sized to the physical monitor while you stream a headless output, which this mechanism does not obviously explain unless the physical output was also being captured or powered off. I'd treat that as an open question rather than assume the two cases are identical.

  3. jeffscottward commented on Sep 29, 2026

    @jeffscottward

    Follow-up measurement on the same rig, this time with the monitor on. It suggests a second path besides the DPMS-off request pile-up.

    • After restarting Sunshine and before any client connected, unattributed VRAM was about 330 MiB.
    • Moonlight reconnected at 22:49. At 22:56:04, Couldn't import RGB Image: 0000300C appeared 10 times, each followed by Creating encoder [h264_nvenc] (a capture/encoder re-init).
    • Unattributed VRAM then stood at about 1,460 MiB. It stayed flat for the following minutes (1,445 → 1,465 MiB in 1-minute samples), with DP-1 dpmsStatus: true and no further bursts.

    That is roughly ~110 MiB retained per re-init after an EGL dma-buf import failure. This is an estimate from coarse samples and was not traced inside Sunshine. The same 0x300C → 5× re-init bursts appear in my first log at 21:21:58 and 21:58:21.

    The in-flight guard in #5748 would not cover this path. The teardown on that re-init path (encoder hw frames / CUDA-GL registrations / capture objects) seems worth checking.

  4. jeffscottward commented on Sep 29, 2026

    @jeffscottward

    Correction to my previous comment: the "~110 MiB retained per re-init" attribution is not supported by finer samples, so please disregard it.

    Minute samples of unattributed VRAM (nvidia-smi --query-gpu=memory.used minus the process-table sum) on the same rig, with a client connected:

    • 23:02–23:06: +3.5 GiB at about 950 MiB/min. Sunshine logged nothing between 23:00 and 23:05:15. At 23:05:16 it logged 32 Couldn't import RGB Image: 0000300C → h264_nvenc re-inits, and the growth stopped shortly after.
    • 23:14–23:17: another +2.3 GiB at the same rate. It ended with 6 more import-failure re-inits and the client disconnecting at 23:17.
    • The memory stayed allocated after the disconnect (about 7.2 GiB unattributed). systemctl --user restart of Sunshine freed about 6.6 GiB.

    Measured growth was 943–965 MiB/min. One 2560×1440 XRGB buffer per second (pitch 10240 × 1440 = 14,745,600 B) would be about 844 MiB/min, so the growth is roughly 13% higher. It is on the order of one frame-sized allocation per second. That is consistent with a capture that stops completing while a new request is issued every second (the path in my first comment), but I haven't measured the per-request allocation, so it is not an exact match. In both windows the growth started before the re-inits (about 23:02 vs 23:05:16, and about 23:14–23:15 vs 23:17), so the re-inits are not what starts it. I did not log DPMS state or the reason frames stopped arriving; the monitor reported on when I checked at 23:08.

  5. jeffscottward commented on Oct 11, 2026

    @jeffscottward

    Follow-up with a longer run of the local build (f53f0f9 + the #5748 in-flight guard + #5764).

    Observed (omarchy-rig, Hyprland 0.56.2, RTX 3080 Ti, driver 610.57.04, capture = wlr, NVENC):

    • Sunshine has run without a restart since 2026-10-05 21:36 EDT (about 5 days), with 11 client connections in its log.
    • Total GPU memory used is 2.0 GiB of 12 GiB; Sunshine's own process shows 258 MiB. Before the patch, an idle session grew to 11.8 GiB until NVENC failed with CUDA out-of-memory.
    • No CUDA out-of-memory or Couldn't import RGB Image errors in the log over this period.

    Limits: I did not confirm that DP-1 was DPMS-off while a client was connected during this period, so this run shows "no leak observed", not a controlled reproduction of the trigger.

    Question about #5883: its next_screencopy_request(pending, …) documentation says asking again on every timeout would pile requests up in the compositor, which matches the mechanism here. Does that pending-request check also apply when capture is not event-driven (output refresh below 1.5× the stream rate, e.g. a 60 Hz output streamed at 60 fps), or only to the copy_with_damage path? If it covers both, #5883 plus #5764 should replace the local guard, and I can test a build from master on this machine with DP-1 powered off.

    — Posted by Claude Opus 5.5 (OMP coding agent) on behalf of @jeffscottward

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    aiPR has signs of heavy ai usage (either indicated by user or assumed)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions