Repository navigation
VRAM leak: orphaned DMA-BUFs sized to physical monitor accumulate while streaming a virtual headless output (wlr capture, NVENC) #5810
Description
Activity
- addedaiPR has signs of heavy ai usage (either indicated by user or assumed)PR has signs of heavy ai usage (either indicated by user or assumed)
on Sep 28, 2026 Independent reproduction on a physical output, with a likely mechanism reproduced outside Sunshine.
Environment: Arch Linux, Hyprland 0.56.2, RTX 3080 Ti 12 GB, NVIDIA 610.57.04, Sunshine 2026.830.165455 (f53f0f9),
capture= wlr (auto), NVENC H.264, physical output DP-1 2560×1440.Symptom: an idle Moonlight session was left open for about 1.5 h. VRAM that
nvidia-smidoes not attribute to any process grew to 11.8 of 12 GB at 0 % GPU utilization, until NVENC failed withCUDA_ERROR_OUT_OF_MEMORY/InitializeEncoder failed: out of memory.systemctl --user restartof Sunshine freed about 10.5 GB immediately. The idle session had let the output power off (hyprctl monitors -j→"dpmsStatus": false). The log also showed bursts ofCouldn't import RGB Image: 0000300C, each followed by five encoder re-creations within about 250 ms.Likely mechanism (reproduced with a standalone client; I have not instrumented the running Sunshine process):
wlr_t::snapshot()callsdmabuf.listen(), which starts a newcapture_outputrequest, on every call.- While the output is off, Hyprland never completes the copy (no
ready/failed). It does still answer every new request immediately withbuffer/linux_dmabuf/buffer_done. - So after each 1000 ms timeout,
create_and_copy_dmabuf()allocates another full-size GBM buffer object andwl_buffer, overwritingcurrent_bo/current_wl_bufferfrom the request still in flight. ~dmabuf_t()also never closes the render-node fd passed togbm_create_device().
Reproduction without Sunshine: a minimal client that mirrors
snapshot(): GBM buffer, linux-dmabuf, screencopy, a 1000 ms dispatch timeout, a new request on every call, and overwritten pointers. Against the powered-off output:- Current behaviour: 30 requests → 30
buffer_done→ 30 × 14,745,600-byte allocations, 0ready, 0failed. Unattributed VRAM rose about 480 MiB in 30 s and dropped back only when the process exited. - With no new request while one is in flight: 30 timeouts → 1 allocation, and VRAM stayed flat.
Both parts already have open PRs:
- perf(linux): capture wlr screencopy frames with damage #5748 adds exactly that guard (
should_request_frame()), bundled with the switch tocopy_with_damage. - fix(linux): close the DRM render-node fd owned by wl::dmabuf_t #5764 closes the render-node fd.
Together they should fix this. The guard also matters with plain
copywhenever the compositor defers a frame, so it could be landed on its own if the damage change needs more review. I haven't reproduced your headless-output setup. Your leaked buffers are sized to the physical monitor while you stream a headless output, which this mechanism does not obviously explain unless the physical output was also being captured or powered off. I'd treat that as an open question rather than assume the two cases are identical.Follow-up measurement on the same rig, this time with the monitor on. It suggests a second path besides the DPMS-off request pile-up.
- After restarting Sunshine and before any client connected, unattributed VRAM was about 330 MiB.
- Moonlight reconnected at 22:49. At 22:56:04,
Couldn't import RGB Image: 0000300Cappeared 10 times, each followed byCreating encoder [h264_nvenc](a capture/encoder re-init). - Unattributed VRAM then stood at about 1,460 MiB. It stayed flat for the following minutes (1,445 → 1,465 MiB in 1-minute samples), with DP-1
dpmsStatus: trueand no further bursts.
That is roughly ~110 MiB retained per re-init after an EGL dma-buf import failure. This is an estimate from coarse samples and was not traced inside Sunshine. The same 0x300C → 5× re-init bursts appear in my first log at 21:21:58 and 21:58:21.
The in-flight guard in #5748 would not cover this path. The teardown on that re-init path (encoder hw frames / CUDA-GL registrations / capture objects) seems worth checking.
Correction to my previous comment: the "~110 MiB retained per re-init" attribution is not supported by finer samples, so please disregard it.
Minute samples of unattributed VRAM (
nvidia-smi --query-gpu=memory.usedminus the process-table sum) on the same rig, with a client connected:- 23:02–23:06: +3.5 GiB at about 950 MiB/min. Sunshine logged nothing between 23:00 and 23:05:15. At 23:05:16 it logged 32
Couldn't import RGB Image: 0000300C→h264_nvencre-inits, and the growth stopped shortly after. - 23:14–23:17: another +2.3 GiB at the same rate. It ended with 6 more import-failure re-inits and the client disconnecting at 23:17.
- The memory stayed allocated after the disconnect (about 7.2 GiB unattributed).
systemctl --user restartof Sunshine freed about 6.6 GiB.
Measured growth was 943–965 MiB/min. One 2560×1440 XRGB buffer per second (pitch 10240 × 1440 = 14,745,600 B) would be about 844 MiB/min, so the growth is roughly 13% higher. It is on the order of one frame-sized allocation per second. That is consistent with a capture that stops completing while a new request is issued every second (the path in my first comment), but I haven't measured the per-request allocation, so it is not an exact match. In both windows the growth started before the re-inits (about 23:02 vs 23:05:16, and about 23:14–23:15 vs 23:17), so the re-inits are not what starts it. I did not log DPMS state or the reason frames stopped arriving; the monitor reported on when I checked at 23:08.
- 23:02–23:06: +3.5 GiB at about 950 MiB/min. Sunshine logged nothing between 23:00 and 23:05:15. At 23:05:16 it logged 32
Follow-up with a longer run of the local build (f53f0f9 + the #5748 in-flight guard + #5764).
Observed (omarchy-rig, Hyprland 0.56.2, RTX 3080 Ti, driver 610.57.04,
capture = wlr, NVENC):- Sunshine has run without a restart since 2026-10-05 21:36 EDT (about 5 days), with 11 client connections in its log.
- Total GPU memory used is 2.0 GiB of 12 GiB; Sunshine's own process shows 258 MiB. Before the patch, an idle session grew to 11.8 GiB until NVENC failed with CUDA out-of-memory.
- No CUDA out-of-memory or
Couldn't import RGB Imageerrors in the log over this period.
Limits: I did not confirm that DP-1 was DPMS-off while a client was connected during this period, so this run shows "no leak observed", not a controlled reproduction of the trigger.
Question about #5883: its
next_screencopy_request(pending, …)documentation says asking again on every timeout would pile requests up in the compositor, which matches the mechanism here. Does that pending-request check also apply when capture is not event-driven (output refresh below 1.5× the stream rate, e.g. a 60 Hz output streamed at 60 fps), or only to thecopy_with_damagepath? If it covers both, #5883 plus #5764 should replace the local guard, and I can test a build from master on this machine with DP-1 powered off.— Posted by Claude Opus 5.5 (OMP coding agent) on behalf of @jeffscottward
Describe the Bug
Streaming from a virtual headless output leaks VRAM as orphaned DMA-BUFs that
are visible in neither
nvidia-smi's process list nor any process's fd table.The leak accumulates continuously (~1.3 GB/h) while a stream is active, and
survives client disconnect — only restarting Sunshine frees it.
Measured on an RTX 5090 32GB over two sessions:
systemctl --user restart sunshine(24893 → 2723 MiB)streaming, stays flat after client disconnect, grows again on reconnect
Key observation: leaked buffers match the physical monitor, not the streamed one
/sys/kernel/debug/dma_buf/bufinfo(as root) shows thousands of leakedbuffers with refcount 2 and no attached devices:
(
output_name = SUNSHINE), i.e. the leaked buffers are the wrong size forthe capture target — they appear to be allocated per capture cycle against
the physical display and never released.
Config
Hyprland (Wayland). Sunshine uses the
zwlr_screencopy_manager_v1path("Screencasting with Wayland's protocol" in the log), NVENC encoder.
Environment
(includes fix(linux/kmsgrab): fix handle leak in update_cursor #4757, fix(linux/wlr): Fix dmabuf buffer params protocol violation/leak #4588, wlr-screencopy capture targets wrong render node on multi-GPU NVIDIA systems, causing FD leak and VRAM exhaustion via failed dmabuf imports #5023 — behavior persists)
the leak is only visible via dma_buf/bufinfo and by diffing
nvidia-smi --query-gpu=memory.usedagainst the attributed process sum.Reproduction
hyprctl output create headless SUNSHINE), Sunshineoutput_name= theheadless output, capture = wlr
nvidia-smi --query-gpu=memory.usedminus the process sumsystemctl --user restart sunshine— leak drops to baseline instantlyExpected behavior
DMA-BUFs used for capture should be released when the frame is consumed; VRAM
usage should stay flat (minus normal encoder buffers) for the session
lifetime and drop on client disconnect.