fix(linux): close the DRM render-node fd owned by wl::dmabuf_t - #5764
raphaelpereira wants to merge 2 commits into
Conversation
f793536 to
56711fd
Compare
`dmabuf_t::init_gbm()` opens the render node and hands the descriptor to `gbm_create_device()`; the destructor destroys the GBM device but never closes the descriptor, on the assumption that GBM owns it. It does not: libgbm leaves the fd to the caller (`gbm_device_destroy()` does not close it). Every capture (re)initialisation constructs a new `dmabuf_t` -- one per streaming session start, plus one per output mode change and per encoder validation pass that reaches a real capture -- so the process keeps one open `/dev/dri/renderD*` descriptor per (re)init for its lifetime. Measured on a long-running host (Hyprland, `capture = wlr`, `encoder = nvenc`, NVIDIA 610.57.04, v2026.516.143833; unchanged on master): `/dev/dri` descriptors went 4 -> 8 across three Moonlight connect / disconnect cycles and only a restart of Sunshine returned them. Own the descriptor together with the device in a small RAII type, `wl::gbm_device_t`, whose `reset()` destroys the device and then closes the descriptor (and which closes it when device creation fails); open it with `O_CLOEXEC` so prep/do commands do not inherit it. The open/create/ destroy calls are injectable (`gbm_device_accessors_t`, same shape as `gbm_bo_accessors_t`) so the lifecycle is covered by unit tests that need no GPU: the descriptor is closed on reset, on the failure path and on destruction. Signed-off-by: Raphael Derosso Pereira <raphael@insignis.com.br>
56711fd to
fe49a19
Compare
| }; | ||
| fake_destroyed_devices() = 0; | ||
|
|
||
| // Learn which descriptor the next open() will return, so its fate can be checked afterwards. |
There was a problem hiding this comment.
Please record the descriptor actually opened by the fake accessor. This test closes probe and assumes the next open returns the same number; another thread can claim that number between the calls. The test can then fail intermittently or check a different fd while the render-node fd remains open. Capturing the fd in open_fake_render_node or fail_to_create_device would make the assertion reliable.
There was a problem hiding this comment.
Done in a36b606: open_fake_render_node() now records the descriptor it returns, and the test asserts on that descriptor; the probe open()/close() is gone.
| this->accessors = accessors; | ||
|
|
||
| drm_fd = accessors.open_render_node(render_path.c_str()); | ||
| if (drm_fd < 0) { |
There was a problem hiding this comment.
Please add a test with open_render_node returning -1 and verify that create_device is never called and the owner remains empty. This new early-return path is not covered by the three added tests.
There was a problem hiding this comment.
Added DoesNotCreateDeviceWhenRenderNodeFailsToOpen in a36b606: open_render_node returns -1, the fake create_device accessors count their calls, and the test asserts zero calls, no device, fd() == -1 and nothing destroyed. It fails if the early return is removed.
ClosesRenderNodeDescriptorWhenDeviceCreationFails no longer predicts the descriptor with a probe open()/close(): open_fake_render_node() records the descriptor it returned, and the test checks that one, so another thread claiming the number between the calls cannot make it flaky or check the wrong fd. DoesNotCreateDeviceWhenRenderNodeFailsToOpen covers the early return in gbm_device_t::init(): with open_render_node returning -1, create_device is never called (counted by the fake accessors), and the owner stays empty (no device, fd() == -1, nothing destroyed).
|
|
I combined this PR with #5748 and am running the result on the host from #5810 (Hyprland 0.56.2, NVIDIA RTX 3080 Ti, driver 610.57.04, Results:
Not verified yet:
Is there a rough timeline for reviewing/merging #5748 and #5764? They fix the VRAM growth in #5810, and for now I'm running a local build to avoid it. |



Description
wl::dmabuf_t::init_gbm()opens the render node and hands the descriptor togbm_create_device(); the destructor destroys the GBM device but never closes the descriptor — the comment says "We should close the DRM FD, but it's owned by GBM", which is not the case: libgbm leaves the fd to the caller (gbm_device_destroy()does not close it).Every capture (re)initialisation constructs a new
dmabuf_t— one per streaming session start, one per output mode change, and one per encoder validation pass that reaches a real capture — so the process accumulates one open/dev/dri/renderD*descriptor per (re)init for its whole lifetime.Change
wl::gbm_device_t: owns the render-node descriptor together with the GBM device.init()opens the node (O_CLOEXEC, so prep/do commands do not inherit it) and creates the device, closing the descriptor if creation fails;reset()destroys the device then closes the descriptor; the destructor callsreset().gbm_device_accessors_t— the same shape as the existinggbm_bo_accessors_tfrom fix(linux): export all Wayland DMA-BUF planes #5699 — so the lifecycle is unit-testable without a GPU.dmabuf_tholds agbm_device_tinstead of a rawgbm_device *;init_gbm()and the destructor delegate to it.Tests
tests/unit/platform/linux/test_wayland.cpp,WaylandGbmDeviceTest(fake accessors open/dev/nulland hand out a dummy device pointer):ClosesRenderNodeDescriptorOnReset— afterreset()the descriptor failsfcntl(F_GETFD)and the device was destroyed exactly onceClosesRenderNodeDescriptorWhenDeviceCreationFails—init()returns false, nothing is held, the opened descriptor is closedDestructorClosesRenderNodeDescriptorAgainst the previous code the first and third fail (descriptor still open after teardown).
Observed
Long-running host: Hyprland,
capture = wlr,encoder = nvenc, RTX 5060 Ti,nvidia-open-dkms 610.57.04, Sunshinev2026.516.143833(the code is unchanged onmaster).ls -l /proc/$(pgrep -x sunshine)/fd | grep -c /dev/driaround three Moonlight connect/disconnect cycles (each connect re-inits capture ~3× on this host because a prep command changes the output mode):/dev/drifds held by Sunshine/dev/drifdsOnly a restart of Sunshine returns them. At this host's reconnect pace that is a few hundred descriptors a day.
Related but different: #5023 / #5030 (descriptors leaked by failed imports on the wrong render node) and #5699 (exported plane fds closed on error). Neither closes the render-node fd the capture object itself opens.
Not in this PR
The same host also retains GPU memory across capture re-inits (~20 MiB per session, not attributed to the process by
nvidia-smi, returned by a Sunshine restart). An isolated test shows a leaked DRM fd alone does not pin memory on this driver, so that is a separate defect (encode-session or GL-object teardown); it will follow with its own measurement.Verification
Built on Arch (
cmake -DBUILD_TESTS=ON -DSUNSHINE_ENABLE_WAYLAND=ON -DSUNSHINE_ENABLE_CUDA=ON);test_sunshine --gtest_filter='Wayland*'→ 12 tests, all passed.clang-format --dry-run --Werrorclean on the three files.Type of Change
Checklist