test(net): preserve control event progress after stale listener readiness - #2319
Merged
fslongjin merged 1 commit intoSep 22, 2026
Conversation
…ness Add deterministic accept and accept4 regression coverage for readiness consumed through a duplicated listener. Enable listener nonblocking mode with FIONBIO, observe EPOLLIN, accept the sole connection through an alias, and require EAGAIN when using the stale readiness before processing a queued control message. This protects the event-loop contract implicated in nginx graceful shutdown hangs. The runtime mode synchronization fix is already merged in DragonOS-Community#2317. Reuse the existing socket ownership helper and bounded child watchdog; no kernel changes, threads, sleeps, new dependencies, or nginx-specific behavior are introduced. Validation: 21 socket_nonblocking cases pass on Linux and DragonOS; the two new cases detect blocking with only the previous FIONBIO mode divergence restored. 64 focused guest cases pass, as do ten default nginx start/quit cycles and graceful completion of an in-flight proxy request. Kernel build, formatting, and independent design/code review pass. Additional validation limitation: all three existing tcp_accept_handshake cases fail to observe SYN-ACK on the unchanged master kernel, including a fresh guest before nginx or the new tests run. That separate baseline failure remains unresolved and is not counted as passing coverage. Signed-off-by: longjin <longjin@dragonos.org>
Member
Author
|
@codex review |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add two deterministic dunitest regressions ensuring that consumed listener readiness cannot block an event-driven worker from processing its control channel.
The nginx graceful shutdown hang is addressed by the socket-mode synchronization fix already merged in #2317. Removing only the synchronization in File::set_flags reproduces the hang: the master receives SIGQUIT, one worker exits, and another remains blocked in accept4. This PR protects that contract without adding another kernel fix or nginx-specific workaround.
Test design
Reuse existing FD ownership and bounded child-process helpers. No racing threads, sleeps, new dependencies, or application framework are needed. The accept4 flag still controls only the new connection, not whether the listener waits.
Validation
Known validation limitation
The three existing tcp_accept_handshake cases fail to observe SYN-ACK in this environment, with synthetic-peer TTL-exceeded diagnostics. Rebuilding the suite and running it first in a fresh snapshot with the unchanged master kernel reproduces all three failures, before nginx or the newly added tests run. That separate baseline/fixture issue is unresolved; this PR does not claim the complete network suite passes and does not alter those tests or routes to hide the failure.