[LXC] RHEL setup is now correctly documented and supported in CI - #1316
Merged
Merged
Conversation
Containers on RHEL 10 came up with an IPv6 address but never an IPv4 lease, so every LXC network test failed its readiness gate. The image runs firewalld with an nftables backend, whose default zone permits router advertisement but rejects IPv4 DHCP on the bridge. lxc-net's own accept rules cannot override that: nftables evaluates every base chain registered at a hook, so a reject in the firewalld chain still applies. Host preparation now assigns the bridge to the trusted zone at runtime, and the LXC guide records the same step for developer machines. The NAT check alongside it derived the subnet from the bridge's host address and looked for a POSTROUTING rule matching 10.0.3.1, which lxc-net never writes -- it matches the network. The check therefore never found the existing rule and appended a duplicate on every run. Separately, the peer-backed network tests started a listener, slept a second, and treated `kill -0` as proof it was serving. A background child that has already exited remains a zombie, so that check succeeds on a listener that never bound, and the fixed sleep turns interpreter start-up on a loaded host into a test failure. They now poll the socket and report what the listener wrote when it stays unreachable. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Every container run logged six iptables failures while installing its inbound chain and two more while removing it, on runs that succeeded. Both sets are the designed outcome. A fresh container has no chain to reset, so the pre-install reset is expected to report one missing, and teardown drains INPUT references until iptables reports none left -- that final error is how the drain knows it is done. Only the caller can tell those apart from a real error, but the runner logged every non-zero exit before the caller ever classified it, so a genuine teardown failure read exactly like the eight routine ones surrounding it. The runner is now silent and each caller logs what it judges to be a failure. A reset that does remove something says so, because a chain present there escaped an earlier teardown and is worth reporting. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The diagnostics block, the standalone workflow and the matrix entry existed only to reproduce the firewalld bridge failure in CI. The fix is verified, so the matrix returns to its committed state. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
lxc-net installs masquerade for the bridge subnet as part of bringing the bridge up, choosing nft or iptables at startup. On a host where nft is usable it writes its own table, which iptables -t nat -S POSTROUTING does not list, so the probe here concluded NAT was absent and appended a rule every run. start_lxc_bridge only ever starts lxc-net, so the bridge cannot come up without the accompanying NAT rule, and the check has nothing left to cover. Document the Red Hat setup the CI host performs, so the firewalld zone step and the EPEL prerequisite are discoverable outside the script. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Restores the standalone workflow and the matrix entry, and reports the bridge NAT state so the run shows which backend lxc-net used. Revert before merging. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Widens the enabled plan to Debian 13 and both Ubuntu pools alongside RHEL 10, so the NAT change is exercised where lxc-net picks iptables as well as where it picks nft. Revert before merging. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The standalone workflow, the enabled matrix entries and the bridge NAT diagnostics existed only to exercise the LXC lane while the fix was in progress. The matrix returns to its committed state. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
Soham Das (SohamDas2021)
approved these changes
Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
📖 Description
This PR updates documentation and CI image setup for RHEL (and similar distros). After investigating LXC failures on RHEL systems, it became clear that
firewalldwasn't properly configured withlxcbr0in a trusted zone. CI now uses the trusted zone, while documentation recommends picking any zone as long aslxcbr0has the desired access.Other changes:
tests/scripts/lib/lxc_peer_listener.sh.🔗 References
Resolves #1274
🔍 Validation
Validated in CI: https://github.com/microsoft/mxc/actions/runs/36478382638/job/109119872066
✅ Checklist
Cargo.lock, thedependency-feed-checkcheck passes (see docs/pull-requests.md)📋 Issue Type
Microsoft Reviewers: Open in CodeFlow