Repository navigation
feat(net): implement nftables core for Docker firewall - #2367
Conversation
Add transactional nftables tables, chains, rules, sets, verdict maps, netlink dumps and notifications with per-network-namespace ownership. Evaluate IPv4, IPv6 and inet rules at ingress, forward and output hooks, including conntrack, defragmentation, NAT and the xt compatibility expressions used by iptables-nft. Carry routed/deferred packet context through smoltcp and pin the corresponding dependency PR revision. Add guest coverage for nft netlink operations, interval sets, packet filtering, bridge forwarding and Docker firewall initialization. Verified with make kernel, make fmt, 106/106 guest netfilter tests, host Linux differential checks and smoltcp CI. Docker still requires dynamic bridge/veth creation (DKC-036); optional Netfilter extensions are out of scope. Signed-off-by: longjin <longjin@dragonos.org>
|
@codex review |
Keep NAPI runnable for LocalOnly traffic after an administrative link-down while leaving virtio physical interrupts masked. This restores delivery to local addresses retained on down interfaces. Run IPv4 local output through the filtered ingress path, advance queued local UDP output after placement and routing locks are released, and require interface group membership before multicast UDP fanout. Use the membership operation resolved ifindex for both interface references and the UDP registry so implicit leave cannot desynchronize them. Add guest regressions for down-link raw UDP delivery and multicast join/leave behavior. Verified the full netlink/netfilter and route semantics suites, the IPv4 UDP loopback gVisor suites, formatting, and kernel build. Signed-off-by: longjin <longjin@dragonos.org>
|
CI failure follow-up (fb79997): the nftables local-output path exposed two existing delivery assumptions. Administrative link-down left NAPI paused even though local routes remain valid, and prepared IPv4 UDP output could return before its local receive path had an opportunity to run. The patch resumes LocalOnly NAPI after the link-state transition, keeps physical virtio interrupts masked while down, runs locally generated IPv4 through the filtered ingress path, and gives queued local UDP output a bounded opportunity to advance. The multicast ingress path now requires membership on the receiving interface; join/leave and the UDP registry use the same resolved interface index. Guest validation: netlink/netfilter 107/107, route semantics 13/13, IPv4 UDP loopback gVisor 86 passed (4 skipped), loopback netlink 4/4, and 100 repeated IP_MULTICAST IP_PKTINFO cases. make fmt and make kernel pass. The new regression tests cover DOWN-interface raw UDP delivery and multicast membership transitions. This is opportunistic output progress, not a guarantee that a separate nonblocking UDP receiver is immediately readable after sendto. I also completed an independent adversarial review of the final diff; no remaining blocking issues were found. |
|
@codex review |
After draining routed local output, recheck the local input queue before reporting the interface idle. The earlier poll_again snapshot can predate enqueueing the looped-back packet, allowing nonblocking UDP receives to observe EAGAIN even after a successful send. Split nftables by packet context, object model, rule evaluation, hook execution, transaction updates, and tests. Move conntrack to its conventional mod.rs layout without changing tracking behavior. Preserve the internal nftables API and existing test cases. Verified with make fmt, make kernel, and the two formerly failing gVisor nonblocking UDP cases repeated 20 times each in the DragonOS guest; their blocking variants also pass. Signed-off-by: longjin <longjin@dragonos.org>
|
Updated this PR with commit ac706ad. The failed Integration Test contained two SOCK_NONBLOCK UDP loopback cases that returned EAGAIN after a successful send. The local-output drain could enqueue an ingress packet after poll_again was sampled, while poll() still reported idle. The fix checks the local input queue after draining, matching the existing NAPI progress condition. I also split nftables.rs by responsibility and moved conntrack.rs into conntrack/mod.rs without changing firewall or tracking behavior. Validation: make fmt and make kernel passed; each previously failing guest gVisor case passed 20 repeated runs, and both blocking variants passed. The new CI run remains the final integration check. |
|
@codex review |
Use DragonOS-Community/smoltcp at the squash-merge commit from PR DragonOS-Community#37 instead of the temporary fork revision. The merged commit and the previous PR-head revision have identical Git trees, so this changes provenance without changing dependency source content. Verified Cargo resolves the locked revision from the community repository and make kernel completes successfully. Signed-off-by: longjin <longjin@dragonos.org>
|
smoltcp PR #37 has merged into dragonos/v0.12.0. This PR now pins DragonOS-Community/smoltcp at the merged commit 4debe658d878193b38c436ec12017a83adbf542c instead of the temporary fork revision. The merged commit and the previous PR-head revision have identical Git trees. Cargo locked-resolution and make kernel both pass. All checks on the preceding DragonOS commit passed; the new dependency-pin commit is awaiting CI. |
|
@codex review |
Summary
Implement the nftables core ABI and the iptables-nft rule subset required for Docker firewall initialization. Rules are transactional and per network namespace; IPv4/IPv6/inet filtering, sets/verdict maps, conntrack, defragmentation, NAT and compatibility expressions act on real packet paths.
This depends on DragonOS-Community/smoltcp#37 and pins its current PR commit. The dependency URL/revision should be updated to the merged Community commit when #37 is merged.
Verification
Scope and limitations
This implements the nftables core and Docker-used firewall behavior, not every optional Linux Netfilter extension. Unsupported operations fail explicitly. Dynamic bridge/veth creation is a separate DKC-036 issue; this PR does not claim that default docker run succeeds yet.