Skip to content

conn, device, tun: connected UDP sockets, direct receive and batched utun I/O (off by default) - #107

Open
keeleysam wants to merge 6 commits into
tailscale:tailscalefrom
keeleysam:connected-sockets
Open

keeleysam wants to merge 6 commits into
tailscale:tailscalefrom
keeleysam:connected-sockets

Conversation

@keeleysam

Copy link
Copy Markdown

Reference code for tailscale/tailscale#21602, which has the measurements and background. Nothing here changes behaviour unless a caller asks for it. The commits are closely related, but each builds and passes tests on its own, so they can be split out or taken separately:

  1. internal/darwinbatch, tun: batched utun I/O through libSystem's sendmsg_x/recvmsg_x, with a loopback self-test before first use. Building with ts_omit_darwin_spi removes every reference to the private calls.
  2. conn: ConnectedSockets, one connected socket per (local, remote) address pair once it has carried 1 MiB, opened from either direction, up to 256, sharing the bind's port. ConnectedDeliveryCheck leaves it off on kernels that do not deliver to the connected socket first (Linux before 5.4). Usable by any Bind.
  3. conn: StdNetBind uses it with WithConnectedSockets(true).
  4. conn, device: ReceiveFuncStarter, so a bind can start a receive routine per socket after Open, reading straight into the device's buffers (capped at 32 MiB).
  5. conn: WithDontFragment, for darwin.
  6. conn: WithBatchedIO, batched I/O on darwin's shared unconnected socket, with its own self-test.

The magicsock side is tailscale/tailscale#21603.

Tested: gofmt, vet and build on 12 targets per commit, darwin and Linux race (1,390 pass), and a build with ts_omit_darwin_spi checked with nm and strings for the private call names.

Updates tailscale/tailscale#21602

…vmsg_x

Darwin has no recvmmsg or sendmmsg, so the utun is read and written one packet per syscall and NativeTun.BatchSize is 1, which makes the whole device pipeline carry one packet per handoff. This adds internal/darwinbatch, a small wrapper around darwin's batched calls, sendmsg_x and recvmsg_x, and uses it for utun reads and writes.

The utun queues only one packet for userspace by default, which leaves a batched read nothing to collect, so the queue is raised to 512. That happens in CreateTUNFromFile, so a utun handed over as a file (a network extension's) gets it too.

The calls are reached through libSystem's exported wrappers without cgo, the way x/sys/unix calls libc. struct msghdr_x is in no public header, so the first use runs a loopback self-test, and a kernel that fails it gets one packet per syscall as before. Building with -tags ts_omit_darwin_spi leaves no reference to either call in the binary, for app review.

Kernel behaviour handled:

- A batched utun write of a packet larger than one page stalls (about 3 ms per call on Intel), so such packets go through write(2).
- recvmsg_x reads each message header back in, so lengths and flags are restored before a slot is reused; macOS 12 sets MSG_TRUNC and then fails the slot with EINVAL until it is cleared.
- The MTU that sizes read slots is cached, so a packet that fills its slot is checked against its IP header length and dropped if it was cut short.
- A short sendmsg_x continues with the remaining packets, since the device ignores Write's count.

Updates tailscale/tailscale#21602

Signed-off-by: Samuel Keeley <samuel@keeley.net>
…ny Bind

StdNetBind sends every peer's traffic through one unconnected socket per address family: the kernel resolves the route on every send, and every peer's receive goes through one queue and one reader. On darwin, which cannot batch that socket's receives, it is the bottleneck.

ConnectedSockets keeps a socket for each (local address, remote address) pair that carries real traffic, connected to the remote and bound to the caller's port on that local address. Packets look the same on the wire, so NAT mappings are unaffected, and the kernel delivers each peer's traffic to its socket. It is exported and keyed on netip.AddrPort so any Bind can use it; the caller passes its port and a Control for whatever its own socket carries, such as a fwmark or an interface binding.

- A pair opens once it has carried ConnectedConfig.OpenAfter bytes (1 MiB by default) within one idle interval, and closes after a whole interval with no traffic, so handshakes and idle peers stay on the caller's socket. MaxSockets caps the set at 256.
- The port is shared without opening it to other users: SO_REUSEPORT alone on Linux, where the kernel limits it to one effective uid, and SO_REUSEADDR alone on darwin and the BSDs, where SO_REUSEPORT on a wildcard socket would let any local user take the port. ReusePortControl gives the caller's own socket the matching option.
- Sends follow the kernel's route selection: the route to each remote is looked up once and rechecked at every idle check, and Rebind redials every socket after the caller's socket moves.
- Batches use sendmmsg with UDP GSO and recvmmsg with GRO on Linux, sendmsg_x and recvmsg_x on darwin, and one datagram per syscall elsewhere. Receive slots start at 2 KiB and grow to 9 KiB the first time a datagram fills one. After a read that finds a partial batch, the reader pauses briefly (500 us on darwin, 200 us elsewhere) so the next read collects more.
- An ICMP error reported on a read (ECONNREFUSED while a peer restarts) does not stop a socket's reader; a source address that disappears closes the socket so the next send dials a new one.
- ConnectedDeliveryCheck tests once over loopback that the kernel delivers a peer's datagrams to its connected socket rather than to the wildcard socket sharing the port. Linux only does so since 5.4 (backported to 4.19.75); if the check fails, no sockets open.

It builds on every unix except AIX, Solaris and illumos, and is a nil-safe stub elsewhere.

Updates tailscale/tailscale#21602

Signed-off-by: Samuel Keeley <samuel@keeley.net>
WithConnectedSockets(true) makes StdNetBind open a ConnectedSockets set beside its own sockets. They are opened with ReusePortControl so the set can share their port, Send tries the set first and falls back to the shared socket, and the set's reads are one more ReceiveFunc. NewStdNetBind and NewDefaultBind now take Options; the Windows Bind accepts and ignores them.

On Linux a send uses the pair from the endpoint's sticky source, so a reply leaves from the address the peer sent to, as it would from the shared socket. SetMark resets the set so its sockets are redialled with the mark.

Connected sockets are off by default on every platform for now, so they are only in use when a caller asks for them.

Updates tailscale/tailscale#21602

Signed-off-by: Samuel Keeley <samuel@keeley.net>
…uffers

ConnectedSockets read each socket in a goroutine of its own and copied what it read into the device's slab. On small cores that copy and handoff cost more than connected sockets saved: an Intel N305 receiver spent 11% more CPU per bit than with the unconnected socket.

A Bind may now start receive routines after Open, through a new optional interface, conn.ReceiveFuncStarter, which the device implements. ConnectedSockets hands each socket to its caller as a ConnectedReadFunc through ConnectedConfig.Reader, and StdNetBind asks the device to run it, so each socket is read with recvmmsg (recvmsg_x on darwin) straight into the device's slab. ReadBatch and the per-socket reader goroutines are gone.

In the device, starts are refused once the bind begins closing, a started routine retires on net.ErrClosed, a batch larger than BatchSize is split across per-peer containers, and the larger slabs a GRO socket reads into (64 slots, 4 MiB) come from a per-size pool capped at 32 MiB that makes a reader wait while all are in flight.

Against the copying design, receive CPU per bit is 18-29% lower on the N305, now below the unconnected socket, and 6-12% lower on Xeons and Apple Silicon, with memory unchanged.

Updates tailscale/tailscale#21602

Signed-off-by: Samuel Keeley <samuel@keeley.net>
WithDontFragment sets IP_DONTFRAG and IPV6_DONTFRAG on a StdNetBind's sockets and its connected sockets. It is off by default and only has an effect on darwin.

Without DF, macOS gives each IPv4 datagram a random identification field; with DF it writes zero. The random ID costs the sender about 0.5 us of kernel time per datagram, and a Linux receiver's GRO only merges datagrams whose IDs are fixed or count up by one, so it receives a Mac's IPv4 datagrams one at a time. With DF, one TCP flow from an M2 Max to a Linux peer went from 3.3-3.8 to 5.8-6.2 Gbit/s over IPv4.

The cost is that a datagram larger than the path MTU is dropped rather than fragmented, as Linux already does for UDP by default.

Updates tailscale/tailscale#21602

Signed-off-by: Samuel Keeley <samuel@keeley.net>
WithBatchedIO makes a StdNetBind read and send on its own unconnected sockets up to IdealBatchSize datagrams per syscall on darwin, with recvmsg_x and sendmsg_x. It is off by default and does nothing elsewhere, since Linux already batches with recvmmsg and sendmmsg.

The work is in an exported type, UnconnectedBatch, whose ReadBatch and WriteBatch take the same ipv6.Message slices as x/net's batch methods, so StdNetBind drives it through its existing batched path and other Binds can use it. sendmsg_x honours a per-datagram destination on an unconnected socket (checked on macOS 12, 15 and 27): darwinbatch gains StageTo and AddrPort for it, and UnconnectedSelfTestErr checks it with the running kernel before use.

It is independent of connected sockets. In a quick measurement it saved some receive CPU without raising throughput, since darwin's unconnected socket is limited by per-datagram routing rather than by syscall count, so it is kept optional.

Updates tailscale/tailscale#21602

Signed-off-by: Samuel Keeley <samuel@keeley.net>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant