Files
slocker-lite/docs/networking-design.md
ceamac 5080895043 Document UDP port-forward support
CLAUDE.md's port_forward.{h,cpp} entry, README.md's -p table row, and
docs/networking-design.md's syntax line all updated for the new
[/tcp|udp] suffix. Includes the local dev-machine (root, via the scoped
doas rule) verification detail: TCP unaffected, UDP confirmed end-to-end
(a raw datagram sent to the forwarded host port was read back inside the
container via -x/--exec), same port pair coexisting on both protocols,
invalid-protocol parse errors, clean teardown, and --clean-processes
sweeping both the old 3-field and new 4-field state-file formats. Real
Android iptables/tetherctrl_FORWARD confirmation for UDP is still open.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-04 06:48:20 +00:00

62 KiB
Raw Permalink Blame History

Network isolation design

Status: all six commits landed (see "Implementation plan: commit sequence" below for what shipped, including corrections found by testing along the way). Rootless networking and an nftables backend remain deliberately out of scope (see below). Captured 2026-08-30 on the networking branch.

Goal

Add persistent, named networks that -r/--run containers can join, similar in spirit to how -v/--volume already creates persistent named volumes:

  • extern networks: reachable from/to the host's real network.
  • intern networks: only reachable by other containers that joined the same network, never from the host or outside.
  • -r/--run takes one or more -n/--network <name> flags to join networks.
  • -p [<network>:]<host-port>:<container-port> forwards a host port into a container on one of its extern networks.

Why not slirp4netns

slirp4netns (and pasta) exist specifically to provide networking to an unprivileged network namespace without CAP_NET_ADMIN or host iptables access — a strict 1-to-1 usermode-NAT translator between the host and exactly one namespace. Neither tool bridges multiple namespaces together; Podman's own rootless multi-container networks get inter-container connectivity by putting every container of that network into one shared namespace holding a real Linux bridge, with only that one namespace getting a slirp4netns/ pasta instance for outside access.

This project's target device is Android and normally has root available (slocker-lite-priv-drop's whole reason to exist is making root-mode -r/--run first-class). Root-only for now — rootless networking is explicitly deferred as its own future task. Given root is assumed, the usermode-NAT workaround buys nothing: native Linux bridge/veth/iptables is faster (kernel-native routing/NAT, no usermode packet copy), simpler (no extra long-lived process per network to babysit), and is exactly what Docker/Podman themselves do when running as root. slirp4netns/pasta are dropped from the design entirely; the target device doesn't need either installed for this feature.

bwrap itself has no "join an existing network namespace" flag (confirmed via bwrap --help): it has --userns FD/--pidns FD (join-by-FD) for user/pid, but only --unshare-net (always create new) for networking. To put a container into a specific persistent namespace, the join happens the same way this project already joins the rootless containers-storage mount's namespace: wrap the bwrap invocation in nsenter --net=<path> first, and don't ask bwrap for --unshare-net on that invocation. bwrap.cpp's existing wrap_for_root_namespace() is precisely this pattern already (for mnt/user) and is directly reusable for net.

Mechanism

Both extern and intern networks use the same core mechanism — a Linux bridge, with one veth pair per joined container (container-side end moved into that container's own, still separately---unshare-net'ed namespace; the bridge-side end attached to the network's bridge). Both kinds' bridges live inside their own dedicated, free-standing namespace (persistent the way ip netns add keeps a namespace alive with no process in it — bind-mounting its ns/net file to a path that outlives any one process). This wasn't always true for extern — see "Resolved: extern had no connectivity at all on the real device" below for why it changed, and why that change was necessary in the first place.

  • intern: no route out exists at all beyond the private namespace: a real structural isolation boundary, not merely "extern but without a NAT rule" (which would still be reachable from the host itself).
  • extern: needs net.ipv4.ip_forward=1 plus one iptables MASQUERADE (SNAT) rule for the bridge's subnet, both applied inside that private namespace — the same pattern docker0 uses, just relocated. Since a namespace genuinely isolated this way has no path outside on its own by construction, it additionally gets an uplink: a second, narrower point-to-point link out to the host's own root namespace, with NAT and routing set up there instead. See "Resolved: extern had no connectivity at all on the real device" below for the uplink's own mechanism and the three distinct host-root-side pieces (a specifically-ordered iptables rule, and two ip rules) it needed to actually carry traffic. If IPv6 is enabled for the network, only the IPv6 forwarding sysctl is set — deliberately no ip6tables MASQUERADE rule, and no IPv6 equivalent of the uplink either: the ULA (fd00::/8) addresses this project allocates are non-globally-routable by design (RFC 4193), so NAT66 for them isn't correct IPv6 practice regardless, and confirmed on the real target device that neither ip6tables nor nftables can even create an IPv6 NAT table on that kernel at all. extern's IPv6 side is thus same-bridge reachability only, exactly what intern's IPv6 side already is — see the "Resolved" IPv6 sections further below.

Each container always keeps its own private net namespace, so its own loopback and interface set never clash with another container's. Joining N networks is just N veth pairs into that one namespace — genuine multi-network membership falls out naturally, no namespace-sharing tricks needed.

Port forwarding (-p)

No slirp4netns hostfwd API to lean on, so this is implemented directly:

  • An iptables DNAT rule in PREROUTING: --dport <host-port> -j DNAT --to-destination <container-veth-ip>:<container-port>.
  • A FORWARD ACCEPT rule for the bridge's subnet — needed because a default FORWARD DROP policy (plausible on a stock Android kernel) would otherwise silently eat the forwarded traffic, the same gap Docker itself works around.

Syntax: -p [<network>:]<host-port>:<container-port>[/tcp|udp]. <network> is optional — when omitted, resolves to whichever single extern network the container joined; an error (not a silent guess) if the container joined more than one extern network and didn't disambiguate. The protocol suffix is optional too, defaulting to tcp; both the DNAT and FORWARD ACCEPT rules above use whichever protocol was resolved, so a UDP forward gets the exact same treatment as TCP, just with -p udp instead of -p tcp on both — the same port pair can be forwarded once per protocol (e.g. -p 53:53/tcp -p 53:53/udp for a DNS-like service), since neither rule set collides with the other.

DNS resolution

Containers on a shared -n <network> can resolve each other by the name given via --hostname on any network they share, plus host.containers.internal (podman's own convention) resolving to the first extern network's own gateway, if any. Implemented via dnsmasq — no resolver of any kind was referenced anywhere in this project before this, and dnsmasq isn't in README.md's runtime-dependency list, so this is a genuinely new, but deliberately optional (best-effort, gated on is_dnsmasq_available()), dependency: a missing dnsmasq degrades with a warning rather than failing -r/--run.

One dnsmasq instance per session, not per network. The "obvious" design — one persistent instance per network, mirroring how the uplink provisions one process per network — was considered and rejected before being built, once a real correctness problem was worked through: a container joined to two networks would list two nameserver lines in /etc/resolv.conf, and standard stub resolvers (glibc, musl, busybox included) don't fall through to the next nameserver on NXDOMAIN, only on timeout — a name that exists only on the container's second network would silently fail to resolve, since the first nameserver's authoritative "no such name" ends the lookup right there. Running one instance per session instead sidesteps the problem entirely: it's entered into the container's own network namespace and bound to 127.0.0.1:53 there, so /etc/resolv.conf is always just nameserver 127.0.0.1 regardless of how many networks were joined, and that one instance is configured (via dnsmasq's own repeatable --hostsdir=<dir>, one per joined network, inotify-watched so no reload signal is ever needed) to already know about every network the specific container asking could possibly mean — there's never a second nameserver to fall through to in the first place.

Each network gets one shared, host-global hosts-directory ($XDG_STATE_HOME/slocker-lite/dns-hosts/<network>); each session that joined it and was given a --hostname writes one record file into that directory (plain <ip> <hostname> syntax) so sibling sessions' own resolvers, watching the same directory, discover it — and since a session's own record lands in the exact directory its own resolver also watches, self-resolution falls out for free, no special-casing needed. host.containers.internal is generated per session instead (a small, private --addn-hosts=<file>, not shared) pointing at the first extern network's own gateway address, omitted entirely for an intern-only session.

Three real bugs found while building this

Each confirmed by direct testing — isolating the change, reproducing the symptom, then confirming the fix — not assumed:

  1. dnsmasq only writes --pid-file while actually daemonizing. -d/--no-daemon was tried first (a simpler model: the forked/exec'd pid stays the real pid throughout, matching how nsenter's own in-place execve() already lets other parts of this project treat a pid as stable across exec()). It suppresses --pid-file entirely — confirmed directly: dnsmasq started and successfully read the hosts file (visible in its own, still-attached stdout at the time) but the pid-file this project polls for (the same bounded-poll shape network_join.cpp's own wait_for_isolated_net_namespace() already uses) never appeared, so every attempt timed out. Fixed by letting dnsmasq daemonize normally instead: the forked/exec'd process is then only the intermediate one (reaped immediately, not tracked as the resolver's own pid), and the real, final daemon pid is read back from the pid-file itself once it appears — which also is the crash-orphan record --clean-processes needs, no separate write step required.
  2. dnsmasq drops root privileges to an unprivileged user by default. Broke reading anything under $XDG_STATE_HOME at all (typically /root/.local/state/slocker-lite/..., mode 0700 — a non-root user can't even traverse into /root), traced directly to dnsmasq's own log: bad dynamic directory .../dns-hosts/<network>: Permission denied, followed by every query coming back REFUSED. Fixed with an explicit --user=root --group=root, matching this project's already-root-only networking model throughout (bridges, iptables, the uplink, all already assume root). Deliberately not the final answer — tracked as a security follow-up in this repo's own TODO.md: run dnsmasq as a real low-privilege user instead, with the state it needs relocated somewhere that user can reach, rather than keeping a resolver process root for its entire lifetime purely to work around a directory permissions mismatch.
  3. An AAAA query for a name with only an A record came back REFUSED. Every record here is IPv4-only (matching JoinedNetwork::container_ip's own existing scope) — but without --filter-AAAA, an AAAA query for a name dnsmasq otherwise knows perfectly well (an A record, just resolved correctly moments earlier via nslookup) also came back REFUSED, not a clean "no data" answer. This broke ping <name> outright: ping, like most getaddrinfo()-based tools, queries both A and AAAA together and treats REFUSED on either as a hard failure for the whole lookup, not merely "no IPv6 available" — confirmed by testing both with and without --filter-AAAA against the identical hosts record. host.containers.internal needed a second, related fix on top: serving it via dnsmasq's own --address=/name/ip option kept returning REFUSED for AAAA even with --filter-AAAA given — confirmed --address records aren't treated the same internally as ordinary hosts-file entries are — so it's instead served through the exact same mechanism as everything else (a plain, session-private --addn-hosts=<file>), which resolved it immediately.

Verified end-to-end, both on this dev machine and on the real Android target device (a fresh reboot, invoked consistently from $HOME — see this repo's own TODO.md for a separate, unresolved issue found along the way where invoking from a different working directory produced a completely separate config/state tree): two containers on a shared network resolve each other by name (including self-resolution) and can ping by name; a container joined to both an intern and an extern network resolves both its intern peer and host.containers.internal simultaneously — the specific scenario the per-network-instance design would have broken.

Subnet / IP allocation

  • IPv4 auto-allocates a /24 starting at 10.168.0.0/24, incrementing per network (10.168.1.0/24, 10.168.2.0/24, ...), with an optional --subnet <cidr> override at -n/--network creation time.
  • IPv6 is a per-network on/off option, defaulting to enabled. When on, also auto-allocates a ULA /64 alongside the IPv4 block: a randomly generated /48 (RFC 4193 — fdf0:f243:f06f::/48, replacing an earlier fd00:168:0::/48 placeholder that was never actually randomly generated; see the "Resolved" IPv6 randomization section further below) with the subnet-id hextet offset by a fixed 168fdf0:f243:f06f:168::/64, fdf0:f243:f06f:169::/64, ... — keeping the same project-recognizable 168 stamp the old scheme's fixed 2nd hextet had, now as a constant offset rather than a numerically-identical index. --subnet6 <cidr> overrides it per network as before. --no-ipv6 at creation time skips both the v6 allocation and the IPv6-forwarding sysctl for that network entirely (there's no ip6tables setup to skip — see "Mechanism" above for why one is never added at all).

CLI surface

-n/--network, dual-purpose like -v/--volume (alone creates/manages a network; combined with -r joins it) — with a simpler token shape than -v's, since a network join has no equivalent of a volume's container-mount-path second argument:

  • Alone: -n <name> --extern|--intern [--subnet <cidr>] [--no-ipv6] [--subnet6 <cidr>]<name> is -n's own required_argument (a single token, standard getopt_long); kind and the subnet/ipv6 options are ordinary separate flags rather than additional positional tokens.
  • With -r: -n <name> is repeatable, one token (just the name) per occurrence — joins that network. No membership limit.
  • --list-networks / --delete-network <name> round out the set, mirroring --list-volumes / --delete-volume.
  • --delete-network-full <name> (landed later, once stale host state actually caused a real problem — see the "Resolved" section below): mirrors --delete-volume-full, but unlike that one it doesn't gate config removal on teardown succeeding — a network's live state is multiple independent pieces (rules, bridge, namespace), each individually best-effort, since a piece "failing" because it was already removed by hand is the expected case this exists to handle, not a reason to leave the network stuck in the config.

Config file schema

New top-level networks section in config.yaml, parallel to volumes (config_file.{h,cpp}) — per entry: name, kind (extern/intern), subnet, ipv6 (bool), subnet6 (if ipv6). create_volume_command() / list_volumes_command() / delete_volume_command() (commands.cpp) are the direct templates to follow, including the shared tab-aligned listing helper for --list-networks.

A network's config-file entry is the durable source of truth, the same way a volume's directory path is. The live host-side state (namespace file for intern, bridge, veths, iptables rules) doesn't survive a reboot and is reconciled lazily and transparently — checked and, if missing/stale, recreated from the config entry — the first time anything needs it after a reboot (no explicit "start" step; the first -r/--run -n <name> or -p after boot just works).

Tooling

Shell out to iptables (via process.h's existing run_process(), matching how containers-storage/bwrap/fuse-overlayfs are already invoked) for MASQUERADE/DNAT/FORWARD rules, and to ip for bridge/veth/namespace management. ip6tables is deliberately never used at all (see "Mechanism" above) — one fewer required tool on the real target device, whose ip6tables build turned out not to support MASQUERADE anyway. Both ip and iptables need an on-device availability check the same way check_required_dependencies() already gates on containers-storage/bwrap.

iptables only, for now — confirmed available on the real target device; nftables is not currently installed there. nft support is real future work (its own task, once it actually matters — e.g. a device that only ships nft), not built preemptively as a dual-backend abstraction here.

Host-global state discipline

Unlike a bwrap session (self-contained via its own namespaces), bridges/veths/iptables rules are host-global, named, persistent resources that outlive any single process. Needs the same create-on-start/ remove-on-stop discipline already used for pid locks (pid_file.{h,cpp}) and cgroups (session_cgroup.{h,cpp}), plus a --clean-processes-style sweep for anything orphaned by a crash, so a dead slocker-lite doesn't leave stray bridges/veths/iptables rules behind forever.

Implementation plan: commit sequence

The whole feature is too large for one commit (unlike, say, --kill, which landed as a single commit despite touching several new files). Split into six commits, each a coherent, independently buildable and manually verifiable unit, in dependency order. Docs (CLAUDE.md/README.md) get updated within each commit, matching this branch's existing practice — not saved for a final pass.

  1. Config schema + subnet/IPv6 allocation + -n/--network create/list/delete (config-only, no host-side effects yet)

    • config_file.{h,cpp}: new NetworkEntry {name, kind, subnet, ipv6, subnet6} (kind = extern/intern), a networks section — directly parallel to VolumeEntry/volumes.
    • New network_subnet.{h,cpp}: IPv4 /24 auto-allocation starting at 10.168.0.0/24 incrementing per existing network, --subnet override; paired IPv6 ULA /64 auto-allocation (fd00:168:0:N::/64 at the time — later replaced with a genuinely randomly generated /48, see the "Resolved" IPv6 randomization section further below) when enabled (default), --subnet6 override, --no-ipv6 to skip. Pure allocation logic against the already-loaded config's existing networks — no kernel/ ip/iptables calls in this commit.
    • cli_args.{h,cpp}: Mode::network/list_networks/delete_network, -n/--network (required_argument, single token = name; separate --extern/--intern/--subnet/--no-ipv6/--subnet6 flags for the create case), --list-networks, --delete-network <name> — directly mirroring -v/--volume's existing three-mode shape in the same file.
    • commands.cpp: create_network_command()/list_networks_command()/ delete_network_command() — mirroring create_volume_command()/ list_volumes_command()/delete_volume_command(), including the shared tab-aligned listing helper.
    • Verify: -n mynet --extern, -n other --intern --no-ipv6, --list-networks shows both with correct kind/subnet, --subnet/ --subnet6 overrides land correctly in config.yaml, --delete-network removes an entry. No bridges/namespaces/iptables rules exist yet — purely config bookkeeping, same as a freshly-created volume before it's ever mounted.
  2. Persistent network namespace primitives (generic infra for intern networks)

    • New persistent_netns.{h,cpp}: create/find/remove a persistent network namespace kept alive with no process in it, the way ip netns add does (bind-mount a fresh namespace's ns/net onto a path that outlives the creating process) — narrow, reusable infra, no intern/extern branching or bridge logic here (parallels how session_cgroup.{h,cpp} stayed narrowly scoped to cgroup mechanics only).
    • Not wired into -n/--network yet in this commit.
    • Verify: a small manual exercise (or a -t/--test addition) creating a persistent namespace, confirming it survives after the creating process exits, then removing it.
  3. Bridge provisioning for a network (idempotent — this is also the reboot- reconciliation mechanism, not a separate later step)

    • New network_bridge.{h,cpp}: given a NetworkEntry, ensure its bridge exists and is configured — creating it if missing (idempotent, so this doubles as "reconcile after reboot" with no separate code path):
      • extern: bridge in the host's own root namespace; assign it the gateway IP from the network's subnet; net.ipv4.ip_forward=1 (+ IPv6 forwarding sysctl if ipv6); one iptables MASQUERADE rule for the subnet. Landed without the ip6tables MASQUERADE rule this bullet originally called for (see "Resolved" IPv6 section further below, and "Mechanism" above): ULA addresses are non-globally-routable by design, so NAT66 for them was never correct IPv6 practice, and the real target device's ip6tables build doesn't support MASQUERADE at all regardless.
      • intern: bridge inside its own dedicated persistent_netns.h namespace (commit 2); gateway IP assigned; no forwarding, no NAT rule — no route out at all.
    • Wire this "ensure provisioned" call into create_network_command() (so creating a network actually stands up its bridge immediately) — later commits also call it lazily before a join, covering the reboot case.
    • check_required_dependencies()-style availability check added for ip/ iptables, alongside the existing containers-storage/bwrap check (no ip6tables check, for the same reason it's never called).
    • Verify: -n mynet --extern produces a real bridge with the expected gateway IP, ip_forward enabled, and a matching MASQUERADE rule (ip link show, iptables -t nat -L); an intern network's bridge exists in its own namespace with no such rule. Delete/recreate a network's config entry, delete its bridge by hand (ip link del), then trigger provisioning again (e.g. re-running --network create or the first join in commit 4) and confirm it comes back.
  4. Joining networks at -r/--run time: veth creation, IP assignment, route

    • Repeatable -n <name> with -r/--run (cli_args.cpp, same repeatable-with--r shape -v/--volume already has).
    • commands.cpp's run_container(): once bwrap's pid (and via resolve_namespace_pid()-style lookup, sandbox_process.h, its actual net namespace) is known — same timing hook on_bwrap_pid_known/ on_start already provides for session locks/cgroups (bwrap.cpp) — for each joined network: ensure it's provisioned (commit 3, covers reboot recreation), create a veth pair, move the container-side end into the container's net namespace, attach the bridge-side end, assign the container's veth an IP from the subnet, and (for an extern join) set it as the default route.
    • This is the core connectivity commit — no veth pairs exist before it, regardless of how many networks are configured/joined.
    • Verify: two containers joined to the same intern network can ping each other and cannot reach the host or outside; a container joined to an extern network can reach the outside (and the host cannot reach it without commit 5's port forwarding); a container joined to both loses neither path (two interfaces, both functional).
  5. -p port forwarding

    • cli_args.cpp: -p [<network>:]<host-port>:<container-port>, <network> optional (resolves to the container's sole extern network; error if ambiguous).
    • New port_forward.{h,cpp}: add/remove the iptables DNAT (PREROUTING) + FORWARD ACCEPT rule pair for one mapping, tied to the container's own session lifecycle the same create-on-start/ remove-on-stop way pid_file.{h,cpp}/session_cgroup.{h,cpp} already are.
    • Verify: -p 8080:80 against a container on an extern network answering on port 80 is reachable via curl localhost:8080 from the host; the rule is gone after the container exits.
    • Landed with two real corrections found by testing (see CLAUDE.md's port_forward.{h,cpp} entry for the full detail): the DNAT rule needs both PREROUTING and OUTPUT (locally-generated traffic never traverses PREROUTING); and curl localhost:<port> specifically still doesn't work even so (NAT hairpinning — the container sees an inbound packet claiming a loopback source on a non-loopback interface and drops it as martian) — verified instead via curl <host's real IP>:<port>, the actually-relevant path for real clients. Also surfaced, unrelated to -p itself but found while testing it: -x/--exec didn't join the net namespace (written back when this project never isolated networking at all), so it saw the host's network stack, not a network-isolated session's own — fixed in a follow-up commit (exec_session.cpp, see CLAUDE.md's own entry for that file).
  6. Crash-orphan cleanup sweep

    • Extend --clean-processes (or add a dedicated --clean-networks, whichever reads better once this is reached) to find and remove bridges/veths/iptables rules left behind by a slocker-lite that died before its own teardown ran — mirroring clean_stale_sessions() (pid_file.cpp)'s existing stale-pid-file sweep, but for host-global network state instead of pid files.
    • Verify: kill -9 a running -r/--run session mid-flight (bypassing its normal cleanup), confirm the orphaned veth/iptables rule is detected and removed by the sweep, and that a still-running session's state is left untouched.
    • Landed narrower in scope than the bullet above once the actual orphan surface was worked out (see CLAUDE.md's port_forward.{h,cpp} entry for the full detail): veths need no sweep at all (the kernel tears down an entire pair once either end's namespace is destroyed — never survives a crash), and bridges/persistent namespaces are deliberately meant to always outlive any one session (that's the whole point of the reboot-reconciliation design, not something a crash changes). Only -p's iptables rules — host-global, named, with no automatic teardown — can actually outlive a crashed session, so that's the entire sweep: extended --clean-processes (not a separate flag) with clean_stale_port_forwards(), cross-referencing a small per-session port-forward record file against list_sessions()'s own liveness check. Verified via a controlled scratch test rather than a literal kill -9 on a root-owned slocker-lite process (not achievable through this session's scoped doas rule, which only permits running slocker-lite itself, not arbitrary commands like kill): a fabricated stale record was correctly detected, its removal attempted, and its file cleaned up, while a record matching a real running session was left untouched.

TUN/TAP fallback for veth-less kernels

Trigger: the real target device's kernel supports tun/tap (CONFIG_TUN — Android needs this for VpnService-based VPN apps) but not veth (CONFIG_VETH, commonly stripped from mobile kernels), so -n --extern/--intern as designed above (a veth pair per join) simply can't work there at all — ip link add ... type veth ... fails outright. Bridge support was separately confirmed working on this same device, which rules out the more complex fallback this section originally considered (see git log on this file for the superseded sketch: a per-network userspace Ethernet switch with no bridge dependency at all) in favor of a much smaller design.

Why tap can't just replace veth 1:1: a veth pair is two real kernel netdevices, switched between (or into a bridge) entirely by the kernel with zero userspace involvement. A tap device only has one kernel-side netdevice — the other "end" is a raw-Ethernet-frame file descriptor that only a userspace process can read/write, so there's no second kernel endpoint to attach to a bridge. This is exactly why slirp4netns/QEMU's own tap networking need a userspace process on the fd side at all.

Design, confirmed and implemented: per network-join, two tap devices + one small relay process that copies bytes 1:1 between them — a direct functional substitute for one veth pair, reusing the existing bridge as the switching fabric so provision_bridge() needs no changes at all:

  • A host-side tap device, created wherever the network's bridge lives and enslaved to it — exactly veth's host-side role.
  • A container-side tap device, created directly inside the container's own namespace, named eth<N> from the start — no peer-name-then- rename dance needed, unlike veth.
  • A relay process holding both fds open, copying raw Ethernet frames bidirectionally between them for as long as it runs. This is the "veth wire," just implemented once in userspace instead of by the kernel.

Once the container-side tap exists as eth<N>, everything downstream — IP assignment, routes, the address handed to -p — is completely unchanged; only the interface-creation step is swapped. Strategy selection is per-network, per-join: should_use_veth(network) = network.veth && probe_veth_support(), mirroring the existing kernel-capability-vs-policy split namespace_policy_enabled() (bwrap.cpp) already uses for --unshare-xxx. --no-veth at network-creation time forces the fallback even on a veth-capable kernel — how this was actually tested, since the real target device wasn't available during development.

Implementation plan: commit sequence

Landed as four commits (a fifth, this doc update, closes it out) — see CLAUDE.md's own entries (network_bridge.{h,cpp}, network_tap_relay.{h,cpp}, network_join.{h,cpp}, self_test.{h,cpp}) for full file-by-file detail:

  1. Veth capability probe + --no-veth flag, no relay yet. probe_veth_support() (fork, unshare(CLONE_NEWNET) into a throwaway namespace, try ip link add ... type veth ... there — the same kernel-capability-probing shape bwrap.cpp's own kernel_supports_namespace() already uses); NetworkEntry::veth + YAML round-trip; --no-veth CLI flag. Verified: probe_veth_support() returns true on this dev machine (a real veth pair is created successfully); --no-veth persists veth: false.
  2. network_tap_relay.{h,cpp}: relay creation/loop/teardown, standalone. Verified via a new -t/--test case: a throwaway bridge + throwaway network namespace, confirming the host-side tap attaches to the bridge, the container-side tap appears inside the target namespace with the requested name. The original assumption here — that both devices disappear on their own once the relay is stopped, since neither was created with IFF_PERSIST — turned out to be wrong on the real target device (see "Resolved: tap devices need to be created persistently, not tied to the relay's own fd lifetime" below); devices are now created persistently via an external ip tuntap add, and the host-side one is explicitly removed (ip link del) when the relay stops.
  3. Wire into join_one_network()/join_networks(). JoinedNetwork gains an optional relay handle; run_container() collects and stops them after run_bwrap() returns, mirroring -p's own active_port_forwards handling exactly. Two real bugs found here, not assumed: the relay child, unlike every other forked child in this project, never exec()s, so it inherited (and never closed) a live copy of daemonize.cpp's own report-pipe write end, hanging -D combined with -n indefinitely until fixed with an explicit close_inherited_fds(); and a first-attempt fix making --kill reach the relay directly (adding its pid to the session's own cgroup) was reverted after it caused a different bug (the cgroup's own removal, which happens before run_container() gets to stop the relay, started failing with EBUSY) — the ordinary flow already stops the relay correctly on its own, so the added complexity wasn't worth it. Verified end-to-end: two containers on a --no-veth extern network got distinct addresses via two tap+relay pairs (no veth at all) and pinged each other with 0% packet loss, repeatably.
  4. Crash-orphan sweep. record_tap_relays()/clean_stale_tap_relays(), the direct structural analog of -p's own sweep, wired into --clean-processes. Verified the same way the port-forward sweep was: a real rootless session's pid alongside a hand-written matching record (left untouched) and a fabricated stale one (correctly swept).

Resolved: gateway/outside reachability was stale host state, not a bug

Originally reported as an unconfirmed gap, now resolved. An earlier pass of testing found gateway/outside reachability through the tap+relay fallback consistently failing (ICMP and TCP both unanswered), while the identical bridge/subnet reached via veth worked — with rp_filter=0 (host-tap, bridge, and global all scope) tried and not fixing it. The actual cause turned out to be accumulated leftover bridges and iptables rules from many earlier rounds of manual testing on this same dev machine (this feature's --delete-network, before --delete-network-full existed, never tore down live host state — see that flag's own section below). After manually clearing every leftover bridge and iptables rule and retesting on a clean host, gateway and outside reachability both work correctly through the tap+relay fallback: a fresh --no-veth extern network's gateway IP and a real external host both answered ICMP with 0% packet loss, and a raw TCP connect (nc) to an external host on port 80 completed cleanly (a wget HTTP request against the same host separately segfaulted — confirmed to be a busybox wget bug unrelated to networking, reproducing identically regardless of which join mechanism was used). Inter-container connectivity was reconfirmed working at the same time. --delete-network-full (see below) exists specifically so this kind of stale-state accumulation can't recur — always prefer it over --delete-network when a network won't be recreated with the same name, or when testing repeatedly against the same name during development.

Resolved: extern had no connectivity at all on the real device

A second, deeper problem — genuinely device-specific, not stale state. After the dev-machine investigation above, the user confirmed eth0 creation itself was reliable on the real Android target device — but a plain extern network there had no connectivity at all, gateway included, in either IPv4 or IPv6, even after a full device reboot ruled out stale state (unlike the dev-machine case above, intern on the same device worked correctly the whole time — a genuine, reproducible structural difference).

Root cause, found by isolating the actual differences between extern and intern provisioning one at a time on a disposable branch, rather than guessing: extern's bridge lived directly in the host's own root network namespace, while intern's always lived inside its own dedicated persistent one. Relocating extern's bridge into the same kind of private namespace fixed gateway reachability immediately, both IPv4 and IPv6 — the "Mechanism" section above already describes the resulting (current) architecture. The underlying cause is almost certainly Android's own netd-managed iptables/routing policy, which applies only in the root namespace and never touches a genuinely isolated one — consistent with everything found in the steps below, all of which turned out to be root-namespace-specific policy, not anything wrong with the bridge/NAT/forwarding mechanism itself.

That fix alone loses outside connectivity by construction — a namespace this isolated has no path to the real network at all. Restoring it needed an uplink: a second, narrower point-to-point tap+relay link (reusing network_tap_relay.h's existing primitive with a new attach_host_side_to_bridge=false mode — a plain routed link, not another bridge port) between the private namespace and the host's root namespace, on its own small deterministic 169.254.0.0/16 transit subnet, with NAT applied only in host root.

Three further, independently necessary pieces were needed for that uplink to actually carry traffic — each found by direct real-device testing (SSH access to the device, live inspection of iptables -L -n -v, ip rule show, and /proc/net/nf_conntrack), not assumed, and each confirmed necessary by dropping it and reproducing the exact original symptom:

  1. The iptables FORWARD accept rule needed -I FORWARD 1 (insert at the front), not -A (append). A first attempt appended the rule and saw no improvement; iptables -L FORWARD -n -v, captured live during a test, showed why: Android's own FORWARD chain unconditionally jumps through several of its own subordinate chains before reaching anything else, and one of them — tetherctrl_FORWARD, its tethering-control chain — contains a single unconditional DROP with no match criteria at all. Every forwarded packet reaches it and dies there regardless of interface; an appended rule sits after that point and is structurally unreachable, since DROP is already a terminal verdict. Inserting at the very front pre-empts the whole chain of subordinate jumps.
  2. An outbound ip rule. Even with #1 fixed, packets still went nowhere — iptables -L FORWARD -n -v showed 0 packets ever reaching the chain at all for a real destination. ip rule show, captured live, explained it: every rule landing in a table with a real, working route requires iif lo (locally-generated traffic only — exactly why an interactive ip route get 8.8.8.8 always looked fine on its own, since that lookup itself has iif lo). A genuinely forwarded packet — arriving via the uplink's own host-root-side tap, not generated locally — instead falls through to a generic fwmark 0/0x10000 catch-all landing in an unrelated, routeless table, and is dropped before a forwarding decision is even made. Fixed by discovering, dynamically (ip route get 8.8.8.8, parsing out the table <N> it names — not hardcoded, since the table number is specific to whichever real network, WiFi or cellular, is currently active), whichever table the host is actually using for its own real traffic right now, and adding ip rule add priority 100 iif <uplink-tap> lookup <table> — the explicit low priority matters too: a first attempt without one was silently placed after the same fwmark 0/0x10000 catch-all and never actually consulted, the identical append-vs-insert mistake as #1, just in ip rule instead of iptables.
  3. A return-path ip rule, mirroring #2 for the reverse direction. Even with #1 and #2 both fixed, outside connectivity still didn't work. Live inspection of /proc/net/nf_conntrack during a real request settled it: the outbound leg was already fully working — a genuine, tracked reply from the real destination, not merely a locally-generated packet succeeding (no [UNREPLIED] marker). But the reply, arriving back on whichever real interface is currently active and correctly de-MASQUERADEd by conntrack back to the uplink's own transit-subnet address, still had nowhere to go: ip route get <transit-addr> from <remote> iif <real-interface> returned "Network unreachable" — the exact same "falls into a routeless table" failure as #2, just for a packet whose destination (not source) is now the transit subnet, arriving on a real interface instead of the uplink's own. Fixed with ip rule add priority 100 to <transit-subnet> lookup main — routing by destination into the plain main table, which already has the directly-connected route to this subnet, regardless of which real interface a reply happens to arrive on.

Verified completely end-to-end on the real device, from a clean state: gateway IPv4 0% loss, gateway IPv6 0% loss, and a real outside destination (8.8.8.8) 0% loss (3/3 replies) through a container on a freshly created extern network — both through the veth path and, separately, through the --no-veth tap+relay fallback.

IPv6 outside connectivity was investigated separately and found not practically achievable on this device, so it stays local-only by deliberate decision (see "IPv6: no NAT (MASQUERADE), by design" below for the existing, still-valid reasoning against NAT66 specifically) — checked directly on the real device: nft add table ip6 ... fails outright ("Not supported" / "cache initialization failed"), so neither ip6tables nor nftables can NAT IPv6 on this kernel at all, ruling out a MASQUERADE- based fix analogous to the IPv4 uplink entirely. The NAT-free alternative (NDP proxying real addresses carved out of the device's own global prefix) was also ruled out: that prefix rotates every ~596 seconds (under 10 minutes) on the network tested against, for both the privacy address and the normally long-lived EUI64 address, too short-lived to build a stable addressing scheme on top of.

Resolved: joining 2+ networks left every network after the first unreachable

Reported from the real target device: joining a container to two or more networks in a single -r/--run left every network after the first permanently unreachable — eth0, eth1, eth2 all appeared inside the container, but only eth0's gateway ever answered a ping. Reproduced regardless of whether the networks involved were extern or intern, so the uplink mechanism above was ruled out as a cause early on.

Several timing-based experiments (explicit delays before/after join, varying join order) gave mixed, sometimes-contradictory results — the device's own background load turned out to vary wildly enough (a plain sleep 5 observed taking anywhere from ~5s to ~60s of real wall-clock time) that "adding a delay fixed it" couldn't be trusted as a signal at all. At the user's own suggestion, installing strace on the device and tracing the whole invocation (strace -f -tt, following every forked/exec'd process with absolute timestamps) turned this from an inconclusive guessing exercise into a five-minute, conclusive one.

Root cause: the trace showed the second network's own relay process dying on its literal first frame:

write(4, "...", 86) = -1 EIO (Input/output error)
exit_group(0)

EIO writing to a tap fd means the device isn't administratively up yet — and it genuinely wasn't. create_tap_relay() (network_tap_relay.cpp) returns, and its relay starts polling immediately, the instant the container-side tap device is created; join_one_network() (network_join.cpp), a separate process, still has its own ip addr add/ ip link set <if> up steps left to run afterward for that same device. The trace's own timestamps confirmed the ordering directly: ip addr add ... dev eth1 ran after the relay's fatal write, not before. For the first network joined, this race is narrow enough in practice that no frame ever arrives before those steps finish; for the second (and any later) network, something reliably delivers a frame before the interface is up — and the relay's frame-forwarding loop treated any write() failure as fatal, exiting silently and permanently on that single EIO, so that network never worked again for the rest of the session.

Fix: retry specifically on EIO/ENETDOWN (both mean "the device isn't up yet," a startup race, not a torn-down namespace) with a short bounded backoff — up to 50 × 20ms = 1s, generous compared to the ~14ms gap actually observed in the trace — instead of exiting immediately. No change was needed to join_one_network()'s own ordering; the relay simply waits out the gap.

A tempting shortcut ruled out: reordering join_one_network() to bring the interface up before creating the container-side tap device isn't possible — the device has to exist before it can be addressed or brought up, so some gap between "device exists" (when the relay starts polling) and "device is up" (when the relay can actually forward into it) is unavoidable by construction. Retrying in the relay is the only fix that doesn't require the relay to somehow block until an unrelated process finishes its own, separate setup steps.

Methodology note: an earlier strace -f attempt, wrapped in timeout 30, produced a misleadingly corrupted trace instead of a clean one — GNU timeout sends its kill signal to the whole process group by default, and strace's own tracing overhead was large enough (mount steps that normally take seconds took 3+ minutes under trace) that the real wall-clock timeout elapsed while still in early setup, silently killing several traced children mid-flight and producing a trace that looked like a crash but was really just the test harness cutting it off. Dropping the timeout wrapper entirely and letting the traced command finish naturally produced a clean, complete trace that correctly captured the real bug.

Verified end-to-end on the real target device: both 2 and 3 intern networks joined simultaneously in one session, all gateways reachable at 0% packet loss, clean teardown, no leftover state.

Resolved: -p/--port-forward had no connectivity at all on extern

Reported after the fixes above had already shipped: a server listening on an extern network wasn't reachable via -p at all — not from the host, and not from a real outside client — even though that same network's own gateway and outside connectivity (the uplink fix earlier in this document) worked correctly. intern was never affected (-p only ever targets extern networks by design, since an intern network's bridge has no path from the host at all).

Root cause: moving extern's bridge into a private namespace (the fix above) restored gateway reachability but, as a direct side effect, also left host root with no route to the container subnet whatsoever — confirmed directly: ip route get <container-ip> from host root fell through to whatever the host's own default route happened to be (the real LAN gateway), not the container's actual namespace. -p's own DNAT rule (port_forward.cpp) is added in host root and targets the container's real IP directly, so the rewritten packet had nowhere to go — the connection simply timed out, both from curl <real LAN IP>:<port> on the host itself and from a separate machine on the same LAN.

Fix, two independently necessary pieces — confirmed by testing, the same "a route alone isn't enough on Android" lesson the uplink's own outbound/return-path ip rules (above) already learned, just for the container subnet instead of the uplink's own transit subnet:

  1. A host-root route to the container subnet, through the uplink's own netns-side address (ip route add <container-subnet> via <uplink-netns- addr> dev <uplink-root-tap>). From there, no second NAT stage is needed: the private namespace already forwards arriving traffic to the bridge on its own (a fresh namespace's FORWARD policy is ACCEPT by default, and ip_forward is already enabled), and the reply's return path is already covered by the private namespace's own default route back out through this same uplink, plus host root's own conntrack correctly reversing the original DNAT on the way back out.
  2. A matching ip rule add priority 100 to <container-subnet> lookup main. Without this, the route added in step 1 is silently never consulted at all: confirmed via ip rule show on the real device that Android's own policy routing has no default "lookup main" rule anywhere in its list (it ends in a catch-all unreachable before priority 32766, where that default would normally live) — a generic fwmark 0/0x10000 lookup 99 rule (matching any untouched/forwarded packet, arriving on a real interface rather than locally generated) intercepts the packet first and routes it into an unrelated table with no route to the container subnet, long before rule evaluation would ever reach main.

A related robustness bug found while testing this, not the original bug itself: reproduced directly during testing when two concurrent test invocations collided (an incidental accident of testing on a live device, not a deliberate scenario) — a first failed ensure_uplink_provisioned() call only stopped the tap relay, leaving every ip rule/iptables piece it had already added (all deterministic, hash-derived names tied to the network name) live on the host. That leftover state then made every subsequent attempt to create a network with the same name fail identically and permanently — ip rule add returning File exists against a rule nothing had ever removed — with no way to recover short of a manual fix or a full device reboot. Fixed by recording the relay's pid to the uplink state file as soon as it's known (before any of the steps that can fail), so a failure can call the exact same teardown_uplink_state() a real --delete-network-full would use to roll back everything already added, instead of a second, partial copy of that cleanup logic.

Verified end-to-end on the real target device: a busybox httpd on a freshly (re-)created extern network, reached via -p 18080:80 — a curl from the device's own shell against its real LAN IP got a genuine HTTP response, and, separately, a curl from a completely different machine on the same LAN got the same result. Reproduced reliably across repeated fresh network creations, including after a full device reboot with no leftover state. curl 127.0.0.1:<port> (from the host, against loopback specifically) still doesn't work — that's the pre-existing, already- documented NAT-hairpinning limitation above (see "Port forwarding (-p)"), unrelated to this fix and out of scope here.

IPv6: no NAT (MASQUERADE), by design

Trigger: the real target device's ip6tables build doesn't support a MASQUERADE target at all, so provision_bridge()'s original IPv6 MASQUERADE rule (added in commit 3 of the implementation plan above) simply fails there.

This isn't worked around — it's dropped entirely, on both the veth and tap+relay paths, unconditionally, because it was never correct IPv6 design to begin with. The fd00::/8 addresses network_subnet.h auto-allocates are ULA (Unique Local Address, RFC 4193) — deliberately non-globally- routable, the IPv6 equivalent of RFC1918 private space (10.0.0.0/8, etc.). NAT66 (masquerading a ULA source to a real global address) is possible in principle and some consumer routers do offer it, but it's explicitly discouraged: one of IPv6's own core design goals was eliminating the need for NAT via its vastly larger address space — the "correct" way for a network to get real outside IPv6 access is a properly delegated, globally-routable prefix (via DHCPv6-PD from an upstream router), not NAT on a private range. This project doesn't do prefix delegation (a materially bigger feature, not currently planned), so attempting NAT66 here would only ever have been a workaround for that gap, not a real solution — dropping it is the more honest design, not a compromise forced by the missing ip6tables target.

Net effect: extern networks' IPv6 side now behaves exactly like intern's already did — real same-bridge reachability between containers over their ULA addresses, no path to the actual internet. IPv4 is unaffected; extern still gets full NAT'd outside access there. IPv6 forwarding (net.ipv6.conf.all.forwarding=1) is still enabled for extern (harmless, global, symmetric with the IPv4 case, and available if a host administrator wants to wire up real inter-network IPv6 routing some other way later) — only the ip6tables MASQUERADE call itself, and the ip6tables dependency check that gated it, were removed (network_bridge.{h,cpp}'s provision_bridge()/teardown_network_state()/ check_network_dependencies()). --no-ipv6 is unaffected by this — it still means "skip IPv6 addressing entirely," an orthogonal decision from whether NAT is ever attempted for the addresses that are assigned.

Resolved: IPv6 base prefix is now genuinely randomly generated

Trigger: a user asked directly, once IPv6 NAT was already dropped (see above) — the original fd00:168:0::/48 base was never actually generated via RFC 4193's randomization procedure, just a memorable placeholder picked to visibly pair with the IPv4 10.168.x.x scheme. Since ULA's whole collision-avoidance property depends on the /48 actually being randomly chosen (not on it being memorable), a real generator was worth using.

Replaced with fdf0:f243:f06f::/48 (user-provided, from an external ULA generator). Confirmed with the user (AskUserQuestion, two options offered) how to keep the old scheme's recognizable 168 stamp: the /48 alone consumes all three "identity" hextets a ULA prefix has room for, leaving only the subnet-id (4th) hextet — the same one the per-network auto- allocation index already lived in — with nowhere left to also place a fixed marker without colliding with either the random prefix or the index. Chosen answer: offset the per-network index by a constant 168 rather than matching IPv4's index number-for-number as before. First auto-allocated network's IPv6 block is now fdf0:f243:f06f:168::/64 (paired with IPv4's 10.168.0.0/24), second is fdf0:f243:f06f:169::/64 (paired with 10.168.1.0/24), and so on — deterministic and still visibly project-stamped, just via a constant offset instead of an identical digit. network_subnet.cpp's ipv6_ula_prefix48/ipv6_subnet_id_base constants hold the new prefix and the 168 offset respectively. Verified as root via the doas rule: two freshly created extern networks got fdf0:f243:f06f:168::/64 and fdf0:f243:f06f:169::/64 exactly as expected, correctly paired with 10.168.0.0/24/10.168.1.0/24.

Resolved: container-side tap device intermittently not immediately visible

Trigger: a user's own log from the real target device showed nsenter --net=/proc/<ns_pid>/ns/net -- ip addr add 10.168.0.2/24 dev eth0 failing with "Cannot find device \"eth0\"" immediately after the relay had already created it — confirmed by hand that retrying the whole session a few times eventually let it succeed.

First fix tried, later superseded (see next section): network_join.cpp's join_one_network() retried (bounded, ~500ms, quiet until final success/give-up) the three steps that touch the just-created container interface — IPv4 address, IPv6 address, bringing it up — via a run_with_retry() instead of the plain run() used elsewhere.

A tempting "fix" investigated and ruled out by direct A/B testing on the dev machine, not just reasoned about: the obvious first instinct — have the relay itself self-verify the device is visible (a same-process check, immediately after creating it, before ever reporting success back) — was tried first, in network_tap_relay.cpp. It made things categorically worse: it made the container-side device permanently invisible to every external nsenter afterward, 100% reproducibly (confirmed with a 10-second retry budget — never once became visible), on a mechanism that had otherwise worked correctly and instantly, with zero retries needed, on every real session tested earlier the same day — including a from-scratch self-test reproduction (self_test.cpp) that had passed reliably many times before this one change, and immediately went back to passing reliably once the change was reverted. Root cause not fully understood (something about forking a subprocess that inherits the tap fd — deliberately not O_CLOEXEC — while still holding it open, immediately after device creation, appears to corrupt the device's external visibility on this kernel specifically), but the lesson that survived into the final fix is unambiguous: never add an internal, same-process/fd-holding self-check to the relay.

Resolved: tap devices need to be created persistently, not tied to the relay's own fd lifetime

Trigger: the retry fix above turned out to be insufficient — a further round of real-device testing (a second run.log) showed a different failure pattern: ip addr add against the container-side device would sometimes succeed, only for the very next command against that same device (ip link set eth0 up) to fail with "Cannot find device", exhausting every retry. The user's own diagnosis, confirmed correct: the device wasn't merely slow to become visible after creation (the earlier theory) — it was actually disappearing on its own, on this kernel, independent of anything this project's own code was doing to it. This is consistent with the underlying device having been created via a plain ioctl(fd, TUNSETIFF, &ifr) with no IFF_PERSIST flag: on that kernel, its lifetime seems to be tied to something more fragile than "the one fd that created it stays open" (the relay process never closes or re-opens its own fds around any of this) — never fully root-caused, and not worth chasing further once a structurally different approach removed the whole class of symptom.

Fix, per the user's own suggestion: stop relying on ioctl(TUNSETIFF) alone to create the device at all. Both the host-side and container-side tap devices are now created ahead of time by an external ip tuntap add dev <name> mode tap command (create_persistent_tap(), network_tap_relay.cpp) — the same technique QEMU/libvirt use to let an unprivileged process attach to a tap device someone else set up — and only attached to afterward via the existing open_tap()'s open("/dev/net/tun")

  • ioctl(TUNSETIFF) call (unchanged; it no longer creates, only opens an fd onto an already-existing device). A device created this way is a first- class, persistent netdevice from the kernel's point of view, with no tie to any single fd or process at all — the same reason ip tuntap add/ip link add ... type veth never need an owning process to stay alive either.

Consequence: teardown is no longer automatic. Since neither device disappears on its own once the relay stops, stop_tap_relay() now explicitly ip link dels the host-side device after reaping the relay process (wrapped via wrap_for_network() to reach wherever it lives — host root for extern, the network's own persistent namespace for intern). The container-side device needs no equivalent step: it lives inside the container's own network namespace, which the kernel already tears down (taking every interface inside it along, persistent or not) once the session itself ends — nothing new required there. The crash-orphan sweep (record_tap_relays()/clean_stale_tap_relays()) was extended the same way: its state file now also records each relay's network kind/name (not just its pid and host-side device name), so a sweep for a crashed session can reconstruct a NetworkEntry and reach the right namespace to remove the orphaned host-side device, not just kill the orphaned relay process.

All retry logic from the previous fix was dropped, per the user's own explicit instruction, once persistent creation removed the underlying disappearing-device problem it was compensating for: network_join.cpp's run_with_retry() is gone (reverted to the plain run() used everywhere else), and self_test.cpp's matching wait_for_container_device_visible() retry helper is gone too.

self_test.cpp's own expectations updated accordingly: the test used to assert that both the host-side and container-side devices vanish on their own once stop_tap_relay() stops the relay — exactly the assumption this fix disproves. It now asserts the host-side device is gone (the new explicit ip link del step) while the container-side device is still present (correctly persistent, since only the relay stopped, not the container's own namespace) — the namespace itself is destroyed moments later, at the very end of the test, when its throwaway holder process is killed.

Verified end-to-end on this dev machine, --no-veth forcing the fallback path (root, via the scoped doas rule): a fresh extern network's container repeatedly used its tap-relay-backed eth0 across several commands in a row (ip link show, ip addr show, ping) with no disappearance; gateway ping (0% loss) and outside/internet ping to 8.8.8.8 (0% loss) both worked; a second round of ip link show/ip addr show after the pings still saw the same device correctly. Session cleanup left no leftover host-side tap/veth devices behind (only the bridge itself, which is deliberately left standing per this project's reboot-reconciliation design). -t/--test's own tap-relay create/attach/teardown case, updated as above, passes reliably across repeated runs.

Explicitly out of scope for now

  • Rootless networking. An earlier draft of this design considered a hybrid strategy (real bridge+veth when root, a simpler shared-network- namespace fallback when rootless, mirroring --kill's multi-strategy pattern). Shelved: root-only is sufficient for the actual target device today, and the rootless fallback has real limitations (single-network membership only, no per-container port isolation on it) not worth building before there's an actual rootless use case.
  • nftables backend. iptables only, see above.