Files
slocker-lite/docs/networking-design.md
T
ceamac 53b859b7bf Document the tap+relay veth fallback in the design doc
Adds a dedicated section to docs/networking-design.md covering the
tap+relay fallback for veth-less kernels: the trigger (real target
device supports tun/tap but not veth), why tap can't 1:1 replace veth,
the confirmed design (two tap devices + a relay reusing the existing
bridge, replacing veth's earlier N-way-switch-daemon sketch once
bridge support was confirmed available), the four-commit
implementation sequence with what testing actually found (the fd-leak
deadlock, the reverted cgroup fix), and an honest writeup of the
unresolved gateway/outside-reachability gap.

README.md's -n/--network row now also flags that gap directly, next
to the existing NAT-hairpinning limitation note for -p.

This closes out the tap+relay fallback work for now: peer-to-peer
connectivity through it is solid and dev-verified; gateway/outside
reachability needs re-verification on the actual veth-less target
device before being relied on.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-08-30 16:36:43 +00:00

25 KiB

Network isolation design

Status: all six commits landed (see "Implementation plan: commit sequence" below for what shipped, including corrections found by testing along the way). Rootless networking and an nftables backend remain deliberately out of scope (see below). Captured 2026-08-30 on the networking branch.

Goal

Add persistent, named networks that -r/--run containers can join, similar in spirit to how -v/--volume already creates persistent named volumes:

  • extern networks: reachable from/to the host's real network.
  • intern networks: only reachable by other containers that joined the same network, never from the host or outside.
  • -r/--run takes one or more -n/--network <name> flags to join networks.
  • -p [<network>:]<host-port>:<container-port> forwards a host port into a container on one of its extern networks.

Why not slirp4netns

slirp4netns (and pasta) exist specifically to provide networking to an unprivileged network namespace without CAP_NET_ADMIN or host iptables access — a strict 1-to-1 usermode-NAT translator between the host and exactly one namespace. Neither tool bridges multiple namespaces together; Podman's own rootless multi-container networks get inter-container connectivity by putting every container of that network into one shared namespace holding a real Linux bridge, with only that one namespace getting a slirp4netns/ pasta instance for outside access.

This project's target device is Android and normally has root available (slocker-lite-priv-drop's whole reason to exist is making root-mode -r/--run first-class). Root-only for now — rootless networking is explicitly deferred as its own future task. Given root is assumed, the usermode-NAT workaround buys nothing: native Linux bridge/veth/iptables is faster (kernel-native routing/NAT, no usermode packet copy), simpler (no extra long-lived process per network to babysit), and is exactly what Docker/Podman themselves do when running as root. slirp4netns/pasta are dropped from the design entirely; the target device doesn't need either installed for this feature.

bwrap itself has no "join an existing network namespace" flag (confirmed via bwrap --help): it has --userns FD/--pidns FD (join-by-FD) for user/pid, but only --unshare-net (always create new) for networking. To put a container into a specific persistent namespace, the join happens the same way this project already joins the rootless containers-storage mount's namespace: wrap the bwrap invocation in nsenter --net=<path> first, and don't ask bwrap for --unshare-net on that invocation. bwrap.cpp's existing wrap_for_root_namespace() is precisely this pattern already (for mnt/user) and is directly reusable for net.

Mechanism

Both extern and intern networks use the same core mechanism — a Linux bridge, with one veth pair per joined container (container-side end moved into that container's own, still separately---unshare-net'ed namespace; the bridge-side end attached to the network's bridge). The only structural difference is where the bridge lives:

  • extern: the bridge lives directly in the host's own root network namespace, so it already has a path outside via the host's real routing. Needs net.ipv4.ip_forward=1 (and the IPv6 forwarding sysctl, if IPv6 is enabled for that network) plus one iptables MASQUERADE (SNAT) rule for the bridge's subnet — the same pattern docker0 uses.
  • intern: the bridge lives inside its own dedicated, free-standing namespace (persistent the way ip netns add keeps a namespace alive with no process in it — bind-mounting its ns/net file to a path that outlives any one process). No route out exists at all: a real structural isolation boundary, not merely "extern but without a NAT rule" (which would still be reachable from the host itself).

Each container always keeps its own private net namespace, so its own loopback and interface set never clash with another container's. Joining N networks is just N veth pairs into that one namespace — genuine multi-network membership falls out naturally, no namespace-sharing tricks needed.

Port forwarding (-p)

No slirp4netns hostfwd API to lean on, so this is implemented directly:

  • An iptables DNAT rule in PREROUTING: --dport <host-port> -j DNAT --to-destination <container-veth-ip>:<container-port>.
  • A FORWARD ACCEPT rule for the bridge's subnet — needed because a default FORWARD DROP policy (plausible on a stock Android kernel) would otherwise silently eat the forwarded traffic, the same gap Docker itself works around.

Syntax: -p [<network>:]<host-port>:<container-port>. <network> is optional — when omitted, resolves to whichever single extern network the container joined; an error (not a silent guess) if the container joined more than one extern network and didn't disambiguate.

Subnet / IP allocation

  • IPv4 auto-allocates a /24 starting at 10.168.0.0/24, incrementing per network (10.168.1.0/24, 10.168.2.0/24, ...), with an optional --subnet <cidr> override at -n/--network creation time.
  • IPv6 is a per-network on/off option, defaulting to enabled. When on, also auto-allocates a ULA /64 alongside the IPv4 block from a matching incrementing base (proposed: fd00:168:0:1::/64, fd00:168:0:2::/64, ... — mirroring the 168 from the v4 base so the two are visibly paired), with a --subnet6 <cidr> override. --no-ipv6 at creation time skips both the v6 allocation and any ip6tables/IPv6-forwarding setup for that network entirely.

CLI surface

-n/--network, dual-purpose like -v/--volume (alone creates/manages a network; combined with -r joins it) — with a simpler token shape than -v's, since a network join has no equivalent of a volume's container-mount-path second argument:

  • Alone: -n <name> --extern|--intern [--subnet <cidr>] [--no-ipv6] [--subnet6 <cidr>]<name> is -n's own required_argument (a single token, standard getopt_long); kind and the subnet/ipv6 options are ordinary separate flags rather than additional positional tokens.
  • With -r: -n <name> is repeatable, one token (just the name) per occurrence — joins that network. No membership limit.
  • --list-networks / --delete-network <name> round out the set, mirroring --list-volumes / --delete-volume.

Config file schema

New top-level networks section in config.yaml, parallel to volumes (config_file.{h,cpp}) — per entry: name, kind (extern/intern), subnet, ipv6 (bool), subnet6 (if ipv6). create_volume_command() / list_volumes_command() / delete_volume_command() (commands.cpp) are the direct templates to follow, including the shared tab-aligned listing helper for --list-networks.

A network's config-file entry is the durable source of truth, the same way a volume's directory path is. The live host-side state (namespace file for intern, bridge, veths, iptables rules) doesn't survive a reboot and is reconciled lazily and transparently — checked and, if missing/stale, recreated from the config entry — the first time anything needs it after a reboot (no explicit "start" step; the first -r/--run -n <name> or -p after boot just works).

Tooling

Shell out to iptables/ip6tables (via process.h's existing run_process(), matching how containers-storage/bwrap/fuse-overlayfs are already invoked) for MASQUERADE/DNAT/FORWARD rules, and to ip for bridge/veth/namespace management. Both need an on-device availability check the same way check_required_dependencies() already gates on containers-storage/bwrap.

iptables only, for now — confirmed available on the real target device; nftables is not currently installed there. nft support is real future work (its own task, once it actually matters — e.g. a device that only ships nft), not built preemptively as a dual-backend abstraction here.

Host-global state discipline

Unlike a bwrap session (self-contained via its own namespaces), bridges/veths/iptables rules are host-global, named, persistent resources that outlive any single process. Needs the same create-on-start/ remove-on-stop discipline already used for pid locks (pid_file.{h,cpp}) and cgroups (session_cgroup.{h,cpp}), plus a --clean-processes-style sweep for anything orphaned by a crash, so a dead slocker-lite doesn't leave stray bridges/veths/iptables rules behind forever.

Implementation plan: commit sequence

The whole feature is too large for one commit (unlike, say, --kill, which landed as a single commit despite touching several new files). Split into six commits, each a coherent, independently buildable and manually verifiable unit, in dependency order. Docs (CLAUDE.md/README.md) get updated within each commit, matching this branch's existing practice — not saved for a final pass.

  1. Config schema + subnet/IPv6 allocation + -n/--network create/list/delete (config-only, no host-side effects yet)

    • config_file.{h,cpp}: new NetworkEntry {name, kind, subnet, ipv6, subnet6} (kind = extern/intern), a networks section — directly parallel to VolumeEntry/volumes.
    • New network_subnet.{h,cpp}: IPv4 /24 auto-allocation starting at 10.168.0.0/24 incrementing per existing network, --subnet override; paired IPv6 ULA /64 auto-allocation (fd00:168:0:N::/64) when enabled (default), --subnet6 override, --no-ipv6 to skip. Pure allocation logic against the already-loaded config's existing networks — no kernel/ ip/iptables calls in this commit.
    • cli_args.{h,cpp}: Mode::network/list_networks/delete_network, -n/--network (required_argument, single token = name; separate --extern/--intern/--subnet/--no-ipv6/--subnet6 flags for the create case), --list-networks, --delete-network <name> — directly mirroring -v/--volume's existing three-mode shape in the same file.
    • commands.cpp: create_network_command()/list_networks_command()/ delete_network_command() — mirroring create_volume_command()/ list_volumes_command()/delete_volume_command(), including the shared tab-aligned listing helper.
    • Verify: -n mynet --extern, -n other --intern --no-ipv6, --list-networks shows both with correct kind/subnet, --subnet/ --subnet6 overrides land correctly in config.yaml, --delete-network removes an entry. No bridges/namespaces/iptables rules exist yet — purely config bookkeeping, same as a freshly-created volume before it's ever mounted.
  2. Persistent network namespace primitives (generic infra for intern networks)

    • New persistent_netns.{h,cpp}: create/find/remove a persistent network namespace kept alive with no process in it, the way ip netns add does (bind-mount a fresh namespace's ns/net onto a path that outlives the creating process) — narrow, reusable infra, no intern/extern branching or bridge logic here (parallels how session_cgroup.{h,cpp} stayed narrowly scoped to cgroup mechanics only).
    • Not wired into -n/--network yet in this commit.
    • Verify: a small manual exercise (or a -t/--test addition) creating a persistent namespace, confirming it survives after the creating process exits, then removing it.
  3. Bridge provisioning for a network (idempotent — this is also the reboot- reconciliation mechanism, not a separate later step)

    • New network_bridge.{h,cpp}: given a NetworkEntry, ensure its bridge exists and is configured — creating it if missing (idempotent, so this doubles as "reconcile after reboot" with no separate code path):
      • extern: bridge in the host's own root namespace; assign it the gateway IP from the network's subnet; net.ipv4.ip_forward=1 (+ IPv6 forwarding sysctl if ipv6); one iptables MASQUERADE rule for the subnet (ip6tables too, if ipv6).
      • intern: bridge inside its own dedicated persistent_netns.h namespace (commit 2); gateway IP assigned; no forwarding, no NAT rule — no route out at all.
    • Wire this "ensure provisioned" call into create_network_command() (so creating a network actually stands up its bridge immediately) — later commits also call it lazily before a join, covering the reboot case.
    • check_required_dependencies()-style availability check added for ip/ iptables (and ip6tables when needed), alongside the existing containers-storage/bwrap check.
    • Verify: -n mynet --extern produces a real bridge with the expected gateway IP, ip_forward enabled, and a matching MASQUERADE rule (ip link show, iptables -t nat -L); an intern network's bridge exists in its own namespace with no such rule. Delete/recreate a network's config entry, delete its bridge by hand (ip link del), then trigger provisioning again (e.g. re-running --network create or the first join in commit 4) and confirm it comes back.
  4. Joining networks at -r/--run time: veth creation, IP assignment, route

    • Repeatable -n <name> with -r/--run (cli_args.cpp, same repeatable-with--r shape -v/--volume already has).
    • commands.cpp's run_container(): once bwrap's pid (and via resolve_namespace_pid()-style lookup, sandbox_process.h, its actual net namespace) is known — same timing hook on_bwrap_pid_known/ on_start already provides for session locks/cgroups (bwrap.cpp) — for each joined network: ensure it's provisioned (commit 3, covers reboot recreation), create a veth pair, move the container-side end into the container's net namespace, attach the bridge-side end, assign the container's veth an IP from the subnet, and (for an extern join) set it as the default route.
    • This is the core connectivity commit — no veth pairs exist before it, regardless of how many networks are configured/joined.
    • Verify: two containers joined to the same intern network can ping each other and cannot reach the host or outside; a container joined to an extern network can reach the outside (and the host cannot reach it without commit 5's port forwarding); a container joined to both loses neither path (two interfaces, both functional).
  5. -p port forwarding

    • cli_args.cpp: -p [<network>:]<host-port>:<container-port>, <network> optional (resolves to the container's sole extern network; error if ambiguous).
    • New port_forward.{h,cpp}: add/remove the iptables DNAT (PREROUTING) + FORWARD ACCEPT rule pair for one mapping, tied to the container's own session lifecycle the same create-on-start/ remove-on-stop way pid_file.{h,cpp}/session_cgroup.{h,cpp} already are.
    • Verify: -p 8080:80 against a container on an extern network answering on port 80 is reachable via curl localhost:8080 from the host; the rule is gone after the container exits.
    • Landed with two real corrections found by testing (see CLAUDE.md's port_forward.{h,cpp} entry for the full detail): the DNAT rule needs both PREROUTING and OUTPUT (locally-generated traffic never traverses PREROUTING); and curl localhost:<port> specifically still doesn't work even so (NAT hairpinning — the container sees an inbound packet claiming a loopback source on a non-loopback interface and drops it as martian) — verified instead via curl <host's real IP>:<port>, the actually-relevant path for real clients. Also surfaced, unrelated to -p itself but found while testing it: -x/--exec didn't join the net namespace (written back when this project never isolated networking at all), so it saw the host's network stack, not a network-isolated session's own — fixed in a follow-up commit (exec_session.cpp, see CLAUDE.md's own entry for that file).
  6. Crash-orphan cleanup sweep

    • Extend --clean-processes (or add a dedicated --clean-networks, whichever reads better once this is reached) to find and remove bridges/veths/iptables rules left behind by a slocker-lite that died before its own teardown ran — mirroring clean_stale_sessions() (pid_file.cpp)'s existing stale-pid-file sweep, but for host-global network state instead of pid files.
    • Verify: kill -9 a running -r/--run session mid-flight (bypassing its normal cleanup), confirm the orphaned veth/iptables rule is detected and removed by the sweep, and that a still-running session's state is left untouched.
    • Landed narrower in scope than the bullet above once the actual orphan surface was worked out (see CLAUDE.md's port_forward.{h,cpp} entry for the full detail): veths need no sweep at all (the kernel tears down an entire pair once either end's namespace is destroyed — never survives a crash), and bridges/persistent namespaces are deliberately meant to always outlive any one session (that's the whole point of the reboot-reconciliation design, not something a crash changes). Only -p's iptables rules — host-global, named, with no automatic teardown — can actually outlive a crashed session, so that's the entire sweep: extended --clean-processes (not a separate flag) with clean_stale_port_forwards(), cross-referencing a small per-session port-forward record file against list_sessions()'s own liveness check. Verified via a controlled scratch test rather than a literal kill -9 on a root-owned slocker-lite process (not achievable through this session's scoped doas rule, which only permits running slocker-lite itself, not arbitrary commands like kill): a fabricated stale record was correctly detected, its removal attempted, and its file cleaned up, while a record matching a real running session was left untouched.

TUN/TAP fallback for veth-less kernels

Trigger: the real target device's kernel supports tun/tap (CONFIG_TUN — Android needs this for VpnService-based VPN apps) but not veth (CONFIG_VETH, commonly stripped from mobile kernels), so -n --extern/--intern as designed above (a veth pair per join) simply can't work there at all — ip link add ... type veth ... fails outright. Bridge support was separately confirmed working on this same device, which rules out the more complex fallback this section originally considered (see git log on this file for the superseded sketch: a per-network userspace Ethernet switch with no bridge dependency at all) in favor of a much smaller design.

Why tap can't just replace veth 1:1: a veth pair is two real kernel netdevices, switched between (or into a bridge) entirely by the kernel with zero userspace involvement. A tap device only has one kernel-side netdevice — the other "end" is a raw-Ethernet-frame file descriptor that only a userspace process can read/write, so there's no second kernel endpoint to attach to a bridge. This is exactly why slirp4netns/QEMU's own tap networking need a userspace process on the fd side at all.

Design, confirmed and implemented: per network-join, two tap devices + one small relay process that copies bytes 1:1 between them — a direct functional substitute for one veth pair, reusing the existing bridge as the switching fabric so provision_bridge() needs no changes at all:

  • A host-side tap device, created wherever the network's bridge lives and enslaved to it — exactly veth's host-side role.
  • A container-side tap device, created directly inside the container's own namespace, named eth<N> from the start — no peer-name-then- rename dance needed, unlike veth.
  • A relay process holding both fds open, copying raw Ethernet frames bidirectionally between them for as long as it runs. This is the "veth wire," just implemented once in userspace instead of by the kernel.

Once the container-side tap exists as eth<N>, everything downstream — IP assignment, routes, the address handed to -p — is completely unchanged; only the interface-creation step is swapped. Strategy selection is per-network, per-join: should_use_veth(network) = network.veth && probe_veth_support(), mirroring the existing kernel-capability-vs-policy split namespace_policy_enabled() (bwrap.cpp) already uses for --unshare-xxx. --no-veth at network-creation time forces the fallback even on a veth-capable kernel — how this was actually tested, since the real target device wasn't available during development.

Implementation plan: commit sequence

Landed as four commits (a fifth, this doc update, closes it out) — see CLAUDE.md's own entries (network_bridge.{h,cpp}, network_tap_relay.{h,cpp}, network_join.{h,cpp}, self_test.{h,cpp}) for full file-by-file detail:

  1. Veth capability probe + --no-veth flag, no relay yet. probe_veth_support() (fork, unshare(CLONE_NEWNET) into a throwaway namespace, try ip link add ... type veth ... there — the same kernel-capability-probing shape bwrap.cpp's own kernel_supports_namespace() already uses); NetworkEntry::veth + YAML round-trip; --no-veth CLI flag. Verified: probe_veth_support() returns true on this dev machine (a real veth pair is created successfully); --no-veth persists veth: false.
  2. network_tap_relay.{h,cpp}: relay creation/loop/teardown, standalone. Verified via a new -t/--test case: a throwaway bridge + throwaway network namespace, confirming the host-side tap attaches to the bridge, the container-side tap appears inside the target namespace with the requested name, and — the biggest assumption going in — both devices disappear on their own once the relay is stopped, no explicit ip link del needed (neither is created with IFF_PERSIST).
  3. Wire into join_one_network()/join_networks(). JoinedNetwork gains an optional relay handle; run_container() collects and stops them after run_bwrap() returns, mirroring -p's own active_port_forwards handling exactly. Two real bugs found here, not assumed: the relay child, unlike every other forked child in this project, never exec()s, so it inherited (and never closed) a live copy of daemonize.cpp's own report-pipe write end, hanging -D combined with -n indefinitely until fixed with an explicit close_inherited_fds(); and a first-attempt fix making --kill reach the relay directly (adding its pid to the session's own cgroup) was reverted after it caused a different bug (the cgroup's own removal, which happens before run_container() gets to stop the relay, started failing with EBUSY) — the ordinary flow already stops the relay correctly on its own, so the added complexity wasn't worth it. Verified end-to-end: two containers on a --no-veth extern network got distinct addresses via two tap+relay pairs (no veth at all) and pinged each other with 0% packet loss, repeatably.
  4. Crash-orphan sweep. record_tap_relays()/clean_stale_tap_relays(), the direct structural analog of -p's own sweep, wired into --clean-processes. Verified the same way the port-forward sweep was: a real rootless session's pid alongside a hand-written matching record (left untouched) and a fabricated stale one (correctly swept).

Known gap: gateway/outside reachability unconfirmed

Confirmed by testing, not yet root-caused. Peer-to-peer connectivity through the tap+relay fallback is solid (verified above). Reaching the network's own gateway IP — and, in turn, the real outside through NAT — is not: neither ICMP nor a TCP wget ever got a response, despite ARP resolving correctly (the container's own ip neigh shows a REACHABLE entry with the gateway's real MAC, ruling out an L2/relay-framing problem). The identical bridge/subnet/host reached via veth instead of this fallback works perfectly — ruling out every environment-level explanation (host firewall, rp_filter, this dev sandbox's own networking) that would otherwise affect both paths equally, since those apply regardless of which mechanism connects the container. rp_filter=0 (tried at the host-tap, bridge, and global all scope) did not fix it. Diagnosing further needs host-level tools (tcpdump, direct iptables/sysctl inspection) this project's doas-scoped root access during development didn't permit (restricted to running slocker-lite itself, no other commands). This needs re-verification — ideally on the actual veth-less target device, where the environment differs and this dev sandbox's own unidentified cause may simply not apply — before extern/outside connectivity through this fallback is relied on. intern networks (peer-to-peer only, no gateway involved at all) are unaffected by this gap.

Explicitly out of scope for now

  • Rootless networking. An earlier draft of this design considered a hybrid strategy (real bridge+veth when root, a simpler shared-network- namespace fallback when rootless, mirroring --kill's multi-strategy pattern). Shelved: root-only is sufficient for the actual target device today, and the rootless fallback has real limitations (single-network membership only, no per-container port isolation on it) not worth building before there's an actual rootless use case.
  • nftables backend. iptables only, see above.