net was deliberately excluded from exec_session.cpp's joinable_namespaces list, written back when this project never isolated networking at all. Now that -r/--run sometimes does (whenever -n/--network was used), -x/--exec'ing into such a session saw the host's own network stack instead of the container's -- confirmed directly: it showed the host's unrelated listening ports and couldn't reach the container's own service on 127.0.0.1. Fixed by joining net the same way -x/--exec already joins mnt/uts/ipc/pid/cgroup/user when they differ from the caller's own -- not required, so a session with no isolated net namespace (never joined any network) is unaffected, the entry is just skipped like any other identical-to-ours namespace. Verified as root (via a scoped doas rule): execing into a session joined to an extern network now correctly shows its own eth0 and reaches its own service on 127.0.0.1; execing into a plain session with no -n is unaffected.
18 KiB
Network isolation design
Status: all six commits landed (see "Implementation plan: commit sequence"
below for what shipped, including corrections found by testing along the
way). Rootless networking and an nftables backend remain deliberately out of
scope (see below). Captured 2026-08-30 on the networking branch.
Goal
Add persistent, named networks that -r/--run containers can join, similar
in spirit to how -v/--volume already creates persistent named volumes:
externnetworks: reachable from/to the host's real network.internnetworks: only reachable by other containers that joined the same network, never from the host or outside.-r/--runtakes one or more-n/--network <name>flags to join networks.-p [<network>:]<host-port>:<container-port>forwards a host port into a container on one of its extern networks.
Why not slirp4netns
slirp4netns (and pasta) exist specifically to provide networking to an
unprivileged network namespace without CAP_NET_ADMIN or host iptables
access — a strict 1-to-1 usermode-NAT translator between the host and exactly
one namespace. Neither tool bridges multiple namespaces together; Podman's
own rootless multi-container networks get inter-container connectivity by
putting every container of that network into one shared namespace holding
a real Linux bridge, with only that one namespace getting a slirp4netns/
pasta instance for outside access.
This project's target device is Android and normally has root available
(slocker-lite-priv-drop's whole reason to exist is making root-mode
-r/--run first-class). Root-only for now — rootless networking is
explicitly deferred as its own future task. Given root is assumed, the
usermode-NAT workaround buys nothing: native Linux bridge/veth/iptables is
faster (kernel-native routing/NAT, no usermode packet copy), simpler (no
extra long-lived process per network to babysit), and is exactly what
Docker/Podman themselves do when running as root. slirp4netns/pasta are
dropped from the design entirely; the target device doesn't need either
installed for this feature.
bwrap itself has no "join an existing network namespace" flag (confirmed
via bwrap --help): it has --userns FD/--pidns FD (join-by-FD) for
user/pid, but only --unshare-net (always create new) for networking. To
put a container into a specific persistent namespace, the join happens the
same way this project already joins the rootless containers-storage
mount's namespace: wrap the bwrap invocation in nsenter --net=<path>
first, and don't ask bwrap for --unshare-net on that invocation.
bwrap.cpp's existing wrap_for_root_namespace() is precisely this pattern
already (for mnt/user) and is directly reusable for net.
Mechanism
Both extern and intern networks use the same core mechanism — a Linux
bridge, with one veth pair per joined container (container-side end moved
into that container's own, still separately---unshare-net'ed namespace; the
bridge-side end attached to the network's bridge). The only structural
difference is where the bridge lives:
- extern: the bridge lives directly in the host's own root network
namespace, so it already has a path outside via the host's real routing.
Needs
net.ipv4.ip_forward=1(and the IPv6 forwarding sysctl, if IPv6 is enabled for that network) plus one iptablesMASQUERADE(SNAT) rule for the bridge's subnet — the same patterndocker0uses. - intern: the bridge lives inside its own dedicated, free-standing
namespace (persistent the way
ip netns addkeeps a namespace alive with no process in it — bind-mounting itsns/netfile to a path that outlives any one process). No route out exists at all: a real structural isolation boundary, not merely "extern but without a NAT rule" (which would still be reachable from the host itself).
Each container always keeps its own private net namespace, so its own loopback and interface set never clash with another container's. Joining N networks is just N veth pairs into that one namespace — genuine multi-network membership falls out naturally, no namespace-sharing tricks needed.
Port forwarding (-p)
No slirp4netns hostfwd API to lean on, so this is implemented directly:
- An iptables
DNATrule inPREROUTING:--dport <host-port> -j DNAT --to-destination <container-veth-ip>:<container-port>. - A
FORWARDACCEPTrule for the bridge's subnet — needed because a defaultFORWARD DROPpolicy (plausible on a stock Android kernel) would otherwise silently eat the forwarded traffic, the same gap Docker itself works around.
Syntax: -p [<network>:]<host-port>:<container-port>. <network> is
optional — when omitted, resolves to whichever single extern network the
container joined; an error (not a silent guess) if the container joined more
than one extern network and didn't disambiguate.
Subnet / IP allocation
- IPv4 auto-allocates a
/24starting at10.168.0.0/24, incrementing per network (10.168.1.0/24,10.168.2.0/24, ...), with an optional--subnet <cidr>override at-n/--networkcreation time. - IPv6 is a per-network on/off option, defaulting to enabled. When on,
also auto-allocates a ULA
/64alongside the IPv4 block from a matching incrementing base (proposed:fd00:168:0:1::/64,fd00:168:0:2::/64, ... — mirroring the168from the v4 base so the two are visibly paired), with a--subnet6 <cidr>override.--no-ipv6at creation time skips both the v6 allocation and anyip6tables/IPv6-forwarding setup for that network entirely.
CLI surface
-n/--network, dual-purpose like -v/--volume (alone creates/manages a
network; combined with -r joins it) — with a simpler token shape than
-v's, since a network join has no equivalent of a volume's
container-mount-path second argument:
- Alone:
-n <name> --extern|--intern [--subnet <cidr>] [--no-ipv6] [--subnet6 <cidr>]—<name>is-n's ownrequired_argument(a single token, standardgetopt_long); kind and the subnet/ipv6 options are ordinary separate flags rather than additional positional tokens. - With
-r:-n <name>is repeatable, one token (just the name) per occurrence — joins that network. No membership limit. --list-networks/--delete-network <name>round out the set, mirroring--list-volumes/--delete-volume.
Config file schema
New top-level networks section in config.yaml, parallel to volumes
(config_file.{h,cpp}) — per entry: name, kind (extern/intern),
subnet, ipv6 (bool), subnet6 (if ipv6). create_volume_command() /
list_volumes_command() / delete_volume_command() (commands.cpp) are the
direct templates to follow, including the shared tab-aligned listing helper
for --list-networks.
A network's config-file entry is the durable source of truth, the same way a
volume's directory path is. The live host-side state (namespace file for
intern, bridge, veths, iptables rules) doesn't survive a reboot and is
reconciled lazily and transparently — checked and, if missing/stale,
recreated from the config entry — the first time anything needs it after a
reboot (no explicit "start" step; the first -r/--run -n <name> or -p
after boot just works).
Tooling
Shell out to iptables/ip6tables (via process.h's existing
run_process(), matching how containers-storage/bwrap/fuse-overlayfs
are already invoked) for MASQUERADE/DNAT/FORWARD rules, and to ip for
bridge/veth/namespace management. Both need an on-device availability check
the same way check_required_dependencies() already gates on
containers-storage/bwrap.
iptables only, for now — confirmed available on the real target device;
nftables is not currently installed there. nft support is real future
work (its own task, once it actually matters — e.g. a device that only ships
nft), not built preemptively as a dual-backend abstraction here.
Host-global state discipline
Unlike a bwrap session (self-contained via its own namespaces),
bridges/veths/iptables rules are host-global, named, persistent resources
that outlive any single process. Needs the same create-on-start/
remove-on-stop discipline already used for pid locks (pid_file.{h,cpp}) and
cgroups (session_cgroup.{h,cpp}), plus a --clean-processes-style sweep for
anything orphaned by a crash, so a dead slocker-lite doesn't leave stray
bridges/veths/iptables rules behind forever.
Implementation plan: commit sequence
The whole feature is too large for one commit (unlike, say, --kill, which
landed as a single commit despite touching several new files). Split into six
commits, each a coherent, independently buildable and manually verifiable
unit, in dependency order. Docs (CLAUDE.md/README.md) get updated within
each commit, matching this branch's existing practice — not saved for a final
pass.
-
Config schema + subnet/IPv6 allocation +
-n/--network create/list/delete(config-only, no host-side effects yet)config_file.{h,cpp}: newNetworkEntry {name, kind, subnet, ipv6, subnet6}(kind=extern/intern), anetworkssection — directly parallel toVolumeEntry/volumes.- New
network_subnet.{h,cpp}: IPv4/24auto-allocation starting at10.168.0.0/24incrementing per existing network,--subnetoverride; paired IPv6 ULA/64auto-allocation (fd00:168:0:N::/64) when enabled (default),--subnet6override,--no-ipv6to skip. Pure allocation logic against the already-loaded config's existing networks — no kernel/ip/iptablescalls in this commit. cli_args.{h,cpp}:Mode::network/list_networks/delete_network,-n/--network(required_argument, single token = name; separate--extern/--intern/--subnet/--no-ipv6/--subnet6flags for the create case),--list-networks,--delete-network <name>— directly mirroring-v/--volume's existing three-mode shape in the same file.commands.cpp:create_network_command()/list_networks_command()/delete_network_command()— mirroringcreate_volume_command()/list_volumes_command()/delete_volume_command(), including the shared tab-aligned listing helper.- Verify:
-n mynet --extern,-n other --intern --no-ipv6,--list-networksshows both with correct kind/subnet,--subnet/--subnet6overrides land correctly inconfig.yaml,--delete-networkremoves an entry. No bridges/namespaces/iptables rules exist yet — purely config bookkeeping, same as a freshly-created volume before it's ever mounted.
-
Persistent network namespace primitives (generic infra for
internnetworks)- New
persistent_netns.{h,cpp}: create/find/remove a persistent network namespace kept alive with no process in it, the wayip netns adddoes (bind-mount a fresh namespace'sns/netonto a path that outlives the creating process) — narrow, reusable infra, nointern/externbranching or bridge logic here (parallels howsession_cgroup.{h,cpp}stayed narrowly scoped to cgroup mechanics only). - Not wired into
-n/--networkyet in this commit. - Verify: a small manual exercise (or a
-t/--testaddition) creating a persistent namespace, confirming it survives after the creating process exits, then removing it.
- New
-
Bridge provisioning for a network (idempotent — this is also the reboot- reconciliation mechanism, not a separate later step)
- New
network_bridge.{h,cpp}: given aNetworkEntry, ensure its bridge exists and is configured — creating it if missing (idempotent, so this doubles as "reconcile after reboot" with no separate code path):extern: bridge in the host's own root namespace; assign it the gateway IP from the network's subnet;net.ipv4.ip_forward=1(+ IPv6 forwarding sysctl ifipv6); one iptablesMASQUERADErule for the subnet (ip6tablestoo, ifipv6).intern: bridge inside its own dedicatedpersistent_netns.hnamespace (commit 2); gateway IP assigned; no forwarding, no NAT rule — no route out at all.
- Wire this "ensure provisioned" call into
create_network_command()(so creating a network actually stands up its bridge immediately) — later commits also call it lazily before a join, covering the reboot case. check_required_dependencies()-style availability check added forip/iptables(andip6tableswhen needed), alongside the existingcontainers-storage/bwrapcheck.- Verify:
-n mynet --externproduces a real bridge with the expected gateway IP,ip_forwardenabled, and a matchingMASQUERADErule (ip link show,iptables -t nat -L); aninternnetwork's bridge exists in its own namespace with no such rule. Delete/recreate a network's config entry, delete its bridge by hand (ip link del), then trigger provisioning again (e.g. re-running--network createor the first join in commit 4) and confirm it comes back.
- New
-
Joining networks at
-r/--runtime: veth creation, IP assignment, route- Repeatable
-n <name>with-r/--run(cli_args.cpp, same repeatable-with--rshape-v/--volumealready has). commands.cpp'srun_container(): oncebwrap's pid (and viaresolve_namespace_pid()-style lookup,sandbox_process.h, its actual net namespace) is known — same timing hookon_bwrap_pid_known/on_startalready provides for session locks/cgroups (bwrap.cpp) — for each joined network: ensure it's provisioned (commit 3, covers reboot recreation), create a veth pair, move the container-side end into the container's net namespace, attach the bridge-side end, assign the container's veth an IP from the subnet, and (for anexternjoin) set it as the default route.- This is the core connectivity commit — no veth pairs exist before it, regardless of how many networks are configured/joined.
- Verify: two containers joined to the same
internnetwork can ping each other and cannot reach the host or outside; a container joined to anexternnetwork can reach the outside (and the host cannot reach it without commit 5's port forwarding); a container joined to both loses neither path (two interfaces, both functional).
- Repeatable
-
-pport forwardingcli_args.cpp:-p [<network>:]<host-port>:<container-port>,<network>optional (resolves to the container's soleexternnetwork; error if ambiguous).- New
port_forward.{h,cpp}: add/remove the iptablesDNAT(PREROUTING) +FORWARD ACCEPTrule pair for one mapping, tied to the container's own session lifecycle the same create-on-start/ remove-on-stop waypid_file.{h,cpp}/session_cgroup.{h,cpp}already are. - Verify:
-p 8080:80against a container on an extern network answering on port 80 is reachable viacurl localhost:8080from the host; the rule is gone after the container exits. - Landed with two real corrections found by testing (see
CLAUDE.md'sport_forward.{h,cpp}entry for the full detail): theDNATrule needs bothPREROUTINGandOUTPUT(locally-generated traffic never traversesPREROUTING); andcurl localhost:<port>specifically still doesn't work even so (NAT hairpinning — the container sees an inbound packet claiming a loopback source on a non-loopback interface and drops it as martian) — verified instead viacurl <host's real IP>:<port>, the actually-relevant path for real clients. Also surfaced, unrelated to-pitself but found while testing it:-x/--execdidn't join thenetnamespace (written back when this project never isolated networking at all), so it saw the host's network stack, not a network-isolated session's own — fixed in a follow-up commit (exec_session.cpp, seeCLAUDE.md's own entry for that file).
-
Crash-orphan cleanup sweep
- Extend
--clean-processes(or add a dedicated--clean-networks, whichever reads better once this is reached) to find and remove bridges/veths/iptables rules left behind by aslocker-litethat died before its own teardown ran — mirroringclean_stale_sessions()(pid_file.cpp)'s existing stale-pid-file sweep, but for host-global network state instead of pid files. - Verify: kill
-9a running-r/--runsession mid-flight (bypassing its normal cleanup), confirm the orphaned veth/iptables rule is detected and removed by the sweep, and that a still-running session's state is left untouched. - Landed narrower in scope than the bullet above once the actual orphan
surface was worked out (see
CLAUDE.md'sport_forward.{h,cpp}entry for the full detail): veths need no sweep at all (the kernel tears down an entire pair once either end's namespace is destroyed — never survives a crash), and bridges/persistent namespaces are deliberately meant to always outlive any one session (that's the whole point of the reboot-reconciliation design, not something a crash changes). Only-p's iptables rules — host-global, named, with no automatic teardown — can actually outlive a crashed session, so that's the entire sweep: extended--clean-processes(not a separate flag) withclean_stale_port_forwards(), cross-referencing a small per-session port-forward record file againstlist_sessions()'s own liveness check. Verified via a controlled scratch test rather than a literalkill -9on a root-ownedslocker-liteprocess (not achievable through this session's scopeddoasrule, which only permits runningslocker-liteitself, not arbitrary commands likekill): a fabricated stale record was correctly detected, its removal attempted, and its file cleaned up, while a record matching a real running session was left untouched.
- Extend
Explicitly out of scope for now
- Rootless networking. An earlier draft of this design considered a
hybrid strategy (real bridge+veth when root, a simpler shared-network-
namespace fallback when rootless, mirroring
--kill's multi-strategy pattern). Shelved: root-only is sufficient for the actual target device today, and the rootless fallback has real limitations (single-network membership only, no per-container port isolation on it) not worth building before there's an actual rootless use case. - nftables backend. iptables only, see above.