Files
slocker-lite/CLAUDE.md
T
ceamac 37ab2fcb1e Document the multi-network join race and its fix
Records the strace-based root-cause investigation for the bug where
joining 2+ networks in one -r/--run left every network after the first
permanently unreachable, and the retry-on-EIO/ENETDOWN fix applied in
network_tap_relay.cpp, matching the level of detail already recorded for
the extern-connectivity investigation above it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-08-31 17:00:25 +00:00

1879 lines
134 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## Project state
`slocker-lite` (C++20, built with Meson) mounts an OCI Image Layout tar (`oci-layout` +
`index.json` + `blobs/sha256/*`, as produced by `skopeo`/`podman save --format
oci-archive`/modern `docker save`) using `containers-storage` and `fuse-overlayfs`, then
(via `-r/--run`) runs a sandboxed command against it with `bwrap`. The real deployment
target is Android with a stock kernel, where `podman`/`docker` don't run (missing
namespace support) and there's no kernel overlayfs (hence `fuse-overlayfs`); `bwrap` is
invoked in "degraded mode" using only whichever `--unshare-xxx` namespaces the running
kernel actually supports. See `README.md` for the human-facing overview (build/usage/
status); this file stays the dense, file-by-file reference. Still early-stage.
Source layout (all under `src/`):
- `main.cpp` — the CLI entry point only, and deliberately tiny (~30 lines):
loads the config file, applies its `global.log-level` (`apply_log_level()`,
`cli_args.h`) before CLI parsing so an explicit `--log-level` can still
override it afterward, calls `parse_args()` (`cli_args.h`) and returns its
exit code immediately if it gives one (covers `-h`/`-V` and every parse
error), otherwise calls `dispatch_command()` (`commands.h`) and returns its
result. All of the actual option-parsing and command logic that used to live
here moved out into `cli_args.{h,cpp}`/`commands.{h,cpp}`/`self_test.{h,cpp}`
(see below) specifically to keep this file from re-growing into a dumping
ground as more commands (docker-compose support, etc.) get added.
- `cli_args.{h,cpp}` — command-line parsing only, nothing else. `Mode` (the
enum of every CLI action) and `ParsedArgs` (everything `parse_args()`
extracts from `argv`) live in the header since `commands.h`'s
`dispatch_command()` consumes them; `print_usage()`/`print_version()`, the
`options` namespace of getopt long-option codes, and the `long_options`
array itself are `.cpp`-local. `parse_args(argc, argv, out)` runs the
`getopt_long` loop plus all of the post-loop validation that used to live at
the top of `main()`: mode/`--volume` interaction (`-v` alone vs. combined
with `-r`, see below), `--group` requires `--user`, no leftover positional
args outside `-r`/`-e`, and (moved here from what used to be inline in the
`Mode::exec` dispatch arm) `-x/--exec <pid>`'s own pid parsing/validation
(`ParsedArgs::exec_pid`, a positive integer or a hard error) and its
trailing-command requirement (`ParsedArgs::command`, required non-empty).
`--kill <pid>` (`ParsedArgs::kill_pid`) shares that same positive-integer
parsing/validation via a small extracted `parse_pid_arg()` helper (`.cpp`-local)
rather than duplicating the `strtol` dance a second time — unlike `-x/--exec`,
it takes no trailing command, so it's simply not added to the leftover-args
exemption list (`Mode::run`/`Mode::exec` only). Returns an exit code `main()`
should return immediately (`0` for `-h`/`-V`,
`1` for any parse error) when set; `nullopt` means `out` is ready for
`dispatch_command()`. `-D/--daemonize` (has a short form; `'D'` was free) is
a plain boolean flag (`ParsedArgs::daemonize_flag`, set in its own `case
'D':`, same pattern as `--no-nsenter`). `--hostname <name>`/`--env
VAR=VALUE`/`--env-file <file>` (all long-option only, `--env`/`--env-file`
both repeatable) are collected here into `ParsedArgs::hostname_flag`/
`env_specs``--env` pushes `{false, optarg}`, `--env-file` pushes `{true,
optarg}` into the *same* ordered `std::vector<EnvSpec>` (not two separate
lists), preserving their exact relative command-line order across both
flags, since `resolve_env_specs()` (`env_spec.{h,cpp}`, see below) needs
that order to let a later one override an earlier one for the same variable
name — actually resolving them happens later, in `commands.cpp`'s
`run_container()`. `-v/--volume` is dual-purpose: used alone it's a
standalone `Mode::volume` request; combined with `-r/--run` it's repeatable
and requests a volume mount instead (resolved later by
`resolve_volume_mount()`, `volume_mount.{h,cpp}`, see below). Since `-v`
must be repeatable with `-r` but each occurrence still takes two
space-separated tokens, the getopt loop doesn't let `'v'` set `Mode`
directly: it accumulates `(spec, path)` pairs into `ParsedArgs::volume_specs`
(consuming the second token manually, with a guard against swallowing the
next flag if there isn't one), and only *after* the loop decides whether
that means one standalone `Mode::volume` call or, together with `-r`,
passes `volume_specs` through unresolved for `dispatch_command()`/
`run_container()` to handle. `-n/--network` (see `docs/networking-design.md`
for the full feature design) is dual-purpose the same way, but simpler: since
a network join has no equivalent of a volume's container-mount-path second
argument, each occurrence is a single `required_argument` token (just the
name) accumulated into `ParsedArgs::network_specs` — no manual second-token
consumption needed, `'n'`'s own `case` just does
`network_specs.push_back(optarg)`. The same post-loop split as `-v` decides
`Mode::network` (standalone, exactly one occurrence) vs. join-with-`-r`
(repeatable, no limit). `--extern`/`--intern`/`--subnet <cidr>`/`--no-ipv6`/
`--subnet6 <cidr>`/`--no-veth` (`ParsedArgs::network_extern_flag`/`network_intern_flag`/
`network_subnet_flag`/`network_no_ipv6_flag`/`network_subnet6_flag`/
`network_no_veth_flag`) only
apply to the standalone (create) case and are rejected with a clear error if
given any other way (e.g. alongside `-r`) — `Mode::network` additionally
requires exactly one of `--extern`/`--intern`. `--no-veth` forces
`NetworkEntry::veth` (`config_file.h`) to `false` at creation time,
overriding the default `true` — see `network_bridge.h`'s
`probe_veth_support()`/`should_use_veth()` for what this controls: lets
the tap+relay fallback (the real target device's kernel lacks `CONFIG_VETH`
— see `docs/networking-design.md`'s addendum) be exercised on a
veth-capable machine like this dev box, without needing the actual
veth-less hardware. **`-n` used to belong to
`--no-nsenter`**: reassigned here since `--network` will be far more
heavily used; `--no-nsenter` moved to long-option-only (`options::no_nsenter`)
rather than hunting for a new letter, matching `--kill`'s own "rare/niche
flag, long-only is no real loss" precedent. `-p/--port-forward` (`'p'` was
free) is repeatable the same accumulate-now, resolve-after-the-loop way as
`-n` (`ParsedArgs::port_forward_specs`, raw `"[<network>:]<host-port>:
<container-port>"` strings — actual parsing happens later, in
`port_forward.h`, since resolving which network a spec refers to needs
runtime join state that doesn't exist yet at parse time), but has no
standalone use at all: rejected post-loop unless combined with `-r`.
- `commands.{h,cpp}` — every command's implementation, plus the dispatcher.
`dispatch_command(args, config_path, config)` (the only externally-linked
function; everything else in this file is `.cpp`-local) is a `switch
(args.mode)` with **one explicit `case` per `Mode` enumerator and no
`default:`**, so `-Wswitch` (this project builds at `warning_level=3`)
forces a compile warning/error if a future `Mode` value is ever added
without a matching dispatch case, instead of silently falling through to
the wrong command — confirmed by testing (temporarily adding an unhandled
enumerator triggered exactly the expected `-Wswitch` warning). `Mode::mount`
has its own explicit case (`mount_command()`) for the same reason: it used
to be handled only by *falling off the end* of a long `if`/`else` chain in
`main()` with no explicit check at all — the very kind of implicit,
easy-to-silently-break behavior this dispatcher redesign exists to close
off, especially with more `Mode` values (docker-compose support) expected
soon. `list_processes_command()` implements `--list-processes` (long-option
only): calls `list_sessions()` (`pid_file.{h,cpp}`, see below) and prints
one tab-aligned `pid`, `container name`, `running`/`exited` row per entry
(same two-column tab-alignment scheme as `list_images_command()`/
`list_volumes_command()`, extended to a third column), no header row, silent
success on an empty list. `clean_processes_command()` implements
`--clean-processes` (also long-option only): calls `clean_stale_sessions()`
(`pid_file.{h,cpp}`) and prints one `removed stale pid file for '<name>' (pid
<pid>)` line per file actually removed, then also calls
`clean_stale_port_forwards()` (`port_forward.h`, see below — commit 6 of
`docs/networking-design.md`'s sequence) and prints one `removed stale
port-forward rules for '<name>-<pid>'` line per record actually swept,
then `clean_stale_tap_relays()` (`network_tap_relay.h`, see below — the
direct tap+relay analog of the port-forward sweep) and prints one
`removed stale tap-relay processes for '<name>-<pid>'` line the same way
— nothing is printed for sessions still running, and an empty result
(nothing stale) is silent success, same convention as the rest of this
file's list/delete commands. `Mode::exec`'s
dispatch case is a one-line call to `exec_in_session(*args.exec_pid,
args.command)` (`exec_session.{h,cpp}`, see below) — the pid/command
parsing and validation now happens in `cli_args.cpp`'s `parse_args()`
instead (see above).
`inspect_image_command()` implements `-i/--inspect
<image.tar>`: prints every `OciImageConfig` field (user/group, exposed ports, env,
volumes, default command) without mounting or running the image — extend it
whenever `OciImageConfig` gains a new field (see `oci_image.{h,cpp}` below).
`run_container()` unconditionally calls `read_oci_image_config()` and reuses the
result for two independent defaults: the command to run (`Entrypoint ++ Cmd`) when
none is given on the command line, and, when `--user` wasn't given, the sandboxed
process's user/group (`config.User`, split into `OciImageConfig::user`/`group`) —
an explicit `--user`/`--group` on the command line always takes precedence.
`create_volume_command()` implements `-v/--volume <name> <directory>`;
`list_volumes_command()` implements `--list-volumes` (same tab-alignment scheme as
`list_images_command()`, reused as-is); `delete_volume_command()` implements both
`--delete-volume <name>` (config entry only) and `--delete-volume-full <name>`
(also `std::filesystem::remove_all()`s the host directory — errors out before
touching the config if that fails, warns instead of failing if the directory was
already gone) — see `config_file.{h,cpp}` below for what a "volume" means here (a
distinct concept from `OciImageConfig::volumes`). `dispatch_command()`'s
`Mode::volume`/`Mode::delete_volume`/`Mode::delete_volume_full` cases call these.
`create_network_command()`/`list_networks_command()`/`delete_network_command()`
are the direct network equivalents (`Mode::network`/`Mode::list_networks`/
`Mode::delete_network`) — see `docs/networking-design.md` for the full feature
design and `config_file.{h,cpp}` below for `NetworkEntry`. Joining a network
from `-r/--run` (repeatable `-n <name>`, `ParsedArgs::network_specs`, see
`cli_args.{h,cpp}` above) is handled by `run_container()`, further below,
via `network_join.{h,cpp}` (see below). `create_network_command()` rejects a
name containing `':'` first (`is_valid_network_name()`, `network_subnet.h`
— needed since `port_forward.h`'s `-p` syntax splits a spec on `':'`; a
network name containing one would make that parse ambiguous), then a
duplicate name, then
resolves `subnet`/`subnet6`: an explicit `--subnet`/`--subnet6` is validated
(`is_valid_ipv4_cidr()`/`is_valid_ipv6_cidr()`) and checked for overlap
against every existing network's subnet (`ipv4_cidrs_overlap()`/
`ipv6_cidrs_overlap()`, `network_subnet.{h,cpp}`, see below); otherwise the
next free block is auto-allocated (`allocate_ipv4_subnet()`/
`allocate_ipv6_subnet()`). Once a `subnet`/`subnet6` is resolved,
`create_network_command()` calls `ensure_network_provisioned()`
(`network_bridge.h`, see below) to actually stand up the network's
host-side state (bridge, sysctls, iptables rules for `extern`; a dedicated
persistent namespace + bridge for `intern`) — only once that succeeds is
the entry appended to `config.networks` and persisted; a network that
fails to provision isn't saved. `list_networks_command()` reuses the same
independently-per-column tab-alignment scheme as `list_processes_command()`
(name/kind/subnet/bridge each aligned — the bridge name recomputed via
`bridge_name(network.name)``network_bridge.h` — rather than stored, since
it's already a pure deterministic function of the name — then the IPv6
subnet — or `"(no ipv6)"` — appended unaligned as the trailing column,
nothing follows it).
`delete_network_command()` takes a `delete_full` bool, same shape as
`delete_volume_command()`'s own: `--delete-network` (`false`) only removes
the config entry, leaving the network's live host-side state untouched;
`--delete-network-full` (`true`) additionally calls
`teardown_network_state()` (`network_bridge.h`, see below) first. Unlike
`delete_volume_command()`'s `-full` variant (which bails out *before*
touching the config if its single `remove_all()` call fails),
`delete_network_command()` doesn't gate the config removal on
`teardown_network_state()`'s success at all — that function is
deliberately best-effort/non-fatal per-step (see its own doc comment), so
a step "failing" because that piece was already gone by hand (exactly
`ensure_network_provisioned()`'s own existence-check caveat, above) is
expected, not a reason to leave a network the user explicitly asked to
delete sitting in the config.
`write_config_command()` implements `-w/--write-config`: unlike
`create_volume_command()`/`delete_volume_command()`'s use of
`write_config_file()` (which only ever persists `AppConfig` fields that are
already set), this fills in *every* field before writing — the six
`unshare-*` bools via `.value_or(true)`, and `log_level` from the actually
active `spdlog::get_level()` (not merely a default for when unset — this
also captures an explicit `--log-level` passed alongside `-w` on the same
command line, overriding whatever an existing config file's own
`log-level` already was, since `main()`/`parse_args()` already applied it
in that precedence order by the time this runs) — so a bare `-w` bootstraps
a complete, fully-populated config file for hand-editing, and `-w` combined
with other flags captures their effective values into it. `volumes` is left
exactly as loaded — an open-ended list with no "default" entry to
materialize. Prints the config file's full path (`write_config_file()`
already creates the parent directory and the file itself if missing, so no
separate existence check is needed here).
`run_container()` (the `Mode::run` dispatch case) resolves each `-v` spec
(erroring out, `ok = false`, same as a failed `--user` resolution — `bwrap`
is skipped but unmount/cleanup still runs) into a `ResolvedVolumeMount`,
rejecting a duplicate or non-absolute container path first, and passes the
resolved list to `run_bwrap()`. `--hostname <name>` is likewise threaded
straight through `run_container()` into `run_bwrap()`/`build_bwrap_args()`
(`bwrap.{h,cpp}`) — see there for how/when it actually takes effect.
`run_container()` also derives a `container_name` for the session-tracking
pid file (see `pid_file.{h,cpp}` below): `read_image_ref()`
(`oci_image.{h,cpp}`) applied to the single image tar being run, formatted
as `name:tag`, falling back to the tar's own filename stem if
`read_image_ref()` can't determine one — passed through to `run_bwrap()`
alongside everything else. `--env`/`--env-file`'s already-ordered
`env_specs` (see `cli_args.{h,cpp}` above) are resolved here, once, via
`resolve_env_specs()` (`env_spec.{h,cpp}`, see below) — same `ok =
false`-on-failure pattern as volume/user resolution — and the resolved list
is passed to `run_bwrap()` as `extra_env`. `-D/--daemonize`'s
`daemonize_flag` is also consumed here: `run_container()` computes
`container_name` *before* `mount_image()` (needed so `daemonize()` below can
use the real container name for the log file from its very first line, not
just after a later rename) and, if daemonizing, calls
`daemonize(container_name)` (`daemonize.{h,cpp}`, see below) immediately
after: a returned value means this is the original (parent) process (or a
hard daemonize failure) — print it and `return` right away; `nullopt` means
this is the now-detached child, which falls through into the rest of
`run_container()`'s existing body completely unchanged, including the
unmount/cleanup that already runs after `run_bwrap()` returns (no separate
watcher/reaper — the daemonized child *is* what runs the whole session,
start to finish). `network_specs` (repeatable `-n`, `cli_args.h`) is
validated up front, before `run_bwrap()` is ever called: joining a network
needs a real, isolated network namespace to attach a veth into (unlike
`--hostname`, this can't just be skipped/degraded when unavailable), so if
any networks were requested, `run_container()` checks both that
`namespace_config.net` is actually enabled (`global.unshare-net`,
`config_file.h`) and that `detect_bwrap_unshare_args()` (`bwrap.h`) reports
the running kernel actually supports `--unshare-net` — either failing sets
`ok = false` with a clear error, the same pattern as a failed `--user`
resolution. If validation passed, `on_bwrap_pid_known` (already used for
`-D/--daemonize`'s `report_daemon_started()`) additionally calls
`join_networks(pid, network_specs, app_config)` (`network_join.h`, see
below) — networks are joined *before* the daemonize report is sent, so a
`-D`-daemonized caller doesn't get control back until network setup has
already had its chance to run. `join_networks()` itself early-returns (no
namespace wait at all) when `network_specs` is empty, so calling it
unconditionally whenever `on_bwrap_pid_known` fires for *any* reason (e.g.
`-D/--daemonize` alone, no `-n`) doesn't cost anything. `-p`'s
`port_forward_specs` (`cli_args.h`) are syntax/range-parsed
(`parse_port_forward_spec()`, `port_forward.h`) up front too — `ok = false`
on a bad spec, same as other validation failures — but *resolving* which
network each targets can only happen after `join_networks()` returns (it
needs to know which networks actually joined, and their assigned IPs), so
that happens in the same `on_bwrap_pid_known` callback, right after the
`join_networks()` call: `add_port_forward()` per spec, collecting the
ones that actually landed into a `std::vector<ActivePortForward>` declared
in `run_container()`'s own scope (captured by reference) — read again
*after* `run_bwrap()` returns to `remove_port_forward()` each one. This
two-places split (add during the callback, remove after `run_bwrap()`
returns) mirrors how `join_networks()`'s own veths don't need an explicit
removal step (the kernel tears them down once the session's namespace
goes away) while port-forward rules — host-global, named, persistent
iptables state — very much do. `on_bwrap_pid_known` itself is only set at
all when `daemonize_flag || !network_specs.empty() ||
!parsed_port_forwards.empty()` — a real gap caught while wiring this up: an
earlier version only checked the first two, so `-p` given *without* `-n`
or `-D` would silently never even attempt to run (no error, nothing
logged) since the callback that resolves/applies it would never fire at
all. At the end of the callback, `record_port_forwards(container_name,
pid, active_port_forwards)` (`port_forward.h`, see below) persists
whatever actually landed to a small state file — so that a later
`--clean-processes` run can find and remove these rules even if *this*
process crashes before ever reaching its own `remove_port_forward()` calls
after `run_bwrap()` returns; those calls are paired with a
`remove_port_forward_record(container_name, bwrap_pid)` (`bwrap_pid`
captured from the same callback, in a variable declared in
`run_container()`'s own scope) so a cleanly-exiting session's own record
doesn't linger for `--clean-processes` to find later. The same callback
also collects each `JoinedNetwork::relay` (`network_join.h`) returned by
`join_networks()` into a `std::vector<TapRelayHandle> active_relays`
declared in `run_container()`'s own scope — the direct tap+relay
(`network_tap_relay.h`) analog of `active_port_forwards` above, since a
relay process is likewise independent host-global state (unlike a veth
pair) that needs an explicit `stop_tap_relay()` call for each, made right
alongside the `remove_port_forward()` loop after `run_bwrap()` returns —
paired the same two-places way with `record_tap_relays(container_name,
pid, active_relays)` (called right after `record_port_forwards()`, same
callback) and `remove_tap_relay_record(container_name, bwrap_pid)` (called
right after the `stop_tap_relay()` loop) for `--clean-processes`'s own
crash-orphan sweep (`network_tap_relay.h`'s `clean_stale_tap_relays()`,
see below).
- `self_test.{h,cpp}``run_self_tests()` implements `-t/--test`, this
project's own built-in self-test mode (distinct from the Meson-driven
fixture smoke test under `tests/`, described in "Build & test commands"
below; previously reported `detect_bwrap_unshare_args()`'s output —
`bwrap.{h,cpp}` — unplugged since that's kernel-capability diagnostics, not
a test). Currently exercises `persistent_netns.{h,cpp}`'s (see below)
create/verify/remove cycle: skipped with a message (not a failure) when not
root, since `create_persistent_netns()` requires it for the bind mount.
Confirms the namespace is missing before creation, exists right after
(checked from this process, *after* the forked child that actually did the
`unshare()`/bind-mount has already exited — the actual claim being tested:
the namespace outlives its creating process), then gone again after
removal. Also exercises `network_tap_relay.{h,cpp}`'s (see below)
create/attach/teardown cycle, same root-only skip: a throwaway bridge
(`ip link add ... type bridge` inside a throwaway persistent namespace
of its own — `create_tap_relay()` now always enters a network's
persistent namespace first, both kinds, see `network_bridge.{h,cpp}`'s
own "Resolved: `extern` had no connectivity" entry, so this test needs
one too even though `network_bridge.h`'s own bridge *provisioning* logic
still isn't otherwise exercised here — a tap device only cares that
*some* bridge exists to attach to) stands in for a real network's bridge,
and a throwaway network namespace (a forked
child that `unshare(CLONE_NEWNET)`s then blocks in `pause()` until
signaled) stands in for a real `-r/--run` session's isolated one. **Real
race caught by testing, not assumed**: `fork()` returning to the parent
doesn't mean the child has actually reached its own `unshare(CLONE_NEWNET)`
call yet — the same race `network_join.cpp`'s own
`wait_for_isolated_net_namespace()` already guards against for a real
session; an earlier version of this test used the child's pid immediately
and the container-side tap device silently ended up created in the *host's*
namespace instead (confirmed: it wasn't visible via `nsenter` into the
child's namespace at all). Fixed the same way — polling (bounded, 1s,
20ms interval) `namespace_isolated()` (`sandbox_process.h`) until the
child's own net namespace actually differs from this process's before
trusting its pid. Confirms the host-side tap gets created and attached to
the bridge (`ip link show` output contains `master <bridge>`), the
container-side tap gets created with the requested name inside the target
namespace (checked via a plain, single-shot
`nsenter --net=/proc/<pid>/ns/net -- ip link show`; an earlier version
wrapped this in a bounded retry, added when tap devices were still created
via bare `ioctl(TUNSETIFF)` and intermittently weren't immediately
visible — see `network_tap_relay.{h,cpp}`'s own entry below for why that
retry, and the whole class of symptom it was compensating for, is gone
now that devices are created persistently instead), and — **updated once
the original "both devices disappear on their own once `stop_tap_relay()`
stops the relay, no explicit `ip link del` needed" assumption turned out
to be wrong on the real target device** (`network_tap_relay.{h,cpp}`'s own
entry below has the full story) — now confirms the *host*-side device is
explicitly gone after `stop_tap_relay()` (the new `ip link del` step)
while the *container*-side device deliberately still exists (correctly
persistent — only the relay stopped, not the container's own network
namespace); the container-side device's actual disappearance, once that
namespace itself is torn down, isn't separately re-checked (`nsenter` has
nothing left to target once the namespace's only holding process has
already exited) — it's destroyed moments later anyway, at the very end of
this test, when its throwaway namespace-holder process is killed.
Deliberately its own small file since more real tests are expected here as
more of the networking feature lands.
- `env_spec.{h,cpp}``resolve_env_specs()` turns an ordered list of
`EnvSpec {is_file, value}` (see `cli_args.{h,cpp}` above) into a flat, ordered list of
`(key, value)` pairs. A literal (`--env`) is split at its *first* `=` (the
value may itself contain `=`; the key must be non-empty). A file (`--env-file`)
is read line by line: blank/whitespace-only lines and lines whose first
non-whitespace character is `#` are skipped (comments), with a trailing `\r`
stripped first for CRLF files; every other line is parsed the same way as a
literal. Logs a specific error and returns `nullopt` on the first hard failure
(malformed line, empty key, or an unreadable file) — deliberately stops at the
first line, not "skip and warn", since an env file with a typo should fail
loudly rather than silently omit a variable a container might depend on.
`build_sandbox_env()` (`bwrap.cpp`, see below) appends the resolved list after
its own built-in `PATH`/`HOME`/`PWD`/`TERM` — no deduplication needed there,
since `run_process_foreground()`'s own `setenv(..., 1)` loop already lets the
later occurrence in iteration order win for a repeated key, so an explicit
`--env PATH=...` still overrides the default.
- `oci_image.{h,cpp}` — validates/parses the OCI Image Layout tar (libarchive +
nlohmann_json) and extracts layer blobs. `list_oci_images()` scans a directory
(non-recursively) for `*.tar`/`*.tar.*` files and, for each valid OCI archive,
derives an image name/tag via `read_image_ref()` from its `index.json` manifest
annotations (`io.containerd.image.name` preferred, else
`org.opencontainers.image.ref.name`), falling back to the archive's filename and
`"latest"` respectively. `read_image_ref()` is public (not just an internal
helper of `list_oci_images()`) precisely so `run_container()` (`commands.cpp`) can
reuse the exact same logic to name a *single* image tar's session pid file (see
`pid_file.{h,cpp}` below) instead of duplicating it.
`read_oci_image_config()` reads the image config blob referenced by the manifest
and extracts `User` (split on `:` into `OciImageConfig::user`/`group`),
`ExposedPorts`, `Env`, `Volumes`, and the effective default command
(`Entrypoint ++ Cmd`). `user`/`group` and the default command are consumed by
`-r/--run`, and every field is displayed by `-i/--inspect` (see `commands.{h,cpp}` above)
`ExposedPorts`/`Env`/`Volumes` are otherwise still just captured for when
networking/volumes are implemented.
- `containers_storage.{h,cpp}` — wraps the `containers-storage` CLI (`import-layer`,
`mount`, `unmount`, `layer --json`, `delete-layer`), forcing `fuse-overlayfs` as the
overlay `mount_program`. `cleanup_layer_chain()` walks a layer's parent chain
(children before parents) deleting each one.
- `bwrap.{h,cpp}``detect_bwrap_unshare_args()` probes the kernel (via a forked
`unshare(2)` per namespace type) for which `--unshare-xxx` flags `bwrap` can actually
use; `build_bwrap_args()`/`run_bwrap()` assemble and run the sandboxed command.
`build_bwrap_args()` also takes a `std::vector<ResolvedVolumeMount>` (see
`volume_mount.h` below) and appends one writable `--bind <host_directory>
<container_path>` per entry. `wrap_for_root_namespace()` is `run_bwrap()`'s own
nsenter-wrapping logic pulled out into a reusable, exported function — it's also
what `volume_mount.cpp`'s copy-into-an-empty-volume step uses to reach the
image's content when running rootless (see below); `run_bwrap()` itself now just
calls it once on the assembled `bwrap` argv.
`build_bwrap_args()`/`run_bwrap()` also take a `NamespaceConfig` (`bwrap.h`) — one
plain `bool` field per `namespace_probes` entry (`user`/`ipc`/`pid`/`net`/`uts`/
`cgroup`, default `true`), resolved by `run_container()` (`commands.cpp`) from
`AppConfig`'s six `global.unshare-*` keys (`config_file.h`, see above) once, up
front — `bwrap.{h,cpp}` itself never touches `AppConfig`/YAML, only this
already-resolved struct. For each flag `detect_bwrap_unshare_args()` finds the
kernel supports, `build_bwrap_args()` additionally requires the matching
`NamespaceConfig` field to be `true` (looked up via a `.cpp`-local
`namespace_policy_enabled()` if-chain over `namespace_probes`' `name`s) before
actually passing it to `bwrap` — kernel support and policy are separate gates,
both must allow a type. This replaced an earlier hardcoded special case that
always dropped `--unshare-net` regardless of policy or kernel support (without
any network setup, e.g. `slirp4netns`, unsharing it just left the sandbox with no
network at all) — `net` now goes through the exact same policy path as every
other type, defaulting to enabled like the rest. **This is a deliberate,
user-acknowledged transitional behavior change**: as of this, a plain `-r/--run`
with no config file override gets a real network namespace and thus no network
access at all, until `slirp4netns` integration (the next task on this same
branch) actually sets one up; `global.unshare-net: off` restores the prior
no-isolation behavior in the meantime. `detect_bwrap_unshare_args()` itself is
untouched by any of this — still an unfiltered kernel-capability probe, unrelated
to policy (no longer surfaced via `-t/--test`, see `self_test.{h,cpp}` below).
`build_bwrap_args()`/`run_bwrap()` also take an optional `hostname` (from
`--hostname`, long-option only): passed through as bwrap's own `--hostname` only
when `--unshare-uts` is actually among the flags `bwrap` is being given (bwrap
itself refuses `--hostname` without it) — otherwise logs a warning and leaves the
sandbox's hostname alone, since a stock Android kernel in degraded mode may not
support a UTS namespace at all.
Never requests `--unshare-user` when running as root: root doesn't need a fresh
user namespace for privilege, and bwrap's own single-mapping uid/gid setup for one
triggers the kernel's unprivileged-userns setgroups() restriction, which showed up
as every other supplementary group collapsing to the overflow gid ("nobody") in
`id`, and `su` inside the sandbox failing with "can't set groups: Operation not
permitted". Because of that, bwrap's own `--uid`/`--gid` (which require
`--unshare-user`) aren't usable when running as root either — `--user`/`--group`
work around this: when set, `build_bwrap_args()`/`run_bwrap()` bind-mount the
separate `slocker-lite-priv-drop` helper (see below) into the sandbox at a fixed
hidden path and route the real command through it as
`<uid>:<gid> -- <command...>`. This only actually works without a user namespace
(i.e. running as root) — under `--unshare-user`, the sandbox's uid map has only
one valid entry, so the helper's own `setuid()` fails cleanly there instead of
silently doing nothing. `run_bwrap()` fails fast (returns -1) if the helper can't
be found next to this binary when `--user` was requested, rather than silently
running the command as root. `run_bwrap()` also takes a `container_name` and
tracks the running session with it: it passes a lambda as
`run_process_foreground()`'s new `on_start` callback (see `process.{h,cpp}`
below) that calls `create_session_lock(container_name, pid)` (`pid_file.{h,cpp}`,
see below) the instant the real `bwrap` pid is known, then calls
`release_session_lock()` once `run_process_foreground()` returns (covering
every exit path — normal, nonzero, or a forwarded-signal exit — since that call
always blocks until the child has actually exited).
`build_bwrap_args()` no longer passes `--clearenv`/`--setenv` to `bwrap` itself;
instead, `build_sandbox_env()` builds the sandboxed command's exact environment
(`PATH`, `HOME`, `PWD` — hardcoded to `"/"`, matching `--chdir`'s own value; note
per bwrap's own man page `--clearenv` never actually unset `PWD` in the first
place, so this isn't a straight port of a prior `--setenv` — and `TERM`, only
if the host process has one, followed by `extra_env` — the resolved
`--env`/`--env-file` list from `resolve_env_specs()` (`env_spec.h`), appended
last so it can override the built-in defaults for the same key) and
`run_bwrap()` passes it straight to `run_process_foreground()`'s own `env`
override (see `process.{h,cpp}` below). This works because `bwrap` (and
`nsenter`, when interposed via `wrap_for_root_namespace()`) doesn't alter its
own inherited environment unless told to, and neither does
`slocker-lite-priv-drop` (just `setgroups()`/`setgid()`/`setuid()`/`execvp()`,
no env manipulation) — so controlling it once, at the outermost exec, is
sufficient for it to reach the final sandboxed command unchanged. **Because
that outermost exec now uses this same explicitly-built environment**,
`build_bwrap_args()`/`wrap_for_root_namespace()` resolve `bwrap`'s and
`nsenter`'s own argv[0] to an absolute path via `find_in_path()` (called from
this process's own, unmodified environment, before `fork()`) instead of
leaving them as bare names — confirmed by direct testing: `--env PATH=...`
used to break `execvp()`'s ability to even *locate* `bwrap`/`nsenter`
(bare-name lookup happens in the child, using the already-overridden PATH),
not just what the sandboxed command itself sees. With the fix, only the
*sandboxed command's own* lookup is affected by a `--env PATH=...` override
(as expected — same as overriding `PATH` in any real shell before running a
bare command name), and `bwrap`/`nsenter` are always found regardless.
`run_bwrap()` also takes an optional `on_bwrap_pid_known` callback, invoked
alongside (not instead of) the session-lock-creation lambda, at the exact
same `on_start` timing — `-D/--daemonize` (`daemonize.{h,cpp}`, see below)
hooks in here via `report_daemon_started()` to learn the real pid at the
same instant everything else that needs it does, rather than needing its own
separate pid-discovery mechanism. That same `on_start` lambda also calls
`create_session_cgroup()` (`session_cgroup.h`, see below), right alongside
`create_session_lock()`, so `--kill` can later find every process the
session ever starts via its dedicated cgroup; `remove_session_cgroup()` is
called from the same post-`run_process_foreground()` spot
`release_session_lock()` already is.
- `priv_drop_helper.cpp` → the separate `slocker-lite-priv-drop` binary (its own
`executable()` target in `meson.build`, **built with `-static`**). Deliberately
has zero dependencies on the rest of this project (no fmt/spdlog/etc.) and is
fully statically linked: it gets bind-mounted *into the container image's own
filesystem*, which won't have `slocker-lite`'s own shared library dependencies —
a dynamically linked binary bind-mounted that way fails outright ("error while
loading shared libraries"), which is exactly what happened before this was split
out (the original approach bind-mounted `slocker-lite`'s own — dynamically linked
— binary via `/proc/self/exe` and reexeced it; kept only as a lesson, not as
working code). Usage: `slocker-lite-priv-drop <uid>:<gid> -- <command> [args...]`;
does `setgroups(0,…)``setgid()``setuid()``execvp()`, in that order
(dropping the group needs `CAP_SETGID`, which is lost once `setuid()` drops root).
`find_priv_drop_helper()` and the `priv_drop::path`/`priv_drop::helper_name`
constants live in `bwrap.h` (not just internal to `bwrap.cpp`) specifically so
`exec_in_session()` (`exec_session.cpp`, see below) can reuse the exact same
already-bind-mounted helper for `-x/--exec`'s own `--user`/`--group` support,
instead of a second copy needing to be bind-mounted for it (which wouldn't even
be possible — `-x/--exec` joins an *already-running* session's mount namespace,
it doesn't get to add bind mounts to it). `find_priv_drop_helper()` itself only
checks this binary's own host-side existence; it says nothing about whether a
given session actually has it bind-mounted (only true when that session's
`-r/--run` resolved a user in the first place).
- `user_spec.{h,cpp}``resolve_user_and_group()` resolves a user/group spec (each
a name or numeric id) against the *container's own* `/etc/passwd`/`/etc/group`
**content** (not the host's, and not a path — callers own reading it, since the
two current callers get that content two different ways: `run_container()`
reads it directly off the merged mount path, while `exec_in_session()` fetches
it over `nsenter`, since a running session's mount namespace isn't otherwise
reachable from this process — see below). `nullopt` content for either file
means "unreadable/absent"; a numeric user with no group still resolves fine
without it (defaults gid to the same numeric value as the uid) but a named one
doesn't. `ResolvedUser` (`bwrap.h`) also carries `home`, looked up by the final
resolved uid's `/etc/passwd` entry (field 5) regardless of whether `user` was
given as a name or a number; falls back to `"/root"` for uid 0 or `"/"`
otherwise when there's no matching row. `build_sandbox_env()` (`bwrap.cpp`)
sets the sandboxed process's `HOME` from this — `"/root"` only when no user
override applies at all (no `--user`, no image-declared `config.User`).
`run_container()` (`commands.cpp`) calls `resolve_user_and_group()` with either
the explicit `--user`/`--group` flags, or, when `--user` wasn't given, the
image's own declared `config.User` (`OciImageConfig::user`/`group`) — so a
container defaults to running as whatever user the image itself declares, not
root, unless the image declares none.
- `process.{h,cpp}` — argv-based subprocess helpers (fork/execvp, no shell):
`run_process()` captures stdout (used for `containers-storage` calls),
`run_process_foreground()` inherits all of stdio (used for the interactive `bwrap`
run). Also `find_in_path()`, a shared `$PATH` lookup. `run_process_foreground()`
installs a SIGINT/SIGTERM handler around its `waitpid()` that forwards the signal
to the running child and keeps waiting instead of letting the default disposition
kill `slocker-lite` itself — without this, Ctrl-C (or `kill`) during `-r`'s `bwrap`
run would skip `run_container()`'s unmount/cleanup entirely, leaving the layer
imported and/or mounted. `run_process_foreground()` also takes an optional
`on_start` callback, invoked with the child's real pid right after `fork()`
succeeds (before the signal handlers go up and it blocks in `waitpid()`) — the
only point where that pid is knowable, and still accurate even when `argv`
itself execs into something else first (e.g. `nsenter` handing off to the final
command via its own in-place `execvp()` — a pid never changes across `exec()`).
`run_bwrap()` (`bwrap.cpp`) is the one caller that uses it, for session pid-file
tracking (see `pid_file.{h,cpp}` below). `run_process_foreground()` also takes
an optional `env` (list of key/value pairs): when set, the forked child
replaces its entire environment via `clearenv()`/`setenv()` (plain POSIX, not
the GNU-only `execvpe()` — the target platform includes musl) before `execvp()`,
instead of inheriting this process's own. `nullopt` (the default) leaves the
child's environment untouched. `run_bwrap()` is again the one caller that uses
this, via `build_sandbox_env()` (`bwrap.cpp`) — see there.
- `pid_file.{h,cpp}` — tracks one running `-r/--run` session (a live `bwrap`
process) as a locked pid file, so an outside process (or a later
`slocker-lite` invocation) can tell whether it's still running.
`sanitize_for_filename()` (anything outside `[A-Za-z0-9._-]``_`, falling
back to `"container"` if that leaves nothing) is exported here (not just
`.cpp`-local) specifically so `session_cgroup.{h,cpp}` (see below) can reuse
the exact same `<name>-<pid>` naming rule for its own per-session cgroup
directory without drifting from this file's own. `xdg_state_dir()`
(`$XDG_STATE_HOME/slocker-lite`, or the `$HOME/.local/state/...` fallback)
is likewise exported (moved out of this file's own anonymous namespace) so
`persistent_netns.{h,cpp}` (see below) and `network_join.cpp`'s own
per-address lease files (`xdg_state_dir() / "net-leases"`) can resolve
their own subdirectories under the same state root without a second,
drifting copy of this resolution logic. `session_pid_file_path()` resolves
`$XDG_STATE_HOME/slocker-lite/run/<container_name>-<pid>` (falling back to
`$HOME/.local/state/...` when `XDG_STATE_HOME` is unset/empty — same
resolution pattern as `config_file_path()` below, for state instead of
config), sanitizing `container_name` first (anything outside `[A-Za-z0-9._-]`
`_`, since an image name/tag can contain `/` or `:`). `session_log_file_path()`
is a sibling resolving to `$XDG_STATE_HOME/slocker-lite/logs/<container_name>-<pid>.log`
instead — same sanitization, same `$XDG_STATE_HOME`/`$HOME` fallback, just a
different subdirectory and a `.log` extension — used by `daemonize.{h,cpp}`
(see below) for `-D/--daemonize`'s log file. `create_session_lock()`
creates the file (`O_CREAT|O_WRONLY|O_TRUNC|O_CLOEXEC`, mode 0644 — `O_CLOEXEC`
matters: this fd must never leak into the sandboxed command's own fd table),
writes the pid as text, and takes an exclusive, non-blocking `flock()` on it —
held only by that fd, so its lifetime tracks `slocker-lite`'s own process
lifetime (released automatically on any exit, including a crash), which lines
up with `bwrap` itself being invoked with `--die-with-parent`. Any external
tool can check liveness the same way: attempt the same exclusive non-blocking
`flock()` on the file — success means nothing holds it anymore (stale, safe to
remove), `EWOULDBLOCK` means a live process still does. `release_session_lock()`
closes the fd (releasing the flock immediately) and removes the file. Every
failure path here (can't create the directory/file, can't lock, can't remove)
is a `spdlog::warn`, never fatal — session tracking is best-effort and must
never block or fail `-r/--run` itself. `list_sessions()` implements
`--list-processes` (`commands.cpp`'s `list_processes_command()`): scans the same
`run/` directory and reports one `SessionInfo {pid, container_name, running}`
per readable pid file. `pid` is read from the file's own contents, not parsed
from the filename (ambiguous for names that themselves contain `-`);
`container_name` is then recovered by stripping that exact `-<pid>` suffix
back off the filename. `running` reuses the same liveness check any external
tool would do — a non-blocking exclusive `flock()` that succeeds means the
file is actually stale, so `running` is false in that case; the lock is always
released again immediately either way, never left held by the check itself. A
file that can't be opened or doesn't parse as a pid (e.g. removed mid-scan) is
silently skipped, not reported as an error — scanning a live directory is
inherently racy. Both `list_sessions()` and `clean_stale_sessions()` (the
latter implements `--clean-processes`) share a private `open_session_file()`
helper for the open/read-pid/recover-name step. `clean_stale_sessions()`
doesn't just remove whatever a separate `list_sessions()` call reported as not
running — it re-takes the same non-blocking `flock()` used to test liveness
and holds it across the `remove()` call itself, per file, so the stale check
and the removal stay atomic against a new session starting in the gap between
a check and a later removal. Only files it actually removes are reported back
(as `SessionInfo`s with `running=false`); still-locked (running) files are
left untouched and not reported.
- `persistent_netns.{h,cpp}` — generic, narrow infrastructure for keeping a
network namespace alive with no process in it, the way `ip netns add` does;
no `intern`/`extern` policy or bridge logic here (that's `network_bridge.{h,cpp}`,
see below, which is what actually calls `create_persistent_netns()` for an
`intern` network). `persistent_netns_path()`
resolves `xdg_state_dir() / "netns" / sanitize_for_filename(name)`
(`pid_file.h`, see above). `persistent_netns_exists()` checks whether that
path is actually a live bind-mounted namespace, not just a stale/never-
mounted file: `stat()`s the path and its parent directory and compares
`st_dev` — a genuine bind mount always has a different device number than
its parent, the same "is this a mountpoint" technique used elsewhere. Never
needs root itself (just `stat()`). `create_persistent_netns()` forks a
child (never touches the caller's own network namespace — `unshare(2)`
affects only the calling process) that `unshare(CLONE_NEWNET)`s its own
fresh namespace, bind-mounts its `/proc/self/ns/net` onto the target path,
then exits immediately — the bind mount itself is what keeps the namespace
alive from then on, independent of the now-exited child, exactly `ip netns
add`'s own technique. Requires `CAP_SYS_ADMIN` (root) for the bind mount,
matching this feature's current root-only scope (see
`docs/networking-design.md`) — best-effort like this project's other
host-state primitives (session locks, cgroups): logs and returns `false` on
any failure (already exists, fork/unshare/mount failure) rather than
throwing. `remove_persistent_netns()` unmounts then removes the file.
- `network_bridge.{h,cpp}` — stands up (or confirms already-standing) a
network's actual host-side state: `ensure_network_provisioned()` is
idempotent by design (checks `ip link show <bridge>` first and does nothing
further if it's already there) — this is deliberately also the reboot-
reconciliation mechanism, not a separate code path: nothing about a
network's live state (bridge, veths, iptables rules; the persistent
namespace itself for `intern`) survives a reboot except its `config.yaml`
entry, so calling this again after one just recreates whatever's missing.
`bridge_name()` derives a stable interface name from the network's own name
via a hand-rolled 32-bit FNV-1a (`"slk" + 8 hex chars`, 11 characters,
comfortably under Linux's `IFNAMSIZ - 1` = 15-character limit regardless of
how long the network name is) — deliberately not `std::hash<std::string>`,
whose exact value is implementation-defined and not guaranteed stable
across a rebuild with a different standard library, which would silently
orphan an already-provisioned bridge a rebuilt binary
can no longer find by the name it now computes. Every command, both kinds,
is wrapped through `nsenter --net=<persistent path>` (`wrap_for_network()`)
into the network's own dedicated namespace (`persistent_netns.h`, created
here first if it doesn't exist yet) — the same pattern
`wrap_for_root_namespace()` (`bwrap.cpp`) already uses for the rootless
`containers-storage` mount's namespace, just targeting a persistent
bind-mounted path instead of a live pid's `/proc` entry. **`extern` used
to run every command directly, unwrapped, since its bridge used to live in
the host's own root namespace** — see this entry's own "Resolved: `extern`
had no connectivity at all on the real device" paragraph further below for
why that changed: confirmed by direct on-device testing to be the actual
cause of a real, reproducible bug, not just an implementation choice.
`bridge_name()`/`wrap_for_network()` are both exported (not just
this file's own internal helpers) specifically so `network_join.{h,cpp}`
(see below) can attach a container's veth to the exact same bridge, in the
exact same place, this file provisioned it in. `provision_bridge()`
(`.cpp`-local): creates the bridge, assigns it
the gateway address from `network_subnet.h`'s `ipv4_gateway_address()`
(and `ipv6_gateway_address()` if `network.ipv6`), brings it up, then —
`extern` only — `sysctl -w net.ipv4.ip_forward=1` and one `iptables`
`POSTROUTING`/`MASQUERADE` rule for the subnet (`! -o <bridge>`, the same
`docker0` shape, so bridge-local inter-container traffic isn't
unnecessarily NAT'd; both idempotent global host sysctls, not per-bridge,
so no separate "already enabled" tracking is needed), plus, if `ipv6`, the
IPv6 forwarding sysctl — but **deliberately no `ip6tables` MASQUERADE
rule**: the `fd00::/8` ULA addresses `network_subnet.h` allocates are
non-globally-routable by design (RFC 4193, the IPv6 equivalent of RFC1918
private space), so NAT66 for them isn't correct IPv6 practice to begin
with (also confirmed not universally supported on the real target device —
its `ip6tables` build lacks a `MASQUERADE` target entirely). `extern`'s
IPv6 side is thus same-bridge reachability only, exactly what `intern`'s
IPv6 already was — and confirmed on the real target device, not just
theorized, that this is the only option regardless: neither `ip6tables`
nor `nftables` can even create an IPv6 NAT table on that kernel at all
("Not supported"). `check_network_dependencies()` gates on the tool set
each kind actually needs (`ip`/`nsenter` always, both now reaching their
bridge through a private namespace; `iptables`/`sysctl` additionally for
`extern` — no `ip6tables`, even when `ipv6`, since none is ever called) —
same shape/spirit as `commands.cpp`'s own `check_required_dependencies()`,
kept separate since the tool set here depends on the network's own
kind/`ipv6` setting.
`create_network_command()` (`commands.cpp`, see above) calls
`ensure_network_provisioned()` after resolving subnets but *before*
persisting the config entry — a network that fails to provision isn't
saved, so a later join doesn't find a config entry for something that
doesn't actually exist on the host. **Historical verification note**: an
earlier verification pass (before the "Resolved" paragraph below) confirmed
a real `extern` network's bridge/gateway/`ip_forward`/`MASQUERADE` rules
all coming up correctly *directly in the host's root namespace* — that
description is now superseded; see below for the current, corrected
architecture and why it changed. `ensure_network_provisioned()` no longer
short-circuits on `bridge_exists()` alone either: the uplink step below
(`extern` only) now runs on every call, even when the bridge already
existed, via its own separate idempotency check.
`teardown_network_state()` is the reverse: `--delete-network-full`
(`commands.cpp`'s `delete_network_command()`) calls it to tear down
exactly what `ensure_network_provisioned()` stood up. For `extern`: first
tears down the uplink (`teardown_uplink_state()`, see below), then removes
the `iptables` MASQUERADE rule for the container subnet (no `ip6tables`
counterpart, since none is ever added — see `ensure_network_provisioned()`'s
own comment above) — via a `.cpp`-local `teardown_step()`, the same shape
as `run_admin_command()` but logging a *warning*, not an error, on failure,
since a step failing because that piece was already gone by hand is the
expected, common case this exists to handle (exactly
`ensure_network_provisioned()`'s own existence-check caveat above — this is
precisely how a manually-removed MASQUERADE rule can go from "won't come
back on recreate" to "cleanly torn down and recreated" once
`--delete-network-full` exists at all). **Both kinds** then remove the
whole persistent namespace (`persistent_netns.h`) in one step, which
destroys everything left inside it — the bridge, and for `extern` the
uplink's own private-namespace-side tap device, both included — with no
separate `ip link del` needed for those. Deliberately never touches the
IPv4/IPv6 forwarding sysctls `provision_bridge()` enables for `extern`
those are global host state shared across every `extern` network, not
per-network, so turning them off here could break others still relying on
them. **Verified end-to-end on this dev machine (root, via the scoped
`doas` rule)**: an `extern` network's bridge and MASQUERADE rule were
both confirmed gone after `--delete-network-full` (`ip link show`
reporting "Device does not exist"), and recreating a network with the
*same name* afterward correctly went through `provision_bridge()` again
from scratch (confirmed via the debug log) instead of short-circuiting
on a stale `bridge_exists()` check — fixing exactly the gap a user
reported (a manually-removed MASQUERADE rule never came back on
`--delete-network` + recreate, since the old bridge was silently still
there); an `intern` network's persistent namespace was similarly
confirmed fully removed and recreatable without conflict.
**Resolved: `extern` had no connectivity at all on the real device.**
Full incident writeup in `docs/networking-design.md`'s own section of the
same name — this entry just covers the resulting code. Root cause: the
bridge lived directly in the host's own root namespace, and (almost
certainly) Android's own `netd`-managed iptables/routing policy applies
only there, never to a genuinely isolated namespace — exactly why `intern`
was unaffected the whole time. `wrap_for_network()` no longer special-cases
`extern` (see its own entry above) — this one change alone fixed gateway
reachability, both IPv4 and IPv6, but also removes `extern`'s only path
outside by construction, restored by a new **uplink**: a second,
point-to-point tap+relay link (reusing `network_tap_relay.h`, see its own
entry below for the new `attach_host_side_to_bridge=false` mode this
needed) between the network's private namespace and the host's root
namespace, on its own small deterministic `169.254.0.0/16` transit subnet
(`uplink_transit_addresses()`/`uplink_transit_subnet()`, `.cpp`-local,
same `fnv1a()`-based derivation `bridge_name()` already uses).
`ensure_uplink_provisioned()` (`.cpp`-local, called from
`ensure_network_provisioned()` for `extern` only): idempotent via its own
`uplink_provisioned()` device-existence check; creates the relay
(`uph<hash>` in host root, `upn<hash>` in the private namespace), assigns
each end an address, sets the private namespace's own default route via
the uplink, enables `ip_forward`, and adds a `MASQUERADE` rule in host
root for the transit subnet — plus three more pieces, each independently
required and each found by real on-device testing, not assumed (full
detail, including exactly how each was diagnosed, in
`docs/networking-design.md`'s own section): an `iptables -I FORWARD 1`
accept rule (inserted at the front — Android's own `tetherctrl_FORWARD`
chain unconditionally drops everything reaching it, so an appended rule is
structurally unreachable), an outbound `ip rule` routing the uplink's own
traffic into whichever table `discover_default_table()` (`.cpp`-local,
parses `table <N>` out of `ip route get 8.8.8.8`'s own output — not
hardcoded, adapts to whichever real network is currently active) names,
and a **return-path** `ip rule` routing traffic *to* the transit subnet
into the plain `main` table regardless of which real interface a reply
arrives on. `TapRelayHandle` gained `root_side_tap_name`
(`network_tap_relay.h`) since, unlike a real container join's own
container-side tap (torn down for free once that session's namespace
goes away), the uplink's host-root-side tap never disappears on its own —
`stop_tap_relay()` removes it explicitly when set. The relay's own pid is
recorded to `$XDG_STATE_HOME/slocker-lite/network-uplinks/<network>`
(`uplink_state_path()`) since it must outlive the single
`ensure_network_provisioned()` call that created it, potentially spanning
many separate `slocker-lite` invocations before `teardown_uplink_state()`
(`.cpp`-local, the reverse of every step above, best-effort throughout)
eventually stops it. **Verified completely end-to-end on the real target
device, from a clean state**: gateway IPv4 0% loss, gateway IPv6 0% loss,
and a real outside destination (`8.8.8.8`) 0% loss (3/3 replies) through a
container on a freshly created `extern` network, both via the veth path
and via `--no-veth`'s tap+relay fallback; also verified on this dev
machine (self-test, a plain veth join, `--no-veth`, and `intern` — still
unaffected, no uplink, `ip route`'s own "Network unreachable" for outside
as intended). IPv6 outside connectivity was investigated separately and
found not practically fixable on this device (confirmed `nft add table
ip6 ...` itself fails — no IPv6 NAT support in this kernel at all, via
either `ip6tables` or `nftables`; the alternative, NDP-proxying real
addresses out of the device's own global prefix, was ruled out too, since
that prefix rotates every ~10 minutes on the network tested against) — so
it stays local-only by deliberate decision, matching what `intern`'s IPv6
side already was.
`probe_veth_support()` (added for the tap+relay fallback, see
`docs/networking-design.md`'s addendum and `network_join.{h,cpp}` below):
the real target device supports `tun`/`tap` but not `veth`
(`CONFIG_VETH` commonly stripped from mobile kernels), so joins there can't
use the veth-pair mechanism this file/`network_join.cpp` otherwise assume.
Probes kernel support the same way `bwrap.cpp`'s
`kernel_supports_namespace()` probes namespace types: forks a child that
`unshare(CLONE_NEWNET)`s into a throwaway namespace and attempts `ip link
add ... type veth peer name ...` there (via the existing `run_process()`)
— the whole namespace, and anything created in it, vanishes with the
child, so no cleanup is needed either way. Cached in a function-local
static (a fixed fact about the running kernel, not something that varies
per network, so a container joining several networks in one run only
probes once). `should_use_veth(network)` combines this with the network's
own `veth` policy flag (`config_file.h`'s `NetworkEntry::veth`, default
`true`) the same "capability and policy are independent gates" way
`namespace_policy_enabled()` (`bwrap.cpp`) already combines kernel support
with `global.unshare-*` policy — both must allow veth for it to actually
be used. **Verified on this dev machine (root, via the scoped `doas`
rule)**: `probe_veth_support()` correctly returns `true` here (a real `ip
link add ... type veth ...` succeeds), and `--no-veth` at network-creation
time correctly persists `NetworkEntry::veth = false`, making
`should_use_veth()` return `false` even though the kernel itself supports
veth — this dev machine's own way to exercise the tap+relay fallback (see
`network_tap_relay.{h,cpp}`, not yet built) without needing the actual
veth-less target device.
- `network_join.{h,cpp}` — joins a just-started `-r/--run` session to each
network named in `-n`. `join_networks()` first waits (bounded, 3s,
20ms-interval `nanosleep()` polling — `wait_for_isolated_net_namespace()`,
`.cpp`-local) for `resolve_namespace_pid()` (`sandbox_process.h`) to name a
child whose net namespace is actually isolated
(`namespace_isolated(outer_pid, ns_pid, "net")`) — necessary because
`run_bwrap()`'s `on_bwrap_pid_known` fires right after `fork()`, before
`bwrap` has done any of its own namespace setup, so that child may not even
exist yet the instant this is called; while it doesn't,
`resolve_namespace_pid()` falls back to returning `outer_pid` itself, so
comparing a namespace to itself naturally keeps the loop going without a
separate "does a child exist yet" check. **Known limitation, not solved
here**: for a very short-lived sandboxed command, the whole session can
exit before this poll ever catches up (confirmed by testing: `-n <net> --
echo hi` reliably timed out) — `bwrap` execs straight into the target
command with no hook point in between namespace creation and exec, so
there's no way for this project to guarantee network setup completes
before a near-instant command already has too. Real (long-running)
networked services are unaffected — confirmed by testing (see below).
For each named network: looked up in `config.networks` (an unknown name is
a per-network error, not fatal to the others); `ensure_network_provisioned()`
(`network_bridge.h`) covers post-reboot recreation; then, per
`should_use_veth(network)` (`network_bridge.h` — combines the network's own
`veth` policy flag with a kernel-capability probe, see that file's own
entry above), either a veth pair is created wherever that network's bridge
lives (`wrap_for_network()`, reused from `network_bridge.h`), the
bridge-side end attached and brought up, the container-side end moved into
the session's own namespace (`ip link set ... netns <ns_pid>`) and renamed
`eth<N>`, **or**, when veth isn't available or the network was created
with `--no-veth`, `create_tap_relay()` (`network_tap_relay.h`, see below)
is used instead, producing the same end state (a ready `eth<N>` in the
container's namespace) via two tap devices and a relay process rather than
a kernel veth pair. `eth<N>`'s `N` is the network's position in the `-n`
list, so multiple joins each get a distinct interface, regardless of which
strategy created it — everything downstream (IP assignment, routes, the
address returned to `port_forward.h`) is identical either way, since it
only ever operates on `eth<N>` by name. Unlike a veth pair (torn down by
the kernel automatically once the session's namespace goes away, whatever
else fails), a tap relay is an independent process with no such automatic
cleanup — if any step after `create_tap_relay()` succeeds fails later in
`join_one_network()` (address exhaustion, a failed `ip addr add`/route
command), a `fail()` helper (`.cpp`-local, only present when a relay was
actually created) calls `stop_tap_relay()` before returning `nullopt`, so
a partial failure doesn't leak the relay process. **Address
allocation, `pick_free_address()`, needed a real fix during testing, not
just design**: an interface's actual assigned IP lives inside its own
private per-container namespace, invisible from the bridge's own namespace
— an earlier version queried `ip -o addr show master <bridge>` (only the
*host* side of each veth, with no address of its own, is visible there) and
always saw nothing, so two concurrently-running containers on the same
network were both handed the identical address (confirmed by testing:
`10.168.0.2` twice). Fixed by giving each candidate address its own tiny
lock file under `xdg_state_dir() / "net-leases"` (`pid_file.h`) and holding
an exclusive, non-blocking `flock()` on it via an intentionally
never-`close()`d fd — the same technique `pid_file.h`'s own `SessionLock`
uses for session liveness, released automatically by the kernel the
instant this process exits for any reason, no explicit release step or
cleanup sweep needed. Picking a free address is then just "the first
candidate (`network_subnet.h`'s `ipv4_host_address()`/`ipv6_host_address()`,
`n = 2, 3, ...`) whose lock file isn't already held." For an `extern` join,
`ip route replace default via <gateway> dev eth<N>` (`replace`, not `add`,
so a container joining a *second* extern network doesn't fail outright with
"File exists" — whichever extern network is joined last ends up as the
effective default route; `intern` gets no default route at all, matching
the design's "no route out exists" intent — the connected route for the
local subnet is already automatic once an address is assigned, no explicit
route command needed for same-bridge reachability regardless of kind).
Every step failure is logged specifically (which command, which network)
and best-effort: `join_networks()` returns one `JoinedNetwork {network,
container_ip, relay}` per network that actually joined (in `-n` order, so
shorter than the request list on any partial failure), never fatal to the
already-running session (network setup can only happen after `bwrap`'s own
namespace exists, i.e. potentially after the sandboxed command is already
running) — this return value exists specifically for `port_forward.h`
(see below) to resolve a `-p` spec against which networks/IPs are actually
usable, not as a pass/fail signal on its own; `relay` (`nullopt` for a
veth-joined network) is what `run_container()` (`commands.cpp`) collects
to call `stop_tap_relay()` on after `run_bwrap()` returns, mirroring how it
already collects `active_port_forwards` for `-p`'s own cleanup. An empty
`network_names` returns immediately (no namespace wait at all), so callers
that always invoke this once `on_bwrap_pid_known` fires for any reason
(`commands.cpp` also fires it for `-D/--daemonize` alone, with no `-n`)
don't pay for a wait that has nothing to do. Veth teardown
needs no explicit code: the kernel destroys an entire veth pair (both
ends, including the one still attached to the bridge) the instant *either*
end's owning namespace is destroyed, so a session's veths disappear on
their own once its namespace does — only the bridge/iptables/persistent-
namespace state is deliberately left behind (`network_bridge.h`'s
reboot-reconciliation design); a tap relay instead needs the explicit
`stop_tap_relay()` call described above, since it's an independent process
with no namespace of its own to be torn down by. **Verified end-to-end on
this dev machine
(root, via a scoped `doas` rule)**: two concurrently-running containers on
the same `intern` network got distinct addresses and could ping each
other; an `intern`-joined container could not reach the outside
(`Network unreachable`); an `extern`-joined container reached the real
internet through the bridge's NAT; a container joining both an `intern`
and an `extern` network simultaneously got two working interfaces
(`eth0`/`eth1`) with neither one breaking the other.
**Real, separate bug found while testing this commit — since fixed**
(`exec_session.{h,cpp}`, see that file's own entry below): `-x/--exec`
deliberately never joined the `net` namespace type, written back when this
project genuinely never isolated networking at all, so there was nothing
to join. Once `-r/--run` sometimes isolates networking (whenever any `-n`
was given), `-x/--exec`'ing into such a session saw the *host's* network
stack instead of the container's — confirmed directly: execing into a
session running a network-isolated `httpd` showed the host's own unrelated
listening ports and failed to reach the container's own service on
`127.0.0.1`. Fixed by joining `net` too, the same way `-x/--exec` already
joins `mnt`/`uts`/`ipc`/`pid`/`cgroup`/`user` when they differ from the
caller's own — reverified afterward: execing into that same session now
correctly shows the container's own `eth0` and reaches its own service on
`127.0.0.1`, while execing into a plain session with no `-n` at all is
unaffected (still just loopback, whether or not the kernel happened to
give it its own otherwise-empty net namespace via the default
`global.unshare-net` policy).
**Real bug reported from the real target device (`-n <extern network> --
/bin/sh`, tap+relay fallback): the container-side tap device wasn't always
immediately visible.** The user's own log showed `nsenter --net=/proc/<ns_pid>
/ns/net -- ip addr add 10.168.0.2/24 dev eth0` failing with `"Cannot find
device \"eth0\""` right after `network_tap_relay.h`'s relay had already
created it — and confirmed by hand that simply retrying the whole session a
few times eventually worked. A first fix added a bounded (~500ms) retry
around the steps that touch the just-created `container_if` — later found
insufficient (see below) and removed again; `join_one_network()` now uses a
plain, single-attempt `run()` for every step, same as before any of this.
**A tempting "fix" investigated and ruled out by direct A/B testing on this
dev machine, not just reasoned about**: the obvious first instinct — have
the relay *itself* self-verify the device is visible (a same-process check
via its own `run_process()` call, immediately after `open_tap()`, before
ever reporting success) — was tried first, in `network_tap_relay.cpp`'s
`relay_child_main()`. It made things *worse*, not better: it made the
container-side device **permanently invisible to every external `nsenter`
afterward, 100% reproducibly** (confirmed with a 10-second retry budget —
never once became visible), on a mechanism that had otherwise worked
correctly and instantly on every single real session tested earlier this
same day, with no retries ever needed. Root cause not fully understood
(something about forking a subprocess that inherits the tap fd — opened
via `open("/dev/net/tun", O_RDWR)`, deliberately *not* `O_CLOEXEC` — while
still holding it open, immediately after device creation, appears to
corrupt the device's *external* visibility specifically on this kernel;
the *same* process's own view of the device it just created stayed correct
throughout). The lesson that survived into the final fix: never add an
internal, same-process/fd-holding self-check to the relay.
**The retry fix above turned out to be insufficient**: a further round of
real-device testing showed a *different* failure — `ip addr add` against
the container-side device would sometimes succeed, only for the very next
command against that same device (`ip link set eth0 up`) to fail with
"Cannot find device", exhausting every retry. The device wasn't merely
slow to become visible after creation; it was **disappearing on its own**,
consistent with the underlying `ioctl(TUNSETIFF)`-created device (no
`IFF_PERSIST`) having a more fragile lifetime on that kernel than "stays
alive as long as the one fd that created it stays open." Per the user's
own suggested direction, the fix was structural, not another retry: both
tap devices are now created ahead of time via an external `ip tuntap add
dev <name> mode tap` (`create_persistent_tap()`, `network_tap_relay.cpp`
see that file's own entry below for the full detail), which sidesteps the
whole class of symptom by making the device a genuinely persistent
netdevice with no tie to any fd or process. All retry logic (`run()` is
used unconditionally, everywhere) was removed as part of this — the
earlier retry was compensating for a problem this fix removes outright,
not one it makes more likely to need retrying.
- `port_forward.{h,cpp}` — implements `-p`. `parse_port_forward_spec()`
splits `"[<network>:]<host-port>:<container-port>"` on `':'` (2 or 3
fields; the network name is deliberately restricted to excluding `':'` --
`is_valid_network_name()`, `network_subnet.h` -- specifically so this
split stays unambiguous) and validates both ports are `1..65535`
(`.cpp`-local `parse_port()`) -- pure syntax/range parsing, no knowledge of
which networks exist or joined; that's `add_port_forward()`'s job, called
later once `join_networks()` (`network_join.h`) has actually run.
`add_port_forward()` resolves `spec.network` against the `JoinedNetwork`
list -- by name if given (erroring if that network wasn't successfully
joined, or isn't `extern`: an `intern` network's bridge has no path from
the host at all, so forwarding into one could never work), or, if unset,
the container's sole joined `extern` network (erroring if none or more
than one, rather than guessing). Then adds one iptables `DNAT` rule to
**both** `nat PREROUTING` *and* `nat OUTPUT` -- a real bug caught by
testing, not assumed: `PREROUTING`-only left `curl <this host's own real
IP>:<host-port>`, run *on this same host*, connection-refused, since
`PREROUTING` only ever sees packets arriving from an actual network
interface, never locally-generated ones (those go through `OUTPUT`
instead) -- the same split Docker's own DNAT setup already accounts for.
Also adds one `FORWARD ACCEPT` rule for the destination (in case of a
default `FORWARD DROP` policy, which would otherwise silently eat the
forwarded traffic even though the `DNAT` itself succeeded); if a later
rule fails after an earlier one already landed, those are removed again so
a failure doesn't leave a half-applied mapping. **Known limitation, not
solved here, also found by testing**: `curl localhost:<host-port>` (or any
`127.0.0.0/8` destination) specifically still doesn't work even with both
`DNAT` chains covered -- confirmed to be a separate problem, NAT
hairpinning: once `DNAT` rewrites the destination to the container's IP,
the packet still carries its *original* source address (`127.0.0.1`); the
container's own kernel sees an inbound packet claiming to be *from*
loopback arriving on a non-loopback interface (`eth<N>`) and drops it as a
martian source. (A `net.ipv4.conf.{all,lo}.route_localnet=1` sysctl was
tried and confirmed *not* to fix this on its own, then removed again
rather than left in as dead/superstitious code.) A full fix needs source
masquerading scoped to exactly this case (matching only host-local
traffic, not genuine external clients -- unconditionally masquerading
would lose the real client IP for those, a regression) or a userland
proxy, the approach Docker itself historically used for the same reason --
out of scope here; `curl <this host's real, externally-reachable IP>:
<host-port>` (verified working) is the actually-relevant path `-p` exists
for. `remove_port_forward()` (`commands.cpp`'s `run_container()`, called
for each `ActivePortForward` collected during `on_bwrap_pid_known`, after
`run_bwrap()` returns) removes the exact same rules `add_port_forward()`
added -- best-effort, logs a warning on failure, never fatal. **Verified
end-to-end on this dev machine (root, via a scoped `doas` rule)**: a
container serving HTTP on an `extern` network with `-p 8080:80` was
reachable via `curl <host's real IP>:8080` from the host; the rule was
confirmed gone (connection refused) after the session was killed.
**Crash-orphan sweep** (commit 6 of `docs/networking-design.md`'s
sequence): unlike `join_networks()`'s veths (torn down automatically by
the kernel once the session's namespace goes away) or the bridges/
persistent namespaces themselves (deliberately meant to outlive any one
session — `network_bridge.h`'s reboot-reconciliation design), a `-p`
mapping's iptables rules are host-global state with no automatic teardown
at all — if `slocker-lite` itself is killed/crashes before reaching its
own `remove_port_forward()` calls, those rules simply outlive the session
forever otherwise (`bwrap` itself dies immediately in that case too, via
`--die-with-parent`, so the *container* never becomes a stray process
needing separate handling — only these rules can). `port_forward_state_path()`
resolves `xdg_state_dir() / "port-forwards" / "<container_name>-<pid>"`
deliberately the *exact* same naming scheme as `session_pid_file_path()`
(`pid_file.h`), so `clean_stale_port_forwards()` can cross-reference this
directory's filenames directly against `list_sessions()`'s own
`SessionInfo::path` to reuse its liveness check, rather than re-deriving
pid liveness a second, drifting way. `record_port_forwards()` writes one
line per mapping (`"<host_port> <container_ip> <container_port>"`) to that
path — a no-op if there's nothing to record. `clean_stale_port_forwards()`
(`commands.cpp`'s `clean_processes_command()`, alongside
`clean_stale_sessions()`) scans that directory: a record whose filename
doesn't match any currently-*running* session is stale — every line is
parsed back into an `ActivePortForward` and removed
(`remove_port_forward()`) before the record file itself is deleted; a
record whose session is still running is left completely untouched.
**Verified via a controlled scratch test** (root wasn't needed for the
logic itself — only the underlying `iptables -D` calls, already proven
working as root above; killing a root-owned `slocker-lite` process
directly, bypassing `--kill`'s own graceful cgroup-based teardown, wasn't
achievable through the scoped `doas` rule this session has, which only
permits running `slocker-lite` itself): a real running session's own pid
file was used to construct a *matching* port-forward record (left
untouched by the sweep, confirmed still present afterward) alongside a
*fabricated* record for a nonexistent pid (correctly identified as stale,
its rule-removal attempted — visibly failing only for lack of root in this
particular rootless test — and its record file actually removed,
reported as `removed stale port-forward rules for 'faketest-999999'`).
- `network_tap_relay.{h,cpp}` — a tap-backed substitute for one veth pair,
used when `should_use_veth()` (`network_bridge.h`) is false: either the
running kernel doesn't support veth at all (the real target device's
kernel supports `tun`/`tap` but lacks `CONFIG_VETH`, commonly stripped
from mobile kernels — see `docs/networking-design.md`'s tap+relay
addendum), or the network was created with `--no-veth` (`cli_args.{h,cpp}`)
specifically to exercise this path on a veth-capable machine. **Why tap
can't just replace veth 1:1**: a veth pair is two real kernel netdevices,
switched between (or into a bridge) entirely by the kernel; a tap device
only has *one* kernel-side netdevice — the other "end" is a raw-Ethernet-
frame file descriptor only a userspace process can read/write, so there's
no second kernel endpoint to attach to a bridge (the same reason
`slirp4netns`/QEMU's own tap networking need a userspace process on the fd
side). `create_tap_relay(network, bridge, host_tap_name, container_ns_pid,
container_if_name)` reproduces a veth pair's role with two tap devices and
one relay process that copies bytes between them, **reusing the existing
bridge as the switching fabric** so `network_bridge.cpp`'s
`provision_bridge()`/NAT setup needs no changes at all: a host-side tap
device (created wherever `network`'s bridge lives — `network_bridge.h`'s
`wrap_for_network()` namespace, reached here via a direct `setns()` rather
than that function's `nsenter`-argv-wrapping, since this whole sequence
must keep running, and later hold onto live fds, across each namespace
switch, not just for the duration of one external command) gets enslaved
to `bridge` exactly like veth's host-side end does in
`network_join.cpp`'s `join_one_network()`; a container-side tap device
gets created directly inside the namespace named by `container_ns_pid`,
**named `container_if_name` (e.g. `eth0`) from the start** — no
peer-name-then-rename dance needed, unlike veth. `host_tap_name` is
caller-provided (not derived here) specifically so a future caller
(`join_one_network()`, once this is wired in) can reuse its own existing
fnv1a-based veth-naming scheme rather than this file growing a second,
drifting copy of that six-line hash. Each device is now created two-step:
`create_persistent_tap(name)` (`.cpp`-local) first runs an external `ip
tuntap add dev <name> mode tap`, then `open_tap()` (`.cpp`-local, mostly
unchanged) `open("/dev/net/tun")` + `ioctl(TUNSETIFF, IFF_TAP |
IFF_NO_PI)`s onto that already-existing device (`IFF_NO_PI` so both ends
agree on raw-frame framing with no extra header) — `open_tap()` now only
*attaches* an fd to a device, it no longer *creates* one. **This split
replaced an earlier, simpler design** where `open_tap()` alone both
created (via the same `ioctl`, with no `IFF_PERSIST`) and attached, on the
assumption the device would then simply disappear on its own once its
one-and-only fd closed, the same "no explicit teardown" property veth
already has — see this entry's own "tap devices need to be created
persistently" paragraph further down for why that assumption turned out
to be wrong on the real target device, and `docs/networking-design.md`'s
matching section for the full incident writeup. The relay child's
entire setup sequence — enter the network's own namespace first, always
now (`persistent_netns_path()`, `persistent_netns.h` — used to be `intern`
only; see `network_bridge.{h,cpp}`'s own "Resolved: `extern` had no
connectivity" entry above for why `extern` needs this too now), create+
attach the host-side tap, `setns()` into the container's namespace (the
already-open host-side fd stays valid across this switch — fds aren't
namespace-scoped, only their *creation* is, the same property
`wrap_for_root_namespace()`, `bwrap.cpp`, already relies on), create the
container-side tap — reports success/failure back to `create_tap_relay()`
over a `pipe2(O_CLOEXEC)` (same handshake shape `daemonize()`,
`daemonize.cpp`, already uses), then falls into an unbounded
`poll()`/`read()`/`write()` loop copying raw frames bidirectionally
between the two fds — this loop *is* the actual "veth wire," just
implemented once in userspace instead of by the kernel. No `SIGTERM`
handler is installed in the relay: default disposition (terminate) already
closes both fds on the way out. `stop_tap_relay()` sends `SIGTERM`, reaps
the process, then explicitly `ip link del`s the host-side device
(`wrap_for_network(handle.network, ...)`, reaching wherever it lives —
the network's own persistent namespace, both kinds — hence
`TapRelayHandle` carrying its own `NetworkEntry`) — now
required since the device is persistent and no longer disappears just
because the relay's fd closed (see below). The container-side device
needs no matching step: it lives inside the container's own network
namespace, which the kernel already tears down (every interface inside
it, persistent or not, along with it) once the session itself ends —
**except** the uplink's own second tap (`network_bridge.cpp`'s
`ensure_uplink_provisioned()`), which lives directly in the host's root
namespace instead and never goes away on its own; `TapRelayHandle` gained
`root_side_tap_name` for exactly this case, and `stop_tap_relay()` removes
it too (unwrapped, always host root by construction) when set. Likewise
`create_tap_relay()` gained an `attach_host_side_to_bridge` parameter
(default `true`, no change for either existing call site) — `false` skips
the `master <bridge>` step entirely and just brings the host-side tap up
plain, used by the uplink since it's deliberately a point-to-point routed
link, not another bridge port.
Once `create_tap_relay()` returns successfully,
`container_if_name` is a completely ordinary interface from the
container's own point of view — `join_one_network()`'s existing IP
assignment/route/DNAT-target-address code (unchanged, not yet wired to
call this) needs no changes at all. See `self_test.{h,cpp}` above for how
this file's own create/attach/teardown cycle was verified end-to-end in
isolation first, before being wired into `join_one_network()`.
**Real fd-leak bug caught by direct testing, not assumed**: the relay
child, unlike every other forked child elsewhere in this project, never
`exec()`s — so `O_CLOEXEC` on fds created *before* this fork (e.g.
`daemonize.cpp`'s own report-pipe write end, still open in the forking
process at this point since `report_daemon_started()` — which closes it —
hasn't run yet when `join_networks()` is called) never takes effect,
since it only closes fds *across `exec()`*, not across a fork that never
execs. Without a fix, the relay child inherited a live copy of that pipe's
write end and never closed it, so `daemonize()`'s read-until-EOF in the
*original, pre-fork* process blocked forever, even after
`report_daemon_started()` closed its own copy — a pipe only reports EOF
once *every* copy of its write end, across every process, is closed.
Confirmed directly: `-r -D -n <no-veth network> -- sleep 600` hung
indefinitely; killing the session (which reaches `stop_tap_relay()` via
`run_container()`'s own post-`run_bwrap()` cleanup, closing the leaked
copy) immediately unblocked the original process. Fixed by
`close_inherited_fds()` (`.cpp`-local): scans `/proc/self/fd` and closes
everything except stdin/stdout/stderr and the report pipe's own write
end, called as the very first thing in `relay_child_main()`. A
first-attempt companion fix — adding the relay's pid to the session's own
cgroup (`session_cgroup.h`) so `--kill` would reach it directly, since a
relay is a *sibling* of bwrap rather than a descendant and so would never
inherit cgroup membership on its own — was tried and then **reverted**:
`remove_session_cgroup()` runs *inside* `run_bwrap()`, before
`run_container()` ever gets to call `stop_tap_relay()`, so the cgroup was
still non-empty (the relay still in it) at removal time, and every
session using this fallback left a stray, never-removed cgroup directory
behind (`rmdir` failing with `EBUSY`, confirmed by testing). Since the
ordinary flow already stops the relay correctly on its own (killing bwrap
unblocks `run_bwrap()`'s own `waitpid()`, letting `run_container()`
finish its normal cleanup, `stop_tap_relay()` included) and the only gap
left by *not* doing this is a benign, self-resolving race in `--kill`'s
own "has it fully stopped" check (a genuine crash of the whole session
process, not just bwrap, is a separate, already-scoped concern — see the
crash-orphan sweep below), the added complexity wasn't worth it.
**Real bug reported from the real target device: tap devices need to be
created persistently, not tied to the relay's own fd lifetime.** Two
rounds of real-device testing (`network_join.cpp`'s own entry above has
the full incident writeup) found the container-side device intermittently
either not immediately visible after creation, or — worse, found on the
second round — visible and usable for one command (e.g. `ip addr add`
succeeding) and then gone for the very next one (`ip link set ... up`
failing with "Cannot find device"), on a kernel where the bare
`ioctl(TUNSETIFF)`-created (no `IFF_PERSIST`) device evidently has a more
fragile lifetime than "stays alive as long as its one creating fd stays
open." Per the user's own suggested direction, the fix (see
`create_persistent_tap()` above) creates both tap devices ahead of time
via an external `ip tuntap add dev <name> mode tap` — the same technique
QEMU/libvirt use to let an unprivileged process attach to a tap device set
up ahead of time — turning each into a genuinely persistent netdevice with
no tie to any fd or process at all, the same as a veth pair already is.
All retry logic from the first round's fix (`network_join.cpp`'s
`run_with_retry()`, `self_test.cpp`'s `wait_for_container_device_visible()`)
was removed once this structural fix made it unnecessary — see both
files' own entries.
**Verified end-to-end on this dev machine (root, via the scoped `doas`
rule), using a `--no-veth` `extern` network specifically to exercise this
path**: two real containers joined the same network, each getting a
distinct address (`10.168.0.2`/`10.168.0.3`) via the tap+relay path with
no veth involved at all, and pinged each other successfully (0% packet
loss, confirmed repeatably). **Gateway/outside reachability — originally
reported as an unconfirmed gap here, since resolved**: neither container
could initially reach the network's own gateway IP, despite ARP resolving
correctly (ruling out an L2/relay-framing problem) and the identical
bridge/subnet working perfectly via veth instead (ruling out every
environment-level explanation — host firewall, `rp_filter`, tried at
several scopes and confirmed not to fix it — since those would affect
both paths identically). **Actual cause, found once retested on a clean
host**: accumulated leftover bridges/iptables rules from many earlier
rounds of manual testing — `--delete-network` (before
`--delete-network-full` existed, see that flag's own entry below) never
tore down live host state, so stale rules/bridges from unrelated earlier
test networks were still present and interfering. After manually clearing
all of it and retesting fresh: a `--no-veth extern` network's gateway and
a real external host both answered ICMP with 0% loss, and a raw TCP
connect (`nc`) to an external host completed cleanly (a separate `wget`
segfault against the same host was confirmed to be an unrelated busybox
bug, reproducing identically regardless of join mechanism). Peer-to-peer
connectivity and gateway/outside reachability are both now confirmed
working through the tap+relay fallback on this dev machine —
`--delete-network-full` exists specifically so this class of
stale-state-masking-as-a-bug can't recur.
**Re-verified end-to-end on this dev machine after the persistent-device
redesign above**, again with `--no-veth` forcing the fallback: a single
container repeatedly used its tap-relay-backed `eth0` across several
commands in a row (`ip link show`, `ip addr show`, two rounds of `ping`)
with no disappearance between commands — the exact symptom the real
device hit — and both gateway ping and outside/internet ping (`8.8.8.8`)
succeeded at 0% loss. Session cleanup left no leftover host-side tap
device behind (only the bridge itself, deliberately left standing per
this project's reboot-reconciliation design); `-t/--test`'s own
`tap-relay create/attach/teardown` case (updated per `self_test.{h,cpp}`'s
own entry above) passes reliably across repeated runs.
**Real bug reported from the real target device: joining 2+ networks in
one `-r/--run` left every network after the first permanently
unreachable, regardless of extern/intern.** Root-caused via `strace -f`
on the real device (the user's own suggestion, after several
inconclusive timing-based experiments): the second network's own relay
died on its literal first frame — `write(fd_container, ..., 86) = -1
EIO`, immediately followed by `exit_group(0)`. `EIO` writing to a tap fd
means the device isn't administratively up yet, and it genuinely wasn't:
`create_tap_relay()` returns, and this relay starts polling, the instant
the container-side tap device is *created*; `join_one_network()`
(`network_join.cpp`), a *different* process, still has its own `ip addr
add`/`ip link set <if> up` steps left to run afterward for that same
device — confirmed via the trace's own timestamps, `ip addr add ... dev
eth1` ran *after* the relay's fatal write. For the first network joined
this race is narrow enough that no frame ever arrives first; for the
second (and any later) network, something reliably delivers a frame
before the interface is up, and the previous code treated any `write()`
failure as fatal — exiting for good on that single `EIO`, so the network
never worked again for the rest of the session. Fixed in the relay's
frame-forwarding loop by retrying specifically on `EIO`/`ENETDOWN` (both
mean "not up yet", a startup race, not a torn-down namespace) with a
short bounded backoff (up to 50 × 20ms = 1s) instead of exiting
immediately — generous compared to the ~14ms gap actually observed in
the trace. **Methodology note**: an earlier `strace -f` attempt, wrapped
in `timeout 30`, produced a misleadingly corrupted trace — GNU `timeout`
sends its kill signal to the whole process group by default, and
`strace`'s own tracing overhead was large enough (mount steps that
normally take seconds took 3+ minutes under trace) that the real
wall-clock timeout elapsed mid-setup, killing several traced children
prematurely; dropping the `timeout` wrapper entirely produced a clean,
complete trace. **Verified end-to-end on the real target device** with
both 2 and 3 `intern` networks joined simultaneously in one session, all
gateways reachable at 0% packet loss, clean teardown, no leftover state.
**Crash-orphan sweep**, the direct tap+relay analog of `port_forward.h`'s
own (see its own entry below): unlike a veth pair or a session's own
bridge/persistent-namespace state, a relay process is host-global state
with no automatic teardown at all if `slocker-lite` itself is
killed/crashes before reaching its own `stop_tap_relay()` calls (`bwrap`
still dies immediately in that case via `--die-with-parent`, so only the
*relay* — never the sandboxed container itself — can actually leak).
`tap_relay_state_path()` resolves `xdg_state_dir() / "tap-relays" /
"<container_name>-<pid>"` — the same naming scheme
`session_pid_file_path()`/`port_forward_state_path()` already use, so
`clean_stale_tap_relays()` can cross-reference filenames directly against
`list_sessions()`'s own `SessionInfo::path`. `record_tap_relays()` writes
one line per relay (`"<relay_pid> <host_tap_name> <kind> <network_name>"`,
`<kind>` = `"extern"`/`"intern"`, `<network_name>` last since it's the one
field that can contain whitespace) to that path — a no-op if there's
nothing to record; the `<kind>`/`<network_name>` fields were added
alongside the persistent-tap-device redesign above, so a later sweep can
reconstruct a `NetworkEntry` and reach the right namespace to remove the
now-persistent host-side device too, not just kill the relay process.
`clean_stale_tap_relays()`
(`commands.cpp`'s `clean_processes_command()`, alongside
`clean_stale_sessions()`/`clean_stale_port_forwards()`) scans that
directory: a record whose filename doesn't match any currently-*running*
session is stale — every relay pid listed is `SIGKILL`ed (best-effort; an
already-dead pid, or one this process was never the parent of, isn't
treated as an error, since this sweep runs from a *separate* later
invocation that can't `waitpid()` an orphan it didn't fork — its true
parent's own exit, or `init` after reparenting, reaps it), the host-side
tap device it named is removed (`ip link del`, via
`wrap_for_network()` against a `NetworkEntry` reconstructed from the
record's own `<kind>`/`<network_name>` fields — best-effort, same as
`stop_tap_relay()`'s own removal) before the record file itself is
deleted; a record whose session is still running is left completely
untouched. **Verified via a controlled scratch test**,
the same shape `port_forward.h`'s own sweep test used: root wasn't needed
for the sweep *logic* itself (only real tap/bridge creation needs it),
so this ran as a plain rootless daemonized session (`-D`, no `-n`) to get
a real, live pid + container name, alongside two hand-written record
files in that same rootless `$XDG_STATE_HOME` — one *matching* the live
session (confirmed left untouched by `--clean-processes`) and one
*fabricated* for a nonexistent pid (confirmed identified as stale, its
`kill()` attempt failing harmlessly with `ESRCH`, and its record file
actually removed, reported as `removed stale tap-relay processes for
'faketest-999999'`); killing the real session and re-running
`--clean-processes` then correctly swept its own now-stale record too.
- `session_cgroup.{h,cpp}` — gives `--kill` (`kill_session.{h,cpp}`, see
below) a reliable way to find every process a session ever started, however
deeply forked/daemonized/reparented, by putting it in a dedicated cgroup v2
group from the moment it starts. `cgroup_v2_available()` checks for
`/sys/fs/cgroup/cgroup.controllers` (the same signal systemd's own
unified-hierarchy detection uses) — only cgroup v2 is supported; v1 (which
splits per-controller into separate hierarchies with no unified
`cgroup.procs` at the top) is deliberately out of scope, since this
project's real target (Android) has used the unified v2 hierarchy by
default since Android 12. `session_cgroup_path()` is deterministic —
`/sys/fs/cgroup/slocker-lite/<name>-<pid>/`, reusing `pid_file.h`'s own
`sanitize_for_filename()` — so no separate lookup state is needed anywhere.
`create_session_cgroup()` is called from `run_bwrap()`'s `on_start` callback
(`bwrap.cpp`, see below), the same spot `create_session_lock()` already
fires from: `create_directories()`'s the leaf directory (also creating the
`slocker-lite/` parent the first time — a plain grouping cgroup, no resource
controllers are ever enabled on it via `cgroup.subtree_control`, so the "no
internal processes" restriction that comes with actually delegating
controllers never applies here) and writes the bwrap pid into its
`cgroup.procs`. From that point on, every process bwrap (or anything it
execs into) forks inherits this cgroup automatically, permanently —
including anything that later daemonizes/double-forks and gets reparented,
unlike pid-namespace child membership (only the processes `clone()` itself
creates) or process-group membership (many daemonizing services explicitly
`setpgid()`/`setsid()` away from it on purpose). Best-effort, mirroring
`create_session_lock()`: returns `nullopt` (logging a warning, never fatal)
if cgroup v2 isn't available, or the directory can't be created/written
(no delegated subtree when running rootless, or an SELinux policy blocking
cgroupfs writes even for a root-euid process, are both real, confirmed-by-
testing causes on the two environments this project actually runs on).
`remove_session_cgroup()` (called from the same post-`run_process_foreground()`
spot `release_session_lock()` already is) only succeeds once the cgroup is
empty — a straggler process still alive at normal exit (e.g. a daemonized
process that outlived the session's own main command, a pre-existing
exposure independent of this feature) leaves it in place with a warning, not
a fatal error. `session_cgroup_pids()` reads `cgroup.procs` — this is the
actual answer to "gather every process running inside the container":
unlike anything derived from `/proc` parent-pid chains or pid namespaces,
cgroup membership reliably includes every process the session ever started.
`session_cgroup_supports_kill()`/`kill_session_cgroup()` wrap the
`cgroup.kill` knob (Linux 5.14+): writing `"1"` to it atomically `SIGKILL`s
every process currently in the cgroup in one step.
- `sandbox_process.{h,cpp}` — process-tree/namespace-resolution utilities
shared by `exec_session.{h,cpp}` and `kill_session.{h,cpp}` (see both
below); pulled into their own file (rather than staying private to
`exec_session.cpp`, where `resolve_namespace_pid()` originally lived) once
`--kill` needed the exact same "find the real sandboxed child" logic, to
avoid a second, drifting copy. `resolve_namespace_pid()` is unchanged from
its original `exec_session.cpp` form (see that entry for the full
reasoning: bwrap's own outer/tracked pid never actually enters the
pid/uts/ipc/cgroup namespaces it creates for its clone()'d child, only that
child does). Two new utilities added alongside it for `kill_session()`:
`pid_namespace_isolated(outer_pid, ns_pid)` compares
`/proc/<outer_pid>/ns/pid` and `/proc/<ns_pid>/ns/pid`'s own `readlink()`
targets directly — true only when bwrap's `clone()` actually created a
separate pid namespace for its child (`--unshare-pid` was requested *and*
the kernel supported it), the precondition for the kernel's own guarantee
that killing a pid namespace's pid 1 forcibly tears down every remaining
process in it. Generalized into `namespace_isolated(outer_pid, ns_pid,
ns_type)` (parametrized over which `/proc/<pid>/ns/<ns_type>` entry to
compare) once `network_join.{h,cpp}` (see below) needed the exact same
check for `"net"` instead of `"pid"`
`pid_namespace_isolated()` is now just `namespace_isolated(outer_pid,
ns_pid, "pid")`, kept as its own function since `kill_session.h` already
depends on that exact name/signature. `collect_descendant_pids(root)` generalizes
`resolve_namespace_pid()`'s own `/proc/<n>/stat` ppid-scanning fallback to
collect a whole transitive tree (root included) instead of just one child,
sharing the actual stat-parsing loop between both via a private
`build_ppid_map()` (one `/proc` pass, used by both the single-child lookup
and the full-tree collection). Reliable specifically when `root` is a
genuinely isolated pid namespace's own pid 1: anything that reparents
within it (e.g. a daemonizing service) is guaranteed by the kernel to land
back on `root` itself, unlike on a kernel without pid namespace support,
where it escapes to the *host's* real pid 1 instead (see
`kill_session.{h,cpp}` below for exactly this scenario, confirmed on a real
target device).
- `daemonize.{h,cpp}` — implements `-D/--daemonize`'s fork/detach mechanics.
`daemonize(container_name)` sets up a `pipe2(..., O_CLOEXEC)` pair (so it
never leaks into `bwrap`/the sandboxed command, same reasoning as the pid
file's own `O_CLOEXEC`) and `fork()`s. The **child** calls `setsid()`
deliberately here, not via re-adding bwrap's own `--new-session` (removed
earlier, see the `build_bwrap_args()` comment): `--new-session` only calls
`setsid()` for the deeply-nested sandboxed command *inside* bwrap's own
namespace setup, leaving the outer `bwrap`/`nsenter`/`slocker-lite` processes
still attached to the *original* session and still receiving its signals
(e.g. a `SIGHUP` when the controlling terminal closes) — not real
daemonization. Calling `setsid()` in our own forked child, *before* it execs
into `nsenter`/`bwrap`, detaches the entire chain at once, since `exec()`
never changes session membership — confirmed by direct testing
(`ps -o pid,sid,pgid,tty`): the daemon child becomes its own session leader
with no controlling tty, and `bwrap` (a later descendant) shares that same
session, also with no tty. The child also `sigaction()`s `SIGHUP` to
`SIG_IGN` (survives the later `exec()` into `nsenter`/`bwrap`, unlike a real
handler, which `exec()` resets to default — confirmed by sending `SIGHUP`
directly to a running daemonized `bwrap` pid and it staying alive), then
redirects stdin to `/dev/null` and stdout/stderr to a log file at
`session_log_file_path(container_name, getpid())` (`pid_file.h`) — named
after its *own* pid since the real session pid (`bwrap`'s) isn't known yet.
If the log directory/file can't be set up at all, that's a **hard** failure
here (`_exit(1)`), not best-effort — silently losing the very output
`--daemonize` was asked to capture would defeat the point of the flag. The
child reports `"LOG <path>\n"` over the pipe immediately (so the parent can
show a useful location even on failure) and returns `nullopt` to its caller
(`run_container()`, `commands.cpp`), which then falls through into the rest of
that function's existing body completely unchanged — **the daemonized child
is what runs the whole rest of `run_container()`, including the unmount/
cleanup that already existed after `run_bwrap()` returns; no separate
watcher/reaper process exists**. The **parent** blocks reading the pipe until
EOF, returning the accumulated `"LOG "`/`"PID "` lines as a `DaemonizeResult`
— the caller then prints it and exits immediately without running any
session logic itself. `report_daemon_started(container_name, pid)` (called
from `run_bwrap()`'s new `on_bwrap_pid_known` callback — see `bwrap.{h,cpp}`
below — the instant the real `bwrap` pid is known) renames the pid-named log
file to `<container_name>-<pid>.log`, re-reports the *updated* `"LOG "` line
(a real bug caught by testing: the parent's first `"LOG "` line names the
pre-rename, daemon-pid-named path — without a second one, the parent would
print a stale filename that doesn't match where the file actually ends up),
then `"PID <pid>\n"` and closes its own end of the pipe — must happen here,
explicitly, rather than waiting for the pipe to close naturally at the end of
the (potentially very long) daemon's lifetime, or the parent would block for
as long as the session runs instead of returning promptly. The pipe's write
fd and the current log path are tracked as private file-scope state in
`daemonize.cpp` (matching `process.cpp`'s own `g_foreground_child_pid`
pattern for "there's only ever one of these per process" runtime state),
since `report_daemon_started()` is called later, from a different function,
not threaded explicitly through every call in between.
- `exec_session.{h,cpp}` — implements `-x/--exec <pid>`: joins an already-running
`-r/--run` session's namespaces via `nsenter` and runs a command inside it in
the foreground. `exec_in_session()` first confirms `pid` is a tracked, running
session via `list_sessions()` (`pid_file.h`) — same liveness check
`--list-processes`/`--clean-processes` already use, no new logic needed there.
**Key discovery, confirmed by direct testing, not assumed**: `pid` (the one
`run_process_foreground()` captured and pid-file-tracked when `-r` launched
`bwrap`) is bwrap's own *outer* process — it sets up the mount and user
namespaces itself, then `clone()`s the actual sandboxed command into fresh
pid/uts/ipc/cgroup namespaces, and `clone()`'s namespace-creation flags only
ever affect the newly created child, never the caller. So the outer process
itself never actually enters those namespaces — comparing
`/proc/<outer_pid>/ns/{pid,uts,ipc,cgroup}` against this process's own showed
them identical, while only `mnt`/`user` differed. `resolve_namespace_pid()`
(`sandbox_process.{h,cpp}` — moved out of this file once `--kill`
needed the exact same logic, see that entry) finds that real inner process
so this can join *its* namespaces instead. For each of `{mnt→--mount, uts→--uts, ipc→--ipc, pid→--pid,
net→--net, cgroup→--cgroup, user→--user}``net` used to be excluded here
(this project never isolated networking at all, back when this comment was
first written), but now that a session started with `-n/--network`
(`network_join.h`) genuinely does get an isolated net namespace, skipping
it left `-x/--exec` seeing the *host's* network stack instead of the
container's — confirmed directly (execing into a network-isolated session
showed the host's own unrelated listening ports and couldn't reach the
container's own service on `127.0.0.1`), fixed by including it the same
way as the other optional types: a session with no isolated net namespace
at all (i.e. `net` identical to ours) just has this entry skipped like any
other, so nothing changes for a session that never joined a network.
`readlink()`s both
`/proc/<ns_pid>/ns/<type>` and `/proc/self/ns/<type>` and only passes
nsenter's corresponding `--type=/proc/<ns_pid>/ns/<type>` flag when they
differ — an identical-namespace re-entry attempt can fail outright
(`setns()`'s own `EINVAL` restriction on re-entering a namespace you're
already in), so skipping is deliberate, not just an optimization. `mnt` is the
one type where a *read* failure (permission denied, or the process vanished)
is treated as fatal, since without it "joining the container" is meaningless;
every other type just degrades to a skip. Always appends
`--preserve-credentials`: without it, `nsenter --user` also tries to
`setuid()`/`setgid()`/`setgroups()` to the target's identity within the new
user namespace, which fails outright (`setgroups failed: Operation not
permitted`) against the `setgroups`-denied unprivileged user namespace bwrap
creates whenever `-r/--run` isn't root — confirmed by hitting this exact
failure during manual testing before adding the flag. Runs the final
`nsenter ... -- <command>` via the existing `run_process_foreground()`
(`process.h`) — same inherited stdio and SIGINT/SIGTERM forwarding as every
other foreground external command, no new process-running logic needed.
`exec_in_session()` also takes optional `--user`/`--group` (mirroring `-r/--run`'s
own): given, they resolve against the session's own `/etc/passwd`/`/etc/group`
(fetched via `cat` run through the same `nsenter` join, since this process can't
otherwise see into that namespace, then handed to `resolve_user_and_group()`
`user_spec.h`, see below); if unset, defaults to whatever uid/gid the session's
own sandboxed command is *already* running as (read from `/proc/<ns_pid>/status`),
rather than root/the caller — fixing a real bug (reported after this project's
own `-x/--exec` and priv-drop features had both shipped separately): without this,
`-x/--exec` always ran as whatever the *host* invocation was, ignoring any
`--user`/`--group` the session itself was started with. Either way, the resolved
identity is applied by running `command` through the session's already
bind-mounted `slocker-lite-priv-drop` helper (`priv_drop::path`, `bwrap.h`) —
reused as-is, not bind-mounted again (`-x/--exec` can't add bind mounts to an
already-running session's namespace anyway). Skipped entirely when the resolved
uid *and* gid are both 0: a session that was never given a resolvable user at
`-r/--run` time never got the helper bind-mounted at all, and dropping to 0:0
would be a no-op regardless; a missing helper for a genuinely non-root
resolution instead surfaces as `nsenter`'s own "No such file or directory" once
it tries to exec `priv_drop::path`, diagnostic enough on its own. **Second real
bug, caught by direct testing on a rootless dev machine before this shipped**:
when the session's own `-r/--run` used `--unshare-user` (i.e. ran rootless —
see the root-vs-rootless paragraph below), "root inside the container" is
achieved purely through the kernel's own uid mapping for that namespace, not a
real privilege drop — so `/proc/<ns_pid>/status`'s uid/gid, read from *outside*
that namespace, shows the host-mapped id (e.g. `1000`), not the
container-relative one (`0`). Treating that as "needs a priv-drop to 1000" is
wrong two ways: the helper is typically never bind-mounted for a session with
no resolved `--user`, and even when it is, `setuid()` fails outright under the
single-entry uid map an unprivileged user namespace gets (confirmed directly:
`failed to drop privileges to 0:0: Operation not permitted`). Fixed by tracking
whether the `user` namespace type was actually one of the ones joined (it only
is when it differs from this process's own, i.e. exactly when `-r/--run` used
`--unshare-user`) and, when so, leaving the default identity unresolved (no
priv-drop) for that case — joining that same user namespace with
`--preserve-credentials` (already done regardless) already reproduces the
container's own view correctly via that same kernel mapping, with nothing
further needed. Verified end-to-end on this same rootless dev machine: a
daemonized `-r --run` busybox session with no declared user, `--exec`'d with no
`--user`, now correctly shows `uid=0(root)` (previously would have attempted,
and failed, a priv-drop to the host-mapped uid); an explicit `--exec --user 0`
against the same session correctly resolves to `0:0` and skips the priv-drop
step; `--exec --user portage` against it correctly resolves the name to its
real `250:250` via the fetched `/etc/passwd` and then fails clearly (helper not
bind-mounted, since the session itself had no declared user) rather than
silently running as the wrong identity.
- `kill_session.{h,cpp}` — implements `--kill <pid>`, stopping a tracked,
running `-r/--run` session and everything it started. `kill_session()`
validates `pid` the same way `exec_in_session()` does (via `list_sessions()`,
`pid_file.h`). **Real bug reported by the user against their own actual
target device, confirmed via a captured session log**: a plain `kill
<tracked_bwrap_pid>` doesn't kill everything a container started — their
`/init` script `php-fpm --daemonize`s (double-forks, detaches) then `exec
caddy ...`s (replaces itself); after killing the tracked pid, both `caddy`
and the `php-fpm` master+workers kept running as orphans. **Root cause**: on
that device, `bwrap`'s `--unshare-pid` isn't actually in effect at all —
`detect_bwrap_unshare_args()` (`bwrap.cpp`) only requests `--unshare-xxx`
flags the kernel actually supports, and that kernel doesn't support pid
namespaces (independently confirmed elsewhere this session, see
`exec_session.{h,cpp}`'s own `CONFIG_CHECKPOINT_RESTORE` bug above) — so
`php-fpm --daemonize` reparents to the *host's own* pid 1, completely
disconnected from the sandboxed session; the classic "kill a pid namespace's
pid 1, the kernel guarantees the whole namespace collapses" trick simply
doesn't apply there. Per the user's own explicit request (they want to
choose the mechanism per host capability, and may need an even more basic
one later for some hypothetical older device), `kill_session()` picks
between three independently-named strategies, selected dynamically per
session (not a single cached host-wide capability flag, since e.g. cgroup
creation can fail for session-specific reasons like permissions even on a
host that generally supports cgroups) — each runs its own complete
`SIGTERM` → wait-up-to-`grace_period_seconds` (10s default, no CLI flag) →
forced-`SIGKILL` escalation internally, with no cross-strategy
fallback-after-failure chaining:
1. `kill_via_cgroup()` — preferred whenever the session has a non-empty
dedicated cgroup (`session_cgroup_pids()`, `session_cgroup.h`): `SIGTERM`
to every pid currently in it, and, if forcing is needed, either the
atomic `cgroup.kill` knob or a fresh re-read-and-`SIGKILL` sweep (fresh,
not the original snapshot, since a process could have forked a new child
after the graceful sweep but before dying). The **only** mechanism that
reliably reaches every process regardless of pid namespace support.
**Critical correctness point, caught during design review before this
shipped**: the "is it stopped yet" poll must gate on the *cgroup being
empty*, not `list_sessions()`'s running flag — that flag only reflects
the pid file's flock, released the moment the tracked outer `bwrap` pid
exits, and the `SIGTERM` sweep necessarily hits `bwrap` itself too (it's
a cgroup member) — `bwrap` dies and gets reaped in well under a second,
long before slower descendants (`caddy` shutting down gracefully,
`php-fpm` finishing in-flight requests) actually exit. Gating on the pid
file instead would make the poll resolve "done" almost immediately, the
forced-kill step would never run, and the original bug would reproduce
with unused machinery around it.
2. `kill_via_pid_namespace()` — used when no cgroup exists for the session,
but `resolve_namespace_pid()`/`pid_namespace_isolated()`
(`sandbox_process.h`) confirm `--unshare-pid` was genuinely in effect for
it. `SIGTERM`s `collect_descendant_pids(ns_pid)` (reliable here
specifically because reparenting within a genuinely isolated pid
namespace always lands back on that namespace's own pid 1); if forcing
is needed, a single `SIGKILL` to `ns_pid` alone is *guaranteed* complete
by the kernel itself, independent of whatever the graceful sweep missed.
"Stopped" is simply `kill(ns_pid, 0)` failing with `ESRCH`. Verified
end-to-end on this project's rootless dev machine (which does support
pid namespaces, unlike the user's real target device): a daemonized
busybox session running `sh -c 'sleep 300 & exec sleep 300'` (mirroring
the daemonize-then-exec shape of the original bug) was fully cleaned up
by `--kill`, including the backgrounded child, with no leftover
processes, mounts, or layers; a second run using `sh -c 'trap "" TERM;
sleep 300'` (ignoring `SIGTERM` entirely) confirmed the forced-`SIGKILL`
escalation path too, taking the full 10s grace period before the pid
namespace's own collapse-on-kill guarantee cleaned it up regardless.
3. `kill_via_tracked_pid()` — fallback when neither of the above applies:
signals the tracked `bwrap` pid directly, `SIGTERM` then `SIGKILL`,
polling `list_sessions()` for "stopped" since that's the only signal
available without a cgroup or an isolated pid namespace to check
directly. Exactly today's manual-`kill` behavior — least complete, but
always available, and strictly no worse than before this feature
existed. This is the path the user's own real target device actually
takes today (no cgroup delegation confirmed working there yet; no pid
namespace support at all) — a future, even more basic strategy (for some
hypothetical still-more-limited device) would slot in here the same way,
per the user's own explicit request to keep this extensible.
`poll_until()`/`sleep_ms()` (`.cpp`-local) use `nanosleep()` in an
`EINTR`-retry loop — matching this project's existing direct-POSIX style
(`process.cpp` already retries `waitpid()` the same way) — rather than
`<thread>`/`<chrono>` (unused anywhere else in this project).
- `config_file.{h,cpp}``load_config_file()` reads and parses (via libyaml's
document API, `<yaml.h>`) the `global`, `volumes`, and `networks` sections of
the local YAML config file located by `config_file_path()`
(`$XDG_CONFIG_HOME/slocker-lite/config.yaml`,
falling back to `$HOME/.config/slocker-lite/config.yaml`). Supported `global` keys:
`log-level`, and six `unshare-<type>` keys (`unshare-user`/`unshare-ipc`/
`unshare-pid`/`unshare-net`/`unshare-uts`/`unshare-cgroup`, one per
`bwrap.cpp`'s own `namespace_probes` entry) controlling whether `-r/--run`
requests each of bwrap's `--unshare-xxx` flags — every other long option is a
one-shot flag, not a setting, so it doesn't belong in a persistent config file.
Each `unshare-*` key accepts `"1"`/`"on"`/`"yes"`/`"true"` (enabled) or
`"0"`/`"off"`/`"no"`/`"false"` (disabled), case-insensitively (`.cpp`-local
`parse_bool_flag()`); an unset key defaults to enabled, and an unrecognized
value logs a `spdlog::warn` and is treated as unset (default enabled) rather
than failing the whole config load — consistent with this file's existing
forward-compatible/ignore-malformed-entries policy (only malformed *YAML
syntax* is a hard error). A missing file returns a default-constructed
(empty) `AppConfig`, not an error; unknown sections/keys (and malformed
individual volume entries) are likewise ignored for forward-compatibility.
`main()` applies `config->log_level` (via the existing `apply_log_level()`)
right after `spdlog::cfg::load_env_levels()` and before parsing CLI options,
so an explicit `--log-level` on the command line always overwrites it
afterward — same precedence pattern already used for `SPDLOG_LEVEL`.
`write_config_file()` writes the whole file back out (via libyaml's
document-building/emitter API, symmetric to the read side) — used by
`-v/--volume` (`create_volume_command()`, `commands.cpp`) to persist a new
`VolumeEntry {name, directory}` into the `volumes` section, preserving
`global` (including any set `unshare-*` keys, re-serialized as canonical
`"true"`/`"false"`) untouched. **`VolumeEntry`/the `volumes` section is a
distinct concept from `OciImageConfig::volumes`**: this is a user-defined
`name -> host directory` mapping created via `-v/--volume`, not an image's
own declared mount points (still unconsumed, see `oci_image.{h,cpp}` above).
`AppConfig`/`config.volumes` is looked up by name in `resolve_volume_mount()`
(`volume_mount.{h,cpp}`, see below), which is how `-r/--run`'s own `-v` usage
finds a named volume's host directory. `run_container()` (`commands.cpp`)
resolves the six `unshare-*` fields (each `value_or(true)`) into a
`NamespaceConfig` (`bwrap.h`, see below) once, up front, and passes it to
`run_bwrap()``bwrap.{h,cpp}` itself has no dependency on this file or on
YAML parsing at all, only on the already-resolved, defaults-applied struct.
**`networks` section** (see `docs/networking-design.md` for the full
feature): unlike `volumes` (a flat `name -> directory` scalar mapping), each
network entry is itself a *nested* mapping (`kind`/`subnet`/`ipv6`/
`subnet6`), since one network needs more than a single value to describe.
`NetworkEntry` (`kind` is `NetworkKind::extern_`/`intern` — trailing
underscore on `extern_` since `extern` is a reserved C++ keyword and can't
be an enumerator name — parsed from the YAML strings `"extern"`/`"intern"`)
round-trips through `AppConfig::networks` the same way `VolumeEntry` does;
an entry with an unrecognized `kind` (or missing `kind`/`subnet`) is skipped
on load, same forward-compatible policy as everything else here. `ipv6`
reuses `parse_bool_flag()`, defaulting to `true` (enabled) if absent or
unparseable; `subnet6` is only read/written when `ipv6` is true.
`write_config_file()` writes each network as its own nested mapping under
`networks`, `ipv6` re-serialized as canonical `"true"`/`"false"` like the
`unshare-*` keys. `veth` (default `true`) round-trips the same way as
`ipv6` (`parse_bool_flag()`, written as canonical `"true"`/`"false"`,
always written regardless of value — unlike `subnet6`, there's no
companion field whose presence depends on it) — see `network_bridge.h`'s
`probe_veth_support()`/`should_use_veth()` above and `--no-veth`
(`cli_args.{h,cpp}`) below for what it controls.
- `network_subnet.{h,cpp}` — pure CIDR arithmetic backing `-n/--network`'s
subnet allocation and `network_bridge.{h,cpp}`'s (see below) gateway-address
computation; no kernel/`ip`/`iptables` calls of its own. `is_valid_network_name()`
(non-empty, no `':'`) mirrors `volume_mount.h`'s `is_valid_volume_name()`
(which rejects `'/'`) — `':'` specifically because `port_forward.h`'s `-p`
syntax splits a spec on it; a network name containing one would make that
parse ambiguous. `is_valid_ipv4_cidr()`/
`is_valid_ipv6_cidr()` and `ipv4_cidrs_overlap()`/`ipv6_cidrs_overlap()` all
build on one `.cpp`-local `parse_cidr()` (via `inet_pton()`, not hand-rolled
parsing) producing a plain byte-vector address (4 bytes for IPv4, 16 for
IPv6) + prefix length, and one shared `bytes_overlap()` byte/bit-mask
comparison generic over that byte length — IPv4 and IPv6 overlap checking
are the same algorithm, not two parallel implementations.
`allocate_ipv4_subnet()`/`allocate_ipv6_subnet()` (`commands.cpp`'s
`create_network_command()`) scan `10.168.<n>.0/24`/
`fdf0:f243:f06f:<168+n>::/64` for `n` in `0..255` and return the first one
that doesn't overlap *any* existing network's subnet (via the overlap
checks above, not just other auto-allocated ones — a manually
`--subnet`-overridden network is checked too). `fdf0:f243:f06f::/48` is a
randomly generated ULA (RFC 4193) — **replaced the original
`fd00:168:0::/48`, which was never actually randomly generated, just a
memorable placeholder, at the user's own request** ("since this is an
ULA, let's use a randomly generated prefix"); the `168` offset on the
subnet-id hextet (confirmed with the user via `AskUserQuestion`, over the
alternative of dropping it and starting both at `0`) keeps the same
project-recognizable stamp the old scheme's fixed 2nd hextet had, just as
a constant offset now rather than a numerically-identical index — the
same `n` range still drives both v4 and v6 allocation, so the common case
(no manual overrides) still allocates deterministically paired blocks per
network, just offset by `168` on the v6 side instead of matching exactly.
Since IPv6 hextets are hexadecimal, `n + 168 >= 10` (i.e. always, given
the offset) renders as a valid but numerically-different-from-`n + 168`
address when read back as hex (e.g. `n=15``183` decimal → renders as
`...:183::/64`, which is hex `0x183`, not `183`) — purely cosmetic,
allocation correctness doesn't depend on this matching numerically at
all. `ipv4_gateway_address()`/
`ipv6_gateway_address()` (`network_bridge.cpp`'s `provision_bridge()`) and
`ipv4_host_address()`/`ipv6_host_address()` (`network_join.cpp`'s
`pick_free_address()`, see below — `n = 2, 3, ...` for individual
containers) are all thin wrappers around one shared `.cpp`-local
`host_address(af, cidr, n)`: masks `cidr` down to its network address first
(`mask_to_network()`, in case it — e.g. a manual `--subnet`/`--subnet6`
wasn't already a canonical network address), then adds `n` as a big-endian
integer into the trailing host-portion bytes with proper carry propagation
(generic over address length, so the same code handles both IPv4's 4 bytes
and IPv6's 16 without two parallel implementations), rejecting `n` outright
if it doesn't fit the address's host-bit width. The gateway functions are
just `host_address(af, cidr, 1)` — the `.1` convention this project's
bridges use. Reuses the same `parse_cidr()` as the validation/overlap
functions above.
- `volume_mount.{h,cpp}``is_valid_volume_name()` (no `/`, checked by both
`create_volume_command()` and to tell a `-v` spec's name/path apart) and
`resolve_volume_mount()`, called once per `-v` occurrence from `run_container()`
when running with `-r`. A spec with no `/` is looked up in `config.volumes` by
name (error if unknown); one with `/` is treated as a host directory path and
`create_directories()`'d if missing. If the resulting host directory is empty,
`initialize_volume_directory()` reconciles it against the image's own directory
at the given container path: if that image directory is non-empty, its contents
are copied in first; then, **whether or not there was content to copy**, the
host directory's own mode/ownership/timestamps (and xattrs/ACLs where
supported) are always set to match the image directory's own — **real bug
fixed by the user, not assumed**: an earlier version only ever copied when the
image directory was non-empty, so an image declaring an *empty* directory with
specific ownership/permissions (e.g. a data directory owned by a non-root
uid/gid) got a host directory with default `create_directories()` permissions
instead, and even the non-empty case never reconciled the directory's *own*
attributes (only each copied entry's). The existence check, content copy, and
attribute reconciliation all run as a single `sh -c` invocation wrapped through
`wrap_for_root_namespace()` (`bwrap.h`) — **not** a plain `std::filesystem`
check — because a rootless `containers-storage mount`'s content isn't visible
to this process at all without `nsenter`, the same constraint `run_bwrap()`
itself works around (see the root-vs-rootless paragraph below). `cp -a
--preserve=mode,ownership,timestamps,links[,xattr] --attributes-only -T` does
the attribute-reconciliation step; whether `,xattr` is included is decided by
a direct `setxattr()`/`removexattr()` probe on the host directory (no new
library dependency — Linux POSIX ACLs are themselves stored as xattrs, so this
one probe stands in for both, logging a single `spdlog::warn` if unsupported).
`-T`/`--no-target-directory` is required on that second `cp` — confirmed by
direct testing: without it, since the host directory already exists, plain
`cp SRC DST` copies `SRC` *into* `DST` as a nested `DST/basename(SRC)`
subdirectory instead of reconciling `DST`'s own attributes, which is exactly
the bug this fix closes. A nonzero `cp` exit is only ever a warning, never
fatal — often just an ownership-preservation shortfall when not running as
root. Verified end-to-end against `images/gitea.tar`'s real declared
`/etc/gitea`/`/var/lib/gitea` volumes under a real rootless mount: the
resulting host directories' mode/ownership matched the image's own declared
values in both cases, and no mounts/layers were left behind afterward.
Errors are logged via `spdlog::error`; every external command is also traced at debug
level in `run_process()`/`run_process_foreground()` (`src/process.cpp`) — visible via
`SPDLOG_LEVEL=debug`, since spdlog's default level is `info` — and a failed external
command additionally logs a `spdlog::warn`, which is visible by default (no env var
needed). The final "mounted image at: ..." success line is direct stdout program
output, not a log.
Because `containers-storage mount` runs rootless, it reexecs itself into a private
user+mount namespace to gain the privilege it needs for the overlay mount — which
leaves the result invisible to a plain shell or child process outside that namespace.
Confirmed `containers-storage unshare` does **not** rejoin an already-running mount's
namespace; only `nsenter` targeting the live `fuse-overlayfs` daemon's PID does.
`-r/--run` handles this automatically by locating that PID and running `bwrap` via
`nsenter` into its namespaces (`wrap_for_root_namespace()`, `src/bwrap.h`) — reused
as-is by `volume_mount.cpp`'s copy-into-an-empty-volume step, since that also needs
to read image content that's otherwise invisible outside the same namespace.
**Running as root sidesteps all of this**: no privilege
reexec is needed, so the mount is already directly visible in the current namespace,
and `nsenter --user=...` into it then fails ("reassociate to namespace 'ns/user'
failed: Invalid argument") since the caller is already in that same user namespace.
`-r/--run` detects `geteuid() == 0` and skips `nsenter` automatically in that case;
`--no-nsenter` forces it off manually for any other situation where the mount turns
out to already be directly visible.
**Mutable global state and multi-container support:** an audit ahead of planned
docker-compose support (running multiple containers at once) found exactly three
pieces of mutable global/file-scope state in `src/`: `g_mount_program`
(`containers_storage.{h,cpp}`, the resolved `fuse-overlayfs` path — genuinely
process-wide, invariant across containers), `g_foreground_child_pid`
(`process.cpp`, plus `run_process_foreground()`'s process-wide `SIGINT`/`SIGTERM`
handler installation — tracks one foreground child at a time), and
`g_report_fd`/`g_log_path` (`daemonize.cpp`, one in-flight `-D/--daemonize`
handshake's report-pipe fd and log path). **Decision, confirmed by the user:**
multi-container/compose support will run each container's session in its own
forked OS process — the same model `-D/--daemonize` already uses — rather than
one process managing multiple containers concurrently without forking. Under
that model, "one OS process" and "one running container" stay the same thing
they already are today, so **none of these globals need to become per-container
state** — each forked child only ever tracks/signals one foreground child and
handles one daemonize handshake, exactly as today. This is a load-bearing
constraint for however the compose orchestrator ends up implemented: it must
fork (not thread, not run an in-process event loop over N containers) one child
per service, each child reusing `run_container()`'s existing single-container
code path unchanged.
## Build & test commands
Build directory is `buildDir/` (already configured).
- Configure (only needed if `buildDir/` is missing or deleted): `meson setup buildDir`
- Build: `meson compile -C buildDir` (or `ninja -C buildDir`) — also builds
`buildDir/slocker-lite-priv-drop`, the statically-linked helper `-r --user` needs
(see `priv_drop_helper.cpp` in "Project state")
- Run the executable: `./buildDir/slocker-lite -m <image.tar>` (see `--help` for the
full flag list: `-m/--mount`, `-r/--run`, `-u/--umount`, `-c/--cleanup`,
`-l/--list-images`, `-i/--inspect`, `-x/--exec`, `--kill`, `--no-nsenter`, `-D/--daemonize`,
`--user`, `--group`, `--hostname`, `--env`, `--env-file`, `-v/--volume`, `--list-volumes`, `--delete-volume`,
`--delete-volume-full`, `-n/--network`, `--extern`, `--intern`, `--subnet`, `--no-ipv6`, `--subnet6`,
`--no-veth`, `--list-networks`, `--delete-network`, `-p/--port-forward`, `--list-processes`, `--clean-processes`,
`-w/--write-config`, `-t/--test`, `--log-level`, `-h/--help`, `-V/--version`)
- Run tests: `meson test -C buildDir`
## Code style
- Null-pointer checks: prefer `if (!ptr)` / `if (ptr)` over `if (ptr == nullptr)` / `if (ptr != nullptr)`.
- Constants: no `k` Hungarian-notation prefix. `enum class` values are already qualified by
the enum's own name (e.g. `Mode::run`, `OciPortProtocol::tcp`), so plain snake_case
enumerators are enough on their own. Free-standing constants also use plain snake_case;
when several are conceptually related, group them under a named `namespace` instead of
relying on a shared prefix to imply the grouping (e.g. `cli_args.cpp`'s `getopt_long` long-option
codes live in `namespace options { constexpr int log_level = ...; }`, and `bwrap.cpp`'s
priv-drop-helper path/binary-name pair live in `namespace priv_drop { ... }`) — nest the named
namespace inside the file's existing anonymous namespace where one is already present, so
internal linkage is unchanged. A `kXxx`-named identifier that turns out not to actually be
`const` (mutable global/static state) instead follows this codebase's existing `g_` prefix
convention (e.g. `containers_storage.cpp`'s `g_mount_program`, matching `process.cpp`'s
`g_foreground_child_pid` and `daemonize.cpp`'s `g_report_fd`/`g_log_path`).
## Licensing
- Every `.c`/`.cpp`/`.h` file under `src/` must start with the GPLv2-or-later copyright
header (see any existing file under `src/` for the exact text).
- After adding a new source file under `src/`, run `./add-license.sh` from the repo
root to prepend the header (it reads `copyright-header` and inserts it via `sed`,
skipping files that already have it, so it's safe to re-run at any time).
## Build configuration notes
- `meson.build` sets `warning_level=3` and `cpp_std=c++20` — keep new code warning-clean under `-Wall -Wextra -Wpedantic`-equivalent settings.
- The single Meson `test()` target runs `slocker-lite` against a fixture OCI image tar generated at build time by `tests/gen_fixture.py` (a `custom_target`) and checks its exit code (no test framework is wired in yet).