# CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. ## Project state `slocker-lite` (C++20, built with Meson) mounts an OCI Image Layout tar (`oci-layout` + `index.json` + `blobs/sha256/*`, as produced by `skopeo`/`podman save --format oci-archive`/modern `docker save`) using `containers-storage` and `fuse-overlayfs`, then (via `-r/--run`) runs a sandboxed command against it with `bwrap`. The real deployment target is Android with a stock kernel, where `podman`/`docker` don't run (missing namespace support) and there's no kernel overlayfs (hence `fuse-overlayfs`); `bwrap` is invoked in "degraded mode" using only whichever `--unshare-xxx` namespaces the running kernel actually supports. See `README.md` for the human-facing overview (build/usage/ status); this file stays the dense, file-by-file reference. Still early-stage. Source layout (all under `src/`): - `main.cpp` — the CLI entry point only, and deliberately tiny (~30 lines): loads the config file, applies its `global.log-level` (`apply_log_level()`, `cli_args.h`) before CLI parsing so an explicit `--log-level` can still override it afterward, calls `parse_args()` (`cli_args.h`) and returns its exit code immediately if it gives one (covers `-h`/`-V` and every parse error), otherwise calls `dispatch_command()` (`commands.h`) and returns its result. All of the actual option-parsing and command logic that used to live here moved out into `cli_args.{h,cpp}`/`commands.{h,cpp}`/`self_test.{h,cpp}` (see below) specifically to keep this file from re-growing into a dumping ground as more commands (docker-compose support, etc.) get added. - `cli_args.{h,cpp}` — command-line parsing only, nothing else. `Mode` (the enum of every CLI action) and `ParsedArgs` (everything `parse_args()` extracts from `argv`) live in the header since `commands.h`'s `dispatch_command()` consumes them; `print_usage()`/`print_version()`, the `options` namespace of getopt long-option codes, and the `long_options` array itself are `.cpp`-local. `parse_args(argc, argv, out)` runs the `getopt_long` loop plus all of the post-loop validation that used to live at the top of `main()`: mode/`--volume` interaction (`-v` alone vs. combined with `-r`, see below), `--group` requires `--user`, no leftover positional args outside `-r`/`-e`, and (moved here from what used to be inline in the `Mode::exec` dispatch arm) `-x/--exec `'s own pid parsing/validation (`ParsedArgs::exec_pid`, a positive integer or a hard error) and its trailing-command requirement (`ParsedArgs::command`, required non-empty). `--kill ` (`ParsedArgs::kill_pid`) shares that same positive-integer parsing/validation via a small extracted `parse_pid_arg()` helper (`.cpp`-local) rather than duplicating the `strtol` dance a second time — unlike `-x/--exec`, it takes no trailing command, so it's simply not added to the leftover-args exemption list (`Mode::run`/`Mode::exec` only). Returns an exit code `main()` should return immediately (`0` for `-h`/`-V`, `1` for any parse error) when set; `nullopt` means `out` is ready for `dispatch_command()`. `-D/--daemonize` (has a short form; `'D'` was free) is a plain boolean flag (`ParsedArgs::daemonize_flag`, set in its own `case 'D':`, same pattern as `--no-nsenter`). `--hostname `/`--env VAR=VALUE`/`--env-file ` (all long-option only, `--env`/`--env-file` both repeatable) are collected here into `ParsedArgs::hostname_flag`/ `env_specs` — `--env` pushes `{false, optarg}`, `--env-file` pushes `{true, optarg}` into the *same* ordered `std::vector` (not two separate lists), preserving their exact relative command-line order across both flags, since `resolve_env_specs()` (`env_spec.{h,cpp}`, see below) needs that order to let a later one override an earlier one for the same variable name — actually resolving them happens later, in `commands.cpp`'s `run_container()`. `-v/--volume` is dual-purpose: used alone it's a standalone `Mode::volume` request; combined with `-r/--run` it's repeatable and requests a volume mount instead (resolved later by `resolve_volume_mount()`, `volume_mount.{h,cpp}`, see below). Since `-v` must be repeatable with `-r` but each occurrence still takes two space-separated tokens, the getopt loop doesn't let `'v'` set `Mode` directly: it accumulates `(spec, path)` pairs into `ParsedArgs::volume_specs` (consuming the second token manually, with a guard against swallowing the next flag if there isn't one), and only *after* the loop decides whether that means one standalone `Mode::volume` call or, together with `-r`, passes `volume_specs` through unresolved for `dispatch_command()`/ `run_container()` to handle. `-n/--network` (see `docs/networking-design.md` for the full feature design) is dual-purpose the same way, but simpler: since a network join has no equivalent of a volume's container-mount-path second argument, each occurrence is a single `required_argument` token (just the name) accumulated into `ParsedArgs::network_specs` — no manual second-token consumption needed, `'n'`'s own `case` just does `network_specs.push_back(optarg)`. The same post-loop split as `-v` decides `Mode::network` (standalone, exactly one occurrence) vs. join-with-`-r` (repeatable, no limit). `--extern`/`--intern`/`--subnet `/`--no-ipv6`/ `--subnet6 `/`--no-veth` (`ParsedArgs::network_extern_flag`/`network_intern_flag`/ `network_subnet_flag`/`network_no_ipv6_flag`/`network_subnet6_flag`/ `network_no_veth_flag`) only apply to the standalone (create) case and are rejected with a clear error if given any other way (e.g. alongside `-r`) — `Mode::network` additionally requires exactly one of `--extern`/`--intern`. `--no-veth` forces `NetworkEntry::veth` (`config_file.h`) to `false` at creation time, overriding the default `true` — see `network_bridge.h`'s `probe_veth_support()`/`should_use_veth()` for what this controls: lets the tap+relay fallback (the real target device's kernel lacks `CONFIG_VETH` — see `docs/networking-design.md`'s addendum) be exercised on a veth-capable machine like this dev box, without needing the actual veth-less hardware. **`-n` used to belong to `--no-nsenter`**: reassigned here since `--network` will be far more heavily used; `--no-nsenter` moved to long-option-only (`options::no_nsenter`) rather than hunting for a new letter, matching `--kill`'s own "rare/niche flag, long-only is no real loss" precedent. `-p/--port-forward` (`'p'` was free) is repeatable the same accumulate-now, resolve-after-the-loop way as `-n` (`ParsedArgs::port_forward_specs`, raw `"[:]: "` strings — actual parsing happens later, in `port_forward.h`, since resolving which network a spec refers to needs runtime join state that doesn't exist yet at parse time), but has no standalone use at all: rejected post-loop unless combined with `-r`. - `commands.{h,cpp}` — every command's implementation, plus the dispatcher. `dispatch_command(args, config_path, config)` (the only externally-linked function; everything else in this file is `.cpp`-local) is a `switch (args.mode)` with **one explicit `case` per `Mode` enumerator and no `default:`**, so `-Wswitch` (this project builds at `warning_level=3`) forces a compile warning/error if a future `Mode` value is ever added without a matching dispatch case, instead of silently falling through to the wrong command — confirmed by testing (temporarily adding an unhandled enumerator triggered exactly the expected `-Wswitch` warning). `Mode::mount` has its own explicit case (`mount_command()`) for the same reason: it used to be handled only by *falling off the end* of a long `if`/`else` chain in `main()` with no explicit check at all — the very kind of implicit, easy-to-silently-break behavior this dispatcher redesign exists to close off, especially with more `Mode` values (docker-compose support) expected soon. `list_processes_command()` implements `--list-processes` (long-option only): calls `list_sessions()` (`pid_file.{h,cpp}`, see below) and prints one tab-aligned `pid`, `container name`, `running`/`exited` row per entry (same two-column tab-alignment scheme as `list_images_command()`/ `list_volumes_command()`, extended to a third column), no header row, silent success on an empty list. `clean_processes_command()` implements `--clean-processes` (also long-option only): calls `clean_stale_sessions()` (`pid_file.{h,cpp}`) and prints one `removed stale pid file for '' (pid )` line per file actually removed, then also calls `clean_stale_port_forwards()` (`port_forward.h`, see below — commit 6 of `docs/networking-design.md`'s sequence) and prints one `removed stale port-forward rules for '-'` line per record actually swept, then `clean_stale_tap_relays()` (`network_tap_relay.h`, see below — the direct tap+relay analog of the port-forward sweep) and prints one `removed stale tap-relay processes for '-'` line the same way — nothing is printed for sessions still running, and an empty result (nothing stale) is silent success, same convention as the rest of this file's list/delete commands. `Mode::exec`'s dispatch case is a one-line call to `exec_in_session(*args.exec_pid, args.command)` (`exec_session.{h,cpp}`, see below) — the pid/command parsing and validation now happens in `cli_args.cpp`'s `parse_args()` instead (see above). `inspect_image_command()` implements `-i/--inspect `: prints every `OciImageConfig` field (user/group, exposed ports, env, volumes, default command) without mounting or running the image — extend it whenever `OciImageConfig` gains a new field (see `oci_image.{h,cpp}` below). `run_container()` unconditionally calls `read_oci_image_config()` and reuses the result for two independent defaults: the command to run (`Entrypoint ++ Cmd`) when none is given on the command line, and, when `--user` wasn't given, the sandboxed process's user/group (`config.User`, split into `OciImageConfig::user`/`group`) — an explicit `--user`/`--group` on the command line always takes precedence. `create_volume_command()` implements `-v/--volume `; `list_volumes_command()` implements `--list-volumes` (same tab-alignment scheme as `list_images_command()`, reused as-is); `delete_volume_command()` implements both `--delete-volume ` (config entry only) and `--delete-volume-full ` (also `std::filesystem::remove_all()`s the host directory — errors out before touching the config if that fails, warns instead of failing if the directory was already gone) — see `config_file.{h,cpp}` below for what a "volume" means here (a distinct concept from `OciImageConfig::volumes`). `dispatch_command()`'s `Mode::volume`/`Mode::delete_volume`/`Mode::delete_volume_full` cases call these. `create_network_command()`/`list_networks_command()`/`delete_network_command()` are the direct network equivalents (`Mode::network`/`Mode::list_networks`/ `Mode::delete_network`) — see `docs/networking-design.md` for the full feature design and `config_file.{h,cpp}` below for `NetworkEntry`. Joining a network from `-r/--run` (repeatable `-n `, `ParsedArgs::network_specs`, see `cli_args.{h,cpp}` above) is handled by `run_container()`, further below, via `network_join.{h,cpp}` (see below). `create_network_command()` rejects a name containing `':'` first (`is_valid_network_name()`, `network_subnet.h` — needed since `port_forward.h`'s `-p` syntax splits a spec on `':'`; a network name containing one would make that parse ambiguous), then a duplicate name, then resolves `subnet`/`subnet6`: an explicit `--subnet`/`--subnet6` is validated (`is_valid_ipv4_cidr()`/`is_valid_ipv6_cidr()`) and checked for overlap against every existing network's subnet (`ipv4_cidrs_overlap()`/ `ipv6_cidrs_overlap()`, `network_subnet.{h,cpp}`, see below); otherwise the next free block is auto-allocated (`allocate_ipv4_subnet()`/ `allocate_ipv6_subnet()`). Once a `subnet`/`subnet6` is resolved, `create_network_command()` calls `ensure_network_provisioned()` (`network_bridge.h`, see below) to actually stand up the network's host-side state (bridge, sysctls, iptables rules for `extern`; a dedicated persistent namespace + bridge for `intern`) — only once that succeeds is the entry appended to `config.networks` and persisted; a network that fails to provision isn't saved. `list_networks_command()` reuses the same independently-per-column tab-alignment scheme as `list_processes_command()` (name/kind/subnet/bridge each aligned — the bridge name recomputed via `bridge_name(network.name)` — `network_bridge.h` — rather than stored, since it's already a pure deterministic function of the name — then the IPv6 subnet — or `"(no ipv6)"` — appended unaligned as the trailing column, nothing follows it). `delete_network_command()` takes a `delete_full` bool, same shape as `delete_volume_command()`'s own: `--delete-network` (`false`) only removes the config entry, leaving the network's live host-side state untouched; `--delete-network-full` (`true`) additionally calls `teardown_network_state()` (`network_bridge.h`, see below) first. Unlike `delete_volume_command()`'s `-full` variant (which bails out *before* touching the config if its single `remove_all()` call fails), `delete_network_command()` doesn't gate the config removal on `teardown_network_state()`'s success at all — that function is deliberately best-effort/non-fatal per-step (see its own doc comment), so a step "failing" because that piece was already gone by hand (exactly `ensure_network_provisioned()`'s own existence-check caveat, above) is expected, not a reason to leave a network the user explicitly asked to delete sitting in the config. `write_config_command()` implements `-w/--write-config`: unlike `create_volume_command()`/`delete_volume_command()`'s use of `write_config_file()` (which only ever persists `AppConfig` fields that are already set), this fills in *every* field before writing — the six `unshare-*` bools via `.value_or(true)`, and `log_level` from the actually active `spdlog::get_level()` (not merely a default for when unset — this also captures an explicit `--log-level` passed alongside `-w` on the same command line, overriding whatever an existing config file's own `log-level` already was, since `main()`/`parse_args()` already applied it in that precedence order by the time this runs) — so a bare `-w` bootstraps a complete, fully-populated config file for hand-editing, and `-w` combined with other flags captures their effective values into it. `volumes` is left exactly as loaded — an open-ended list with no "default" entry to materialize. Prints the config file's full path (`write_config_file()` already creates the parent directory and the file itself if missing, so no separate existence check is needed here). `run_container()` (the `Mode::run` dispatch case) resolves each `-v` spec (erroring out, `ok = false`, same as a failed `--user` resolution — `bwrap` is skipped but unmount/cleanup still runs) into a `ResolvedVolumeMount`, rejecting a duplicate or non-absolute container path first, and passes the resolved list to `run_bwrap()`. `--hostname ` is likewise threaded straight through `run_container()` into `run_bwrap()`/`build_bwrap_args()` (`bwrap.{h,cpp}`) — see there for how/when it actually takes effect. `run_container()` also derives a `container_name` for the session-tracking pid file (see `pid_file.{h,cpp}` below): `read_image_ref()` (`oci_image.{h,cpp}`) applied to the single image tar being run, formatted as `name:tag`, falling back to the tar's own filename stem if `read_image_ref()` can't determine one — passed through to `run_bwrap()` alongside everything else. `--env`/`--env-file`'s already-ordered `env_specs` (see `cli_args.{h,cpp}` above) are resolved here, once, via `resolve_env_specs()` (`env_spec.{h,cpp}`, see below) — same `ok = false`-on-failure pattern as volume/user resolution — and the resolved list is passed to `run_bwrap()` as `extra_env`. `-D/--daemonize`'s `daemonize_flag` is also consumed here: `run_container()` computes `container_name` *before* `mount_image()` (needed so `daemonize()` below can use the real container name for the log file from its very first line, not just after a later rename) and, if daemonizing, calls `daemonize(container_name)` (`daemonize.{h,cpp}`, see below) immediately after: a returned value means this is the original (parent) process (or a hard daemonize failure) — print it and `return` right away; `nullopt` means this is the now-detached child, which falls through into the rest of `run_container()`'s existing body completely unchanged, including the unmount/cleanup that already runs after `run_bwrap()` returns (no separate watcher/reaper — the daemonized child *is* what runs the whole session, start to finish). `network_specs` (repeatable `-n`, `cli_args.h`) is validated up front, before `run_bwrap()` is ever called: joining a network needs a real, isolated network namespace to attach a veth into (unlike `--hostname`, this can't just be skipped/degraded when unavailable), so if any networks were requested, `run_container()` checks both that `namespace_config.net` is actually enabled (`global.unshare-net`, `config_file.h`) and that `detect_bwrap_unshare_args()` (`bwrap.h`) reports the running kernel actually supports `--unshare-net` — either failing sets `ok = false` with a clear error, the same pattern as a failed `--user` resolution. If validation passed, `on_bwrap_pid_known` (already used for `-D/--daemonize`'s `report_daemon_started()`) additionally calls `join_networks(pid, network_specs, app_config)` (`network_join.h`, see below) — networks are joined *before* the daemonize report is sent, so a `-D`-daemonized caller doesn't get control back until network setup has already had its chance to run. `join_networks()` itself early-returns (no namespace wait at all) when `network_specs` is empty, so calling it unconditionally whenever `on_bwrap_pid_known` fires for *any* reason (e.g. `-D/--daemonize` alone, no `-n`) doesn't cost anything. `-p`'s `port_forward_specs` (`cli_args.h`) are syntax/range-parsed (`parse_port_forward_spec()`, `port_forward.h`) up front too — `ok = false` on a bad spec, same as other validation failures — but *resolving* which network each targets can only happen after `join_networks()` returns (it needs to know which networks actually joined, and their assigned IPs), so that happens in the same `on_bwrap_pid_known` callback, right after the `join_networks()` call: `add_port_forward()` per spec, collecting the ones that actually landed into a `std::vector` declared in `run_container()`'s own scope (captured by reference) — read again *after* `run_bwrap()` returns to `remove_port_forward()` each one. This two-places split (add during the callback, remove after `run_bwrap()` returns) mirrors how `join_networks()`'s own veths don't need an explicit removal step (the kernel tears them down once the session's namespace goes away) while port-forward rules — host-global, named, persistent iptables state — very much do. `on_bwrap_pid_known` itself is only set at all when `daemonize_flag || !network_specs.empty() || !parsed_port_forwards.empty()` — a real gap caught while wiring this up: an earlier version only checked the first two, so `-p` given *without* `-n` or `-D` would silently never even attempt to run (no error, nothing logged) since the callback that resolves/applies it would never fire at all. At the end of the callback, `record_port_forwards(container_name, pid, active_port_forwards)` (`port_forward.h`, see below) persists whatever actually landed to a small state file — so that a later `--clean-processes` run can find and remove these rules even if *this* process crashes before ever reaching its own `remove_port_forward()` calls after `run_bwrap()` returns; those calls are paired with a `remove_port_forward_record(container_name, bwrap_pid)` (`bwrap_pid` captured from the same callback, in a variable declared in `run_container()`'s own scope) so a cleanly-exiting session's own record doesn't linger for `--clean-processes` to find later. The same callback also collects each `JoinedNetwork::relay` (`network_join.h`) returned by `join_networks()` into a `std::vector active_relays` declared in `run_container()`'s own scope — the direct tap+relay (`network_tap_relay.h`) analog of `active_port_forwards` above, since a relay process is likewise independent host-global state (unlike a veth pair) that needs an explicit `stop_tap_relay()` call for each, made right alongside the `remove_port_forward()` loop after `run_bwrap()` returns — paired the same two-places way with `record_tap_relays(container_name, pid, active_relays)` (called right after `record_port_forwards()`, same callback) and `remove_tap_relay_record(container_name, bwrap_pid)` (called right after the `stop_tap_relay()` loop) for `--clean-processes`'s own crash-orphan sweep (`network_tap_relay.h`'s `clean_stale_tap_relays()`, see below). - `self_test.{h,cpp}` — `run_self_tests()` implements `-t/--test`, this project's own built-in self-test mode (distinct from the Meson-driven fixture smoke test under `tests/`, described in "Build & test commands" below; previously reported `detect_bwrap_unshare_args()`'s output — `bwrap.{h,cpp}` — unplugged since that's kernel-capability diagnostics, not a test). Currently exercises `persistent_netns.{h,cpp}`'s (see below) create/verify/remove cycle: skipped with a message (not a failure) when not root, since `create_persistent_netns()` requires it for the bind mount. Confirms the namespace is missing before creation, exists right after (checked from this process, *after* the forked child that actually did the `unshare()`/bind-mount has already exited — the actual claim being tested: the namespace outlives its creating process), then gone again after removal. Also exercises `network_tap_relay.{h,cpp}`'s (see below) create/attach/teardown cycle, same root-only skip: a throwaway bridge (`ip link add ... type bridge` inside a throwaway persistent namespace of its own — `create_tap_relay()` now always enters a network's persistent namespace first, both kinds, see `network_bridge.{h,cpp}`'s own "Resolved: `extern` had no connectivity" entry, so this test needs one too even though `network_bridge.h`'s own bridge *provisioning* logic still isn't otherwise exercised here — a tap device only cares that *some* bridge exists to attach to) stands in for a real network's bridge, and a throwaway network namespace (a forked child that `unshare(CLONE_NEWNET)`s then blocks in `pause()` until signaled) stands in for a real `-r/--run` session's isolated one. **Real race caught by testing, not assumed**: `fork()` returning to the parent doesn't mean the child has actually reached its own `unshare(CLONE_NEWNET)` call yet — the same race `network_join.cpp`'s own `wait_for_isolated_net_namespace()` already guards against for a real session; an earlier version of this test used the child's pid immediately and the container-side tap device silently ended up created in the *host's* namespace instead (confirmed: it wasn't visible via `nsenter` into the child's namespace at all). Fixed the same way — polling (bounded, 1s, 20ms interval) `namespace_isolated()` (`sandbox_process.h`) until the child's own net namespace actually differs from this process's before trusting its pid. Confirms the host-side tap gets created and attached to the bridge (`ip link show` output contains `master `), the container-side tap gets created with the requested name inside the target namespace (checked via a plain, single-shot `nsenter --net=/proc//ns/net -- ip link show`; an earlier version wrapped this in a bounded retry, added when tap devices were still created via bare `ioctl(TUNSETIFF)` and intermittently weren't immediately visible — see `network_tap_relay.{h,cpp}`'s own entry below for why that retry, and the whole class of symptom it was compensating for, is gone now that devices are created persistently instead), and — **updated once the original "both devices disappear on their own once `stop_tap_relay()` stops the relay, no explicit `ip link del` needed" assumption turned out to be wrong on the real target device** (`network_tap_relay.{h,cpp}`'s own entry below has the full story) — now confirms the *host*-side device is explicitly gone after `stop_tap_relay()` (the new `ip link del` step) while the *container*-side device deliberately still exists (correctly persistent — only the relay stopped, not the container's own network namespace); the container-side device's actual disappearance, once that namespace itself is torn down, isn't separately re-checked (`nsenter` has nothing left to target once the namespace's only holding process has already exited) — it's destroyed moments later anyway, at the very end of this test, when its throwaway namespace-holder process is killed. Deliberately its own small file since more real tests are expected here as more of the networking feature lands. - `env_spec.{h,cpp}` — `resolve_env_specs()` turns an ordered list of `EnvSpec {is_file, value}` (see `cli_args.{h,cpp}` above) into a flat, ordered list of `(key, value)` pairs. A literal (`--env`) is split at its *first* `=` (the value may itself contain `=`; the key must be non-empty). A file (`--env-file`) is read line by line: blank/whitespace-only lines and lines whose first non-whitespace character is `#` are skipped (comments), with a trailing `\r` stripped first for CRLF files; every other line is parsed the same way as a literal. Logs a specific error and returns `nullopt` on the first hard failure (malformed line, empty key, or an unreadable file) — deliberately stops at the first line, not "skip and warn", since an env file with a typo should fail loudly rather than silently omit a variable a container might depend on. `build_sandbox_env()` (`bwrap.cpp`, see below) appends the resolved list after its own built-in `PATH`/`HOME`/`PWD`/`TERM` — no deduplication needed there, since `run_process_foreground()`'s own `setenv(..., 1)` loop already lets the later occurrence in iteration order win for a repeated key, so an explicit `--env PATH=...` still overrides the default. - `oci_image.{h,cpp}` — validates/parses the OCI Image Layout tar (libarchive + nlohmann_json) and extracts layer blobs. `list_oci_images()` scans a directory (non-recursively) for `*.tar`/`*.tar.*` files and, for each valid OCI archive, derives an image name/tag via `read_image_ref()` from its `index.json` manifest annotations (`io.containerd.image.name` preferred, else `org.opencontainers.image.ref.name`), falling back to the archive's filename and `"latest"` respectively. `read_image_ref()` is public (not just an internal helper of `list_oci_images()`) precisely so `run_container()` (`commands.cpp`) can reuse the exact same logic to name a *single* image tar's session pid file (see `pid_file.{h,cpp}` below) instead of duplicating it. `read_oci_image_config()` reads the image config blob referenced by the manifest and extracts `User` (split on `:` into `OciImageConfig::user`/`group`), `ExposedPorts`, `Env`, `Volumes`, and the effective default command (`Entrypoint ++ Cmd`). `user`/`group` and the default command are consumed by `-r/--run`, and every field is displayed by `-i/--inspect` (see `commands.{h,cpp}` above) — `ExposedPorts`/`Env`/`Volumes` are otherwise still just captured for when networking/volumes are implemented. - `containers_storage.{h,cpp}` — wraps the `containers-storage` CLI (`import-layer`, `mount`, `unmount`, `layer --json`, `delete-layer`), forcing `fuse-overlayfs` as the overlay `mount_program`. `cleanup_layer_chain()` walks a layer's parent chain (children before parents) deleting each one. - `bwrap.{h,cpp}` — `detect_bwrap_unshare_args()` probes the kernel (via a forked `unshare(2)` per namespace type) for which `--unshare-xxx` flags `bwrap` can actually use; `build_bwrap_args()`/`run_bwrap()` assemble and run the sandboxed command. `build_bwrap_args()` also takes a `std::vector` (see `volume_mount.h` below) and appends one writable `--bind ` per entry. `wrap_for_root_namespace()` is `run_bwrap()`'s own nsenter-wrapping logic pulled out into a reusable, exported function — it's also what `volume_mount.cpp`'s copy-into-an-empty-volume step uses to reach the image's content when running rootless (see below); `run_bwrap()` itself now just calls it once on the assembled `bwrap` argv. `build_bwrap_args()`/`run_bwrap()` also take a `NamespaceConfig` (`bwrap.h`) — one plain `bool` field per `namespace_probes` entry (`user`/`ipc`/`pid`/`net`/`uts`/ `cgroup`, default `true`), resolved by `run_container()` (`commands.cpp`) from `AppConfig`'s six `global.unshare-*` keys (`config_file.h`, see above) once, up front — `bwrap.{h,cpp}` itself never touches `AppConfig`/YAML, only this already-resolved struct. For each flag `detect_bwrap_unshare_args()` finds the kernel supports, `build_bwrap_args()` additionally requires the matching `NamespaceConfig` field to be `true` (looked up via a `.cpp`-local `namespace_policy_enabled()` if-chain over `namespace_probes`' `name`s) before actually passing it to `bwrap` — kernel support and policy are separate gates, both must allow a type. This replaced an earlier hardcoded special case that always dropped `--unshare-net` regardless of policy or kernel support (without any network setup, e.g. `slirp4netns`, unsharing it just left the sandbox with no network at all) — `net` now goes through the exact same policy path as every other type, defaulting to enabled like the rest. **This is a deliberate, user-acknowledged transitional behavior change**: as of this, a plain `-r/--run` with no config file override gets a real network namespace and thus no network access at all, until `slirp4netns` integration (the next task on this same branch) actually sets one up; `global.unshare-net: off` restores the prior no-isolation behavior in the meantime. `detect_bwrap_unshare_args()` itself is untouched by any of this — still an unfiltered kernel-capability probe, unrelated to policy (no longer surfaced via `-t/--test`, see `self_test.{h,cpp}` below). `build_bwrap_args()`/`run_bwrap()` also take an optional `hostname` (from `--hostname`, long-option only): passed through as bwrap's own `--hostname` only when `--unshare-uts` is actually among the flags `bwrap` is being given (bwrap itself refuses `--hostname` without it) — otherwise logs a warning and leaves the sandbox's hostname alone, since a stock Android kernel in degraded mode may not support a UTS namespace at all. Never requests `--unshare-user` when running as root: root doesn't need a fresh user namespace for privilege, and bwrap's own single-mapping uid/gid setup for one triggers the kernel's unprivileged-userns setgroups() restriction, which showed up as every other supplementary group collapsing to the overflow gid ("nobody") in `id`, and `su` inside the sandbox failing with "can't set groups: Operation not permitted". Because of that, bwrap's own `--uid`/`--gid` (which require `--unshare-user`) aren't usable when running as root either — `--user`/`--group` work around this: when set, `build_bwrap_args()`/`run_bwrap()` bind-mount the separate `slocker-lite-priv-drop` helper (see below) into the sandbox at a fixed hidden path and route the real command through it as `: -- `. This only actually works without a user namespace (i.e. running as root) — under `--unshare-user`, the sandbox's uid map has only one valid entry, so the helper's own `setuid()` fails cleanly there instead of silently doing nothing. `run_bwrap()` fails fast (returns -1) if the helper can't be found next to this binary when `--user` was requested, rather than silently running the command as root. `run_bwrap()` also takes a `container_name` and tracks the running session with it: it passes a lambda as `run_process_foreground()`'s new `on_start` callback (see `process.{h,cpp}` below) that calls `create_session_lock(container_name, pid)` (`pid_file.{h,cpp}`, see below) the instant the real `bwrap` pid is known, then calls `release_session_lock()` once `run_process_foreground()` returns (covering every exit path — normal, nonzero, or a forwarded-signal exit — since that call always blocks until the child has actually exited). `build_bwrap_args()` no longer passes `--clearenv`/`--setenv` to `bwrap` itself; instead, `build_sandbox_env()` builds the sandboxed command's exact environment (`PATH`, `HOME`, `PWD` — hardcoded to `"/"`, matching `--chdir`'s own value; note per bwrap's own man page `--clearenv` never actually unset `PWD` in the first place, so this isn't a straight port of a prior `--setenv` — and `TERM`, only if the host process has one, followed by `extra_env` — the resolved `--env`/`--env-file` list from `resolve_env_specs()` (`env_spec.h`), appended last so it can override the built-in defaults for the same key) and `run_bwrap()` passes it straight to `run_process_foreground()`'s own `env` override (see `process.{h,cpp}` below). This works because `bwrap` (and `nsenter`, when interposed via `wrap_for_root_namespace()`) doesn't alter its own inherited environment unless told to, and neither does `slocker-lite-priv-drop` (just `setgroups()`/`setgid()`/`setuid()`/`execvp()`, no env manipulation) — so controlling it once, at the outermost exec, is sufficient for it to reach the final sandboxed command unchanged. **Because that outermost exec now uses this same explicitly-built environment**, `build_bwrap_args()`/`wrap_for_root_namespace()` resolve `bwrap`'s and `nsenter`'s own argv[0] to an absolute path via `find_in_path()` (called from this process's own, unmodified environment, before `fork()`) instead of leaving them as bare names — confirmed by direct testing: `--env PATH=...` used to break `execvp()`'s ability to even *locate* `bwrap`/`nsenter` (bare-name lookup happens in the child, using the already-overridden PATH), not just what the sandboxed command itself sees. With the fix, only the *sandboxed command's own* lookup is affected by a `--env PATH=...` override (as expected — same as overriding `PATH` in any real shell before running a bare command name), and `bwrap`/`nsenter` are always found regardless. `run_bwrap()` also takes an optional `on_bwrap_pid_known` callback, invoked alongside (not instead of) the session-lock-creation lambda, at the exact same `on_start` timing — `-D/--daemonize` (`daemonize.{h,cpp}`, see below) hooks in here via `report_daemon_started()` to learn the real pid at the same instant everything else that needs it does, rather than needing its own separate pid-discovery mechanism. That same `on_start` lambda also calls `create_session_cgroup()` (`session_cgroup.h`, see below), right alongside `create_session_lock()`, so `--kill` can later find every process the session ever starts via its dedicated cgroup; `remove_session_cgroup()` is called from the same post-`run_process_foreground()` spot `release_session_lock()` already is. - `priv_drop_helper.cpp` → the separate `slocker-lite-priv-drop` binary (its own `executable()` target in `meson.build`, **built with `-static`**). Deliberately has zero dependencies on the rest of this project (no fmt/spdlog/etc.) and is fully statically linked: it gets bind-mounted *into the container image's own filesystem*, which won't have `slocker-lite`'s own shared library dependencies — a dynamically linked binary bind-mounted that way fails outright ("error while loading shared libraries"), which is exactly what happened before this was split out (the original approach bind-mounted `slocker-lite`'s own — dynamically linked — binary via `/proc/self/exe` and reexeced it; kept only as a lesson, not as working code). Usage: `slocker-lite-priv-drop : -- [args...]`; does `setgroups(0,…)` → `setgid()` → `setuid()` → `execvp()`, in that order (dropping the group needs `CAP_SETGID`, which is lost once `setuid()` drops root). `find_priv_drop_helper()` and the `priv_drop::path`/`priv_drop::helper_name` constants live in `bwrap.h` (not just internal to `bwrap.cpp`) specifically so `exec_in_session()` (`exec_session.cpp`, see below) can reuse the exact same already-bind-mounted helper for `-x/--exec`'s own `--user`/`--group` support, instead of a second copy needing to be bind-mounted for it (which wouldn't even be possible — `-x/--exec` joins an *already-running* session's mount namespace, it doesn't get to add bind mounts to it). `find_priv_drop_helper()` itself only checks this binary's own host-side existence; it says nothing about whether a given session actually has it bind-mounted (only true when that session's `-r/--run` resolved a user in the first place). - `user_spec.{h,cpp}` — `resolve_user_and_group()` resolves a user/group spec (each a name or numeric id) against the *container's own* `/etc/passwd`/`/etc/group` **content** (not the host's, and not a path — callers own reading it, since the two current callers get that content two different ways: `run_container()` reads it directly off the merged mount path, while `exec_in_session()` fetches it over `nsenter`, since a running session's mount namespace isn't otherwise reachable from this process — see below). `nullopt` content for either file means "unreadable/absent"; a numeric user with no group still resolves fine without it (defaults gid to the same numeric value as the uid) but a named one doesn't. `ResolvedUser` (`bwrap.h`) also carries `home`, looked up by the final resolved uid's `/etc/passwd` entry (field 5) regardless of whether `user` was given as a name or a number; falls back to `"/root"` for uid 0 or `"/"` otherwise when there's no matching row. `build_sandbox_env()` (`bwrap.cpp`) sets the sandboxed process's `HOME` from this — `"/root"` only when no user override applies at all (no `--user`, no image-declared `config.User`). `run_container()` (`commands.cpp`) calls `resolve_user_and_group()` with either the explicit `--user`/`--group` flags, or, when `--user` wasn't given, the image's own declared `config.User` (`OciImageConfig::user`/`group`) — so a container defaults to running as whatever user the image itself declares, not root, unless the image declares none. - `process.{h,cpp}` — argv-based subprocess helpers (fork/execvp, no shell): `run_process()` captures stdout (used for `containers-storage` calls), `run_process_foreground()` inherits all of stdio (used for the interactive `bwrap` run). Also `find_in_path()`, a shared `$PATH` lookup. `run_process_foreground()` installs a SIGINT/SIGTERM handler around its `waitpid()` that forwards the signal to the running child and keeps waiting instead of letting the default disposition kill `slocker-lite` itself — without this, Ctrl-C (or `kill`) during `-r`'s `bwrap` run would skip `run_container()`'s unmount/cleanup entirely, leaving the layer imported and/or mounted. `run_process_foreground()` also takes an optional `on_start` callback, invoked with the child's real pid right after `fork()` succeeds (before the signal handlers go up and it blocks in `waitpid()`) — the only point where that pid is knowable, and still accurate even when `argv` itself execs into something else first (e.g. `nsenter` handing off to the final command via its own in-place `execvp()` — a pid never changes across `exec()`). `run_bwrap()` (`bwrap.cpp`) is the one caller that uses it, for session pid-file tracking (see `pid_file.{h,cpp}` below). `run_process_foreground()` also takes an optional `env` (list of key/value pairs): when set, the forked child replaces its entire environment via `clearenv()`/`setenv()` (plain POSIX, not the GNU-only `execvpe()` — the target platform includes musl) before `execvp()`, instead of inheriting this process's own. `nullopt` (the default) leaves the child's environment untouched. `run_bwrap()` is again the one caller that uses this, via `build_sandbox_env()` (`bwrap.cpp`) — see there. - `pid_file.{h,cpp}` — tracks one running `-r/--run` session (a live `bwrap` process) as a locked pid file, so an outside process (or a later `slocker-lite` invocation) can tell whether it's still running. `sanitize_for_filename()` (anything outside `[A-Za-z0-9._-]` → `_`, falling back to `"container"` if that leaves nothing) is exported here (not just `.cpp`-local) specifically so `session_cgroup.{h,cpp}` (see below) can reuse the exact same `-` naming rule for its own per-session cgroup directory without drifting from this file's own. `xdg_state_dir()` (`$XDG_STATE_HOME/slocker-lite`, or the `$HOME/.local/state/...` fallback) is likewise exported (moved out of this file's own anonymous namespace) so `persistent_netns.{h,cpp}` (see below) and `network_join.cpp`'s own per-address lease files (`xdg_state_dir() / "net-leases"`) can resolve their own subdirectories under the same state root without a second, drifting copy of this resolution logic. `session_pid_file_path()` resolves `$XDG_STATE_HOME/slocker-lite/run/-` (falling back to `$HOME/.local/state/...` when `XDG_STATE_HOME` is unset/empty — same resolution pattern as `config_file_path()` below, for state instead of config), sanitizing `container_name` first (anything outside `[A-Za-z0-9._-]` → `_`, since an image name/tag can contain `/` or `:`). `session_log_file_path()` is a sibling resolving to `$XDG_STATE_HOME/slocker-lite/logs/-.log` instead — same sanitization, same `$XDG_STATE_HOME`/`$HOME` fallback, just a different subdirectory and a `.log` extension — used by `daemonize.{h,cpp}` (see below) for `-D/--daemonize`'s log file. `create_session_lock()` creates the file (`O_CREAT|O_WRONLY|O_TRUNC|O_CLOEXEC`, mode 0644 — `O_CLOEXEC` matters: this fd must never leak into the sandboxed command's own fd table), writes the pid as text, and takes an exclusive, non-blocking `flock()` on it — held only by that fd, so its lifetime tracks `slocker-lite`'s own process lifetime (released automatically on any exit, including a crash), which lines up with `bwrap` itself being invoked with `--die-with-parent`. Any external tool can check liveness the same way: attempt the same exclusive non-blocking `flock()` on the file — success means nothing holds it anymore (stale, safe to remove), `EWOULDBLOCK` means a live process still does. `release_session_lock()` closes the fd (releasing the flock immediately) and removes the file. Every failure path here (can't create the directory/file, can't lock, can't remove) is a `spdlog::warn`, never fatal — session tracking is best-effort and must never block or fail `-r/--run` itself. `list_sessions()` implements `--list-processes` (`commands.cpp`'s `list_processes_command()`): scans the same `run/` directory and reports one `SessionInfo {pid, container_name, running}` per readable pid file. `pid` is read from the file's own contents, not parsed from the filename (ambiguous for names that themselves contain `-`); `container_name` is then recovered by stripping that exact `-` suffix back off the filename. `running` reuses the same liveness check any external tool would do — a non-blocking exclusive `flock()` that succeeds means the file is actually stale, so `running` is false in that case; the lock is always released again immediately either way, never left held by the check itself. A file that can't be opened or doesn't parse as a pid (e.g. removed mid-scan) is silently skipped, not reported as an error — scanning a live directory is inherently racy. Both `list_sessions()` and `clean_stale_sessions()` (the latter implements `--clean-processes`) share a private `open_session_file()` helper for the open/read-pid/recover-name step. `clean_stale_sessions()` doesn't just remove whatever a separate `list_sessions()` call reported as not running — it re-takes the same non-blocking `flock()` used to test liveness and holds it across the `remove()` call itself, per file, so the stale check and the removal stay atomic against a new session starting in the gap between a check and a later removal. Only files it actually removes are reported back (as `SessionInfo`s with `running=false`); still-locked (running) files are left untouched and not reported. - `persistent_netns.{h,cpp}` — generic, narrow infrastructure for keeping a network namespace alive with no process in it, the way `ip netns add` does; no `intern`/`extern` policy or bridge logic here (that's `network_bridge.{h,cpp}`, see below, which is what actually calls `create_persistent_netns()` for an `intern` network). `persistent_netns_path()` resolves `xdg_state_dir() / "netns" / sanitize_for_filename(name)` (`pid_file.h`, see above). `persistent_netns_exists()` checks whether that path is actually a live bind-mounted namespace, not just a stale/never- mounted file: `stat()`s the path and its parent directory and compares `st_dev` — a genuine bind mount always has a different device number than its parent, the same "is this a mountpoint" technique used elsewhere. Never needs root itself (just `stat()`). `create_persistent_netns()` forks a child (never touches the caller's own network namespace — `unshare(2)` affects only the calling process) that `unshare(CLONE_NEWNET)`s its own fresh namespace, bind-mounts its `/proc/self/ns/net` onto the target path, then exits immediately — the bind mount itself is what keeps the namespace alive from then on, independent of the now-exited child, exactly `ip netns add`'s own technique. Requires `CAP_SYS_ADMIN` (root) for the bind mount, matching this feature's current root-only scope (see `docs/networking-design.md`) — best-effort like this project's other host-state primitives (session locks, cgroups): logs and returns `false` on any failure (already exists, fork/unshare/mount failure) rather than throwing. `remove_persistent_netns()` unmounts then removes the file. - `network_bridge.{h,cpp}` — stands up (or confirms already-standing) a network's actual host-side state: `ensure_network_provisioned()` is idempotent by design (checks `ip link show ` first and does nothing further if it's already there) — this is deliberately also the reboot- reconciliation mechanism, not a separate code path: nothing about a network's live state (bridge, veths, iptables rules; the persistent namespace itself for `intern`) survives a reboot except its `config.yaml` entry, so calling this again after one just recreates whatever's missing. `bridge_name()` derives a stable interface name from the network's own name via a hand-rolled 32-bit FNV-1a (`"slk" + 8 hex chars`, 11 characters, comfortably under Linux's `IFNAMSIZ - 1` = 15-character limit regardless of how long the network name is) — deliberately not `std::hash`, whose exact value is implementation-defined and not guaranteed stable across a rebuild with a different standard library, which would silently orphan an already-provisioned bridge a rebuilt binary can no longer find by the name it now computes. Every command, both kinds, is wrapped through `nsenter --net=` (`wrap_for_network()`) into the network's own dedicated namespace (`persistent_netns.h`, created here first if it doesn't exist yet) — the same pattern `wrap_for_root_namespace()` (`bwrap.cpp`) already uses for the rootless `containers-storage` mount's namespace, just targeting a persistent bind-mounted path instead of a live pid's `/proc` entry. **`extern` used to run every command directly, unwrapped, since its bridge used to live in the host's own root namespace** — see this entry's own "Resolved: `extern` had no connectivity at all on the real device" paragraph further below for why that changed: confirmed by direct on-device testing to be the actual cause of a real, reproducible bug, not just an implementation choice. `bridge_name()`/`wrap_for_network()` are both exported (not just this file's own internal helpers) specifically so `network_join.{h,cpp}` (see below) can attach a container's veth to the exact same bridge, in the exact same place, this file provisioned it in. `provision_bridge()` (`.cpp`-local): creates the bridge, assigns it the gateway address from `network_subnet.h`'s `ipv4_gateway_address()` (and `ipv6_gateway_address()` if `network.ipv6`), brings it up, then — `extern` only — `sysctl -w net.ipv4.ip_forward=1` and one `iptables` `POSTROUTING`/`MASQUERADE` rule for the subnet (`! -o `, the same `docker0` shape, so bridge-local inter-container traffic isn't unnecessarily NAT'd; both idempotent global host sysctls, not per-bridge, so no separate "already enabled" tracking is needed), plus, if `ipv6`, the IPv6 forwarding sysctl — but **deliberately no `ip6tables` MASQUERADE rule**: the `fd00::/8` ULA addresses `network_subnet.h` allocates are non-globally-routable by design (RFC 4193, the IPv6 equivalent of RFC1918 private space), so NAT66 for them isn't correct IPv6 practice to begin with (also confirmed not universally supported on the real target device — its `ip6tables` build lacks a `MASQUERADE` target entirely). `extern`'s IPv6 side is thus same-bridge reachability only, exactly what `intern`'s IPv6 already was — and confirmed on the real target device, not just theorized, that this is the only option regardless: neither `ip6tables` nor `nftables` can even create an IPv6 NAT table on that kernel at all ("Not supported"). `check_network_dependencies()` gates on the tool set each kind actually needs (`ip`/`nsenter` always, both now reaching their bridge through a private namespace; `iptables`/`sysctl` additionally for `extern` — no `ip6tables`, even when `ipv6`, since none is ever called) — same shape/spirit as `commands.cpp`'s own `check_required_dependencies()`, kept separate since the tool set here depends on the network's own kind/`ipv6` setting. `create_network_command()` (`commands.cpp`, see above) calls `ensure_network_provisioned()` after resolving subnets but *before* persisting the config entry — a network that fails to provision isn't saved, so a later join doesn't find a config entry for something that doesn't actually exist on the host. **Historical verification note**: an earlier verification pass (before the "Resolved" paragraph below) confirmed a real `extern` network's bridge/gateway/`ip_forward`/`MASQUERADE` rules all coming up correctly *directly in the host's root namespace* — that description is now superseded; see below for the current, corrected architecture and why it changed. `ensure_network_provisioned()` no longer short-circuits on `bridge_exists()` alone either: the uplink step below (`extern` only) now runs on every call, even when the bridge already existed, via its own separate idempotency check. `teardown_network_state()` is the reverse: `--delete-network-full` (`commands.cpp`'s `delete_network_command()`) calls it to tear down exactly what `ensure_network_provisioned()` stood up. For `extern`: first tears down the uplink (`teardown_uplink_state()`, see below), then removes the `iptables` MASQUERADE rule for the container subnet (no `ip6tables` counterpart, since none is ever added — see `ensure_network_provisioned()`'s own comment above) — via a `.cpp`-local `teardown_step()`, the same shape as `run_admin_command()` but logging a *warning*, not an error, on failure, since a step failing because that piece was already gone by hand is the expected, common case this exists to handle (exactly `ensure_network_provisioned()`'s own existence-check caveat above — this is precisely how a manually-removed MASQUERADE rule can go from "won't come back on recreate" to "cleanly torn down and recreated" once `--delete-network-full` exists at all). **Both kinds** then remove the whole persistent namespace (`persistent_netns.h`) in one step, which destroys everything left inside it — the bridge, and for `extern` the uplink's own private-namespace-side tap device, both included — with no separate `ip link del` needed for those. Deliberately never touches the IPv4/IPv6 forwarding sysctls `provision_bridge()` enables for `extern` — those are global host state shared across every `extern` network, not per-network, so turning them off here could break others still relying on them. **Verified end-to-end on this dev machine (root, via the scoped `doas` rule)**: an `extern` network's bridge and MASQUERADE rule were both confirmed gone after `--delete-network-full` (`ip link show` reporting "Device does not exist"), and recreating a network with the *same name* afterward correctly went through `provision_bridge()` again from scratch (confirmed via the debug log) instead of short-circuiting on a stale `bridge_exists()` check — fixing exactly the gap a user reported (a manually-removed MASQUERADE rule never came back on `--delete-network` + recreate, since the old bridge was silently still there); an `intern` network's persistent namespace was similarly confirmed fully removed and recreatable without conflict. **Resolved: `extern` had no connectivity at all on the real device.** Full incident writeup in `docs/networking-design.md`'s own section of the same name — this entry just covers the resulting code. Root cause: the bridge lived directly in the host's own root namespace, and (almost certainly) Android's own `netd`-managed iptables/routing policy applies only there, never to a genuinely isolated namespace — exactly why `intern` was unaffected the whole time. `wrap_for_network()` no longer special-cases `extern` (see its own entry above) — this one change alone fixed gateway reachability, both IPv4 and IPv6, but also removes `extern`'s only path outside by construction, restored by a new **uplink**: a second, point-to-point tap+relay link (reusing `network_tap_relay.h`, see its own entry below for the new `attach_host_side_to_bridge=false` mode this needed) between the network's private namespace and the host's root namespace, on its own small deterministic `169.254.0.0/16` transit subnet (`uplink_transit_addresses()`/`uplink_transit_subnet()`, `.cpp`-local, same `fnv1a()`-based derivation `bridge_name()` already uses). `ensure_uplink_provisioned()` (`.cpp`-local, called from `ensure_network_provisioned()` for `extern` only): idempotent via its own `uplink_provisioned()` device-existence check; creates the relay (`uph` in host root, `upn` in the private namespace), assigns each end an address, sets the private namespace's own default route via the uplink, enables `ip_forward`, and adds a `MASQUERADE` rule in host root for the transit subnet — plus three more pieces, each independently required and each found by real on-device testing, not assumed (full detail, including exactly how each was diagnosed, in `docs/networking-design.md`'s own section): an `iptables -I FORWARD 1` accept rule (inserted at the front — Android's own `tetherctrl_FORWARD` chain unconditionally drops everything reaching it, so an appended rule is structurally unreachable), an outbound `ip rule` routing the uplink's own traffic into whichever table `discover_default_table()` (`.cpp`-local, parses `table ` out of `ip route get 8.8.8.8`'s own output — not hardcoded, adapts to whichever real network is currently active) names, and a **return-path** `ip rule` routing traffic *to* the transit subnet into the plain `main` table regardless of which real interface a reply arrives on. `TapRelayHandle` gained `root_side_tap_name` (`network_tap_relay.h`) since, unlike a real container join's own container-side tap (torn down for free once that session's namespace goes away), the uplink's host-root-side tap never disappears on its own — `stop_tap_relay()` removes it explicitly when set. The relay's own pid is recorded to `$XDG_STATE_HOME/slocker-lite/network-uplinks/` (`uplink_state_path()`) since it must outlive the single `ensure_network_provisioned()` call that created it, potentially spanning many separate `slocker-lite` invocations before `teardown_uplink_state()` (`.cpp`-local, the reverse of every step above, best-effort throughout) eventually stops it. **Verified completely end-to-end on the real target device, from a clean state**: gateway IPv4 0% loss, gateway IPv6 0% loss, and a real outside destination (`8.8.8.8`) 0% loss (3/3 replies) through a container on a freshly created `extern` network, both via the veth path and via `--no-veth`'s tap+relay fallback; also verified on this dev machine (self-test, a plain veth join, `--no-veth`, and `intern` — still unaffected, no uplink, `ip route`'s own "Network unreachable" for outside as intended). IPv6 outside connectivity was investigated separately and found not practically fixable on this device (confirmed `nft add table ip6 ...` itself fails — no IPv6 NAT support in this kernel at all, via either `ip6tables` or `nftables`; the alternative, NDP-proxying real addresses out of the device's own global prefix, was ruled out too, since that prefix rotates every ~10 minutes on the network tested against) — so it stays local-only by deliberate decision, matching what `intern`'s IPv6 side already was. **Resolved: `-p/--port-forward` had no connectivity at all on an `extern` network.** Reported after the connectivity/multi-network-join fixes above had already shipped: a server listening on an `extern` network wasn't reachable via `-p`, from the host or from a real outside client, even though the network's own gateway/outside connectivity (verified above) worked fine. Root cause, confirmed by direct testing (`ip route get ` from host root): moving `extern`'s bridge into a private namespace fixed gateway reachability but also meant host root had no route to the container subnet *at all* — it fell through to whatever the host's own default route happened to be (the real LAN gateway) — so `-p`'s own `DNAT` rule (added in host root, targeting the container's real IP directly — `port_forward.cpp`) had nowhere to send the rewritten packet. Fixed with two pieces in `ensure_uplink_provisioned()`, both confirmed independently necessary by direct testing — the exact same "a route alone isn't enough on Android" lesson the uplink's own outbound/return-path `ip rule`s above already learned, just for the container subnet instead of the transit subnet: a host-root route to the container subnet through the uplink's own netns-side address, plus a matching `ip rule add priority 100 to lookup main`. Without the second piece, the route added by the first is silently never consulted: confirmed via `ip rule show` on the real device that Android's own lower-priority-number policy rules — a generic `fwmark 0/0x10000 lookup 99` catch-all among them, matching any untouched/forwarded packet — intercept the packet and route it into an unrelated table with no route to the container subnet, long before rule evaluation would ever reach `main`. **A related robustness bug found while testing this fix, not the original bug itself**: a failed `ensure_uplink_provisioned()` used to only call `stop_tap_relay()`, leaving every `ip rule`/`iptables` piece already added (all deterministic, hash-derived names tied to the network name) live on the host — reproduced directly during testing (an incidental collision between two concurrent test invocations triggered a first failure, whose leftover state then made every subsequent attempt for the same network name fail identically and permanently, "File exists" on an `ip rule add` that was never removed, until fixed by hand or a full device reboot). Fixed by recording the relay's pid to the uplink state file as soon as it's known (before any of the steps that can fail), so a failure can call the exact same `teardown_uplink_state()` a real `--delete-network-full` would use to roll back everything already added, instead of a second, partial, drifting copy of that cleanup logic. **Verified end-to-end on the real target device**: a busybox `httpd` on a freshly (re-)created `extern` network, reached via `-p 18080:80` both from the device's own shell (against its real LAN IP, not `127.0.0.1` — see `port_forward.h`'s own already-documented NAT-hairpinning limitation below for why that specific case still doesn't work, unrelated to this fix) and from a genuinely separate external machine on the same LAN, got a real HTTP response back both times, reproducibly across repeated fresh network creations. `probe_veth_support()` (added for the tap+relay fallback, see `docs/networking-design.md`'s addendum and `network_join.{h,cpp}` below): the real target device supports `tun`/`tap` but not `veth` (`CONFIG_VETH` commonly stripped from mobile kernels), so joins there can't use the veth-pair mechanism this file/`network_join.cpp` otherwise assume. Probes kernel support the same way `bwrap.cpp`'s `kernel_supports_namespace()` probes namespace types: forks a child that `unshare(CLONE_NEWNET)`s into a throwaway namespace and attempts `ip link add ... type veth peer name ...` there (via the existing `run_process()`) — the whole namespace, and anything created in it, vanishes with the child, so no cleanup is needed either way. Cached in a function-local static (a fixed fact about the running kernel, not something that varies per network, so a container joining several networks in one run only probes once). `should_use_veth(network)` combines this with the network's own `veth` policy flag (`config_file.h`'s `NetworkEntry::veth`, default `true`) the same "capability and policy are independent gates" way `namespace_policy_enabled()` (`bwrap.cpp`) already combines kernel support with `global.unshare-*` policy — both must allow veth for it to actually be used. **Verified on this dev machine (root, via the scoped `doas` rule)**: `probe_veth_support()` correctly returns `true` here (a real `ip link add ... type veth ...` succeeds), and `--no-veth` at network-creation time correctly persists `NetworkEntry::veth = false`, making `should_use_veth()` return `false` even though the kernel itself supports veth — this dev machine's own way to exercise the tap+relay fallback (see `network_tap_relay.{h,cpp}`, not yet built) without needing the actual veth-less target device. - `network_join.{h,cpp}` — joins a just-started `-r/--run` session to each network named in `-n`. `join_networks()` first waits (bounded, 3s, 20ms-interval `nanosleep()` polling — `wait_for_isolated_net_namespace()`, `.cpp`-local) for `resolve_namespace_pid()` (`sandbox_process.h`) to name a child whose net namespace is actually isolated (`namespace_isolated(outer_pid, ns_pid, "net")`) — necessary because `run_bwrap()`'s `on_bwrap_pid_known` fires right after `fork()`, before `bwrap` has done any of its own namespace setup, so that child may not even exist yet the instant this is called; while it doesn't, `resolve_namespace_pid()` falls back to returning `outer_pid` itself, so comparing a namespace to itself naturally keeps the loop going without a separate "does a child exist yet" check. **Known limitation, not solved here**: for a very short-lived sandboxed command, the whole session can exit before this poll ever catches up (confirmed by testing: `-n -- echo hi` reliably timed out) — `bwrap` execs straight into the target command with no hook point in between namespace creation and exec, so there's no way for this project to guarantee network setup completes before a near-instant command already has too. Real (long-running) networked services are unaffected — confirmed by testing (see below). For each named network: looked up in `config.networks` (an unknown name is a per-network error, not fatal to the others); `ensure_network_provisioned()` (`network_bridge.h`) covers post-reboot recreation; then, per `should_use_veth(network)` (`network_bridge.h` — combines the network's own `veth` policy flag with a kernel-capability probe, see that file's own entry above), either a veth pair is created wherever that network's bridge lives (`wrap_for_network()`, reused from `network_bridge.h`), the bridge-side end attached and brought up, the container-side end moved into the session's own namespace (`ip link set ... netns `) and renamed `eth`, **or**, when veth isn't available or the network was created with `--no-veth`, `create_tap_relay()` (`network_tap_relay.h`, see below) is used instead, producing the same end state (a ready `eth` in the container's namespace) via two tap devices and a relay process rather than a kernel veth pair. `eth`'s `N` is the network's position in the `-n` list, so multiple joins each get a distinct interface, regardless of which strategy created it — everything downstream (IP assignment, routes, the address returned to `port_forward.h`) is identical either way, since it only ever operates on `eth` by name. Unlike a veth pair (torn down by the kernel automatically once the session's namespace goes away, whatever else fails), a tap relay is an independent process with no such automatic cleanup — if any step after `create_tap_relay()` succeeds fails later in `join_one_network()` (address exhaustion, a failed `ip addr add`/route command), a `fail()` helper (`.cpp`-local, only present when a relay was actually created) calls `stop_tap_relay()` before returning `nullopt`, so a partial failure doesn't leak the relay process. **Address allocation, `pick_free_address()`, needed a real fix during testing, not just design**: an interface's actual assigned IP lives inside its own private per-container namespace, invisible from the bridge's own namespace — an earlier version queried `ip -o addr show master ` (only the *host* side of each veth, with no address of its own, is visible there) and always saw nothing, so two concurrently-running containers on the same network were both handed the identical address (confirmed by testing: `10.168.0.2` twice). Fixed by giving each candidate address its own tiny lock file under `xdg_state_dir() / "net-leases"` (`pid_file.h`) and holding an exclusive, non-blocking `flock()` on it via an intentionally never-`close()`d fd — the same technique `pid_file.h`'s own `SessionLock` uses for session liveness, released automatically by the kernel the instant this process exits for any reason, no explicit release step or cleanup sweep needed. Picking a free address is then just "the first candidate (`network_subnet.h`'s `ipv4_host_address()`/`ipv6_host_address()`, `n = 2, 3, ...`) whose lock file isn't already held." For an `extern` join, `ip route replace default via dev eth` (`replace`, not `add`, so a container joining a *second* extern network doesn't fail outright with "File exists" — whichever extern network is joined last ends up as the effective default route; `intern` gets no default route at all, matching the design's "no route out exists" intent — the connected route for the local subnet is already automatic once an address is assigned, no explicit route command needed for same-bridge reachability regardless of kind). Every step failure is logged specifically (which command, which network) and best-effort: `join_networks()` returns one `JoinedNetwork {network, container_ip, relay}` per network that actually joined (in `-n` order, so shorter than the request list on any partial failure), never fatal to the already-running session (network setup can only happen after `bwrap`'s own namespace exists, i.e. potentially after the sandboxed command is already running) — this return value exists specifically for `port_forward.h` (see below) to resolve a `-p` spec against which networks/IPs are actually usable, not as a pass/fail signal on its own; `relay` (`nullopt` for a veth-joined network) is what `run_container()` (`commands.cpp`) collects to call `stop_tap_relay()` on after `run_bwrap()` returns, mirroring how it already collects `active_port_forwards` for `-p`'s own cleanup. An empty `network_names` returns immediately (no namespace wait at all), so callers that always invoke this once `on_bwrap_pid_known` fires for any reason (`commands.cpp` also fires it for `-D/--daemonize` alone, with no `-n`) don't pay for a wait that has nothing to do. Veth teardown needs no explicit code: the kernel destroys an entire veth pair (both ends, including the one still attached to the bridge) the instant *either* end's owning namespace is destroyed, so a session's veths disappear on their own once its namespace does — only the bridge/iptables/persistent- namespace state is deliberately left behind (`network_bridge.h`'s reboot-reconciliation design); a tap relay instead needs the explicit `stop_tap_relay()` call described above, since it's an independent process with no namespace of its own to be torn down by. **Verified end-to-end on this dev machine (root, via a scoped `doas` rule)**: two concurrently-running containers on the same `intern` network got distinct addresses and could ping each other; an `intern`-joined container could not reach the outside (`Network unreachable`); an `extern`-joined container reached the real internet through the bridge's NAT; a container joining both an `intern` and an `extern` network simultaneously got two working interfaces (`eth0`/`eth1`) with neither one breaking the other. **Real, separate bug found while testing this commit — since fixed** (`exec_session.{h,cpp}`, see that file's own entry below): `-x/--exec` deliberately never joined the `net` namespace type, written back when this project genuinely never isolated networking at all, so there was nothing to join. Once `-r/--run` sometimes isolates networking (whenever any `-n` was given), `-x/--exec`'ing into such a session saw the *host's* network stack instead of the container's — confirmed directly: execing into a session running a network-isolated `httpd` showed the host's own unrelated listening ports and failed to reach the container's own service on `127.0.0.1`. Fixed by joining `net` too, the same way `-x/--exec` already joins `mnt`/`uts`/`ipc`/`pid`/`cgroup`/`user` when they differ from the caller's own — reverified afterward: execing into that same session now correctly shows the container's own `eth0` and reaches its own service on `127.0.0.1`, while execing into a plain session with no `-n` at all is unaffected (still just loopback, whether or not the kernel happened to give it its own otherwise-empty net namespace via the default `global.unshare-net` policy). **Real bug reported from the real target device (`-n -- /bin/sh`, tap+relay fallback): the container-side tap device wasn't always immediately visible.** The user's own log showed `nsenter --net=/proc/ /ns/net -- ip addr add 10.168.0.2/24 dev eth0` failing with `"Cannot find device \"eth0\""` right after `network_tap_relay.h`'s relay had already created it — and confirmed by hand that simply retrying the whole session a few times eventually worked. A first fix added a bounded (~500ms) retry around the steps that touch the just-created `container_if` — later found insufficient (see below) and removed again; `join_one_network()` now uses a plain, single-attempt `run()` for every step, same as before any of this. **A tempting "fix" investigated and ruled out by direct A/B testing on this dev machine, not just reasoned about**: the obvious first instinct — have the relay *itself* self-verify the device is visible (a same-process check via its own `run_process()` call, immediately after `open_tap()`, before ever reporting success) — was tried first, in `network_tap_relay.cpp`'s `relay_child_main()`. It made things *worse*, not better: it made the container-side device **permanently invisible to every external `nsenter` afterward, 100% reproducibly** (confirmed with a 10-second retry budget — never once became visible), on a mechanism that had otherwise worked correctly and instantly on every single real session tested earlier this same day, with no retries ever needed. Root cause not fully understood (something about forking a subprocess that inherits the tap fd — opened via `open("/dev/net/tun", O_RDWR)`, deliberately *not* `O_CLOEXEC` — while still holding it open, immediately after device creation, appears to corrupt the device's *external* visibility specifically on this kernel; the *same* process's own view of the device it just created stayed correct throughout). The lesson that survived into the final fix: never add an internal, same-process/fd-holding self-check to the relay. **The retry fix above turned out to be insufficient**: a further round of real-device testing showed a *different* failure — `ip addr add` against the container-side device would sometimes succeed, only for the very next command against that same device (`ip link set eth0 up`) to fail with "Cannot find device", exhausting every retry. The device wasn't merely slow to become visible after creation; it was **disappearing on its own**, consistent with the underlying `ioctl(TUNSETIFF)`-created device (no `IFF_PERSIST`) having a more fragile lifetime on that kernel than "stays alive as long as the one fd that created it stays open." Per the user's own suggested direction, the fix was structural, not another retry: both tap devices are now created ahead of time via an external `ip tuntap add dev mode tap` (`create_persistent_tap()`, `network_tap_relay.cpp` — see that file's own entry below for the full detail), which sidesteps the whole class of symptom by making the device a genuinely persistent netdevice with no tie to any fd or process. All retry logic (`run()` is used unconditionally, everywhere) was removed as part of this — the earlier retry was compensating for a problem this fix removes outright, not one it makes more likely to need retrying. - `port_forward.{h,cpp}` — implements `-p`. `parse_port_forward_spec()` splits `"[:]:[/tcp|udp]"` on `':'` (2 or 3 fields; the network name is deliberately restricted to excluding `':'` -- `is_valid_network_name()`, `network_subnet.h` -- specifically so this split stays unambiguous) and validates both ports are `1..65535` (`.cpp`-local `parse_port()`) -- pure syntax/range parsing, no knowledge of which networks exist or joined; that's `add_port_forward()`'s job, called later once `join_networks()` (`network_join.h`) has actually run. The optional `/tcp`/`/udp` suffix (`PortForwardProtocol`, `port_forward.h` -- plain `enum class`, no prefix, same convention as `NetworkKind`) is stripped off the trailing container-port field *before* `parse_port()` ever sees it (`container_port_str` is an owned copy for exactly this, unlike `host_port_str`, which stays an alias since it never carries a suffix) -- an unrecognized value is a hard parse error (exact lowercase match only, `"tcp"`/`"udp"`, same case-sensitivity precedent as `apply_log_level()`'s `valid_levels` check, not `config_file.cpp`'s case-insensitive YAML parsing, a different context); omitted defaults to `tcp`, so every pre-existing `-p` spec keeps working unchanged. Both `PortForwardSpec` and `ActivePortForward` carry the resolved `protocol` field, and a `.cpp`-local `iptables_proto()` converts it to the exact string `iptables`'s own `-p` flag expects. `add_port_forward()` resolves `spec.network` against the `JoinedNetwork` list -- by name if given (erroring if that network wasn't successfully joined, or isn't `extern`: an `intern` network's bridge has no path from the host at all, so forwarding into one could never work), or, if unset, the container's sole joined `extern` network (erroring if none or more than one, rather than guessing). Then adds one iptables `DNAT` rule, matching `spec.protocol` (`-p tcp` or `-p udp`), to **both** `nat PREROUTING` *and* `nat OUTPUT` -- a real bug caught by testing, not assumed: `PREROUTING`-only left `curl :`, run *on this same host*, connection-refused, since `PREROUTING` only ever sees packets arriving from an actual network interface, never locally-generated ones (those go through `OUTPUT` instead) -- the same split Docker's own DNAT setup already accounts for. Also adds one `FORWARD ACCEPT` rule for the destination, matching the same protocol (in case of a default `FORWARD DROP` policy, which would otherwise silently eat the forwarded traffic even though the `DNAT` itself succeeded); if a later rule fails after an earlier one already landed, those are removed again so a failure doesn't leave a half-applied mapping. **Known limitation, not solved here, also found by testing**: `curl localhost:` (or any `127.0.0.0/8` destination) specifically still doesn't work even with both `DNAT` chains covered -- confirmed to be a separate problem, NAT hairpinning, and equally true for UDP (the martian- source check is at the IP layer, not the TCP state machine): once `DNAT` rewrites the destination to the container's IP, the packet still carries its *original* source address (`127.0.0.1`); the container's own kernel sees an inbound packet claiming to be *from* loopback arriving on a non-loopback interface (`eth`) and drops it as a martian source. (A `net.ipv4.conf.{all,lo}.route_localnet=1` sysctl was tried and confirmed *not* to fix this on its own, then removed again rather than left in as dead/superstitious code.) A full fix needs source masquerading scoped to exactly this case (matching only host-local traffic, not genuine external clients -- unconditionally masquerading would lose the real client IP for those, a regression) or a userland proxy, the approach Docker itself historically used for the same reason -- out of scope here; `curl : ` (verified working) is the actually-relevant path `-p` exists for. `remove_port_forward()` (`commands.cpp`'s `run_container()`, called for each `ActivePortForward` collected during `on_bwrap_pid_known`, after `run_bwrap()` returns) removes the exact same rules `add_port_forward()` added -- best-effort, logs a warning on failure, never fatal. **Verified end-to-end on this dev machine (root, via a scoped `doas` rule)**: a container serving HTTP on an `extern` network with `-p 8080:80` was reachable via `curl :8080` from the host; the rule was confirmed gone (connection refused) after the session was killed. **UDP support added.** `-p :/udp` produces the exact same three-rule shape (2 DNAT + 1 FORWARD ACCEPT) with `-p udp` instead of `-p tcp` throughout -- no other argv-shape change needed, since `iptables`'s own `-p`/`--dport`/`-j DNAT --to-destination` and `-d ... --dport ... -j ACCEPT` forms are identical between the two protocol modules. No dedup/coexistence tracking exists in `add_port_forward()` (it unconditionally issues `iptables -A` on every call), so the same port pair can be forwarded once per protocol without collision -- e.g. `-p 53:53/tcp -p 53:53/udp` for a DNS-like service produces two textually distinct rule sets that add/remove independently. **Verified end-to-end on this dev machine (root, via the scoped `doas` rule)**: on a freshly created `extern` network, `-p 18080:80/tcp -p 18081:80/udp` against one busybox container running both `httpd` (TCP) and `nc -u -l -p 80 > /tmp/udp-received.txt` (UDP) simultaneously -- `curl :18080` got the expected TCP response, and a raw UDP datagram sent to `:18081` (via a small Python `socket.SOCK_DGRAM` script, no `nc` available on this dev host) was confirmed to have actually reached the container by `-x/--exec`ing in and `cat`ing the received-data file afterward. Session teardown removed exactly the four rules that were added (confirmed via the debug log's matching `-D` lines for both protocols). An invalid protocol suffix (`-p 8080:80/xyz`) errored cleanly with the new parse-time message, before touching iptables or even attempting the mount/run, and cleaned up the layer import it had already done. `--clean-processes` against hand-planted stale port-forward records (a rootless test, same shape as the original crash-orphan sweep test below -- root isn't needed for the sweep *logic*, only the underlying `iptables -D` calls) correctly swept both an old-format 3-field record (defaulting to `tcp`, per the parsing note above) and a new-format 4-field `udp` record, each attempting the right `-p tcp`/`-p udp` removal and reporting `removed stale port-forward rules for ''`, while leaving a third record matching a still-running session completely untouched. **Still to verify**: real-device confirmation that UDP DNAT behaves the same way through Android's iptables/`tetherctrl_FORWARD` chain as TCP already does (assumed protocol-agnostic by the rule mechanics, not yet proven there) -- see `TODO.md`/this file's own status tracking for whether that's landed yet. **Crash-orphan sweep** (commit 6 of `docs/networking-design.md`'s sequence): unlike `join_networks()`'s veths (torn down automatically by the kernel once the session's namespace goes away) or the bridges/ persistent namespaces themselves (deliberately meant to outlive any one session — `network_bridge.h`'s reboot-reconciliation design), a `-p` mapping's iptables rules are host-global state with no automatic teardown at all — if `slocker-lite` itself is killed/crashes before reaching its own `remove_port_forward()` calls, those rules simply outlive the session forever otherwise (`bwrap` itself dies immediately in that case too, via `--die-with-parent`, so the *container* never becomes a stray process needing separate handling — only these rules can). `port_forward_state_path()` resolves `xdg_state_dir() / "port-forwards" / "-"` — deliberately the *exact* same naming scheme as `session_pid_file_path()` (`pid_file.h`), so `clean_stale_port_forwards()` can cross-reference this directory's filenames directly against `list_sessions()`'s own `SessionInfo::path` to reuse its liveness check, rather than re-deriving pid liveness a second, drifting way. `record_port_forwards()` writes one line per mapping (`" "`, `` the same `"tcp"`/`"udp"` string `iptables_proto()` produces) to that path — a no-op if there's nothing to record. `clean_stale_port_forwards()` (`commands.cpp`'s `clean_processes_command()`, alongside `clean_stale_sessions()`) scans that directory: a record whose filename doesn't match any currently-*running* session is stale — every line is parsed back into an `ActivePortForward` and removed (`remove_port_forward()`) before the record file itself is deleted; a record whose session is still running is left completely untouched. **Line parsing is deliberately per-line (`std::getline` + `std::istringstream`), not one big `while (in >> a >> b >> c >> proto)`**: chaining a 4th `>>` directly would fail-and-short-circuit the whole `while` condition on an older, pre-UDP-support 3-field line *before* the loop body (the actual removal) ever ran for it, silently leaking that rule forever with no error logged. The per-line parse instead reads the first three fields, skips the line entirely only if *those* fail, and otherwise defaults an absent or unrecognized 4th token to `tcp` — same forward-compatible posture as `config_file.cpp`'s "unknown keys ignored" policy. **Verified via a controlled scratch test** (root wasn't needed for the logic itself — only the underlying `iptables -D` calls, already proven working as root above; killing a root-owned `slocker-lite` process directly, bypassing `--kill`'s own graceful cgroup-based teardown, wasn't achievable through the scoped `doas` rule this session has, which only permits running `slocker-lite` itself): a real running session's own pid file was used to construct a *matching* port-forward record (left untouched by the sweep, confirmed still present afterward) alongside a *fabricated* record for a nonexistent pid (correctly identified as stale, its rule-removal attempted — visibly failing only for lack of root in this particular rootless test — and its record file actually removed, reported as `removed stale port-forward rules for 'faketest-999999'`). - `network_tap_relay.{h,cpp}` — a tap-backed substitute for one veth pair, used when `should_use_veth()` (`network_bridge.h`) is false: either the running kernel doesn't support veth at all (the real target device's kernel supports `tun`/`tap` but lacks `CONFIG_VETH`, commonly stripped from mobile kernels — see `docs/networking-design.md`'s tap+relay addendum), or the network was created with `--no-veth` (`cli_args.{h,cpp}`) specifically to exercise this path on a veth-capable machine. **Why tap can't just replace veth 1:1**: a veth pair is two real kernel netdevices, switched between (or into a bridge) entirely by the kernel; a tap device only has *one* kernel-side netdevice — the other "end" is a raw-Ethernet- frame file descriptor only a userspace process can read/write, so there's no second kernel endpoint to attach to a bridge (the same reason `slirp4netns`/QEMU's own tap networking need a userspace process on the fd side). `create_tap_relay(network, bridge, host_tap_name, container_ns_pid, container_if_name)` reproduces a veth pair's role with two tap devices and one relay process that copies bytes between them, **reusing the existing bridge as the switching fabric** so `network_bridge.cpp`'s `provision_bridge()`/NAT setup needs no changes at all: a host-side tap device (created wherever `network`'s bridge lives — `network_bridge.h`'s `wrap_for_network()` namespace, reached here via a direct `setns()` rather than that function's `nsenter`-argv-wrapping, since this whole sequence must keep running, and later hold onto live fds, across each namespace switch, not just for the duration of one external command) gets enslaved to `bridge` exactly like veth's host-side end does in `network_join.cpp`'s `join_one_network()`; a container-side tap device gets created directly inside the namespace named by `container_ns_pid`, **named `container_if_name` (e.g. `eth0`) from the start** — no peer-name-then-rename dance needed, unlike veth. `host_tap_name` is caller-provided (not derived here) specifically so a future caller (`join_one_network()`, once this is wired in) can reuse its own existing fnv1a-based veth-naming scheme rather than this file growing a second, drifting copy of that six-line hash. Each device is now created two-step: `create_persistent_tap(name)` (`.cpp`-local) first runs an external `ip tuntap add dev mode tap`, then `open_tap()` (`.cpp`-local, mostly unchanged) `open("/dev/net/tun")` + `ioctl(TUNSETIFF, IFF_TAP | IFF_NO_PI)`s onto that already-existing device (`IFF_NO_PI` so both ends agree on raw-frame framing with no extra header) — `open_tap()` now only *attaches* an fd to a device, it no longer *creates* one. **This split replaced an earlier, simpler design** where `open_tap()` alone both created (via the same `ioctl`, with no `IFF_PERSIST`) and attached, on the assumption the device would then simply disappear on its own once its one-and-only fd closed, the same "no explicit teardown" property veth already has — see this entry's own "tap devices need to be created persistently" paragraph further down for why that assumption turned out to be wrong on the real target device, and `docs/networking-design.md`'s matching section for the full incident writeup. The relay child's entire setup sequence — enter the network's own namespace first, always now (`persistent_netns_path()`, `persistent_netns.h` — used to be `intern` only; see `network_bridge.{h,cpp}`'s own "Resolved: `extern` had no connectivity" entry above for why `extern` needs this too now), create+ attach the host-side tap, `setns()` into the container's namespace (the already-open host-side fd stays valid across this switch — fds aren't namespace-scoped, only their *creation* is, the same property `wrap_for_root_namespace()`, `bwrap.cpp`, already relies on), create the container-side tap — reports success/failure back to `create_tap_relay()` over a `pipe2(O_CLOEXEC)` (same handshake shape `daemonize()`, `daemonize.cpp`, already uses), then falls into an unbounded `poll()`/`read()`/`write()` loop copying raw frames bidirectionally between the two fds — this loop *is* the actual "veth wire," just implemented once in userspace instead of by the kernel. No `SIGTERM` handler is installed in the relay: default disposition (terminate) already closes both fds on the way out. `stop_tap_relay()` sends `SIGTERM`, reaps the process, then explicitly `ip link del`s the host-side device (`wrap_for_network(handle.network, ...)`, reaching wherever it lives — the network's own persistent namespace, both kinds — hence `TapRelayHandle` carrying its own `NetworkEntry`) — now required since the device is persistent and no longer disappears just because the relay's fd closed (see below). The container-side device needs no matching step: it lives inside the container's own network namespace, which the kernel already tears down (every interface inside it, persistent or not, along with it) once the session itself ends — **except** the uplink's own second tap (`network_bridge.cpp`'s `ensure_uplink_provisioned()`), which lives directly in the host's root namespace instead and never goes away on its own; `TapRelayHandle` gained `root_side_tap_name` for exactly this case, and `stop_tap_relay()` removes it too (unwrapped, always host root by construction) when set. Likewise `create_tap_relay()` gained an `attach_host_side_to_bridge` parameter (default `true`, no change for either existing call site) — `false` skips the `master ` step entirely and just brings the host-side tap up plain, used by the uplink since it's deliberately a point-to-point routed link, not another bridge port. Once `create_tap_relay()` returns successfully, `container_if_name` is a completely ordinary interface from the container's own point of view — `join_one_network()`'s existing IP assignment/route/DNAT-target-address code (unchanged, not yet wired to call this) needs no changes at all. See `self_test.{h,cpp}` above for how this file's own create/attach/teardown cycle was verified end-to-end in isolation first, before being wired into `join_one_network()`. **Real fd-leak bug caught by direct testing, not assumed**: the relay child, unlike every other forked child elsewhere in this project, never `exec()`s — so `O_CLOEXEC` on fds created *before* this fork (e.g. `daemonize.cpp`'s own report-pipe write end, still open in the forking process at this point since `report_daemon_started()` — which closes it — hasn't run yet when `join_networks()` is called) never takes effect, since it only closes fds *across `exec()`*, not across a fork that never execs. Without a fix, the relay child inherited a live copy of that pipe's write end and never closed it, so `daemonize()`'s read-until-EOF in the *original, pre-fork* process blocked forever, even after `report_daemon_started()` closed its own copy — a pipe only reports EOF once *every* copy of its write end, across every process, is closed. Confirmed directly: `-r -D -n -- sleep 600` hung indefinitely; killing the session (which reaches `stop_tap_relay()` via `run_container()`'s own post-`run_bwrap()` cleanup, closing the leaked copy) immediately unblocked the original process. Fixed by `close_inherited_fds()` (`.cpp`-local): scans `/proc/self/fd` and closes everything except stdin/stdout/stderr and the report pipe's own write end, called as the very first thing in `relay_child_main()`. A first-attempt companion fix — adding the relay's pid to the session's own cgroup (`session_cgroup.h`) so `--kill` would reach it directly, since a relay is a *sibling* of bwrap rather than a descendant and so would never inherit cgroup membership on its own — was tried and then **reverted**: `remove_session_cgroup()` runs *inside* `run_bwrap()`, before `run_container()` ever gets to call `stop_tap_relay()`, so the cgroup was still non-empty (the relay still in it) at removal time, and every session using this fallback left a stray, never-removed cgroup directory behind (`rmdir` failing with `EBUSY`, confirmed by testing). Since the ordinary flow already stops the relay correctly on its own (killing bwrap unblocks `run_bwrap()`'s own `waitpid()`, letting `run_container()` finish its normal cleanup, `stop_tap_relay()` included) and the only gap left by *not* doing this is a benign, self-resolving race in `--kill`'s own "has it fully stopped" check (a genuine crash of the whole session process, not just bwrap, is a separate, already-scoped concern — see the crash-orphan sweep below), the added complexity wasn't worth it. **Real bug reported from the real target device: tap devices need to be created persistently, not tied to the relay's own fd lifetime.** Two rounds of real-device testing (`network_join.cpp`'s own entry above has the full incident writeup) found the container-side device intermittently either not immediately visible after creation, or — worse, found on the second round — visible and usable for one command (e.g. `ip addr add` succeeding) and then gone for the very next one (`ip link set ... up` failing with "Cannot find device"), on a kernel where the bare `ioctl(TUNSETIFF)`-created (no `IFF_PERSIST`) device evidently has a more fragile lifetime than "stays alive as long as its one creating fd stays open." Per the user's own suggested direction, the fix (see `create_persistent_tap()` above) creates both tap devices ahead of time via an external `ip tuntap add dev mode tap` — the same technique QEMU/libvirt use to let an unprivileged process attach to a tap device set up ahead of time — turning each into a genuinely persistent netdevice with no tie to any fd or process at all, the same as a veth pair already is. All retry logic from the first round's fix (`network_join.cpp`'s `run_with_retry()`, `self_test.cpp`'s `wait_for_container_device_visible()`) was removed once this structural fix made it unnecessary — see both files' own entries. **Verified end-to-end on this dev machine (root, via the scoped `doas` rule), using a `--no-veth` `extern` network specifically to exercise this path**: two real containers joined the same network, each getting a distinct address (`10.168.0.2`/`10.168.0.3`) via the tap+relay path with no veth involved at all, and pinged each other successfully (0% packet loss, confirmed repeatably). **Gateway/outside reachability — originally reported as an unconfirmed gap here, since resolved**: neither container could initially reach the network's own gateway IP, despite ARP resolving correctly (ruling out an L2/relay-framing problem) and the identical bridge/subnet working perfectly via veth instead (ruling out every environment-level explanation — host firewall, `rp_filter`, tried at several scopes and confirmed not to fix it — since those would affect both paths identically). **Actual cause, found once retested on a clean host**: accumulated leftover bridges/iptables rules from many earlier rounds of manual testing — `--delete-network` (before `--delete-network-full` existed, see that flag's own entry below) never tore down live host state, so stale rules/bridges from unrelated earlier test networks were still present and interfering. After manually clearing all of it and retesting fresh: a `--no-veth extern` network's gateway and a real external host both answered ICMP with 0% loss, and a raw TCP connect (`nc`) to an external host completed cleanly (a separate `wget` segfault against the same host was confirmed to be an unrelated busybox bug, reproducing identically regardless of join mechanism). Peer-to-peer connectivity and gateway/outside reachability are both now confirmed working through the tap+relay fallback on this dev machine — `--delete-network-full` exists specifically so this class of stale-state-masking-as-a-bug can't recur. **Re-verified end-to-end on this dev machine after the persistent-device redesign above**, again with `--no-veth` forcing the fallback: a single container repeatedly used its tap-relay-backed `eth0` across several commands in a row (`ip link show`, `ip addr show`, two rounds of `ping`) with no disappearance between commands — the exact symptom the real device hit — and both gateway ping and outside/internet ping (`8.8.8.8`) succeeded at 0% loss. Session cleanup left no leftover host-side tap device behind (only the bridge itself, deliberately left standing per this project's reboot-reconciliation design); `-t/--test`'s own `tap-relay create/attach/teardown` case (updated per `self_test.{h,cpp}`'s own entry above) passes reliably across repeated runs. **Real bug reported from the real target device: joining 2+ networks in one `-r/--run` left every network after the first permanently unreachable, regardless of extern/intern.** Root-caused via `strace -f` on the real device (the user's own suggestion, after several inconclusive timing-based experiments): the second network's own relay died on its literal first frame — `write(fd_container, ..., 86) = -1 EIO`, immediately followed by `exit_group(0)`. `EIO` writing to a tap fd means the device isn't administratively up yet, and it genuinely wasn't: `create_tap_relay()` returns, and this relay starts polling, the instant the container-side tap device is *created*; `join_one_network()` (`network_join.cpp`), a *different* process, still has its own `ip addr add`/`ip link set up` steps left to run afterward for that same device — confirmed via the trace's own timestamps, `ip addr add ... dev eth1` ran *after* the relay's fatal write. For the first network joined this race is narrow enough that no frame ever arrives first; for the second (and any later) network, something reliably delivers a frame before the interface is up, and the previous code treated any `write()` failure as fatal — exiting for good on that single `EIO`, so the network never worked again for the rest of the session. Fixed in the relay's frame-forwarding loop by retrying specifically on `EIO`/`ENETDOWN` (both mean "not up yet", a startup race, not a torn-down namespace) with a short bounded backoff (up to 50 × 20ms = 1s) instead of exiting immediately — generous compared to the ~14ms gap actually observed in the trace. **Methodology note**: an earlier `strace -f` attempt, wrapped in `timeout 30`, produced a misleadingly corrupted trace — GNU `timeout` sends its kill signal to the whole process group by default, and `strace`'s own tracing overhead was large enough (mount steps that normally take seconds took 3+ minutes under trace) that the real wall-clock timeout elapsed mid-setup, killing several traced children prematurely; dropping the `timeout` wrapper entirely produced a clean, complete trace. **Verified end-to-end on the real target device** with both 2 and 3 `intern` networks joined simultaneously in one session, all gateways reachable at 0% packet loss, clean teardown, no leftover state. **Crash-orphan sweep**, the direct tap+relay analog of `port_forward.h`'s own (see its own entry below): unlike a veth pair or a session's own bridge/persistent-namespace state, a relay process is host-global state with no automatic teardown at all if `slocker-lite` itself is killed/crashes before reaching its own `stop_tap_relay()` calls (`bwrap` still dies immediately in that case via `--die-with-parent`, so only the *relay* — never the sandboxed container itself — can actually leak). `tap_relay_state_path()` resolves `xdg_state_dir() / "tap-relays" / "-"` — the same naming scheme `session_pid_file_path()`/`port_forward_state_path()` already use, so `clean_stale_tap_relays()` can cross-reference filenames directly against `list_sessions()`'s own `SessionInfo::path`. `record_tap_relays()` writes one line per relay (`" "`, `` = `"extern"`/`"intern"`, `` last since it's the one field that can contain whitespace) to that path — a no-op if there's nothing to record; the ``/`` fields were added alongside the persistent-tap-device redesign above, so a later sweep can reconstruct a `NetworkEntry` and reach the right namespace to remove the now-persistent host-side device too, not just kill the relay process. `clean_stale_tap_relays()` (`commands.cpp`'s `clean_processes_command()`, alongside `clean_stale_sessions()`/`clean_stale_port_forwards()`) scans that directory: a record whose filename doesn't match any currently-*running* session is stale — every relay pid listed is `SIGKILL`ed (best-effort; an already-dead pid, or one this process was never the parent of, isn't treated as an error, since this sweep runs from a *separate* later invocation that can't `waitpid()` an orphan it didn't fork — its true parent's own exit, or `init` after reparenting, reaps it), the host-side tap device it named is removed (`ip link del`, via `wrap_for_network()` against a `NetworkEntry` reconstructed from the record's own ``/`` fields — best-effort, same as `stop_tap_relay()`'s own removal) before the record file itself is deleted; a record whose session is still running is left completely untouched. **Verified via a controlled scratch test**, the same shape `port_forward.h`'s own sweep test used: root wasn't needed for the sweep *logic* itself (only real tap/bridge creation needs it), so this ran as a plain rootless daemonized session (`-D`, no `-n`) to get a real, live pid + container name, alongside two hand-written record files in that same rootless `$XDG_STATE_HOME` — one *matching* the live session (confirmed left untouched by `--clean-processes`) and one *fabricated* for a nonexistent pid (confirmed identified as stale, its `kill()` attempt failing harmlessly with `ESRCH`, and its record file actually removed, reported as `removed stale tap-relay processes for 'faketest-999999'`); killing the real session and re-running `--clean-processes` then correctly swept its own now-stale record too. - `network_dns.{h,cpp}` — per-session DNS resolution via `dnsmasq`: a container joined to a network can resolve any other container's `--hostname` on any network the two share, plus `host.containers.internal` (podman's own convention) resolving to the first `extern` network's own gateway, if any. `is_dnsmasq_available()` (`find_in_path("dnsmasq")`) gates everything here — DNS resolution is best-effort, not a hard dependency the way `ip`/`nsenter` are; missing `dnsmasq` degrades with a warning rather than failing `-r/--run`. **One `dnsmasq` instance per session, deliberately not per network**: an earlier per-network design (mirroring how the uplink provisions one persistent process per network) was rejected before being built, once a real correctness gap was worked out — a container joined to two networks would list two `nameserver` lines in `/etc/resolv.conf`, and standard stub resolvers (glibc/musl/busybox) don't fall through to the next nameserver on NXDOMAIN, only on timeout, so a name that exists only on the *second* network would silently fail to resolve. Running one instance per session instead — entered into the *container's own* network namespace (`resolve_namespace_pid()`, `sandbox_process.h`) and bound to `127.0.0.1:53` inside it, so `/etc/resolv.conf` is always just `nameserver 127.0.0.1` regardless of which/how many networks were joined — sidesteps the whole problem: there's only ever one nameserver to ask, and it already knows about every network that specific container joined via dnsmasq's own repeatable `--hostsdir=` (inotify-based, no reload signal needed). `dns_hosts_dir(network_name)` is `$XDG_STATE_HOME/slocker-lite/dns-hosts/ ` — one shared directory per network; `record_dns_host()`/`remove_dns_host_record()` write/remove one `-` file per (network, session) into it (a plain `/etc/hosts`-syntax line, ` `, only written when `--hostname` was actually given), so sibling sessions on the same network can discover each other — and, since a session's own record lands in the same directory its own resolver watches, self-resolution works for free, no special-casing needed. `start_dns_resolver()` builds the `nsenter --net=/proc//ns/net -- dnsmasq ...` argv (one `--hostsdir` per joined network) and forks/execs it directly (no `process.h` helper fits: `run_process()`/`run_process_foreground()` both block until exit, wrong for a long-running daemon, and `close_inherited_fds()`-style fd hygiene isn't needed here the way `network_tap_relay.cpp`'s relay needs it, since this process *does* `exec()`, so `O_CLOEXEC` just works normally). **Three real bugs found via direct testing while building this, not assumed**, each confirmed by isolating the change and re-testing: 1. dnsmasq only writes `--pid-file` while actually daemonizing — `-d`/`--no-daemon` (tried first, for a simpler "the forked pid is the real pid" model) suppresses it entirely, confirmed directly: dnsmasq started and successfully read the hosts file (visible in its own log output) but the pid-file this function polls for (same bounded-poll shape as `network_join.cpp`'s `wait_for_isolated_net_namespace()`) never appeared. Fixed by letting dnsmasq daemonize normally — the forked/exec'd process is then only the *intermediate* one (reaped immediately, not tracked), and the real, final daemon pid is read back from the pid-file itself once it appears. 2. dnsmasq drops root privileges to an unprivileged user by default, which then couldn't read anything under `$XDG_STATE_HOME` at all (typically `/root/.local/state/slocker-lite/...`, mode `0700` — a non-root user can't even traverse into `/root`) — every query came back `REFUSED`, traced to `bad dynamic directory ...: Permission denied` in dnsmasq's own log. Fixed with an explicit `--user=root --group=root`, matching this project's existing root-only networking model (`bwrap.cpp` never requests `--unshare-user` when already root, for the same "no privilege drop needed/wanted here" reasoning) — tracked as a security follow-up in `TODO.md` (run it as a low-privilege user instead, with the relevant state relocated somewhere that user can reach). 3. an AAAA query for a name with only an A record (every record here is IPv4-only, matching `JoinedNetwork::container_ip`'s own existing scope) came back `REFUSED` rather than a clean "no data" answer — confirmed to break `ping ` even though the exact same name's A record had just resolved correctly via `nslookup` moments earlier, since `ping` (like most `getaddrinfo()`-based tools) queries both A and AAAA together and treats `REFUSED` on either as a hard failure for the whole lookup, not merely "no IPv6 available." Fixed with `--filter-AAAA` (turns an AAAA answer into a clean empty one instead). `host.containers.internal` needed a *second*, related fix: serving it via dnsmasq's own `--address=/name/ip` option kept returning `REFUSED` for AAAA even with `--filter-AAAA` given — confirmed `--address` records aren't treated the same internally as ordinary hosts-file entries — so it's instead served via a small, session-private `--addn-hosts=` (plain hosts-file syntax, generated once per session, removed by `stop_dns_resolver()` alongside the process itself) exactly like every other record here. `stop_dns_resolver()` (`SIGTERM` + `waitpid()`, plus removing the `host.containers.internal` file if one was generated) and `remove_dns_resolver_record()` (removes the pid-file — the same file dnsmasq itself wrote, so no separate "record" write step was ever needed, just a removal one) are called as a pair from `run_container()` (`commands.cpp`) after `run_bwrap()` returns, the same two-calls shape `stop_tap_relay()`/`remove_tap_relay_record()` already use — **a real gap caught by testing this exact pairing, not assumed**: an earlier version only called `stop_dns_resolver()`, and since dnsmasq's own pid-file was never separately removed, `--clean-processes` (harmlessly, but not cleanly) picked up every finished session's own leftover record on its next run instead of finding nothing to sweep. `clean_stale_dns_resolvers()` covers three independent crash-orphan sweeps — the resolver process itself (`dns-resolvers/-`, `SIGKILL`, same `list_sessions()`-cross-reference staleness check every other sweep in this project uses), each network's own per-session hosts record (`dns-hosts//-`, just deleted — dnsmasq's own inotify watch, wherever some other still-running session's resolver is watching that directory, notices the removal on its own), and the `host.containers.internal` file (`dns-internal-hosts/-`, likewise just deleted) — called from `clean_processes_command()` (`commands.cpp`) alongside the three existing sweeps. `--no-dns` (`cli_args.{h,cpp}`, long-option only, same precedent as `--no-veth`) opts out even when `dnsmasq` is available. `build_bwrap_args()`/`run_bwrap()` (`bwrap.{h,cpp}`) gained a `inject_dns_resolv_conf` bool, resolved once in `run_container()` from `!network_specs.empty() && is_dnsmasq_available() && !no_dns_flag` (keeping the policy decision out of `bwrap.cpp`, the same pattern `NamespaceConfig` already uses) — when true, a single static, idempotently-generated file (`ensure_generated_resolv_conf()`, content always just `nameserver 127.0.0.1\n`, since every session's resolver binds there regardless of which networks it joined) is `--ro-bind`-mounted over `/etc/resolv.conf`, which `build_bwrap_args()` otherwise never touches at all (confirmed: whatever's baked into the OCI image is what's live in the sandbox by default, no prior bind-mount or generation existed here). No special-casing needed for `-x/--exec` (already joins the same net namespace as the original session when one was requested, so it shares the same loopback and thus the same running resolver for free) or `-D/--daemonize` (spawned/stopped by the exact same callback/cleanup path as any other run). `-t/--test` exercises the full create/answer/teardown cycle in isolation (root-only skip, same as the persistent-netns/tap-relay tests): a throwaway network namespace stands in for a real session's own (same technique `test_tap_relay()` already uses), a hand-rolled minimal DNS query (`build_dns_a_query()`/`query_dns_a_record()`, `self_test.cpp` — real UDP wire format, not just "the process started") confirms the resolver actually answers a hand-written hosts record correctly. **Verified end-to-end** both on this dev machine and on the real Android target device: two containers on a shared network resolve each other (including self-resolution) and can `ping` by name; a container joined to both an `intern` and an `extern` network resolves both its `intern` peer and `host.containers.internal` simultaneously — the specific scenario the per-network-instance design would have broken. - `session_cgroup.{h,cpp}` — gives `--kill` (`kill_session.{h,cpp}`, see below) a reliable way to find every process a session ever started, however deeply forked/daemonized/reparented, by putting it in a dedicated cgroup v2 group from the moment it starts. `cgroup_v2_available()` checks for `/sys/fs/cgroup/cgroup.controllers` (the same signal systemd's own unified-hierarchy detection uses) — only cgroup v2 is supported; v1 (which splits per-controller into separate hierarchies with no unified `cgroup.procs` at the top) is deliberately out of scope, since this project's real target (Android) has used the unified v2 hierarchy by default since Android 12. `session_cgroup_path()` is deterministic — `/sys/fs/cgroup/slocker-lite/-/`, reusing `pid_file.h`'s own `sanitize_for_filename()` — so no separate lookup state is needed anywhere. `create_session_cgroup()` is called from `run_bwrap()`'s `on_start` callback (`bwrap.cpp`, see below), the same spot `create_session_lock()` already fires from: `create_directories()`'s the leaf directory (also creating the `slocker-lite/` parent the first time — a plain grouping cgroup, no resource controllers are ever enabled on it via `cgroup.subtree_control`, so the "no internal processes" restriction that comes with actually delegating controllers never applies here) and writes the bwrap pid into its `cgroup.procs`. From that point on, every process bwrap (or anything it execs into) forks inherits this cgroup automatically, permanently — including anything that later daemonizes/double-forks and gets reparented, unlike pid-namespace child membership (only the processes `clone()` itself creates) or process-group membership (many daemonizing services explicitly `setpgid()`/`setsid()` away from it on purpose). Best-effort, mirroring `create_session_lock()`: returns `nullopt` (logging a warning, never fatal) if cgroup v2 isn't available, or the directory can't be created/written (no delegated subtree when running rootless, or an SELinux policy blocking cgroupfs writes even for a root-euid process, are both real, confirmed-by- testing causes on the two environments this project actually runs on). `remove_session_cgroup()` (called from the same post-`run_process_foreground()` spot `release_session_lock()` already is) only succeeds once the cgroup is empty — a straggler process still alive at normal exit (e.g. a daemonized process that outlived the session's own main command, a pre-existing exposure independent of this feature) leaves it in place with a warning, not a fatal error. `session_cgroup_pids()` reads `cgroup.procs` — this is the actual answer to "gather every process running inside the container": unlike anything derived from `/proc` parent-pid chains or pid namespaces, cgroup membership reliably includes every process the session ever started. `session_cgroup_supports_kill()`/`kill_session_cgroup()` wrap the `cgroup.kill` knob (Linux 5.14+): writing `"1"` to it atomically `SIGKILL`s every process currently in the cgroup in one step. - `sandbox_process.{h,cpp}` — process-tree/namespace-resolution utilities shared by `exec_session.{h,cpp}` and `kill_session.{h,cpp}` (see both below); pulled into their own file (rather than staying private to `exec_session.cpp`, where `resolve_namespace_pid()` originally lived) once `--kill` needed the exact same "find the real sandboxed child" logic, to avoid a second, drifting copy. `resolve_namespace_pid()` is unchanged from its original `exec_session.cpp` form (see that entry for the full reasoning: bwrap's own outer/tracked pid never actually enters the pid/uts/ipc/cgroup namespaces it creates for its clone()'d child, only that child does). Two new utilities added alongside it for `kill_session()`: `pid_namespace_isolated(outer_pid, ns_pid)` compares `/proc//ns/pid` and `/proc//ns/pid`'s own `readlink()` targets directly — true only when bwrap's `clone()` actually created a separate pid namespace for its child (`--unshare-pid` was requested *and* the kernel supported it), the precondition for the kernel's own guarantee that killing a pid namespace's pid 1 forcibly tears down every remaining process in it. Generalized into `namespace_isolated(outer_pid, ns_pid, ns_type)` (parametrized over which `/proc//ns/` entry to compare) once `network_join.{h,cpp}` (see below) needed the exact same check for `"net"` instead of `"pid"` — `pid_namespace_isolated()` is now just `namespace_isolated(outer_pid, ns_pid, "pid")`, kept as its own function since `kill_session.h` already depends on that exact name/signature. `collect_descendant_pids(root)` generalizes `resolve_namespace_pid()`'s own `/proc//stat` ppid-scanning fallback to collect a whole transitive tree (root included) instead of just one child, sharing the actual stat-parsing loop between both via a private `build_ppid_map()` (one `/proc` pass, used by both the single-child lookup and the full-tree collection). Reliable specifically when `root` is a genuinely isolated pid namespace's own pid 1: anything that reparents within it (e.g. a daemonizing service) is guaranteed by the kernel to land back on `root` itself, unlike on a kernel without pid namespace support, where it escapes to the *host's* real pid 1 instead (see `kill_session.{h,cpp}` below for exactly this scenario, confirmed on a real target device). - `daemonize.{h,cpp}` — implements `-D/--daemonize`'s fork/detach mechanics. `daemonize(container_name)` sets up a `pipe2(..., O_CLOEXEC)` pair (so it never leaks into `bwrap`/the sandboxed command, same reasoning as the pid file's own `O_CLOEXEC`) and `fork()`s. The **child** calls `setsid()` — deliberately here, not via re-adding bwrap's own `--new-session` (removed earlier, see the `build_bwrap_args()` comment): `--new-session` only calls `setsid()` for the deeply-nested sandboxed command *inside* bwrap's own namespace setup, leaving the outer `bwrap`/`nsenter`/`slocker-lite` processes still attached to the *original* session and still receiving its signals (e.g. a `SIGHUP` when the controlling terminal closes) — not real daemonization. Calling `setsid()` in our own forked child, *before* it execs into `nsenter`/`bwrap`, detaches the entire chain at once, since `exec()` never changes session membership — confirmed by direct testing (`ps -o pid,sid,pgid,tty`): the daemon child becomes its own session leader with no controlling tty, and `bwrap` (a later descendant) shares that same session, also with no tty. The child also `sigaction()`s `SIGHUP` to `SIG_IGN` (survives the later `exec()` into `nsenter`/`bwrap`, unlike a real handler, which `exec()` resets to default — confirmed by sending `SIGHUP` directly to a running daemonized `bwrap` pid and it staying alive), then redirects stdin to `/dev/null` and stdout/stderr to a log file at `session_log_file_path(container_name, getpid())` (`pid_file.h`) — named after its *own* pid since the real session pid (`bwrap`'s) isn't known yet. If the log directory/file can't be set up at all, that's a **hard** failure here (`_exit(1)`), not best-effort — silently losing the very output `--daemonize` was asked to capture would defeat the point of the flag. The child reports `"LOG \n"` over the pipe immediately (so the parent can show a useful location even on failure) and returns `nullopt` to its caller (`run_container()`, `commands.cpp`), which then falls through into the rest of that function's existing body completely unchanged — **the daemonized child is what runs the whole rest of `run_container()`, including the unmount/ cleanup that already existed after `run_bwrap()` returns; no separate watcher/reaper process exists**. The **parent** blocks reading the pipe until EOF, returning the accumulated `"LOG "`/`"PID "` lines as a `DaemonizeResult` — the caller then prints it and exits immediately without running any session logic itself. `report_daemon_started(container_name, pid)` (called from `run_bwrap()`'s new `on_bwrap_pid_known` callback — see `bwrap.{h,cpp}` below — the instant the real `bwrap` pid is known) renames the pid-named log file to `-.log`, re-reports the *updated* `"LOG "` line (a real bug caught by testing: the parent's first `"LOG "` line names the pre-rename, daemon-pid-named path — without a second one, the parent would print a stale filename that doesn't match where the file actually ends up), then `"PID \n"` and closes its own end of the pipe — must happen here, explicitly, rather than waiting for the pipe to close naturally at the end of the (potentially very long) daemon's lifetime, or the parent would block for as long as the session runs instead of returning promptly. The pipe's write fd and the current log path are tracked as private file-scope state in `daemonize.cpp` (matching `process.cpp`'s own `g_foreground_child_pid` pattern for "there's only ever one of these per process" runtime state), since `report_daemon_started()` is called later, from a different function, not threaded explicitly through every call in between. - `exec_session.{h,cpp}` — implements `-x/--exec `: joins an already-running `-r/--run` session's namespaces via `nsenter` and runs a command inside it in the foreground. `exec_in_session()` first confirms `pid` is a tracked, running session via `list_sessions()` (`pid_file.h`) — same liveness check `--list-processes`/`--clean-processes` already use, no new logic needed there. **Key discovery, confirmed by direct testing, not assumed**: `pid` (the one `run_process_foreground()` captured and pid-file-tracked when `-r` launched `bwrap`) is bwrap's own *outer* process — it sets up the mount and user namespaces itself, then `clone()`s the actual sandboxed command into fresh pid/uts/ipc/cgroup namespaces, and `clone()`'s namespace-creation flags only ever affect the newly created child, never the caller. So the outer process itself never actually enters those namespaces — comparing `/proc//ns/{pid,uts,ipc,cgroup}` against this process's own showed them identical, while only `mnt`/`user` differed. `resolve_namespace_pid()` (`sandbox_process.{h,cpp}` — moved out of this file once `--kill` needed the exact same logic, see that entry) finds that real inner process so this can join *its* namespaces instead. For each of `{mnt→--mount, uts→--uts, ipc→--ipc, pid→--pid, net→--net, cgroup→--cgroup, user→--user}` — `net` used to be excluded here (this project never isolated networking at all, back when this comment was first written), but now that a session started with `-n/--network` (`network_join.h`) genuinely does get an isolated net namespace, skipping it left `-x/--exec` seeing the *host's* network stack instead of the container's — confirmed directly (execing into a network-isolated session showed the host's own unrelated listening ports and couldn't reach the container's own service on `127.0.0.1`), fixed by including it the same way as the other optional types: a session with no isolated net namespace at all (i.e. `net` identical to ours) just has this entry skipped like any other, so nothing changes for a session that never joined a network. `readlink()`s both `/proc//ns/` and `/proc/self/ns/` and only passes nsenter's corresponding `--type=/proc//ns/` flag when they differ — an identical-namespace re-entry attempt can fail outright (`setns()`'s own `EINVAL` restriction on re-entering a namespace you're already in), so skipping is deliberate, not just an optimization. `mnt` is the one type where a *read* failure (permission denied, or the process vanished) is treated as fatal, since without it "joining the container" is meaningless; every other type just degrades to a skip. Always appends `--preserve-credentials`: without it, `nsenter --user` also tries to `setuid()`/`setgid()`/`setgroups()` to the target's identity within the new user namespace, which fails outright (`setgroups failed: Operation not permitted`) against the `setgroups`-denied unprivileged user namespace bwrap creates whenever `-r/--run` isn't root — confirmed by hitting this exact failure during manual testing before adding the flag. Runs the final `nsenter ... -- ` via the existing `run_process_foreground()` (`process.h`) — same inherited stdio and SIGINT/SIGTERM forwarding as every other foreground external command, no new process-running logic needed. `exec_in_session()` also takes optional `--user`/`--group` (mirroring `-r/--run`'s own): given, they resolve against the session's own `/etc/passwd`/`/etc/group` (fetched via `cat` run through the same `nsenter` join, since this process can't otherwise see into that namespace, then handed to `resolve_user_and_group()` — `user_spec.h`, see below); if unset, defaults to whatever uid/gid the session's own sandboxed command is *already* running as (read from `/proc//status`), rather than root/the caller — fixing a real bug (reported after this project's own `-x/--exec` and priv-drop features had both shipped separately): without this, `-x/--exec` always ran as whatever the *host* invocation was, ignoring any `--user`/`--group` the session itself was started with. Either way, the resolved identity is applied by running `command` through the session's already bind-mounted `slocker-lite-priv-drop` helper (`priv_drop::path`, `bwrap.h`) — reused as-is, not bind-mounted again (`-x/--exec` can't add bind mounts to an already-running session's namespace anyway). Skipped entirely when the resolved uid *and* gid are both 0: a session that was never given a resolvable user at `-r/--run` time never got the helper bind-mounted at all, and dropping to 0:0 would be a no-op regardless; a missing helper for a genuinely non-root resolution instead surfaces as `nsenter`'s own "No such file or directory" once it tries to exec `priv_drop::path`, diagnostic enough on its own. **Second real bug, caught by direct testing on a rootless dev machine before this shipped**: when the session's own `-r/--run` used `--unshare-user` (i.e. ran rootless — see the root-vs-rootless paragraph below), "root inside the container" is achieved purely through the kernel's own uid mapping for that namespace, not a real privilege drop — so `/proc//status`'s uid/gid, read from *outside* that namespace, shows the host-mapped id (e.g. `1000`), not the container-relative one (`0`). Treating that as "needs a priv-drop to 1000" is wrong two ways: the helper is typically never bind-mounted for a session with no resolved `--user`, and even when it is, `setuid()` fails outright under the single-entry uid map an unprivileged user namespace gets (confirmed directly: `failed to drop privileges to 0:0: Operation not permitted`). Fixed by tracking whether the `user` namespace type was actually one of the ones joined (it only is when it differs from this process's own, i.e. exactly when `-r/--run` used `--unshare-user`) and, when so, leaving the default identity unresolved (no priv-drop) for that case — joining that same user namespace with `--preserve-credentials` (already done regardless) already reproduces the container's own view correctly via that same kernel mapping, with nothing further needed. Verified end-to-end on this same rootless dev machine: a daemonized `-r --run` busybox session with no declared user, `--exec`'d with no `--user`, now correctly shows `uid=0(root)` (previously would have attempted, and failed, a priv-drop to the host-mapped uid); an explicit `--exec --user 0` against the same session correctly resolves to `0:0` and skips the priv-drop step; `--exec --user portage` against it correctly resolves the name to its real `250:250` via the fetched `/etc/passwd` and then fails clearly (helper not bind-mounted, since the session itself had no declared user) rather than silently running as the wrong identity. - `kill_session.{h,cpp}` — implements `--kill `, stopping a tracked, running `-r/--run` session and everything it started. `kill_session()` validates `pid` the same way `exec_in_session()` does (via `list_sessions()`, `pid_file.h`). **Real bug reported by the user against their own actual target device, confirmed via a captured session log**: a plain `kill ` doesn't kill everything a container started — their `/init` script `php-fpm --daemonize`s (double-forks, detaches) then `exec caddy ...`s (replaces itself); after killing the tracked pid, both `caddy` and the `php-fpm` master+workers kept running as orphans. **Root cause**: on that device, `bwrap`'s `--unshare-pid` isn't actually in effect at all — `detect_bwrap_unshare_args()` (`bwrap.cpp`) only requests `--unshare-xxx` flags the kernel actually supports, and that kernel doesn't support pid namespaces (independently confirmed elsewhere this session, see `exec_session.{h,cpp}`'s own `CONFIG_CHECKPOINT_RESTORE` bug above) — so `php-fpm --daemonize` reparents to the *host's own* pid 1, completely disconnected from the sandboxed session; the classic "kill a pid namespace's pid 1, the kernel guarantees the whole namespace collapses" trick simply doesn't apply there. Per the user's own explicit request (they want to choose the mechanism per host capability, and may need an even more basic one later for some hypothetical older device), `kill_session()` picks between three independently-named strategies, selected dynamically per session (not a single cached host-wide capability flag, since e.g. cgroup creation can fail for session-specific reasons like permissions even on a host that generally supports cgroups) — each runs its own complete `SIGTERM` → wait-up-to-`grace_period_seconds` (10s default, no CLI flag) → forced-`SIGKILL` escalation internally, with no cross-strategy fallback-after-failure chaining: 1. `kill_via_cgroup()` — preferred whenever the session has a non-empty dedicated cgroup (`session_cgroup_pids()`, `session_cgroup.h`): `SIGTERM` to every pid currently in it, and, if forcing is needed, either the atomic `cgroup.kill` knob or a fresh re-read-and-`SIGKILL` sweep (fresh, not the original snapshot, since a process could have forked a new child after the graceful sweep but before dying). The **only** mechanism that reliably reaches every process regardless of pid namespace support. **Critical correctness point, caught during design review before this shipped**: the "is it stopped yet" poll must gate on the *cgroup being empty*, not `list_sessions()`'s running flag — that flag only reflects the pid file's flock, released the moment the tracked outer `bwrap` pid exits, and the `SIGTERM` sweep necessarily hits `bwrap` itself too (it's a cgroup member) — `bwrap` dies and gets reaped in well under a second, long before slower descendants (`caddy` shutting down gracefully, `php-fpm` finishing in-flight requests) actually exit. Gating on the pid file instead would make the poll resolve "done" almost immediately, the forced-kill step would never run, and the original bug would reproduce with unused machinery around it. 2. `kill_via_pid_namespace()` — used when no cgroup exists for the session, but `resolve_namespace_pid()`/`pid_namespace_isolated()` (`sandbox_process.h`) confirm `--unshare-pid` was genuinely in effect for it. `SIGTERM`s `collect_descendant_pids(ns_pid)` (reliable here specifically because reparenting within a genuinely isolated pid namespace always lands back on that namespace's own pid 1); if forcing is needed, a single `SIGKILL` to `ns_pid` alone is *guaranteed* complete by the kernel itself, independent of whatever the graceful sweep missed. "Stopped" is simply `kill(ns_pid, 0)` failing with `ESRCH`. Verified end-to-end on this project's rootless dev machine (which does support pid namespaces, unlike the user's real target device): a daemonized busybox session running `sh -c 'sleep 300 & exec sleep 300'` (mirroring the daemonize-then-exec shape of the original bug) was fully cleaned up by `--kill`, including the backgrounded child, with no leftover processes, mounts, or layers; a second run using `sh -c 'trap "" TERM; sleep 300'` (ignoring `SIGTERM` entirely) confirmed the forced-`SIGKILL` escalation path too, taking the full 10s grace period before the pid namespace's own collapse-on-kill guarantee cleaned it up regardless. 3. `kill_via_tracked_pid()` — fallback when neither of the above applies: signals the tracked `bwrap` pid directly, `SIGTERM` then `SIGKILL`, polling `list_sessions()` for "stopped" since that's the only signal available without a cgroup or an isolated pid namespace to check directly. Exactly today's manual-`kill` behavior — least complete, but always available, and strictly no worse than before this feature existed. This is the path the user's own real target device actually takes today (no cgroup delegation confirmed working there yet; no pid namespace support at all) — a future, even more basic strategy (for some hypothetical still-more-limited device) would slot in here the same way, per the user's own explicit request to keep this extensible. `poll_until()`/`sleep_ms()` (`.cpp`-local) use `nanosleep()` in an `EINTR`-retry loop — matching this project's existing direct-POSIX style (`process.cpp` already retries `waitpid()` the same way) — rather than ``/`` (unused anywhere else in this project). - `config_file.{h,cpp}` — `load_config_file()` reads and parses (via libyaml's document API, ``) the `global`, `volumes`, and `networks` sections of the local YAML config file located by `config_file_path()` (`$XDG_CONFIG_HOME/slocker-lite/config.yaml`, falling back to `$HOME/.config/slocker-lite/config.yaml`). Supported `global` keys: `log-level`, and six `unshare-` keys (`unshare-user`/`unshare-ipc`/ `unshare-pid`/`unshare-net`/`unshare-uts`/`unshare-cgroup`, one per `bwrap.cpp`'s own `namespace_probes` entry) controlling whether `-r/--run` requests each of bwrap's `--unshare-xxx` flags — every other long option is a one-shot flag, not a setting, so it doesn't belong in a persistent config file. Each `unshare-*` key accepts `"1"`/`"on"`/`"yes"`/`"true"` (enabled) or `"0"`/`"off"`/`"no"`/`"false"` (disabled), case-insensitively (`.cpp`-local `parse_bool_flag()`); an unset key defaults to enabled, and an unrecognized value logs a `spdlog::warn` and is treated as unset (default enabled) rather than failing the whole config load — consistent with this file's existing forward-compatible/ignore-malformed-entries policy (only malformed *YAML syntax* is a hard error). A missing file returns a default-constructed (empty) `AppConfig`, not an error; unknown sections/keys (and malformed individual volume entries) are likewise ignored for forward-compatibility. `main()` applies `config->log_level` (via the existing `apply_log_level()`) right after `spdlog::cfg::load_env_levels()` and before parsing CLI options, so an explicit `--log-level` on the command line always overwrites it afterward — same precedence pattern already used for `SPDLOG_LEVEL`. `write_config_file()` writes the whole file back out (via libyaml's document-building/emitter API, symmetric to the read side) — used by `-v/--volume` (`create_volume_command()`, `commands.cpp`) to persist a new `VolumeEntry {name, directory}` into the `volumes` section, preserving `global` (including any set `unshare-*` keys, re-serialized as canonical `"true"`/`"false"`) untouched. **`VolumeEntry`/the `volumes` section is a distinct concept from `OciImageConfig::volumes`**: this is a user-defined `name -> host directory` mapping created via `-v/--volume`, not an image's own declared mount points (still unconsumed, see `oci_image.{h,cpp}` above). `AppConfig`/`config.volumes` is looked up by name in `resolve_volume_mount()` (`volume_mount.{h,cpp}`, see below), which is how `-r/--run`'s own `-v` usage finds a named volume's host directory. `run_container()` (`commands.cpp`) resolves the six `unshare-*` fields (each `value_or(true)`) into a `NamespaceConfig` (`bwrap.h`, see below) once, up front, and passes it to `run_bwrap()` — `bwrap.{h,cpp}` itself has no dependency on this file or on YAML parsing at all, only on the already-resolved, defaults-applied struct. **`networks` section** (see `docs/networking-design.md` for the full feature): unlike `volumes` (a flat `name -> directory` scalar mapping), each network entry is itself a *nested* mapping (`kind`/`subnet`/`ipv6`/ `subnet6`), since one network needs more than a single value to describe. `NetworkEntry` (`kind` is `NetworkKind::extern_`/`intern` — trailing underscore on `extern_` since `extern` is a reserved C++ keyword and can't be an enumerator name — parsed from the YAML strings `"extern"`/`"intern"`) round-trips through `AppConfig::networks` the same way `VolumeEntry` does; an entry with an unrecognized `kind` (or missing `kind`/`subnet`) is skipped on load, same forward-compatible policy as everything else here. `ipv6` reuses `parse_bool_flag()`, defaulting to `true` (enabled) if absent or unparseable; `subnet6` is only read/written when `ipv6` is true. `write_config_file()` writes each network as its own nested mapping under `networks`, `ipv6` re-serialized as canonical `"true"`/`"false"` like the `unshare-*` keys. `veth` (default `true`) round-trips the same way as `ipv6` (`parse_bool_flag()`, written as canonical `"true"`/`"false"`, always written regardless of value — unlike `subnet6`, there's no companion field whose presence depends on it) — see `network_bridge.h`'s `probe_veth_support()`/`should_use_veth()` above and `--no-veth` (`cli_args.{h,cpp}`) below for what it controls. - `network_subnet.{h,cpp}` — pure CIDR arithmetic backing `-n/--network`'s subnet allocation and `network_bridge.{h,cpp}`'s (see below) gateway-address computation; no kernel/`ip`/`iptables` calls of its own. `is_valid_network_name()` (non-empty, no `':'`) mirrors `volume_mount.h`'s `is_valid_volume_name()` (which rejects `'/'`) — `':'` specifically because `port_forward.h`'s `-p` syntax splits a spec on it; a network name containing one would make that parse ambiguous. `is_valid_ipv4_cidr()`/ `is_valid_ipv6_cidr()` and `ipv4_cidrs_overlap()`/`ipv6_cidrs_overlap()` all build on one `.cpp`-local `parse_cidr()` (via `inet_pton()`, not hand-rolled parsing) producing a plain byte-vector address (4 bytes for IPv4, 16 for IPv6) + prefix length, and one shared `bytes_overlap()` byte/bit-mask comparison generic over that byte length — IPv4 and IPv6 overlap checking are the same algorithm, not two parallel implementations. `allocate_ipv4_subnet()`/`allocate_ipv6_subnet()` (`commands.cpp`'s `create_network_command()`) scan `10.168..0/24`/ `fdf0:f243:f06f:<168+n>::/64` for `n` in `0..255` and return the first one that doesn't overlap *any* existing network's subnet (via the overlap checks above, not just other auto-allocated ones — a manually `--subnet`-overridden network is checked too). `fdf0:f243:f06f::/48` is a randomly generated ULA (RFC 4193) — **replaced the original `fd00:168:0::/48`, which was never actually randomly generated, just a memorable placeholder, at the user's own request** ("since this is an ULA, let's use a randomly generated prefix"); the `168` offset on the subnet-id hextet (confirmed with the user via `AskUserQuestion`, over the alternative of dropping it and starting both at `0`) keeps the same project-recognizable stamp the old scheme's fixed 2nd hextet had, just as a constant offset now rather than a numerically-identical index — the same `n` range still drives both v4 and v6 allocation, so the common case (no manual overrides) still allocates deterministically paired blocks per network, just offset by `168` on the v6 side instead of matching exactly. Since IPv6 hextets are hexadecimal, `n + 168 >= 10` (i.e. always, given the offset) renders as a valid but numerically-different-from-`n + 168` address when read back as hex (e.g. `n=15` → `183` decimal → renders as `...:183::/64`, which is hex `0x183`, not `183`) — purely cosmetic, allocation correctness doesn't depend on this matching numerically at all. `ipv4_gateway_address()`/ `ipv6_gateway_address()` (`network_bridge.cpp`'s `provision_bridge()`) and `ipv4_host_address()`/`ipv6_host_address()` (`network_join.cpp`'s `pick_free_address()`, see below — `n = 2, 3, ...` for individual containers) are all thin wrappers around one shared `.cpp`-local `host_address(af, cidr, n)`: masks `cidr` down to its network address first (`mask_to_network()`, in case it — e.g. a manual `--subnet`/`--subnet6` — wasn't already a canonical network address), then adds `n` as a big-endian integer into the trailing host-portion bytes with proper carry propagation (generic over address length, so the same code handles both IPv4's 4 bytes and IPv6's 16 without two parallel implementations), rejecting `n` outright if it doesn't fit the address's host-bit width. The gateway functions are just `host_address(af, cidr, 1)` — the `.1` convention this project's bridges use. Reuses the same `parse_cidr()` as the validation/overlap functions above. - `volume_mount.{h,cpp}` — `is_valid_volume_name()` (no `/`, checked by both `create_volume_command()` and to tell a `-v` spec's name/path apart) and `resolve_volume_mount()`, called once per `-v` occurrence from `run_container()` when running with `-r`. A spec with no `/` is looked up in `config.volumes` by name (error if unknown); one with `/` is treated as a host directory path and `create_directories()`'d if missing. If the resulting host directory is empty, `initialize_volume_directory()` reconciles it against the image's own directory at the given container path: if that image directory is non-empty, its contents are copied in first; then, **whether or not there was content to copy**, the host directory's own mode/ownership/timestamps (and xattrs/ACLs where supported) are always set to match the image directory's own — **real bug fixed by the user, not assumed**: an earlier version only ever copied when the image directory was non-empty, so an image declaring an *empty* directory with specific ownership/permissions (e.g. a data directory owned by a non-root uid/gid) got a host directory with default `create_directories()` permissions instead, and even the non-empty case never reconciled the directory's *own* attributes (only each copied entry's). The existence check, content copy, and attribute reconciliation all run as a single `sh -c` invocation wrapped through `wrap_for_root_namespace()` (`bwrap.h`) — **not** a plain `std::filesystem` check — because a rootless `containers-storage mount`'s content isn't visible to this process at all without `nsenter`, the same constraint `run_bwrap()` itself works around (see the root-vs-rootless paragraph below). `cp -a --preserve=mode,ownership,timestamps,links[,xattr] --attributes-only -T` does the attribute-reconciliation step; whether `,xattr` is included is decided by a direct `setxattr()`/`removexattr()` probe on the host directory (no new library dependency — Linux POSIX ACLs are themselves stored as xattrs, so this one probe stands in for both, logging a single `spdlog::warn` if unsupported). `-T`/`--no-target-directory` is required on that second `cp` — confirmed by direct testing: without it, since the host directory already exists, plain `cp SRC DST` copies `SRC` *into* `DST` as a nested `DST/basename(SRC)` subdirectory instead of reconciling `DST`'s own attributes, which is exactly the bug this fix closes. A nonzero `cp` exit is only ever a warning, never fatal — often just an ownership-preservation shortfall when not running as root. Verified end-to-end against `images/gitea.tar`'s real declared `/etc/gitea`/`/var/lib/gitea` volumes under a real rootless mount: the resulting host directories' mode/ownership matched the image's own declared values in both cases, and no mounts/layers were left behind afterward. Errors are logged via `spdlog::error`; every external command is also traced at debug level in `run_process()`/`run_process_foreground()` (`src/process.cpp`) — visible via `SPDLOG_LEVEL=debug`, since spdlog's default level is `info` — and a failed external command additionally logs a `spdlog::warn`, which is visible by default (no env var needed). The final "mounted image at: ..." success line is direct stdout program output, not a log. Because `containers-storage mount` runs rootless, it reexecs itself into a private user+mount namespace to gain the privilege it needs for the overlay mount — which leaves the result invisible to a plain shell or child process outside that namespace. Confirmed `containers-storage unshare` does **not** rejoin an already-running mount's namespace; only `nsenter` targeting the live `fuse-overlayfs` daemon's PID does. `-r/--run` handles this automatically by locating that PID and running `bwrap` via `nsenter` into its namespaces (`wrap_for_root_namespace()`, `src/bwrap.h`) — reused as-is by `volume_mount.cpp`'s copy-into-an-empty-volume step, since that also needs to read image content that's otherwise invisible outside the same namespace. **Running as root sidesteps all of this**: no privilege reexec is needed, so the mount is already directly visible in the current namespace, and `nsenter --user=...` into it then fails ("reassociate to namespace 'ns/user' failed: Invalid argument") since the caller is already in that same user namespace. `-r/--run` detects `geteuid() == 0` and skips `nsenter` automatically in that case; `--no-nsenter` forces it off manually for any other situation where the mount turns out to already be directly visible. **Mutable global state and multi-container support:** an audit ahead of planned docker-compose support (running multiple containers at once) found exactly three pieces of mutable global/file-scope state in `src/`: `g_mount_program` (`containers_storage.{h,cpp}`, the resolved `fuse-overlayfs` path — genuinely process-wide, invariant across containers), `g_foreground_child_pid` (`process.cpp`, plus `run_process_foreground()`'s process-wide `SIGINT`/`SIGTERM` handler installation — tracks one foreground child at a time), and `g_report_fd`/`g_log_path` (`daemonize.cpp`, one in-flight `-D/--daemonize` handshake's report-pipe fd and log path). **Decision, confirmed by the user:** multi-container/compose support will run each container's session in its own forked OS process — the same model `-D/--daemonize` already uses — rather than one process managing multiple containers concurrently without forking. Under that model, "one OS process" and "one running container" stay the same thing they already are today, so **none of these globals need to become per-container state** — each forked child only ever tracks/signals one foreground child and handles one daemonize handshake, exactly as today. This is a load-bearing constraint for however the compose orchestrator ends up implemented: it must fork (not thread, not run an in-process event loop over N containers) one child per service, each child reusing `run_container()`'s existing single-container code path unchanged. ## Build & test commands Build directory is `buildDir/` (already configured). - Configure (only needed if `buildDir/` is missing or deleted): `meson setup buildDir` - Build: `meson compile -C buildDir` (or `ninja -C buildDir`) — also builds `buildDir/slocker-lite-priv-drop`, the statically-linked helper `-r --user` needs (see `priv_drop_helper.cpp` in "Project state") - Run the executable: `./buildDir/slocker-lite -m ` (see `--help` for the full flag list: `-m/--mount`, `-r/--run`, `-u/--umount`, `-c/--cleanup`, `-l/--list-images`, `-i/--inspect`, `-x/--exec`, `--kill`, `--no-nsenter`, `-D/--daemonize`, `--user`, `--group`, `--hostname`, `--env`, `--env-file`, `-v/--volume`, `--list-volumes`, `--delete-volume`, `--delete-volume-full`, `-n/--network`, `--extern`, `--intern`, `--subnet`, `--no-ipv6`, `--subnet6`, `--no-veth`, `--list-networks`, `--delete-network`, `-p/--port-forward`, `--no-dns`, `--list-processes`, `--clean-processes`, `-w/--write-config`, `-t/--test`, `--log-level`, `-h/--help`, `-V/--version`) - Run tests: `meson test -C buildDir` ## Code style - Null-pointer checks: prefer `if (!ptr)` / `if (ptr)` over `if (ptr == nullptr)` / `if (ptr != nullptr)`. - Constants: no `k` Hungarian-notation prefix. `enum class` values are already qualified by the enum's own name (e.g. `Mode::run`, `OciPortProtocol::tcp`), so plain snake_case enumerators are enough on their own. Free-standing constants also use plain snake_case; when several are conceptually related, group them under a named `namespace` instead of relying on a shared prefix to imply the grouping (e.g. `cli_args.cpp`'s `getopt_long` long-option codes live in `namespace options { constexpr int log_level = ...; }`, and `bwrap.cpp`'s priv-drop-helper path/binary-name pair live in `namespace priv_drop { ... }`) — nest the named namespace inside the file's existing anonymous namespace where one is already present, so internal linkage is unchanged. A `kXxx`-named identifier that turns out not to actually be `const` (mutable global/static state) instead follows this codebase's existing `g_` prefix convention (e.g. `containers_storage.cpp`'s `g_mount_program`, matching `process.cpp`'s `g_foreground_child_pid` and `daemonize.cpp`'s `g_report_fd`/`g_log_path`). ## Licensing - Every `.c`/`.cpp`/`.h` file under `src/` must start with the GPLv2-or-later copyright header (see any existing file under `src/` for the exact text). - After adding a new source file under `src/`, run `./add-license.sh` from the repo root to prepend the header (it reads `copyright-header` and inserts it via `sed`, skipping files that already have it, so it's safe to re-run at any time). ## Build configuration notes - `meson.build` sets `warning_level=3` and `cpp_std=c++20` — keep new code warning-clean under `-Wall -Wextra -Wpedantic`-equivalent settings. - The single Meson `test()` target runs `slocker-lite` against a fixture OCI image tar generated at build time by `tests/gen_fixture.py` (a `custom_target`) and checks its exit code (no test framework is wired in yet).