Commit 2/6 of the network isolation feature (docs/networking-design.md).
persistent_netns.{h,cpp}: create/verify/remove a network namespace kept
alive with no process in it, the way `ip netns add` does (fork a child,
unshare(CLONE_NEWNET), bind-mount its /proc/self/ns/net onto a
persistent path, exit -- the bind mount keeps it alive). Root-only
(CAP_SYS_ADMIN for the bind mount), best-effort like this project's
other host-state primitives. Not wired into -n/--network yet.
xdg_state_dir() (pid_file.cpp) moved out of its anonymous namespace so
this file can reuse the same $XDG_STATE_HOME resolution rather than a
second, drifting copy.
-t/--test now exercises the create/verify/remove cycle (skipped with a
message, not a failure, when not root) -- confirmed working via doas.
71 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Project state
slocker-lite (C++20, built with Meson) mounts an OCI Image Layout tar (oci-layout +
index.json + blobs/sha256/*, as produced by skopeo/podman save --format oci-archive/modern docker save) using containers-storage and fuse-overlayfs, then
(via -r/--run) runs a sandboxed command against it with bwrap. The real deployment
target is Android with a stock kernel, where podman/docker don't run (missing
namespace support) and there's no kernel overlayfs (hence fuse-overlayfs); bwrap is
invoked in "degraded mode" using only whichever --unshare-xxx namespaces the running
kernel actually supports. See README.md for the human-facing overview (build/usage/
status); this file stays the dense, file-by-file reference. Still early-stage.
Source layout (all under src/):
-
main.cpp— the CLI entry point only, and deliberately tiny (~30 lines): loads the config file, applies itsglobal.log-level(apply_log_level(),cli_args.h) before CLI parsing so an explicit--log-levelcan still override it afterward, callsparse_args()(cli_args.h) and returns its exit code immediately if it gives one (covers-h/-Vand every parse error), otherwise callsdispatch_command()(commands.h) and returns its result. All of the actual option-parsing and command logic that used to live here moved out intocli_args.{h,cpp}/commands.{h,cpp}/self_test.{h,cpp}(see below) specifically to keep this file from re-growing into a dumping ground as more commands (docker-compose support, etc.) get added. -
cli_args.{h,cpp}— command-line parsing only, nothing else.Mode(the enum of every CLI action) andParsedArgs(everythingparse_args()extracts fromargv) live in the header sincecommands.h'sdispatch_command()consumes them;print_usage()/print_version(), theoptionsnamespace of getopt long-option codes, and thelong_optionsarray itself are.cpp-local.parse_args(argc, argv, out)runs thegetopt_longloop plus all of the post-loop validation that used to live at the top ofmain(): mode/--volumeinteraction (-valone vs. combined with-r, see below),--grouprequires--user, no leftover positional args outside-r/-e, and (moved here from what used to be inline in theMode::execdispatch arm)-x/--exec <pid>'s own pid parsing/validation (ParsedArgs::exec_pid, a positive integer or a hard error) and its trailing-command requirement (ParsedArgs::command, required non-empty).--kill <pid>(ParsedArgs::kill_pid) shares that same positive-integer parsing/validation via a small extractedparse_pid_arg()helper (.cpp-local) rather than duplicating thestrtoldance a second time — unlike-x/--exec, it takes no trailing command, so it's simply not added to the leftover-args exemption list (Mode::run/Mode::execonly). Returns an exit codemain()should return immediately (0for-h/-V,1for any parse error) when set;nulloptmeansoutis ready fordispatch_command().-D/--daemonize(has a short form;'D'was free) is a plain boolean flag (ParsedArgs::daemonize_flag, set in its owncase 'D':, same pattern as--no-nsenter).--hostname <name>/--env VAR=VALUE/--env-file <file>(all long-option only,--env/--env-fileboth repeatable) are collected here intoParsedArgs::hostname_flag/env_specs—--envpushes{false, optarg},--env-filepushes{true, optarg}into the same orderedstd::vector<EnvSpec>(not two separate lists), preserving their exact relative command-line order across both flags, sinceresolve_env_specs()(env_spec.{h,cpp}, see below) needs that order to let a later one override an earlier one for the same variable name — actually resolving them happens later, incommands.cpp'srun_container().-v/--volumeis dual-purpose: used alone it's a standaloneMode::volumerequest; combined with-r/--runit's repeatable and requests a volume mount instead (resolved later byresolve_volume_mount(),volume_mount.{h,cpp}, see below). Since-vmust be repeatable with-rbut each occurrence still takes two space-separated tokens, the getopt loop doesn't let'v'setModedirectly: it accumulates(spec, path)pairs intoParsedArgs::volume_specs(consuming the second token manually, with a guard against swallowing the next flag if there isn't one), and only after the loop decides whether that means one standaloneMode::volumecall or, together with-r, passesvolume_specsthrough unresolved fordispatch_command()/run_container()to handle.-n/--network(seedocs/networking-design.mdfor the full feature design) is dual-purpose the same way, but simpler: since a network join has no equivalent of a volume's container-mount-path second argument, each occurrence is a singlerequired_argumenttoken (just the name) accumulated intoParsedArgs::network_specs— no manual second-token consumption needed,'n''s owncasejust doesnetwork_specs.push_back(optarg). The same post-loop split as-vdecidesMode::network(standalone, exactly one occurrence) vs. join-with--r(repeatable, no limit).--extern/--intern/--subnet <cidr>/--no-ipv6/--subnet6 <cidr>(ParsedArgs::network_extern_flag/network_intern_flag/network_subnet_flag/network_no_ipv6_flag/network_subnet6_flag) only apply to the standalone (create) case and are rejected with a clear error if given any other way (e.g. alongside-r) —Mode::networkadditionally requires exactly one of--extern/--intern.-nused to belong to--no-nsenter: reassigned here since--networkwill be far more heavily used;--no-nsentermoved to long-option-only (options::no_nsenter) rather than hunting for a new letter, matching--kill's own "rare/niche flag, long-only is no real loss" precedent. -
commands.{h,cpp}— every command's implementation, plus the dispatcher.dispatch_command(args, config_path, config)(the only externally-linked function; everything else in this file is.cpp-local) is aswitch (args.mode)with one explicitcaseperModeenumerator and nodefault:, so-Wswitch(this project builds atwarning_level=3) forces a compile warning/error if a futureModevalue is ever added without a matching dispatch case, instead of silently falling through to the wrong command — confirmed by testing (temporarily adding an unhandled enumerator triggered exactly the expected-Wswitchwarning).Mode::mounthas its own explicit case (mount_command()) for the same reason: it used to be handled only by falling off the end of a longif/elsechain inmain()with no explicit check at all — the very kind of implicit, easy-to-silently-break behavior this dispatcher redesign exists to close off, especially with moreModevalues (docker-compose support) expected soon.list_processes_command()implements--list-processes(long-option only): callslist_sessions()(pid_file.{h,cpp}, see below) and prints one tab-alignedpid,container name,running/exitedrow per entry (same two-column tab-alignment scheme aslist_images_command()/list_volumes_command(), extended to a third column), no header row, silent success on an empty list.clean_processes_command()implements--clean-processes(also long-option only): callsclean_stale_sessions()(pid_file.{h,cpp}) and prints oneremoved stale pid file for '<name>' (pid <pid>)line per file actually removed — nothing is printed for sessions still running, and an empty result (nothing stale) is silent success, same convention as the rest of this file's list/delete commands.Mode::exec's dispatch case is a one-line call toexec_in_session(*args.exec_pid, args.command)(exec_session.{h,cpp}, see below) — the pid/command parsing and validation now happens incli_args.cpp'sparse_args()instead (see above).inspect_image_command()implements-i/--inspect <image.tar>: prints everyOciImageConfigfield (user/group, exposed ports, env, volumes, default command) without mounting or running the image — extend it wheneverOciImageConfiggains a new field (seeoci_image.{h,cpp}below).run_container()unconditionally callsread_oci_image_config()and reuses the result for two independent defaults: the command to run (Entrypoint ++ Cmd) when none is given on the command line, and, when--userwasn't given, the sandboxed process's user/group (config.User, split intoOciImageConfig::user/group) — an explicit--user/--groupon the command line always takes precedence.create_volume_command()implements-v/--volume <name> <directory>;list_volumes_command()implements--list-volumes(same tab-alignment scheme aslist_images_command(), reused as-is);delete_volume_command()implements both--delete-volume <name>(config entry only) and--delete-volume-full <name>(alsostd::filesystem::remove_all()s the host directory — errors out before touching the config if that fails, warns instead of failing if the directory was already gone) — seeconfig_file.{h,cpp}below for what a "volume" means here (a distinct concept fromOciImageConfig::volumes).dispatch_command()'sMode::volume/Mode::delete_volume/Mode::delete_volume_fullcases call these.create_network_command()/list_networks_command()/delete_network_command()are the direct network equivalents (Mode::network/Mode::list_networks/Mode::delete_network) — seedocs/networking-design.mdfor the full feature design andconfig_file.{h,cpp}below forNetworkEntry. This commit is config-only: no bridge, namespace, or iptables state is created yet, only the config entry (later commits in the design doc's sequence wire up the actual host-side networking).create_network_command()rejects a duplicate name first, then resolvessubnet/subnet6: an explicit--subnet/--subnet6is validated (is_valid_ipv4_cidr()/is_valid_ipv6_cidr()) and checked for overlap against every existing network's subnet (ipv4_cidrs_overlap()/ipv6_cidrs_overlap(),network_subnet.{h,cpp}, see below); otherwise the next free block is auto-allocated (allocate_ipv4_subnet()/allocate_ipv6_subnet()).list_networks_command()reuses the same independently-per-column tab-alignment scheme aslist_processes_command()(name/kind/subnet each aligned, then the IPv6 subnet — or"(no ipv6)"— appended unaligned as the trailing column, nothing follows it).write_config_command()implements-w/--write-config: unlikecreate_volume_command()/delete_volume_command()'s use ofwrite_config_file()(which only ever persistsAppConfigfields that are already set), this fills in every field before writing — the sixunshare-*bools via.value_or(true), andlog_levelfrom the actually activespdlog::get_level()(not merely a default for when unset — this also captures an explicit--log-levelpassed alongside-won the same command line, overriding whatever an existing config file's ownlog-levelalready was, sincemain()/parse_args()already applied it in that precedence order by the time this runs) — so a bare-wbootstraps a complete, fully-populated config file for hand-editing, and-wcombined with other flags captures their effective values into it.volumesis left exactly as loaded — an open-ended list with no "default" entry to materialize. Prints the config file's full path (write_config_file()already creates the parent directory and the file itself if missing, so no separate existence check is needed here).run_container()(theMode::rundispatch case) resolves each-vspec (erroring out,ok = false, same as a failed--userresolution —bwrapis skipped but unmount/cleanup still runs) into aResolvedVolumeMount, rejecting a duplicate or non-absolute container path first, and passes the resolved list torun_bwrap().--hostname <name>is likewise threaded straight throughrun_container()intorun_bwrap()/build_bwrap_args()(bwrap.{h,cpp}) — see there for how/when it actually takes effect.run_container()also derives acontainer_namefor the session-tracking pid file (seepid_file.{h,cpp}below):read_image_ref()(oci_image.{h,cpp}) applied to the single image tar being run, formatted asname:tag, falling back to the tar's own filename stem ifread_image_ref()can't determine one — passed through torun_bwrap()alongside everything else.--env/--env-file's already-orderedenv_specs(seecli_args.{h,cpp}above) are resolved here, once, viaresolve_env_specs()(env_spec.{h,cpp}, see below) — sameok = false-on-failure pattern as volume/user resolution — and the resolved list is passed torun_bwrap()asextra_env.-D/--daemonize'sdaemonize_flagis also consumed here:run_container()computescontainer_namebeforemount_image()(needed sodaemonize()below can use the real container name for the log file from its very first line, not just after a later rename) and, if daemonizing, callsdaemonize(container_name)(daemonize.{h,cpp}, see below) immediately after: a returned value means this is the original (parent) process (or a hard daemonize failure) — print it andreturnright away;nulloptmeans this is the now-detached child, which falls through into the rest ofrun_container()'s existing body completely unchanged, including the unmount/cleanup that already runs afterrun_bwrap()returns (no separate watcher/reaper — the daemonized child is what runs the whole session, start to finish). -
self_test.{h,cpp}—run_self_tests()implements-t/--test, this project's own built-in self-test mode (distinct from the Meson-driven fixture smoke test undertests/, described in "Build & test commands" below; previously reporteddetect_bwrap_unshare_args()'s output —bwrap.{h,cpp}— unplugged since that's kernel-capability diagnostics, not a test). Currently exercisespersistent_netns.{h,cpp}'s (see below) create/verify/remove cycle: skipped with a message (not a failure) when not root, sincecreate_persistent_netns()requires it for the bind mount. Confirms the namespace is missing before creation, exists right after (checked from this process, after the forked child that actually did theunshare()/bind-mount has already exited — the actual claim being tested: the namespace outlives its creating process), then gone again after removal. Deliberately its own small file since more real tests are expected here as more of the networking feature lands. -
env_spec.{h,cpp}—resolve_env_specs()turns an ordered list ofEnvSpec {is_file, value}(seecli_args.{h,cpp}above) into a flat, ordered list of(key, value)pairs. A literal (--env) is split at its first=(the value may itself contain=; the key must be non-empty). A file (--env-file) is read line by line: blank/whitespace-only lines and lines whose first non-whitespace character is#are skipped (comments), with a trailing\rstripped first for CRLF files; every other line is parsed the same way as a literal. Logs a specific error and returnsnullopton the first hard failure (malformed line, empty key, or an unreadable file) — deliberately stops at the first line, not "skip and warn", since an env file with a typo should fail loudly rather than silently omit a variable a container might depend on.build_sandbox_env()(bwrap.cpp, see below) appends the resolved list after its own built-inPATH/HOME/PWD/TERM— no deduplication needed there, sincerun_process_foreground()'s ownsetenv(..., 1)loop already lets the later occurrence in iteration order win for a repeated key, so an explicit--env PATH=...still overrides the default. -
oci_image.{h,cpp}— validates/parses the OCI Image Layout tar (libarchive + nlohmann_json) and extracts layer blobs.list_oci_images()scans a directory (non-recursively) for*.tar/*.tar.*files and, for each valid OCI archive, derives an image name/tag viaread_image_ref()from itsindex.jsonmanifest annotations (io.containerd.image.namepreferred, elseorg.opencontainers.image.ref.name), falling back to the archive's filename and"latest"respectively.read_image_ref()is public (not just an internal helper oflist_oci_images()) precisely sorun_container()(commands.cpp) can reuse the exact same logic to name a single image tar's session pid file (seepid_file.{h,cpp}below) instead of duplicating it.read_oci_image_config()reads the image config blob referenced by the manifest and extractsUser(split on:intoOciImageConfig::user/group),ExposedPorts,Env,Volumes, and the effective default command (Entrypoint ++ Cmd).user/groupand the default command are consumed by-r/--run, and every field is displayed by-i/--inspect(seecommands.{h,cpp}above) —ExposedPorts/Env/Volumesare otherwise still just captured for when networking/volumes are implemented. -
containers_storage.{h,cpp}— wraps thecontainers-storageCLI (import-layer,mount,unmount,layer --json,delete-layer), forcingfuse-overlayfsas the overlaymount_program.cleanup_layer_chain()walks a layer's parent chain (children before parents) deleting each one. -
bwrap.{h,cpp}—detect_bwrap_unshare_args()probes the kernel (via a forkedunshare(2)per namespace type) for which--unshare-xxxflagsbwrapcan actually use;build_bwrap_args()/run_bwrap()assemble and run the sandboxed command.build_bwrap_args()also takes astd::vector<ResolvedVolumeMount>(seevolume_mount.hbelow) and appends one writable--bind <host_directory> <container_path>per entry.wrap_for_root_namespace()isrun_bwrap()'s own nsenter-wrapping logic pulled out into a reusable, exported function — it's also whatvolume_mount.cpp's copy-into-an-empty-volume step uses to reach the image's content when running rootless (see below);run_bwrap()itself now just calls it once on the assembledbwrapargv.build_bwrap_args()/run_bwrap()also take aNamespaceConfig(bwrap.h) — one plainboolfield pernamespace_probesentry (user/ipc/pid/net/uts/cgroup, defaulttrue), resolved byrun_container()(commands.cpp) fromAppConfig's sixglobal.unshare-*keys (config_file.h, see above) once, up front —bwrap.{h,cpp}itself never touchesAppConfig/YAML, only this already-resolved struct. For each flagdetect_bwrap_unshare_args()finds the kernel supports,build_bwrap_args()additionally requires the matchingNamespaceConfigfield to betrue(looked up via a.cpp-localnamespace_policy_enabled()if-chain overnamespace_probes'names) before actually passing it tobwrap— kernel support and policy are separate gates, both must allow a type. This replaced an earlier hardcoded special case that always dropped--unshare-netregardless of policy or kernel support (without any network setup, e.g.slirp4netns, unsharing it just left the sandbox with no network at all) —netnow goes through the exact same policy path as every other type, defaulting to enabled like the rest. This is a deliberate, user-acknowledged transitional behavior change: as of this, a plain-r/--runwith no config file override gets a real network namespace and thus no network access at all, untilslirp4netnsintegration (the next task on this same branch) actually sets one up;global.unshare-net: offrestores the prior no-isolation behavior in the meantime.detect_bwrap_unshare_args()itself is untouched by any of this — still an unfiltered kernel-capability probe, unrelated to policy (no longer surfaced via-t/--test, seeself_test.{h,cpp}below).build_bwrap_args()/run_bwrap()also take an optionalhostname(from--hostname, long-option only): passed through as bwrap's own--hostnameonly when--unshare-utsis actually among the flagsbwrapis being given (bwrap itself refuses--hostnamewithout it) — otherwise logs a warning and leaves the sandbox's hostname alone, since a stock Android kernel in degraded mode may not support a UTS namespace at all. Never requests--unshare-userwhen running as root: root doesn't need a fresh user namespace for privilege, and bwrap's own single-mapping uid/gid setup for one triggers the kernel's unprivileged-userns setgroups() restriction, which showed up as every other supplementary group collapsing to the overflow gid ("nobody") inid, andsuinside the sandbox failing with "can't set groups: Operation not permitted". Because of that, bwrap's own--uid/--gid(which require--unshare-user) aren't usable when running as root either —--user/--groupwork around this: when set,build_bwrap_args()/run_bwrap()bind-mount the separateslocker-lite-priv-drophelper (see below) into the sandbox at a fixed hidden path and route the real command through it as<uid>:<gid> -- <command...>. This only actually works without a user namespace (i.e. running as root) — under--unshare-user, the sandbox's uid map has only one valid entry, so the helper's ownsetuid()fails cleanly there instead of silently doing nothing.run_bwrap()fails fast (returns -1) if the helper can't be found next to this binary when--userwas requested, rather than silently running the command as root.run_bwrap()also takes acontainer_nameand tracks the running session with it: it passes a lambda asrun_process_foreground()'s newon_startcallback (seeprocess.{h,cpp}below) that callscreate_session_lock(container_name, pid)(pid_file.{h,cpp}, see below) the instant the realbwrappid is known, then callsrelease_session_lock()oncerun_process_foreground()returns (covering every exit path — normal, nonzero, or a forwarded-signal exit — since that call always blocks until the child has actually exited).build_bwrap_args()no longer passes--clearenv/--setenvtobwrapitself; instead,build_sandbox_env()builds the sandboxed command's exact environment (PATH,HOME,PWD— hardcoded to"/", matching--chdir's own value; note per bwrap's own man page--clearenvnever actually unsetPWDin the first place, so this isn't a straight port of a prior--setenv— andTERM, only if the host process has one, followed byextra_env— the resolved--env/--env-filelist fromresolve_env_specs()(env_spec.h), appended last so it can override the built-in defaults for the same key) andrun_bwrap()passes it straight torun_process_foreground()'s ownenvoverride (seeprocess.{h,cpp}below). This works becausebwrap(andnsenter, when interposed viawrap_for_root_namespace()) doesn't alter its own inherited environment unless told to, and neither doesslocker-lite-priv-drop(justsetgroups()/setgid()/setuid()/execvp(), no env manipulation) — so controlling it once, at the outermost exec, is sufficient for it to reach the final sandboxed command unchanged. Because that outermost exec now uses this same explicitly-built environment,build_bwrap_args()/wrap_for_root_namespace()resolvebwrap's andnsenter's own argv[0] to an absolute path viafind_in_path()(called from this process's own, unmodified environment, beforefork()) instead of leaving them as bare names — confirmed by direct testing:--env PATH=...used to breakexecvp()'s ability to even locatebwrap/nsenter(bare-name lookup happens in the child, using the already-overridden PATH), not just what the sandboxed command itself sees. With the fix, only the sandboxed command's own lookup is affected by a--env PATH=...override (as expected — same as overridingPATHin any real shell before running a bare command name), andbwrap/nsenterare always found regardless.run_bwrap()also takes an optionalon_bwrap_pid_knowncallback, invoked alongside (not instead of) the session-lock-creation lambda, at the exact sameon_starttiming —-D/--daemonize(daemonize.{h,cpp}, see below) hooks in here viareport_daemon_started()to learn the real pid at the same instant everything else that needs it does, rather than needing its own separate pid-discovery mechanism. That sameon_startlambda also callscreate_session_cgroup()(session_cgroup.h, see below), right alongsidecreate_session_lock(), so--killcan later find every process the session ever starts via its dedicated cgroup;remove_session_cgroup()is called from the same post-run_process_foreground()spotrelease_session_lock()already is. -
priv_drop_helper.cpp→ the separateslocker-lite-priv-dropbinary (its ownexecutable()target inmeson.build, built with-static). Deliberately has zero dependencies on the rest of this project (no fmt/spdlog/etc.) and is fully statically linked: it gets bind-mounted into the container image's own filesystem, which won't haveslocker-lite's own shared library dependencies — a dynamically linked binary bind-mounted that way fails outright ("error while loading shared libraries"), which is exactly what happened before this was split out (the original approach bind-mountedslocker-lite's own — dynamically linked — binary via/proc/self/exeand reexeced it; kept only as a lesson, not as working code). Usage:slocker-lite-priv-drop <uid>:<gid> -- <command> [args...]; doessetgroups(0,…)→setgid()→setuid()→execvp(), in that order (dropping the group needsCAP_SETGID, which is lost oncesetuid()drops root).find_priv_drop_helper()and thepriv_drop::path/priv_drop::helper_nameconstants live inbwrap.h(not just internal tobwrap.cpp) specifically soexec_in_session()(exec_session.cpp, see below) can reuse the exact same already-bind-mounted helper for-x/--exec's own--user/--groupsupport, instead of a second copy needing to be bind-mounted for it (which wouldn't even be possible —-x/--execjoins an already-running session's mount namespace, it doesn't get to add bind mounts to it).find_priv_drop_helper()itself only checks this binary's own host-side existence; it says nothing about whether a given session actually has it bind-mounted (only true when that session's-r/--runresolved a user in the first place). -
user_spec.{h,cpp}—resolve_user_and_group()resolves a user/group spec (each a name or numeric id) against the container's own/etc/passwd//etc/groupcontent (not the host's, and not a path — callers own reading it, since the two current callers get that content two different ways:run_container()reads it directly off the merged mount path, whileexec_in_session()fetches it overnsenter, since a running session's mount namespace isn't otherwise reachable from this process — see below).nulloptcontent for either file means "unreadable/absent"; a numeric user with no group still resolves fine without it (defaults gid to the same numeric value as the uid) but a named one doesn't.ResolvedUser(bwrap.h) also carrieshome, looked up by the final resolved uid's/etc/passwdentry (field 5) regardless of whetheruserwas given as a name or a number; falls back to"/root"for uid 0 or"/"otherwise when there's no matching row.build_sandbox_env()(bwrap.cpp) sets the sandboxed process'sHOMEfrom this —"/root"only when no user override applies at all (no--user, no image-declaredconfig.User).run_container()(commands.cpp) callsresolve_user_and_group()with either the explicit--user/--groupflags, or, when--userwasn't given, the image's own declaredconfig.User(OciImageConfig::user/group) — so a container defaults to running as whatever user the image itself declares, not root, unless the image declares none. -
process.{h,cpp}— argv-based subprocess helpers (fork/execvp, no shell):run_process()captures stdout (used forcontainers-storagecalls),run_process_foreground()inherits all of stdio (used for the interactivebwraprun). Alsofind_in_path(), a shared$PATHlookup.run_process_foreground()installs a SIGINT/SIGTERM handler around itswaitpid()that forwards the signal to the running child and keeps waiting instead of letting the default disposition killslocker-liteitself — without this, Ctrl-C (orkill) during-r'sbwraprun would skiprun_container()'s unmount/cleanup entirely, leaving the layer imported and/or mounted.run_process_foreground()also takes an optionalon_startcallback, invoked with the child's real pid right afterfork()succeeds (before the signal handlers go up and it blocks inwaitpid()) — the only point where that pid is knowable, and still accurate even whenargvitself execs into something else first (e.g.nsenterhanding off to the final command via its own in-placeexecvp()— a pid never changes acrossexec()).run_bwrap()(bwrap.cpp) is the one caller that uses it, for session pid-file tracking (seepid_file.{h,cpp}below).run_process_foreground()also takes an optionalenv(list of key/value pairs): when set, the forked child replaces its entire environment viaclearenv()/setenv()(plain POSIX, not the GNU-onlyexecvpe()— the target platform includes musl) beforeexecvp(), instead of inheriting this process's own.nullopt(the default) leaves the child's environment untouched.run_bwrap()is again the one caller that uses this, viabuild_sandbox_env()(bwrap.cpp) — see there. -
pid_file.{h,cpp}— tracks one running-r/--runsession (a livebwrapprocess) as a locked pid file, so an outside process (or a laterslocker-liteinvocation) can tell whether it's still running.sanitize_for_filename()(anything outside[A-Za-z0-9._-]→_, falling back to"container"if that leaves nothing) is exported here (not just.cpp-local) specifically sosession_cgroup.{h,cpp}(see below) can reuse the exact same<name>-<pid>naming rule for its own per-session cgroup directory without drifting from this file's own.xdg_state_dir()($XDG_STATE_HOME/slocker-lite, or the$HOME/.local/state/...fallback) is likewise exported (moved out of this file's own anonymous namespace) sopersistent_netns.{h,cpp}(see below) can resolve its own subdirectory under the same state root without a second, drifting copy of this resolution logic.session_pid_file_path()resolves$XDG_STATE_HOME/slocker-lite/run/<container_name>-<pid>(falling back to$HOME/.local/state/...whenXDG_STATE_HOMEis unset/empty — same resolution pattern asconfig_file_path()below, for state instead of config), sanitizingcontainer_namefirst (anything outside[A-Za-z0-9._-]→_, since an image name/tag can contain/or:).session_log_file_path()is a sibling resolving to$XDG_STATE_HOME/slocker-lite/logs/<container_name>-<pid>.loginstead — same sanitization, same$XDG_STATE_HOME/$HOMEfallback, just a different subdirectory and a.logextension — used bydaemonize.{h,cpp}(see below) for-D/--daemonize's log file.create_session_lock()creates the file (O_CREAT|O_WRONLY|O_TRUNC|O_CLOEXEC, mode 0644 —O_CLOEXECmatters: this fd must never leak into the sandboxed command's own fd table), writes the pid as text, and takes an exclusive, non-blockingflock()on it — held only by that fd, so its lifetime tracksslocker-lite's own process lifetime (released automatically on any exit, including a crash), which lines up withbwrapitself being invoked with--die-with-parent. Any external tool can check liveness the same way: attempt the same exclusive non-blockingflock()on the file — success means nothing holds it anymore (stale, safe to remove),EWOULDBLOCKmeans a live process still does.release_session_lock()closes the fd (releasing the flock immediately) and removes the file. Every failure path here (can't create the directory/file, can't lock, can't remove) is aspdlog::warn, never fatal — session tracking is best-effort and must never block or fail-r/--runitself.list_sessions()implements--list-processes(commands.cpp'slist_processes_command()): scans the samerun/directory and reports oneSessionInfo {pid, container_name, running}per readable pid file.pidis read from the file's own contents, not parsed from the filename (ambiguous for names that themselves contain-);container_nameis then recovered by stripping that exact-<pid>suffix back off the filename.runningreuses the same liveness check any external tool would do — a non-blocking exclusiveflock()that succeeds means the file is actually stale, sorunningis false in that case; the lock is always released again immediately either way, never left held by the check itself. A file that can't be opened or doesn't parse as a pid (e.g. removed mid-scan) is silently skipped, not reported as an error — scanning a live directory is inherently racy. Bothlist_sessions()andclean_stale_sessions()(the latter implements--clean-processes) share a privateopen_session_file()helper for the open/read-pid/recover-name step.clean_stale_sessions()doesn't just remove whatever a separatelist_sessions()call reported as not running — it re-takes the same non-blockingflock()used to test liveness and holds it across theremove()call itself, per file, so the stale check and the removal stay atomic against a new session starting in the gap between a check and a later removal. Only files it actually removes are reported back (asSessionInfos withrunning=false); still-locked (running) files are left untouched and not reported. -
persistent_netns.{h,cpp}— generic, narrow infrastructure for keeping a network namespace alive with no process in it, the wayip netns adddoes; nointern/externpolicy or bridge logic here (that's a later commit,network_bridge.{h,cpp}, perdocs/networking-design.md's commit sequence), and not yet wired into-n/--networkat all.persistent_netns_path()resolvesxdg_state_dir() / "netns" / sanitize_for_filename(name)(pid_file.h, see above).persistent_netns_exists()checks whether that path is actually a live bind-mounted namespace, not just a stale/never- mounted file:stat()s the path and its parent directory and comparesst_dev— a genuine bind mount always has a different device number than its parent, the same "is this a mountpoint" technique used elsewhere. Never needs root itself (juststat()).create_persistent_netns()forks a child (never touches the caller's own network namespace —unshare(2)affects only the calling process) thatunshare(CLONE_NEWNET)s its own fresh namespace, bind-mounts its/proc/self/ns/netonto the target path, then exits immediately — the bind mount itself is what keeps the namespace alive from then on, independent of the now-exited child, exactlyip netns add's own technique. RequiresCAP_SYS_ADMIN(root) for the bind mount, matching this feature's current root-only scope (seedocs/networking-design.md) — best-effort like this project's other host-state primitives (session locks, cgroups): logs and returnsfalseon any failure (already exists, fork/unshare/mount failure) rather than throwing.remove_persistent_netns()unmounts then removes the file. -
session_cgroup.{h,cpp}— gives--kill(kill_session.{h,cpp}, see below) a reliable way to find every process a session ever started, however deeply forked/daemonized/reparented, by putting it in a dedicated cgroup v2 group from the moment it starts.cgroup_v2_available()checks for/sys/fs/cgroup/cgroup.controllers(the same signal systemd's own unified-hierarchy detection uses) — only cgroup v2 is supported; v1 (which splits per-controller into separate hierarchies with no unifiedcgroup.procsat the top) is deliberately out of scope, since this project's real target (Android) has used the unified v2 hierarchy by default since Android 12.session_cgroup_path()is deterministic —/sys/fs/cgroup/slocker-lite/<name>-<pid>/, reusingpid_file.h's ownsanitize_for_filename()— so no separate lookup state is needed anywhere.create_session_cgroup()is called fromrun_bwrap()'son_startcallback (bwrap.cpp, see below), the same spotcreate_session_lock()already fires from:create_directories()'s the leaf directory (also creating theslocker-lite/parent the first time — a plain grouping cgroup, no resource controllers are ever enabled on it viacgroup.subtree_control, so the "no internal processes" restriction that comes with actually delegating controllers never applies here) and writes the bwrap pid into itscgroup.procs. From that point on, every process bwrap (or anything it execs into) forks inherits this cgroup automatically, permanently — including anything that later daemonizes/double-forks and gets reparented, unlike pid-namespace child membership (only the processesclone()itself creates) or process-group membership (many daemonizing services explicitlysetpgid()/setsid()away from it on purpose). Best-effort, mirroringcreate_session_lock(): returnsnullopt(logging a warning, never fatal) if cgroup v2 isn't available, or the directory can't be created/written (no delegated subtree when running rootless, or an SELinux policy blocking cgroupfs writes even for a root-euid process, are both real, confirmed-by- testing causes on the two environments this project actually runs on).remove_session_cgroup()(called from the same post-run_process_foreground()spotrelease_session_lock()already is) only succeeds once the cgroup is empty — a straggler process still alive at normal exit (e.g. a daemonized process that outlived the session's own main command, a pre-existing exposure independent of this feature) leaves it in place with a warning, not a fatal error.session_cgroup_pids()readscgroup.procs— this is the actual answer to "gather every process running inside the container": unlike anything derived from/procparent-pid chains or pid namespaces, cgroup membership reliably includes every process the session ever started.session_cgroup_supports_kill()/kill_session_cgroup()wrap thecgroup.killknob (Linux 5.14+): writing"1"to it atomicallySIGKILLs every process currently in the cgroup in one step. -
sandbox_process.{h,cpp}— process-tree/namespace-resolution utilities shared byexec_session.{h,cpp}andkill_session.{h,cpp}(see both below); pulled into their own file (rather than staying private toexec_session.cpp, whereresolve_namespace_pid()originally lived) once--killneeded the exact same "find the real sandboxed child" logic, to avoid a second, drifting copy.resolve_namespace_pid()is unchanged from its originalexec_session.cppform (see that entry for the full reasoning: bwrap's own outer/tracked pid never actually enters the pid/uts/ipc/cgroup namespaces it creates for its clone()'d child, only that child does). Two new utilities added alongside it forkill_session():pid_namespace_isolated(outer_pid, ns_pid)compares/proc/<outer_pid>/ns/pidand/proc/<ns_pid>/ns/pid's ownreadlink()targets directly — true only when bwrap'sclone()actually created a separate pid namespace for its child (--unshare-pidwas requested and the kernel supported it), the precondition for the kernel's own guarantee that killing a pid namespace's pid 1 forcibly tears down every remaining process in it.collect_descendant_pids(root)generalizesresolve_namespace_pid()'s own/proc/<n>/statppid-scanning fallback to collect a whole transitive tree (root included) instead of just one child, sharing the actual stat-parsing loop between both via a privatebuild_ppid_map()(one/procpass, used by both the single-child lookup and the full-tree collection). Reliable specifically whenrootis a genuinely isolated pid namespace's own pid 1: anything that reparents within it (e.g. a daemonizing service) is guaranteed by the kernel to land back onrootitself, unlike on a kernel without pid namespace support, where it escapes to the host's real pid 1 instead (seekill_session.{h,cpp}below for exactly this scenario, confirmed on a real target device). -
daemonize.{h,cpp}— implements-D/--daemonize's fork/detach mechanics.daemonize(container_name)sets up apipe2(..., O_CLOEXEC)pair (so it never leaks intobwrap/the sandboxed command, same reasoning as the pid file's ownO_CLOEXEC) andfork()s. The child callssetsid()— deliberately here, not via re-adding bwrap's own--new-session(removed earlier, see thebuild_bwrap_args()comment):--new-sessiononly callssetsid()for the deeply-nested sandboxed command inside bwrap's own namespace setup, leaving the outerbwrap/nsenter/slocker-liteprocesses still attached to the original session and still receiving its signals (e.g. aSIGHUPwhen the controlling terminal closes) — not real daemonization. Callingsetsid()in our own forked child, before it execs intonsenter/bwrap, detaches the entire chain at once, sinceexec()never changes session membership — confirmed by direct testing (ps -o pid,sid,pgid,tty): the daemon child becomes its own session leader with no controlling tty, andbwrap(a later descendant) shares that same session, also with no tty. The child alsosigaction()sSIGHUPtoSIG_IGN(survives the laterexec()intonsenter/bwrap, unlike a real handler, whichexec()resets to default — confirmed by sendingSIGHUPdirectly to a running daemonizedbwrappid and it staying alive), then redirects stdin to/dev/nulland stdout/stderr to a log file atsession_log_file_path(container_name, getpid())(pid_file.h) — named after its own pid since the real session pid (bwrap's) isn't known yet. If the log directory/file can't be set up at all, that's a hard failure here (_exit(1)), not best-effort — silently losing the very output--daemonizewas asked to capture would defeat the point of the flag. The child reports"LOG <path>\n"over the pipe immediately (so the parent can show a useful location even on failure) and returnsnulloptto its caller (run_container(),commands.cpp), which then falls through into the rest of that function's existing body completely unchanged — the daemonized child is what runs the whole rest ofrun_container(), including the unmount/ cleanup that already existed afterrun_bwrap()returns; no separate watcher/reaper process exists. The parent blocks reading the pipe until EOF, returning the accumulated"LOG "/"PID "lines as aDaemonizeResult— the caller then prints it and exits immediately without running any session logic itself.report_daemon_started(container_name, pid)(called fromrun_bwrap()'s newon_bwrap_pid_knowncallback — seebwrap.{h,cpp}below — the instant the realbwrappid is known) renames the pid-named log file to<container_name>-<pid>.log, re-reports the updated"LOG "line (a real bug caught by testing: the parent's first"LOG "line names the pre-rename, daemon-pid-named path — without a second one, the parent would print a stale filename that doesn't match where the file actually ends up), then"PID <pid>\n"and closes its own end of the pipe — must happen here, explicitly, rather than waiting for the pipe to close naturally at the end of the (potentially very long) daemon's lifetime, or the parent would block for as long as the session runs instead of returning promptly. The pipe's write fd and the current log path are tracked as private file-scope state indaemonize.cpp(matchingprocess.cpp's owng_foreground_child_pidpattern for "there's only ever one of these per process" runtime state), sincereport_daemon_started()is called later, from a different function, not threaded explicitly through every call in between. -
exec_session.{h,cpp}— implements-x/--exec <pid>: joins an already-running-r/--runsession's namespaces viansenterand runs a command inside it in the foreground.exec_in_session()first confirmspidis a tracked, running session vialist_sessions()(pid_file.h) — same liveness check--list-processes/--clean-processesalready use, no new logic needed there. Key discovery, confirmed by direct testing, not assumed:pid(the onerun_process_foreground()captured and pid-file-tracked when-rlaunchedbwrap) is bwrap's own outer process — it sets up the mount and user namespaces itself, thenclone()s the actual sandboxed command into fresh pid/uts/ipc/cgroup namespaces, andclone()'s namespace-creation flags only ever affect the newly created child, never the caller. So the outer process itself never actually enters those namespaces — comparing/proc/<outer_pid>/ns/{pid,uts,ipc,cgroup}against this process's own showed them identical, while onlymnt/userdiffered.resolve_namespace_pid()(sandbox_process.{h,cpp}— moved out of this file once--killneeded the exact same logic, see that entry) finds that real inner process so this can join its namespaces instead. For each of{mnt→--mount, uts→--uts, ipc→--ipc, pid→--pid, cgroup→--cgroup, user→--user}(netdeliberately excluded — this project never isolates networking, seebwrap.cppbelow),readlink()s both/proc/<ns_pid>/ns/<type>and/proc/self/ns/<type>and only passes nsenter's corresponding--type=/proc/<ns_pid>/ns/<type>flag when they differ — an identical-namespace re-entry attempt can fail outright (setns()'s ownEINVALrestriction on re-entering a namespace you're already in), so skipping is deliberate, not just an optimization.mntis the one type where a read failure (permission denied, or the process vanished) is treated as fatal, since without it "joining the container" is meaningless; every other type just degrades to a skip. Always appends--preserve-credentials: without it,nsenter --useralso tries tosetuid()/setgid()/setgroups()to the target's identity within the new user namespace, which fails outright (setgroups failed: Operation not permitted) against thesetgroups-denied unprivileged user namespace bwrap creates whenever-r/--runisn't root — confirmed by hitting this exact failure during manual testing before adding the flag. Runs the finalnsenter ... -- <command>via the existingrun_process_foreground()(process.h) — same inherited stdio and SIGINT/SIGTERM forwarding as every other foreground external command, no new process-running logic needed.exec_in_session()also takes optional--user/--group(mirroring-r/--run's own): given, they resolve against the session's own/etc/passwd//etc/group(fetched viacatrun through the samensenterjoin, since this process can't otherwise see into that namespace, then handed toresolve_user_and_group()—user_spec.h, see below); if unset, defaults to whatever uid/gid the session's own sandboxed command is already running as (read from/proc/<ns_pid>/status), rather than root/the caller — fixing a real bug (reported after this project's own-x/--execand priv-drop features had both shipped separately): without this,-x/--execalways ran as whatever the host invocation was, ignoring any--user/--groupthe session itself was started with. Either way, the resolved identity is applied by runningcommandthrough the session's already bind-mountedslocker-lite-priv-drophelper (priv_drop::path,bwrap.h) — reused as-is, not bind-mounted again (-x/--execcan't add bind mounts to an already-running session's namespace anyway). Skipped entirely when the resolved uid and gid are both 0: a session that was never given a resolvable user at-r/--runtime never got the helper bind-mounted at all, and dropping to 0:0 would be a no-op regardless; a missing helper for a genuinely non-root resolution instead surfaces asnsenter's own "No such file or directory" once it tries to execpriv_drop::path, diagnostic enough on its own. Second real bug, caught by direct testing on a rootless dev machine before this shipped: when the session's own-r/--runused--unshare-user(i.e. ran rootless — see the root-vs-rootless paragraph below), "root inside the container" is achieved purely through the kernel's own uid mapping for that namespace, not a real privilege drop — so/proc/<ns_pid>/status's uid/gid, read from outside that namespace, shows the host-mapped id (e.g.1000), not the container-relative one (0). Treating that as "needs a priv-drop to 1000" is wrong two ways: the helper is typically never bind-mounted for a session with no resolved--user, and even when it is,setuid()fails outright under the single-entry uid map an unprivileged user namespace gets (confirmed directly:failed to drop privileges to 0:0: Operation not permitted). Fixed by tracking whether theusernamespace type was actually one of the ones joined (it only is when it differs from this process's own, i.e. exactly when-r/--runused--unshare-user) and, when so, leaving the default identity unresolved (no priv-drop) for that case — joining that same user namespace with--preserve-credentials(already done regardless) already reproduces the container's own view correctly via that same kernel mapping, with nothing further needed. Verified end-to-end on this same rootless dev machine: a daemonized-r --runbusybox session with no declared user,--exec'd with no--user, now correctly showsuid=0(root)(previously would have attempted, and failed, a priv-drop to the host-mapped uid); an explicit--exec --user 0against the same session correctly resolves to0:0and skips the priv-drop step;--exec --user portageagainst it correctly resolves the name to its real250:250via the fetched/etc/passwdand then fails clearly (helper not bind-mounted, since the session itself had no declared user) rather than silently running as the wrong identity. -
kill_session.{h,cpp}— implements--kill <pid>, stopping a tracked, running-r/--runsession and everything it started.kill_session()validatespidthe same wayexec_in_session()does (vialist_sessions(),pid_file.h). Real bug reported by the user against their own actual target device, confirmed via a captured session log: a plainkill <tracked_bwrap_pid>doesn't kill everything a container started — their/initscriptphp-fpm --daemonizes (double-forks, detaches) thenexec caddy ...s (replaces itself); after killing the tracked pid, bothcaddyand thephp-fpmmaster+workers kept running as orphans. Root cause: on that device,bwrap's--unshare-pidisn't actually in effect at all —detect_bwrap_unshare_args()(bwrap.cpp) only requests--unshare-xxxflags the kernel actually supports, and that kernel doesn't support pid namespaces (independently confirmed elsewhere this session, seeexec_session.{h,cpp}'s ownCONFIG_CHECKPOINT_RESTOREbug above) — sophp-fpm --daemonizereparents to the host's own pid 1, completely disconnected from the sandboxed session; the classic "kill a pid namespace's pid 1, the kernel guarantees the whole namespace collapses" trick simply doesn't apply there. Per the user's own explicit request (they want to choose the mechanism per host capability, and may need an even more basic one later for some hypothetical older device),kill_session()picks between three independently-named strategies, selected dynamically per session (not a single cached host-wide capability flag, since e.g. cgroup creation can fail for session-specific reasons like permissions even on a host that generally supports cgroups) — each runs its own completeSIGTERM→ wait-up-to-grace_period_seconds(10s default, no CLI flag) → forced-SIGKILLescalation internally, with no cross-strategy fallback-after-failure chaining:kill_via_cgroup()— preferred whenever the session has a non-empty dedicated cgroup (session_cgroup_pids(),session_cgroup.h):SIGTERMto every pid currently in it, and, if forcing is needed, either the atomiccgroup.killknob or a fresh re-read-and-SIGKILLsweep (fresh, not the original snapshot, since a process could have forked a new child after the graceful sweep but before dying). The only mechanism that reliably reaches every process regardless of pid namespace support. Critical correctness point, caught during design review before this shipped: the "is it stopped yet" poll must gate on the cgroup being empty, notlist_sessions()'s running flag — that flag only reflects the pid file's flock, released the moment the tracked outerbwrappid exits, and theSIGTERMsweep necessarily hitsbwrapitself too (it's a cgroup member) —bwrapdies and gets reaped in well under a second, long before slower descendants (caddyshutting down gracefully,php-fpmfinishing in-flight requests) actually exit. Gating on the pid file instead would make the poll resolve "done" almost immediately, the forced-kill step would never run, and the original bug would reproduce with unused machinery around it.kill_via_pid_namespace()— used when no cgroup exists for the session, butresolve_namespace_pid()/pid_namespace_isolated()(sandbox_process.h) confirm--unshare-pidwas genuinely in effect for it.SIGTERMscollect_descendant_pids(ns_pid)(reliable here specifically because reparenting within a genuinely isolated pid namespace always lands back on that namespace's own pid 1); if forcing is needed, a singleSIGKILLtons_pidalone is guaranteed complete by the kernel itself, independent of whatever the graceful sweep missed. "Stopped" is simplykill(ns_pid, 0)failing withESRCH. Verified end-to-end on this project's rootless dev machine (which does support pid namespaces, unlike the user's real target device): a daemonized busybox session runningsh -c 'sleep 300 & exec sleep 300'(mirroring the daemonize-then-exec shape of the original bug) was fully cleaned up by--kill, including the backgrounded child, with no leftover processes, mounts, or layers; a second run usingsh -c 'trap "" TERM; sleep 300'(ignoringSIGTERMentirely) confirmed the forced-SIGKILLescalation path too, taking the full 10s grace period before the pid namespace's own collapse-on-kill guarantee cleaned it up regardless.kill_via_tracked_pid()— fallback when neither of the above applies: signals the trackedbwrappid directly,SIGTERMthenSIGKILL, pollinglist_sessions()for "stopped" since that's the only signal available without a cgroup or an isolated pid namespace to check directly. Exactly today's manual-killbehavior — least complete, but always available, and strictly no worse than before this feature existed. This is the path the user's own real target device actually takes today (no cgroup delegation confirmed working there yet; no pid namespace support at all) — a future, even more basic strategy (for some hypothetical still-more-limited device) would slot in here the same way, per the user's own explicit request to keep this extensible.
poll_until()/sleep_ms()(.cpp-local) usenanosleep()in anEINTR-retry loop — matching this project's existing direct-POSIX style (process.cppalready retrieswaitpid()the same way) — rather than<thread>/<chrono>(unused anywhere else in this project). -
config_file.{h,cpp}—load_config_file()reads and parses (via libyaml's document API,<yaml.h>) theglobal,volumes, andnetworkssections of the local YAML config file located byconfig_file_path()($XDG_CONFIG_HOME/slocker-lite/config.yaml, falling back to$HOME/.config/slocker-lite/config.yaml). Supportedglobalkeys:log-level, and sixunshare-<type>keys (unshare-user/unshare-ipc/unshare-pid/unshare-net/unshare-uts/unshare-cgroup, one perbwrap.cpp's ownnamespace_probesentry) controlling whether-r/--runrequests each of bwrap's--unshare-xxxflags — every other long option is a one-shot flag, not a setting, so it doesn't belong in a persistent config file. Eachunshare-*key accepts"1"/"on"/"yes"/"true"(enabled) or"0"/"off"/"no"/"false"(disabled), case-insensitively (.cpp-localparse_bool_flag()); an unset key defaults to enabled, and an unrecognized value logs aspdlog::warnand is treated as unset (default enabled) rather than failing the whole config load — consistent with this file's existing forward-compatible/ignore-malformed-entries policy (only malformed YAML syntax is a hard error). A missing file returns a default-constructed (empty)AppConfig, not an error; unknown sections/keys (and malformed individual volume entries) are likewise ignored for forward-compatibility.main()appliesconfig->log_level(via the existingapply_log_level()) right afterspdlog::cfg::load_env_levels()and before parsing CLI options, so an explicit--log-levelon the command line always overwrites it afterward — same precedence pattern already used forSPDLOG_LEVEL.write_config_file()writes the whole file back out (via libyaml's document-building/emitter API, symmetric to the read side) — used by-v/--volume(create_volume_command(),commands.cpp) to persist a newVolumeEntry {name, directory}into thevolumessection, preservingglobal(including any setunshare-*keys, re-serialized as canonical"true"/"false") untouched.VolumeEntry/thevolumessection is a distinct concept fromOciImageConfig::volumes: this is a user-definedname -> host directorymapping created via-v/--volume, not an image's own declared mount points (still unconsumed, seeoci_image.{h,cpp}above).AppConfig/config.volumesis looked up by name inresolve_volume_mount()(volume_mount.{h,cpp}, see below), which is how-r/--run's own-vusage finds a named volume's host directory.run_container()(commands.cpp) resolves the sixunshare-*fields (eachvalue_or(true)) into aNamespaceConfig(bwrap.h, see below) once, up front, and passes it torun_bwrap()—bwrap.{h,cpp}itself has no dependency on this file or on YAML parsing at all, only on the already-resolved, defaults-applied struct.networkssection (seedocs/networking-design.mdfor the full feature): unlikevolumes(a flatname -> directoryscalar mapping), each network entry is itself a nested mapping (kind/subnet/ipv6/subnet6), since one network needs more than a single value to describe.NetworkEntry(kindisNetworkKind::extern_/intern— trailing underscore onextern_sinceexternis a reserved C++ keyword and can't be an enumerator name — parsed from the YAML strings"extern"/"intern") round-trips throughAppConfig::networksthe same wayVolumeEntrydoes; an entry with an unrecognizedkind(or missingkind/subnet) is skipped on load, same forward-compatible policy as everything else here.ipv6reusesparse_bool_flag(), defaulting totrue(enabled) if absent or unparseable;subnet6is only read/written whenipv6is true.write_config_file()writes each network as its own nested mapping undernetworks,ipv6re-serialized as canonical"true"/"false"like theunshare-*keys. -
network_subnet.{h,cpp}— pure CIDR arithmetic backing-n/--network's subnet allocation, no kernel/ip/iptablescalls (those come in a later commit perdocs/networking-design.md's sequence).is_valid_ipv4_cidr()/is_valid_ipv6_cidr()andipv4_cidrs_overlap()/ipv6_cidrs_overlap()all build on one.cpp-localparse_cidr()(viainet_pton(), not hand-rolled parsing) producing a plain byte-vector address (4 bytes for IPv4, 16 for IPv6) + prefix length, and one sharedbytes_overlap()byte/bit-mask comparison generic over that byte length — IPv4 and IPv6 overlap checking are the same algorithm, not two parallel implementations.allocate_ipv4_subnet()/allocate_ipv6_subnet()(commands.cpp'screate_network_command()) scan10.168.<n>.0/24/fd00:168:0:<n>::/64fornin0..255and return the first one that doesn't overlap any existing network's subnet (via the overlap checks above, not just other auto-allocated ones — a manually--subnet-overridden network is checked too). The samenrange for both is deliberate, so the common case (no manual overrides) allocates visibly paired v4/v6 blocks per network — though since IPv6 hextets are hexadecimal,n >= 10renders as a valid but numerically-different-from-naddress (e.g.n=15becomes...:15::/64, which is hex0x15= 21) — purely cosmetic, allocation correctness doesn't depend on the two matching numerically. -
volume_mount.{h,cpp}—is_valid_volume_name()(no/, checked by bothcreate_volume_command()and to tell a-vspec's name/path apart) andresolve_volume_mount(), called once per-voccurrence fromrun_container()when running with-r. A spec with no/is looked up inconfig.volumesby name (error if unknown); one with/is treated as a host directory path andcreate_directories()'d if missing. If the resulting host directory is empty,initialize_volume_directory()reconciles it against the image's own directory at the given container path: if that image directory is non-empty, its contents are copied in first; then, whether or not there was content to copy, the host directory's own mode/ownership/timestamps (and xattrs/ACLs where supported) are always set to match the image directory's own — real bug fixed by the user, not assumed: an earlier version only ever copied when the image directory was non-empty, so an image declaring an empty directory with specific ownership/permissions (e.g. a data directory owned by a non-root uid/gid) got a host directory with defaultcreate_directories()permissions instead, and even the non-empty case never reconciled the directory's own attributes (only each copied entry's). The existence check, content copy, and attribute reconciliation all run as a singlesh -cinvocation wrapped throughwrap_for_root_namespace()(bwrap.h) — not a plainstd::filesystemcheck — because a rootlesscontainers-storage mount's content isn't visible to this process at all withoutnsenter, the same constraintrun_bwrap()itself works around (see the root-vs-rootless paragraph below).cp -a --preserve=mode,ownership,timestamps,links[,xattr] --attributes-only -Tdoes the attribute-reconciliation step; whether,xattris included is decided by a directsetxattr()/removexattr()probe on the host directory (no new library dependency — Linux POSIX ACLs are themselves stored as xattrs, so this one probe stands in for both, logging a singlespdlog::warnif unsupported).-T/--no-target-directoryis required on that secondcp— confirmed by direct testing: without it, since the host directory already exists, plaincp SRC DSTcopiesSRCintoDSTas a nestedDST/basename(SRC)subdirectory instead of reconcilingDST's own attributes, which is exactly the bug this fix closes. A nonzerocpexit is only ever a warning, never fatal — often just an ownership-preservation shortfall when not running as root. Verified end-to-end againstimages/gitea.tar's real declared/etc/gitea//var/lib/giteavolumes under a real rootless mount: the resulting host directories' mode/ownership matched the image's own declared values in both cases, and no mounts/layers were left behind afterward.
Errors are logged via spdlog::error; every external command is also traced at debug
level in run_process()/run_process_foreground() (src/process.cpp) — visible via
SPDLOG_LEVEL=debug, since spdlog's default level is info — and a failed external
command additionally logs a spdlog::warn, which is visible by default (no env var
needed). The final "mounted image at: ..." success line is direct stdout program
output, not a log.
Because containers-storage mount runs rootless, it reexecs itself into a private
user+mount namespace to gain the privilege it needs for the overlay mount — which
leaves the result invisible to a plain shell or child process outside that namespace.
Confirmed containers-storage unshare does not rejoin an already-running mount's
namespace; only nsenter targeting the live fuse-overlayfs daemon's PID does.
-r/--run handles this automatically by locating that PID and running bwrap via
nsenter into its namespaces (wrap_for_root_namespace(), src/bwrap.h) — reused
as-is by volume_mount.cpp's copy-into-an-empty-volume step, since that also needs
to read image content that's otherwise invisible outside the same namespace.
Running as root sidesteps all of this: no privilege
reexec is needed, so the mount is already directly visible in the current namespace,
and nsenter --user=... into it then fails ("reassociate to namespace 'ns/user'
failed: Invalid argument") since the caller is already in that same user namespace.
-r/--run detects geteuid() == 0 and skips nsenter automatically in that case;
--no-nsenter forces it off manually for any other situation where the mount turns
out to already be directly visible.
Mutable global state and multi-container support: an audit ahead of planned
docker-compose support (running multiple containers at once) found exactly three
pieces of mutable global/file-scope state in src/: g_mount_program
(containers_storage.{h,cpp}, the resolved fuse-overlayfs path — genuinely
process-wide, invariant across containers), g_foreground_child_pid
(process.cpp, plus run_process_foreground()'s process-wide SIGINT/SIGTERM
handler installation — tracks one foreground child at a time), and
g_report_fd/g_log_path (daemonize.cpp, one in-flight -D/--daemonize
handshake's report-pipe fd and log path). Decision, confirmed by the user:
multi-container/compose support will run each container's session in its own
forked OS process — the same model -D/--daemonize already uses — rather than
one process managing multiple containers concurrently without forking. Under
that model, "one OS process" and "one running container" stay the same thing
they already are today, so none of these globals need to become per-container
state — each forked child only ever tracks/signals one foreground child and
handles one daemonize handshake, exactly as today. This is a load-bearing
constraint for however the compose orchestrator ends up implemented: it must
fork (not thread, not run an in-process event loop over N containers) one child
per service, each child reusing run_container()'s existing single-container
code path unchanged.
Build & test commands
Build directory is buildDir/ (already configured).
- Configure (only needed if
buildDir/is missing or deleted):meson setup buildDir - Build:
meson compile -C buildDir(orninja -C buildDir) — also buildsbuildDir/slocker-lite-priv-drop, the statically-linked helper-r --userneeds (seepriv_drop_helper.cppin "Project state") - Run the executable:
./buildDir/slocker-lite -m <image.tar>(see--helpfor the full flag list:-m/--mount,-r/--run,-u/--umount,-c/--cleanup,-l/--list-images,-i/--inspect,-x/--exec,--kill,--no-nsenter,-D/--daemonize,--user,--group,--hostname,--env,--env-file,-v/--volume,--list-volumes,--delete-volume,--delete-volume-full,-n/--network,--extern,--intern,--subnet,--no-ipv6,--subnet6,--list-networks,--delete-network,--list-processes,--clean-processes,-w/--write-config,-t/--test,--log-level,-h/--help,-V/--version) - Run tests:
meson test -C buildDir
Code style
- Null-pointer checks: prefer
if (!ptr)/if (ptr)overif (ptr == nullptr)/if (ptr != nullptr). - Constants: no
kHungarian-notation prefix.enum classvalues are already qualified by the enum's own name (e.g.Mode::run,OciPortProtocol::tcp), so plain snake_case enumerators are enough on their own. Free-standing constants also use plain snake_case; when several are conceptually related, group them under a namednamespaceinstead of relying on a shared prefix to imply the grouping (e.g.cli_args.cpp'sgetopt_longlong-option codes live innamespace options { constexpr int log_level = ...; }, andbwrap.cpp's priv-drop-helper path/binary-name pair live innamespace priv_drop { ... }) — nest the named namespace inside the file's existing anonymous namespace where one is already present, so internal linkage is unchanged. AkXxx-named identifier that turns out not to actually beconst(mutable global/static state) instead follows this codebase's existingg_prefix convention (e.g.containers_storage.cpp'sg_mount_program, matchingprocess.cpp'sg_foreground_child_pidanddaemonize.cpp'sg_report_fd/g_log_path).
Licensing
- Every
.c/.cpp/.hfile undersrc/must start with the GPLv2-or-later copyright header (see any existing file undersrc/for the exact text). - After adding a new source file under
src/, run./add-license.shfrom the repo root to prepend the header (it readscopyright-headerand inserts it viased, skipping files that already have it, so it's safe to re-run at any time).
Build configuration notes
meson.buildsetswarning_level=3andcpp_std=c++20— keep new code warning-clean under-Wall -Wextra -Wpedantic-equivalent settings.- The single Meson
test()target runsslocker-liteagainst a fixture OCI image tar generated at build time bytests/gen_fixture.py(acustom_target) and checks its exit code (no test framework is wired in yet).