139 Commits

Author SHA1 Message Date
ceamac 4f2c7d4af2 Bump version to 0.2.0
The first version with a genuinely usable feature set (mount/run,
volumes, networking, exec/kill/daemonize, and compose orchestration) --
picking a round number for it rather than continuing to increment from
the original placeholder 0.0.1.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 17:10:44 +00:00
ceamac 586c20f341 Add a man page, built optionally via scdoc
docs/slocker-lite.1.scd covers the full CLI (grouped like README.md's
flag table, including the DEBUGGING section for --mount/--umount/
--cleanup and the reassigned-short-option note), FILES, EXIT STATUS,
and SEE ALSO. meson.build builds it into slocker-lite.1 and installs it
under man1 only when scdoc is found on the host -- configuration still
succeeds without it, matching this project's existing "degrade
gracefully when an optional tool is missing" policy. Verified the
rendered page with `man --warnings` (clean, no troff warnings) and a
DESTDIR install.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 17:08:08 +00:00
ceamac b9d79baf74 Update README for compose support, networking, and reassigned short flags
Documents -u/--up, -d/--down, --list-containers, and the new Compose
support section; mentions dnsmasq/ip/iptables/sysctl as runtime deps;
notes that -m/-u/-c no longer mean --mount/--umount/--cleanup and moves
those debug-only commands to the end of the flag docs; refreshes Status,
Configuration, and How it works to reflect networking/compose support;
fixes two stale -m/-e examples that no longer match the actual CLI.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 16:59:10 +00:00
ceamac 083d358dae Document the two real device bugs found running the compose test on-device
Both fixes (is_network_in_use()'s ip-exit-code misread, and the
kill_session()-vs-async-cleanup race) plus the INFO()-based diagnostics
that made finding them possible are now documented in
test_compose_orchestrator.cpp's own CLAUDE.md entry, alongside
confirmation that the full [integration][root][net] suite (73 assertions,
14 cases) passes cleanly on the real Android target device.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 15:19:18 +00:00
ceamac 5bec5c5e53 Fix a real race: -d checked is_network_in_use() before async cleanup ran
Even after fixing is_network_in_use()'s own exit-code bug, the real device
still failed the same way: both managed networks reported "still in use"
milliseconds after their only session was confirmed stopped. Root cause:
kill_session() only waits for the *sandboxed process itself* (via its
cgroup) to die -- it says nothing about the separate, independently
scheduled daemonized process that started it, which still has its own
post-waitpid() cleanup left to run (stop_tap_relay() among it) before a
tap+relay join's host-side device is actually detached from the bridge.
Checking is_network_in_use() the instant kill_session() returns can
genuinely still see it attached; on the real device, the kill+check pair
landed single-digit milliseconds apart in the log.

network_becomes_unused() retries is_network_in_use() for up to 10s
(nanosleep()-based, matching this project's existing polling style)
instead of giving up on the first still-attached answer.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 14:46:02 +00:00
ceamac d3108845a6 Fix is_network_in_use() misreading the real device's own ip exit code
Found on the real Android target device (root, over SSH): -d/--down left
every managed network standing, reporting each as "still in use" right
after their only attached session had already been confirmed stopped.
Root cause, confirmed by direct on-device inspection: is_network_in_use()
treated any nonzero exit from `ip -o link show master <bridge>` as "the
check itself failed" and failed closed (assumed in use) -- but the
device's own minimal `ip` build exits 1, not 0, for the "bridge exists,
nothing attached" case this project's own dev machine reports as exit 0
with identical (empty) output. The exit code was never a reliable signal
across ip implementations; the output content is.

Fixed by splitting into two checks: first confirm the bridge device itself
is reachable at all (fail closed only if *that* fails or is empty --
namespace/bridge genuinely missing), then decide "in use" purely from
whether the filtered membership query's own output is non-empty,
regardless of its exit code.

Verified: dev machine (root, both with and without the device-mimicking
-c config) still passes cleanly; the real device re-run (queued
separately) will confirm the actual fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 14:17:28 +00:00
ceamac ac4194c96b Surface -u/-d's captured stdout via INFO() on test failure
CapturedStdout swallows compose_up_command()/compose_down_command()'s own
fmt::print()/spdlog output (spdlog's default sink is stdout) so it doesn't
pollute Catch2's console reporting -- but that also meant a failure had no
diagnostic trail at all. Both captures are now kept and attached via
INFO(), which Catch2 only actually prints alongside a failing assertion in
the same scope, staying silent on a normal passing run.

Also softens the port-forward reply REQUIRE to CHECK: a flaky reply on one
run must never skip the -d/--down cleanup below it, which is exactly what
a REQUIRE there would do -- found while investigating a real failure on
the actual target device (managed networks not torn down by -d), where
losing that diagnostic trail (and, if this were interactive, the -d skip
compounding into a real leaked uplink relay) would have made root-causing
much harder.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 13:34:07 +00:00
ceamac 73b9c58b8b Add a full -u/-d integration test against test-compose/compose.yaml
tests/integration/test_compose_orchestrator.cpp ([integration][root][net])
exercises the real, checked-in test-compose/compose.yaml skeleton end to
end: creates test-preexisting-net if needed, -u, confirms both services'
environment/env_file values via their own log files, confirms the server's
port-forward actually relays the worker's reply over a real TCP
connection, -d, confirms managed networks were torn down while the
external network and the managed volume both survived, then deletes the
volume too so a later run can exercise its creation again.

Deliberately doesn't use ScratchXdgDirs (unlike every other [root][net]
test in this suite): test-preexisting-net is external: true, meaning it's
the user's own responsibility to set up once and keep reusing, matching
real Compose's own convention for that field. A daemonized service's own
dispatch_command() call can return before the sandboxed script has
produced any output, so log-content and port-forward checks poll rather
than asserting immediately; the port-forward check connects via the
host's own real global IPv4 (discovered via `ip`, not hardcoded) instead
of 127.0.0.1, since loopback never reaches the container regardless of
whether forwarding itself works (a documented NAT-hairpinning limitation).

Verified end to end on this dev machine (root, via the scoped doas rule),
twice in a row: 27 assertions passed both times, and
--list-containers/--list-networks/--list-volumes/ps all confirmed clean
afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 12:36:49 +00:00
ceamac e9cbe49a5c Add a testing-only image-resolution fallback for -u/--up
resolve_compose_images() now falls back to matching by a tar file's
literal filename (stem, .tar/.tar.<compression> stripped) when no image's
own declared name:tag matches the compose file's image: reference --
logged as a warning, not silent, since it's a deliberately loose
name-shaped guess rather than real image resolution. Added specifically so
a locally built/renamed image (one whose own embedded manifest name:tag
doesn't match its filename at all -- e.g. a musl-built busybox variant
tagged under a different registry path) can still be used for local
testing without editing the compose file to match.

Verified manually: a compose file requesting "busybox:latest" against a
busybox.tar whose own embedded ref is a completely different name:tag now
resolves and starts correctly, with the expected warning logged; the
session ran and exited cleanly with no leftover state.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 12:24:10 +00:00
ceamac 6624fd3ab5 Add user: support to compose services
ComposeService::user/group (compose_file.{h,cpp}) parse a "user[:group]"
key, split on the first ':' the same way an image's own declared USER is
split (oci_image.cpp). start_compose_services() passes both straight
through to run_mounted_container()'s existing user/group parameters, which
already fall back to the image's own declared user when unset -- the same
default -r/--run itself has when --user isn't given.

Verified end to end on the real target machine (root, via the scoped doas
rule), checked via `ps -eo pid,ppid,uid,cmd` (not -x/--exec, see below):
the actual sandboxed command runs as the resolved uid/gid, matching plain
-r --user's own already-working behavior.

Also recorded in TODO.md: verifying this surfaced a real but unrelated
bug in resolve_namespace_pid() (sandbox_process.cpp), which -x/--exec's
own default-identity resolution uses -- it stops at bwrap's own pid-1
namespace supervisor instead of walking one level deeper to the real
(correctly priv-dropped) target, so `-x/--exec <pid> -- id` with no
explicit --user misreports root for a session that's actually running as
a non-root user the whole time. Not a regression from this change and not
fixed here.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 12:16:52 +00:00
ceamac 129e4d89cd Wire up port-forwarding for compose services
start_compose_services() was passing {} for run_mounted_container()'s
existing port_forward_specs parameter -- the same one -r/--run's own -p
already uses. Each service's already-parsed ComposeService::ports is now
reserialized back into the raw "<host-port>:<container-port>[/tcp|udp]"
strings that parameter expects, reusing the existing parse-then-resolve
pipeline (including its "no network: prefix means the sole extern network"
default) unchanged.

Verified end to end on the real target machine (root, via the scoped doas
rule): a compose service on an extern network with `ports: - "18099:80"`
running busybox httpd was reachable via curl against the host's real LAN
IP once -u started it, and -d correctly tore down the now-unused network
(including its extern uplink) afterward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:55:37 +00:00
ceamac 8603556d12 -d/--down: tear down unused managed networks, always keep volumes
New is_network_in_use() (network_bridge.{h,cpp}) checks whether anything
still has a live interface on a network's own bridge (ip -o link show
master <bridge>, run in the network's own namespace) -- a veth- or
tap+relay-joined container's host-side end is enslaved to the bridge the
same way regardless of join mechanism, so this catches any attached user,
not just this compose file's own services. Fails closed (assumes in use)
if the check itself can't run.

compose_down_command() now, after stopping every recorded session, loads
the compose file again (best-effort -- skipped, not an error, if it's
gone) to find each managed network's project-prefixed name and tear it
down via the newly-exported delete_network_command(), but only when
is_network_in_use() confirms nothing is still attached. external: true
networks and all volumes are never touched, matching real `docker compose
down`'s own default plus the user's own explicit direction that volumes
should always persist.

Verified end to end on the real target machine (root, via the scoped doas
rule): a network is correctly deleted once its only service stops, but
correctly left alone when an unrelated ad-hoc -r --run session is also
still joined to it -- until that session is killed by hand.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:42:54 +00:00
ceamac 8b33afe03b Implement -d/--down, taking only an optional compose file name
Splits -u/-d's previously-shared CLI parsing: -u/--up keeps its two-token
(required images directory + optional compose file) shape, while -d/--down
becomes its own no_argument option with a single manually-peeked optional
trailing token -- no images directory needed at all, since stopping a
stack doesn't mount or resolve any image.

stop_compose_services() (compose_orchestrator.{h,cpp}) reads the state
file a previous -u/--up wrote for the same compose path, stops every
recorded session via kill_session() (the same graceful mechanism --kill
already uses), then removes the state file. If the compose file still
exists and parses, each service's own stop_grace_period_seconds (parsed
since compose_file.cpp's first commit but unused until now) is honored as
that service's own grace period instead of kill_session()'s 10s default,
falling back to it otherwise -- the state file alone already has
everything strictly required.

Known scope limitation: doesn't tear down the compose file's own managed
networks/volumes, matching real `docker compose down`'s own default.

Verified manually end to end: -u followed by a bare -d correctly finds and
stops both services via the recorded state file, removes it, and leaves no
processes behind; -d against a compose file with nothing recorded exits
cleanly (not an error).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:27:20 +00:00
ceamac 24de80e058 Add --list-containers to list compose-started containers
list_compose_containers() (compose_orchestrator.{h,cpp}) scans every
compose state file under xdg_state_dir()/"compose"/ and reports one row
per recorded service: compose file, service, container name, pid, and a
real liveness status cross-referenced against list_sessions() -- not just
"this line exists in the state file". record_compose_services() now also
writes the compose file's own real path as a header line, since the state
file's own name only encodes a lossy, sanitized version of it.

commands.cpp's own pad_column() factors out the per-column tab-alignment
scheme every other list command in this file already duplicates inline,
since this one needs it across four columns.

Verified manually: --list-containers against a real running 2-service
compose stack shows both running with correct fields; killing one flips
just that row to exited on a re-run, confirming real liveness checking.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:17:58 +00:00
ceamac 778bc9c130 Document the compose orchestrator in CLAUDE.md
Covers all five -u/--up steps (compose_orchestrator.{h,cpp}), the
run_container()/run_mounted_container() split, and the newly-exported
commands.cpp functions the orchestrator reuses.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:08:26 +00:00
ceamac c80ebb4e28 compose orchestrator step 5: record started services' pids for -d/--down
compose_state_file_path()/record_compose_services() write one line per
started service ("<service_name> <container_name> <pid>") to
$XDG_STATE_HOME/slocker-lite/compose/<sanitized-absolute-compose-path>,
keyed by the compose file's own resolved path -- the same state-file
naming pattern port_forward.h's/network_tap_relay.h's own crash-orphan
records already use, so a future -d/--down implementation can find exactly
which sessions a given -u/--up started. Overwrites any previous run's own
record for the same file, a known limitation until -d/--down itself exists
to keep the two in sync.

Verified manually end to end: -u against a 2-service, depends_on-linked
compose file mounts, provisions, starts both daemonized in order, and
writes a correct state file; --kill against the recorded pids leaves no
processes behind.

This completes the -u/--up sequence requested (collect images, mount all
before starting any, provision networks/volumes, start in dependency
order backgrounded, record pids). Not yet implemented: -d/--down itself
(still a stub), port-forwarding/--user/--group for compose services, and
re-verification of network provisioning specifically against real root
(the volume/single-network-free path was verified rootless; network
provisioning reuses already-proven create_network_command()/
ensure_network_provisioned() as-is).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:07:22 +00:00
ceamac c93160acb4 compose orchestrator step 4: start services in dependency order, daemonized
start_compose_services() topologically orders services by depends_on
(Kahn's algorithm -- load_compose_file() already guarantees the graph is
acyclic) and starts each one daemonized (as if -D/--daemonize had been
given), against its own already-mounted image, provisioned networks
(project-prefixed names resolved via the mapping from step 3), and
provisioned named volumes. A service whose dependency failed to start (or
was itself skipped) is skipped too, logged clearly, and never started
against a dependency that isn't actually running.

A service's session identity (pid-file/log/cgroup naming) is its explicit
container_name: if given, else "<project>_<service>"; its DNS hostname
(what sibling services resolve it by) is instead its explicit
container_name: or bare service name, never project-prefixed -- matching
real Compose's own service-name resolution. Port-forwarding and explicit
--user/--group aren't wired up for compose services yet.

Verified manually (rootless): a 2-service compose file with a depends_on
edge starts both daemonized, in order, with correct --hostname each;
--list-processes/`ps` confirm both genuinely running; --kill cleanly stops
both with no leftover processes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:06:16 +00:00
ceamac 9d1f9cb88b Extract run_mounted_container() out of run_container() (pure refactor)
run_container() (-r/--run) still decides container_name, handles the
-D/--daemonize fork (which must happen before mounting so the log file can
be named from its first line), and calls mount_image() itself. Everything
after that -- volume/env/user resolution, namespace policy, network/
port-forward/DNS setup, running bwrap, and unmount/cleanup -- moves into a
new, exported run_mounted_container(), taking an already-mounted image
instead of mounting its own. No behavior change for -r/--run itself; this
is prep for the compose orchestrator's own dependency-ordered starting
(next commit), which mounts every service's image up front and needs to
run each one against its own already-mounted image rather than a second
copy of ~200 lines of this logic.

Verified: meson test clean, [integration][net]~[root] clean (exercises
this exact path via test_rootless_run.cpp), and a direct manual `-r
images/busybox.tar -- echo` smoke test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:04:15 +00:00
ceamac 0cc89f9e09 compose orchestrator step 3: provision managed networks and volumes
Exports create_volume_command()/create_network_command() (pure refactor)
so the orchestrator can reuse -v/--volume's and -n/--network's own
provisioning logic directly. New provision_compose_networks_and_volumes()
creates every managed (non-external) network/volume the compose file
declares, project-name-prefixed (compose_project_name(), derived the same
way real Docker Compose derives its own default project name -- the
compose file's parent directory basename) so two different compose
projects can't collide in slocker-lite's single, flat networks/volumes
namespace. Finding one that already exists is deliberately not an error --
reused as-is, per the user's own explicit direction (e.g. left over from an
interrupted previous -u/--up run). external: true networks map to
themselves, unprefixed, since they're real pre-existing networks by
definition.

Verified manually (rootless): volume provisioning creates the
project-prefixed directory/config entry, and a second -u run against the
same compose file correctly reuses it without error. Network provisioning
reuses ensure_network_provisioned()/create_network_command() as-is (no new
logic there) but wasn't separately re-verified live, since it needs root.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:02:33 +00:00
ceamac c1cc564419 compose orchestrator step 2: mount every image before starting anything
mount_compose_images() mounts every resolved image up front, deliberately
before starting any service -- mounting can be slow on some devices, so
this avoids leaving earlier services already running while a later one is
still mounting. On any failure partway through, unmounts+cleans up every
image already mounted in the same call before returning, so a failed
-u/--up never leaves a partial mount set behind.

Verified manually: single and multi-service mounts each produce distinct
layer IDs; the successful path leaves images mounted deliberately, since
the next step's dependency-ordered starting still needs them mounted.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 11:00:01 +00:00
ceamac 0c8e144e3d compose orchestrator step 1: resolve every service's image
Exports MountedImage/mount_image() from commands.cpp's own anonymous
namespace (pure refactor, no behavior change) so the new orchestrator can
reuse them. New src/compose_orchestrator.{h,cpp} adds
resolve_compose_images(), matching each service's image: reference against
list_oci_images() -- the same function -l/--list-images already uses --
before anything is mounted or started, wired into -u/--up.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 10:58:45 +00:00
ceamac f354463b40 Add -u/--up and -d/--down, wired to compose-parser stub commands
Both take a required OCI images directory plus an optional compose file
name (defaulting to "compose.yaml", resolved relative to the cwd) via the
same manual two-token consumption -v/--volume already uses, just with the
second token optional. The stub commands aren't no-ops: they resolve the
images directory, load and validate the compose file through the existing
load_compose_file()/validate_compose_external_state(), and print a summary
-- confirming the CLI wiring and parser work end to end -- before logging
that actual orchestration isn't implemented yet.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 10:45:48 +00:00
ceamac a626ff2cdd Add a unit test for duplicate top-level compose volume names
The check itself already existed in load_compose_file() (same shape as the
already-tested duplicate network name); this was just missing coverage.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 10:45:42 +00:00
ceamac 4ffc68a00e Add validate_compose_external_state() for env_file/external-network checks
Separate from load_compose_file() (which stays pure YAML validation with no
host-state dependency): checks every env_file exists as a readable regular
file, and every network marked external: true already exists in the real
persistent.yaml. Fail-fast, same convention as load_compose_file()'s own
cross-validation. Bind-mount host directories are deliberately not checked
here, since resolve_volume_mount() already auto-creates a missing one.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 09:39:12 +00:00
ceamac 410faff79c Add three more compose-file sanity checks
- depends_on cycles (not just direct self-reference): a three-color DFS
  over the dependency graph reports the actual cycle path.
- Duplicate host-port/protocol across (or within) services: two services
  both publishing the same (host_port, protocol) would only ever leave one
  reachable, even though both DNAT rules would get added later.
- container_name colliding with another service's own implicit name, not
  just two explicit container_names matching each other.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 09:31:10 +00:00
ceamac a6130aa69d Add a Compose YAML parser for the subset of fields slocker-lite supports
load_compose_file() (src/compose_file.{h,cpp}) parses and validates
services/networks/volumes -- unrecognized keys are silently ignored, but a
malformed value for a supported key is a hard parse error, since a Compose
file describes an actual deployment rather than being a version-spanning
settings file. Confirmed no dedicated C++ library for this exists, so it's
hand-written against the already-present libyaml dependency rather than
pulling in the official JSON Schema plus a validator library.

scalar_value()/find_in_mapping() move out of config_file.cpp's own
.cpp-local pair into a new shared src/yaml_util.{h,cpp} (plus a new
sequence_items(), for Compose's list-valued keys) so both files share one
YAML-traversal implementation instead of drifting copies.

tests/unit/test_compose_file.cpp covers every supported field/form and
validation error against small hand-written snippets -- the checked-in
test-compose/compose.yaml skeleton is reserved for later integration tests
once an orchestrator exists, not these unit tests.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 09:20:46 +00:00
ceamac 9f7af0e45a Add a skeleton test-compose.yaml for docker-compose parsing work
Two busybox services (a worker echoing a fixed reply, a server that relays
it) exercising the compose fields we intend to support: services/image/
command, container_name, environment/env_file, depends_on (condition only),
stop_grace_period, networks (internal/external), ports, and bind + named
volumes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-07 08:58:46 +00:00
ceamac 174b6b4c29 Use explicit test-only subnets and test- prefixed names in the test suite
Running the full device test suite on the real Android target found a
genuine test-harness gap: every extern-network test in
test_network_join_scenarios.cpp failed with "RTNETLINK answers: File
exists" adding its own uplink route. Root cause: each test runs under its
own ScratchXdgDirs, so its own persistent.yaml starts empty every time --
allocate_ipv4_subnet()'s own auto-allocation always picks the very first
slot (10.168.0.0/24) with no way to see whatever real, non-test networks
already exist on the host. The device happens to have a real, long-lived
"extern" network already occupying exactly that subnet from prior manual
testing, and since extern's uplink adds a route back into the (necessarily
shared, host-root) routing table -- unlike intern, whose routing lives
entirely inside its own isolated per-network namespace -- every extern
test collided with it. Not a production bug: a real end user only ever has
one persistent.yaml where allocation correctly sees every existing entry.

Fixed by giving each of the 10 test networks in
test_network_join_scenarios.cpp its own fixed, explicit subnet (--subnet)
in 10.169.0.0/16 -- a different /16 than production's own default
10.168.0.0/16 range, so a test run can't collide with a real network
regardless of how many the host already has.

Also renamed every network/hostname/container-name string literal used
across the test suite (test_network_join_scenarios.cpp,
test_root_networking.cpp, test_rootless_run.cpp,
test_session_cleanup.cpp) from "selftest*" to "test-*", to further reduce
the chance of colliding with anything a real invocation might already be
using. Low-level interface device literals (slkselftest0/thselftest0/
ethselftest in test_root_networking.cpp) are left as-is -- they're
internal identifiers for a throwaway unit test, not network or container
names.

Verified: clean rebuild, meson test, and 3 consecutive
[integration][root] suite runs (61 assertions, 14 test cases) with no
failures.
2026-09-05 14:03:14 +00:00
ceamac 7c33375d4d Fix flaky network-join tests: wait for an assigned address, not just link existence
The tap+relay intern-ping test and the DNS hostname-ping test both
flaked intermittently (roughly 1 in 13-20 runs), always in a "found_success
== false" shape with no assertion-level clue as to why.

Root cause, confirmed via a temporary production-code diagnostic (since
reverted) and cross-referenced against network_join.cpp's own code: the
readiness poll only waited for eth0 to *exist* (`ip link show eth0`), but
an interface can become visible before join_one_network()'s own later
`ip addr add`/`ip link set ... up` steps for it have actually run. A
script that used the interface as soon as it merely existed could ping,
get "Network unreachable" (no address yet), and exit almost immediately
-- and since a session's own sandboxed process is the sole occupant of
its pid/net namespace (--unshare-pid/--unshare-net), its exit destroys
that namespace outright. That, in turn, made whichever other nsenter
call was still in flight against the same namespace -- join_one_network()'s
own remaining steps, or the entirely separate per-session DNS resolver
(network_dns.cpp's start_dns_resolver(), which enters every networked
session's namespace regardless of whether --hostname was given) -- fail
with "No such file or directory" against a namespace that had already
collapsed underneath it.

This is the same general class of race network_join.{h,cpp}'s own
CLAUDE.md entry already documents (a very-short-lived sandboxed command
can outrun its own concurrent network setup), just one step further than
the eth0-existence race already fixed earlier in this file -- a test-code
issue, not a production bug. Fixed by polling for an actually assigned
address on eth0 instead of mere existence, in both wait_for_eth0_then()
and BackgroundPeer's own inline readiness script.

Verified with the diagnostic in place that the DNS resolver's own
namespace lookup was never itself stale, isolating the cause to the
script's own premature exit. Re-verified extensively after the fix:
8/8 isolated repeats of the previously-flaky tap+relay test, and 6
consecutive full [integration][root][net] suite runs (78 test-case
executions total) with no failures.
2026-09-05 12:00:29 +00:00
ceamac 65c6d8fc03 Add DNS hostname resolution ping tests (veth + tap+relay variants)
Two peers on the same intern network, one started with --hostname
peer-a, resolve and ping each other by name via the per-session dnsmasq
resolver (network_dns.cpp). Skips cleanly when dnsmasq isn't available.

wait_for_hostname_then() retries the whole ping-by-name probe (not just
an eth0 existence check) until it succeeds or the bound is hit, since
this races against two independent things starting concurrently with
the sandboxed command: the interface coming up, and the per-session
resolver picking up the peer's own hosts record -- same underlying
join_networks() timing limitation the earlier wait_for_eth0_then() fix
was for. Verified end-to-end as root, both variants.
2026-09-05 11:42:06 +00:00
ceamac 4b264d3df3 Add extern-network outside-reachability tests (veth + tap+relay variants)
The positive counterpart to the intern-network isolation tests: a
container joined to an extern network gets a default route through its
uplink and can reach a real outside address. Verified end-to-end as
root, both variants.
2026-09-05 11:40:16 +00:00
ceamac 441ccf1419 Add extern-network peer IP ping tests (veth + tap+relay variants)
Same shape as the intern-network peer-ping pair, on an extern network
instead: confirms same-bridge peer reachability still works once the
extern uplink (network_bridge.cpp's ensure_uplink_provisioned()) is also
provisioned alongside the bridge. Verified end-to-end as root, both
veth and tap+relay variants.
2026-09-05 11:39:40 +00:00
ceamac b42073c48e Add intern-network outside-isolation tests (veth + tap+relay variants)
A container joined only to an intern network gets no default route
(network_join.cpp), so pinging a real outside address (8.8.8.8) must
fail outright, not merely succeed slower than an extern join would.
Verified end-to-end as root: both variants correctly report
"Network unreachable" and the tests assert a RESULT= line was seen
(the command actually ran) that isn't RESULT=0.
2026-09-05 11:39:01 +00:00
ceamac f9e68d48d9 Add intern-network IP ping tests (veth + tap+relay variants)
First of a planned series of end-to-end -n/--network join tests
(tests/integration/test_network_join_scenarios.cpp, [integration][root][net]):
two containers joined to the same intern network ping each other by IP,
run once with a real veth pair and once forced onto the tap+relay
fallback, since the two are genuinely different implementations. IPv6 is
deliberately excluded pending a known device-specific peculiarity.

split_lines_trimmed()/extract_marked_lines() moved from
test_rootless_run.cpp into tests/support/fixtures.{h,cpp} for reuse here.

wait_for_eth0_then() wraps a sandboxed command's own network-touching
script in a poll for eth0 to exist first: join_networks() runs
concurrently with, not before, the sandboxed command starting, so a
near-instant command can otherwise exit before its own join finishes --
the exact limitation already documented in network_join.{h,cpp}'s own
CLAUDE.md entry. Confirmed by testing (not assumed): without this, the
container's own immediate ping-and-exit sometimes raced ahead of the
veth-move step, which then failed outright ("Invalid netns value")
against an already-exited pid.
2026-09-05 11:38:21 +00:00
ceamac 31e81c88d0 Skip individual namespace-type checks that policy disables
Same class of test gap as the net-namespace test just fixed: this test
filtered which of pid/uts/ipc/cgroup to check by kernel support alone,
never by the config's own policy, so running the full suite as root with
every unshare-* flag forced off (via -c/--config-file) correctly made
bwrap skip requesting all four -- production code working as configured,
not a bug -- while the test still asserted each was isolated and failed.

Reuses namespace_type_would_isolate() (introduced in the previous commit)
to filter to_check by both gates instead of kernel support alone.

Full suite as root with everything disabled: 78 test cases, 76 passed, 2
skipped (this one and the net-namespace test), 0 failed -- the complete
picture requested.
2026-09-05 10:34:33 +00:00
ceamac 4f988c1bbc Skip the net-namespace-isolation test when policy disables it
Found by running the full suite as root with every unshare-* flag forced
off via -c/--config-file: this test only ever checked kernel support for
--unshare-net, never the config's own policy, so disabling it via
global.unshare-net correctly makes bwrap skip requesting the namespace --
production code working exactly as configured, not a bug -- while the test
still asserted isolation and failed.

Generalized the existing pid-only two-gate check into
namespace_type_would_isolate(type), covering any of bwrap.cpp's own
namespace_probes names, and used it here instead of a kernel-support-only
check.
2026-09-05 10:33:18 +00:00
ceamac e047d243f2 Fix a race in create_session_cgroup(); skip the straggler test when unreachable
create_session_cgroup() used to be called from run_bwrap()'s on_start
callback, in the parent, concurrently with the just-forked child execing
into bwrap and bwrap then doing its own internal clone() of the sandboxed
target. Without --unshare-pid, bwrap has little enough setup work to do
that it could reliably win that race, cloning its target before the
parent's own write into cgroup.procs completed -- leaving that target, and
everything it later spawns, permanently outside the tracked cgroup, so the
post-exit sweep found nothing to reap. Found via the user's own request to
test "pid namespace off, cgroup on" as root: confirmed directly by
inspecting cgroup.procs mid-session, showing only bwrap's own pid.

Fixed by giving run_process_foreground() a new before_exec hook, invoked in
the child synchronously right before execvp() -- the child cannot proceed
to exec (and thus cannot trigger any of bwrap's own internal forking) until
this has already returned, closing the race structurally rather than by
timing luck. run_bwrap() now creates the session cgroup there instead of in
on_start; the parent side just reconstructs the deterministic path
unconditionally, since the downstream sweep/cleanup functions already
tolerate a nonexistent directory gracefully either way.

Also fixes a false positive found while verifying this: the regression
test's own process-matching did a substring search across a whole cmdline
blob, which matched an unrelated manual `pkill -f 'sleep 137'` diagnostic
command run by hand during the investigation. Tightened to an exact
argv[0]/argv[1] match.

Finally, the regression test now SKIP()s (instead of failing) when neither
a pid namespace nor a working session cgroup is available for the current
effective config -- a documented, known residual limitation, not a
regression -- checked directly via two new helpers rather than assumed from
e.g. geteuid().
2026-09-05 10:28:07 +00:00
ceamac 4dd0f462a5 Wire -c/--config-file's effective config into the self-test suite
run_self_tests() now takes the same effective AppConfig any other command
gets and stashes it into a new g_test_app_config global (tests/support/
fixtures.h) before Catch2 runs anything. test_rootless_run.cpp's own
run_in_fixture() reads a copy of it instead of a hardcoded default
AppConfig{}, so -c actually reaches that test's container creation.

Verified: a config with every unshare-*/with-* key set to false, used via
-c <file> -t -- "[integration][net]~[root]", flips all 3 of that file's
container-creating tests to failing -- including the nohup-straggler
regression test, since with neither a pid namespace nor a cgroup nothing
reaps the backgrounded process -- while [unit] and non-networking
[integration] tests are completely unaffected, as expected.
2026-09-05 10:02:12 +00:00
ceamac a22024fcf8 Add regression test: a nohup-backgrounded process doesn't survive a session
Reproduces the user's reported real-world shape end to end with a real
busybox container ("nohup sleep 137 & exit"), verifying via the host's own
/proc that the backgrounded process is actually gone afterward -- checked
by cmdline substring, not pid, since a pid seen inside an isolated pid
namespace doesn't correspond to the same-numbered host pid. A bounded 2s
poll guards against the pid namespace's own kernel collapse-on-pid-1-exit
timing (which already covers this case for free on this dev machine).

Complements test_session_cleanup.cpp's existing [integration][root] test,
which exercises kill_via_cgroup() directly -- this one instead proves the
outward, visible contract holds through the real -r/--run path.
2026-09-05 09:50:49 +00:00
ceamac cbec986e78 Split config.yaml into global+persistent files; add -c/--config-file
config.yaml now holds only the global section (log-level, unshare-*,
with-veth, with-ipv6); a new persistent.yaml holds volumes/networks.
load_config_file()/write_config_file() are replaced by
load_global_config()/load_persistent_config()/write_global_config()/
write_persistent_config(), each touching only their own file.

-c/--config-file <path> lets one invocation use an alternate file for the
global section only -- persistent.yaml is always the one fixed path,
regardless of -c, so an experiment can never affect real volumes/networks
(a -c file's own volumes/networks, if any, are simply never read either).
A -c path that doesn't exist is a hard error, unlike the default path's
existing missing-file leniency.

migrate_legacy_config_if_needed() moves volumes/networks out of an
old-format config.yaml into persistent.yaml on first run after upgrading,
always against the fixed default paths regardless of -c. A name collision
aborts the migration for that run (touching neither file) rather than
risking data loss.

Required reordering main() to parse CLI args before loading config (so
-c's value is known first) -- ParsedArgs::log_level_flag_given tracks
whether --log-level was already given so the config file's own log-level
doesn't clobber it despite the reversed call order.
2026-09-05 09:32:36 +00:00
ceamac d2d631117d Drop -m/-u/-c short options for --mount/--umount/--cleanup
Now that -r/--run and -x/--exec cover normal use, --mount/--umount/--cleanup
are debug-only escape hatches not worth a short letter. Reassigned their
long_options codes to long-option-only constants (options::mount/umount/
cleanup) and dropped m:/u:/c: from getopt_long's own short-options string.

Updated the fixture smoke test (tests/run_test.py) and docs, which invoked
-m/-u/-c directly.
2026-09-05 08:43:58 +00:00
ceamac cdf9dcd210 Add global.with-veth/with-ipv6 config defaults for -n/--network creation
Replaces --no-ipv6/--no-veth (plain flags) with --with-ipv6/--with-veth,
each taking an explicit true/false value (e.g. --with-veth=false), parsed
via the same parse_bool_flag() the config file itself already uses (now
exported from config_file.h so cli_args.cpp can reuse it).

create_network_command() now resolves ipv6/veth as CLI flag -> config's own
global.with-ipv6/global.with-veth -> true, so a host that always wants the
tap+relay fallback (or no IPv6) can set it once in the config instead of
passing the flag on every network creation. -w/--write-config fills in both
new keys like the existing six unshare-* bools.
2026-09-05 08:36:12 +00:00
ceamac 1f52bc5f6f Add regression test for the session-straggler sweep; resolve TODO entry
test_session_cleanup.cpp exercises kill_via_cgroup() directly against two
plain forked processes (one setsid()-ing away from the other before it
exits), confirming a reparented straggler is actually reaped -- reproducing
the real escape shape (no pid namespace support at all) through a full
mount/bwrap session isn't possible from the CLI on a single run, since
--unshare-pid is a config-file-only setting, not a flag.

Also resolves TODO.md's SIGINT/SIGTERM entry and extends the relevant
CLAUDE.md sections (bwrap.{h,cpp}, session_cgroup.{h,cpp}, kill_session.{h,cpp})
with the fix's rationale and its known residual limitation (a kernel with
neither cgroup v2 nor pid namespace support still can't be reached
automatically).
2026-09-05 07:20:51 +00:00
ceamac aea4d90463 Automatically reap session stragglers after run_bwrap() exits
A container process that daemonizes/double-forks and setsid()'s away can
escape bwrap's own pid tree and outlive the session, whether it ends via a
normal exit, a forwarded SIGINT/SIGTERM, or -D/--daemonize -- --kill's own
cgroup-based sweep already reaches such a straggler for a still-running
session, but nothing ran that sweep automatically once the session itself
ended.

Export kill_via_cgroup() (previously kill_session.cpp-local) and call it
from run_bwrap(), right after run_process_foreground() returns and before
the session cgroup is removed, whenever anything is still left in it -- runs
unconditionally regardless of why bwrap exited, and covers -D for free since
it re-enters this same run_bwrap() call from within the daemonized child.
2026-09-05 07:18:41 +00:00
ceamac cadea945e4 Fix [integration][net] rootless tests to not assume kernel capability
Found via real-device testing (the actual Android target), not assumed:
both tests hardcoded assumptions that don't hold on every kernel.

1. "a fresh network namespace has only loopback" assumed exactly 3 lines
   of /proc/net/dev (2-line header + one "lo" entry) -- the real target
   device's kernel auto-creates several harmless placeholder tunnel
   interfaces (sit0, ip6tnl0, ip_vti0, ip6_vti0) in *every* fresh network
   namespace, alongside loopback. The namespace is still genuinely
   isolated (confirmed: none of the *host's* real interfaces leak in) --
   the test's assumption was just wrong for this kernel. Replaced the
   exact-count check with a readlink-based /proc/self/ns/net identity
   comparison (proves genuine isolation regardless of kernel config) plus
   a simple "loopback is present" check, dropping the brittle count
   assertion entirely.

2. "pid/uts/ipc namespaces differ from this process's own" assumed all
   three are always readable via /proc/self/ns/<type>. The real target
   device has neither PID nor IPC namespace support *as a kernel feature
   at all* -- confirmed directly: even this test process's own `readlink
   /proc/self/ns/pid`, run completely outside any container, fails
   outright there. This matches this project's own already-documented
   standing lesson (neither CONFIG_CHECKPOINT_RESTORE nor pid namespace
   support on this target). Comparing against a namespace type the kernel
   doesn't expose at all wouldn't prove anything either way.

Both tests now build their expectations from detect_bwrap_unshare_args()
(bwrap.h) -- this host's own live kernel-capability probe, the exact same
one build_bwrap_args() itself already gates on -- rather than assuming a
fixed set of namespace types is always available. "user" is deliberately
excluded from the generic per-type check: build_bwrap_args() never
requests --unshare-user when running as root, so asserting on it would be
wrong specifically when these tests are run as root (as they are on the
real device).

Verified: passes repeatably on this dev machine (all 6 namespace types
supported, 8 assertions/2 test cases either way -- same coverage as
before, just derived instead of hardcoded), full combined suite and
meson test both still clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-04 15:09:08 +00:00
ceamac 0c3a4bae99 TODO: Ctrl-C/SIGTERM on a foreground session may leave orphans
forward_signal_to_foreground_child() (process.cpp) does a plain kill() on
only the one tracked bwrap pid -- unlike --kill's kill_session(), which
picks a strategy (cgroup, pid-namespace, or tracked-pid) to reach every
process the session started. Anything inside the sandbox that
daemonizes/double-forks into a new session escapes the simple forward and
can be left running after Ctrl-C, even though --kill against the same
session would reach it. Reported by the user during real-device testing;
not yet reproduced with a specific repro, just the architectural gap.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-04 14:46:39 +00:00
ceamac b602de8e7b Document the new Catch2 test suite
README.md: new "Testing" section -- the 4 category/tag-expression table,
what meson test covers vs. what stays manual, and tests/setup-tests.py.

CLAUDE.md: rewrote the self_test.{h,cpp} entry to describe its new role
(pure Catch2 Session::run() plumbing, ENABLE_TESTS-guarded) instead of the
hand-rolled tests it used to contain directly, and added a full
per-file breakdown of the new tests/unit, tests/integration, and
tests/support infrastructure -- including the real bugs found building it
(the two parse_args()/getopt_long state-reset bugs, the ScratchXdgDirs
mixed-iterator UB, the Catch2-inherited-SIGTERM-handler artifact, the
missing /sys mount and spdlog-writes-to-stdout findings), all in the same
narrative depth this file already uses throughout. Also updated the
"Build & test commands" flag list and meson test description.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-04 11:23:24 +00:00
ceamac 865862762d Wire [unit] and safe [integration] tests into meson test
Two new test() entries alongside the existing fixture smoke test:
'unit-tests' (-t -- "[unit]") and 'integration-tests'
(-t -- "[integration]~[net]") -- both safe to run unprivileged with no
network setup, so meson test -C buildDir now catches regressions in those
categories automatically. [net] and [root] tests stay manual-only
(developer-run on a real/root-capable machine), matching how this
project's own self-tests were never part of meson test either. Only
registered when enable_tests is on, matching test_sources' own guard --
verified a -Denable_tests=false build still runs cleanly with just the
original smoke test.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-04 11:23:24 +00:00
ceamac 80cc49d898 Add [integration][net] rootless container-run tests
tests/integration/test_rootless_run.cpp runs a real busybox image through
the exact real -r/--run dispatch path (dispatch_command(), commands.h) --
mount, resolve, run_bwrap, unmount, cleanup, in-process rather than via a
subprocess -- and confirms bwrap's *default* sandboxing (no -n/-p at all)
is genuinely isolating: a fresh network namespace with nothing but
loopback, and pid/uts/ipc namespaces that differ from this test process's
own. No root needed, same as a plain `-r image.tar -- <command>` already
isn't.

New tests/support helpers: CapturedStdout (RAII, redirects this process's
own fd 1 -- and anything a forked/exec'd child inherits from it -- to a
throwaway temp file for its lifetime) so the sandboxed command's own
output can actually be asserted on.

Two real, non-obvious findings from getting this working, not assumed:
1. bwrap's own sandbox mounts --proc /proc and --dev /dev, but *not*
   /sys -- confirmed directly (`ls /sys/class/net` inside the sandbox:
   "No such file or directory", reproduced identically via the real CLI,
   not just this test). Switched the loopback-only check to
   /proc/net/dev instead (two header lines + one "<iface>: ..." line per
   interface), which correctly shows only "lo".
2. spdlog's default sink writes to stdout, not stderr, same as the plain
   "mounted image at: ..." success line (see CLAUDE.md) -- so a naive
   capture-and-line-split mixed slocker-lite's own status/log output in
   with the sandboxed command's real output. Fixed by having the
   sandboxed command bracket its own output between two unique markers
   and extracting only what's strictly between them.

Verified: both tests pass repeatably, 15 stress-test runs of the full
combined [unit]+[integration] suite with zero failures, plus a full run
as root.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-04 11:23:24 +00:00
ceamac dd66886de8 Add tests/setup-tests.py: fetch a real busybox fixture for [net] tests
Idempotent: does nothing if images/busybox.tar already exists (your own
unofficial build, or a previous fetch). Otherwise tries skopeo, then
podman, then docker, in that order -- skopeo/podman both reliably produce
a genuine OCI Image Layout tar; a plain `docker save` only does if the
containerd image store happens to be enabled, so the result is verified
(oci-layout/index.json actually present at the tar root) regardless of
which tool produced it, falling through to the next option otherwise.
Clear instructions + nonzero exit if none of the three are available and
no fixture already exists.

Verified the no-op (already-present) path and the no-tools-available error
path directly (temporarily moved the existing images/busybox.tar aside and
back) -- this dev machine has none of skopeo/podman/docker installed, so
the actual fetch path itself is unverified here; the format-verification
step (looks_like_oci_layout()) is what protects against a `docker save`
that produced the legacy Docker tar format on some other machine.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-09-04 11:23:24 +00:00