ceamac 1c186b0365 Fix -e/--exec namespace resolution on kernels without CONFIG_CHECKPOINT_RESTORE
/proc/<pid>/task/<pid>/children doesn't exist on every kernel (confirmed
missing on a real Android target), so resolve_namespace_pid() fell back to
the outer bwrap pid itself and nsenter ended up with no namespace flags at
all. Add a portable fallback that scans /proc/<n>/stat for the child whose
ppid matches, the same information pstree uses to build its tree.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-08-29 07:42:02 +00:00

slocker-lite

slocker-lite mounts an OCI Image Layout tar (the format produced by skopeo, podman save --format oci-archive, or a modern docker save) and runs a sandboxed command against it — without podman, docker, or kernel overlayfs.

Why

The target environment is Android with a stock kernel: podman/docker don't run there (missing namespace support), and there's no kernel overlayfs. slocker-lite works around both: it imports image layers into containers-storage and mounts them with fuse-overlayfs (userspace, no kernel overlayfs needed), then sandboxes the run with bwrap in "degraded mode" — using only whichever --unshare-xxx namespaces the running kernel actually supports, instead of requiring the full set.

Status

Early-stage. Mounting, running, and dropping privileges to a specific user/group all work. Named volumes (-v/--volume) can be created and are persisted in the config file, and can be mounted into -r/--run (repeatably), along with ad hoc host directories. Image-declared networking (ExposedPorts/Env from the image config, and the image's own separately-declared Volumes) are parsed but not yet applied, and there's no background/daemonized run mode yet.

Requirements

Build-time:

  • Meson + a C++20 compiler
  • fmt, libarchive, nlohmann_json, yaml-0.1 (libyaml)
  • spdlog (uses the system package if found, otherwise fetched automatically via the vendored subprojects/spdlog.wrap)
  • catch2, only if the enable_tests Meson option is on (default: on)

Runtime:

  • containers-storage
  • fuse-overlayfs
  • bwrap (bubblewrap)
  • nsenter (only needed for non-root runs — see "How it works" below)

Build

meson setup buildDir
meson compile -C buildDir
meson test -C buildDir

This also builds buildDir/slocker-lite-priv-drop, a small statically-linked helper that -r --user/--group needs at runtime (see "How it works").

Usage

slocker-lite -m|--mount <image.tar>
slocker-lite -r|--run <image.tar> [-v <name-or-dir> <container-path>]... [-- <command> [args...]]
slocker-lite -u|--umount <layer-id>
slocker-lite -c|--cleanup <layer-id>
slocker-lite -l|--list-images <directory>
slocker-lite -i|--inspect <image.tar>
slocker-lite -e|--exec <pid> [-- <command> [args...]]
slocker-lite -v|--volume <name> <directory>
slocker-lite --list-volumes
slocker-lite --delete-volume <name>
slocker-lite --delete-volume-full <name>
slocker-lite --list-processes
slocker-lite --clean-processes
slocker-lite -t|--test
slocker-lite -h|--help
slocker-lite -V|--version
Flag Description
-m, --mount <image.tar> Validate and mount an OCI Image Layout tar.
-r, --run <image.tar> Mount, run bwrap in the foreground, then unmount and clean up on exit. Defaults to the image's own Entrypoint/Cmd (or /bin/sh if neither is set); pass -- <command> [args...] to override.
-u, --umount <layer-id> Unmount a previously mounted layer (the ID printed by --mount/--run, or from containers-storage layers).
-c, --cleanup <layer-id> Delete a layer and its ancestor chain from local storage (unmount it first).
-n, --no-nsenter With --run, bind the mount directly instead of nsenter-ing into fuse-overlayfs's namespace. Automatic when running as root; use this to force it off otherwise.
-D, --daemonize With --run, fork into the background: detaches from the controlling terminal (setsid()), ignores SIGHUP, and redirects stdin from /dev/null and stdout/stderr to a log file under $XDG_STATE_HOME/slocker-lite/logs/. Prints the session's pid and log path, then returns — the same pid --list-processes/-e/--exec use.
--user <user> With --run, run the command as this user (name or numeric uid) instead of the image's own declared user (or root, if it declares none). Resolved against the image's own /etc/passwd. Only takes effect when --run executes as root.
--group <group> With --user, use this group (name or numeric gid) instead of the user's primary group.
--hostname <name> With --run, set the sandbox's hostname. Only takes effect if the running kernel supports --unshare-uts; ignored with a warning otherwise.
--env VAR=VALUE With --run, set an environment variable in the sandbox (overrides the default PATH/HOME/PWD/TERM if given the same name). Repeatable; combined with --env-file in command-line order, each later one winning over an earlier one for the same name.
--env-file <file> With --run, load environment variables from <file> — one VAR=VALUE per line; blank lines and #-comments are skipped. Repeatable.
-l, --list-images <dir> List OCI Image Layout tars (*.tar, *.tar.*) found directly in <dir>, with their name:tag.
-i, --inspect <image.tar> Print an image's declared user, exposed ports, env, volumes, and default command, without mounting or running it.
-e, --exec <pid> Join an already-running --run session (<pid> must be one --list-processes shows as running) and run a command inside its container. Pass -- <command> [args...] to specify it.
-v, --volume <name> <dir> Create a named volume mapped to a host directory (created if missing), recorded in the config file's volumes section. Fails if the name or directory is already used by an existing volume. Volume names can't contain /. With --run, instead mounts a volume into the sandbox (repeatable): <name> is an existing named volume, or, if it contains /, a host directory path (created if missing); <dir> is the absolute path inside the container to mount it at. If the host directory is empty and the image already has content there, that content is copied in first, preserving numeric ownership/permissions/links and, where the host filesystem supports them, extended attributes/ACLs (skipped with a warning otherwise).
--list-volumes List all named volumes (see -v/--volume) with their host directory.
--delete-volume <name> Remove a named volume from the config. The host directory is left untouched.
--delete-volume-full <name> Like --delete-volume, but also recursively deletes the volume's host directory.
--list-processes List running --run sessions found by their pid files under $XDG_STATE_HOME/slocker-lite/run/, with their pid, container name, and status (running or exited).
--clean-processes Remove stale pid files (see --list-processes) left behind by sessions that are no longer running.
-t, --test Print which bwrap --unshare-xxx namespaces the running kernel supports.
--log-level <level> Set log verbosity (trace, debug, info, warn, error, critical, off).
-h, --help Print usage and exit.
-V, --version Print version information and exit.

Examples

# Mount an image and inspect it (prints the merged mount path)
./buildDir/slocker-lite -m myimage.tar

# Mount, run the image's default command, then unmount and clean up
./buildDir/slocker-lite -r myimage.tar

# Run a specific command instead
./buildDir/slocker-lite -r myimage.tar -- /bin/sh -c 'echo hello'

# Run as a specific user (as root only)
sudo ./buildDir/slocker-lite -r myimage.tar --user git

# Run with a custom hostname inside the sandbox
./buildDir/slocker-lite -r myimage.tar --hostname mybox

# Run with extra environment variables, from flags and/or a file
./buildDir/slocker-lite -r myimage.tar --env FOO=bar --env-file ./app.env

# Run in the background; prints its pid and log file, then returns
./buildDir/slocker-lite -r myimage.tar -D

# List every OCI image tar in a directory
./buildDir/slocker-lite -l ./images

# Inspect an image's declared config without mounting or running it
./buildDir/slocker-lite -i myimage.tar

# Start a busybox container in the background, then get a shell inside it from
# another terminal (find its pid with --list-processes)
./buildDir/slocker-lite -r busybox.tar &
./buildDir/slocker-lite --list-processes
./buildDir/slocker-lite -e 12345 -- /bin/sh

# Create a named volume backed by a host directory
./buildDir/slocker-lite -v mydata ~/slocker-volumes/mydata

# Run, mounting that named volume plus an ad hoc host directory
./buildDir/slocker-lite -r myimage.tar -v mydata /data -v ~/scratch /scratch

# List all named volumes
./buildDir/slocker-lite --list-volumes

# Remove a named volume (keeps its host directory)
./buildDir/slocker-lite --delete-volume mydata

# Remove a named volume and delete its host directory too
./buildDir/slocker-lite --delete-volume-full mydata

# List currently running (and any leftover, exited) --run sessions
./buildDir/slocker-lite --list-processes

# Remove any leftover, stale pid files
./buildDir/slocker-lite --clean-processes

Configuration

Persistent settings can be kept in a local YAML config file at $XDG_CONFIG_HOME/slocker-lite/config.yaml (falling back to $HOME/.config/slocker-lite/config.yaml if XDG_CONFIG_HOME isn't set). The file is organized into sections:

global:
  log-level: debug
volumes:
  mydata: /home/user/slocker-volumes/mydata

global.log-level is the only standing preference supported today (one-shot commands like --mount/--run/--user don't belong in a config file). An explicit --log-level on the command line always overrides the config file. The volumes section is managed by -v/--volume (see above) rather than hand-edited — it's what -r/--run's own -v usage looks named volumes up in. A missing config file is fine either way (nothing is overridden, and one gets created the first time -v/--volume is used).

How it works

Image layers are imported into containers-storage (parent-chained) and the resulting top layer is mounted via fuse-overlayfs; -r/--run then sandboxes the requested command with bwrap. Because containers-storage mount runs rootless by re-execing into a private user+mount namespace, -r/--run normally has to nsenter into that namespace to reach the mount — except when running as root, where the mount is already directly visible and --unshare-user is skipped entirely (a fresh user namespace isn't needed for root's own privilege, and forces an unrelated supplementary-group bug in that case). Running as root also unlocks --user/ --group: since bwrap --uid/--gid require a user namespace that isn't available there, slocker-lite instead bind-mounts a separate, statically-linked helper (slocker-lite-priv-drop) into the sandbox and routes the command through it to drop privileges before exec.

While a -r/--run session is active, its bwrap process is tracked as a locked PID file under $XDG_STATE_HOME/slocker-lite/run/ (falling back to $HOME/.local/state/...), named after the image and its PID so the same image can be run concurrently without collisions. The file is removed automatically once the run ends; any tool can check whether a session is still alive by attempting the same exclusive, non-blocking flock() on its file. --list-processes does exactly that for every pid file it finds, reporting each one's pid, container name, and running/exited status. Normally the file is removed automatically when its own session ends, but --clean-processes removes any stale ones left behind (e.g. after a crash) using that same check, atomically per file, so it never removes one that's still genuinely running.

-e/--exec <pid> joins a running session's namespaces with nsenter and runs a command there. Because bwrap itself sets up the sandbox's mount/user namespaces and then hands the actual sandboxed command off to a child process in fresh pid/uts/ipc/cgroup namespaces, -e resolves that real child first (via /proc) rather than joining the outer bwrap process's own namespaces, so the joined command sees the container's process tree and hostname too, not just its filesystem.

-D/--daemonize forks and detaches into the background by calling setsid() itself, rather than re-enabling bwrap's own --new-session — that flag only detaches the deeply-nested sandboxed command, leaving bwrap/nsenter still attached to the original session. Calling setsid() in slocker-lite's own forked child, before it execs into nsenter/bwrap, detaches the whole chain at once (exec() never changes session membership), and correctly scopes bwrap's own --die-with-parent to that child. The child ignores SIGHUP and redirects output to a log file before doing anything else; the original, still-foreground process waits only long enough to learn the real session pid (the same one --list-processes/-e/--exec use) before printing it and returning — the detached child is what runs the entire session afterward, including the same unmount/cleanup that always ran once the sandboxed command exits.

See CLAUDE.md for the full architecture writeup (file-by-file breakdown, the reasoning behind each of the above, and known gaps).

License

GPL-2.0-or-later. See COPYING.

S
Description
Run a docker image with bubblewrap on low end devices
Readme 772 KiB
Languages
C++ 97.9%
Python 1%
Meson 0.7%
C 0.3%