Files
slocker-lite/README.md
T
ceamac 00482b6b42 Add -e/--exec to join a running -r/--run session
Validates the given pid against the same tracked-session liveness
check --list-processes/--clean-processes already use, then joins its
namespaces via nsenter and runs a command there in the foreground.

Two things discovered only by testing against a live session, not
assumed up front:

- The tracked pid is bwrap's own outer process. It sets up the
  mount/user namespaces itself, then clone()s the actual sandboxed
  command into fresh pid/uts/ipc/cgroup namespaces -- clone()'s
  namespace flags only ever affect the new child, never the caller,
  so the outer process itself never enters those namespaces at all.
  exec_in_session() resolves that real child via
  /proc/<pid>/task/<pid>/children and joins its namespaces instead,
  falling back to the outer pid if that can't be read.

- Rather than nsenter -a (which would hit a known "Invalid argument"
  failure re-entering an identical namespace -- this project already
  worked around exactly that once, for the containers-storage mount
  path), each namespace type is only joined if
  /proc/<pid>/ns/<type> actually differs from this process's own.
  nsenter also needs --preserve-credentials, or it tries to
  setuid/setgid/setgroups to the target's identity, which fails
  outright against the setgroups-denied unprivileged user namespace
  bwrap creates whenever -r/--run isn't root.

Verified end-to-end: joined shell gets the container's own hostname,
process tree (ps shows only container processes), and root
filesystem; untracked/stale pids error out cleanly without touching
nsenter; Ctrl-C during the joined command doesn't disturb the
original session; no leftover mounts after either exits.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
2026-08-22 09:41:32 +00:00

212 lines
11 KiB
Markdown

# slocker-lite
`slocker-lite` mounts an [OCI Image Layout](https://github.com/opencontainers/image-spec/blob/main/image-layout.md)
tar (the format produced by `skopeo`, `podman save --format oci-archive`, or a modern
`docker save`) and runs a sandboxed command against it — without `podman`, `docker`,
or kernel overlayfs.
## Why
The target environment is Android with a stock kernel: `podman`/`docker` don't run
there (missing namespace support), and there's no kernel overlayfs. `slocker-lite`
works around both: it imports image layers into `containers-storage` and mounts them
with `fuse-overlayfs` (userspace, no kernel overlayfs needed), then sandboxes the run
with `bwrap` in "degraded mode" — using only whichever `--unshare-xxx` namespaces the
running kernel actually supports, instead of requiring the full set.
## Status
Early-stage. Mounting, running, and dropping privileges to a specific user/group all
work. Named volumes (`-v/--volume`) can be created and are persisted in the config
file, and can be mounted into `-r/--run` (repeatably), along with ad hoc host
directories. Image-declared networking (`ExposedPorts`/`Env` from the image config,
and the image's own separately-declared `Volumes`) are parsed but not yet applied,
and there's no background/daemonized run mode yet.
## Requirements
Build-time:
- Meson + a C++20 compiler
- `fmt`, `libarchive`, `nlohmann_json`, `yaml-0.1` (libyaml)
- `spdlog` (uses the system package if found, otherwise fetched automatically via
the vendored `subprojects/spdlog.wrap`)
- `catch2`, only if the `enable_tests` Meson option is on (default: on)
Runtime:
- `containers-storage`
- `fuse-overlayfs`
- `bwrap` (bubblewrap)
- `nsenter` (only needed for non-root runs — see "How it works" below)
## Build
```sh
meson setup buildDir
meson compile -C buildDir
meson test -C buildDir
```
This also builds `buildDir/slocker-lite-priv-drop`, a small statically-linked helper
that `-r --user`/`--group` needs at runtime (see "How it works").
## Usage
```
slocker-lite -m|--mount <image.tar>
slocker-lite -r|--run <image.tar> [-v <name-or-dir> <container-path>]... [-- <command> [args...]]
slocker-lite -u|--umount <layer-id>
slocker-lite -c|--cleanup <layer-id>
slocker-lite -l|--list-images <directory>
slocker-lite -i|--inspect <image.tar>
slocker-lite -e|--exec <pid> [-- <command> [args...]]
slocker-lite -v|--volume <name> <directory>
slocker-lite --list-volumes
slocker-lite --delete-volume <name>
slocker-lite --delete-volume-full <name>
slocker-lite --list-processes
slocker-lite --clean-processes
slocker-lite -t|--test
slocker-lite -h|--help
slocker-lite -V|--version
```
| Flag | Description |
| --- | --- |
| `-m, --mount <image.tar>` | Validate and mount an OCI Image Layout tar. |
| `-r, --run <image.tar>` | Mount, run `bwrap` in the foreground, then unmount and clean up on exit. Defaults to the image's own `Entrypoint`/`Cmd` (or `/bin/sh` if neither is set); pass `-- <command> [args...]` to override. |
| `-u, --umount <layer-id>` | Unmount a previously mounted layer (the ID printed by `--mount`/`--run`, or from `containers-storage layers`). |
| `-c, --cleanup <layer-id>` | Delete a layer and its ancestor chain from local storage (unmount it first). |
| `-n, --no-nsenter` | With `--run`, bind the mount directly instead of `nsenter`-ing into `fuse-overlayfs`'s namespace. Automatic when running as root; use this to force it off otherwise. |
| `--user <user>` | With `--run`, run the command as this user (name or numeric uid) instead of the image's own declared user (or root, if it declares none). Resolved against the image's own `/etc/passwd`. Only takes effect when `--run` executes as root. |
| `--group <group>` | With `--user`, use this group (name or numeric gid) instead of the user's primary group. |
| `--hostname <name>` | With `--run`, set the sandbox's hostname. Only takes effect if the running kernel supports `--unshare-uts`; ignored with a warning otherwise. |
| `-l, --list-images <dir>` | List OCI Image Layout tars (`*.tar`, `*.tar.*`) found directly in `<dir>`, with their `name:tag`. |
| `-i, --inspect <image.tar>` | Print an image's declared user, exposed ports, env, volumes, and default command, without mounting or running it. |
| `-e, --exec <pid>` | Join an already-running `--run` session (`<pid>` must be one `--list-processes` shows as `running`) and run a command inside its container. Pass `-- <command> [args...]` to specify it. |
| `-v, --volume <name> <dir>` | Create a named volume mapped to a host directory (created if missing), recorded in the config file's `volumes` section. Fails if the name or directory is already used by an existing volume. Volume names can't contain `/`. With `--run`, instead mounts a volume into the sandbox (repeatable): `<name>` is an existing named volume, or, if it contains `/`, a host directory path (created if missing); `<dir>` is the absolute path inside the container to mount it at. If the host directory is empty and the image already has content there, that content is copied in first, preserving numeric ownership/permissions/links and, where the host filesystem supports them, extended attributes/ACLs (skipped with a warning otherwise). |
| `--list-volumes` | List all named volumes (see `-v/--volume`) with their host directory. |
| `--delete-volume <name>` | Remove a named volume from the config. The host directory is left untouched. |
| `--delete-volume-full <name>` | Like `--delete-volume`, but also recursively deletes the volume's host directory. |
| `--list-processes` | List running `--run` sessions found by their pid files under `$XDG_STATE_HOME/slocker-lite/run/`, with their pid, container name, and status (`running` or `exited`). |
| `--clean-processes` | Remove stale pid files (see `--list-processes`) left behind by sessions that are no longer running. |
| `-t, --test` | Print which `bwrap --unshare-xxx` namespaces the running kernel supports. |
| `--log-level <level>` | Set log verbosity (`trace`, `debug`, `info`, `warn`, `error`, `critical`, `off`). |
| `-h, --help` | Print usage and exit. |
| `-V, --version` | Print version information and exit. |
### Examples
```sh
# Mount an image and inspect it (prints the merged mount path)
./buildDir/slocker-lite -m myimage.tar
# Mount, run the image's default command, then unmount and clean up
./buildDir/slocker-lite -r myimage.tar
# Run a specific command instead
./buildDir/slocker-lite -r myimage.tar -- /bin/sh -c 'echo hello'
# Run as a specific user (as root only)
sudo ./buildDir/slocker-lite -r myimage.tar --user git
# Run with a custom hostname inside the sandbox
./buildDir/slocker-lite -r myimage.tar --hostname mybox
# List every OCI image tar in a directory
./buildDir/slocker-lite -l ./images
# Inspect an image's declared config without mounting or running it
./buildDir/slocker-lite -i myimage.tar
# Start a busybox container in the background, then get a shell inside it from
# another terminal (find its pid with --list-processes)
./buildDir/slocker-lite -r busybox.tar &
./buildDir/slocker-lite --list-processes
./buildDir/slocker-lite -e 12345 -- /bin/sh
# Create a named volume backed by a host directory
./buildDir/slocker-lite -v mydata ~/slocker-volumes/mydata
# Run, mounting that named volume plus an ad hoc host directory
./buildDir/slocker-lite -r myimage.tar -v mydata /data -v ~/scratch /scratch
# List all named volumes
./buildDir/slocker-lite --list-volumes
# Remove a named volume (keeps its host directory)
./buildDir/slocker-lite --delete-volume mydata
# Remove a named volume and delete its host directory too
./buildDir/slocker-lite --delete-volume-full mydata
# List currently running (and any leftover, exited) --run sessions
./buildDir/slocker-lite --list-processes
# Remove any leftover, stale pid files
./buildDir/slocker-lite --clean-processes
```
## Configuration
Persistent settings can be kept in a local YAML config file at
`$XDG_CONFIG_HOME/slocker-lite/config.yaml` (falling back to
`$HOME/.config/slocker-lite/config.yaml` if `XDG_CONFIG_HOME` isn't set). The file is
organized into sections:
```yaml
global:
log-level: debug
volumes:
mydata: /home/user/slocker-volumes/mydata
```
`global.log-level` is the only standing preference supported today (one-shot
commands like `--mount`/`--run`/`--user` don't belong in a config file). An explicit
`--log-level` on the command line always overrides the config file. The `volumes`
section is managed by `-v/--volume` (see above) rather than hand-edited — it's what
`-r/--run`'s own `-v` usage looks named volumes up in. A missing config file is fine
either way (nothing is overridden, and one gets created the first time
`-v/--volume` is used).
## How it works
Image layers are imported into `containers-storage` (parent-chained) and the
resulting top layer is mounted via `fuse-overlayfs`; `-r/--run` then sandboxes the
requested command with `bwrap`. Because `containers-storage mount` runs rootless by
re-execing into a private user+mount namespace, `-r/--run` normally has to `nsenter`
into that namespace to reach the mount — except when running as root, where the mount
is already directly visible and `--unshare-user` is skipped entirely (a fresh user
namespace isn't needed for root's own privilege, and forces an unrelated
supplementary-group bug in that case). Running as root also unlocks `--user`/
`--group`: since `bwrap --uid`/`--gid` require a user namespace that isn't available
there, `slocker-lite` instead bind-mounts a separate, statically-linked helper
(`slocker-lite-priv-drop`) into the sandbox and routes the command through it to drop
privileges before exec.
While a `-r/--run` session is active, its `bwrap` process is tracked as a locked
PID file under `$XDG_STATE_HOME/slocker-lite/run/` (falling back to
`$HOME/.local/state/...`), named after the image and its PID so the same image can
be run concurrently without collisions. The file is removed automatically once the
run ends; any tool can check whether a session is still alive by attempting the
same exclusive, non-blocking `flock()` on its file. `--list-processes` does
exactly that for every pid file it finds, reporting each one's pid, container
name, and `running`/`exited` status. Normally the file is removed automatically
when its own session ends, but `--clean-processes` removes any stale ones left
behind (e.g. after a crash) using that same check, atomically per file, so it
never removes one that's still genuinely running.
`-e/--exec <pid>` joins a running session's namespaces with `nsenter` and runs a
command there. Because `bwrap` itself sets up the sandbox's mount/user namespaces
and then hands the actual sandboxed command off to a child process in fresh
pid/uts/ipc/cgroup namespaces, `-e` resolves that real child first (via `/proc`)
rather than joining the outer `bwrap` process's own namespaces, so the joined
command sees the container's process tree and hostname too, not just its
filesystem.
See `CLAUDE.md` for the full architecture writeup (file-by-file breakdown, the
reasoning behind each of the above, and known gaps).
## License
GPL-2.0-or-later. See [`COPYING`](COPYING).