00482b6b42
Validates the given pid against the same tracked-session liveness check --list-processes/--clean-processes already use, then joins its namespaces via nsenter and runs a command there in the foreground. Two things discovered only by testing against a live session, not assumed up front: - The tracked pid is bwrap's own outer process. It sets up the mount/user namespaces itself, then clone()s the actual sandboxed command into fresh pid/uts/ipc/cgroup namespaces -- clone()'s namespace flags only ever affect the new child, never the caller, so the outer process itself never enters those namespaces at all. exec_in_session() resolves that real child via /proc/<pid>/task/<pid>/children and joins its namespaces instead, falling back to the outer pid if that can't be read. - Rather than nsenter -a (which would hit a known "Invalid argument" failure re-entering an identical namespace -- this project already worked around exactly that once, for the containers-storage mount path), each namespace type is only joined if /proc/<pid>/ns/<type> actually differs from this process's own. nsenter also needs --preserve-credentials, or it tries to setuid/setgid/setgroups to the target's identity, which fails outright against the setgroups-denied unprivileged user namespace bwrap creates whenever -r/--run isn't root. Verified end-to-end: joined shell gets the container's own hostname, process tree (ps shows only container processes), and root filesystem; untracked/stale pids error out cleanly without touching nsenter; Ctrl-C during the joined command doesn't disturb the original session; no leftover mounts after either exits. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gv3s5jckJKzh6JkMoi2Akz
212 lines
11 KiB
Markdown
212 lines
11 KiB
Markdown
# slocker-lite
|
|
|
|
`slocker-lite` mounts an [OCI Image Layout](https://github.com/opencontainers/image-spec/blob/main/image-layout.md)
|
|
tar (the format produced by `skopeo`, `podman save --format oci-archive`, or a modern
|
|
`docker save`) and runs a sandboxed command against it — without `podman`, `docker`,
|
|
or kernel overlayfs.
|
|
|
|
## Why
|
|
|
|
The target environment is Android with a stock kernel: `podman`/`docker` don't run
|
|
there (missing namespace support), and there's no kernel overlayfs. `slocker-lite`
|
|
works around both: it imports image layers into `containers-storage` and mounts them
|
|
with `fuse-overlayfs` (userspace, no kernel overlayfs needed), then sandboxes the run
|
|
with `bwrap` in "degraded mode" — using only whichever `--unshare-xxx` namespaces the
|
|
running kernel actually supports, instead of requiring the full set.
|
|
|
|
## Status
|
|
|
|
Early-stage. Mounting, running, and dropping privileges to a specific user/group all
|
|
work. Named volumes (`-v/--volume`) can be created and are persisted in the config
|
|
file, and can be mounted into `-r/--run` (repeatably), along with ad hoc host
|
|
directories. Image-declared networking (`ExposedPorts`/`Env` from the image config,
|
|
and the image's own separately-declared `Volumes`) are parsed but not yet applied,
|
|
and there's no background/daemonized run mode yet.
|
|
|
|
## Requirements
|
|
|
|
Build-time:
|
|
- Meson + a C++20 compiler
|
|
- `fmt`, `libarchive`, `nlohmann_json`, `yaml-0.1` (libyaml)
|
|
- `spdlog` (uses the system package if found, otherwise fetched automatically via
|
|
the vendored `subprojects/spdlog.wrap`)
|
|
- `catch2`, only if the `enable_tests` Meson option is on (default: on)
|
|
|
|
Runtime:
|
|
- `containers-storage`
|
|
- `fuse-overlayfs`
|
|
- `bwrap` (bubblewrap)
|
|
- `nsenter` (only needed for non-root runs — see "How it works" below)
|
|
|
|
## Build
|
|
|
|
```sh
|
|
meson setup buildDir
|
|
meson compile -C buildDir
|
|
meson test -C buildDir
|
|
```
|
|
|
|
This also builds `buildDir/slocker-lite-priv-drop`, a small statically-linked helper
|
|
that `-r --user`/`--group` needs at runtime (see "How it works").
|
|
|
|
## Usage
|
|
|
|
```
|
|
slocker-lite -m|--mount <image.tar>
|
|
slocker-lite -r|--run <image.tar> [-v <name-or-dir> <container-path>]... [-- <command> [args...]]
|
|
slocker-lite -u|--umount <layer-id>
|
|
slocker-lite -c|--cleanup <layer-id>
|
|
slocker-lite -l|--list-images <directory>
|
|
slocker-lite -i|--inspect <image.tar>
|
|
slocker-lite -e|--exec <pid> [-- <command> [args...]]
|
|
slocker-lite -v|--volume <name> <directory>
|
|
slocker-lite --list-volumes
|
|
slocker-lite --delete-volume <name>
|
|
slocker-lite --delete-volume-full <name>
|
|
slocker-lite --list-processes
|
|
slocker-lite --clean-processes
|
|
slocker-lite -t|--test
|
|
slocker-lite -h|--help
|
|
slocker-lite -V|--version
|
|
```
|
|
|
|
| Flag | Description |
|
|
| --- | --- |
|
|
| `-m, --mount <image.tar>` | Validate and mount an OCI Image Layout tar. |
|
|
| `-r, --run <image.tar>` | Mount, run `bwrap` in the foreground, then unmount and clean up on exit. Defaults to the image's own `Entrypoint`/`Cmd` (or `/bin/sh` if neither is set); pass `-- <command> [args...]` to override. |
|
|
| `-u, --umount <layer-id>` | Unmount a previously mounted layer (the ID printed by `--mount`/`--run`, or from `containers-storage layers`). |
|
|
| `-c, --cleanup <layer-id>` | Delete a layer and its ancestor chain from local storage (unmount it first). |
|
|
| `-n, --no-nsenter` | With `--run`, bind the mount directly instead of `nsenter`-ing into `fuse-overlayfs`'s namespace. Automatic when running as root; use this to force it off otherwise. |
|
|
| `--user <user>` | With `--run`, run the command as this user (name or numeric uid) instead of the image's own declared user (or root, if it declares none). Resolved against the image's own `/etc/passwd`. Only takes effect when `--run` executes as root. |
|
|
| `--group <group>` | With `--user`, use this group (name or numeric gid) instead of the user's primary group. |
|
|
| `--hostname <name>` | With `--run`, set the sandbox's hostname. Only takes effect if the running kernel supports `--unshare-uts`; ignored with a warning otherwise. |
|
|
| `-l, --list-images <dir>` | List OCI Image Layout tars (`*.tar`, `*.tar.*`) found directly in `<dir>`, with their `name:tag`. |
|
|
| `-i, --inspect <image.tar>` | Print an image's declared user, exposed ports, env, volumes, and default command, without mounting or running it. |
|
|
| `-e, --exec <pid>` | Join an already-running `--run` session (`<pid>` must be one `--list-processes` shows as `running`) and run a command inside its container. Pass `-- <command> [args...]` to specify it. |
|
|
| `-v, --volume <name> <dir>` | Create a named volume mapped to a host directory (created if missing), recorded in the config file's `volumes` section. Fails if the name or directory is already used by an existing volume. Volume names can't contain `/`. With `--run`, instead mounts a volume into the sandbox (repeatable): `<name>` is an existing named volume, or, if it contains `/`, a host directory path (created if missing); `<dir>` is the absolute path inside the container to mount it at. If the host directory is empty and the image already has content there, that content is copied in first, preserving numeric ownership/permissions/links and, where the host filesystem supports them, extended attributes/ACLs (skipped with a warning otherwise). |
|
|
| `--list-volumes` | List all named volumes (see `-v/--volume`) with their host directory. |
|
|
| `--delete-volume <name>` | Remove a named volume from the config. The host directory is left untouched. |
|
|
| `--delete-volume-full <name>` | Like `--delete-volume`, but also recursively deletes the volume's host directory. |
|
|
| `--list-processes` | List running `--run` sessions found by their pid files under `$XDG_STATE_HOME/slocker-lite/run/`, with their pid, container name, and status (`running` or `exited`). |
|
|
| `--clean-processes` | Remove stale pid files (see `--list-processes`) left behind by sessions that are no longer running. |
|
|
| `-t, --test` | Print which `bwrap --unshare-xxx` namespaces the running kernel supports. |
|
|
| `--log-level <level>` | Set log verbosity (`trace`, `debug`, `info`, `warn`, `error`, `critical`, `off`). |
|
|
| `-h, --help` | Print usage and exit. |
|
|
| `-V, --version` | Print version information and exit. |
|
|
|
|
### Examples
|
|
|
|
```sh
|
|
# Mount an image and inspect it (prints the merged mount path)
|
|
./buildDir/slocker-lite -m myimage.tar
|
|
|
|
# Mount, run the image's default command, then unmount and clean up
|
|
./buildDir/slocker-lite -r myimage.tar
|
|
|
|
# Run a specific command instead
|
|
./buildDir/slocker-lite -r myimage.tar -- /bin/sh -c 'echo hello'
|
|
|
|
# Run as a specific user (as root only)
|
|
sudo ./buildDir/slocker-lite -r myimage.tar --user git
|
|
|
|
# Run with a custom hostname inside the sandbox
|
|
./buildDir/slocker-lite -r myimage.tar --hostname mybox
|
|
|
|
# List every OCI image tar in a directory
|
|
./buildDir/slocker-lite -l ./images
|
|
|
|
# Inspect an image's declared config without mounting or running it
|
|
./buildDir/slocker-lite -i myimage.tar
|
|
|
|
# Start a busybox container in the background, then get a shell inside it from
|
|
# another terminal (find its pid with --list-processes)
|
|
./buildDir/slocker-lite -r busybox.tar &
|
|
./buildDir/slocker-lite --list-processes
|
|
./buildDir/slocker-lite -e 12345 -- /bin/sh
|
|
|
|
# Create a named volume backed by a host directory
|
|
./buildDir/slocker-lite -v mydata ~/slocker-volumes/mydata
|
|
|
|
# Run, mounting that named volume plus an ad hoc host directory
|
|
./buildDir/slocker-lite -r myimage.tar -v mydata /data -v ~/scratch /scratch
|
|
|
|
# List all named volumes
|
|
./buildDir/slocker-lite --list-volumes
|
|
|
|
# Remove a named volume (keeps its host directory)
|
|
./buildDir/slocker-lite --delete-volume mydata
|
|
|
|
# Remove a named volume and delete its host directory too
|
|
./buildDir/slocker-lite --delete-volume-full mydata
|
|
|
|
# List currently running (and any leftover, exited) --run sessions
|
|
./buildDir/slocker-lite --list-processes
|
|
|
|
# Remove any leftover, stale pid files
|
|
./buildDir/slocker-lite --clean-processes
|
|
```
|
|
|
|
## Configuration
|
|
|
|
Persistent settings can be kept in a local YAML config file at
|
|
`$XDG_CONFIG_HOME/slocker-lite/config.yaml` (falling back to
|
|
`$HOME/.config/slocker-lite/config.yaml` if `XDG_CONFIG_HOME` isn't set). The file is
|
|
organized into sections:
|
|
|
|
```yaml
|
|
global:
|
|
log-level: debug
|
|
volumes:
|
|
mydata: /home/user/slocker-volumes/mydata
|
|
```
|
|
|
|
`global.log-level` is the only standing preference supported today (one-shot
|
|
commands like `--mount`/`--run`/`--user` don't belong in a config file). An explicit
|
|
`--log-level` on the command line always overrides the config file. The `volumes`
|
|
section is managed by `-v/--volume` (see above) rather than hand-edited — it's what
|
|
`-r/--run`'s own `-v` usage looks named volumes up in. A missing config file is fine
|
|
either way (nothing is overridden, and one gets created the first time
|
|
`-v/--volume` is used).
|
|
|
|
## How it works
|
|
|
|
Image layers are imported into `containers-storage` (parent-chained) and the
|
|
resulting top layer is mounted via `fuse-overlayfs`; `-r/--run` then sandboxes the
|
|
requested command with `bwrap`. Because `containers-storage mount` runs rootless by
|
|
re-execing into a private user+mount namespace, `-r/--run` normally has to `nsenter`
|
|
into that namespace to reach the mount — except when running as root, where the mount
|
|
is already directly visible and `--unshare-user` is skipped entirely (a fresh user
|
|
namespace isn't needed for root's own privilege, and forces an unrelated
|
|
supplementary-group bug in that case). Running as root also unlocks `--user`/
|
|
`--group`: since `bwrap --uid`/`--gid` require a user namespace that isn't available
|
|
there, `slocker-lite` instead bind-mounts a separate, statically-linked helper
|
|
(`slocker-lite-priv-drop`) into the sandbox and routes the command through it to drop
|
|
privileges before exec.
|
|
|
|
While a `-r/--run` session is active, its `bwrap` process is tracked as a locked
|
|
PID file under `$XDG_STATE_HOME/slocker-lite/run/` (falling back to
|
|
`$HOME/.local/state/...`), named after the image and its PID so the same image can
|
|
be run concurrently without collisions. The file is removed automatically once the
|
|
run ends; any tool can check whether a session is still alive by attempting the
|
|
same exclusive, non-blocking `flock()` on its file. `--list-processes` does
|
|
exactly that for every pid file it finds, reporting each one's pid, container
|
|
name, and `running`/`exited` status. Normally the file is removed automatically
|
|
when its own session ends, but `--clean-processes` removes any stale ones left
|
|
behind (e.g. after a crash) using that same check, atomically per file, so it
|
|
never removes one that's still genuinely running.
|
|
|
|
`-e/--exec <pid>` joins a running session's namespaces with `nsenter` and runs a
|
|
command there. Because `bwrap` itself sets up the sandbox's mount/user namespaces
|
|
and then hands the actual sandboxed command off to a child process in fresh
|
|
pid/uts/ipc/cgroup namespaces, `-e` resolves that real child first (via `/proc`)
|
|
rather than joining the outer `bwrap` process's own namespaces, so the joined
|
|
command sees the container's process tree and hostname too, not just its
|
|
filesystem.
|
|
|
|
See `CLAUDE.md` for the full architecture writeup (file-by-file breakdown, the
|
|
reasoning behind each of the above, and known gaps).
|
|
|
|
## License
|
|
|
|
GPL-2.0-or-later. See [`COPYING`](COPYING).
|