Fix a race in create_session_cgroup(); skip the straggler test when unreachable
create_session_cgroup() used to be called from run_bwrap()'s on_start callback, in the parent, concurrently with the just-forked child execing into bwrap and bwrap then doing its own internal clone() of the sandboxed target. Without --unshare-pid, bwrap has little enough setup work to do that it could reliably win that race, cloning its target before the parent's own write into cgroup.procs completed -- leaving that target, and everything it later spawns, permanently outside the tracked cgroup, so the post-exit sweep found nothing to reap. Found via the user's own request to test "pid namespace off, cgroup on" as root: confirmed directly by inspecting cgroup.procs mid-session, showing only bwrap's own pid. Fixed by giving run_process_foreground() a new before_exec hook, invoked in the child synchronously right before execvp() -- the child cannot proceed to exec (and thus cannot trigger any of bwrap's own internal forking) until this has already returned, closing the race structurally rather than by timing luck. run_bwrap() now creates the session cgroup there instead of in on_start; the parent side just reconstructs the deterministic path unconditionally, since the downstream sweep/cleanup functions already tolerate a nonexistent directory gracefully either way. Also fixes a false positive found while verifying this: the regression test's own process-matching did a substring search across a whole cmdline blob, which matched an unrelated manual `pkill -f 'sleep 137'` diagnostic command run by hand during the investigation. Tightened to an exact argv[0]/argv[1] match. Finally, the regression test now SKIP()s (instead of failing) when neither a pid namespace nor a working session cgroup is available for the current effective config -- a documented, known residual limitation, not a regression -- checked directly via two new helpers rather than assumed from e.g. geteuid().
This commit is contained in:
@@ -458,7 +458,7 @@ Source layout (all under `src/`):
|
|||||||
own automatic post-exit straggler sweep** (`bwrap.cpp`/`session_cgroup.h`'s
|
own automatic post-exit straggler sweep** (`bwrap.cpp`/`session_cgroup.h`'s
|
||||||
"Resolved" entries), reproducing the user's own reported shape end to
|
"Resolved" entries), reproducing the user's own reported shape end to
|
||||||
end with a real container: `nohup sleep 137 & exit` inside the sandbox,
|
end with a real container: `nohup sleep 137 & exit` inside the sandbox,
|
||||||
then `any_process_cmdline_contains()`/`process_cmdline_gone_within()`
|
then `sleep_process_running()`/`sleep_process_gone_within()`
|
||||||
(`.cpp`-local, same `/proc`-scanning shape as `bwrap.cpp`'s own
|
(`.cpp`-local, same `/proc`-scanning shape as `bwrap.cpp`'s own
|
||||||
`find_fuse_overlayfs_pid()`) confirm the backgrounded process is gone
|
`find_fuse_overlayfs_pid()`) confirm the backgrounded process is gone
|
||||||
from the *host's* own process table afterward — checked by cmdline
|
from the *host's* own process table afterward — checked by cmdline
|
||||||
@@ -475,6 +475,33 @@ Source layout (all under `src/`):
|
|||||||
`[integration][root]` test (below) for one that exercises
|
`[integration][root]` test (below) for one that exercises
|
||||||
`kill_via_cgroup()` directly, since the real target device's own
|
`kill_via_cgroup()` directly, since the real target device's own
|
||||||
no-pid-namespace escape shape can't be forced via the CLI at all.
|
no-pid-namespace escape shape can't be forced via the CLI at all.
|
||||||
|
**Named `sleep_process_*` (not the original `any_process_cmdline_*`) since
|
||||||
|
a real false positive was found via this exact test session**: the
|
||||||
|
original version searched for the literal text `"sleep 137"` anywhere in
|
||||||
|
a candidate process's cmdline blob, which matched a manual diagnostic
|
||||||
|
`pkill -f 'sleep 137'` cleanup command run by hand while investigating a
|
||||||
|
separate issue — that command's own cmdline literally contains the
|
||||||
|
search text as a `pkill` pattern argument, despite not being a sleep
|
||||||
|
process at all. Fixed by requiring an exact match instead: `argv[0]`'s
|
||||||
|
basename is exactly `sleep` and `argv[1]` exactly equals the expected
|
||||||
|
duration. **`SKIP()`s when neither a pid namespace nor a working session
|
||||||
|
cgroup is available** under the current effective config
|
||||||
|
(`g_test_app_config`, `fixtures.h`) and live kernel capability — checked
|
||||||
|
via `pid_namespace_would_isolate()` (the same two-gate policy-and-kernel-
|
||||||
|
support check `build_bwrap_args()` itself applies) and
|
||||||
|
`session_cgroup_would_work()` (tries `create_directories()` +
|
||||||
|
`access(..., W_OK)` against a throwaway, never-joined probe path under
|
||||||
|
`session_cgroup_path()` — tried directly rather than guessed from e.g.
|
||||||
|
`geteuid()`, since the real failure mode this needs to detect, no
|
||||||
|
delegated subtree when rootless, is exactly a permissions question those
|
||||||
|
calls can answer directly; never writes this process's own pid into the
|
||||||
|
probe directory, so cleanup is just removing the still-empty directory).
|
||||||
|
In that combination there is genuinely no mechanism left that could reap
|
||||||
|
a reparented straggler — the documented, known residual limitation
|
||||||
|
(`session_cgroup.h`'s own entry above), not a regression — added
|
||||||
|
specifically per the user's own follow-up request after using `-c` to
|
||||||
|
force exactly that combination and correctly getting a failure instead
|
||||||
|
of a skip.
|
||||||
- `tests/integration/test_root_networking.cpp` (`[integration][root][net]`)
|
- `tests/integration/test_root_networking.cpp` (`[integration][root][net]`)
|
||||||
— the persistent-netns/tap-relay/dns-resolver tests that originally
|
— the persistent-netns/tap-relay/dns-resolver tests that originally
|
||||||
lived directly in this file, ported to tagged `TEST_CASE`s (every
|
lived directly in this file, ported to tagged `TEST_CASE`s (every
|
||||||
@@ -1901,6 +1928,73 @@ Source layout (all under `src/`):
|
|||||||
still has no automatic way to reach a reparented straggler — the same
|
still has no automatic way to reach a reparented straggler — the same
|
||||||
fundamental gap `kill_via_tracked_pid()` (the weakest of `--kill`'s own
|
fundamental gap `kill_via_tracked_pid()` (the weakest of `--kill`'s own
|
||||||
three strategies, below) already represents.
|
three strategies, below) already represents.
|
||||||
|
|
||||||
|
**Resolved: a genuine race could leave the sweep unable to reach the
|
||||||
|
straggler at all, specifically without a pid namespace.** Found by the
|
||||||
|
user's own explicit follow-up request to test the "pid namespace off,
|
||||||
|
cgroup on" combination as root (via `-c`/`--config-file`'s now-wired-in
|
||||||
|
effect on `-t`, see `tests/support/fixtures.h`'s own entry below) — the
|
||||||
|
nohup-straggler regression test (`test_rootless_run.cpp`) failed
|
||||||
|
*consistently*, even as root with cgroup v2 genuinely available and
|
||||||
|
writable. Root-caused directly, not assumed: inspecting
|
||||||
|
`/sys/fs/cgroup/slocker-lite/<session>/cgroup.procs` *while the session was
|
||||||
|
still running* showed only bwrap's own outer pid — its own child (the
|
||||||
|
actual sandboxed `sh` process, and the `sleep` it later backgrounded) was
|
||||||
|
never a member at all. Cause: `create_session_cgroup()` used to be called
|
||||||
|
from `run_bwrap()`'s `on_start` callback, which runs in the *parent*,
|
||||||
|
concurrently with the just-`fork()`'d child execing into bwrap and bwrap
|
||||||
|
then doing its *own* internal `clone()` of the sandboxed target — a genuine
|
||||||
|
race between the parent's own directory-create-plus-write (several
|
||||||
|
syscalls) and bwrap's own setup. Without `--unshare-pid`, bwrap has
|
||||||
|
meaningfully less setup work to do (no new pid namespace to create), making
|
||||||
|
it reliably fast enough to win that race and clone its target *before* the
|
||||||
|
parent's own write into `cgroup.procs` ever completed — leaving that
|
||||||
|
target, and everything it later spawns, permanently outside the tracked
|
||||||
|
cgroup. This apparently didn't manifest with `--unshare-pid` requested
|
||||||
|
(bwrap's extra setup work there was consistently slow enough for the
|
||||||
|
parent to win instead) or on the real target device (confirmed working
|
||||||
|
there earlier, likely for the same reason, or because cgroup delegation
|
||||||
|
wasn't actually the mechanism catching it there) — but was never a
|
||||||
|
guaranteed property either way, just a timing coincidence.
|
||||||
|
|
||||||
|
**Fix**: `run_process_foreground()` (`process.{h,cpp}`) gained a new
|
||||||
|
`before_exec` parameter — a callback invoked *in the child*, synchronously,
|
||||||
|
immediately before `execvp()` — alongside the existing `on_start` (which
|
||||||
|
still runs in the parent, unchanged, for the session lock and
|
||||||
|
`on_bwrap_pid_known`). `run_bwrap()` now creates the session cgroup from
|
||||||
|
*inside* `before_exec` (`create_session_cgroup(container_name, getpid())`)
|
||||||
|
instead of from `on_start` — since the child cannot proceed to `execvp()`
|
||||||
|
(and thus cannot start any of bwrap's own internal forking) until this
|
||||||
|
call has already returned, the race is closed structurally, not by timing
|
||||||
|
luck. The parent's own `on_start` callback still reconstructs the
|
||||||
|
(fully deterministic) `SessionCgroup{session_cgroup_path(container_name,
|
||||||
|
pid)}` unconditionally, regardless of whether the child-side creation
|
||||||
|
actually succeeded — `session_cgroup_pids()`/`remove_session_cgroup()`
|
||||||
|
already tolerate a nonexistent directory gracefully either way, the same
|
||||||
|
as they already did for any other `create_session_cgroup()` failure (e.g.
|
||||||
|
no cgroup v2 delegation when rootless). Ordinary (non-async-signal-safe)
|
||||||
|
work in a post-`fork()`/pre-`exec()` child is safe here since this project
|
||||||
|
has no threads — the same reasoning `daemonize()`'s own child branch
|
||||||
|
already relies on for its own, more extensive pre-exec setup.
|
||||||
|
|
||||||
|
**Verified end-to-end on this dev machine, root, via the scoped `doas`
|
||||||
|
rule**: reproduced the failure consistently (3/3 trials, each from a
|
||||||
|
freshly-cleaned state) against a `-c`-supplied config with
|
||||||
|
`unshare-pid: false` and every other field at its default, using the
|
||||||
|
live `cgroup.procs` inspection above to confirm the exact mechanism; after
|
||||||
|
the fix, the identical reproduction passed consistently (3/3 isolated
|
||||||
|
trials, plus repeated full-category runs). **A second, unrelated bug
|
||||||
|
surfaced and was fixed while verifying this**: the regression test's own
|
||||||
|
`any_process_cmdline_contains()` did a substring search across a
|
||||||
|
candidate process's *entire* cmdline blob, which could false-positive
|
||||||
|
against an unrelated process that merely *mentions* the marker text
|
||||||
|
somewhere in its own arguments — confirmed directly: a manual diagnostic
|
||||||
|
`pkill -f 'sleep 137'` cleanup command run by hand during this same
|
||||||
|
investigation was itself briefly matched as if it were the sleep process.
|
||||||
|
Renamed to `sleep_process_running()`/`sleep_process_gone_within()` and
|
||||||
|
tightened to require `argv[0]`'s basename to be exactly `sleep` and
|
||||||
|
`argv[1]` to exactly match the expected duration, eliminating that class
|
||||||
|
of false positive.
|
||||||
- `sandbox_process.{h,cpp}` — process-tree/namespace-resolution utilities
|
- `sandbox_process.{h,cpp}` — process-tree/namespace-resolution utilities
|
||||||
shared by `exec_session.{h,cpp}` and `kill_session.{h,cpp}` (see both
|
shared by `exec_session.{h,cpp}` and `kill_session.{h,cpp}` (see both
|
||||||
below); pulled into their own file (rather than staying private to
|
below); pulled into their own file (rather than staying private to
|
||||||
|
|||||||
@@ -45,13 +45,23 @@ free from the kernel itself (killing a pid namespace's own pid 1 -- whether
|
|||||||
via a normal exit or a forced kill -- collapses the whole namespace
|
via a normal exit or a forced kill -- collapses the whole namespace
|
||||||
regardless of this fix); the scenario this fix specifically targets --
|
regardless of this fix); the scenario this fix specifically targets --
|
||||||
`--unshare-pid` unavailable or disabled, so a daemonizing process reparents
|
`--unshare-pid` unavailable or disabled, so a daemonizing process reparents
|
||||||
completely outside any namespace -- couldn't be exercised end-to-end
|
completely outside any namespace -- originally couldn't be exercised
|
||||||
locally (no CLI-level way to force that off for a single run; it's a
|
end-to-end locally (no CLI-level way to force that off for a single run;
|
||||||
config-file-only `NamespaceConfig` field), and needs on-device
|
it was a config-file-only `NamespaceConfig` field).
|
||||||
confirmation, same as everything else in this project that depends on the
|
|
||||||
real target's own kernel capabilities. Whether the real device's cgroup v2
|
**Since resolved**: `-c/--config-file` (added later) made this directly
|
||||||
is actually writable as root (the case that matters there) also hasn't
|
testable locally after all, and doing exactly that (per the user's own
|
||||||
been specifically re-verified yet.
|
explicit request) found a real, separate bug -- a race between this fix's
|
||||||
|
own `create_session_cgroup()` call and bwrap's own internal forking, which
|
||||||
|
could leave the sweep unable to reach the straggler at all specifically
|
||||||
|
without a pid namespace, even as root with cgroup v2 genuinely available.
|
||||||
|
See `CLAUDE.md`'s `session_cgroup.{h,cpp}` entry for the full root-cause
|
||||||
|
and fix (`process.h`'s new `before_exec` hook). Re-verified end-to-end
|
||||||
|
(root, this dev machine, `-c <unshare-pid: false>`) after the fix: 3/3
|
||||||
|
isolated trials and repeated full-category runs all correctly reaped the
|
||||||
|
straggler. On-device re-confirmation (the real target device has neither
|
||||||
|
pid namespace nor previously-confirmed cgroup delegation) is still
|
||||||
|
worthwhile but no longer the only way to exercise this path locally.
|
||||||
|
|
||||||
**Known residual limitation**: this only helps when `create_session_cgroup()`
|
**Known residual limitation**: this only helps when `create_session_cgroup()`
|
||||||
actually succeeded (cgroup v2 mounted and writable). A kernel with *neither*
|
actually succeeded (cgroup v2 mounted and writable). A kernel with *neither*
|
||||||
|
|||||||
+30
-7
@@ -397,17 +397,40 @@ int run_bwrap(const std::string& root, const std::vector<std::string>& command,
|
|||||||
*argv,
|
*argv,
|
||||||
[&](pid_t pid) {
|
[&](pid_t pid) {
|
||||||
session_lock = create_session_lock(container_name, pid);
|
session_lock = create_session_lock(container_name, pid);
|
||||||
// Best-effort, like the session lock above -- lets kill_session()
|
// The actual cgroup creation/membership write happens
|
||||||
// (kill_session.cpp) reach every process this session ever starts,
|
// synchronously in the child itself, via before_exec below,
|
||||||
// including anything that later daemonizes/double-forks and gets
|
// *before* it execs into bwrap -- doing it here instead (in the
|
||||||
// reparented, regardless of whether the kernel supports pid
|
// parent, concurrently with the child) raced against bwrap's own
|
||||||
// namespaces at all. See session_cgroup.h.
|
// internal fork of the sandboxed target and could lose: confirmed
|
||||||
session_cgroup = create_session_cgroup(container_name, pid);
|
// by testing that without --unshare-pid, bwrap has little enough
|
||||||
|
// setup work to do that it can clone its target before this
|
||||||
|
// callback's own write into cgroup.procs ever completed, leaving
|
||||||
|
// that target -- and everything it later spawns -- outside the
|
||||||
|
// tracked cgroup entirely (session_cgroup_pids() showing only
|
||||||
|
// bwrap's own pid, never its child's). The path itself is fully
|
||||||
|
// deterministic from (container_name, pid), so it's reconstructed
|
||||||
|
// here unconditionally regardless of whether the child-side
|
||||||
|
// creation actually succeeded -- session_cgroup_pids()/
|
||||||
|
// remove_session_cgroup() below already tolerate a nonexistent
|
||||||
|
// directory gracefully either way, the same as they already did
|
||||||
|
// for a session_lock/session_cgroup that failed to create for
|
||||||
|
// some other reason (e.g. no cgroup v2 delegation when rootless).
|
||||||
|
session_cgroup = SessionCgroup{session_cgroup_path(container_name, pid)};
|
||||||
if (on_bwrap_pid_known) {
|
if (on_bwrap_pid_known) {
|
||||||
on_bwrap_pid_known(pid);
|
on_bwrap_pid_known(pid);
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
build_sandbox_env(user, extra_env));
|
build_sandbox_env(user, extra_env),
|
||||||
|
[&]() {
|
||||||
|
// Runs in the child, before execvp() -- see this function's own
|
||||||
|
// before_exec doc comment (process.h) for why this must happen
|
||||||
|
// here rather than in the on_start callback above. getpid() here
|
||||||
|
// is this same child's own pid, matching exactly the `pid`
|
||||||
|
// on_start was (or will be) called with -- exec() never changes
|
||||||
|
// a process's pid, so this is the same identity bwrap itself
|
||||||
|
// will have.
|
||||||
|
create_session_cgroup(container_name, getpid());
|
||||||
|
});
|
||||||
|
|
||||||
// Runs unconditionally, regardless of *why* run_process_foreground() just
|
// Runs unconditionally, regardless of *why* run_process_foreground() just
|
||||||
// returned (normal exit, or bwrap forwarded a SIGINT/SIGTERM) -- covers a
|
// returned (normal exit, or bwrap forwarded a SIGINT/SIGTERM) -- covers a
|
||||||
|
|||||||
+5
-1
@@ -109,7 +109,8 @@ ProcessResult run_process(const std::vector<std::string>& argv) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
int run_process_foreground(const std::vector<std::string>& argv, const std::function<void(pid_t)>& on_start,
|
int run_process_foreground(const std::vector<std::string>& argv, const std::function<void(pid_t)>& on_start,
|
||||||
const std::optional<std::vector<std::pair<std::string, std::string>>>& env) {
|
const std::optional<std::vector<std::pair<std::string, std::string>>>& env,
|
||||||
|
const std::function<void()>& before_exec) {
|
||||||
spdlog::debug("running external command: {}", fmt::join(argv, " "));
|
spdlog::debug("running external command: {}", fmt::join(argv, " "));
|
||||||
|
|
||||||
pid_t pid = fork();
|
pid_t pid = fork();
|
||||||
@@ -124,6 +125,9 @@ int run_process_foreground(const std::vector<std::string>& argv, const std::func
|
|||||||
setenv(key.c_str(), value.c_str(), 1);
|
setenv(key.c_str(), value.c_str(), 1);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
if (before_exec) {
|
||||||
|
before_exec();
|
||||||
|
}
|
||||||
auto c_argv = to_c_argv(argv);
|
auto c_argv = to_c_argv(argv);
|
||||||
execvp(c_argv[0], c_argv.data());
|
execvp(c_argv[0], c_argv.data());
|
||||||
const char* msg = "run_process_foreground: execvp failed\n";
|
const char* msg = "run_process_foreground: execvp failed\n";
|
||||||
|
|||||||
+17
-3
@@ -45,7 +45,21 @@ ProcessResult run_process(const std::vector<std::string>& argv);
|
|||||||
// child replaces its entire environment with exactly these key/value pairs (via
|
// child replaces its entire environment with exactly these key/value pairs (via
|
||||||
// clearenv()/setenv(), before exec) instead of inheriting this process's own --
|
// clearenv()/setenv(), before exec) instead of inheriting this process's own --
|
||||||
// e.g. run_bwrap() uses this to give the sandboxed command a minimal, controlled
|
// e.g. run_bwrap() uses this to give the sandboxed command a minimal, controlled
|
||||||
// environment without relying on bwrap's own --clearenv/--setenv. Returns the
|
// environment without relying on bwrap's own --clearenv/--setenv. If `before_exec`
|
||||||
|
// is set, it's called *in the child*, synchronously, immediately before execvp()
|
||||||
|
// -- for setup that must be guaranteed complete before the target program starts
|
||||||
|
// doing anything of its own, which `on_start` (running concurrently, in the
|
||||||
|
// parent) can't guarantee: e.g. run_bwrap() uses this to put the child into its
|
||||||
|
// own session cgroup (session_cgroup.h) before it execs into bwrap, since doing
|
||||||
|
// that from the parent's own `on_start` instead raced against bwrap's own
|
||||||
|
// internal fork of the sandboxed target and could lose (confirmed by testing:
|
||||||
|
// without --unshare-pid, bwrap can clone its target before the parent's own
|
||||||
|
// write into cgroup.procs completes, leaving that target -- and everything it
|
||||||
|
// later spawns -- outside the tracked cgroup entirely). Only async-signal-safe-ish
|
||||||
|
// work belongs here in the strictest POSIX sense, but this project has no
|
||||||
|
// threads, so ordinary calls (filesystem I/O, etc.) are safe in practice --
|
||||||
|
// the same reasoning daemonize()'s own child branch (daemonize.cpp) already
|
||||||
|
// relies on. Returns the
|
||||||
// exit code if the child exited normally; -1 only if fork() itself failed
|
// exit code if the child exited normally; -1 only if fork() itself failed
|
||||||
// (on_start is not called in that case) -- a failed execvp inside the child
|
// (on_start is not called in that case) -- a failed execvp inside the child
|
||||||
// is a normal exit with code 127, not -1. If the child was instead killed by
|
// is a normal exit with code 127, not -1. If the child was instead killed by
|
||||||
@@ -61,8 +75,8 @@ ProcessResult run_process(const std::vector<std::string>& argv);
|
|||||||
// unexpected crash.
|
// unexpected crash.
|
||||||
int run_process_foreground(const std::vector<std::string>& argv,
|
int run_process_foreground(const std::vector<std::string>& argv,
|
||||||
const std::function<void(pid_t)>& on_start = nullptr,
|
const std::function<void(pid_t)>& on_start = nullptr,
|
||||||
const std::optional<std::vector<std::pair<std::string, std::string>>>& env =
|
const std::optional<std::vector<std::pair<std::string, std::string>>>& env = std::nullopt,
|
||||||
std::nullopt);
|
const std::function<void()>& before_exec = nullptr);
|
||||||
|
|
||||||
// Searches $PATH for an executable regular file named `name`, in PATH order.
|
// Searches $PATH for an executable regular file named `name`, in PATH order.
|
||||||
// Returns its full path, or nullopt if not found.
|
// Returns its full path, or nullopt if not found.
|
||||||
|
|||||||
@@ -66,6 +66,7 @@
|
|||||||
#include "commands.h"
|
#include "commands.h"
|
||||||
#include "config_file.h"
|
#include "config_file.h"
|
||||||
#include "fixtures.h"
|
#include "fixtures.h"
|
||||||
|
#include "session_cgroup.h"
|
||||||
|
|
||||||
namespace {
|
namespace {
|
||||||
|
|
||||||
@@ -141,17 +142,39 @@ std::string run_in_fixture(const std::filesystem::path& image, const std::vector
|
|||||||
return capture.contents();
|
return capture.contents();
|
||||||
}
|
}
|
||||||
|
|
||||||
// Scans /proc for any process whose cmdline contains `needle` -- used below
|
// Splits a raw /proc/<pid>/cmdline blob (NUL-separated, per Linux's own
|
||||||
// to confirm a backgrounded process detached inside a container doesn't
|
// documented format) into its individual arguments.
|
||||||
// survive the session, checked from the *host's* own process table. This
|
std::vector<std::string> split_cmdline_args(const std::string& raw) {
|
||||||
// works regardless of whether the session's own pid namespace is isolated
|
std::vector<std::string> args;
|
||||||
// or not: every process, wherever it lives namespace-wise, is still a
|
size_t start = 0;
|
||||||
// perfectly ordinary task on the host with its own real pid and
|
while (start < raw.size()) {
|
||||||
// /proc/<pid>/cmdline entry -- only the pid *number* a process sees for
|
size_t end = raw.find('\0', start);
|
||||||
// itself differs inside an isolated namespace, not whether it shows up
|
if (end == std::string::npos) {
|
||||||
// here at all. Same directory-scanning shape as bwrap.cpp's own
|
end = raw.size();
|
||||||
// find_fuse_overlayfs_pid().
|
}
|
||||||
bool any_process_cmdline_contains(const std::string& needle) {
|
args.push_back(raw.substr(start, end - start));
|
||||||
|
start = end + 1;
|
||||||
|
}
|
||||||
|
return args;
|
||||||
|
}
|
||||||
|
|
||||||
|
// True if some process on the *host* is genuinely running `sleep <arg>` --
|
||||||
|
// checked from the host's own process table (works regardless of whether
|
||||||
|
// the session's own pid namespace is isolated or not: every process,
|
||||||
|
// wherever it lives namespace-wise, is still a perfectly ordinary task on
|
||||||
|
// the host with its own real pid and /proc/<pid>/cmdline entry -- only the
|
||||||
|
// pid *number* a process sees for itself differs inside an isolated
|
||||||
|
// namespace, not whether it shows up here at all). Deliberately checks
|
||||||
|
// argv[0]'s basename and argv[1] exactly, not a substring search across the
|
||||||
|
// whole cmdline -- a real bug found via this exact test session: an early
|
||||||
|
// version searched for the literal text "sleep 137" anywhere in the
|
||||||
|
// cmdline, which produced a false positive by matching this test's *own*
|
||||||
|
// diagnostic `pkill -f 'sleep 137'` cleanup commands (run by hand, outside
|
||||||
|
// the test, while investigating a separate issue) -- their own cmdline
|
||||||
|
// literally contains that text as a pkill pattern argument, despite not
|
||||||
|
// being a sleep process at all. Same directory-scanning shape as
|
||||||
|
// bwrap.cpp's own find_fuse_overlayfs_pid().
|
||||||
|
bool sleep_process_running(const std::string& arg) {
|
||||||
std::error_code ec;
|
std::error_code ec;
|
||||||
auto it = std::filesystem::directory_iterator("/proc", ec);
|
auto it = std::filesystem::directory_iterator("/proc", ec);
|
||||||
if (ec) {
|
if (ec) {
|
||||||
@@ -163,27 +186,28 @@ bool any_process_cmdline_contains(const std::string& needle) {
|
|||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
std::ifstream cmdline_file(entry.path() / "cmdline", std::ios::binary);
|
std::ifstream cmdline_file(entry.path() / "cmdline", std::ios::binary);
|
||||||
std::string cmdline((std::istreambuf_iterator<char>(cmdline_file)), std::istreambuf_iterator<char>());
|
std::string raw((std::istreambuf_iterator<char>(cmdline_file)), std::istreambuf_iterator<char>());
|
||||||
std::replace(cmdline.begin(), cmdline.end(), '\0', ' ');
|
auto args = split_cmdline_args(raw);
|
||||||
if (cmdline.find(needle) != std::string::npos) {
|
if (args.size() == 2 && std::filesystem::path(args[0]).filename() == "sleep" && args[1] == arg) {
|
||||||
return true;
|
return true;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return false;
|
return false;
|
||||||
}
|
}
|
||||||
|
|
||||||
// Bounded (2s, 50ms interval) poll for `needle` to disappear from the
|
// Bounded (2s, 50ms interval) poll for the `sleep <arg>` process to
|
||||||
// host's own process table -- a session ending via a pid namespace's own
|
// disappear from the host's own process table -- a session ending via a
|
||||||
// kernel-guaranteed collapse-on-pid-1-exit isn't necessarily synchronously
|
// pid namespace's own kernel-guaranteed collapse-on-pid-1-exit isn't
|
||||||
// complete by the instant run_in_fixture() above returns (that guarantee is
|
// necessarily synchronously complete by the instant run_in_fixture() above
|
||||||
// about eventual termination, not that every other task in the namespace
|
// returns (that guarantee is about eventual termination, not that every
|
||||||
// has already been fully reaped), so a single instantaneous check right
|
// other task in the namespace has already been fully reaped), so a single
|
||||||
// after could be flaky. The straggler sweep this test actually targets
|
// instantaneous check right after could be flaky. The straggler sweep this
|
||||||
// (run_bwrap()'s own post-exit cgroup sweep, bwrap.cpp) already applies a
|
// test actually targets (run_bwrap()'s own post-exit cgroup sweep,
|
||||||
// real grace period internally for the same reason.
|
// bwrap.cpp) already applies a real grace period internally for the same
|
||||||
bool process_cmdline_gone_within(const std::string& needle, int timeout_ms) {
|
// reason.
|
||||||
|
bool sleep_process_gone_within(const std::string& arg, int timeout_ms) {
|
||||||
for (int waited = 0; waited < timeout_ms; waited += 50) {
|
for (int waited = 0; waited < timeout_ms; waited += 50) {
|
||||||
if (!any_process_cmdline_contains(needle)) {
|
if (!sleep_process_running(arg)) {
|
||||||
return true;
|
return true;
|
||||||
}
|
}
|
||||||
struct timespec ts {
|
struct timespec ts {
|
||||||
@@ -192,7 +216,44 @@ bool process_cmdline_gone_within(const std::string& needle, int timeout_ms) {
|
|||||||
while (nanosleep(&ts, &ts) != 0 && errno == EINTR) {
|
while (nanosleep(&ts, &ts) != 0 && errno == EINTR) {
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
return !any_process_cmdline_contains(needle);
|
return !sleep_process_running(arg);
|
||||||
|
}
|
||||||
|
|
||||||
|
// Whether the current -t run's effective config (g_test_app_config,
|
||||||
|
// fixtures.h) would actually get bwrap to request --unshare-pid for a
|
||||||
|
// container it creates -- both the config's own policy *and* live kernel
|
||||||
|
// support (detect_bwrap_unshare_args(), bwrap.h) have to allow it, the same
|
||||||
|
// two-gate check build_bwrap_args() itself applies.
|
||||||
|
bool pid_namespace_would_isolate() {
|
||||||
|
if (!g_test_app_config.unshare_pid.value_or(true)) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
auto supported = detect_bwrap_unshare_args();
|
||||||
|
return std::find(supported.begin(), supported.end(), "--unshare-pid") != supported.end();
|
||||||
|
}
|
||||||
|
|
||||||
|
// Whether run_bwrap()'s own session cgroup (session_cgroup.h) would actually
|
||||||
|
// be creatable/writable right now -- tried directly against a throwaway,
|
||||||
|
// never-joined path rather than guessing from e.g. geteuid(), since the real
|
||||||
|
// failure mode this needs to detect (no delegated subtree when rootless) is
|
||||||
|
// exactly a permissions question create_directories()/access() can answer
|
||||||
|
// directly. Deliberately never writes this process's own pid into the
|
||||||
|
// probe directory's cgroup.procs (unlike the real create_session_cgroup()),
|
||||||
|
// so cleanup is just removing the still-empty directory -- no need to
|
||||||
|
// un-join anything first.
|
||||||
|
bool session_cgroup_would_work() {
|
||||||
|
if (!cgroup_v2_available()) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
auto probe_path = session_cgroup_path("selftest-cgroup-probe", getpid());
|
||||||
|
std::error_code ec;
|
||||||
|
std::filesystem::create_directories(probe_path, ec);
|
||||||
|
if (ec) {
|
||||||
|
return false;
|
||||||
|
}
|
||||||
|
bool writable = access((probe_path / "cgroup.procs").c_str(), W_OK) == 0;
|
||||||
|
std::filesystem::remove(probe_path, ec);
|
||||||
|
return writable;
|
||||||
}
|
}
|
||||||
|
|
||||||
} // namespace
|
} // namespace
|
||||||
@@ -295,14 +356,25 @@ TEST_CASE("rootless -r/--run: a nohup-backgrounded process does not survive the
|
|||||||
SKIP("no busybox fixture (images/busybox.tar) -- see tests/setup-tests.py");
|
SKIP("no busybox fixture (images/busybox.tar) -- see tests/setup-tests.py");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// With neither a pid namespace nor a working session cgroup available,
|
||||||
|
// there is genuinely no mechanism left that could reap a reparented
|
||||||
|
// straggler -- a documented, known limitation (CLAUDE.md's
|
||||||
|
// session_cgroup.{h,cpp} entry, "Known residual limitation"), not a
|
||||||
|
// regression this test should be flagging. Skip rather than fail so a
|
||||||
|
// deliberately-crippled config (e.g. via -c/--config-file, g_test_app_config
|
||||||
|
// above) doesn't produce a misleading failure here.
|
||||||
|
if (!pid_namespace_would_isolate() && !session_cgroup_would_work()) {
|
||||||
|
SKIP("neither a pid namespace nor a working session cgroup is available under the "
|
||||||
|
"current config/kernel -- no mechanism exists to reap a reparented straggler here");
|
||||||
|
}
|
||||||
|
|
||||||
// A distinctive duration -- not a realistic value anything else on this
|
// A distinctive duration -- not a realistic value anything else on this
|
||||||
// host would coincidentally already be sleeping for -- so scanning
|
// host would coincidentally already be sleeping for.
|
||||||
// /proc for it can't produce a false positive either way.
|
const std::string marker = "137";
|
||||||
const std::string marker = "sleep 137";
|
REQUIRE_FALSE(sleep_process_running(marker));
|
||||||
REQUIRE_FALSE(any_process_cmdline_contains(marker));
|
|
||||||
|
|
||||||
ScratchXdgDirs scratch;
|
ScratchXdgDirs scratch;
|
||||||
run_in_fixture(*image, {"sh", "-c", "nohup sleep 137 >/dev/null 2>&1 & exit"});
|
run_in_fixture(*image, {"sh", "-c", "nohup sleep 137 >/dev/null 2>&1 & exit"});
|
||||||
|
|
||||||
CHECK(process_cmdline_gone_within(marker, 2000));
|
CHECK(sleep_process_gone_within(marker, 2000));
|
||||||
}
|
}
|
||||||
|
|||||||
Reference in New Issue
Block a user