supervise() blocks SIGCHLD across the fork but unblocks it before returning, and the caller
installs the pid afterwards:
($mon_respawn, $pid_MON) = xCAT::RespawnUtils::supervise { ... } ...;
so the assignment is outside the blocked region -- the same unprotected window that existed
before e0b0ac6, moved from xcatd into the helper that was meant to make it impossible to get
wrong. ssl_reaper matches the dead child against $pid_MON and folds the death into
$mon_respawn; a monitor dying in that gap is compared against a pid still holding 0, missed,
and the caller then overwrites both with a pid that no longer exists. !$pid_MON never fires
again, so the respawn loop never runs and xcatiport stays dead until xcatd is restarted --
the failure this PR exists to remove.
Have supervise() install them itself, which is why `state` and `pid` are now passed by
reference: the pacing state is recorded and the pid assigned while SIGCHLD is still blocked,
and only then is it unblocked, so there is no point at which a reaper can run and see either
of them stale. Nothing is left for the caller to do afterwards, so both call sites become
plain statements that read $pid_MON when they need it. The child unblocks before running its
body, as it did when the unblock sat ahead of the fork's branch. The new pid is returned as
well, for a caller that wants it inline.
Verified on a live MN (xcat54-mn, AlmaLinux 10.2, xCAT 2.19.0): the startup fork produces a
monitor holding xcatiport 3002; killing it is recovered in 5s, killing the replacement at
once in 11s -- the backoff -- and killing one that had served past the healthy interval is
recovered in 1s, with the port reclaimed and xcatd active throughout. The unit test's window
subtest, red in the preceding commit, now passes.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
xCAT-test/unit
Unit tests. These run against the source tree only -- no xCAT installation, no running daemons, no management node.
They are executed on every pull request by the xcat_test GitHub Actions workflow,
which calls run_unit_tests() in github_action_xcat_test.pl:
prove -r xCAT-test/unit
You can run exactly the same thing from a clean checkout:
cd <xcat-core checkout>
prove -r xCAT-test/unit
What belongs here
A test belongs in unit/ when everything it needs is in the checkout: plugin and
library sources, kickstart/preseed/subiquity templates, postscripts, packaging
metadata. Such a test asserts on rendered output or module logic and reaches the
repository root through FindBin:
use FindBin;
use lib "$FindBin::Bin/../../perl-xCAT";
use lib "$FindBin::Bin/../../xCAT-server/lib/perl";
Because of those FindBin paths the tests only work from a source tree. The copy
installed under /opt/xcat/share/xcat/tools/autotest/unit is not a substitute --
../.. resolves to /opt/xcat/share/xcat/tools there and the tests die or silently
skip. The CI takes a copy of the checkout before the build for this reason; see
preserve_source_tree().
What does not belong here
Anything that needs an installed xCAT, a populated /install, a real service binary
or a live daemon. Those go in ../integration and run on
a management node through xcattest. Both suites run on every pull request -- the
workflow installs xCAT on the runner and then runs the ci_test cases against it --
so putting a test in integration/ does not cost it CI coverage. What differs is what
each suite is allowed to depend on, and that unit tests also run standalone from a
bare checkout with no xCAT at all.
The distinction matters because a test that needs an absent environment does not fail
-- it calls plan skip_all and reports as skipped. A handful of those in a suite of
several hundred assertions is easy to stop reading. Keeping the two kinds in separate
directories means a skip in unit/ is a real signal rather than routine noise.
Guarding on a source file, on the other hand, is fine and common here:
plan skip_all => "compute.subiquity.tmpl not found" unless -f $tmpl_path;
That guard never fires when the tree is intact.