2
0
mirror of https://github.com/xcat2/xcat-core.git synced 2026-09-05 04:27:55 +00:00
Files
xcat-core/xCAT-test/unit
Daniel Hilst ce709cbeb0 test(xcat-core): nothing catches the reaper being handed a stale pid
supervise() unblocks SIGCHLD and then hands the pid back for the caller to assign, so the
caller's ($mon_respawn, $pid_MON) = ... still lands after the signal is let back in -- the
exact window e0b0ac6 closed, reopened by moving the sequence into a function. A monitor that
dies in that gap is compared by ssl_reaper against a $pid_MON still holding 0, missed, and
the caller then writes the dead pid back over the reaper's work: !$pid_MON never fires again
and the monitor is never respawned. That is this PR's own failure, reached through the
respawn rather than through startup, and nothing in the file notices it.

Add a subtest that watches the window from inside. Racing a real death into it is not
something a test can arrange reliably, so it arranges a certainty instead: a decoy child is
forked and exits with SIGCHLD blocked, leaving the signal pending, so the handler is
guaranteed to run the moment supervise() unblocks -- inside supervise(), before it has
returned. What the handler sees there is what ssl_reaper would see: it must find the live pid
and a pacing state that already knows a child was forked.

Separately, the fork-and-port subtest's closing assertion could not fail. It asks that the
respawn after a healthy monitor's death lands within 2 seconds, with max_interval set to 2 --
so a backoff pinned at the ceiling satisfies it too, and deleting the healthy-run reset from
exited() leaves the subtest green. Raise the ceiling to 8, where only the reset can produce a
prompt respawn, and assert first that the monitor being killed had actually been up long
enough to count as having served, which is the premise the assertion rests on.

Verified: against the current supervise() the new subtest fails on both assertions ("got 0,
expected <pid>"), and the raised ceiling turns the closing assertion red (8 <= 2) when the
healthy branch of exited() is removed -- where at a ceiling of 2 it stayed green.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-09-01 20:02:18 -03:00
..

xCAT-test/unit

Unit tests. These run against the source tree only -- no xCAT installation, no running daemons, no management node.

They are executed on every pull request by the xcat_test GitHub Actions workflow, which calls run_unit_tests() in github_action_xcat_test.pl:

prove -r xCAT-test/unit

You can run exactly the same thing from a clean checkout:

cd <xcat-core checkout>
prove -r xCAT-test/unit

What belongs here

A test belongs in unit/ when everything it needs is in the checkout: plugin and library sources, kickstart/preseed/subiquity templates, postscripts, packaging metadata. Such a test asserts on rendered output or module logic and reaches the repository root through FindBin:

use FindBin;
use lib "$FindBin::Bin/../../perl-xCAT";
use lib "$FindBin::Bin/../../xCAT-server/lib/perl";

Because of those FindBin paths the tests only work from a source tree. The copy installed under /opt/xcat/share/xcat/tools/autotest/unit is not a substitute -- ../.. resolves to /opt/xcat/share/xcat/tools there and the tests die or silently skip. The CI takes a copy of the checkout before the build for this reason; see preserve_source_tree().

What does not belong here

Anything that needs an installed xCAT, a populated /install, a real service binary or a live daemon. Those go in ../integration and run on a management node through xcattest. Both suites run on every pull request -- the workflow installs xCAT on the runner and then runs the ci_test cases against it -- so putting a test in integration/ does not cost it CI coverage. What differs is what each suite is allowed to depend on, and that unit tests also run standalone from a bare checkout with no xCAT at all.

The distinction matters because a test that needs an absent environment does not fail -- it calls plan skip_all and reports as skipped. A handful of those in a suite of several hundred assertions is easy to stop reading. Keeping the two kinds in separate directories means a skip in unit/ is a real signal rather than routine noise.

Guarding on a source file, on the other hand, is fine and common here:

plan skip_all => "compute.subiquity.tmpl not found" unless -f $tmpl_path;

That guard never fires when the tree is intact.