The respawn added for the install monitor is paced by a retry budget, and once that budget is spent the daemon stops trying. That reintroduces the failure the respawn exists to remove: with no monitor alive there is nothing left to reset the counter, so xcatiport stays dead until the whole daemon is restarted, and a port that becomes free a minute later is never picked back up. Pacing the retries is necessary -- an unguarded re-fork spins as fast as fork allows while the port is held, and keeps re-entering do_installm_service's USR2 socket-takeover handshake -- but pacing must not decay into giving up. The property that matters is therefore behavioural, not structural: the monitor comes back on its own, at a bounded rate, no matter how long it has been failing. Assert it by driving xcatd's real pacing code rather than grepping for it -- extract the marked mon-respawn-policy region from the script verbatim, the way build_ubunturepo_lock.t drives build-ubunturepo's real lock, and run it. Over a virtual clock, check that the delay backs off to a ceiling and holds there, that the daemon is still forking monitors three hours into a failure, and that a monitor which stayed up long enough to serve resets the pacing when it later dies. Then do it for real against a genuinely held TCP port: fail several times, release the port, and require that a respawned monitor binds it and stays up without the daemon being restarted. Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
xCAT-test/unit
Unit tests. These run against the source tree only -- no xCAT installation, no running daemons, no management node.
They are executed on every pull request by the xcat_test GitHub Actions workflow,
which calls run_unit_tests() in github_action_xcat_test.pl:
prove -r xCAT-test/unit
You can run exactly the same thing from a clean checkout:
cd <xcat-core checkout>
prove -r xCAT-test/unit
What belongs here
A test belongs in unit/ when everything it needs is in the checkout: plugin and
library sources, kickstart/preseed/subiquity templates, postscripts, packaging
metadata. Such a test asserts on rendered output or module logic and reaches the
repository root through FindBin:
use FindBin;
use lib "$FindBin::Bin/../../perl-xCAT";
use lib "$FindBin::Bin/../../xCAT-server/lib/perl";
Because of those FindBin paths the tests only work from a source tree. The copy
installed under /opt/xcat/share/xcat/tools/autotest/unit is not a substitute --
../.. resolves to /opt/xcat/share/xcat/tools there and the tests die or silently
skip. The CI takes a copy of the checkout before the build for this reason; see
preserve_source_tree().
What does not belong here
Anything that needs an installed xCAT, a populated /install, a real service binary
or a live daemon. Those go in ../integration and run on
a management node through xcattest. Both suites run on every pull request -- the
workflow installs xCAT on the runner and then runs the ci_test cases against it --
so putting a test in integration/ does not cost it CI coverage. What differs is what
each suite is allowed to depend on, and that unit tests also run standalone from a
bare checkout with no xCAT at all.
The distinction matters because a test that needs an absent environment does not fail
-- it calls plan skip_all and reports as skipped. A handful of those in a suite of
several hundred assertions is easy to stop reading. Keeping the two kinds in separate
directories means a skip in unit/ is a real signal rather than routine noise.
Guarding on a source file, on the other hand, is fine and common here:
plan skip_all => "compute.subiquity.tmpl not found" unless -f $tmpl_path;
That guard never fires when the tree is intact.