2
0
mirror of https://github.com/xcat2/xcat-core.git synced 2026-09-05 20:47:55 +00:00
Commit Graph

2572 Commits

Author SHA1 Message Date
Daniel Hilst 6a9983abef test(xcat-core): the respawn test asserts xcatd's shape, not its behaviour
The pacing that keeps the install monitor alive is written inline in xcatd, and xcatd
cannot be run in a unit test: it needs the database, SSL, the plugin tree and
/var/run/xcat before it will start at all. So the test reached for the only thing left and
matched regular expressions against the script's source -- that a respawn branch exists,
that it mentions an interval, that it names a cap. Every one of those assertions passes
against pacing that is subtly wrong, and none of them would notice the retry budget
running out and never being refilled, which is the actual defect under review. Grepping
the implementation also pins its shape, so the code cannot be rearranged without editing
the test that is supposed to be guarding it.

State the pacing instead as an interface a test can execute: xCAT::RespawnUtils, pure
functions that take a state and a time and return the next state, with no clock, no
globals and no I/O of their own. Passing the time in is what lets the schedule be checked
over a virtual clock rather than in real seconds.

Drive it for the delay backing off to a ceiling and holding there, for the never-give-up
property (three hours into a continuous failure the daemon is still forking monitors), for
the reset (a monitor that stayed up long enough to serve clears the backoff when it later
dies), for the guards on a policy that could not back off, and for purity itself. Then
drive it for real against a genuinely held TCP port: fail several times, release the port,
and require that a respawned monitor binds it and stays up without the daemon being
restarted.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-08-26 17:01:05 -03:00
Daniel Hilst 5970c4efd1 test(xcat-core): capture the install monitor giving up on xcatiport for good
The respawn added for the install monitor is paced by a retry budget, and once that
budget is spent the daemon stops trying. That reintroduces the failure the respawn
exists to remove: with no monitor alive there is nothing left to reset the counter, so
xcatiport stays dead until the whole daemon is restarted, and a port that becomes free
a minute later is never picked back up. Pacing the retries is necessary -- an unguarded
re-fork spins as fast as fork allows while the port is held, and keeps re-entering
do_installm_service's USR2 socket-takeover handshake -- but pacing must not decay into
giving up.

The property that matters is therefore behavioural, not structural: the monitor comes
back on its own, at a bounded rate, no matter how long it has been failing. Assert it by
driving xcatd's real pacing code rather than grepping for it -- extract the marked
mon-respawn-policy region from the script verbatim, the way build_ubunturepo_lock.t
drives build-ubunturepo's real lock, and run it. Over a virtual clock, check that the
delay backs off to a ceiling and holds there, that the daemon is still forking monitors
three hours into a failure, and that a monitor which stayed up long enough to serve
resets the pacing when it later dies. Then do it for real against a genuinely held TCP
port: fail several times, release the port, and require that a respawned monitor binds
it and stays up without the daemon being restarted.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-08-25 18:05:59 -03:00
Daniel Hilst 195ab4243b fix(xcat-core): respawn the xcatd install monitor when it dies
Re-fork the install monitor from the main service loop when $pid_MON has been cleared
and xcatiport is still configured, so a single death of that child no longer leaves
the port dead until the whole daemon is restarted. The forked child closes the SSL
listener and the UDP control socket before re-entering do_installm_service, so it
serves only the install monitor.

Rate limit the respawn. do_installm_service dies when it cannot bind the port after
its own retries, which is exactly the case where an unguarded re-fork would spin as
fast as fork allows and keep re-entering that function's USR2 socket-takeover
handshake against whatever still holds the socket. Consecutive attempts are separated
by XCATD_MON_RESPAWN_INTERVAL seconds (default 5) and capped at XCATD_MON_RESPAWN_MAX
(default 10), after which xcatd logs that it is giving up on the port rather than
retrying forever. A monitor that stayed up long enough to outlast the whole retry
budget resets the counter, so an unrelated death much later gets a full budget again.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-08-24 14:44:49 -03:00
Daniel Hilst d73cd41415 test(xcat-core): capture xcatd never respawning its install monitor
The install monitor -- the child listening on xcatiport for node install-status
updates and the "next" boot-flip request -- is forked exactly once at daemon startup.
When it dies the SIGCHLD reaper only clears $pid_MON and nothing re-forks it, so a
single death of that child (a stray signal, or a lost socket takeover during an xcatd
restart) leaves xcatiport permanently dead while the main daemon keeps running.
Installing nodes can then no longer report booted or request the boot flip until the
whole daemon is restarted, which is disruptive to any concurrent operation.

A respawn must also be rate limited. do_installm_service dies when it cannot bind the
port after its own retries, so an unguarded re-fork in the main loop would spin as
fast as fork allows for as long as the port stays held, and would collide with that
same function's USR2 socket-takeover handshake.

Assert that the main loop re-forks the monitor, that the child re-enters
do_installm_service, and that respawns are spaced, capped, and reported on exhaustion.

Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
2026-08-24 14:43:46 -03:00
Daniel Hilst f2f96b67fc Merge pull request #7728 from VersatusHPC/fix/xml-external-entity
fix(xcatd): block XML external entities on the legacy parser path
2026-08-24 14:23:48 -03:00
Daniel Hilst 5ca148889c Merge pull request #7749 from VersatusHPC/fix/nodestat-usefping-option
fix(nodestat): accept the fping option that the usage message gives
2026-08-24 12:42:21 -03:00
Daniel Hilst bcf6f9059a Merge pull request #7750 from VersatusHPC/fix/dbobjutils-exact-only-if-values
fix(dbobjutils): match exact only-if values
2026-08-24 12:39:34 -03:00
Daniel Hilst 14feebce2f Merge pull request #7753 from VersatusHPC/fix/genesis-lzma-via-xz
fix(mknb): compress the genesis image with xz when lzma is absent
2026-08-24 12:36:26 -03:00
Daniel Hilst ca5d1cfa86 Merge pull request #7751 from VersatusHPC/refactor/dbobjutils-remove-legacy-group-matcher
refactor(dbobjutils): remove redundant group matcher
2026-08-24 12:29:02 -03:00
Daniel Hilst 8c3aaa4471 Merge pull request #7727 from VersatusHPC/refactor/kea-shared-service-mapping
refactor(kea): reuse shared service mapping
2026-08-24 12:27:38 -03:00
Daniel Hilst a7f4c770b5 Merge pull request #7631 from VersatusHPC/refactor/ipmi-rmcp-response-identity
refactor(ipmi): centralize RMCP response identity check
2026-08-24 12:27:17 -03:00
Vinícius Ferrão eaa1e94a32 test(nodestat): pin the fping option and the usemon abbreviations
Add a unit test for the option that selects fping instead of nmap.

The test takes the specification out of the plugin source and gives it to
Getopt::Long with the settings that the daemon uses, so it drives the
specification that the plugin ships.

It shows that -f, --usefping and the older --useping each select fping, that
--use and --us still select usemon and do not select fping, that the bundles
-mf and -fm select both options, and that both places parse through the one
specification.
2026-08-23 22:38:41 -03:00
Vinícius Ferrão 73b145d7ff test(mknb): pin which program compresses the genesis image
Add a unit test for the routine that chooses the compression program. The
test lifts the routine out of the plugin source, because the plugin needs a
management node to load.

The test shows that lzma is used when it is there, that xz stands in when it
is not, and that xz is asked for the lzma container rather than its own. It
also shows that the caller takes the command from the routine, that the file
keeps its name and its suffix, and that the gzip fallback and the rename into
place both remain.
2026-08-23 22:38:41 -03:00
Vinícius Ferrão d3d0f86abf test(network): cover shared netmask behavior 2026-08-23 13:51:39 -03:00
Vinícius Ferrão 7afa152f8d test(dhcp): prepare shared network helper stubs 2026-08-23 13:49:15 -03:00
Vinícius Ferrão 1e8917ebe0 test(dbobjutils): guard only-if routing invariants 2026-08-23 13:13:16 -03:00
Vinícius Ferrão 21755f8f93 test(dbobjutils): cover exact only-if value matching 2026-08-23 11:38:55 -03:00
Vinícius Ferrão 5d39fc30d4 test(utils): cover comma-list membership 2026-08-23 11:08:59 -03:00
Daniel Hilst 7733d16f8d Merge pull request #7721 from VersatusHPC/feature/genesis-openembedded
feat(genesis): build images with OpenEmbedded
2026-08-21 20:36:20 -03:00
Vinícius Ferrão 9da2d71b90 test(genesis): cover legacy destiny retries 2026-08-21 20:11:31 -03:00
Vinícius Ferrão c6c11d70a1 test(genesis): cover the action boot gate 2026-08-21 02:13:19 -03:00
Vinícius Ferrão 4c50b4da50 test(genesis): cover disappearing devices 2026-08-21 02:12:56 -03:00
Vinícius Ferrão 53c8538e4a test(genesis): cover logging failures 2026-08-21 02:12:22 -03:00
Vinícius Ferrão a47c7f8e7c test(genesis): cover long kernel command lines 2026-08-21 02:10:23 -03:00
Vinícius Ferrão 9fc3d52b81 test(getinstdisk): follow the log lines of the new selection
The install disk autotests read the log of a provisioned node. The
choice files no longer carry the identifier in their name, and the
selection message names the driver group and the identifier instead of
the previous wording, so read the new lines. The reinstall case reads
the record of its disk without naming a group, as it did before.
2026-08-21 01:13:50 -03:00
Vinícius Ferrão 492f9171a2 test(getinstdisk): cover the driver group against the identifier
Cover a RAID volume that reports a WWN against a direct attached disk
that reports none, in both scan orders, which the previous readback
decided by identifier. Keep the identifier rules of one group under
test as well: the disk that reports a WWN wins, the lower WWN wins
between two, and a path wins over no identifier at all.
2026-08-21 01:13:16 -03:00
Vinícius Ferrão bbd1a7e9d3 test(getinstdisk): cover the Xen disk scan
Cover a guest whose only disk is a Xen disk, which the scan has to
select rather than leave to the fallback, and a guest with two Xen
disks, where the driver group decides. Against the previous filter both
cases fail.
2026-08-21 01:13:15 -03:00
Vinícius Ferrão 7e8994e1b9 test(getinstdisk): pin the single script layout
Assert that the RHEL 10 copy is gone, that the RHEL 10 installer
includes the common script, and that the common script keeps the VROC
fallback, the Xen fallback and the guarded failure log.
2026-08-21 01:13:15 -03:00
Vinícius Ferrão 046c97ba2a test(getinstdisk): target the common script only
The RHEL 10 copy of the script is about to go away, so stop naming it
here first. The cases keep running against the common script, so the
coverage does not change.
2026-08-21 01:13:15 -03:00
Vinícius Ferrão 549034d758 test(getinstdisk): follow the SAS host adapters to the third group
The install disk autotest reads the log of a node whose disks sit
behind a SAS host adapter, and that driver group moved from the second
choice to the third. Read the third group instead.
2026-08-21 01:13:14 -03:00
Vinícius Ferrão 59219181c9 test(genesis): use Yocto extension filenames 2026-08-20 23:23:19 -03:00
Vinícius Ferrão baa68945cc test(genesis): use Yocto release filenames 2026-08-20 23:21:39 -03:00
Vinícius Ferrão cba24cc4c0 test(genesis): reject unsafe network state 2026-08-20 23:07:31 -03:00
Vinícius Ferrão cef35ba5e1 test(genesis): cover signed extension bundles 2026-08-20 23:03:53 -03:00
Vinícius Ferrão 56040cb4f0 test(genesis): cover release compliance output 2026-08-20 23:00:15 -03:00
Vinícius Ferrão 321595fb62 test(getinstdisk): cover the install disk selection order
Run the real scripts in a sandbox. A stub udevadm serves the device
properties from fixture files, and the partition list and the output
paths move into the sandbox. Every case runs against the common script
and against the copy the RHEL 10 installer includes.

Cover the direct attached disk against a RAID volume, a RAID only
server, the host adapter against a direct attached disk and against an
unknown driver, an NVMe device from the last group, and the default
fallback. Against the previous scripts the RAID cases fail, so they
discriminate.
2026-08-20 22:59:11 -03:00
Vinícius Ferrão a903a12463 test(genesis): cover sequential BMC setup 2026-08-20 22:56:07 -03:00
Vinícius Ferrão 43ca1c6363 test(genesis): cover fatal action service policy 2026-08-20 22:52:44 -03:00
Vinícius Ferrão c0af2535f8 test(genesis): cover network refresh rollback 2026-08-20 22:48:13 -03:00
Vinícius Ferrão 63f2326b6e test(genesis): cover BMCs without SOL 2026-08-20 22:43:13 -03:00
Vinícius Ferrão 0da0828913 test(genesis): cover narrow console headers 2026-08-20 22:41:53 -03:00
Vinícius Ferrão 8442dba2bf test(genesis): cover destiny client termination 2026-08-20 22:40:37 -03:00
Vinícius Ferrão 6562a4e9d7 test(genesis): cover export manifest 2026-08-20 22:29:57 -03:00
Vinícius Ferrão 1f083299f6 test(genesis): cover image contract 2026-08-20 22:29:57 -03:00
Vinícius Ferrão cc5c2192ce test(genesis): cover hardware providers 2026-08-20 22:29:56 -03:00
Vinícius Ferrão f6acf56ef2 test(genesis): cover signed extensions 2026-08-20 22:29:56 -03:00
Vinícius Ferrão 71cee46751 test(genesis): cover status console 2026-08-20 22:29:56 -03:00
Vinícius Ferrão e1c5d1565c test(genesis): cover provisioning runtime 2026-08-20 22:29:56 -03:00
Vinícius Ferrão bae92ffc47 fix(credentials): audit delegated certificate signing 2026-08-20 17:37:19 -03:00
Vinícius Ferrão 1f68e97f9e test(credentials): cover service node certificate delegation 2026-08-20 16:50:31 -03:00