The pacing that keeps the install monitor alive is written inline in xcatd, and xcatd
cannot be run in a unit test: it needs the database, SSL, the plugin tree and
/var/run/xcat before it will start at all. So the test reached for the only thing left and
matched regular expressions against the script's source -- that a respawn branch exists,
that it mentions an interval, that it names a cap. Every one of those assertions passes
against pacing that is subtly wrong, and none of them would notice the retry budget
running out and never being refilled, which is the actual defect under review. Grepping
the implementation also pins its shape, so the code cannot be rearranged without editing
the test that is supposed to be guarding it.
State the pacing instead as an interface a test can execute: xCAT::RespawnUtils, pure
functions that take a state and a time and return the next state, with no clock, no
globals and no I/O of their own. Passing the time in is what lets the schedule be checked
over a virtual clock rather than in real seconds.
Drive it for the delay backing off to a ceiling and holding there, for the never-give-up
property (three hours into a continuous failure the daemon is still forking monitors), for
the reset (a monitor that stayed up long enough to serve clears the backoff when it later
dies), for the guards on a policy that could not back off, and for purity itself. Then
drive it for real against a genuinely held TCP port: fail several times, release the port,
and require that a respawned monitor binds it and stays up without the daemon being
restarted.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The respawn added for the install monitor is paced by a retry budget, and once that
budget is spent the daemon stops trying. That reintroduces the failure the respawn
exists to remove: with no monitor alive there is nothing left to reset the counter, so
xcatiport stays dead until the whole daemon is restarted, and a port that becomes free
a minute later is never picked back up. Pacing the retries is necessary -- an unguarded
re-fork spins as fast as fork allows while the port is held, and keeps re-entering
do_installm_service's USR2 socket-takeover handshake -- but pacing must not decay into
giving up.
The property that matters is therefore behavioural, not structural: the monitor comes
back on its own, at a bounded rate, no matter how long it has been failing. Assert it by
driving xcatd's real pacing code rather than grepping for it -- extract the marked
mon-respawn-policy region from the script verbatim, the way build_ubunturepo_lock.t
drives build-ubunturepo's real lock, and run it. Over a virtual clock, check that the
delay backs off to a ceiling and holds there, that the daemon is still forking monitors
three hours into a failure, and that a monitor which stayed up long enough to serve
resets the pacing when it later dies. Then do it for real against a genuinely held TCP
port: fail several times, release the port, and require that a respawned monitor binds
it and stays up without the daemon being restarted.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Re-fork the install monitor from the main service loop when $pid_MON has been cleared
and xcatiport is still configured, so a single death of that child no longer leaves
the port dead until the whole daemon is restarted. The forked child closes the SSL
listener and the UDP control socket before re-entering do_installm_service, so it
serves only the install monitor.
Rate limit the respawn. do_installm_service dies when it cannot bind the port after
its own retries, which is exactly the case where an unguarded re-fork would spin as
fast as fork allows and keep re-entering that function's USR2 socket-takeover
handshake against whatever still holds the socket. Consecutive attempts are separated
by XCATD_MON_RESPAWN_INTERVAL seconds (default 5) and capped at XCATD_MON_RESPAWN_MAX
(default 10), after which xcatd logs that it is giving up on the port rather than
retrying forever. A monitor that stayed up long enough to outlast the whole retry
budget resets the counter, so an unrelated death much later gets a full budget again.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The install monitor -- the child listening on xcatiport for node install-status
updates and the "next" boot-flip request -- is forked exactly once at daemon startup.
When it dies the SIGCHLD reaper only clears $pid_MON and nothing re-forks it, so a
single death of that child (a stray signal, or a lost socket takeover during an xcatd
restart) leaves xcatiport permanently dead while the main daemon keeps running.
Installing nodes can then no longer report booted or request the boot flip until the
whole daemon is restarted, which is disruptive to any concurrent operation.
A respawn must also be rate limited. do_installm_service dies when it cannot bind the
port after its own retries, so an unguarded re-fork in the main loop would spin as
fast as fork allows for as long as the port stays held, and would collide with that
same function's USR2 socket-takeover handshake.
Assert that the main loop re-forks the monitor, that the child re-enters
do_installm_service, and that respawns are spaced, capped, and reported on exhaustion.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Add a unit test for the option that selects fping instead of nmap.
The test takes the specification out of the plugin source and gives it to
Getopt::Long with the settings that the daemon uses, so it drives the
specification that the plugin ships.
It shows that -f, --usefping and the older --useping each select fping, that
--use and --us still select usemon and do not select fping, that the bundles
-mf and -fm select both options, and that both places parse through the one
specification.
Add a unit test for the routine that chooses the compression program. The
test lifts the routine out of the plugin source, because the plugin needs a
management node to load.
The test shows that lzma is used when it is there, that xz stands in when it
is not, and that xz is asked for the lzma container rather than its own. It
also shows that the caller takes the command from the routine, that the file
keeps its name and its suffix, and that the gzip fallback and the rename into
place both remain.
The install disk autotests read the log of a provisioned node. The
choice files no longer carry the identifier in their name, and the
selection message names the driver group and the identifier instead of
the previous wording, so read the new lines. The reinstall case reads
the record of its disk without naming a group, as it did before.
Cover a RAID volume that reports a WWN against a direct attached disk
that reports none, in both scan orders, which the previous readback
decided by identifier. Keep the identifier rules of one group under
test as well: the disk that reports a WWN wins, the lower WWN wins
between two, and a path wins over no identifier at all.
Cover a guest whose only disk is a Xen disk, which the scan has to
select rather than leave to the fallback, and a guest with two Xen
disks, where the driver group decides. Against the previous filter both
cases fail.
Assert that the RHEL 10 copy is gone, that the RHEL 10 installer
includes the common script, and that the common script keeps the VROC
fallback, the Xen fallback and the guarded failure log.
The RHEL 10 copy of the script is about to go away, so stop naming it
here first. The cases keep running against the common script, so the
coverage does not change.
The install disk autotest reads the log of a node whose disks sit
behind a SAS host adapter, and that driver group moved from the second
choice to the third. Read the third group instead.
Run the real scripts in a sandbox. A stub udevadm serves the device
properties from fixture files, and the partition list and the output
paths move into the sandbox. Every case runs against the common script
and against the copy the RHEL 10 installer includes.
Cover the direct attached disk against a RAID volume, a RAID only
server, the host adapter against a direct attached disk and against an
unknown driver, an NVMe device from the last group, and the default
fallback. Against the previous scripts the RAID cases fail, so they
discriminate.