After nodepurge removes a Subiquity node, /install/autoinst/<node> is still on
disk with meta-data, user-data and vendor-data in it. user-data carries the root
password hash of a node that no longer exists.
remove_node_config_files removed each path with unlink. unlink cannot remove a
directory, and mkinstall in debian.pm calls mkpath for a Subiquity node, so the
node configuration is a directory there and a plain file on the preseed and
kickstart paths. The routine now removes a directory with rmtree.
nodepurge_autoinst_cleanup.t fails without this change and passes with it.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
nodepurge removes the autoinstall configuration of each node it deletes. The
cleanup loop was inline in the nodepurge sub of profilednodes.pm, which no test
can load, so the loop moves to xCAT::ProfiledNodeUtils->remove_node_config_files
with its behaviour unchanged.
nodepurge_autoinst_cleanup.t drives that routine against a scratch directory. It
fails here: the Subiquity node keeps its directory, and the preseed file and the
.pre and .post scripts are removed.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The GitHub check installs no rpm, so riscv64_packaging.t and
genesis_spec_target_arch.t skip every check that runs rpmspec. The
workflow now installs rpm, which provides rpmspec on Ubuntu.
The GitHub check of master installs the xcat-dep packages of the latest
channel, which serves the stable release. A master change that needs a
new xcat-dep package then fails the check until that package is in a
stable release. The check now reads the devel channel, which carries
the xcat-dep packages of the next release. The 2.19 branch keeps the
latest channel.
The release information page lists every release up to 2.18.0. 2.18.2,
2.19.0 and 2.19.1 are published and absent from it, so a reader cannot
tell from the documentation which releases exist, when each one shipped,
or where its notes are.
docs/source/overview/_files/2.19.x.csv is new and holds the 2.19.0 and
2.19.1 rows, xcat2_release.rst gains its section above 2.18.x, and
2.18.x.csv gains the 2.18.2 row. Each date is the date of that release
on GitHub. 2.18.1 has no GitHub release of its own, so it has no row.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The release checklist said that the published .repo files point at
latest. latest is a link into the newest series, so a file that names
it gives the next series to users of this one as soon as that series
ships. Each file under repos/yum/X.Y/ now names repos/yum/X.Y/, and the
installation check uses the files as published.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Read the Docs takes the version in the title of each page from release
in docs/source/conf.py, which was last set for 2.17.0. Every build
since, 2.18.x and 2.19.0 included, is titled "xCAT 2.17.0
documentation". Set it to 2.20.0, the Version of master.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
xcat2/xcat-core#7864 adds openEuler 20.03 LTS SP4, 22.03 LTS SP4 and
24.03 LTS SP4 on x86_64, for the management node, service nodes and
stateful and stateless compute nodes. The support matrix does not list
openEuler at all.
Add one row per release, x86_64 only, and say which node roles the
support covers. Widen the Version column to fit the release names.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
genesis_ubuntu_build_root.t cut REQUIRED_PACKAGES out of
builddeb-genesis-base with a regex and evaluated it through bash -c. It
then matched the apt-cache fallback calls as text. Reformatting the
script broke it, and a wrong choice between renamed packages passed it.
genesis_payload_verification.t ran the verifier script and parsed its
stderr.
genesis_ubuntu_build_root.t now calls required_packages() with a chosen
set of carried packages and asserts the exact result: the amd64 extras,
the name picked for each renamed package, the order apt is asked,
tzdata-legacy, and the error for a release that carries neither name.
It reads the mandatory commands from XCAT::GenesisPayload, the code the
build uses. genesis_payload_verification.t calls the XCAT::GenesisPayload
functions with chosen payload trees and asserts the exact missing paths
and results.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
builddeb-genesis-base built its package list inline, between apt-get
update and apt-get install. A test could reach the list only by cutting
the REQUIRED_PACKAGES assignment out of the script and evaluating it.
It could not reach the choice between renamed packages (bind9-dnsutils
or dnsutils, util-linux-extra or util-linux) at all. The payload check
in verify-genesis-payload was bash, so a test could only run the script
and read its stderr.
XCAT::GenesisBuildRoot::required_packages() now returns the list for a
dpkg architecture. A code ref says which packages the release carries;
the default asks apt-cache. builddeb-genesis-base calls it at the same
point in the build and installs the same packages.
XCAT::GenesisPayload holds the mandatory-command list and the payload
check. verify-genesis-payload runs its main() and keeps the same
arguments, exit codes and messages. buildrpms.pl stages the module
beside the script. Both modules use core Perl only, because the build
root has perl-base and nothing more.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The error-command test ended with
[ ! -e /run/testnode-logs.tar ]
meant to show the archive had not gone to the default path. It cannot show that.
The file may exist for reasons that have nothing to do with this test, in which case
the assertion fails while nothing is wrong; and its absence would be equally true if
the override had never worked at all. It answers a question about the host, not
about the run.
The positive assertion above it already carries the proof: the tar shadow writes a
marker, and the test greps for that marker in the path it passed. The output being
there is what shows the redirection went there.
Removing it changes nothing about what the test catches. With the override taken out
of the template, so the archive path is hard-coded again, the remaining assertion
still fails.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Brings in #7866, which removes --database from the createrepo call. Without it the
branch cannot index a repository on the shared build tree: createrepo_c tries to
write primary.sqlite there and the NFS re-export answers
Cannot open .repodata/primary.sqlite: Can not create db_info table: disk I/O error
which failed every EL build in xcat-ci #44 while the Ubuntu builds, which do not use
createrepo, passed.
Killing the command is not killing the build. sh() ran the command through /bin/sh,
and cancellation signalled that shell alone -- but dpkg-buildpackage starts workers
of its own, and those survive their shell. The lock was then released while they
were still writing debian/changelog and debian/control, which is the state the lock
exists to prevent: the next build takes the checkout and the two rewrite it
together.
The command now runs in its own process group, so cancellation can take all of it.
Both sides call setpgid, so neither depends on which runs first, and INT and TERM
are blocked across the fork so cancellation cannot land in the window before the
group exists.
Cancellation escalates from the caught signal to KILL, and then CHECKS: a shell that
has exited is not a build that has stopped, so it waits for the whole group to
disappear rather than for the leader to be reaped. If the group is still there after
that, the locks are RETAINED and the process exits non-zero. Releasing a lock while
a worker may still be writing is worse than leaving a lock behind for a person to
clear -- the first corrupts a build, the second stops one.
cancel_build ignores INT and TERM while it runs, so a second Ctrl-C cannot interrupt
the cleanup half way and release the lock early.
sh() also reports a signalled command as 128+signal instead of 0. $? >> 8 is zero
for a child killed by a signal, so a build stopped mid-way looked to its caller like
one that had succeeded.
Two cases added to builddebs_lock_cancellation.t: a build whose worker is a
grandchild, and a command killed by a signal. Verified by signalling the pid instead
of the group, which leaves the worker running and turns the first red.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
mknb wrote BOOTIF=01-${netX/machyp} into the BIOS Genesis script.
machyp is an xNBA setting, and stock iPXE expands it to an empty string.
Without the boot MAC, the legacy Genesis stops with "Unable to find boot
device" after at least 10 minutes. The OpenEmbedded Genesis fails at once
with BOOT_INTERFACE_NOT_FOUND.
Use ${netX/mac:hexhyp}, as the UEFI Genesis script and xnba.pm do. xNBA
expands both settings to the same value, so boot with xNBA is unchanged.
The install monitor holds the later connections for a node whose handler
is still running, and forks the next one when it reaps that handler.
SIGCHLD is what brings the parent back to look: it interrupts the accept.
A handler can exit after the parent checks its queue and before the
accept begins. The signal is then handled where there is no accept to
interrupt, and the parent blocks in accept with a connection already
queued and ready to run. That connection waits until some other node
calls in. A single node retrying on its own waits until it times out.
do_installm_service now waits through wait_for_installm_connection, which
selects on the listening socket. The wait is bounded by
$installm_wakeup_seconds while connections are queued, so the parent
looks at its queue again instead of waiting for another client. An idle
monitor with an empty queue still waits without a bound, because a
handler that exits then leaves nothing to do.
The new case asserts the wait ends on its own bound with nothing to
accept, and ends at once when a connection is already there. It fails
when the bound is ignored.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The install monitor answers a node's requests in the order they arrived.
Two ways to lose that order had no test.
The first is a queued client that gives up. The parent holds the later
connections for a busy node unread, so a client that closes its socket
must not let the request behind it overtake the request that is running.
The new case runs one request, queues two more, drops the middle client,
and asserts the last request starts after the running one ends. It fails
when the per-node queue is removed.
The second is a fork that fails. The monitor then answers the node
itself, in line. The new case makes one fork fail and asserts the request
is answered and that the node's next request starts only after it. It
fails when the fallback drops the connection instead.
The file also asserts the lifted service holds no literal /var/run path.
Every access to the pid file goes through $installm_pidfile, so pointing
that variable at the scratch tree redirects all of them, and a path
written out again inside the routine would reach the host file.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
xcatd_install_monitor_concurrency.t failed about one run in five. The
assertion "the node is released only after the destiny advance finished"
compares the time the test read "done" from the socket against the time
the plugin stand-in recorded when it finished. The stand-in wrote that
time with %.3f, which rounds up, so a recorded time can be later than the
moment it was taken. The test then failed on the rounding and not on the
order:
'1790220468.11892' >= '1790220468.119'
The events file now holds whole microseconds from gettimeofday, and the
comparisons read the same clock. gettimeofday rounds nothing.
Each case also writes an events file of its own. A handler forked by one
case outlives the monitor that forked it, so it could append to the case
that runs next.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The build lock is released in DESTROY, and perl does not run DESTROY when a signal
ends the process. A build stopped with SIGTERM or SIGINT therefore left its lock
directory behind, and the next build of that checkout died on
FATAL: another build of <path> already holds <dir> (held by [pid=NNNN])
naming a pid that had already exited. Nothing clears it but a person. One such
directory blocked an openSUSE target across three consecutive CI runs before anyone
looked at what the lock actually said.
buildrpms.pl has released its lock on cancellation for some time, through an END
block and an abort handler. This is the Debian builder catching up.
The order matters, and is the reason this is not simply an END block. The command in
flight is stopped BEFORE the lock is released: handing the checkout to a second build
while dpkg-buildpackage is still rewriting debian/changelog and debian/control in it
is worse than holding the lock a moment longer. The wait for that command is bounded,
so a subprocess that ignores the signal cannot hold the lock for ever either.
sh() now forks and execs rather than calling system(), because system() gives no pid
and a handler cannot stop what it cannot name. The child _exits rather than exits, so
it never runs the parent's END block and releases a lock the parent still holds.
The handler is installed by XCAT::BuildUtils::install_build_cancellation rather than
written inline in the builder, so a test can use the same wiring the builder uses. A
test that installs an equivalent handler of its own proves the helper works while
saying nothing about whether anything calls it -- the first version of this test did
exactly that, and passed with the wiring removed.
Release is idempotent: a signal handler and then DESTROY both reach it, and the
second must not remove a directory a LATER build has since taken.
builddebs_lock_cancellation.t terminates the holder, then takes the lock again, and
checks no build subprocess was orphaned. Verified by removing the wiring: assertions
9 through 12 fail, naming the leaked lock, the refused build and the stray process.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The project had no written release process, and the 2.18 and 2.19
releases missed steps: the release table in the documentation, the
docs version, signed tags, and the wiki and website index pages.
A new Releases section describes the branch and version model, the
label and milestone that each pull request needs, and a checklist for
release candidates, publishing, the signed tag, the GitHub release, the
release notes, the website and the announcement. The build hosts, the
signing key and the publish procedure stay in the private repository of
the maintainers.
The OpenEmbedded metadata job ran on every pull request and took about
20 minutes. It reads only xCAT-genesis-builder/oe and
xCAT-genesis-scripts, and it pins its upstream sources to fixed commits.
A pull request that changes neither directory gets the same result each
time.
The job moves unchanged to its own workflow, which runs only when those
paths or the workflow file change. xcat_pr_test stays in xcat_test.yml
and runs on every pull request, because a required check that does not
run blocks the merge.
2.19.0 is released, and master still builds packages that call themselves
2.19.0. Every snapshot built from master since then carries the released
version, so a candidate cannot be told from the release it follows.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>