master added the common repository gate to the workflow and rewrote the
riscv64 comment of packages-manifest.conf after this branch started.
The workflow keeps both new prove lines. The two tests are unrelated.
The riscv64 comment takes master's text. This branch said the x86 boot
components are not built for riscv64; master then listed elilo-xcat,
syslinux-xcat and xnba-undi in the cell, because a riscv64 management
node serves the x86 nodes of a mixed cluster, so the branch sentence is
no longer true. The last line keeps this branch's statement: the perl
set is the EL10 one plus perl-HTML-Form. The pin itself merged without
a conflict, and t/riscv64_perl_cell.t passes on the merged cell.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
A noarch builder of a forcearch target runs in the native chroot of the
release, but its cleanup step was registered against the target
configuration. The step scrubbed a chroot that did not exist and left the
native bootstrap behind on every run, one per noarch step.
Compute the configuration once per step and scrub the same one it built in.
collect_rpms copied every binary rpm a builder produced. The syslinux builder
also produces syslinux, syslinux-extlinux and their debug rpms for the chroot
it runs in, so a forcearch target that builds the noarch boot loaders in the
native x86_64 chroot would have published x86_64 rpms in the riscv64 cell.
The completeness gate checks names and pins only, so it would have passed.
Keep an rpm only when it is noarch or carries the cell's architecture, read
from the rpm header. The rule lives in MockBuildUtils as rpm_in_cell.
The riscv64 cell must carry elilo-xcat, grub2-xcat, syslinux-xcat and
xnba-undi at the pins of the EL10 ppc64le cell. 3 assertions fail against
the previous manifest.
The forcearch riscv64 profile built grub2-xcat but not elilo-xcat,
syslinux-xcat or xnba-undi, and the rocky-10-riscv64-xcat cell did not list
them, so a riscv64 management node could not serve the x86 nodes of a mixed
cluster, unlike a ppc64le one.
Add the three to the profile and the cell, at the pins the EL10 ppc64le cell
uses. They are noarch and are built in the native x86_64 chroot, like
grub2-xcat. syslinux-xcat is marked noarch in the builder table, as its spec
declares, so the forcearch target does not try it in the emulated chroot,
where its ExclusiveArch excludes it. The target is now cross-built on x86_64
only, which the mock config states: a native riscv64 host could not build
syslinux-xcat either.
The manifest consistency check covered the amd64 and ppc64el sections. It
now covers the riscv64 sections too, so the 4 boot components must be listed
there. 12 assertions fail against the previous manifest.
The riscv64 sections of debs-manifest.conf listed grub2-xcat only. A riscv64
management node serves the x86 nodes of a mixed cluster, so its repository
must carry syslinux-xcat, elilo-xcat and xnba-undi, as the ppc64el sections
already require.
List the three in every riscv64 section. They are Architecture: all, built
once on amd64 and assembled into every index, so the build phase is
unchanged and the publish gate now verifies the riscv64 index carries them.
The assertion fails against the previous manifest, naming perl-HTML-Form as
the package the cell lacked. It reads the set from the builder, so a package
added to the builder later fails here until the cell names it.
perl-xCAT hard-requires perl(HTML::Form). EPEL supplies it on x86_64 and
ppc64le, so the EL9 and EL10 cells omit it, and the riscv64 cell copied that
omission. There is no EPEL for riscv64 and Rocky 10 riscv64 carries no
perl-HTML-Form in BaseOS, AppStream, CRB or extras, so dnf install xCAT could
not resolve on a riscv64 management node.
--list-packages prints the set a run selects, after --epel-gap and --packages,
and exits before the root and mock checks. The riscv64 manifest cell has no
other repository behind it, so a test can now hold the cell to this list
without a build host.
The probe starts a descendant, kills the leader with SIGKILL, and asserts the
descendant is gone when the call returns. It fails against the previous
run_bounded, where the descendant survived.
run_bounded signalled the process group on a timeout and on a cancellation, but
not when the leader itself died. A SIGKILL, or the OOM killer, takes the shell
and leaves schroot and qemu in the group, holding the chroot and writing into
staging after the call reports the build finished.
Those processes are not children of this one, so there is nothing to wait for.
Signal the group on the way out.
The assertions cover a normal exit, the highest exit code, and death by TERM,
KILL and a core-dumping signal. run_bounded is driven through the same helper,
so the decoding and its caller cannot disagree.
A child killed by a signal leaves 0 in the high byte of its wait status. The
per-codename worker loop read only that byte, so a cancelled or OOM-killed
codename was counted as built, and the run could reach validation with the
staging tree that worker never finished.
One helper decodes a wait status for both the worker loop and run_bounded, and
reports 128 plus the signal for a child the kernel killed.
ForkManager records a worker in run_on_start, which runs after the fork, so a
signal arriving in between reached a handler that did not know the worker. Both
worker pools now hold INT, TERM and HUP across start() and release them once the
parent has the pid.
The wait for a free slot happens before the mask is taken, or a cancellation
would stay pending for as long as the pool is full. With a single worker
ForkManager does not fork at all, so only a real child drops the inherited
forwarder, whose copy names siblings the parent already signals.
The probe sends itself a signal while the mask is held and asserts it arrives
only once the mask is restored, which is what makes the fork and the pid
registration a single uninterruptible step.
The forwarding handler was in place before the fork, but the parent recorded the
worker pid after it. A signal in between reached a handler that did not know the
worker, so the orchestrator died while the worker kept building for hours and
held the per-architecture lock.
The mask discipline run_bounded already used is now a pair of helpers, and the
worker loop blocks the handled signals across both the fork and the
registration. The worker resets the inherited handlers before restoring the mask.
The build wrote its debs straight into the directory a publish assembles from,
and the smoke ran afterwards. A cancellation during the smoke, which reaches the
builder as a signal, left an unverified deb there with no cleanup, and the
publish gate gets no further than names and versions.
Each package now builds into a private directory under the result directory and
its debs are moved out only after every check passes. A leftover directory from
an interrupted run is removed before the next build of that package.
The race itself is not reproducible without a hook in the production path, so
the assertions pin the invariants the fix establishes: the build starts with the
handled signals unblocked, and the caller keeps them unblocked afterwards.
The handlers were installed after fork(), so a signal arriving in between killed
the wrapper under its inherited handler and left the new process group running.
INT, TERM and HUP are now blocked around the fork and delivered once the handler
is in place. The child restores the mask before exec, or the build would inherit
a blocked TERM and ignore the signal the forwarding depends on.
The debs land in staging before the smoke runs, and the publish gate checks
names and versions only. A deb whose binary could not run therefore stayed in
staging and was eligible for publication, which is the case the smoke exists to
catch. The failure now removes the debs it rejected and says so.
The worker spawned the builder with system(), so the builder, its schroot
session and qemu were in the worker's process group but owned by nobody: a
cancelled worker died and left them running, holding the chroot and writing
into staging. Only the build inside the builder was protected, and nothing
signalled the builder.
The builder now runs through run_bounded, which gives it its own process group
and forwards the signal, so one cancellation unwinds the whole chain. The
wall-clock bound stays with the builder, which derives it from the chroot arch.
A forked worker also drops the parent's forwarder, whose copy names siblings the
parent already signals.
Two real processes stand in for a worker and the build it runs. The probe fails
when the handler forwards to nobody, which is what the orchestrators did.
Both orchestrators fork a worker per codename or per build step, so a signal
sent to the orchestrator never reached the builds: run_bounded's forwarding
covers the build inside one worker, not the workers themselves. The
orchestrator exited and released its locks while its workers kept building,
and the next run raced processes it could not see.
One handler now passes INT, TERM and HUP to the live workers, waits for them
and re-raises the signal, shared by the sbuild loop and both ForkManager pools.
With no deadline the helper called system(), which leaves the build in the
orchestrator's process group and installs no signal handler. A cancellation
then killed the orchestrator, which released its locks while the build kept
writing into staging, and a build killed by a signal was reported as rc=0.
Both paths now fork; a timeout of 0 removes the deadline, nothing else.
BUILD.md covered the EL10 riscv64 cross-build but said nothing about the apt
side, which now builds riscv64 too. Record the binfmt prerequisite, the ports
mirror, what each package does on that arch, the post-build smoke and the
measured build times.
The package declared libc6 and libssl by hand and never used
${shlibs:Depends}, so it named neither readline nor ncurses. The riscv64 deb
installed and then failed with "error while loading shared libraries:
libreadline.so.8". dpkg-shlibdeps now supplies the list from the built binary,
and the manifest pins follow the new revision.
Six of the assertions fail against the previous build_deb_in_chroot, which
accepted a smoke argument it never acted on: a wrong version, a binary that
exits non-zero, and a smoke naming a deb the build never produced all passed.
A cross-built deb links against the target's loader and libraries, neither of
which exists on the build host, so a binary that cannot run still produces a
green build. The rpm side installs the package into the mock chroot and runs -V
there; the deb side had no smoke at all.
build_deb_in_chroot now takes an optional smoke: it installs the produced deb
with apt-get inside the chroot that built it, runs the named command and matches
its output. apt-get rather than dpkg -i, so a wrong or missing Depends fails
here instead of on a node. --skip-install drops the smoke.
The two assertions fail against the previous package list and pass with it,
so a later edit cannot drop the binfmt handler that a foreign chroot needs.