collect_rpms copied every binary rpm a builder produced. The syslinux builder
also produces syslinux, syslinux-extlinux and their debug rpms for the chroot
it runs in, so a forcearch target that builds the noarch boot loaders in the
native x86_64 chroot would have published x86_64 rpms in the riscv64 cell.
The completeness gate checks names and pins only, so it would have passed.
Keep an rpm only when it is noarch or carries the cell's architecture, read
from the rpm header. The rule lives in MockBuildUtils as rpm_in_cell.
The riscv64 cell must carry elilo-xcat, grub2-xcat, syslinux-xcat and
xnba-undi at the pins of the EL10 ppc64le cell. 3 assertions fail against
the previous manifest.
The forcearch riscv64 profile built grub2-xcat but not elilo-xcat,
syslinux-xcat or xnba-undi, and the rocky-10-riscv64-xcat cell did not list
them, so a riscv64 management node could not serve the x86 nodes of a mixed
cluster, unlike a ppc64le one.
Add the three to the profile and the cell, at the pins the EL10 ppc64le cell
uses. They are noarch and are built in the native x86_64 chroot, like
grub2-xcat. syslinux-xcat is marked noarch in the builder table, as its spec
declares, so the forcearch target does not try it in the emulated chroot,
where its ExclusiveArch excludes it. The target is now cross-built on x86_64
only, which the mock config states: a native riscv64 host could not build
syslinux-xcat either.
The manifest consistency check covered the amd64 and ppc64el sections. It
now covers the riscv64 sections too, so the 4 boot components must be listed
there. 12 assertions fail against the previous manifest.
The riscv64 sections of debs-manifest.conf listed grub2-xcat only. A riscv64
management node serves the x86 nodes of a mixed cluster, so its repository
must carry syslinux-xcat, elilo-xcat and xnba-undi, as the ppc64el sections
already require.
List the three in every riscv64 section. They are Architecture: all, built
once on amd64 and assembled into every index, so the build phase is
unchanged and the publish gate now verifies the riscv64 index carries them.
The probe starts a descendant, kills the leader with SIGKILL, and asserts the
descendant is gone when the call returns. It fails against the previous
run_bounded, where the descendant survived.
run_bounded signalled the process group on a timeout and on a cancellation, but
not when the leader itself died. A SIGKILL, or the OOM killer, takes the shell
and leaves schroot and qemu in the group, holding the chroot and writing into
staging after the call reports the build finished.
Those processes are not children of this one, so there is nothing to wait for.
Signal the group on the way out.
The assertions cover a normal exit, the highest exit code, and death by TERM,
KILL and a core-dumping signal. run_bounded is driven through the same helper,
so the decoding and its caller cannot disagree.
A child killed by a signal leaves 0 in the high byte of its wait status. The
per-codename worker loop read only that byte, so a cancelled or OOM-killed
codename was counted as built, and the run could reach validation with the
staging tree that worker never finished.
One helper decodes a wait status for both the worker loop and run_bounded, and
reports 128 plus the signal for a child the kernel killed.
ForkManager records a worker in run_on_start, which runs after the fork, so a
signal arriving in between reached a handler that did not know the worker. Both
worker pools now hold INT, TERM and HUP across start() and release them once the
parent has the pid.
The wait for a free slot happens before the mask is taken, or a cancellation
would stay pending for as long as the pool is full. With a single worker
ForkManager does not fork at all, so only a real child drops the inherited
forwarder, whose copy names siblings the parent already signals.
The probe sends itself a signal while the mask is held and asserts it arrives
only once the mask is restored, which is what makes the fork and the pid
registration a single uninterruptible step.
The forwarding handler was in place before the fork, but the parent recorded the
worker pid after it. A signal in between reached a handler that did not know the
worker, so the orchestrator died while the worker kept building for hours and
held the per-architecture lock.
The mask discipline run_bounded already used is now a pair of helpers, and the
worker loop blocks the handled signals across both the fork and the
registration. The worker resets the inherited handlers before restoring the mask.
The build wrote its debs straight into the directory a publish assembles from,
and the smoke ran afterwards. A cancellation during the smoke, which reaches the
builder as a signal, left an unverified deb there with no cleanup, and the
publish gate gets no further than names and versions.
Each package now builds into a private directory under the result directory and
its debs are moved out only after every check passes. A leftover directory from
an interrupted run is removed before the next build of that package.
The race itself is not reproducible without a hook in the production path, so
the assertions pin the invariants the fix establishes: the build starts with the
handled signals unblocked, and the caller keeps them unblocked afterwards.
The handlers were installed after fork(), so a signal arriving in between killed
the wrapper under its inherited handler and left the new process group running.
INT, TERM and HUP are now blocked around the fork and delivered once the handler
is in place. The child restores the mask before exec, or the build would inherit
a blocked TERM and ignore the signal the forwarding depends on.
The debs land in staging before the smoke runs, and the publish gate checks
names and versions only. A deb whose binary could not run therefore stayed in
staging and was eligible for publication, which is the case the smoke exists to
catch. The failure now removes the debs it rejected and says so.
The worker spawned the builder with system(), so the builder, its schroot
session and qemu were in the worker's process group but owned by nobody: a
cancelled worker died and left them running, holding the chroot and writing
into staging. Only the build inside the builder was protected, and nothing
signalled the builder.
The builder now runs through run_bounded, which gives it its own process group
and forwards the signal, so one cancellation unwinds the whole chain. The
wall-clock bound stays with the builder, which derives it from the chroot arch.
A forked worker also drops the parent's forwarder, whose copy names siblings the
parent already signals.
Two real processes stand in for a worker and the build it runs. The probe fails
when the handler forwards to nobody, which is what the orchestrators did.
Both orchestrators fork a worker per codename or per build step, so a signal
sent to the orchestrator never reached the builds: run_bounded's forwarding
covers the build inside one worker, not the workers themselves. The
orchestrator exited and released its locks while its workers kept building,
and the next run raced processes it could not see.
One handler now passes INT, TERM and HUP to the live workers, waits for them
and re-raises the signal, shared by the sbuild loop and both ForkManager pools.
With no deadline the helper called system(), which leaves the build in the
orchestrator's process group and installs no signal handler. A cancellation
then killed the orchestrator, which released its locks while the build kept
writing into staging, and a build killed by a signal was reported as rc=0.
Both paths now fork; a timeout of 0 removes the deadline, nothing else.
BUILD.md covered the EL10 riscv64 cross-build but said nothing about the apt
side, which now builds riscv64 too. Record the binfmt prerequisite, the ports
mirror, what each package does on that arch, the post-build smoke and the
measured build times.
The package declared libc6 and libssl by hand and never used
${shlibs:Depends}, so it named neither readline nor ncurses. The riscv64 deb
installed and then failed with "error while loading shared libraries:
libreadline.so.8". dpkg-shlibdeps now supplies the list from the built binary,
and the manifest pins follow the new revision.
Six of the assertions fail against the previous build_deb_in_chroot, which
accepted a smoke argument it never acted on: a wrong version, a binary that
exits non-zero, and a smoke naming a deb the build never produced all passed.
A cross-built deb links against the target's loader and libraries, neither of
which exists on the build host, so a binary that cannot run still produces a
green build. The rpm side installs the package into the mock chroot and runs -V
there; the deb side had no smoke at all.
build_deb_in_chroot now takes an optional smoke: it installs the produced deb
with apt-get inside the chroot that built it, runs the named command and matches
its output. apt-get rather than dpkg -i, so a wrong or missing Depends fails
here instead of on a node. --skip-install drops the smoke.
The two assertions fail against the previous package list and pass with it,
so a later edit cannot drop the binfmt handler that a foreign chroot needs.
--install-deps named no qemu-user-static and no binfmt-support, so on a host
without them the mode reported success and the next command still refused to
create the riscv64 chroot. The message tells the operator to install those two
packages, which the mode meant to do.
The command records its own pid and sleeps well inside the budget, so the
bound cannot be what ends it; the wrapper is then terminated and the pid
probed. Without the forwarding the build survives.
The bounded build ran in its own process group, but a signal to the
orchestrator was not forwarded to it. On cancellation the orchestrator
exited and released its locks while schroot, mock and qemu kept running and
writing into staging, so the next run raced an orphan it could not see.
INT, TERM and HUP now reap the group before the process dies by the same
signal, which keeps the exit status honest.
Drives the check extracted from sbuild-all.pl with the handler node
redirected into a temporary tree, covering the native case, a foreign
target with no handler, and the same target once one is registered.
A chroot for another architecture is bootstrapped and built through
qemu-user: debootstrap's second stage and every later build run the
target's own binaries. Without a registered handler that failed deep inside
debootstrap, on a builder that had never been prepared for the new riscv64
target. The handler is now checked before the chroot is created, and the
message names what to install. The release recipe in the manual gains the
riscv64 staging run and names the architecture where it publishes and
verifies, because an explicit --expect-arch bypasses staged discovery.
The build timeout helper was imported at compile time, and it pulls in
XCAT::BuildUtils, which needs File::Slurper. The Ubuntu build hosts do not
all carry that module, so every caller began to depend on it -- including
--install-deps, whose whole job is to install it on a host that lacks it.
The helper is now loaded where it is used.
The build guides published with --expect-arch "amd64 ppc64el". An explicit
set bypasses staged-architecture discovery, and the assembly writes the
Packages indexes and Release metadata only for that set, so following them
produced a repository with no riscv64 index even when the build staged it.
The legacy Genesis deb is named per target in the manifest, and riscv64
does not name it: its Genesis is the OpenEmbedded package published once
into the shared pool. The phase ran for every architecture regardless, so
a plain --arch riscv64 run died asking for a --genesis-deb it can never
have, and only after the dependency builds had finished. The phase is now
taken from the manifest, so it runs exactly where a Genesis deb is
expected. --skip-genesis still skips it everywhere.
Drives the scan extracted from sbuild-all.pl against a staged tree holding
every supported architecture plus an unsupported one, so the assertions
read the set publish would carry forward rather than a copy of the rule.
In publish mode with no --expect-arch, the expected set is discovered by
scanning the staged tree, and that scan admitted a hardcoded amd64 or
ppc64el. A staged riscv64 tree was therefore dropped, and only expected
architectures get a binary-<arch> index written and named in the Release
file, so the packages built for riscv64 never reached the repository. The
scan and the verification probe now use the architecture set the rest of
the script already reads from supported_arches().
Every per-package build ran through a bare system() call with no wall-clock
bound. Under qemu-user a build can deadlock -- a riscv64 goconserver `go build`
held both Go pids in futex_wait for 26 minutes with no CPU ticks and no open
socket -- and the step then never returns. The pipeline does not go red; it
stops, and a stopped run reads as "still running".
XCAT::BuildUtils::run_bounded runs the command in its own process group, kills
that group when a budget expires, and prints first what the manual
investigation had to collect by hand: the process tree, each pid's kernel wchan
and stack, its open socket count, and the CPU ticks the group used across a
20-second sample. Zero ticks names a deadlock; ticks name a build that is only
slow. BuildUtils::build_deb_in_chroot bounds every Ubuntu package build, and
mockbuild-all.pl bounds the dep and perl steps of a forcearch target.
The budget is 900 seconds for a native build and ten times that for a foreign
architecture, because qemu-user under TCG runs at roughly a tenth of native
speed. Both sit about five times above the slowest build measured on
xcat-master-ub: 3 minutes native, and 26 minutes for riscv64 ipmitool-xcat on
resolute with four codenames building at once. Native mock steps stay unbounded
-- no measurement of them exists, and a guessed budget would turn a trusted
cell red. sbuild-all.pl --build-timeout and mockbuild-all.pl --build-timeout
override the default; 0 removes the bound.
t/build_timeout.t fails without this change: run_bounded never returns and the
test reports the hang instead of blocking.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
A riscv64 goconserver `go build` sat 26 minutes with zero CPU ticks across a
20-second sample, both Go pids in futex_wait and no socket open. Nothing bounds
a build step, so the cell did not fail -- it hung, and a hung run reads as
"still running" rather than as a defect.
t/build_timeout.t drives the bounded path with a command that hangs and asserts
that the call returns, reports a timeout, and prints the process tree, each
pid's wchan, the open socket count and a CPU-tick sample. A second case drives a
spinning command and asserts the report calls it slow, not deadlocked, so the
sample means something.
Each call under test runs in a forked child whose stdio is detached to a file,
and the parent bounds that child. An unbounded run must fail this test, not
block prove.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
Every Ubuntu build fails its manifest validation:
[noble-amd64] grub2-xcat: built 2.12-2, manifest pins 2.12-1
Adding the EL10 riscv64 grub2 UEFI image bumped grub2-xcat/debian/changelog to
2.12-2 and left debs-manifest.conf pinning 2.12-1, in all twelve sections. The
package builds; only the pin is wrong.
t/sbuild-all.t compares every non-glob pin with its changelog, and fails on this
one without the change.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The Ubuntu build fails at the end, after compiling every package:
FATAL: manifest validation failed:
[noble-amd64] grub2-xcat: built 2.12-2, manifest pins 2.12-1
debs-manifest.conf pins the exact deb version each package must produce, and that
version comes from the package's own debian/changelog. Bumping the changelog
without the pin costs a whole build to discover a one-line edit.
The test compares every non-glob pin with the first line of that package's
debian/changelog. Globbed pins are deliberate -- goconserver's revision is the CD
stamp and xcat-genesis-base is not versioned by xcat-dep -- and are skipped. It
fails on grub2-xcat today.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The architecture-coverage block was appended at the end of t/sbuild-all.t, past
done_testing. Test::More had already declared the plan, so the run ended with
"planned 194 tests but ran 197" and the three assertions counted for nothing.
Move the block above done_testing.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
The 2.19 release needs riscv64 debs, and sbuild-all.pl accepted only amd64 and
ppc64el, so no target produced them.
The supported architecture set moves into BuildUtils as one source of truth
(supported_arches/is_supported_arch), which --arch, --target and --expect-arch now
consult; the mirror rule becomes "anything but amd64 is on ubuntu-ports", which is
what ppc64el already needed and riscv64 needs too; and the genesis control remap
stops naming ppc64el.
ipmitool-xcat also could not build there. Its debian/control names architectures
explicitly and omitted riscv64, so debhelper reported "No packages to build.
Possible architecture mismatch" and the build died at ./configure. conserver and
goconserver say Architecture: any and needed nothing.
debs-manifest.conf gains a [<codename>-riscv64] section for each codename. It
lists neither the x86 boot loaders -- a riscv64 node netboots UEFI grub2 -- nor
xcat-genesis-base, whose riscv64 flavour is the OpenEmbedded one in the shared
pool.
t/sbuild-all.t covers the control-file defect: it fails on ipmitool without this
change.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
ipmitool-xcat fails to build on riscv64:
dh: warning: No packages to build. Possible architecture mismatch:
riscv64, want: i386 amd64 ia64 ppc64el
make: ./configure: No such file or directory
Its debian/control names architectures explicitly, and riscv64 is not in the
list, so debhelper builds nothing and the build dies at configure. conserver and
goconserver say Architecture: any and are unaffected.
The test reads each compiled dep's debian/control and asserts that an explicit
architecture list covers every arch BuildUtils::supported_arches names. It fails
on ipmitool today.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>