master added the common repository gate to the workflow and rewrote the
riscv64 comment of packages-manifest.conf after this branch started.
The workflow keeps both new prove lines. The two tests are unrelated.
The riscv64 comment takes master's text. This branch said the x86 boot
components are not built for riscv64; master then listed elilo-xcat,
syslinux-xcat and xnba-undi in the cell, because a riscv64 management
node serves the x86 nodes of a mixed cluster, so the branch sentence is
no longer true. The last line keeps this branch's statement: the perl
set is the EL10 one plus perl-HTML-Form. The pin itself merged without
a conflict, and t/riscv64_perl_cell.t passes on the merged cell.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
master added the per-cell architecture rule (rpm_arch, rpm_in_cell) after
this branch started, so the two sides collide in three places.
The MockBuildUtils and t/mockbuild-all.t import lists take both sets of
names. Neither side removes a name the other needs.
MockBuildUtils.pm now held two definitions of rpm_arch, one from each
side, and Perl kept the later one. master's definition stands: it reads
the header of a file and falls back to the name suffix otherwise, which
is what its own tests and rpm_in_cell need. The branch definition is
removed. carry_over_rpms takes rpm_arch as an argument, so it keeps its
own architecture gate and its own message.
Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>
t/mockbuild-all.t covers the manifest gate and the skip-run carry-over and
was not in the package test job. Its rpm fixtures skip where rpmbuild is
absent.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
skipped_builder must place a source package with the builder required_pkgs
would skip and never claim the OpenEmbedded Genesis. carry_over_rpms must
keep every binary rpm of a skipped build whose source package the target
manifest still names, leave out whole any build the run already carries a
member of, and
die on an unreadable header, an unsigned or foreign-architecture member of
a selected build, or a package published at two versions. Paths with
spaces are covered. Both fail to import against the previous module.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
Each run collects the rpms it built into its run repository, and deploy
stages the cell from that repository and gates it on the whole manifest,
which a skip run cannot satisfy: with --skip-perl there are no perl
packages, the gate reports them missing, and the deploy fails after the
build succeeded. The gate is right to check the whole manifest, since the
flags describe what this invocation built and not what the cell may lack.
After collection, the run repository now takes the binary rpms a skipped
builder published in the cell: whole builds, so subpackages the manifest
does not name stay too; only builds whose source package the target
manifest still names, so a dropped package is not republished; and only
builds of which the run carries no member yet, so generations of one build
never mix. The bump check, createrepo, the tarball and the deploy gate
then see the same complete set. The carry-over stops the run rather than
publish a partial or doubtful set: an rpm in the cell whose header cannot
be read, apart from the OpenEmbedded Genesis family that is pruned by
name, a member of a selected build the signing key did not sign, by signer
id and by rpmkeys --checksig, or whose digests do not verify when no key
is configured, a member of another architecture than the cell or noarch,
or a package published at more than one version. A package that was never
published is still reported missing, and a full run is unchanged. The
zero-artifact check applies only when a builder ran, so a run that skips
both package builders reaches the deploy.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
verify_rpms_checksig built the isolated keyring and ran rpmkeys on every
rpm in one body. The keyring setup and the per-rpm verdict are now their
own helpers, so another caller can verify a single rpm against the signing
key. The gate behaves as before.
Signed-off-by: Vinícius Ferrão <2031761+viniciusferrao@users.noreply.github.com>
A noarch builder of a forcearch target runs in the native chroot of the
release, but its cleanup step was registered against the target
configuration. The step scrubbed a chroot that did not exist and left the
native bootstrap behind on every run, one per noarch step.
Compute the configuration once per step and scrub the same one it built in.
collect_rpms copied every binary rpm a builder produced. The syslinux builder
also produces syslinux, syslinux-extlinux and their debug rpms for the chroot
it runs in, so a forcearch target that builds the noarch boot loaders in the
native x86_64 chroot would have published x86_64 rpms in the riscv64 cell.
The completeness gate checks names and pins only, so it would have passed.
Keep an rpm only when it is noarch or carries the cell's architecture, read
from the rpm header. The rule lives in MockBuildUtils as rpm_in_cell.
The riscv64 cell must carry elilo-xcat, grub2-xcat, syslinux-xcat and
xnba-undi at the pins of the EL10 ppc64le cell. 3 assertions fail against
the previous manifest.
The forcearch riscv64 profile built grub2-xcat but not elilo-xcat,
syslinux-xcat or xnba-undi, and the rocky-10-riscv64-xcat cell did not list
them, so a riscv64 management node could not serve the x86 nodes of a mixed
cluster, unlike a ppc64le one.
Add the three to the profile and the cell, at the pins the EL10 ppc64le cell
uses. They are noarch and are built in the native x86_64 chroot, like
grub2-xcat. syslinux-xcat is marked noarch in the builder table, as its spec
declares, so the forcearch target does not try it in the emulated chroot,
where its ExclusiveArch excludes it. The target is now cross-built on x86_64
only, which the mock config states: a native riscv64 host could not build
syslinux-xcat either.
The manifest consistency check covered the amd64 and ppc64el sections. It
now covers the riscv64 sections too, so the 4 boot components must be listed
there. 12 assertions fail against the previous manifest.
The riscv64 sections of debs-manifest.conf listed grub2-xcat only. A riscv64
management node serves the x86 nodes of a mixed cluster, so its repository
must carry syslinux-xcat, elilo-xcat and xnba-undi, as the ppc64el sections
already require.
List the three in every riscv64 section. They are Architecture: all, built
once on amd64 and assembled into every index, so the build phase is
unchanged and the publish gate now verifies the riscv64 index carries them.
The assertion fails against the previous manifest, naming perl-HTML-Form as
the package the cell lacked. It reads the set from the builder, so a package
added to the builder later fails here until the cell names it.
perl-xCAT hard-requires perl(HTML::Form). EPEL supplies it on x86_64 and
ppc64le, so the EL9 and EL10 cells omit it, and the riscv64 cell copied that
omission. There is no EPEL for riscv64 and Rocky 10 riscv64 carries no
perl-HTML-Form in BaseOS, AppStream, CRB or extras, so dnf install xCAT could
not resolve on a riscv64 management node.
--list-packages prints the set a run selects, after --epel-gap and --packages,
and exits before the root and mock checks. The riscv64 manifest cell has no
other repository behind it, so a test can now hold the cell to this list
without a build host.
The probe starts a descendant, kills the leader with SIGKILL, and asserts the
descendant is gone when the call returns. It fails against the previous
run_bounded, where the descendant survived.
run_bounded signalled the process group on a timeout and on a cancellation, but
not when the leader itself died. A SIGKILL, or the OOM killer, takes the shell
and leaves schroot and qemu in the group, holding the chroot and writing into
staging after the call reports the build finished.
Those processes are not children of this one, so there is nothing to wait for.
Signal the group on the way out.
The assertions cover a normal exit, the highest exit code, and death by TERM,
KILL and a core-dumping signal. run_bounded is driven through the same helper,
so the decoding and its caller cannot disagree.
A child killed by a signal leaves 0 in the high byte of its wait status. The
per-codename worker loop read only that byte, so a cancelled or OOM-killed
codename was counted as built, and the run could reach validation with the
staging tree that worker never finished.
One helper decodes a wait status for both the worker loop and run_bounded, and
reports 128 plus the signal for a child the kernel killed.
ForkManager records a worker in run_on_start, which runs after the fork, so a
signal arriving in between reached a handler that did not know the worker. Both
worker pools now hold INT, TERM and HUP across start() and release them once the
parent has the pid.
The wait for a free slot happens before the mask is taken, or a cancellation
would stay pending for as long as the pool is full. With a single worker
ForkManager does not fork at all, so only a real child drops the inherited
forwarder, whose copy names siblings the parent already signals.
The probe sends itself a signal while the mask is held and asserts it arrives
only once the mask is restored, which is what makes the fork and the pid
registration a single uninterruptible step.
The forwarding handler was in place before the fork, but the parent recorded the
worker pid after it. A signal in between reached a handler that did not know the
worker, so the orchestrator died while the worker kept building for hours and
held the per-architecture lock.
The mask discipline run_bounded already used is now a pair of helpers, and the
worker loop blocks the handled signals across both the fork and the
registration. The worker resets the inherited handlers before restoring the mask.
The build wrote its debs straight into the directory a publish assembles from,
and the smoke ran afterwards. A cancellation during the smoke, which reaches the
builder as a signal, left an unverified deb there with no cleanup, and the
publish gate gets no further than names and versions.
Each package now builds into a private directory under the result directory and
its debs are moved out only after every check passes. A leftover directory from
an interrupted run is removed before the next build of that package.
The race itself is not reproducible without a hook in the production path, so
the assertions pin the invariants the fix establishes: the build starts with the
handled signals unblocked, and the caller keeps them unblocked afterwards.
The handlers were installed after fork(), so a signal arriving in between killed
the wrapper under its inherited handler and left the new process group running.
INT, TERM and HUP are now blocked around the fork and delivered once the handler
is in place. The child restores the mask before exec, or the build would inherit
a blocked TERM and ignore the signal the forwarding depends on.
The debs land in staging before the smoke runs, and the publish gate checks
names and versions only. A deb whose binary could not run therefore stayed in
staging and was eligible for publication, which is the case the smoke exists to
catch. The failure now removes the debs it rejected and says so.
The worker spawned the builder with system(), so the builder, its schroot
session and qemu were in the worker's process group but owned by nobody: a
cancelled worker died and left them running, holding the chroot and writing
into staging. Only the build inside the builder was protected, and nothing
signalled the builder.
The builder now runs through run_bounded, which gives it its own process group
and forwards the signal, so one cancellation unwinds the whole chain. The
wall-clock bound stays with the builder, which derives it from the chroot arch.
A forked worker also drops the parent's forwarder, whose copy names siblings the
parent already signals.
Two real processes stand in for a worker and the build it runs. The probe fails
when the handler forwards to nobody, which is what the orchestrators did.
Both orchestrators fork a worker per codename or per build step, so a signal
sent to the orchestrator never reached the builds: run_bounded's forwarding
covers the build inside one worker, not the workers themselves. The
orchestrator exited and released its locks while its workers kept building,
and the next run raced processes it could not see.
One handler now passes INT, TERM and HUP to the live workers, waits for them
and re-raises the signal, shared by the sbuild loop and both ForkManager pools.