The password change and the account update set If-Match: * on the
connection itself, and after a forced password change that connection
is handed back, so every later request carried it. Pass the header
with the PATCH alone. Also drop an unused asyncio import.
get_health looked only at each node's State, so an Enabled node whose
Health was Warning or Critical left the chassis reported as ok. It also
skipped a node whose resource could not be read, counting it as
healthy while an unreadable chassis was flagged. Report the node's
Health when it is not OK, and a node that cannot be read as
Unreachable.
The sensor collection has 668 members, and reading each one on its own
takes about 92 seconds and rebuilds on every nodesensors call. Opt in
to $expand=. for that collection, checking once whether the firmware
really inlines the members, so firmware that ignores $expand keeps the
per-member reads.
ForceRestart only restarts the host, so the node BMC kept running and
a reseat did not recover a hung one. The EUREKA firmware provides a
Reseat reset type that removes all power from the slot.
A node whose BMC is not reporting returns Reading 0 from its
temperature sensors, dragging averages down with meaningless
values. The ComputerSystem Oem data flags this via HasBMCMetrics;
skip such nodes, and keep collecting when the flag is absent.
get_health collected an issues list but returned it nowhere, and
flattened every problem to Warning. Return the findings as
badreadings using SensorReading, honor verbose, and map chassis
health through _healthmap so Critical is no longer downgraded.
The sensor URL was built as BMC{N}CpuCPU{X}Temp instead of
BMC{N}CPU{X}Temp, so every request returned 404. The entries also
referenced const.SensorUnits, which does not exist, and used a
dict shape the only consumer, get_average_processor_temperature,
cannot read - it expects thermal-style dicts with ReadingCelsius.
util.json_loads does not exist, so the PasswordChangeRequired flow
raised AttributeError on every 401, silently swallowed by the blanket
except. Use json.loads, which accepts the bytes body directly.
ManagedBy is optional, and when a system does not link a manager
get_bmcurl() answers None. Most callers passed that straight into a
web request, which failed with "Constructor parameter should be str"
from the url library. Route those callers through bmcinfo(), which now
raises UnsupportedFunctionality so confluent reports it plainly. The
event log falls through to its system and chassis fallback, and
list_media lists nothing, since neither needs a manager.
SensorReading computed health and states with .get() and then
overwrote both with a strict lookup, so a sensor whose Status lacks
Health, or carries a value outside the health map, raised KeyError
and aborted the whole sensor listing. Keep the tolerant lookup, and
drop the states of a reading that is OK, as the copy in command.py
already does.
Have redfishbmc be able to attempt a redfish onboarding.
User must supply initial user and password, since we have no idea about the vendor choices at this level.
This implements a secureboot compatible flow, even for PXE.
Non secureboot environments suffer one useless transfer, but otherwise should be unaffected.
It was possible for the name resolution to steal an address from another section.
Fix by having explicit IP addressing consume and then
purge any violaters after concurrent evaluation completes.
Specify connect and login timeouts
to avoid sessions being held open.
Also, in blocking_scan, wrap everything so that finally can ensure the scan is recognized as complete.
For auto discovery, banner extraction, and interactive ssh, refactor to commen sshclient class.
This hooks the pubkeys.ssh in a manner compatible with confluent 3.x
Going from generic-redfish as most specific, then generic-https, and generic-ssh being for ssh-only targets.
For generic-https and generic-ssh, the available ports are specified so code can know if https *also* has port 22 available.
Some experimentation shows that 15 fps is more than enough for the vast majority of console activity, so cut back for reduced file size.
Also, the calculation for last frame was incorrect, tracking duration based on when exiting caught up to the queue. Now add the end time explicitly and use that to reduce last frame lingering.
VP9 is still uncomfortably slow in a default setup, so stick with MP4V despite larger size, user may transcode if they want to make it smaller.
Refactor functions to let netutil depend on neighutil.
Add a suite of functions to take a mac and try to figure out some viable ip for the mac.
Provide a ping6 mainly to support a ping to ff02::1, and follow up with unicast UDP discard packets to trigger neighbor table population.
Try to use the neighbor table to figure out an ip and scope for a mac, preferring LLA.
If this fails, go for a try of converting the mac to lla arithmetically, then try all the nics to see which one seems to work.
This reverts commit be6c7a3794.
They do not exist in Leap/SLE 16 installer initrd.
The diskless hook keeps its copies. imgutil's installkernel instmods
all four, so they are in that image and the calls do work there.
Building a SUSE 16 image from SLE media failed every package with
"key ID fec28eaf09d9ea69: NOKEY". Leap publishes that key as
gpg-pubkey-*.asc, which the existing glob picks up; SLE publishes the
same key only as repodata/repomd.xml.key, so nothing was imported.
15 media carries both spellings, so this changes nothing there.
udev's kmod builtin dlopens libkmod, and dracut installs it from an
inst_libdir_file line in a module-setup.sh rather than by following
NEEDED. Which module carries that line moved: the dracut on SLE 16
media declares it only in 00systemd, which the diskless module set
never loads, so the image came up with no libkmod, udev autoloaded
nothing, and the guest reached the network scan with only loopback.
Leap's newer dracut also declares it in 95udev-rules, which base
depends on, which is why Leap was unaffected.
OpenSSH 10 splits each connection into sshd-session and that into
sshd-auth, so the initramfs sshd on 2222 could not serve a single
session. el10 added sshd-session for the same reason.
SLES 16 ships no ISC dhclient and nothing provides dhcp-client, so the
image could not be built from SLES media at all. Leap carries dhcpcd
too, so one client covers both. el10 made the same move when RHEL
dropped dhclient.
The deploycfg carries the literal 'null', which went into the profile as
a password hash, so the installed root account reported a usable
password instead of a locked one. 15 substitutes '!' for this; the sed
delimiter has to move off '!' to carry it.
The default Etc/UTC is not in the tzdata list agama validates against,
so the load failed and the install fell through to agama defaults and
still reported completion. Normalize it and halt if the load fails.
The attached-media branch could never be taken: the label pattern built
from os-release is opensuse-leap-16.0 or sles-16.0, while the media is
labelled Install-Leap-16.0-x86_64 and Install-SUSE-SLE-16-x86_64. Had it
matched, it would have written an inst.repo to the dracut cmdline that
agama does not read, and skipped inst.script entirely.
They stayed commented out when the hook was forked from el8, so an
IPoIB-only node had no path to the deploy server. The diskless hook
loads them already.
prechroot.sh had setupssh.sh commented out and copied only the keys
inline, so nodes came up without shosts.equiv, without the CA in
ssh_known_hosts and without a setuid ssh-keysign.
A trailing comma made args.cmd a tuple holding the argv list, so
fancy_chroot called startswith on a list and the child died before exec.
permissions.local was written but never applied, leaving ssh-keysign
0755 and hostbased auth inoperative on SUSE diskless images.
Newer systemd prints '(unset)' where it used to print 'n/a', so an
unset console keymap was handed to the installer verbatim. An unset
System Locale has no '=' either, and the whole line was being taken
as the locale.
- the aarch64 osdeploy spec builds the stateful suse16 addons but its
diskless loop was never extended, so the aarch64 rpm shipped suse16
without suse16-diskless and a packed image got a dangling addons.cpio.
- imgutil's builddeb keeps its own copy of the directory list that
confluent_imgutil.spec.tmpl has, and it had learned about neither suse16
nor el10.
- gather_bootloader gained a /usr/share/efi fallback for shim on both
architectures but only for x86_64 on grub, so an aarch64 root found a
shim and then died copying grub.
Finally, rewriting repos.d file by file rather than copying the tree meant
a subdirectory or a file that is not valid UTF-8 aborted the build before
any package was installed, which also regressed SUSE 15. Pass anything
that is not a plain text repo definition through untouched and restore the
modes on the ones that are rewritten.
SuseHandler refused anything but 15.x. What 16 needed beyond widening it:
- its repo urls are written in terms of ${releasever}, which zypper
resolves from the target root's os-release, a file that does not exist
yet when the first packages go in
- its repos name a zypper service backed by a package-provided directory
the target root does not have, so zypper discarded every one of them
as an orphan
- there is no mkinitrd to work out which kernel to build for, and bare
dracut would build for the build host's running kernel
- the efi payloads moved out of /usr/lib64/efi, arping out of /usr/sbin,
nsswitch.conf and protocols under /usr/etc, and the presets enable
sshd already
- the module list predated virtio, so an image built for a KVM guest had
no network at all, and dm-crypt could not allocate a transform for the
encrypted image without the aes-xts modules
- urlmount still links libpthread, an empty stub since glibc 2.34 that
nothing else in the initramfs pulls in
Ported from suse15-diskless, with the differences 16 forces:
- dracut symlinks /lib/dracut/hooks to /var/lib/dracut/hooks, so a hook
shipped at the old path replaces the symlink with a directory
- there is no netconfig or /etc/sysconfig/network to hand the running
address to, so the initramfs writes a NetworkManager keyfile instead.
Without it NetworkManager claims the interface on its own terms and the
tethered root filesystem goes away with the old address, and confignet
never gets the chance to refine anything
- the discovery loop retries without a delay, so a link that takes a
moment to come up can exhaust all 30 tries before the first packet can
go anywhere. Keep asking, as the el9 hook already does
Without it osdeploy import cannot generate a profile at all:
generate_stock_profiles opens profile.yaml unguarded, and initprofile.sh
seds the label into it. The label substitution also still looked for
'sle 15'.
Linux-PAM reads vendor defaults from /usr/lib/pam.d and distributions are
migrating there package by package: systemd and polkit already ship into
it on both EL and Debian, and on SUSE 16 openssh has followed. There the
old code left a dangling /etc/pam.d/confluent and every pam authentication
against it failed.
The deb postinst carries the same logic, so fix it in step. ln -sf rather
than ln -s because -e is false for a dangling link, so the old code retried
the symlink and failed with 'File exists' instead of repairing it.
init_sdr assigned self._sdr before initialize() ran, so a failure left the
half built object in the cache. The next call saw a non-None _sdr and handed
back that partial repository rather than trying again.
The visible symptom is a first call raising and the second appearing to
succeed. The real cost is on a bmc where the read fails once: the client
keeps the incomplete sdr for the life of the session and every later sensor
lookup answers from it without complaint.
get_webclient falls off its end when /api/login answers anything but 200, so
it returned None. wc() passes that back, and thirty of the thirty-four call
sites use it unchecked, so a refused login arrived as "'NoneType' object has
no attribute 'grab_json_response'" from wherever it landed.
Raised where the failure is known, and with the status: 404 is firmware with
no web api, or a Redfish-only capture of one, while 401 is credentials it
will not take. Neither was distinguishable before.
The other four call sites are the inventory reads, and they always did check.
They answer partially when the web interface is out of reach, which is why
nodeinventory still says something useful. They ask through wc_if_available
now, so that tolerance is stated rather than resting on a None.
Two places where the code that exists to explain a failure fails instead, and
the caller is shown the second failure rather than the first.
_do_web_request builds its message from the response body when that body is
not the JSON error document the spec asks for. An html 404 page is exactly
that, the body is bytes, and str + bytes raised TypeError. The status and the
body never reached anyone.
LenovoFirmwareConfig raised a bare Exception when python-lxml and
python-eficompressor are absent. Confluent has no handler for one, so a
missing dependency showed as "Unexpected Error" and hid a message that
already said what to install.
Neither changes what fails, only what the caller is told.
Same function, one line above. EthernetInterfaces is optional, and when it is
absent the None went into a request and raised TypeError from the url library
rather than saying what was missing.
Kept beside the ambiguous case because the two are one question asked twice.
Dedicated plus shared is how BMCs are built, so several NICs is ordinary.
_get_bmc_nic_url only reaches the count when the address the session came in
on matches none of them, which is what a tunnel or a NAT does, and it then
raised the bare PyghmiException base class.
Confluent has no handler for the base class, so it fell through to the
generic one and showed "Unexpected Error" plus a traceback, for a machine
doing nothing wrong. UnsupportedFunctionality now, which the redfish plugin
already reports plainly and which stays inside PyghmiException.
The message said "does not have exactly one interface" without saying how
many, which ones, or what to do. Every caller takes a name, so it lists the
candidates, and the empty case reads differently from the ambiguous one.
A Redfish service may publish no Systems collection, and power and cooling
equipment does exactly that. sysurl is then None, and get_power and
get_bootdev handed it straight to a request, raising TypeError from inside
the url library.
sysinfo already guarded the same field. Both now ask through _system_url and
get a refusal naming what is missing. DMTF publish three services of this
shape, which is why they sit commented out in inventory-dmtf.yaml.
PCIeDevices and PCIeFunctions are optional, and get_health indexed both
without checking. _get_adp_urls in the same file already spells it
.get('PCIeDevices', []), so these two were the outliers.
Power and cooling equipment has no PCIe at all. The KeyError escaped the
health read and reached the user as "Unexpected Error" with a traceback
behind it. Reproduces offline against DMTF's public-rackmount1 mockup.
nodesensors and nodeconfig printed an error and exited 0, so
`nodesensors n1 && next-step` ran next-step after the read it guarded had
already failed.
nodesensors had three of these: the per-node branch never set the exit code,
the top-level branch beside it read `exitcode |= exitcode`, and a normal
return from main() fell off the end of the file. nodeconfig accumulates with
|= all through its read path except the last line, which assigned, so a
failed bmc configuration read was discarded by the system read after it.
The fixed branches now match how nodehealth and client.py spell the same
thing.
parse_time read the digits after the decimal point as whole milliseconds, so
'.5' became 5ms instead of 500ms. Only a three digit fraction came out right,
and a BMC may write either.
'.5', '.50' and '.500' are all 500ms now, '.125' is 125ms. Every other format
parse_time accepts is untouched.
nodeidentify against an IPMI BMC printed the node name, nothing after it, and
exited 0. A script checking the exit code carries on with an empty value,
which is worse than being turned down.
IPMI can set the identify light and has no command to read it back, so aiohmi
has no get_identify. The empty state was a way of not saying so.
The comment above that branch called identify "read-only", which is the
opposite of the truth.
crypt left the standard library in 3.13 and both imports of it here are
unconditional, so the server does not start on a 3.13 distro that ships no
crypt shim of its own.
legacycrypt and crypt_r both reach the same libcrypt call, tried in that
order because legacycrypt is pure ctypes while crypt_r wants a compiler.
Both were checked byte for byte against the stdlib for the $6$ salts used
here, and a hash written under the stdlib verifies under either, so stored
crypted.* attributes keep working.
Recommends rather than Requires, and only above 3.12: el9 and el10 still
ship the stdlib module, and some 3.13 distros package no candidate at all,
where a hard dependency would make the rpm uninstallable.
State of offloader was never checked after acquiring the lock.
Fix by checking with the lock held.
Also, neaten up by putting all the offload startup inside the function to start it up.
Try to propogate form factor of member disk to array.
Also, even if cannot detect m.2, assume a two-member vroc array of nvme is m.2. Not guaranteed, but most likely. This is to deal with lack of DMI information indicating the physical form factor.
Directories are left as 'boring' confluent directories to enable staging.
Then the ownership/permssions on directories are fixed up.
Then after completion, make sure ownership is back to boring before asking rmtree.
By default, ansible prefers to try host based authentication, which is good.
But when it doesn't work, it tries every key attempt, which is normally fine.
However, SSH counts key attempts the same as passwords, so hardening that restirct password attempts are fouled before it can even get to try a public key. Thus let host based only consume one attempt.
By copying, we leave the /etc/hosts with original ownership/permissions/etc.
Otherwise we create a new /etc/hosts, which is subject to new permissions and such.
Confluent didn't act 'healthy' toward systemd delaying restart needlessly.
Worse, in a collective it could never show as started if the quorum didn't come back.
Pull the startup to before collective init, and indicate watchdog liveness during that time.
SATA drives do not directly have a busaddr.
However, at some point the PCI bus comes up in the udev hierarchy as a KERNELS value.
If that matches a detected M.2 slot, then accept the storage as M.2.
Every error branch set exitcode and the script then ran off the end, so
a node that could not be read, or a log source that matched nothing,
still exited 0.
Every entry already says which log it came from, so -s narrows the output to
the ones asked for and "-s list" names what a node offers. Some of what a
platform keeps is noise: an AMI MegaRAC's event log is a list of redfish
sessions being opened and closed, while its useful records are elsewhere.
A selection cannot be cleared, since the platforms offer no such thing and
clearing more than was asked for is not something to do quietly.
A read only ever looked at the manager's log services, and fell back to the
system's when the manager published none. A platform that keeps an event log
in both places had the second one invisible: an AMI MegaRAC keeps power unit
and thermal events in a chassis log that nothing read, 103 records that no
command could reach.
Clearing deliberately does not follow. It stays where it was, so a log that
only a read reaches is never destroyed by one, and clearing a platform that
keeps its only event log on the system still works.
The name test now ignores spacing, since a build that calls its post code log
"BIOS POST Code Log" was read as an event log and merged 2719 post codes in.
Reading an event log skipped every log service whose id or name said journal,
dump, post code, host logger or crash, on every implementation. Those names
are bmcweb's: the AMI and Lenovo bmcs call theirs SEL, EventLog, AuditLog and
PlatformLog. bmcweb does need the distinction, keeping its event log under the
system while a clear would destroy its dumps, so it gets a handler that names
the words and generic names none, reading whatever a platform publishes.
Generic added an UpdateParameters part to every multipart firmware push,
because the specification has one carried. Only the AMI firmware was seen to
insist on it, so name it in that handler and let generic send what the caller
passed.
A bmc behind a forward answers Activate Payload with the port it listens on
itself, and the advertised-port check refused that, so a console failed where
command traffic worked. The advertised port is never sent to, so the check now
applies only on the default port. A bmc on another port advertising a third one
is no longer refused outright, which nothing here could have served anyway.
The plugin connected to 623 whatever the address said, with a TODO in place of
the parsing. It now reads one as the redfish plugin does, minus the brackets,
which getaddrinfo rejects.
IpmiConsole keys its endpoint mapping on host and port so several bmcs behind
one address stay distinct, and unregisters only an entry it actually claimed.
_unixdomainhandler hardcoded /var/run/confluent/api.sock in four places and
derived its directory from a fifth. It is now threaded through SockApi like
the other bind settings, defaulting to the same path. That lets a service run
on a temp socket as an ordinary user, which is what a test needs.
The deadline was set when monitoring began, so a flash that kept
answering for longer than the grace period had already spent it by the
time the reboot it covers arrived, and reported a working update as a
failure.
A device holds its sensor records per lun and scopes the reservation the
same way, so a token taken on lun 0 can be refused for lun 1 with 0xc5.
The retry then took the same wrong token again without end.
Firmware for a supply or a fan matched no fragment and so was called
core, and nothing could ever answer for misc.
Only the collections that mean one thing are matched by url. A Storage
resource is the subsystem, so its firmware is the controller rather than
a drive, and a Processor is a cpu as readily as an accelerator, so that
one is decided by asking the processor what it is.
A member of a collective that cannot reach quorum stalls in the startup loop
until quorum returns, and there was nothing in that loop that noticed a
shutdown. That was survivable while SIGTERM raised SystemExit out of the
signal handler, since that escaped the loop from wherever it happened to be.
Having the event loop deliver the signal instead leaves the stop event set with
nobody reading it until quorum is reached, so stopping the service waits out
the systemd timeout and ends in a kill.
Check the event in the loop condition, and wait on it rather than sleeping
through it, so the answer comes in milliseconds rather than whenever quorum
returns. A service stopped at this point has served nothing yet, so it goes
straight to the same configuration flush the normal exit does.
pyghmi asks for this parameter with xraw_command and catches the completion
code for a bmc that does not have it, and folding aiohmi in renamed that call
to oldraw_command rather than raw_command, so the handler could no longer fire.
Answering the code out of the returned dictionary repaired the crash but kept
the call on the older contract, which is now the only one left in the tree.
Catch it again instead: raw_command puts the completion code on the exception
as ipmicode, and nothing here reads the payload of a reply that carries a code,
which is the one thing catching gives up.
No behaviour change, checked against the previous version over the same fake
session for a good reply, an empty one, 0x80 and 0xC9 with and without a stray
payload, four other completion codes, a timeout, a lost session and a reply
with no data at all: same return value, same exception type, text and ipmicode,
same bytes on the wire.
_read_device_sdr_lun negotiates the read size down when the bmc answers 0xCA,
but the size > 5 guard leaves a size of 5 alone, so a bmc that will not serve
5 bytes at once was asked the same question for ever. Give up once the
request cannot get any smaller, and once a header read would go under the 5
bytes the record length sits in, by falling through to the raise already
there.
The stale reservation retry could not end on its own either: it cleared the
id and left taking a new one to the top of the loop, which only reserves for
a partial read, so the very first request repeated unchanged. Take one where
the code is handled.
Completion codes 201 and 202 mean the chunk asked for was too big, and the
check for them sat after a call that raises, so a bmc that cannot serve 224
bytes at once failed the fru read rather than being asked for less.
The retry could not terminate either: chunksize // 2 + 2 is 4 for a chunksize
of 4, so the chunksize == 3 guard was unreachable and a bmc that kept refusing
would have been asked for 4 bytes for ever.
Completion code 203 on a sensor reading means the sensor is not present, which
is expected of an optional device, but the check for it sat after a call that
raises first, so one absent sensor ended the whole sensor sweep.
_supports_standard_ipv6 read rsp['code'] the same way, so it could only ever
answer True; a platform without the standard parameters raised instead. A
completion code is that platform's answer, while a timeout or a lost session is
not and must not be cached as one.
raw_command's docstring still described itself as the other call it was renamed
from, which is how these checks came to be written against the wrong contract.
raw_command raises on any nonzero completion code, so the 0xCA and 0xC5 checks
in get_sdr could not run: a bmc that will not return a whole record in one go,
or whose reservation went stale, failed the sensor load outright instead of
being retried. Read the code from the exception, which carries it.
The back off also had a fixed point, size // 2 + 2 being 3 for a size of 3, so
a bmc that kept refusing would have been asked the same question for ever.
Give up when the request cannot get any smaller. A header read cannot go
under 5 bytes either, since that is where the record length sits, so give up
there rather than parse a reply too short to index.
The stale reservation retry could not end on its own either: it cleared the id
and left taking a new one to the top of the loop, which only reserves for a
partial read. Take one where the code is handled.
The labels are worked out across the whole inventory, since a platform may give
every entry the same Name, so an entry that came back empty would be asked for a
name it does not have and take the naming of the others with it.
A parameter file that is not json, or that holds something other than an
object, reached the update as a raw parser message or as a TypeError from the
handler that unpacked it.
A writable directory only says the user could have created a file there, and
/tmp lets anyone do that. If the target exists, it has to be writable by the
requesting user too, or confluent would overwrite it as root.
get_user_name documents that it answers None when reading a slot fails, but
raw_command raises before the check that would return it, so that branch has
never run and one refused slot aborted the whole user list. Answer None where
the docstring says to, and let get_users drop the blanket except it grew to
work around it.
Only a completion code counts as the bmc answering about the slot. A timeout
carries the fabricated 0xffff from the session layer and a lost session carries
no code at all, and swallowing either of those reports a list truncated at the
point of failure as a complete one.
Some will not take an image without being told which kind it is, and the
only way to find out was to attempt an update and read the error, which
writes to the bmc before it gets that far.
The sdr carries unit fields for every sensor record, and they were read into
the reading whether or not the sensor has a number for them to describe. A
discrete sensor reports which of its states are asserted and no value at all,
so a watchdog came back with units of "% ", and a discrete sensor on a full
record picks up a base unit the same way, reporting degrees celsius for a
sensor that never has a temperature. A caller that shows the unit alongside
whatever value it was handed then prints a unit with nothing to apply it to,
which is what the client was taught to skip in 8c8c32cc. The client is not
the only consumer of a reading, so answer the question in the library.
The record itself says which sensors those are, and the reading path already
works it out to decide whether to decode a number: a sensor has one only if
the numeric format says how to read it, or the format is unsigned and the
record either supports thresholds or is of reading type 1. Ask that once,
where the units are assembled, and give the sensors that fail it no unit, so
the unit and the value cannot come to different conclusions about whether the
sensor has a reading at all.
A threshold or numeric sensor reports exactly what it did before, including
the ones whose unit is a percentage, a combination of two units, or nothing
because the record names no unit. On the Lenovo XCC this was checked against
nothing changes: its 305 readable sensors decode identically, discrete ones
included, because they are all compact records naming no unit in the first
place. The sensors that change are the ones the finding came from, which
name a unit on a record that has no number to put it on.
A path a caller wants confluent to save something into was being run through
the check meant for a file confluent is asked to read. That check forks, drops
to the calling user and asks os.access for R_OK, which is false for every file
that does not exist yet, so nodesupport servicedata and save_licenses could
only be given a path that was already there. Handed a name to create, they
refused, and refused in a way no caller was looking for, so the command printed
nothing and exited zero.
The previous commit worked around it by skipping the check for a download
target, which fixed the symptom by removing the guard rather than by asking the
right question. Ask the right question instead: whether the user could have
created the file in that directory themselves. A path that is already a
directory is a destination directory, anything else names the file, which is
the same rule the code that goes on to write the file follows.
So a caller can still only make confluent write where they could have written,
and this now also catches an unwritable destination at the point the request is
made rather than several layers further in.
This bmc gives all three of its firmware entries the same Name, "Software
Inventory", and puts what they actually are in the description. The first
entry took that name, and the two after it fell back to their ids, so
nodefirmware answered with "Software Inventory", "cpld_active" and
"d1dc9e4b" for what are the host, cpld and bmc images.
Decide the labels across the collection rather than one entry at a time,
so a name the platform repeats can be recognised as no name at all. Where
that happens, use a description that does tell them apart, and the id when
even that is shared. A platform whose names are already distinct keeps
exactly the names it had.
The labels are what a caller addresses an entry by, so this also turns
inventory/firmware/all/d1dc9e4b into inventory/firmware/all/bmc_image.
Processor inventory carried a single field, the model, so a platform that
does not give one had a processor in the listing with nothing in it, and
the client, which skips empty values, showed no processor at all. This
bmc names the manufacturer, the socket and the core and thread counts, and
gives no model.
Carry those, along with the speed, serial and part number where a platform
offers them, and treat a processor as missing only when the bmc says its
state is absent, rather than whenever it does not describe a state.
The top bit of the major revision byte of Get Device ID says the device is
still initialising or taking a firmware update. It was read as part of
the revision, so immediately after a bmc reset nodefirmware reported "BMC
Version: 131.11" for what is 3.11, and settled down only once the bit
cleared.
Mask it as sdr.py already does for the same byte, so the two agree about
the same field.
Reading the identify state says plainly when a platform describes no
indicator, but writing it went ahead and patched IndicatorLED regardless.
This bmc has neither that property nor the boolean that replaced it, and
answered the write with an internal service error, which reached the user
as one and the log as a traceback.
Ask the same question the read asks. With neither property present there
is nothing to write, so say so in the same words instead of finding out
from the bmc.
This bmc advertises a Bios resource on its system and answers 404 for it.
Confluent followed the link and passed the bmc's complaint on as an
unexpected error, so a nodeconfig read printed every bmc setting and then
ended with "The requested resource of type named 'Bios' was not found",
and the system half of the configuration was a 500 saying the same.
There is already a good answer for a system that offers no bios settings,
and a link that is advertised and not served is the same thing as far as a
caller is concerned, so give it the same one. The result is checked once
and remembered, including the negative, so this costs one request on the
first ask and nothing after.
Redfish identifies an account by a string, and an implementation is free
to use the account name, which this one does. The handler converted the
last element of the path to an integer, so every per user read, update and
delete answered "invalid literal for int() with base 10: 'root'" as an
unexpected error, with a traceback to match. Confluent offered the id
itself, listing the account as "root", and then could not accept it back.
Take the element as given. Everything below already compares ids as
strings, and the input parsing already keeps a non numeric uid, so only
this conversion stood in the way. nodebmcpassword goes through exactly
this path, reading users/all for the id and then writing to that account,
so it could not work at all on such a bmc.
ipmi users really are numbered slots, so the conversion is right there and
stays, but say so when it fails rather than letting a ValueError surface
as an unexpected error.
A caller asking for fans or energy got nothing from any bmc that serves
the Sensors collection. Those sensors were filed under their redfish
reading type, Rotational for a fan, while the categories are named after
the ipmi sensor types the rest of the code uses, so nothing matched.
Power appeared to work only by coincidence, Power and Current happening to
be spelled the same in both vocabularies.
Translate the reading type as the sensor is mapped, so a sensor means the
same thing whether it came from the Sensors collection, from the older
Thermal and Power documents, or from ipmi. On the bmc this was found on,
fans go from nothing to the 24 tachometers, and temperature and power
already agreed with what the same hardware reports over ipmi.
The fan controls stay out, and cannot be brought in. Their reading type
is Percent, which is also what a battery state of health reports, and this
bmc fills in no PhysicalContext to tell them apart, so there is nothing to
classify them by that would not also drag in unrelated percentages.
The loop decoded the first three bytes of a name and never consumed them,
so any sensor or fru name a bmc encodes as 6 bit packed ascii spins at
full speed, appending the same four characters until the process runs out
of memory. Measured at about 10 MB a second, so a bmc using an encoding
the spec gives its own worked example of costs a pinned core and, before
long, the daemon.
Consume each group, and decode a trailing group of one or two bytes rather
than dropping it, since those carry a character each and the name would
otherwise come back short.
The arithmetic was already right, it was only never reached a second time.
Verified by encoding names per the packing and reading them back: exact for
every length except those leaving three characters in a three byte group,
where the byte count cannot say whether three or four were meant and a
trailing space is unavoidable.
A bmc may keep its sensor data records on the sensor device instead of in
a repository, and this one does, so it had no sensors, no health and only
a partial inventory over ipmi.
The records themselves are identical, version 0x51 and the same types, so
everything that decodes them is reused as is. Only the fetching differs:
a command of its own, a reservation of its own, and records held per lun
rather than in one place. The luns to ask, and a change indicator to
cache on, come from Get Device SDR Info.
The fetch is written out rather than shared with the repository one. The
loops are alike, but nothing available here has a repository to test
against, and the price of factoring them together is that a mistake would
land on every bmc that works today rather than only on those that do not
work at all.
Records are cached in memory and on disk exactly as repository records
are, keyed on the change indicator, and a device that offers no such
indicator is read afresh each time rather than cached wrongly.
Names are stripped of the nulls that pad a fixed width field. This bmc
pads every name out to sixteen bytes, and a name carrying them cannot be
matched by a caller asking for a sensor by name.
On the bmc this was written for: 163 sensors and 9 frus, against the 163
the device reports it has. The 54 temperatures and 36 fans it now reads
match what the same machine reports over redfish, to within the precision
each side gives.
The oem lookup answers whether it found a handler for the vendor, and that
answer was being stored as whether the lookup had been done at all. On
anything the map does not name, which is every bmc that is not a Lenovo,
the flag stayed false and each oem_init issued another Get Device ID and
built another handler.
Almost everything goes through oem_init, so this is a round trip added to
almost every operation. Where those calls are close together it is far
worse than that: reading the sensor data records asks for the event
constants once per record, so a run of 172 records fired 176 Get Device ID
commands back to back, which was enough to make the bmc stop answering and
the read fail with a timeout. The same sequence now takes 3 commands.
Settling for the generic handler is an answer. The device id cannot
change within a session, so asking again buys nothing, and the handler it
throws away each time is the one holding the sensor names it had cached.
A bmc that keeps its sensor data records on the sensor device rather than
in a repository was answered with a bare NotImplementedError. With no
text of its own it reached the user as "Unexpected Error:
NotImplementedError" and was logged with a traceback, for a capability the
library had simply never implemented rather than anything having gone
wrong.
Say so instead, in all three places that give up: the two branches for a
bmc without an sdr repository, and the version check that only understands
records of version 0x51, which now names the version it was given.
This makes the inventory usable on such a bmc as a side effect. The
system fru is gathered before the records are, and an unsupported
operation is tolerated where an unexpected error was not, so nodeinventory
answers with the board, chassis and product data instead of one line of
error.
The redfish event log was taken from every log service the manager
advertises, whatever those turned out to be. On a bmc that keeps its
systemd journal there, nodeeventlog answered with a thousand lines of
kernel probe failures and daemon chatter, and the log the user asked for
was never read at all, because this implementation keeps it under the
system. Clearing was worse: of the services it did find, the ones with a
clear action were the dumps, so a clear destroyed diagnostic data, left
the event log untouched, and reported success.
Judge a log service before reading it. A service whose id or name says
journal, dump, post code, host logger or crash is not an event log, and
both reading and clearing skip it, so a clear can no longer take out
something that was never asked for.
If that leaves the manager with no event log at all, look under the
system, where such an implementation keeps it. Only then: a bmc that has
one under the manager is served exactly as before, from the same requests,
so this cannot change what an implementation that already worked reports.
The list of services was also being extended in place, and it belongs to
whatever the url cache is holding, so an extra log added by an oem handler
accumulated on every call within the cache window.
Reading the alert destinations of a bmc that has none reported "Unknown
code 0x80 encountered", which is the fallback text for a completion code
the library has no name for. 0x80 on this parameter is not a failure, it
is the platform saying it does not have alert destinations, and the
redfish side of the same resource has said so in words for a while.
The lan parameter fetch already knew how to tell those apart, so build the
alert reads on it rather than on a raw command that raises on any non-zero
code, and raise UnsupportedFunctionality with something to read. Both the
count and an individual destination are covered, so a platform that offers
one and not the other says the same thing instead of failing differently.
Splitting the completion code handling out of the parameter fetch is what
makes that reuse possible; the interpretation of the payload, and every
answer it gives, is unchanged.
The oem hook for the destination count was passing its byte through ord(),
which raises TypeError on the bytearray it is given. No handler in tree
implements the hook, so it had never been called; hand it the integer.
A bmc that does not implement a lan configuration parameter says so in the
completion code and answers with no data at all. The helper reached
straight into the payload, so such a parameter raised IndexError, and with
it went the whole of nodeconfig over ipmi: the bmc group, the plain, the
detailed, the extended and the advanced reads all ended in "bytearray
index out of range". It also took the attribute enumeration with it, so
the client then rejected names it should have accepted.
The guard that was there caught an exception carrying the completion code,
but oldraw_command reports the code in its response rather than raising on
it, so nothing was ever caught. Read the code from the response instead:
parameter not supported and parameter out of range mean the platform does
not have it, and anything else is a real failure that should say what it
was rather than be mistaken for absence.
The address configuration method was looked up in a table with no regard
for whether it had been read at all, so a bmc that does not report it
would have traded the IndexError for a KeyError. Answer None when it is
absent, as the address above it already does, and name the value when it
is present but unfamiliar.
A console whose bmc had gone away reported "Unexpected error - None", and
the api answered 504 with an error of None. The redfish plugin took the
text for an unreachable target from the strerror of the socket error it
caught, guarded by a hasattr that is always true: every OSError has a
strerror attribute, and it is None on most of the ones a bmc going away
produces, TimeoutError and gaierror among them. Ask for the text the same
way as everywhere else instead, which also keeps the errno on the errors
that do carry one.
The same applies to an unreachable target raised with no message at all,
so use the same helper there, on both transports.
Underneath that, give the node error messages a default to fall back on
rather than carrying whatever they were handed. Each subclass already had
one, in an __init__ that an explicit None went straight past; making it a
class attribute the base class applies means it holds however the message
was built, and removes five copies of the same constructor.
Also repair an affluent handler that put a closing parenthesis in the
wrong place, passing its error text to Queue.put_nowait as a second
argument. Any OSError there other than "no route to host" raised
TypeError from inside the except clause instead of reporting anything.
The receive loop treated only WSMsgType.CLOSE as the end of a session, but
aiohttp reports a peer that has gone away as CLOSED, and it does so
immediately and for every subsequent call. Everything that was not CLOSE
fell to an else branch that printed a line and went round again, so a
console whose bmc restarted became a full speed loop writing one line per
iteration: measured at 2.7 million iterations a second, and observed
filling 15 GB of log in a quarter of an hour while the daemon stopped
answering requests.
Treat every message that is not data as the end of the session, clear the
connected flag and report the disconnect once. A session that ended any
other way than a clean close is recorded in the trace log, unbuffered so
that it survives a daemon that does not, rather than printed.
Both websocket console plugins carried the same loop. While here, give
the openbmc one the parts tsmsol already had: text frames are data rather
than a surprise, and the client session is closed when the upgrade fails
and when the console does, instead of being leaked.
nodefirmware <node> disks reported the bmc version, and so did adapters and
misc. The generic handler takes a category and ignores it, and the ipmi plugin
does not filter either, so every category answered with the one entry ipmi can
report.
Apply the rule R13 established for a redfish inventory that does not categorise
itself: the bmc's own firmware is system firmware, so it answers for core and
for nothing else. Filtering in the handler that produces the entry leaves an oem
handler that does categorise its own firmware free to answer for more.
There is no ipmi command for a bmc hostname, so get_hostname falls back to the
DCMI management controller identifier. A bmc that does not implement the DCMI
group at all rejects that with "invalid command", which was handed to the caller
as if the request had been bad: nodeconfig <node> bmc read seven fields
correctly and then reported "Error: Invalid command", and the api answered 500
Unexpected error.
Route every DCMI request through one helper that turns "invalid command" and
"command disabled or unavailable" into UnsupportedFunctionality, so the
identifier, the asset tag and the hostname all report a platform that cannot do
this rather than a fault. Where the caller asked about a hostname, say that
rather than naming DCMI.
The whole group read still ends in one error line, because an operation a
platform cannot perform and one that failed are the same message to the client.
That is worth separating, but not here.
The unhandled tails of handle_request, handle_configuration and handle_alerts
raised a bare Exception('Not implemented'), so asking for a resource the
transport has no code for was reported as an unexpected error and logged with a
traceback. management_controller/location over ipmi is one such resource: R5
implemented it for redfish only.
Raise UnsupportedFunctionality naming the resource instead, which the plugins
already treat as its own case rather than a fault, and give decode_alert over
redfish the same treatment. Any resource added to the tree without an
implementation on one transport now reports that plainly.
nodereseat printed "Error: " and nothing else against a bmc that refused the
credentials. The redfish plugin reports it properly, but the message it emits
re-raises TargetEndpointBadCredentials with no arguments when a single node is
addressed, and the enclosure plugin renders that with str(e), which is empty.
Both hardwaremanagement plugins already had a helper for exactly this, one copy
each. Keep one in confluent.exceptions instead, teach it to fall back to the
description a confluent exception carries by class before falling back to the
exception name, and use it in the enclosure plugin too.
get_error_body had the mirror image of the same bug, joining the class
description and the message unconditionally and so answering "Bad Credentials -"
with a separator and nothing after it. The apierrorstr property beside it
already gets this right, so use it.
The exit callback opened the pid file unguarded, so when it was already gone the
atexit handler raised FileNotFoundError and python reported an exception ignored
in an atexit callback. The removal of the debug socket immediately above is
guarded, so this was an oversight rather than an intent. Verified by stopping
the service with the pid file deleted first.
terminate() called sys.exit(0) from a signal handler while the asyncio loop was
running. The SystemExit escaped run_forever, and closing the loop afterwards
then failed with "Cannot close a running event loop", so every clean stop wrote
a cascade of tracebacks and left a pending task behind.
Ask the loop to stop instead: deliver the signals through add_signal_handler,
which is the signal safe route, and set an event the main coroutine waits on so
that it returns and asyncio can unwind itself. Flush configuration on the way
out, as the client requested shutdown has always done.
That client requested shutdown went the same way, calling sys.exit from inside a
request coroutine, and it is the route the systemd unit uses to stop the
service. It now asks for the same orderly stop through a hook the running
service registers, keeping the old behaviour when nothing is registered.
Measured on all three routes, with redfish and ipmi sessions to a bmc in flight:
under a third of a second and no tracebacks, where before each one wrote a
cascade.
The web session helper sent X-CSRF-Token. MegaRAC checks X-CSRFTOKEN, so the
login succeeded and then every request answered Invalid Authentication, which is
why this helper has never worked. Confirmed both ways against a bmc: the same
request answers 401 with the old spelling and 200 with the new one.
IndicatorLED is deprecated in redfish in favour of the boolean
LocationIndicatorActive, so a platform that only implements the newer property
reported no identify state and could not be told to light up. Read either one,
preferring the older where both appear, and write the boolean when that is what
the resource offers. The boolean has no way to express blinking, so a request to
blink lights it steadily.
The led resource shares the same reader, so it gains this as well.
Asking for the mac addresses of a node whose inventory does not describe any,
which is every node reached over ipmi, printed absolutely nothing and exited
successfully, leaving no way to tell an empty answer from a broken command. Name
what was asked for instead. The exit code stays successful, since an inventory
that does not mention something is a valid answer rather than a failure.
get_ntp_enabled returns None to mean the platform cannot tell us, and that went
to the caller as the literal text "None", which says nothing at all. Report it as
unsupported instead, which the tooling already renders plainly.
A discrete sensor reports no value, and the unit was appended regardless, so a
watchdog came out as "Watchdog:% " and an event log sensor as "SEL:". The unit
belongs to a reading, so only print it when there is one. On the platform this
was seen on the units field is itself meaningless for such a sensor, carrying a
percent sign and a trailing space from the sdr.
The lookup fell back to the generic handler whenever it was given the service
root, which is the early call during connection setup, so a bmc that names its
vendor only in the manager document was served by two different handlers on
one connection. Read the manager during that early call too.
A stray trailing comma made the update detail a one element tuple, so a
firmware error printed as a python tuple. A missing status printed the whole
response dict. A failure that named no node was dropped entirely, which is how
a service data request that the server refused came out as silence and a
success exit code, and nodestorage, nodelicense and nodesupport exited zero
even when they had reported an error.
nodeconsole crashed decoding an absent screenshot, and again on the terminal
calls behind a pipe, where a log replay crashed too; refuse the terminal only
modes cleanly and dump the log when there is no terminal to replay into.
nodedefine raised a ValueError on an argument without an equals sign, and
firmware for a category the target does not describe printed usage as though
the question had been malformed.
On the server side the readability check was applied to the path a download is
saved to, so asking for service data or saved licences at a path that does not
exist yet failed claiming the destination was not readable.
Deleting a user gave up after one attempt and reported why the fallback of
blanking the name failed rather than why the delete did. MegaRAC reports a
timeout for a delete that a second ask completes, so retry, check whether the
account went away, and keep the original error.
Attaching media judged a device free by ConnectedVia, which describes how the
device is wired to the host rather than whether anything is in it. Every
device on this bmc reports a fixed value there, so nothing was ever selected
and the attach reported success having done nothing. Judge by whether an image
is loaded, fall back to the properties when an advertised insert action is not
served, and say so when no device would take the image.
The firmware category was passed to the library and ignored, so core,
adapters, disks and misc all returned the same full list. Classify by what
each entry is related to, and answer only for core when a platform says
nothing about where its firmware belongs.
reseat_bay reached for a hardcoded Nvidia action on Chassis_0, so it failed
with a not found for that url instead of saying reseat is unsupported.
Five resources called methods that do not exist on the redfish client, so
each answered with an internal error naming the missing attribute: the leds,
the management controller identifier, the domain name, the remote kvm licence
and the alert destinations. Implement the first four from the manager network
protocol, the graphical console and the chassis indicator.
Alert destinations stay unimplemented on purpose. Redfish describes where to
send events with EventService subscriptions, which is a different model from
the numbered PET destinations this resource was built around, so say so and
drop the code that could never run.
The location resource fetched its data and discarded it, so a read produced
no output whatsoever.
Reads are cached for thirty seconds and a write did not invalidate anything,
so setting a boot device or the identify state and then reading it back
reported the value from before the write for up to half a minute. A write can
change documents other than the one written, an action url not being the
resource it acts on, so drop the cache rather than one entry.
A bare UnsupportedFunctionality() left the user with an error containing no
text at all, or with no output and a success exit code, so asking a platform
for something it does not implement looked like nothing had happened.
Name what is unsupported at each raise, treat it as its own case in the
plugins so it reads as a limitation rather than an unexpected error and does
not log a traceback, and fall back to naming the exception when an exception
still arrives with nothing to say. The generic redfish
get_extended_bmc_configuration was also declared without async while the
caller awaits it.
Three things stopped a redfish firmware update on MegaRAC. The AMI handler
opened by asking the bmc to preserve fourteen named settings, and a build that
knows a different set rejects the whole request, which aborted the update
before anything was uploaded; send only the keys the bmc advertises. The
multipart push carried the image alone, and the specification has it carry an
UpdateParameters part too, which this firmware enforces. AMI also wants an
OemParameters part naming the kind of image, and nothing was supplying one.
The kind of image is asked for rather than worked out from the file. The
extension is vendor habit rather than format, and the leading bytes answer just
as confidently about an image they have never seen, while being wrong means a
bmc flashed with a bios image. So nodefirmware takes --type, it travels as far
as the handler that wants it, and where the bmc publishes the types it accepts,
an unknown one is refused with the list, as is asking with none. A platform
that reads the kind of firmware out of the image itself refuses the option
rather than dropping it, so nobody aims an update somewhere they did not mean
to. A parameter file still wins, since it can carry more than the image type.
Updating the bmc takes the bmc, and the task being watched, away for minutes.
That is the update working rather than the monitoring failing, so wait a
bounded while for it to answer again instead of reporting a successful flash
as an error.
Closing a console gives up its claim on the session, and that talks to the
bmc, so let it fail the same way the console's own deactivate is already
allowed to.
kg was left as the caller passed it in both the register key and the reuse
check, so the mismatch fixed for the name and password still applied to it.
The count for a new socket was taken before binding it and before the io
task was known to be up. Take it last, once nothing is left that can still
fail.
Telling a keepalive that the session is gone can end up back in logout,
because reporting it is how a console gives up its claim. The inner pass
finished by clearing the register of keepalives while the outer was still
walking it, so the next entry was looked up on None. Two entries is what an
XCC has, the console's and the oem handler's.
Take the callbacks and give up the register before notifying anyone.
The check for sharing a live session compared the credentials a session
keeps encoded against the strings every caller passes, so it never matched
and each caller built another session beside the one it could not see.
Normalise both sides. Three commands to one bmc went from three sessions on
three sockets to one, and from five sessions open on the bmc to three.
Not from the port to asyncio: upstream compares the same two things the same
way.
One session is now routinely handed to a console and a command at once, and
logout closed it for both, leaving whoever was left holding one that
answered as though it had been lost. Count the holders and give up a claim
instead, unless the session is no longer usable, which logout is told by
sessionok.
A console had no way to give a claim back: close deactivated its sol payload
and left the session alone, which was right when closing meant closing it
for everybody and is a leak now. Both of its exits release it.
__init__ opened by checking for an initialized attribute, meaning it had
been handed a session someone else was establishing. That attribute is only
ever set further down in the same method, so the check cannot be true, and
the port lost the return that made it work upstream. Waiting for someone
else's login is done in __new__ now, so this is dead code claiming to
protect something.
logout decremented the count of sessions on a socket twice over, and
_mark_broken again for the case logout had not, so it went negative and kept
falling. _assignsocket picks the least used socket and refuses one at
MAX_BMCS_PER_SOCKET, and both of those read that number.
initting_sessions exists so a caller can share a session already on its way,
but the entry was added only once the login had finished, leaving the login
itself uncovered. Two callers asking at once each built a session, both on
the socket the other had not claimed yet, and replies route by bmc address
and local port, so only the last to transmit was ever answered. That is the
console session failing about one attempt in three.
Register before the login and remove the entry in a finally. The two old
removals keyed on the encoded credentials while the register is keyed on the
caller's strings, so they never matched and an entry outlived its session.
A session handed over mid login is now waited for rather than returned as
one that answers as though it had been lost.
A session that failed raised with no message at all, so a failed console
read "IpmiException: None". Record the reason wherever a session is marked
broken.
handle_users iterated get_users with async for while the same call is awaited
a few lines below, so listing the users collection, and creating a user,
raised a TypeError about a coroutine having no __aiter__. list_inventory in
the redfish plugin had the mirror of it, awaiting an async generator.
get_extended_bmc_configuration is called with hideadvanced but the ipmi chain
never accepted it, so the extra and extra_advanced resources raised a
TypeError; thread the argument through instead.
A user slot the bmc refuses to describe no longer takes the whole user list
with it: one MegaRAC slot answered Invalid data field for good after an
account was deleted, which broke every user operation.
MegaRAC refuses a PATCH with no If-Match header, so setting the bmc hostname,
ntp, the bmc network configuration, a location and ejecting media all failed
with a precondition error. set_identify already passed etag=*; do the same at
the call sites that did not, including the firmware push busy flag.
get_identify indexed the sysinfo method object rather than awaiting it, so
reading the identify state raised a KeyError naming a bound method on every
redfish system. Read the indicator from the chassis that owns the physical
led, since some implementations leave the copy on the system stale, and say
plainly when a platform does not report one.
The generic get_description took no fishclient while every other handler and
the caller pass one, so the description resource raised a TypeError on any
non Lenovo bmc.
The asyncio rework mistakenly follows up a long waiting acquire with blanking and starting over.
Now use 'True' as a sentinel value to trigger a fetch, otherwise, assume srp is ripe for processing.
Immediately discard srp after unpack and replace with sentinal value.
Fifteen error kinds beyond the two async ones report nothing on this tree
today and have something in it to bite on, so turning them on costs no
findings and keeps it that way.
invalid-syntax is the one that closes a gap rather than covering ground
another check already holds: CI compiles under a modern interpreter, which
accepts syntax the el8 and sles15 interpreters cannot parse. Checking
against python-version rejects it instead, which makes that setting load
bearing for the first time.
Kinds whose subject matter this tree does not contain stay off, among them
everything reached only through typing: the module is never imported, so
TypeVar and namedtuple naming and stale `# type: ignore` have nothing to
find here.
Every rule added here is at zero once the previous commit lands, so it costs
no cleanup: the point is that a future patch cannot introduce one without the
ruff job failing. They are the rest of pyflakes' format-string checks,
flake8-2020, most of bugbear, the pylint warnings that describe bugs rather
than style, four flake8-async rules for blocking calls in coroutines, and
some RUF, LOG, PGH, PIE, ISC and EXE rules in the same spirit. Each was
confirmed to fire on a synthetic violation, so none is silently inert under
the py37 target.
Rules are named by group wherever the group is already clean, and each prefix
stops short of a rule that is not: PLW150 rather than PLW15, which would pull
in PLW1510. Bugbear is listed rule by rule apart from B02 and B03, since
B006, B007 and B018 are all left out on purpose.
bin/confluentsrv.py is a python2 script that setup.py never lists in
scripts, so no package has ever installed it, and both the systemd unit and
the sysvinit script start bin/confluent instead. It had also drifted out of
step with what it calls: main.run takes the argument vector and this passed
none.
confluentsrv.spec goes with it. It is a PyInstaller spec whose only input
is c:/Python27/Scripts/confluentsrv.py, a path that has never existed in
this tree, left over from the Windows compatibility work.
WebConnection requires a port, and every other caller passes one, twice in
this very file. Following a relay URL during discovery raised TypeError
instead.
Four calls reached a base method with an argument list it does not accept,
so they raised TypeError about the argument count.
Three of them would have failed either way, since the base only raises
UnsupportedFunctionality. What changes there is that the failure becomes
the intended, catchable one rather than an argument count error the caller
cannot interpret. get_diagnostic_data grew an autosuffix argument
everywhere except the ipmi generic handler, which is the handler used for
unrecognized hardware. The redfish generic handler already had it. The two
storage super() calls dropped the cfgspec they were given.
The fourth is a real fallback rather than a message: the XCC user_delete
dropped the fishclient it receives from redfish/command.py, so deleting a
uid the XCC does not list raised TypeError instead of attempting the
generic Redfish delete.
ConfluentTargetNotFound takes the node as its first argument, and every
other caller passes it. The two inventory plugins constructed it with no
arguments at all, so asking for a component that is not in the inventory
map raised TypeError instead of returning the 404 the path was written to
return.
Both now follow the pattern used a few lines further down for volumes and
name the component that was not found.
The retry around the firmware progress poll named socket.socket, which is not
an exception class, so the moment the request it guards actually failed Python
raised "catching classes that do not inherit from BaseException is not
allowed" in place of the error.
socket.error is OSError, which is what a failed poll raises and what the retry
below was written for.
The loop resolver takes only host and port positionally, so these three calls
raised "BaseEventLoop.getaddrinfo() takes 3 positional arguments but 5 were
given" every time they ran.
get_ipaddr and _find_service have no handler above them, so link local XCC
discovery and a targeted SSDP search both died outright. The snoop copy sits
under an except Exception, which swallowed it and left the MGTIFACE reply
unanswered instead.
Each is the only thing keeping its rule group from being selectable whole.
userutil.py imported ctypes with a star; the names it uses are POINTER,
byref, c_char_p, c_int, c_int32, c_uint and cdll. The oem lookup loop had an
else with no break, so the else always ran. The alert parameter table wrapped
int in a lambda that only forwards to it. And the watchdog interval passed 0
where os.environ.get documents a string, which worked because int(0) is 0.
get_sensor_names and get_sensor_descriptions reach get_psu_count for any
sensor whose table entry carries elementsfun, and get_psu_count is a
coroutine. As plain generators they could not await it, so range() was handed
the coroutine object and enumeration died with "'coroutine' object cannot be
interpreted as an integer".
Every DW612S has such entries, so nodesensors returned nothing for the
enclosure. get_sensor_descriptions was doubly broken: the Lenovo handler
already iterated it with async for, which a plain generator cannot satisfy.
Verified against a DW612S SMM (FPC variant 38). Before, descriptions raised at
the async for and readings raised partway through enumeration; after, both
return all 34 sensors, 19 of which are the PSU entries that never enumerated.
RUF006 catches a create_task whose result is discarded. The loop holds only a
weak reference, so such a task can be collected while still pending and the
work disappears without a trace.
Selected last, once the three existing offenders are gone, so the tree stays
clean under it from this commit on.
run_handler scheduled the coroutine that serves an async HTTP request and
dropped the returned task. The event loop only keeps a weak reference, so the
task could be collected while still pending, leaving the request unanswered
and "Task was destroyed but it is pending!" in the log.
The session already outlives the request in _asyncsessions, so it holds the
task in a set and discards it from a done callback.
nodeconsole spawned a task per input byte from the stdin reader callback
and kept no reference to it. Two of those tasks overlap as soon as one
parks in relay_keypresses waiting on the VNC connection, so keystrokes
can reach the node out of order and the escape sequence state (buffer,
inputcontext, modkeys) is mutated by more than one task at a time. With
the first keystroke relaying slowly, typing abcdef arrives as bcdefa.
Those tasks were also unreferenced, which asyncio documents as
collectable while still pending, so a keypress could be dropped.
Queue the bytes in the reader callback and process them from one
long-lived task instead. Keep a reference to the watch_input task as
well, since collecting that one takes the whole input handler with it.
local_node_trust_setup() called get_cluster_list() and sign_host_key() without
awaiting them, so "osdeploy initialize -l" aborted with "TypeError: cannot
unpack non-iterable coroutine object" before doing any work.
Both awaits have to land together: sign_host_key() is called in a loop that
unlinks the existing ssh_host_*_key-cert.pub before writing the new one, so
fixing only the unpack would delete every host certificate and then fail.
If someone wants to seal to a PCR
explicitly to prevent booting rescue, the PCR is likely to
extend differently during install.
Leave the volume sealed to the tpm without any PCRs until first boot.
Then wipe the bindings without PCR specified, and seal according to user preferred values.
It has been observed there are times where an ethernet switch is partially working with MLD snoop/IGMP snoop. A workaround for the unreliable behavior seems to be to reassert multicast joins ever so often.
Give it a try to restart the SSDP sockets every minute.
The tree is clean under it now, so the job can fail the run and catch the next
async regression instead of only reporting one.
Nothing is suppressed beyond the two ignores that state their reason at the
line, for an async __new__ and an untyped callback registry, neither of which
pyrefly can model.
The plugin was written against the http.client based SecureHTTPConnection, and
when that went away the reference was pointed at the aiohttp WebConnection,
which shares the name and nothing else. Nothing in it could run: the transport
called an async request() without awaiting it and then reached for a
getresponse() the new class does not have, and three PDUClient methods that
were never coroutines were awaited by the entry points.
Two transports now, both local to this plugin. https is aiohttp and stays on
the event loop, since the cert verifier records new fingerprints through
tasks.spawn. http is http.client in a thread, with its own socket so it can
still ask for a smaller segment size before connect: aiohttp only takes a
socket factory from 3.12 on, newer than el9, el10, ubuntu 24.04 or Leap 16
ship. That side has no cert to verify and its credentials arrive already read,
so the thread touches nothing.
connect() establishes and authenticates, wc is just the accessor now, and
logout() no longer sends a session id it never obtained. update() reports an
unsupported element instead of raising NameError.
On the https side cookies follow aiohttp's domain rules and the one POST with
a body goes out as text/plain, where http.client replayed every cookie and
sent no content type. The http side is as before, and neither can be settled
without an Eaton PDU on the bench. Both transports were exercised against a
stand-in: login, outlet read and set, sensors, logout, and a clamped segment
size on the plaintext path.
virEventRunDefaultImpl waits for an event that an idle domain need not
produce, so the thread could outlive a deactivation that reported success, and
every later activation was refused while it did. Registering a timeout is what
makes it return: measured, a thread with nothing registered was still running
four seconds after being asked to stop, and with a half second timer it came
out at once.
Deactivation dropped its reference once the wait expired, whether or not the
thread had stopped, so the next activation started a second one and revived
the first by setting run_console again. The reference is cleared only when the
thread is really gone, and activation refuses with 0x80 while one is alive.
That makes the wait a courtesy rather than a correctness measure, so it drops
to a second.
Activation started an event thread whichever way the base handler had just
answered, so a refusal started one anyway and an already active console got a
second. activated alone cannot tell the two refusals apart, being true
already on the already active path, so the value from before the call decides.
Deactivation joined that thread on the event loop, where it could stall every
other session and its own response. The wait moves off the loop and is
bounded, and the thread is a daemon.
The console awaits its output handler, but both sample BMCs supplied a plain
function, and both dropped the send_data coroutine. virshbmc additionally
receives its stream callback on a libvirt thread, so the send goes through
run_coroutine_threadsafe against the loop captured at activation, called
asyncloop because the class already has a loop method it uses as a thread
target.
ServerConsole asked the session layer to retry, which a ServerSession cannot
do: it never runs Session.__init__, so it has no timeout, and its _timedout
does nothing. IpmiServer.logout was synchronous and an argument short while
_cleanup awaits logout(False). The boot options handler answered, then read an
unbound name and answered again with 0xff.
bmc.py, serversession.py, fakebmc.py and virshbmc.py were byte identical to
upstream pyghmi: the async port went through the session layer beneath them
and left the server side alone. So every response created a coroutine and
dropped it, and the overrides the now async parent awaits returned None.
Running fakebmc bound no socket, spun a core, and answered nothing.
Everything that sends a response is a coroutine now, and so is the dispatch
that reaches it; the hooks a subclass implements stay ordinary functions, so
an out of tree Bmc is unaffected unless it overrides the payload handlers.
Two things had to leave their constructors, both being coroutines: assigning
the server socket, into bind(), which is why listen() is no longer a
classmethod, and answering the open session request, into
send_open_session_response.
Verified against upstream with ipmitool over nine commands, with identical
output. SOL is not covered: fakebmc reports the payload disabled on both.
Housekeeper dates from when this work was a blocking select loop. Under
asyncio the event loop is already that place, and a thread around it takes a
captured loop reference, a scheduling step before the thread starts, and a
daemon flag, only to sit waiting on a coroutine that never returns with no way
to stop it. A task runs on the right loop by construction and cancels. The
loop is looked up before the coroutine is built, so calling this without one
raises rather than stranding it.
Nothing in this tree used the class. Out of tree callers need
start_housekeeping() instead, and gain the ability to stop it.
TsmConsole created an aiohttp ClientSession and never closed it, and leaked it
again when ws_connect failed. Neither was reachable before the connection path
was repaired. It is closed on both paths, and starts as None so that closing
before a connect does not trip over a missing attribute.
A task per read swallowed failures, could deliver keystrokes out of order, and
at end of input returned with the reader still registered, so a level
triggered selector called it again for the same EOF. The reader queues now,
one consumer sends in order, and it is gathered with the main loop so a
failure reaches the caller.
Three faults in the same few lines. The redfish Command lost its constructor
for an async create, so building one raised TypeError, which the except below
reported as TargetEndpointUnreachable. await_redirect is defined nowhere in
this repository's history, so that call raised too; create performs the
session setup it was meant to trigger. And oem is a coroutine method rather
than an attribute, with its web connection coming from get_wc, which is what
performs the login that sets csrftok.
The retry pass calls apply_configuration with lastchance=True, which only
NetworkManager accepts. It is unreachable today, since only NetworkManager
returns the 1 that fills the retry list, but it springs the moment either of
the others grows a return, or the retry selection is brought in line with the
first pass. Matching the signatures costs nothing.
ConsoleSession grew an async create and lost its constructor, and ShellSession
inherits that. sockapi was updated for the console branch but not the shell
branch immediately below it, so opening a shell session raised TypeError. It
is the only place in the tree that builds one.
Command.eventloop called wait_for_rsp with no timeout. With nothing waiting or
being kept alive there is no deadline to derive one from, so it returns
without suspending and the loop runs flat out, measured at over 100000
iterations in two tenths of a second.
MAX_IDLE gives it something to wait on, as a ceiling rather than a fixed
delay: real deadlines still shorten it and an arriving packet still ends it
early.
The Console was never connected, so main_loop ran against a session that had
never been established. Input arrived on a thread that called send_data and
dropped the coroutine; the thread is gone and the loop watches stdin with
add_reader instead. The output handler was a plain function that
Console._print_data awaits.
Command lost its constructor for an async create classmethod, so the utility
raised TypeError before connecting. The onlogon callback it was built around
is gone as well: create establishes the session itself, so both the callback
and the eventloop that waited for it are unnecessary.
docommand awaited nothing, so every operation produced a coroutine that was
printed and dropped, and three of its calls are async generators. Each BMC is
handled in turn now rather than only the last.
Console.main_loop drives Session.wait_for_rsp, which is a coroutine, so it
spun without ever waiting for a packet. It is a coroutine now, and pyghmicons
runs its main under asyncio.run. pyghmiutil had the same shape around
Command.eventloop.
sockapi sent its collective refusal without awaiting tlvdata.send. redfish
handle_sensors returned a coroutine from most branches and None from the short
ones, so the caller's await raised TypeError; it is a coroutine throughout
now. console send_payload waited for a response without awaiting the wait.
Two suppressions are pyrefly limitations rather than bugs: Session defines an
async __new__, and the keepalive registry holds coroutine functions in an
untyped dict.
Taken from fix/asyncio-port-critical, limited to what pyrefly reports.
The SMM handler still used the httplib style connect/request/getresponse that
the async webclient does not have, so nothing was ever sent. It goes through
grab_response_with_status now, with allow_redirects=False to preserve httplib
behaviour, hence the new webclient parameter. That also fixes a login passing
its headers as urlencode's second positional argument.
Delta PDU's logout was a plain function both callers awaited, so the power
paths raised TypeError. XCC's set_system_configuration was the last
synchronous implementation of a method every caller awaits.
These __main__ blocks called coroutines as if they were functions, so they did
nothing at all. Single calls go through asyncio.run; sshutil, proxmox and
vcenter needed an _selftest coroutine. Two were invisible to pyrefly because
repr() and list() count as using the result: vcenter's get_vm_serial needs an
await, and proxmox's get_vm_inventory is an async generator.
lldp called _extract_neighbor_data twice, once correctly, so the bare call is
dropped. xcc3.remote_nodecfg is not a self test: every other handler defines
it as a coroutine and selfservice.py awaits it.
Catches unawaited coroutines and awaits on non-awaitables, which ruff and
compileall cannot see. Only those two kinds are enabled, and in pyrefly.toml
rather than a --only flag so CI, local runs and editors agree; a full check
reports thousands of errors on this largely unannotated tree.
search-path is what lets imports resolve, and without it a large share of the
findings, including everything in the SMM handler, goes unreported. Pyrefly
walks *.py only and a glob does not lift that, so extensionless tools are
named individually. No baseline: it matches by file, kind and column, so a new
mistake at the same indentation as an old one would pass unnoticed.
Advisory until the existing findings are dealt with.
lxml refuses a str carrying an encoding declaration, and the SMM declares one
when it answers /data/login, so _webconfigcreds has raised ValueError on the
first thing it does after logging in ever since the switch to lxml. stdlib
ElementTree took the same input, which is why it went unnoticed. fromstring
already means to take either shape, so encode there. Confirmed against a
DW612S.
compileall only ever compiles *.py. Handed anything else, even by name on the
command line, it skips the file and still exits 0, so the job has been
checking 218 of the 298 Python files in this tree and reporting success for
the rest. Everything without the extension went unchecked: the whole of
confluent_client/bin and confluent_server/bin, the osdeploy deploy scripts,
the loose misc utilities and the setup.py.tmpl templates.
Those files now go to py_compile, which compiles what it is given. The list
is built from python shebangs, read with the shell builtin rather than by
forking head and grep per file, plus four patterns for the files that carry
no shebang at all and so cannot be detected: the setup templates, configbmc,
add_local_repositories and misc/filterpasswd. It comes to the same 298 files
ruff.toml arrives at through extend-include, and wants keeping in step with
it.
The shebang test matches python anywhere in the line rather than after a
slash or space, because several tools use /usr/libexec/platform-python.
Adds the checks whose findings were cleared in the two preceding commits
(F401, F541, E701, E711, E712, E713, PLC0414) plus three that were already
at zero and cost nothing to lock in: E401, B015 and B023.
B905 is deliberately left out even though it also reads as clean: it only
reports on py310+, and satisfying it would mean adding a keyword the oldest
interpreters this tree runs on cannot parse.
Hand written rather than autofixed, since three of the four need the
surrounding code read to be sure they are equivalent:
- confetty: `powerstate == None` -> `is None`.
- nodeconfig: `setmode != True` / `!= False` -> `not setmode` / `setmode`.
Safe because setmode only ever holds None, True or False, and the two
lines above each test normalise None away first.
- pam: split two `if cond: stmt` one-liners.
- imgutil: `from shutil import copytree as copytree`, an alias that renames
nothing. Not a re-export marker, this is a script.
Entirely mechanical, produced by `ruff check --fix --select F401,F541,E713`
and reviewed rather than taken on faith: deleting an import is only safe if
nothing imports it for its side effects or re-exports it. None of the 19
removed names is referenced anywhere in its file, none appears in any string
literal, and none of the touched files uses eval, exec, globals() or
__import__, so there is no dynamic lookup that could reach them.
The enforced rule set is deliberately narrow: undefined names, statements
in impossible positions, duplicate definitions, invalid escapes and a
couple of bugbear checks that only fire on genuine defects. No style
rules, and the tree is clean under it as of the preceding commits.
Discovery needs help. Ruff only walks *.py, and about a quarter of the
Python here has no extension: every node* CLI tool, the server bin tools,
the osdeploy scripts (some of which carry no shebang either) and the
setup.py templates. extend-include lists them, and *.sh is excluded so
the shell scripts sharing those directories are not parsed as Python.
The CI job pins both the action and the ruff version, since there is no
pyproject.toml for the action to read a version from and an unpinned
`latest` would let a new ruff release fail an unchanged branch.
The guard evaluated the `next` builtin and discarded it, which does
nothing, so a pending node with no handler fell through to
None.NodeHandler(...). The AttributeError was caught by the enclosing
except and logged as "Unexpected error during discovery", turning a node
that should have been quietly skipped into a spurious error in the log.
load_plugins() used `plugin` as the loop variable for plugin file names,
which shadows `import confluent.plugin as plugin` for the whole function.
Nothing in the function needed the module, so this was latent rather than
broken, but the next line that does need it would have failed oddly.
The DedicatedSpareDrives payload was built as a set containing a list
containing a dict comprehension with a constant key, so it collapsed to a
single entry and then raised TypeError on the unhashable list. Build a
list of drive references, the same shape as the Drives list just above it.
Three names were defined twice in the same scope, so the first definition
was unreachable:
- lenovo OEM handler: two set_user_access methods, the second silently
replacing the first. That made the SMM privilege update dead code. The
conditions are mutually exclusive (is_fpc returns None once has_xcc is
true), so merge both into the surviving method.
- redfish plugin handle_cert_authorities and prepfish
disable_host_interface: byte identical copies, drop the redundant one.
Each of these loops rebinds the name that holds the iterable. They work
today because the iterable is evaluated once before the loop starts, but
the name is then gone, so any later use reads a loop item instead of the
collection.
- nodeinventory: `for arg in args` / `for arg in arg.split(',')`.
- confignet (common and debian copies): iname holds the comma separated
interface list and is then reused for each interface in it.
- xcc _get_agentless_firmware: adata holds the adapter query response and
is then reused for each adapter.
No behaviour change, just distinct names for distinct things.
Every one of these raises NameError if its code path is reached:
- nodeapply: run_automation accumulated into an exitcode that only existed
in run(), so any automation error crashed instead of being reported. It
now keeps and returns its own, tracked separately from the exit code of
the ssh commands: the early exit after the spawn loop tests that one,
and folding automation failures into it would exit with children already
running and their pipes abandoned. Both are reported at the real exits.
- nodeconsole: redraw() reads firstnodename, which was local to
do_screenshot(); promote it to a module global like the other drawing
state.
- nodedeploy: the redeploy path appended to a lockednodes list that did not
exist yet. The block that follows re-reads the same lock state and acts
on it, so drop the dead duplicate.
- samples/nodeattrib_from_switch.py, misc/filterpasswd: missing import sys.
- xcc3: fixuuid was never imported. xcc imports xcc3, so take a local copy
the way the smm handler does instead of creating an import cycle.
- httpapi: the async session call still passed the WSGI-era env and an
extra argument to handle_async(), which has taken only querydict since
the aiohttp port. Calling it correctly exposed that handle_async()
registers an AsyncSession before raising on the discontinued long poll
path, so every request to it would leak a session that is never reaped.
It now only creates one when there is a websocket handler to yield it to.
- messages: the InputFirmwareUpdate.filename property checked
self.filebynode[node] with no node in scope. __init__ already validates
every expanded path and nodefile() rechecks per node, so drop the checks.
- pam: drop the python2 branches referencing unicode and raw_input. The
server has been python3 only since the asyncio port.
- cooltera: the sensor-name listing referenced a nonexistent sensors dict.
The available sensors depend on the model, which is only known after
reading the device, so list them from the same status data the readings
use.
- deltapdu, eatonpdu, geist: the not-implemented response in update() used
node outside the loop, unlike retrieve() in the same files and unlike
raritan/enlogic.
- confluentdbgcli: stray self. on a module-level socket connect.
This allows a user to opt into pcrs if they understand what they are doing.
Some PCRs are sensitive to firmware updates and some are sensitive to boot loader, kernel, boot config, or initramfs. All of these are an opportunity for an unsuspecting update to remove access to the boot volume. There are update processes that can be put into place to make this work,
but it is up to the OS update process to address that, and
OS update processes are likely not to address that at this time.
reply_dhcp4 logs the insecure mode remediation hint on every DHCP
discover it refuses. A node in this state never receives a reply, so it
retries for as long as it is powered on and the same message repeats
every few seconds.
Rate limit it per hardware address the way the neighbouring boot attempt
messages already do, reusing the ignoremacs window that check_reply uses
for the missing profile hint.
The per-MAC 90 second log throttle in proxydhcp has been inert: the
`skiplogging = True` reset sat in relay_proxydhcp, where it is a dead
local, while the loop in proxydhcp only ever assigns False. Once the
first packet is handled the flag stays False for the life of the
process, so every retransmitted boot request logs again even though
ignoredisco is updated to suppress it.
Reset the flag at the top of each loop iteration instead, next to the
timestamp check it belongs to, and drop the dead assignment.
reply_dhcp4 declines to answer a PXE boot request unless
deployment.useinsecureprotocols is set to firmware or always, but
proxydhcp had no such check. A node left at the default of never was
therefore still offered a TFTP bootfile and a plain http boot.ipxe URL
whenever the request arrived on port 4011 rather than port 67, so the
attribute silently did nothing in ProxyDHCP deployments alongside an
independent DHCP server.
Apply the same gate, including the UEFI HTTP boot exemption, and log the
same remediation hint. The node attributes are now fetched once and
passed through to get_deployment_profile instead of being looked up
again there.
Requests whose architecture could not be determined are ignored rather
than falling through to the reply. opts_to_dict stops parsing before the
client architecture option whenever the message type is not a request,
and such a packet would otherwise reach the iPXE branch and be handed a
plain http boot.ipxe URL without ever passing the gate.
The default 'cp' is /lbin/cp, but that fails, use full path to the cp that works.
Perform the hmac registration of api key that was missing.
Remove assumption that the ip will be ipv6, wrapping it only if a : is present in address.
The rsync push carried no preservation flags, so files arrived with their
special permission bits explicitly disabled and a setuid/setgid entry could
only be honored by the permissions= chmod on the client side.
Preservation was turned on once before in e52a9ff70f ("Have syncfiles
attempt to preserve more") and rolled back the same day in c0287e93ed
("Roll back rsync ownership"), because rsync also applied the staging copy's
attributes to the parent directories it merely traversed on the way to the
synced files, clobbering the permissions of system directories such as /etc.
Naming every staged file explicitly through --files-from and adding
--no-implied-dirs confines preservation to the content actually being
synchronized, leaving traversed directories alone and creating missing ones
with default attributes.
Two details follow from the way the staging tree is built. Files are staged as
symlinks, so rsync reads their attributes through to the real file, but
directories are staged as directories and need the source attributes copied
onto them for the otherwise empty ones that have to be named explicitly.
Ownership is mapped from the account the daemon runs as to root, since that
account generally does not exist on the node and would otherwise arrive as a
meaningless numeric id.
--xattrs from that earlier attempt is deliberately left out: with --copy-links
rsync reads xattrs off the symlink rather than its referent, so it transfers
nothing here while adding a failure mode on hosts without xattr support.
chown() clears the setuid bit of a file on Linux (and its setgid bit, if
the file is group-executable), even when run by root and even when the
owner/group are unchanged. Since the owner/group chown ran after the
permissions chmod, any syncfiles entry combining owner=/group= with a
setuid/setgid permissions= value silently lost the special bits.
get_syncresult() caught the sync task's exception, logged a repr server
side and returned 200 OK with a null body. The node then called
.get('options') on that null resulted in:
c1: 'NoneType' object has no attribute 'get'
and syncfileclient still exited 0 as if syncing had succeeded.
Return the error to the requestor as a 500 with an error payload. On
the node, unwrap the body that grab_url_with_status raises for a
non-success status, print it once and exit non-zero. Only a failure the
server deliberately reported for this sync is terminal. Anything else,
such as a dropped connection, is re-raised so the existing retry loop
handles it as before. The same case now reports
c1: Error performing syncfiles: Syncing failed due to unreadable files: /etc/dangling.conf
c1: 'syncfileclient' exited with code 1
nodeattrib/nodegroupattrib -e replaced '.' with '_' in the attribute name
before handing it to the server, not just when looking up the environment
variable. Any attribute with a dot in it was therefore rejected, e.g.
$ export info_note=test
$ nodeattrib -e gpu1 info.note
Traceback (most recent call last):
File "/opt/confluent/bin/nodeattrib", line 97, in <module>
exitcode=client.updateattrib(session,args,nodetype, noderange, options, argassign)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/opt/confluent/lib/python/confluent/client.py", line 688, in updateattrib
key, os.environ[key.upper()])
~~~~~~~~~~^^^^^^^^^^^^^
File "<frozen os>", line 714, in __getitem__
KeyError: 'INFO_NOTE'
$ export INFO_NOTE=test
$ nodeattrib -e gpu1 info.note
Error: Bad Request - info_note attribute on node gpu1 is invalid
Keep the attribute name intact and derive the environment variable name
from it separately. A missing environment variable now reports which
variable names were looked for instead of raising a bare KeyError
traceback.
source_remote imageboot.sh is the last thing the diskless cmdline hook
runs, so returning early on a failed extraction ended the hook and left
dracut to time out. Falling through instead reaches the existing
/sysroot/sbin/init guard, which reports the failure and holds the node so
it stays reachable over ssh, as it did before extraction was checked.
imageboot falls back to cp when unsquashfs is unavailable, but the build
side did not: a bare dracut_install/copy_exec aborts initramfs generation
when the binary is absent, and the capture prerequisite check refused to
capture the image at all.
Mark the initramfs copies optional and report the missing package as an
advisory rather than a hard prerequisite, so such images still build and
capture, just without the faster extraction path. Widen the EL check to
every release past el8 so future ones inherit it.
The copy ran after the cd to the repo root, so it read ../LICENSE from
outside the checkout and never placed the file in confluent_osdeploy/. The
tarball went out without it and the spec's %install, which does
"cp LICENSE" after %setup cds into the unpacked directory, failed.
imgutil/buildrpm already copies before its cd; do the same here.
The aarch64 spec has its LICENSE lines commented out, so only the x86_64
build broke, but the copy was equally wrong in both scripts.
Release tags do not live on master: 3.15.2 through 3.15.6 were tagged on branch
3.15, so git describe reaches only 3.15.1 and dev builds were stamped
3.15.2.dev<n>. Besides being confusing, rpm and dpkg both rank the released
3.15.6 above that, so a dev package will not install over a released one.
Add a top-level VERSION file naming the release the branch is working toward
(4.0.0 on master) and a mkversion helper that stamps packages from it, keeping
the tag-derived value as a floor so a forgotten bump cannot go backwards.
mkversion also replaces the block copy-pasted into seven build scripts, and
makesetup no longer writes a per-package VERSION file, so the stale checked-in
confluent_common/VERSION goes with it.
confluent_imgutil does not depend on confluent_server, and the yaml
import is optional too, yet both capture and pack dereference osimage
and yaml unconditionally when writing manifest.yaml. With
confluent_osdeploy present but the server absent that raises rather
than producing a profile.
Guard the manifest on both being importable and say so, since rebase
is what the manifest exists for. The yaml fallback now binds None
instead of leaving the name undefined.
The two call sites carried the manifest write verbatim in both, so fold
them into one function rather than duplicate the guard as well.
osimport polls import progress from a coroutine, so a blocking sleep
between reads stalls the whole client loop. It was the only use of time
in the script, so the import goes with it.
Both loops that read the importer's output test for a percentage first,
so an ERROR: line whose text carries a % takes the percentage branch and
float() raises instead of the error being reported. The import target
name can carry one too, and that one is user supplied. importmedia runs
as a bare task, so the exception is swallowed and the client polls a
phase that never advances.
Test for ERROR: first and treat an unparsable percentage as no
percentage. Set percent on the error path of the second loop as well,
as the first already does.
The loop that drains the importer's remaining output reads a byte at a
time but clears currline on every iteration, one level out from where
the earlier loop clears it. currline is therefore never longer than a
single byte, so the percentage and ERROR: branches can never match and
the tail of an import is silently discarded.
Clear it only once a line has been consumed, as the earlier loop does.
scan_iso walks an entire ISO with blocking libarchive reads, yielding
only once per entry, and the header-sum branch of fingerprint reads the
whole file with no yield at all. Both run in the daemon, reached from
MediaImporter.init on every fingerprint and importing request.
The scan costs about 8us per entry and is indifferent to media size,
since libarchive seeks past file data rather than reading it: measured
at 80ms for 10k entries whether the image is 0.2 GB or 8.8 GB, and at
310ms for 40k. The header-sum branch is the one that scales with size,
reading a multi-gigabyte image end to end.
Make the pair plain functions and hand them to a thread instead.
Nothing unwound the loop device and dm-crypt mapping when the copy loop
raised, so interrupting a pack stranded both, still holding the profile's
rootimg.sfs.
Tear them down from a finally. The retry loop moves with them, so also
honour its tries counter, as unpack_image already does; spinning forever
inside a finally would hang the interrupt it is meant to clean up after.
A bounded retry loop can also give up, and the detach that follows would
then fail with EBUSY and, raising from a finally, replace the exception
that brought us here. Warn and leave both in place instead.
Both functions became coroutines solely to await one get_hashes call,
but their bodies are long stretches of blocking work: mksquashfs, the
encrypt_image copy loop, rsync, ssh and osdeploy.
From Python 3.11 on, asyncio.run installs a SIGINT handler that cancels
the main task and returns rather than raising, so an interrupt is only
noticed at the next await. Interrupting a pack during mksquashfs
surfaced as a CalledProcessError from the dying child instead of a
KeyboardInterrupt, and with a base profile, where nothing is ever
awaited, pack carried on and published the profile before exiting.
Run the loop only around the call that needs it.
check_output unwraps a single list argument, check_call never did, so
callers passing a list hit a TypeError out of create_subprocess_exec.
Two callers do: the genisoimage run behind Windows profile imports,
where an except Exception swallows the failure and the boot.iso is
silently missing, and the nodeconfig run in discovery, which takes out
automatic node configuration on discovery outright.
capture and pack hash the whole profile directory, which by that point
holds rootimg.sfs, the kernel and the distribution initramfs. rebase
only ever looks up entries that came from the profile source directory,
so the image blobs cost gigabytes of hashing for nothing.
Pass the source directory as the filter, as generate_stock_profiles
already does. Older manifests keep working, since rebase reads their
entries with a default.
The asyncio port added an await between every 2048 byte read, which
roughly doubled the cost of hashing. imgutil runs entire packed images
through this, and the server pays it on rebase and media import.
Read a megabyte per iteration instead. That still yields hundreds of
times per gigabyte, so the event loop stays responsive, and sha512 is
independent of the read size, so existing manifests remain valid.
grab_response_with_status hands back bytes, so every failure path put a
b'<status>error</status>' repr in front of the operator rather than what
the SMM said. Decode at the raise, replacing rather than failing on a
body that is not valid utf8. The bodies still reach fromstring() as
bytes, which is what lxml wants when the xml carries an encoding
declaration.
The poll loop spends its retry budget on a poll that goes unanswered but
aborted the update on the first non-200, even though an SMM restarting
its web service part way through the apply keeps answering, with
whatever its httpd has to say, before it stops answering at all. Give a
bad status the same budget as a dead connection.
Every other /data call checks the status, this one discarded the
response, so an SMM that refused the user record was reported to the
caller as a successful privilege change.
Staleness is judged by age alone, so a session the SMM ended on its own
reached the operator as a raw error body instead of being retried.
Route the /data calls through a helper that logs back in and retries
once on a 401, which is how the SMM answers once a session is gone.
A hostname or domain write does not end the session, measured on a
DW612S at firmware 1.18, so this covers what the chassis drops by
itself, not a self-inflicted loss.
The counter is there to ride out a few unanswered progress polls, but
nothing ever cleared it, so three failures spread across a long apply
exhausted it and aborted an update that was still making progress.
A firmware update posts on one session for the minutes its apply loop
runs, and an FFDC collection downloads on the session it acquired, but
wc() judges a session by its age alone, so a settings call arriving
thirty seconds in logged that session out from underneath them. Flag
the long operations the way the IMM and XCC handlers already do and
leave their session in place.
gather propagates the first exception and leaves its siblings running,
so a transport level failure against one node ends the import with a
traceback while the rest of the batch is cancelled at loop shutdown.
The forked children used to contain such a failure to their own node.
assign_macs already reports an error response itself, so this is the
connection dropping rather than the server refusing the assignment.
Collect the exceptions instead, report each one and count it towards
the exit code. Schedule the assignments as tasks while doing so, since
the plain coroutines are left unawaited if defining a later node raises
before the gather is reached.
Replacing the forked children with a gather kept their fan-out: every
row of the import file gets a session of its own and they all start at
once, so a large file opens a local socket and a server side session
task per node simultaneously.
Hold a semaphore for the duration of each node's assignment instead, so
a finished node's session is dropped before the next one starts. Also
build that session once per node rather than once per MAC, and say why
the caller's session is not reused, which was self evident while this
ran in a forked child.
register_endpoint primes current but not total, so a first response
without a count field goes straight to
UnboundLocalError: local variable 'total' referenced before assignment
on the elif. Start at zero, which skips the progress line until the
server does report a count.
Every getter and setter logged out on the way out, which nulled the
cached client and made the session cache inert on exactly the paths it
was meant to serve: a single nodeconfig walk of ntp costs two full
logins for the read and one per server for the write, each of them a
fresh TLS handshake plus, on firmware that omits st2, two extra page
fetches to scrape the tokens.
Leave the session in place and let wc() dispose of it once it expires.
This also stops one coroutine's logout from invalidating the session
another coroutine just fetched and is about to post with.
The retry counter is there to ride out a few unanswered polls, but
exhausting it broke out of the loop with complete still unset and fell
through to the 'complete' return, so an SMM that stopped answering
part way through an apply was reported to the operator as updated.
Raise instead; a genuine finish still leaves the loop on the progress
reaching 100.
Now that the expiry comparison actually caches a client, the session it
holds is shared, so tearing it down and replacing it needs the same care
the IMM handler already takes:
Dispose of an expired session with a logout instead of dropping the
reference, otherwise every refresh leaves an authenticated session
behind on an SMM that only has a handful of slots. That logout has to
tolerate a session the SMM has already reaped, hence the except.
Guard the login itself, so two coroutines arriving at an empty or
expired cache do not both log in and orphan one of the two sessions.
Stamp the vintage once the login round trips are done rather than
before, so a slow SMM cannot hand back a client that is already expired.
The old WebConnection.request() added a
'Content-Type: application/x-www-form-urlencoded' header to any POST
carrying a body, but grab_response_with_status() only sets a content
type for dict payloads, so the SMM login and every /data form POST now
go out as text/plain. This is not a fix for an observed failure: an SMM
running FPC variant 38 was measured accepting a text/plain login exactly
as readily as a urlencoded one. It restores the header the synchronous
code always sent and that the TSM and IMM handlers still set explicitly,
rather than relying on every SMM firmware level being equally lax about
what it will parse.
Also stop hard failing on responses the synchronous code discarded on
purpose. 'set=securityrollback:1' is only understood by newer SMM2
firmware. And /data/logout answers 401 once the session is gone, as
measured on that same SMM, so raising on a non-200 there turns a
completed hostname, domain or NTP operation into a spurious error.
The SMM hostname, domain, and NTP helpers looked synchronous even though their web transport is asynchronous. Removing awaits in the Lenovo OEM handler therefore returned unresolved coroutine work instead of completed settings results.
Convert the SMM settings and logout helpers to the asynchronous web interface, validate HTTP status responses, and await each operation from the OEM handler so callers only observe completed results.
The NextScale SMM path still used the removed http.client-style interface against the asynchronous WebConnection implementation. Login, configuration, diagnostic, and firmware operations consequently called unavailable methods or left request coroutines unresolved.
Make web-client creation asynchronous, migrate the affected requests to grab_response_with_status(), and await the cached client accessor. Correct the cache expiry comparison so fresh authenticated clients are reused and stale clients are renewed.
nodediscover's rescan poll and nodeconsole's screenshot refresh both slept
with time.sleep inside a coroutine. In nodeconsole --video that stops the
input handler and the VNC streaming tasks for the whole interval.
_connect_unix left the socket blocking, while _connect_tls sets a zero
timeout, so every loop.sock_recv and sock_sendall against the local
socket ran the blocking call inline and stalled the whole event loop.
nodeconsole --video showed this most clearly: a power action opens its
own session, so the tiles stopped refreshing and keystrokes went
unhandled until the BMC finished.
asyncio only enforces this in debug mode, where the client failed
outright with ValueError: the socket must be non-blocking. The
descriptor passing retries in asynctlvdata also assume a non-blocking
socket, since they wait for BlockingIOError.
simple_nodegroups_command awaited the async generators returned by read
and update, which raises
TypeError: 'async_generator' object can't be awaited
and printgroupattributes was left synchronous, iterating one of those
generators with plain for. Neither is reachable yet, since nodeattrib
only ever passes a noderange and nodegroupattrib still uses the
traditional client, but they are the paths nodegroupattrib will use once
it is ported.
When sendmsg() reports EAGAIN, _sendmsg rescheduled itself with
loop.add_reader(fd, _sendmsg, loop, fut, sock, fd)
which waits for the socket to become readable rather than writable, and
passes four of the six required arguments, so the callback raised
TypeError once it did fire. Wait for writability and pass the message
and descriptors through.
Also skip the work in _recvmsg if the future was cancelled while waiting
for data, as _sendmsg already does, so a cancelled read does not end in
InvalidStateError from set_result.
This module is imported by the server as well, so both paths are reached
by the daemon whenever a descriptor is passed over the local socket.
import_csv was left with several synchronous idioms:
- search_record is a coroutine function, but was called without await.
The returned coroutine is always truthy, so the rescan on incomplete
discovery data never happened, and iterating the result raised
TypeError: 'coroutine' object is not iterable
- the node creation loop iterated an async generator with plain for
- the per-node discovery assignment was forked off with os.fork() while
the event loop was running, and the child then built a fresh session
on the inherited selector
Assign discovery entries with asyncio.gather instead of a forked child,
which keeps the assignments concurrent and lets their exit codes
propagate. The forked child always ended in sys.exit(0), so its
accumulated errorcode was discarded.
register_endpoint and subscribe_discovery were left as plain functions
iterating the async client generators, so nodediscover register,
subscribe and unsubscribe all failed immediately with
TypeError: 'async_generator' object is not iterable
A node imported by a merge was assigned a free index and then had it
overwritten by the index carried in the backup, which may already belong
to a node in the target database. Keep the allocated index instead; a
full restore still honors the dumped index.
The -x/--exclude option drops matching node and node group attributes
from a dump, restore, or merge, so a backup can leave out dynamic state
such as deployment.state_last_updated or data that should not travel with it.
Patterns use shell-style wildcards, and a bare namespace such as net
excludes every attribute below it. The node "groups" and "id.index"
attributes and the node group "noderange" attribute are always retained
so that a restore can still reconstruct group membership and node index
assignments.
Unfortunately, the problem of urlmount's selinux context is left open.
urlmount starts before policy load, preventing transition.
However the policy blocks access urlmount needs when loaded.
os.path.join(ConfigManager._cfgdir, '/tenants/') discards the cfgdir
because the second component is absolute, so it resolves to /tenants/
rather than <cfgdir>/tenants.
Tenants are not used yet but let's not face this issue in the future.
Rename the format parameter to fmt to stop shadowing the builtin,
pass the already parsed key data dict directly to _restore_keys
instead of reserializing it, use yaml.safe_dump for symmetry with
the safe loader, and consolidate the five repeated per-format dump
blocks into one helper.
Diskless profiles had the updatestatus callback in onboot.sh commented
out because no suitable status existed: 'complete' clears
deployment.pendingprofile, which the PXE responder requires to answer
the next network boot of a diskless node.
Add a 'booted' status that records the pending profile as
deployment.profile while leaving pendingprofile armed and skipping
autolock, and enable the onboot.sh callback in all diskless profiles.
nodedeploy now shows 'pending: <profile> (booted)' for a running
stateless node.
PyYAML implements YAML 1.1 implicit typing, so hand edited values like
'yes', '52:54:00:12:34:56', or '2026-07-17' in a YAML dump would be
restored as bool, sexagesimal int, or date instead of strings (the
date additionally crashing the JSON re-serialization). Load with a
SafeLoader subclass that only implicitly types scalars the PyYAML
dumper would have quoted when emitting strings, keeping dump/restore
round trips faithful.
Restoring a YAML dump without --yaml (or vice versa) previously
reported 'Cannot restore without keys, this may be a redacted dump'.
Point at the actual format of the dump instead when the keys file
exists in the other format.
BUILDSRC is only set if imgutil build is run with --source, otherwise
the build host repos are used. If --source is not used, there is no
distribution symlink and add_local_repositores failed with 404.
Check if BUILDSRC is set and skip add_local_repositories if this is the
case.
`unsquashfs` can use multiple CPU cores during image extraction, significantly reducing boot time.
For example the whole boot time from PXE to shell on a 8-core VM, from approximately 45 seconds to 20 seconds.
This PR adds `squashfs-tools` as a dependency. Since the package is smaller than 1 MB, the additional image size is justified by the performance improvement.
For backward compatibility, the existing `cp`-based extraction method is used when `unsquashfs` is unavailable, such as with images built before this change.
The extraction logic has also been moved into the common functions and is now shared between EL9, EL10, and Ubuntu.
Both untethered `squashfs` images and `confluent_multisquash` images are supported.
Images must be rebuilt to include `unsquashfs` and benefit from the faster extraction path.
EL10 diskless networking uses dhcpcd and no longer dhclient and ipcalc.
Keep the existing dhclient and ipcalc requirements for older EL capture targets.
Centralize the SELinux chcon helper and use it for downloaded
systemd units, onboot hooks, and apiclient files across EL7 through EL10.
Include chcon in captured EL initramfs images.
Without this fix the onboot services failed to start on SELinux enabled
captured image.
If LLDP is uncooperative, maybe the mac was learned.
If no mac apparently learned, then we ping_everywhere in hopes of soliciting traffic, and then rescan the switches.
Then get all mac addresses, try to determine zone from generated mac, and print on success.
The following post install steps were missing on Ubuntu builds:
- Permission fixes
- sysctl load
- Service restart
- confluent PAM symlink to /etc/pam.d/sshd. It works without it on
Ubuntu because it falls back to other which allows login on Ubuntu,
but the behaviour should be the same on every OS. Furtheremore, an
admin might implement additional steps to sshd PAM and would like to
have this in Confluent, too
The CA database under /etc/confluent/tls/ca is typically created by a root context such as osdeploy initialize -t, but the confluent service runs as the owner of /etc/confluent, and openssl ca rewrites the database (index, serial) as the invoking user on every issuance. Certificate issuance through the running service (e.g. the /self/tlscert deployment API) then fails on the root-owned database until packaging happens to repair the ownership.
Run the CA creation (full CA and the currently unused simple CA variant) and the openssl ca invocation under normalize_uid, the convention already used when publishing the CA certificate. The issued certificate is staged through a temporary file since the destination may only be writable by the invoking user, e.g. the web server certificate paths.
Existing root-owned CA databases are repaired by packaging or manually via: chown -R --reference=/etc/confluent /etc/confluent/tls
Apply ruff's safe autofixes.
The changes are mechanical and behaviour-preserving. Issues fixed:
- F401: remove unused imports.
- F841: drop unused local variables and assignments, including discarded
await/return values, unused "except ... as e" bindings, and unused
"with ... as name" targets.
- F541: remove the f prefix from f-strings that contain no placeholders.
- E711: compare against None with "is"/"is not" instead of "=="/"!=".
- E712: test truthiness directly instead of comparing to True.
- E713: use "x not in y" instead of "not x in y".
- E714: use "is not" instead of "not ... is".
- E731: convert lambdas bound to a name into def statements.
- W291/W293: trim trailing whitespace on touched lines.
Before the async port the bmc var was used for a "host=bmc" parameter
which does not exist anymore.
Now add [] around IPv6 addresses with missing brackets but keep the scope zone like %eth0
as this is needed in current code.
The IPMI 1.5 session-challenge callback returned coroutine objects from error reporting and session activation instead of expressing an asynchronous callback contract directly. That made completion depend on the dispatch path noticing and awaiting the returned object.
Make the callback asynchronous and explicitly await both onlogon() and _activate_session(), ensuring failure notification and activation finish before callback dispatch continues.
Fix SC2045 (error): Iterating over ls output is fragile. Use globs.
Add exception SC2068 exception for confluent_client/confluent_env.sh as this is intended.
SC2068 (error): Double quote array expansions to avoid re-splitting elements.
Await the channel access raw command, the TSMA remote media settings
requests, and the XCC3 volume creation responses. These calls returned
or unpacked coroutine objects, breaking set_channel_access, TSMA
virtual media attach, and RAID volume creation at runtime.
Await OEM sensor, NTP, retry, and firmware-update operations that otherwise returned or discarded coroutine objects. Return the initialized energy manager for FAPM systems and update stale utility entry points to use asynchronous Command factories.
Fallback paths called OEM handler constructors directly, but these handlers are initialized through async create factories. This caused generic, TSMA, and SMM3 selection to fail with 'OEMHandler() takes no arguments'. Use and await the factories consistently.
Also await the asynchronous bmcinfo lookup and forward the TSMA pool argument correctly.
Allow arbitrary per-connection network settings, such as static routes
or a firewalld zone, to be specified as semicolon-delimited key=value
pairs on a net.*.extra_settings attribute. The keys are passed through
to the network backend of the deployed OS in its native syntax: nmcli
properties on NetworkManager systems, netplan YAML paths on netplan
systems, and ifcfg variables on wicked systems.
Neither tool detected when the attribute database produces conflicting
name/IP data, silently emitting the conflicts.
confluent2hosts now warns when the same hostname is generated for
multiple different addresses within one address family (dual-stack
IPv4+IPv6 pairs stay silent), which happens naturally in -a mode when a
node has several networks without distinct per-net hostnames.
confluent2dnsmasq now warns when generated reservations share a
hostname across different IPs, reserve the same IP more than once
(dnsmasq refuses to start on a duplicate dhcp-host IP), or reuse a MAC.
A type 0x02 record whose body cannot be decoded (e.g. a bogus EvM
revision) would raise and abort retrieval of the entire event log.
ipmitool and freeipmi print such entries with whatever fields they can
extract rather than failing; degrade to the same raw passthrough used
for undecodable reserved types instead of raising.
The Linux kernel ipmi panic logger stores panic strings in SEL records
of type 0xf0, with a chunk sequence number in byte 4 and up to 11
characters of the message in bytes 5-15. ipmitool and freeipmi both
recognize this convention; do the same rather than presenting such
records as opaque non-timestamped OEM data.
Some BMCs (e.g. AMI) log events using spec-reserved record types like
0x04 with a standard system event record layout. Previously only type
0x02 was decoded, leaving such entries with no usable data and tripping
the generic OEM handler. Follow ipmitool and treat all types below 0xc0
as standard format. If the body of a reserved type turns out not to
follow the standard layout, fall back to passing it through raw instead
of aborting the whole log retrieval.
Allow building EL and Ubuntu diskless images for a foreign architecture (e.g.
aarch64 on an x86_64 host) by leveraging qemu-user-static. The target
architecture is detected automatically from a -s source tree (for EL),
or may be requested explicitly with the new --arch option.
When the target differs from the host, dnf/debootstrap is invoked with
--forcearch/--arch and the presence of an enabled binfmt_misc handler
with the F (fix-binary) flag is verified up front, so emulation keeps working
inside the installroot chroot and a missing setup yields an actionable
error instead of a confusing exec failure mid-build.
The image architecture is recorded in confluentimg.buildinfo so that
pack selects the initramfs addons for the image architecture rather
than the build host, and exec of a foreign-arch root performs the same
binfmt check.
Adds `confluentdbutil showattrib <noderange> <attribute>...` to print the
node attribute.
In contrast to nodeattrib it can shows secrets and crypted values with -u flag.
It's server-side only: reads the config store and master key directly, never over
the API.
It's read-only and works without confluentd running.
Instead of returning a sessionless authdata if webauthn loaded and validation requested, raise a not found indicating missing webauthn module.
Fix str being passod to rsp_write for the 403 return.
Ensure the console and shell session logic triggers only for subordinates of nodes or noderange.
Fix str being passed to rsp.write for the successful console session
The block on user modification shorted out webauthn hooks.
Further, be more picky about the prefix before the username in webauthn registered credentials and validation.
While curl and agama download are happy with the CA bundle, zypper was not. Have pre.sh properly set up the CA certs.
Additionally, indicate the install subdirectory of the repository to agama via it's cmdline conf.
- deployment.lock: add missing 'unlocked' (messages.py's
InputDeploymentLock/DeploymentLock already accept and persist it).
- hardwaremanagement.method: correct stale "ipmi is used if not
specified" claim. Was changed to null in
c14165e2bd.
- snmp.privacyprotocol: document that unset is treated as 'des'
(snmputil.py explicitly groups None with 'des').
validvalues listed 'automatic'/'manual', but that was outdated.
Commit 454e1b8267 and cc70dcfa2b
implemented unset/'tofu' (trust-on-first-use, the default), 'manual', 'ca-only',
and an implicit 'ca' (any value that isn't otherwise handled falls
through to the standard CA-verification path, keying an already
pinned match without a full CA reverify).
The validvalues fix in ecaa75d967 rejected
these new values. Add new valid values with proper documentation.
Since it turns out we already incurred lxml dependency, use lxml etree instead of xml and mitigate risky xml features beyond blocking the word '!entity'
These commands ship in confluent_client/bin but had no .ronn man page, so
they did not appear in the generated documentation. Add man pages matching
the existing style, with synopsis and options taken from each command's
argument parser:
- confluent2ansible: export node inventory to an Ansible hosts file
- confluent2lxca: export nodes to a Lenovo XClarity Administrator bulk import CSV
- confluent2xcat: export nodes to an xCAT stanza definition (and optional macs.csv)
- dir2img: build a FAT image from a directory for nodemedia upload
- nodecertutil: manage BMC CA certificates and sign BMC certificates
- nodegrouprename: rename a node group
- noderename: rename nodes
Match autocons
Skip unless EFI x86_64.
If SPCR, trust it and use that unconditionally.
Otherwise, if only one can respond to TIOCMGET, then use that one.
If multiple can respond, but exactly one shows carrier, use that.
If a confluent user is a system user, do not allow them to
upload paths that their user would not have access to otherwise.
For non-system users, continue with the path based banned behavior.
Agama ignores the dracut cmdline.d and uses only /proc/cmdline
Good news is they have their own agama.conf and they append to it rather than rewrite it.
For one, ensure a unique machine-id. Broadly should be done, but critical for bonds to have unique mac
addressses in event of booting a captured image.
Fix ubuntu slow boot due to waiting forever for a network config. Have the transient network config bake into netplan for
a first pass before confignet comes along to do full configuration.
Clean up spurious error messages about grep true: and device busy on mounting overlay.
Copy Debian apt sources and keyrings into the target before apt runs. Run apt with DEBIAN_FRONTEND=noninteractive.
Return constrained child status to callers, and make pack fail clearly when no kernel was installed.
Do not continue waiting when session is broken.
Do not call _timedout without releasing the lock first.
Properly await on relog with bad rakp4
If an accounting issue pushes logontries too far without touching zero, then still recognize retries were exhausted.
Timeout on missing RAKP2 if retries were already exhausted.
Try to find various layers of network config and normalize.
Ultimately, after post subiquity will do some things and easiest to fix in firstboot instead.
IB VFs have the following "ip l" output:
4: ibp129s0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 2044 qdisc mq state UP mode DEFAULT group default qlen 1000
link/infiniband 00:00:00:8d:fe:80:00:00:00:00:00:00:60:5e:65:03:00:2c:43:c8 brd 00:ff:ff:ff:ff:12:40:1b:ff:ff:00:00:00:00:00:00:ff:ff:ff:ff
vf 0 link/infiniband 00:00:00:8d:fe:80:00:00:00:00:00:00:60:5e:65:03:00:2c:43:c8 brd 00:ff:ff:ff:ff:12:40:1b:ff:ff:00:00:00:00:00:00:ff:ff:ff:ff, spoof checking off, NODE_GUID 00:00:00:00:00:00:00:00, PORT_GUID 00:00:00:00:00:00:00:00, link-state enable, trust off, query_rss off
5: eno1: <NO-CARRIER,BROADCAST,MULTICAST,UP> mtu 1500 qdisc mq state DOWN mode DEFAULT group default qlen 1000
link/ether 30:56:0f:17:c0:b4 brd ff:ff:ff:ff:ff:ff
altname enp196s0
altname enx30560f17c0b4
This breaks the detection script because index 0 of the "vf 0 ..." line is not link/<type> anymore.
This commit improves the detection logic to fix this.
Newer mdadm versions may load arrays under names like
`/dev/md/<hostname>:raid` after reboot instead of `/dev/md/raid`. Detect both
naming schemes when waiting for the array device and use the resolved
path consistently when determining the underlying md device name.
Also clear existing md superblocks before wiping signatures to avoid
stale RAID metadata interfering with array creation or assembly.
Recent mdadm versions introduced an interactive prompt when creating RAID arrays
without an explicit bitmap configuration:
"To optimize recovery speed, it is recommended to enable write-intent bitmap,
do you want to enable it now? [y/N]?"
This behavior was introduced by upstream change:
https://github.com/md-raid-utilities/mdadm/commit/e97c4e18c847803016aa60066cb6e57c528d83a6
In non-interactive environments such as Anaconda, this prompt blocks installation
and causes RAID creation to hang.
Fix this by explicitly enabling the internal bitmap when creating RAID arrays.
Ensure all input order is preserved as it is processed.
Institute a 10ms delay after a key transmit. This is to prevent overwhelming QEMU console, and improve it for other consoles.
Tested by pasting a large volume of text and seeing that it was intact.
If an environment manually manages all materials,
provide -r to let
them request packing of those materials
without trying to generate any of the content.
Replace input handling with an async, this
permitts screen updates while doing commands.
Implement 'send break' (sysrq) and focus move.
Indicate not-yet-active focus with titlebar color.
Drop python3-eventlet from the Ubuntu Noble build Dockerfile. Clean up
remaining greenthread/greenlet terminology in comments across aiohmi
IPMI modules, consoleserver, macmap, and the IPMI plugin. Remove a
commented-out GreenPool reference in macmap.
Replace eventlet.greenpool with concurrent.futures.ThreadPoolExecutor
in the BMC discovery script, using as_completed() for proper exception
propagation and main-thread result aggregation to avoid race conditions.
Remove dead eventlet socket compatibility code (.fd attribute checks)
from the IPMI session layer, and clean up stale eventlet references
in comments across the codebase.
Closes: xcat2/confluent#197
For the manifest, only things that *could* be package updated matter.
So add a parameter to let get_hashes skip files that couldn't be related.
This speeds up packimage and rebase dramatically.
In confluent_osdeploy-aarch64.spec.tmpl, el10 was created as a symlink
to el8, so the subsequent `mv el10/initramfs/usr el10/initramfs/var`
inadvertently renamed el8's usr directory, leaving el8 and el9 (also
symlinked to el8) with hooks at var/lib/dracut/hooks/ instead of
usr/lib/dracut/hooks/. Rocky 9 dracut never found the hooks and dropped
to the emergency shell on all aarch64 nodes.
Use `cp -a el8 el10` as the x86_64 spec already does, so the rename
only affects the el10 copy.
Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Timothy Middelkoop <tmiddelkoop@internet2.edu>
This allows better redirection.
In python3, must write to sys.stdout.buffer. AttributeError for the unlikely event of a python2 based node being deployed.
Add support for a confluent=<host> kernel argument in init-premount: configure networking, flush interfaces, autodetect the primary NIC (saved to /tmp/autodetectnic), verify TLS connectivity to the provided server, call the whoami endpoint over TLS to obtain the node name, and write results to /custom-installation/confluent/confluent.info (with fallback to copernicus on failure).
Also update casper-bottom logic to handle IPv4 manager addresses: for IPv6 the manager is still bracketed and scoped interface resolved as before; for IPv4 the script now uses the previously detected NIC (/tmp/autodetectnic) or falls back to an `ip route get <mgr>` lookup to determine DEVICE. This ensures routed IPv4 deployments work correctly.
Study of the web interface showed that bulk requests are a key component.
This takes a sensor sweep from about 3 minutes to about 7 seconds in a test environment.
Become a notify type service. Reserve the ability to be normal should it come up, but with
the complications around fork, easiest
to punt on daemonize for now.
Create a generic redfish discovery and a MegaRAC specific
variant.
This should open the door for more generic common base redfish discovery
for vaguely compatible implementations. For now, MegaRAC only
overrides the default username and password (which is undefined
in the redfish spec).
Also, have SSDP recognize the variant, and tolerate odd nonsense
like SSDP replies coming from all manner of odd port numbers (no
way to make a sane firewall rule to capture that odd behavior,
but at application level we have a chance).
Add a mechanism to close a session the right way
in tlvdata
Fix confluentdbutil/configmanager to restore/dump db to directory
Move auth to asyncio away from eventlet
Fix some issues with httpapi, enable reading body via aiohttp
Fix health from ipmi plugin
Fix user creation across a collective.
Some subprocess calls were reworked to use asyncio friendly
variants.
Also, osdeploy initialize was checked, and reworked the ssh and tls
handling.
osdeploy import was also reworked to functional with async only.
Have util retain tasks that are 'fire and forget', to avoid
garbage collection trying to delete the background tasks.
Move some utilities explicitly over to asynclient/asynctlvdata that
had previously been reworked.
Implement terminal resize in new asyncssh backend.
Since a lot of the traditional client did not need async,
make life easier by just having them in parallel for now.
The server must use the async client, but the client applications can
stick with the somewhat more straightforward synchronous client.
One issue is that there are multiple networkmanager connections,
clean this up, though this seems not to be a functional issue.
However, sometimes the lldpad usage screws up network configuration,
disable the facility by forcibly disabling fcoe sincec that is what triggers lldpad.
wq
With asyncio, we must close the writer half of a pair
Also rework the get_next_msg to work better.
Still need to allow stop_following to interrupt get_next_msg
Make process_peer async, with socket connection being async,
and dependency.
Have getaddrinfo use the asyncio version.
Rework the snoop to be more effective.
Rework the scan to be less convoluted.
Offer a function in core to normalize plugin return.
A plugin might return an async generator, a traditional generator,
or might even return an awaitable wrapping a traditional generator.
Replace eventlet spawn with util spawn in discover core
Have node attribute update await the set_node_attributes appropriately
Restore 'as available' behavior to noderange over socket
Bring the httpapi to the point where the webui is able to start working,
notably bringing the asynchttp online with the websocket.
Fix a flaw in the async ipmi that would cause hangups.
Purge sockapi of remaining eventlet call
Extend asyncio into the credserver to finish out sockapi.
Have client and sockapi complete TLS connection including password checking
Fix confetty ability to 'create'.
Since we are rebasing to at least Python 3.6, and with
some extra ctypes wranging of the ssl context, we can likely
remove PyOpenSSL. Take first steps by removing it from 'sockapi'.
Have confluent executable become the 'top level' for eventlet, to allow
work on 'de-eventleting' on 'main.py'.
Rework tlvdata to deal with either a socket or a reader, writer tuple.
Using TLS with asyncio is easiest with the 'open_connection'
semantics, which force either a Protocol handler (callback based) or
dual streams. While protocol approach ends with a more socket-like
'transport', the 'protocol' half is a bit unwieldy. So reader and writer
streams instead.
Reap ssh-agent to avoid stale agents lying around.
Remove nuisance warnings about virbr0 when present.
Do a full runthrough as the confluent user to ssh to a node when user
requests with '-a', marking known_hosts and automation key issues.
ap.add_argument('-p', '--prepareonly', help='Prepare only, skip any interaction with a BMC associated with this deployment action', action='store_true')
ap.add_argument('-m', '--maxnodes', help='Specifiy a maximum nodes to be deployed')
ap.add_argument('-m', '--maxnodes', help='Specify a maximum nodes to be deployed')
ap.add_argument('-r', '--redeploy', help='Redeploy nodes with the current or pending profile', action='store_true')
ap.add_argument('noderange', help='Set of nodes to deploy')
ap.add_argument('profile', nargs='?', help='Profile name to deploy')
@@ -131,11 +133,6 @@ def main(args):
curr = nodeinfo[attr].get('value', '')
if curr and node not in profilebynode:
profilebynode[node] = curr
for lockinfo in c.read('/noderange/{0}/deployment/lock'.format(args.noderange)):
sys.stderr.write('The -r/--redeploy option cannot be used with a profile, it redeploys the current or pending profile\n')
return 1
@@ -178,7 +175,7 @@ def main(args):
if node not in databynode:
databynode[node] = {}
for attr in dbn[node]:
if attr in ('deployment.pendingprofile', 'deployment.apiarmed', 'deployment.stagedprofile', 'deployment.profile', 'deployment.state', 'deployment.state_detail'):
if attr in ('deployment.pendingprofile', 'deployment.apiarmed', 'deployment.stagedprofile', 'deployment.profile', 'deployment.state', 'deployment.state_detail', 'deployment.state_last_updated'):
Directory to pull installation from, typically a subdirectory of `/var/lib/confluent/distributions`. By default, the repositories for the build system are used. For Ubuntu, this is not supported; the build system repositories are always used.
* `--arch` <architecture>:
Target architecture to build for (`x86_64` or `aarch64`). For Ubuntu, the
Debian-style names `amd64` and `arm64` are accepted as aliases. Building
for a foreign architecture is supported on EL and Ubuntu build hosts and
requires a statically linked qemu-user emulator registered in binfmt_misc
with the F (fix-binary) flag; on Ubuntu this is provided by the
qemu-user-static package (qemu-user-binfmt as of Ubuntu 26.04). For EL,
the architecture is also detected automatically when building from a `-s`
source tree.
* `-y`, `--non-interactive`:
Avoid prompting for confirmation.
@@ -99,6 +109,15 @@ Build a diskless image from a distribution:
imgutil build -s alma-9.6-x86_64 /tmp/myimage
Build an aarch64 EL diskless image on an x86_64 host (the architecture is
detected from the source tree):
imgutil build -s alma-9.6-aarch64 /tmp/myimage
Build an aarch64 Ubuntu diskless image on an amd64 Ubuntu host:
@@ -35,8 +35,8 @@ devices. Generally every effort is made to passively detect devices as they
become available (as they boot or are plugged in), however sometimes an active
scan is the best approach to catch something that appears to be missing.
**nodedsicover clear** requests the server forget about a selection of
detected device. It takes the same arguments as **nodediscover list**.
**nodediscover clear** requests the server to forget about a selection of
detected devices. It takes the same arguments as **nodediscover list**.
**nodediscover subscribe** and **unsubscribe** instructs confluent to subscribe to or
unsubscribe from the designated switch running affluent with system discovery support.
@@ -75,6 +75,9 @@ the nodes.
## OPTIONS
* `-a`, `--aggressive`:
Perform more aggressive scanning and fingerprinting. This will probe as many addresses as it can find with ssh and https connections. By default more targeted measures are used
that will only tend to probe things that are easily detectable as devices implementing explicit mechanisms for detection and identification.
* `-m MODEL`, `--model=MODEL`:
Operate with nodes matching the specified model number
#If invoked as a command, use the arguments to actually run a function
(return 0 2>/dev/null) || $1 "${@:2}"
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.