2
0
mirror of https://github.com/xcat2/confluent.git synced 2026-09-28 16:20:54 +00:00

Compare commits

...

1236 Commits

Author SHA1 Message Date
Jarrod Johnson a1df0c6939 Users can request and revoke API keys
Proivde support for API keys.  Users can register by doing a POST like:
/confluent-api/sessions/current/apikey/create with body of '{"expiration":null}

API keys can be reset by doing a POST to /confluent-api/sessions/current/apikey/revokeall

A secret is provisioned per user to allow them to clear all keys, while remaining mostly stateless to avoid having to manage every key server side.
2026-09-28 11:34:17 -04:00
Jarrod Johnson dbe194d6fb Supersede generic behavior for SMM3
SMM3 puts everything on chassis, expose that.
2026-09-25 17:03:14 -04:00
Jarrod Johnson ae1afcabf8 Replace generic unsupported error with more specific
The generic error was unclear for a fairly common attempt to check power state.
2026-09-25 16:50:49 -04:00
Jarrod Johnson 03a3786f93 Fix webauthid keyname in user 2026-09-25 16:18:39 -04:00
Jarrod Johnson 2759404f4c Merge remote-tracking branch 'xcat/master' 2026-09-25 16:14:34 -04:00
Jarrod Johnson a31d1a9efa Stub out useless 'system' for SMMv3
SMMv3 is really a BMC without a 'system', reflect that by stubbing out the system information.
2026-09-25 16:13:40 -04:00
Jarrod Johnson 14dd48b026 Merge pull request #305 from Obihoernchen/fix/eureka-oem
Fix MEGWARE EUREKA chassis support
2026-09-25 13:28:11 -04:00
Markus Hilger 913029e19b Send If-Match only with the EUREKA discovery PATCH requests
The password change and the account update set If-Match: * on the
connection itself, and after a forced password change that connection
is handed back, so every later request carried it. Pass the header
with the PATCH alone. Also drop an unused asyncio import.
2026-09-25 19:22:38 +02:00
Markus Hilger 1bf8c01867 Report node health and unreachable nodes from EUREKA health
get_health looked only at each node's State, so an Enabled node whose
Health was Warning or Critical left the chassis reported as ok. It also
skipped a node whose resource could not be read, counting it as
healthy while an unreadable chassis was flagged. Report the node's
Health when it is not OK, and a node that cannot be read as
Unreachable.
2026-09-25 19:22:38 +02:00
Markus Hilger 38e94c73d6 Read EUREKA sensors with $expand when the firmware supports it
The sensor collection has 668 members, and reading each one on its own
takes about 92 seconds and rebuilds on every nodesensors call. Opt in
to $expand=. for that collection, checking once whether the firmware
really inlines the members, so firmware that ignores $expand keeps the
per-member reads.
2026-09-25 19:02:39 +02:00
Markus Hilger 859a8db873 Reseat EUREKA nodes with the Reseat reset type
ForceRestart only restarts the host, so the node BMC kept running and
a reseat did not recover a hung one. The EUREKA firmware provides a
Reseat reset type that removes all power from the slot.
2026-09-25 19:02:39 +02:00
Markus Hilger 7cc8972170 Skip CPU temperatures from unbooted EUREKA node BMCs
A node whose BMC is not reporting returns Reading 0 from its
temperature sensors, dragging averages down with meaningless
values. The ComputerSystem Oem data flags this via HasBMCMetrics;
skip such nodes, and keep collecting when the flag is absent.
2026-09-25 19:02:39 +02:00
Markus Hilger e59230c63a Pass rootinfo through to EUREKA OEM handler
The service root was already fetched by the caller; dropping it
forced OEMHandler.create to request /redfish/v1/ again.
2026-09-25 19:02:39 +02:00
Markus Hilger 4d952892b1 Report detail and severity from EUREKA health
get_health collected an issues list but returned it nowhere, and
flattened every problem to Warning. Return the findings as
badreadings using SensorReading, honor verbose, and map chassis
health through _healthmap so Critical is no longer downgraded.
2026-09-25 19:02:39 +02:00
Markus Hilger 5399ba25a0 Fix EUREKA CPU temperature readings
The sensor URL was built as BMC{N}CpuCPU{X}Temp instead of
BMC{N}CPU{X}Temp, so every request returned 404. The entries also
referenced const.SensorUnits, which does not exist, and used a
dict shape the only consumer, get_average_processor_temperature,
cannot read - it expects thermal-style dicts with ReadingCelsius.
2026-09-25 19:02:39 +02:00
Markus Hilger d8f4049115 Fix password change handling in EUREKA discovery
util.json_loads does not exist, so the PasswordChangeRequired flow
raised AttributeError on every 401, silently swallowed by the blanket
except. Use json.loads, which accepts the bytes body directly.
2026-09-25 19:02:39 +02:00
Jarrod Johnson 8057755b15 oMerge remote-tracking branch 'xcat/master' 2026-09-25 13:00:42 -04:00
Jarrod Johnson a657c7d64c Merge pull request #304 from Obihoernchen/fix/redfish-health-and-manager
Redfish: tolerate a missing sensor Health and a missing manager
2026-09-25 13:00:16 -04:00
Jarrod Johnson ec964e477c Change message on nodedeploy clear
When clearing the pending profile, indicate that action may have taken place.
2026-09-25 12:37:13 -04:00
Markus Hilger 40b1bf5e3b Report a missing redfish manager instead of raising TypeError
ManagedBy is optional, and when a system does not link a manager
get_bmcurl() answers None. Most callers passed that straight into a
web request, which failed with "Constructor parameter should be str"
from the url library. Route those callers through bmcinfo(), which now
raises UnsupportedFunctionality so confluent reports it plainly. The
event log falls through to its system and chassis fallback, and
list_media lists nothing, since neither needs a manager.
2026-09-25 17:55:05 +02:00
Jarrod Johnson 4d2bb02ea1 Fix cooltera pdu sensor read 2026-09-25 09:33:01 -04:00
Markus Hilger 5050580bb8 Stop redfish sensor health from raising on a missing Health
SensorReading computed health and states with .get() and then
overwrote both with a strict lookup, so a sensor whose Status lacks
Health, or carries a value outside the health map, raised KeyError
and aborted the whole sensor listing. Keep the tolerant lookup, and
drop the states of a reading that is OK, as the copy in command.py
already does.
2026-09-24 21:47:41 +02:00
Jarrod Johnson 9ad123a0e2 Merge pull request #303 from Obihoernchen/fix-aggressive-typo
Fix the spelling of nodediscover --aggressive
2026-09-24 11:25:11 -04:00
Markus Hilger c768e0df92 Fix the spelling of nodediscover --aggressive 2026-09-24 17:23:00 +02:00
Jarrod Johnson 5aecc483e0 Fixup malformed UUID in redfishbmc
Now that redfishbmc is wired up, have it fixup odd formatting of UUID.
2026-09-23 16:03:29 -04:00
Jarrod Johnson 6650a33150 Fix generic-ssh for a number of scenarios 2026-09-23 15:20:39 -04:00
Jarrod Johnson 23e6c70449 Wire up generic redfish device configuration
Have redfishbmc be able to attempt a redfish onboarding.

User must supply initial user and password, since we have no idea about the vendor choices at this level.
2026-09-23 15:08:55 -04:00
Jarrod Johnson c72821cdfa Restore genesis specific functions
Common functions depend on curl, but genesis cannot
2026-09-23 10:45:37 -04:00
Jarrod Johnson 69b0b8462f Fixes for ARM genesis
Do not autocons if we have ttyAMA0

Make sure to use the proper architecture for addons.
2026-09-23 10:41:28 -04:00
Jarrod Johnson 8fa19ada11 Have capitilazition of shim be consistent in genesis 2026-09-23 10:19:59 -04:00
Jarrod Johnson 397df5278f Add ARM genesis 2026-09-23 10:01:17 -04:00
Jarrod Johnson 5f321a6612 Update man page for nodeconfig. 2026-09-23 09:53:06 -04:00
Jarrod Johnson 9f5873fe19 Implement better indication of pending when available 2026-09-23 09:51:35 -04:00
Jarrod Johnson 6f8920d47e Fix incorrect name of starmap function 2026-09-22 16:16:52 -04:00
Jarrod Johnson 084cda09d1 Add iPXE spec file
Pull prebuilt from ipxe upstream
2026-09-22 15:42:46 -04:00
Jarrod Johnson 2f6b30f7eb Implement preference for iPXE shim
This implements a secureboot compatible flow, even for PXE.

Non secureboot environments suffer one useless transfer, but otherwise should be unaffected.
2026-09-22 15:06:28 -04:00
Jarrod Johnson 0128ee7df6 Properly handle lease renewal
When a fixed address assignment comes for renewal,
do not make it go to rebinding.
2026-09-22 15:02:33 -04:00
Jarrod Johnson fd7ee32204 Fix switch/port display in nodediscover 2026-09-21 15:31:25 -04:00
Jarrod Johnson dea8b71043 Fix uninitialized niccfg 2026-09-21 13:47:48 -04:00
Jarrod Johnson 57f6c18ebd Fix early handling in Lenovo OEM handling
If mgrinfo is needed, have logic local, to avoid
the mess during early initialization.
2026-09-21 10:16:48 -04:00
Jarrod Johnson ab5e47e6af Fix query to support mac or id as key 2026-09-18 16:59:39 -04:00
Jarrod Johnson fa2232ccff Add auto-register on IP assign 2026-09-18 16:40:10 -04:00
Jarrod Johnson 58e9a11f4c Add nvos onboarding 2026-09-18 16:37:35 -04:00
Jarrod Johnson 043ede7464 Implement lease_time opt in for non-boot dhcp offers 2026-09-18 15:31:12 -04:00
Jarrod Johnson 3dad193926 Fix erroneous cache retention
Do not refresh cache vintage an every access.

Also, give callers finer grained control over cache.
2026-09-18 13:58:05 -04:00
Jarrod Johnson 046ace6dba Improve redfish sensor reading performance
Move to the oem handler and leverage the expansion facility
to speed up supported redfish BMCs.
2026-09-18 13:35:15 -04:00
Jarrod Johnson 79ddd17826 Swallow async timeout in check_fish 2026-09-18 09:22:37 -04:00
Jarrod Johnson bc6b3b5b4c Fix inconsistency in expected confluent directory permissions 2026-09-18 09:02:28 -04:00
Jarrod Johnson e02fe647ac Add lease_time
In preparation for fixed-address DHCP offer support, provide attributes to enable it.
2026-09-17 09:42:52 -04:00
Jarrod Johnson 5dcb1a140c Avoid duplicate of name resolved ip in net config
It was possible for the name resolution to steal an address from another section.

Fix by having explicit IP addressing consume and then
purge any violaters after concurrent evaluation completes.
2026-09-17 08:25:58 -04:00
Jarrod Johnson d793c521fe Decrease max input of ssh connection
Specify connect and login timeouts
to avoid sessions being held open.

Also, in blocking_scan, wrap everything so that finally can ensure the scan is recognized as complete.
2026-09-16 15:41:58 -04:00
Jarrod Johnson 10c822fcd2 Remove disused hashlib 2026-09-16 12:33:09 -04:00
Jarrod Johnson 5803dc8ac5 Change to allow non-mac ids in discovery
Inventing mac addresses for routed discovery is misleading.

Change to a more accurate indication.
2026-09-16 12:30:53 -04:00
Jarrod Johnson 069d010944 Fixup ssh host key CA handling.
Use /etc/ssh/ssh_known_hosts more globally.

Actually store the TOFU behavior by backgrounding the awaitable.
2026-09-15 17:05:46 -04:00
Jarrod Johnson 6cabea665f Start wrapping asyncssh
For auto discovery, banner extraction, and interactive ssh, refactor to commen sshclient class.

This hooks the pubkeys.ssh in a manner compatible with confluent 3.x
2026-09-15 15:45:38 -04:00
Jarrod Johnson 57b74c3d14 Add '-t' for tabular CSV output from collate 2026-09-15 09:40:43 -04:00
Jarrod Johnson e26f94c5e0 Add nvos-switch deteciton
Also, wire up the new generic_eval to the 'register' for routed discovery enhancements.  Inventing a mac address for lack of a better idea.
2026-09-11 19:37:12 -04:00
Jarrod Johnson 23556f4698 Fix potential uninitialized variable in get_switchcreds 2026-09-11 19:34:21 -04:00
Jarrod Johnson e59e274bb4 Prepare to accept custom initial user/password
This allows user extension of systems with initial user/passwords that aren't known to the codebase.
2026-09-11 14:59:54 -04:00
Jarrod Johnson e6aefd39b8 Remove pyrefly directive
Evidently it goes from needing this to needing it not to be ther.
2026-09-11 14:38:27 -04:00
Jarrod Johnson 4f4eeefddb Have rescan -a find generic-ssh, generic-https, and generic-redfish
Going from generic-redfish as most specific, then generic-https, and generic-ssh being for ssh-only targets.

For generic-https and generic-ssh, the available ports are specified so code can know if https *also* has port 22 available.
2026-09-11 14:06:16 -04:00
Jarrod Johnson fc1790ed80 Fix failing to reference converted LVM name 2026-09-11 10:23:38 -04:00
Jarrod Johnson 208abe5084 Refine vnc recording
Some experimentation shows that 15 fps is more than enough for the vast majority of console activity, so cut back for reduced file size.

Also, the calculation for last frame was incorrect, tracking duration based on when exiting caught up to the queue.  Now add the end time explicitly and use that to reduce last frame lingering.

VP9 is still uncomfortably slow in a default setup, so stick with MP4V despite larger size, user may transcode if they want to make it smaller.
2026-09-11 10:21:15 -04:00
Jarrod Johnson a01f4ef877 Give indication of starting exit process
Particularly if recording the session, exit no longer is instant.  Provide feedback to user that the exit process is being worked.
2026-09-11 09:23:48 -04:00
Jarrod Johnson 65960df93e Begin work on aggressive discovery rescan
This begins to ping everywhere and start evaluating everything.
2026-09-10 16:00:46 -04:00
Jarrod Johnson c7d2e17545 Avoid failure on receiving malformed SOL payload packets
Some devices can emit invalid IPMI packets with missing payloads.  Do not be overly burdened by such packets.
2026-09-10 15:09:32 -04:00
Jarrod Johnson 6e7ea6c8b8 Add ability to record graphics console to video files
Use old mp4v so that it is quick, tolerating huge files
2026-09-10 12:00:42 -04:00
Jarrod Johnson f77f0b89a2 Fix broken invocation of get_my_addresses 2026-09-10 10:44:16 -04:00
Jarrod Johnson 9af74d9c44 Change Ctrl-Alt-Delete to use <Del> key 2026-09-10 09:40:22 -04:00
Jarrod Johnson af36db4824 Add a quick utility function to just use ipv6 multicast to get all possible local peers. 2026-09-09 15:18:35 -04:00
Jarrod Johnson 55b1786a68 Add functions to support broader network scanning
Refactor functions to let netutil depend on neighutil.

Add a suite of functions to take a mac and try to figure out some viable ip for the mac.

Provide a ping6 mainly to support a ping to ff02::1, and follow up with unicast UDP discard packets to trigger neighbor table population.

Try to use the neighbor table to figure out an ip and scope for a mac, preferring LLA.

If this fails, go for a try of converting the mac to lla arithmetically, then try all the nics to see which one seems to work.
2026-09-09 15:13:37 -04:00
Jarrod Johnson ba99818334 Add a ping6 function
This allows ping6 without forking a child.

Particularly useful for getting all link local addresses on a segment.
2026-09-09 11:28:42 -04:00
Jarrod Johnson 131a1b7350 Rename switch util to switchutil instead of netutil
netutil was a confusing name, since two modules had the same name.

Fix it by moving switch related stuff to 'switchutil'.
2026-09-09 10:33:41 -04:00
Jarrod Johnson f3f1e692ea Merge pull request #300 from Obihoernchen/fix/imgutil-chkstat
Actually run chkstat on the SUSE image
2026-09-09 08:32:29 -04:00
Jarrod Johnson 4f7228a8d4 Merge pull request #299 from Obihoernchen/fix/localectl-unset
Keep the defaults when localectl reports nothing set
2026-09-09 08:31:43 -04:00
Jarrod Johnson e5b3b8ce65 Merge pull request #298 from Obihoernchen/fix/suse15-setupssh-chmod
Fix the chmod typo in the SUSE 15 ssh setup
2026-09-09 08:31:02 -04:00
Jarrod Johnson ca6682c3cd Merge pull request #297 from Obihoernchen/sles16-support
Sles16 support
2026-09-09 08:30:38 -04:00
Jarrod Johnson b4b86ae909 Have auth protocol be selectable
Also, add AES256 to privacy choices.
2026-09-08 16:31:40 -04:00
Markus Hilger 399ce08680 Stop loading InfiniBand modules the SUSE 16 media lacks
This reverts commit be6c7a3794.

They do not exist in Leap/SLE 16 installer initrd.
The diskless hook keeps its copies. imgutil's installkernel instmods
all four, so they are in that image and the calls do work there.
2026-09-08 16:03:13 +02:00
Jarrod Johnson d12fc1bc8b Merge pull request #296 from Obihoernchen/fix/suse-packaging
Fix SUSE package dependencies and shebangs
2026-09-08 08:41:46 -04:00
Jarrod Johnson cc689fc9ad Merge pull request #295 from Obihoernchen/fix/pam-service
Link the confluent pam service to wherever sshd's config lives
2026-09-08 08:40:22 -04:00
Markus Hilger 583e2fa321 Import the signing key SLE 16 media ships
Building a SUSE 16 image from SLE media failed every package with
"key ID fec28eaf09d9ea69: NOKEY". Leap publishes that key as
gpg-pubkey-*.asc, which the existing glob picks up; SLE publishes the
same key only as repodata/repomd.xml.key, so nothing was imported.

15 media carries both spellings, so this changes nothing there.
2026-09-07 13:30:11 +02:00
Markus Hilger c66d7bcb7e Give the SUSE 16 initramfs libkmod
udev's kmod builtin dlopens libkmod, and dracut installs it from an
inst_libdir_file line in a module-setup.sh rather than by following
NEEDED. Which module carries that line moved: the dracut on SLE 16
media declares it only in 00systemd, which the diskless module set
never loads, so the image came up with no libkmod, udev autoloaded
nothing, and the guest reached the network scan with only loopback.
Leap's newer dracut also declares it in 95udev-rules, which base
depends on, which is why Leap was unaffected.
2026-09-07 13:30:11 +02:00
Markus Hilger d069328dd2 Put the sshd helpers in the SUSE 16 diskless initramfs
OpenSSH 10 splits each connection into sshd-session and that into
sshd-auth, so the initramfs sshd on 2222 could not serve a single
session. el10 added sshd-session for the same reason.
2026-09-07 07:07:03 +02:00
Markus Hilger d7a25a9933 Use dhcpcd for SUSE 16 diskless
SLES 16 ships no ISC dhclient and nothing provides dhcp-client, so the
image could not be built from SLES media at all. Leap carries dhcpcd
too, so one client covers both. el10 made the same move when RHEL
dropped dhclient.
2026-09-07 07:07:03 +02:00
Markus Hilger edf98e3177 Lock root on SUSE 16 when no password is set
The deploycfg carries the literal 'null', which went into the profile as
a password hash, so the installed root account reported a usable
password instead of a locked one. 15 substitutes '!' for this; the sed
delimiter has to move off '!' to carry it.
2026-09-07 06:14:54 +02:00
Markus Hilger 2e0323e333 Run the SUSE 16 firstboot service only once
15 got this from AutoYaST init-scripts. 16 enables its own unit, so it
has to disable it the way el8 and the diskless profiles do.
2026-09-07 06:14:54 +02:00
Markus Hilger 20563d3c3a Apply installedargs on SUSE 16
Every other profile feeds it to the bootloader; agama takes it as
bootloader.extraKernelParams.
2026-09-07 06:14:54 +02:00
Markus Hilger f742626226 Stop the SUSE 16 install when agama rejects the config
The default Etc/UTC is not in the tzdata list agama validates against,
so the load failed and the install fell through to agama defaults and
still reported completion. Normalize it and halt if the load fails.
2026-09-07 06:14:54 +02:00
Markus Hilger f3ea32bcd3 Say autoinstall where the SUSE 16 pre script means it
The comment came from 15, which rewrote an autoyast profile.
2026-09-07 06:14:54 +02:00
Markus Hilger 571076b211 Always netboot the SUSE 16 installer
The attached-media branch could never be taken: the label pattern built
from os-release is opensuse-leap-16.0 or sles-16.0, while the media is
labelled Install-Leap-16.0-x86_64 and Install-SUSE-SLE-16-x86_64. Had it
matched, it would have written an inst.repo to the dracut cmdline that
agama does not read, and skipped inst.script entirely.
2026-09-07 06:14:54 +02:00
Markus Hilger 6cd3805058 Drop the EL vendor names from the SUSE 16 hook
Neither Oracle nor Red Hat can be the first word of a SUSE PRETTY_NAME.
2026-09-07 06:14:54 +02:00
Markus Hilger a0cd0ac5f2 Verify the deploy server when registering a SUSE 16 node
The -k made the --capath on the same line pointless. el8 makes the
identical call without it.
2026-09-07 06:14:54 +02:00
Markus Hilger be6c7a3794 Load the InfiniBand modules for SUSE 16 installs
They stayed commented out when the hook was forked from el8, so an
IPoIB-only node had no path to the deploy server. The diskless hook
loads them already.
2026-09-07 06:14:54 +02:00
Markus Hilger 9e8fc106a1 Drop the anaconda leftovers from the SUSE 16 hook
agama takes inst.install_url and inst.script, set just above.
2026-09-07 06:14:54 +02:00
Markus Hilger 737c761be2 Set up hostbased ssh on installed SUSE 16 systems
prechroot.sh had setupssh.sh commented out and copied only the keys
inline, so nodes came up without shosts.equiv, without the CA in
ssh_known_hosts and without a setuid ssh-keysign.
2026-09-07 06:14:54 +02:00
Markus Hilger ace72d428a Ship the syncfiles template for SUSE 16
post.sh runs syncfileclient, but there was no template to edit.
2026-09-07 06:14:54 +02:00
Markus Hilger 6a77e91bf9 Give SUSE 16 the customization hooks SUSE 15 has
firstboot.custom and the empty pre.d, post.d, firstboot.d and ansible
directories were left out, so the documented drop-in points did not exist.
2026-09-07 06:14:54 +02:00
Markus Hilger fd9bdc00c3 Ship the post.custom stub for SUSE 16
post.sh already runs it, so every install logged a 404 for it.
2026-09-07 06:14:54 +02:00
Markus Hilger dfa80306b5 Install timezone data in SUSE 16 images
Without it onboot.sh cannot apply deployment.timezone.
2026-09-07 06:14:54 +02:00
Markus Hilger 9089d3700b Find ssh-keysign where SUSE 16 puts it
The permissions.local rule named /usr/lib/ssh, so keysign kept mode 0755
and hostbased auth failed with 'could not open any host key'.
2026-09-07 06:14:54 +02:00
Markus Hilger 5ac3880764 Actually run chkstat on the SUSE image
A trailing comma made args.cmd a tuple holding the argv list, so
fancy_chroot called startswith on a list and the child died before exec.
permissions.local was written but never applied, leaving ssh-keysign
0755 and hostbased auth inoperative on SUSE diskless images.
2026-09-07 05:54:16 +02:00
Markus Hilger 46920e8436 Keep the defaults when localectl reports nothing set
Newer systemd prints '(unset)' where it used to print 'n/a', so an
unset console keymap was handed to the installer verbatim. An unset
System Locale has no '=' either, and the whole line was being taken
as the locale.
2026-09-07 05:18:43 +02:00
Markus Hilger 2634c582e7 Fix the chmod typo in the SUSE 15 ssh setup
'chmd' left the mode of the installed root authorized_keys to whatever cp
gave it, and put a command not found in every install log.
2026-09-07 02:40:43 +02:00
Markus Hilger 939a46d8d1 Fix SUSE 16 diskless boot 2026-09-05 02:12:18 +02:00
Markus Hilger 9dd5802698 Fill in the spots the SUSE 16 work missed
- the aarch64 osdeploy spec builds the stateful suse16 addons but its
  diskless loop was never extended, so the aarch64 rpm shipped suse16
  without suse16-diskless and a packed image got a dangling addons.cpio.
- imgutil's builddeb keeps its own copy of the directory list that
  confluent_imgutil.spec.tmpl has, and it had learned about neither suse16
  nor el10.
- gather_bootloader gained a /usr/share/efi fallback for shim on both
  architectures but only for x86_64 on grub, so an aarch64 root found a
  shim and then died copying grub.

Finally, rewriting repos.d file by file rather than copying the tree meant
a subdirectory or a file that is not valid UTF-8 aborted the build before
any package was installed, which also regressed SUSE 15. Pass anything
that is not a plain text repo definition through untouched and restore the
modes on the ones that are rewritten.
2026-09-05 02:12:18 +02:00
Markus Hilger 5645fb5cfd Build SUSE 16 images with imgutil
SuseHandler refused anything but 15.x. What 16 needed beyond widening it:

- its repo urls are written in terms of ${releasever}, which zypper
  resolves from the target root's os-release, a file that does not exist
  yet when the first packages go in
- its repos name a zypper service backed by a package-provided directory
  the target root does not have, so zypper discarded every one of them
  as an orphan
- there is no mkinitrd to work out which kernel to build for, and bare
  dracut would build for the build host's running kernel
- the efi payloads moved out of /usr/lib64/efi, arping out of /usr/sbin,
  nsswitch.conf and protocols under /usr/etc, and the presets enable
  sshd already
- the module list predated virtio, so an image built for a KVM guest had
  no network at all, and dm-crypt could not allocate a transform for the
  encrypted image without the aes-xts modules
- urlmount still links libpthread, an empty stub since glibc 2.34 that
  nothing else in the initramfs pulls in
2026-09-05 02:11:53 +02:00
Markus Hilger 398211a6ed Support diskless boot on SUSE 16
Ported from suse15-diskless, with the differences 16 forces:

- dracut symlinks /lib/dracut/hooks to /var/lib/dracut/hooks, so a hook
  shipped at the old path replaces the symlink with a directory
- there is no netconfig or /etc/sysconfig/network to hand the running
  address to, so the initramfs writes a NetworkManager keyfile instead.
  Without it NetworkManager claims the interface on its own terms and the
  tethered root filesystem goes away with the old address, and confignet
  never gets the chance to refine anything
- the discovery loop retries without a delay, so a link that takes a
  moment to come up can exhaust all 30 tries before the first packet can
  go anywhere. Keep asking, as the el9 hook already does
2026-09-05 02:11:53 +02:00
Markus Hilger b5e24333ce Add the missing suse16 profile.yaml
Without it osdeploy import cannot generate a profile at all:
generate_stock_profiles opens profile.yaml unguarded, and initprofile.sh
seds the label into it. The label substitution also still looked for
'sle 15'.
2026-09-05 02:11:53 +02:00
Markus Hilger ea435d00a2 Fix SUSE package dependencies and shebangs 2026-09-05 01:05:15 +02:00
Markus Hilger 811e5fe03d Link the confluent pam service to wherever sshd's config lives
Linux-PAM reads vendor defaults from /usr/lib/pam.d and distributions are
migrating there package by package: systemd and polkit already ship into
it on both EL and Debian, and on SUSE 16 openssh has followed. There the
old code left a dangling /etc/pam.d/confluent and every pam authentication
against it failed.

The deb postinst carries the same logic, so fix it in step. ln -sf rather
than ln -s because -e is false for a dangling link, so the old code retried
the symlink and failed with 'File exists' instead of repairing it.
2026-09-05 01:03:42 +02:00
Jarrod Johnson 42ca73eb21 Merge pull request #293 from Obihoernchen/fix/drop-sysvinit
Drop the sysvinit script
2026-09-04 13:13:17 -04:00
Jarrod Johnson 6a63d95df6 Merge pull request #294 from Obihoernchen/fix/loop-debug-default
Leave asyncio debug mode off by default
2026-09-04 13:12:38 -04:00
Markus Hilger 2f00b2ff05 Leave asyncio debug mode off by default
set_debug(True) put per-callback overhead and slow callback logging into
every run.
Use PYTHONASYNCIODEBUG=1 or -X dev instead.
2026-09-04 19:04:42 +02:00
Jarrod Johnson 0d88d1ae28 Replace asyncio Locks with re-entrant behavior
If connect_to_leader calls itself, let it use it's own lock.
2026-09-04 12:43:18 -04:00
Markus Hilger 106905be38 Drop the sysvinit script 2026-09-04 18:01:35 +02:00
Jarrod Johnson 5025c904e8 Ensure keepalive are sent while following
keepalives would be postponed by incoming keepalives.

Fix this by tracking keepalive on transmit only, not on receive.
2026-09-04 12:01:01 -04:00
Jarrod Johnson 8fcb5aec9a Avoid exiting on tail failure
tail can fail in certain scenarios.  Switch to wait if that should occur.
2026-09-04 08:39:05 -04:00
Jarrod Johnson 2f49ab602d Fix incorrect attribute on asyncio task 2026-09-03 13:42:43 -04:00
Jarrod Johnson 5df20a92ac Merge pull request #292 from Obihoernchen/fix/sdr-init-caching
Do not cache an SDR that failed to build
2026-09-03 08:57:03 -04:00
Markus Hilger 497a7abffb Do not cache an SDR that failed to build
init_sdr assigned self._sdr before initialize() ran, so a failure left the
half built object in the cache. The next call saw a non-None _sdr and handed
back that partial repository rather than trying again.

The visible symptom is a first call raising and the second appearing to
succeed. The real cost is on a bmc where the read fails once: the client
keeps the incomplete sdr for the life of the session and every later sensor
lookup answers from it without complaint.
2026-09-03 02:15:22 +02:00
Jarrod Johnson 7afbbca691 Merge pull request #291 from Obihoernchen/fix/multi-nic-refusal
Say which interfaces a multi-homed BMC has, rather than crash
2026-09-02 10:29:29 -04:00
Jarrod Johnson d4c5fd850a Merge pull request #290 from Obihoernchen/fix/report-failures-not-crashes
Let a failed request report itself
2026-09-02 10:16:16 -04:00
Jarrod Johnson 26c69e6cc0 Merge pull request #289 from Obihoernchen/fix/health-optional-collections
Handle Redfish services that omit optional collections
2026-09-02 10:13:26 -04:00
Jarrod Johnson fb0ecb50f7 Merge pull request #288 from Obihoernchen/fix/parse-fractional-seconds
Read a fractional second as a fraction
2026-09-02 10:12:03 -04:00
Jarrod Johnson 626d9ba15f Merge pull request #287 from Obihoernchen/fix/ipmi-identify-refusal
Say that IPMI cannot read an identify state
2026-09-02 10:10:35 -04:00
Jarrod Johnson 42ac4d064b Merge pull request #286 from Obihoernchen/fix/cli-exit-codes
Set the exit code when a read fails
2026-09-02 10:09:17 -04:00
Jarrod Johnson 514dc32499 Merge pull request #285 from Obihoernchen/fix/crypt-without-stdlib
Run on a Python that has no crypt module
2026-09-02 10:08:14 -04:00
Markus Hilger 9012888cc0 Report a refused XCC web login instead of returning None
get_webclient falls off its end when /api/login answers anything but 200, so
it returned None. wc() passes that back, and thirty of the thirty-four call
sites use it unchecked, so a refused login arrived as "'NoneType' object has
no attribute 'grab_json_response'" from wherever it landed.

Raised where the failure is known, and with the status: 404 is firmware with
no web api, or a Redfish-only capture of one, while 401 is credentials it
will not take. Neither was distinguishable before.

The other four call sites are the inventory reads, and they always did check.
They answer partially when the web interface is out of reach, which is why
nodeinventory still says something useful. They ask through wc_if_available
now, so that tolerance is stated rather than resting on a None.
2026-09-01 23:52:49 +02:00
Markus Hilger 7a7bd758a4 Let a failed request report itself
Two places where the code that exists to explain a failure fails instead, and
the caller is shown the second failure rather than the first.

_do_web_request builds its message from the response body when that body is
not the JSON error document the spec asks for. An html 404 page is exactly
that, the body is bytes, and str + bytes raised TypeError. The status and the
body never reached anyone.

LenovoFirmwareConfig raised a bare Exception when python-lxml and
python-eficompressor are absent. Confluent has no handler for one, so a
missing dependency showed as "Unexpected Error" and hid a message that
already said what to install.

Neither changes what fails, only what the caller is told.
2026-09-01 23:52:49 +02:00
Markus Hilger 1d8dcae6b0 Refuse plainly when a BMC lists no interfaces at all
Same function, one line above. EthernetInterfaces is optional, and when it is
absent the None went into a request and raised TypeError from the url library
rather than saying what was missing.

Kept beside the ambiguous case because the two are one question asked twice.
2026-09-01 23:52:25 +02:00
Markus Hilger 6909d1b5df Say which interfaces a multi-homed BMC has, rather than crash
Dedicated plus shared is how BMCs are built, so several NICs is ordinary.
_get_bmc_nic_url only reaches the count when the address the session came in
on matches none of them, which is what a tunnel or a NAT does, and it then
raised the bare PyghmiException base class.

Confluent has no handler for the base class, so it fell through to the
generic one and showed "Unexpected Error" plus a traceback, for a machine
doing nothing wrong. UnsupportedFunctionality now, which the redfish plugin
already reports plainly and which stays inside PyghmiException.

The message said "does not have exactly one interface" without saying how
many, which ones, or what to do. Every caller takes a name, so it lists the
candidates, and the empty case reads differently from the ambiguous one.
2026-09-01 23:52:25 +02:00
Markus Hilger 26151c506f Read power and boot from a service that has no system
A Redfish service may publish no Systems collection, and power and cooling
equipment does exactly that. sysurl is then None, and get_power and
get_bootdev handed it straight to a request, raising TypeError from inside
the url library.

sysinfo already guarded the same field. Both now ask through _system_url and
get a refusal naming what is missing. DMTF publish three services of this
shape, which is why they sit commented out in inventory-dmtf.yaml.
2026-09-01 23:51:32 +02:00
Markus Hilger 5d5ad821f4 Read health from a service that publishes no PCIe
PCIeDevices and PCIeFunctions are optional, and get_health indexed both
without checking. _get_adp_urls in the same file already spells it
.get('PCIeDevices', []), so these two were the outliers.

Power and cooling equipment has no PCIe at all. The KeyError escaped the
health read and reached the user as "Unexpected Error" with a traceback
behind it. Reproduces offline against DMTF's public-rackmount1 mockup.
2026-09-01 23:51:32 +02:00
Markus Hilger 5d7fadf615 Set the exit code when a read fails
nodesensors and nodeconfig printed an error and exited 0, so
`nodesensors n1 && next-step` ran next-step after the read it guarded had
already failed.

nodesensors had three of these: the per-node branch never set the exit code,
the top-level branch beside it read `exitcode |= exitcode`, and a normal
return from main() fell off the end of the file. nodeconfig accumulates with
|= all through its read path except the last line, which assigned, so a
failed bmc configuration read was discarded by the system read after it.

The fixed branches now match how nodehealth and client.py spell the same
thing.
2026-09-01 23:51:10 +02:00
Markus Hilger 2e7ce63b51 Read a fractional second as a fraction
parse_time read the digits after the decimal point as whole milliseconds, so
'.5' became 5ms instead of 500ms. Only a three digit fraction came out right,
and a BMC may write either.

'.5', '.50' and '.500' are all 500ms now, '.125' is 125ms. Every other format
parse_time accepts is untouched.
2026-09-01 23:51:00 +02:00
Markus Hilger 492bfeb974 Say that IPMI cannot read an identify state
nodeidentify against an IPMI BMC printed the node name, nothing after it, and
exited 0. A script checking the exit code carries on with an empty value,
which is worse than being turned down.

IPMI can set the identify light and has no command to read it back, so aiohmi
has no get_identify. The empty state was a way of not saying so.

The comment above that branch called identify "read-only", which is the
opposite of the truth.
2026-09-01 23:51:00 +02:00
Jarrod Johnson 31d1d38ac6 Make sure that ls reads the actual size without being confsued by symbolic links 2026-09-01 16:35:56 -04:00
Jarrod Johnson 6d0e364f0e Fixes for media based image boot 2026-09-01 16:23:52 -04:00
Jarrod Johnson c8bc59472f Fix missing close quote 2026-09-01 10:41:50 -04:00
Jarrod Johnson 0b4123fcc3 Add keystrokes for nodeconsole -tv
This allows ctrl-alt and windows key.
2026-09-01 10:33:57 -04:00
Jarrod Johnson 6186d940b6 Move hostname setting to pre.sh in suse16
hostname command is not in the initramfs
2026-09-01 10:00:27 -04:00
Markus Hilger 0d2b23de2f Run on a Python that has no crypt module
crypt left the standard library in 3.13 and both imports of it here are
unconditional, so the server does not start on a 3.13 distro that ships no
crypt shim of its own.

legacycrypt and crypt_r both reach the same libcrypt call, tried in that
order because legacycrypt is pure ctypes while crypt_r wants a compiler.
Both were checked byte for byte against the stdlib for the $6$ salts used
here, and a hash written under the stdlib verifies under either, so stored
crypted.* attributes keep working.

Recommends rather than Requires, and only above 3.12: el9 and el10 still
ship the stdlib module, and some 3.13 distros package no candidate at all,
where a hard dependency would make the rpm uninstallable.
2026-09-01 06:03:16 +02:00
Jarrod Johnson 7ba1798848 Fix SUSE16 identity image deployment
First, give an additional grace period in case identity image is slow to appear compared to other disks.

Also, set hostname according to nodename.
2026-08-31 12:05:45 -04:00
Jarrod Johnson 3fdda3adac Merge remote-tracking branch 'lenovo/master' 2026-08-28 15:58:55 -04:00
Jarrod Johnson aec404e134 Normalize scan to async
Also, for lots of pending nodes, yield between nodes for responsiveness
2026-08-28 15:51:13 -04:00
Jarrod Johnson 306fa83fd7 Merge remote-tracking branch 'xcat/master' 2026-08-28 15:50:27 -04:00
Jarrod Johnson e9649dcb0a Normalize async def scan, also break up pending_nodes by yielding 2026-08-28 15:49:51 -04:00
Jarrod Johnson 054ca63600 Fix macmap offload startup concurrency
State of offloader was never checked after acquiring the lock.

Fix by checking with the lock held.

Also, neaten up by putting all the offload startup inside the function to start it up.
2026-08-28 13:23:16 -04:00
Jarrod Johnson 66488f3041 Merge pull request #284 from robertcaliman/fix-linux-7mm-regex
Update string for 7mm to the correct form and change 7mm identification
2026-08-28 09:51:15 -04:00
Robert Caliman 1e57488f0d Add changes for 7mm disk without a RAID volume 2026-08-28 14:30:32 +03:00
Robert Caliman 84cf69ac4f Update string for 7mm to the correct form 2026-08-28 12:44:55 +03:00
Jarrod Johnson 48280ee5d3 Enhance VROC member criteria
Try to propogate form factor of member disk to array.

Also, even if cannot detect m.2, assume a two-member vroc array of nvme is m.2.  Not guaranteed, but most likely.  This is to deal with lack of DMI information indicating the physical form factor.
2026-08-27 16:55:18 -04:00
Jarrod Johnson 2cf9162d6c Relocate onboot.d scripts in media based image 2026-08-27 16:08:59 -04:00
Jarrod Johnson 724818111c Merge pull request #283 from robertcaliman/rcaliman/add-more-m2-models
Add some more RAID models to the PRIORITY_MODELS
2026-08-27 15:49:45 -04:00
Jarrod Johnson b6d10e5e9b Continue to tweak the Media based image boot of a confluent diskless 2026-08-27 15:47:53 -04:00
Robert Caliman 252b644f34 Remove the 16i models 2026-08-27 22:45:08 +03:00
Jarrod Johnson 34e7a10cfc Refer to self instead of non-existant disk 2026-08-27 15:43:17 -04:00
Jarrod Johnson 116d4aeb04 Make m.2 form factor drives a high priority as a matter of course for OS install 2026-08-27 15:41:49 -04:00
Robert Caliman 9c51577652 Add some more RAID models to the PRIORITY_MODELS 2026-08-27 22:39:56 +03:00
Jarrod Johnson 0c89f041e1 Merge pull request #282 from robertcaliman/rcaliman/m2-esxi-fix
Change the getinstalldisk logic to account for M2 standalone disks
2026-08-27 15:30:50 -04:00
Robert Caliman bdfd1ddfe8 Change the getinstalldisk logic to account for M2 standalone disks 2026-08-27 21:27:46 +03:00
Jarrod Johnson 687b8e47e4 Rework ownership/permission staging and cleanup
Directories are left as 'boring' confluent directories to enable staging.

Then the ownership/permssions on directories are fixed up.

Then after completion, make sure ownership is back to boring before asking rmtree.
2026-08-27 11:46:33 -04:00
Jarrod Johnson 1da744fcab Draft media based images
Prepare for media based diskless images
2026-08-26 16:13:34 -04:00
Jarrod Johnson bbad91369f Tighten permissions on netplan files 2026-08-26 15:36:27 -04:00
Jarrod Johnson 6d1733817e Fix empty kernelargs handling 2026-08-25 15:04:47 -04:00
Jarrod Johnson fff6875d33 Limit host based key types used by ansible
By default, ansible prefers to try host based authentication, which is good.

But when it doesn't work, it tries every key attempt, which is normally fine.

However, SSH counts key attempts the same as passwords, so hardening that restirct password attempts are fouled before it can even get to try a public key.  Thus let host based only consume one attempt.
2026-08-25 10:59:29 -04:00
Jarrod Johnson 8a57216803 Block user from naming group/node the same thing
If a node and a group have the same name, things can get very confusing and hard to deal with.
2026-08-25 10:47:39 -04:00
Jarrod Johnson 7a0a2d0e48 Create a copy instead of rename of /etc/hosts
By copying, we leave the /etc/hosts with original ownership/permissions/etc.

Otherwise we create a new /etc/hosts, which is subject to new permissions and such.
2026-08-25 10:35:35 -04:00
Jarrod Johnson 3496f90fb7 Make sure a child can't except out of a forked child
If an exception were incurred in child in fork, the code could break out.
2026-08-25 09:05:04 -04:00
Jarrod Johnson aa9b7ea4ca Yield cooperatively during nodelist change 2026-08-24 16:52:13 -04:00
Jarrod Johnson 9d6c6495bf Fix SLES 16 install when no serial console
autoinstall is run without a tty, so detect that and use tty1 instead.

Also, clean up agama config by avoiding too many duplicated network entries.
2026-08-24 11:31:31 -04:00
Jarrod Johnson 2308ee57fe Try to add altname for connection name. 2026-08-21 17:27:36 -04:00
Jarrod Johnson 79d4261303 Make dir2img usage look a little nicer 2026-08-21 15:15:33 -04:00
Jarrod Johnson f1961400ec Correct type of maxnodes in some scenarios 2026-08-21 14:57:24 -04:00
Jarrod Johnson f5d7c271b5 Fix issue with non-systemd startup 2026-08-21 07:15:52 -04:00
Jarrod Johnson 8a2d5016f3 Rework the systemd notification.
Confluent didn't act 'healthy' toward systemd delaying restart needlessly.

Worse, in a collective it could never show as started if the quorum didn't come back.

Pull the startup to before collective init, and indicate watchdog liveness during that time.
2026-08-20 12:40:55 -04:00
Jarrod Johnson 138b776c9c Recognize SATA m.2 drives
SATA drives do not directly have a busaddr.

However, at some point the PCI bus comes up in the udev hierarchy as a KERNELS value.

If that matches a detected M.2 slot, then accept the storage as M.2.
2026-08-20 12:03:09 -04:00
Jarrod Johnson 23445b3a5e Rework vtbuffer management
Only start vtbufferd on first use.

Do not restart more frequently than once every 30 seconds.
2026-08-20 08:33:35 -04:00
Jarrod Johnson 9fe79ce08c Merge pull request #280 from Obihoernchen/fixadvanced
Fix advanced XCC BMC settings
2026-08-20 07:34:57 -04:00
Markus Hilger ef9afa7d72 Fix advanced settings
Disabled this by accident when doing the async fixes.
Upstream pyghmi passess this as well.
2026-08-20 00:41:34 +02:00
Jarrod Johnson 8b67b20e8a Remove stray print statement 2026-08-19 14:11:02 -04:00
Jarrod Johnson 58d36a2b14 When closing connection, don't raise issues about connection being closed in unexpected ways. 2026-08-19 14:01:40 -04:00
Jarrod Johnson cf14bda3b7 Wrap up SOL exception more neatly 2026-08-19 13:46:44 -04:00
Jarrod Johnson 3463bd24c1 Avoid spawning redundant macmap workers 2026-08-19 13:43:29 -04:00
Jarrod Johnson 42fd63763e Defer more asyncio Lock creation
Must be created after loop is running for python 3.9.
2026-08-19 13:21:32 -04:00
Jarrod Johnson 2bd931658c Defer creating locks until loop is running
In python 3.9, this pattern was causing issues.
2026-08-19 13:13:20 -04:00
Jarrod Johnson a9da8ab3d3 Fix httpapi getaddrinfo
Another area that failed to get the positional arguments converted to keyword.
2026-08-19 11:24:41 -04:00
Jarrod Johnson 193cced1a6 Remove disused socket imports 2026-08-19 11:21:42 -04:00
Jarrod Johnson a0d99a15ce Fix use of positional arguments in getadrrinfo
The new getaddrinfo needs this as a keyvalue pair.
2026-08-19 11:05:41 -04:00
Jarrod Johnson ec64a6d46c Change to asyncio based name lookup in various places 2026-08-19 10:28:55 -04:00
Jarrod Johnson 31647a52e1 Remove disused RLock and simplify code. Fix name lookup stalls on connection. 2026-08-19 09:56:28 -04:00
Jarrod Johnson 18e8aecf96 Merge pull request #275 from Obihoernchen/feature/eventlog-sources
Add nodeeventlog source filtering for redfish
2026-08-19 09:25:55 -04:00
Markus Hilger 57b3fae0ed Return the status nodeeventlog worked out
Every error branch set exitcode and the script then ran off the end, so
a node that could not be read, or a log source that matched nothing,
still exited 0.
2026-08-19 15:10:55 +02:00
Markus Hilger c13a434246 Say when a named log source answered with nothing
An ipmi node named no source at all, so -s dropped every event, and a
log with no entries can never be named by one.
2026-08-19 15:10:55 +02:00
Markus Hilger 9cdcbe4046 Look elsewhere when the manager publishes no log services
The early return sat before the fallback, so the one layout it was
written for was the one it could not reach.
2026-08-19 15:10:55 +02:00
Markus Hilger 43d0706d6a Drop trailing whitespace from the synopsis line 2026-08-19 15:10:55 +02:00
Markus Hilger 5caf0451fd Let nodeeventlog show one log rather than all of them
Every entry already says which log it came from, so -s narrows the output to
the ones asked for and "-s list" names what a node offers.  Some of what a
platform keeps is noise: an AMI MegaRAC's event log is a list of redfish
sessions being opened and closed, while its useful records are elsewhere.

A selection cannot be cleared, since the platforms offer no such thing and
clearing more than was asked for is not something to do quietly.
2026-08-19 15:10:55 +02:00
Markus Hilger caa857ead0 Gather every event log a platform keeps, wherever it keeps it
A read only ever looked at the manager's log services, and fell back to the
system's when the manager published none.  A platform that keeps an event log
in both places had the second one invisible: an AMI MegaRAC keeps power unit
and thermal events in a chassis log that nothing read, 103 records that no
command could reach.

Clearing deliberately does not follow.  It stays where it was, so a log that
only a read reaches is never destroyed by one, and clearing a platform that
keeps its only event log on the system still works.

The name test now ignores spacing, since a build that calls its post code log
"BIOS POST Code Log" was read as an event log and merged 2719 post codes in.
2026-08-19 15:10:55 +02:00
Jarrod Johnson 3fa1236d34 Merge pull request #278 from Obihoernchen/fix/configurable-endpoints
configurable sockapi path and ipmi port
2026-08-19 08:33:06 -04:00
Jarrod Johnson c87bdc5a40 Merge pull request #274 from Obihoernchen/openbmc-support
Various fixes and features for AMI MegaRAC and OpenBMC BMCs
2026-08-19 08:22:26 -04:00
Markus Hilger 9e811dd81c Leave it to OpenBMC to say which of its logs are not events
Reading an event log skipped every log service whose id or name said journal,
dump, post code, host logger or crash, on every implementation.  Those names
are bmcweb's: the AMI and Lenovo bmcs call theirs SEL, EventLog, AuditLog and
PlatformLog.  bmcweb does need the distinction, keeping its event log under the
system while a clear would destroy its dumps, so it gets a handler that names
the words and generic names none, reading whatever a platform publishes.
2026-08-18 22:58:48 +02:00
Markus Hilger 811d48ed42 Ask only MegaRAC for the parameters part it insists on
Generic added an UpdateParameters part to every multipart firmware push,
because the specification has one carried.  Only the AMI firmware was seen to
insist on it, so name it in that handler and let generic send what the caller
passed.
2026-08-18 22:58:48 +02:00
Jarrod Johnson 985aff2c1a Allow ipv4 address extraction on web open with fe80::
If using fe80::, ask for more viable global addresses automatically.
2026-08-18 16:44:47 -04:00
Markus Hilger eaa1a8bfc7 Let a console work through a forwarded ipmi port
A bmc behind a forward answers Activate Payload with the port it listens on
itself, and the advertised-port check refused that, so a console failed where
command traffic worked. The advertised port is never sent to, so the check now
applies only on the default port. A bmc on another port advertising a third one
is no longer refused outright, which nothing here could have served anyway.
2026-08-16 21:37:50 +02:00
Markus Hilger fd92209d4d Read a port off the manager address over ipmi too
The plugin connected to 623 whatever the address said, with a TODO in place of
the parsing. It now reads one as the redfish plugin does, minus the brackets,
which getaddrinfo rejects.

IpmiConsole keys its endpoint mapping on host and port so several bmcs behind
one address stay distinct, and unregisters only an entry it actually claimed.
2026-08-16 21:37:50 +02:00
Markus Hilger 5b132438f5 Let the api socket path be set rather than fixed
_unixdomainhandler hardcoded /var/run/confluent/api.sock in four places and
derived its directory from a fifth. It is now threaded through SockApi like
the other bind settings, defaulting to the same path. That lets a service run
on a temp socket as an ordinary user, which is what a test needs.
2026-08-16 21:37:50 +02:00
Jarrod Johnson ebab4af396 Merge pull request #277 from Obihoernchen/fix/ipmi-completion-code-checks
Act on the ipmi completion codes again
2026-08-15 10:28:20 -04:00
Markus Hilger 9228ef53de Ask for search access on a download directory too
A directory the caller can write but not enter is one they could not
have created the file in, and write access alone said they could.
2026-08-15 13:46:30 +02:00
Markus Hilger c5f831e589 Start the bmc reset grace when the bmc actually goes
The deadline was set when monitoring began, so a flash that kept
answering for longer than the grace period had already spent it by the
time the reboot it covers arrived, and reported a working update as a
failure.
2026-08-15 13:46:18 +02:00
Markus Hilger 3ace6af075 Reserve the sensor records of the lun they are read from
A device holds its sensor records per lun and scopes the reservation the
same way, so a token taken on lun 0 can be refused for lun 1 with 0xc5.
The retry then took the same wrong token again without end.
2026-08-15 13:45:49 +02:00
Markus Hilger 274b1b310d Record why identify writes the system and not the chassis
An SD665-N V3 refuses every IndicatorLED value on its chassis and takes
all of them on the system, so making the write follow the read breaks it.
2026-08-15 13:44:49 +02:00
Markus Hilger c9804368f2 Only read a user slot as empty when the bmc says it is
Any completion code counted as an absent slot, so a bmc that was busy
or still starting up quietly shortened the user list.
2026-08-15 13:44:49 +02:00
Markus Hilger 803c4f2ceb Skip a shared enclosure when reading the leds
get_identify already ignores a chassis several systems share, so that
one node does not report the enclosure's indicator as its own.
2026-08-15 13:44:49 +02:00
Markus Hilger be1304e560 Name the firmware categories beyond core, adapters and disks
Firmware for a supply or a fan matched no fragment and so was called
core, and nothing could ever answer for misc.

Only the collections that mean one thing are matched by url.  A Storage
resource is the subsystem, so its firmware is the controller rather than
a drive, and a Processor is a cpu as readily as an accelerator, so that
one is decided by asking the processor what it is.
2026-08-15 13:44:37 +02:00
Markus Hilger 8c884e2ece Keep unrelated firmware in core rather than nowhere
An entry naming no RelatedItem was dropped from every category as soon
as any other entry named one, so core lost the bmc and uefi versions.
2026-08-15 12:58:47 +02:00
Markus Hilger e69f82a6a2 Do not let a failed lookup pass for a failed delete
The check for whether the account went sat outside the try, so an error
reading it escaped instead of falling back to blanking the account.
2026-08-15 12:58:32 +02:00
Markus Hilger 24e8cd7e00 Check for a deleted account without the cache
The delete that failed left the account collection cached as it was, so
asking whether the account is gone could only ever answer no.
2026-08-15 12:58:16 +02:00
Markus Hilger 307a1020a2 Set one bmc contact rather than one per letter
A contact name arrives from the client as a string, and handing it to
set_location_information made a Contacts entry of every character.
2026-08-15 12:57:52 +02:00
Markus Hilger 78048d01a1 Answer a stop request while waiting for quorum
A member of a collective that cannot reach quorum stalls in the startup loop
until quorum returns, and there was nothing in that loop that noticed a
shutdown.  That was survivable while SIGTERM raised SystemExit out of the
signal handler, since that escaped the loop from wherever it happened to be.
Having the event loop deliver the signal instead leaves the stop event set with
nobody reading it until quorum is reached, so stopping the service waits out
the systemd timeout and ends in a kill.

Check the event in the loop condition, and wait on it rather than sleeping
through it, so the answer comes in milliseconds rather than whenever quorum
returns.  A service stopped at this point has served nothing yet, so it goes
straight to the same configuration flush the normal exit does.
2026-08-15 12:44:35 +02:00
Markus Hilger fd84d38bbd Read the lan config parameter through raw_command
pyghmi asks for this parameter with xraw_command and catches the completion
code for a bmc that does not have it, and folding aiohmi in renamed that call
to oldraw_command rather than raw_command, so the handler could no longer fire.
Answering the code out of the returned dictionary repaired the crash but kept
the call on the older contract, which is now the only one left in the tree.

Catch it again instead: raw_command puts the completion code on the exception
as ipmicode, and nothing here reads the payload of a reply that carries a code,
which is the one thing catching gives up.

No behaviour change, checked against the previous version over the same fake
session for a good reply, an empty one, 0x80 and 0xC9 with and without a stray
payload, four other completion codes, a timeout, a lost session and a reply
with no data at all: same return value, same exception type, text and ipmicode,
same bytes on the wire.
2026-08-15 05:08:45 +02:00
Markus Hilger 2becb424fc End the device sdr retries a bmc will not satisfy
_read_device_sdr_lun negotiates the read size down when the bmc answers 0xCA,
but the size > 5 guard leaves a size of 5 alone, so a bmc that will not serve
5 bytes at once was asked the same question for ever.  Give up once the
request cannot get any smaller, and once a header read would go under the 5
bytes the record length sits in, by falling through to the raise already
there.

The stale reservation retry could not end on its own either: it cleared the
id and left taking a new one to the top of the loop, which only reserves for
a partial read, so the very first request repeated unchanged.  Take one where
the code is handled.
2026-08-15 04:44:43 +02:00
Markus Hilger 7c68758761 Back off a fru read the bmc will not serve in one piece
Completion codes 201 and 202 mean the chunk asked for was too big, and the
check for them sat after a call that raises, so a bmc that cannot serve 224
bytes at once failed the fru read rather than being asked for less.

The retry could not terminate either: chunksize // 2 + 2 is 4 for a chunksize
of 4, so the chunksize == 3 guard was unreachable and a bmc that kept refusing
would have been asked for 4 bytes for ever.
2026-08-15 04:33:28 +02:00
Markus Hilger d5e9be5abb Skip an absent optional sensor again, and read the ipv6 answer
Completion code 203 on a sensor reading means the sensor is not present, which
is expected of an optional device, but the check for it sat after a call that
raises first, so one absent sensor ended the whole sensor sweep.

_supports_standard_ipv6 read rsp['code'] the same way, so it could only ever
answer True; a platform without the standard parameters raised instead.  A
completion code is that platform's answer, while a timeout or a lost session is
not and must not be cached as one.

raw_command's docstring still described itself as the other call it was renamed
from, which is how these checks came to be written against the wrong contract.
2026-08-15 04:33:28 +02:00
Markus Hilger a513bfc04f Read the sdr partial read codes from the exception
raw_command raises on any nonzero completion code, so the 0xCA and 0xC5 checks
in get_sdr could not run: a bmc that will not return a whole record in one go,
or whose reservation went stale, failed the sensor load outright instead of
being retried.  Read the code from the exception, which carries it.

The back off also had a fixed point, size // 2 + 2 being 3 for a size of 3, so
a bmc that kept refusing would have been asked the same question for ever.
Give up when the request cannot get any smaller.  A header read cannot go
under 5 bytes either, since that is where the record length sits, so give up
there rather than parse a reply too short to index.

The stale reservation retry could not end on its own either: it cleared the id
and left taking a new one to the top of the loop, which only reserves for a
partial read.  Take one where the code is handled.
2026-08-15 04:33:28 +02:00
Markus Hilger f1fddd89b9 Drop a firmware entry that answered with nothing
The labels are worked out across the whole inventory, since a platform may give
every entry the same Name, so an entry that came back empty would be asked for a
name it does not have and take the naming of the others with it.
2026-08-15 04:07:23 +02:00
Markus Hilger 9755c8b0a8 Say what is wrong with an unusable parameter file
A parameter file that is not json, or that holds something other than an
object, reached the update as a raw parser message or as a TypeError from the
handler that unpacked it.
2026-08-14 22:31:21 +02:00
Markus Hilger 264576bddd Check a download target that already exists on its own
A writable directory only says the user could have created a file there, and
/tmp lets anyone do that.  If the target exists, it has to be writable by the
requesting user too, or confluent would overwrite it as root.
2026-08-14 22:12:08 +02:00
Markus Hilger 93a6ee554b Tell a refused user slot apart from a session that went away
get_user_name documents that it answers None when reading a slot fails, but
raw_command raises before the check that would return it, so that branch has
never run and one refused slot aborted the whole user list.  Answer None where
the docstring says to, and let get_users drop the blanket except it grew to
work around it.

Only a completion code counts as the bmc answering about the slot.  A timeout
carries the fabricated 0xffff from the session layer and a lost session carries
no code at all, and swallowing either of those reports a list truncated at the
point of failure as a complete one.
2026-08-14 22:04:59 +02:00
Markus Hilger dda9b47a51 Ask a platform which firmware image types it takes
Some will not take an image without being told which kind it is, and the
only way to find out was to attempt an update and read the error, which
writes to the bmc before it gets that far.
2026-08-14 21:26:29 +02:00
Markus Hilger 055434a862 Only ask a second time when the bmc might answer differently
Retrying a refusal three times over nine seconds only delayed the fallback
meant for it.
2026-08-14 21:26:29 +02:00
Markus Hilger 2ba5119f44 Judge a bios link by the status it answered
A link that is not served need not answer with a redfish error, so the
message id could not decide.
2026-08-14 21:26:29 +02:00
Markus Hilger eb742c5701 Let a virtual media insert report why it failed
Any failure fell back to setting the properties, so the property set's
complaint replaced the real reason.
2026-08-14 21:26:29 +02:00
Markus Hilger dd41aed9cc Write the identify indicator where reading finds it
Reading prefers the chassis, writing looked only at the system.
2026-08-14 21:26:29 +02:00
Markus Hilger 43531c3a10 Name ipmi user link relations with a string
The uids are dict keys and went out as JSON numbers.
2026-08-14 21:26:29 +02:00
Markus Hilger 6576c25f77 Require a bmc to have gone away before an update counts as applied
A bmc that keeps answering while the task read fails is a fault, not the
update landing.
2026-08-14 21:26:29 +02:00
Markus Hilger fe3e98a426 Read device sensor records without raising on the retry codes
raw_command raises on 0xCA and 0xC5 before the partial read loop can act
on them.
2026-08-14 21:26:29 +02:00
Markus Hilger e9cd14c80a Route the location resource to an input handler
Nothing matched the path, so an update was rejected with 400 before any
plugin saw it.
2026-08-14 21:26:29 +02:00
Markus Hilger bf9aba73d3 Report no unit for a sensor that has no reading
The sdr carries unit fields for every sensor record, and they were read into
the reading whether or not the sensor has a number for them to describe.  A
discrete sensor reports which of its states are asserted and no value at all,
so a watchdog came back with units of "% ", and a discrete sensor on a full
record picks up a base unit the same way, reporting degrees celsius for a
sensor that never has a temperature.  A caller that shows the unit alongside
whatever value it was handed then prints a unit with nothing to apply it to,
which is what the client was taught to skip in 8c8c32cc.  The client is not
the only consumer of a reading, so answer the question in the library.

The record itself says which sensors those are, and the reading path already
works it out to decide whether to decode a number: a sensor has one only if
the numeric format says how to read it, or the format is unsigned and the
record either supports thresholds or is of reading type 1.  Ask that once,
where the units are assembled, and give the sensors that fail it no unit, so
the unit and the value cannot come to different conclusions about whether the
sensor has a reading at all.

A threshold or numeric sensor reports exactly what it did before, including
the ones whose unit is a percentage, a combination of two units, or nothing
because the record names no unit.  On the Lenovo XCC this was checked against
nothing changes: its 305 readable sensors decode identically, discrete ones
included, because they are all compact records naming no unit in the first
place.  The sensors that change are the ones the finding came from, which
name a unit on a record that has no number to put it on.
2026-08-14 21:26:29 +02:00
Markus Hilger ce870f4c30 Ask whether a download target can be written, not read
A path a caller wants confluent to save something into was being run through
the check meant for a file confluent is asked to read.  That check forks, drops
to the calling user and asks os.access for R_OK, which is false for every file
that does not exist yet, so nodesupport servicedata and save_licenses could
only be given a path that was already there.  Handed a name to create, they
refused, and refused in a way no caller was looking for, so the command printed
nothing and exited zero.

The previous commit worked around it by skipping the check for a download
target, which fixed the symptom by removing the guard rather than by asking the
right question.  Ask the right question instead: whether the user could have
created the file in that directory themselves.  A path that is already a
directory is a destination directory, anything else names the file, which is
the same rule the code that goes on to write the file follows.

So a caller can still only make confluent write where they could have written,
and this now also catches an unwritable destination at the point the request is
made rather than several layers further in.
2026-08-14 21:26:29 +02:00
Markus Hilger 6ce03f5081 Call a firmware entry something that identifies it
This bmc gives all three of its firmware entries the same Name, "Software
Inventory", and puts what they actually are in the description.  The first
entry took that name, and the two after it fell back to their ids, so
nodefirmware answered with "Software Inventory", "cpld_active" and
"d1dc9e4b" for what are the host, cpld and bmc images.

Decide the labels across the collection rather than one entry at a time,
so a name the platform repeats can be recognised as no name at all.  Where
that happens, use a description that does tell them apart, and the id when
even that is shared.  A platform whose names are already distinct keeps
exactly the names it had.

The labels are what a caller addresses an entry by, so this also turns
inventory/firmware/all/d1dc9e4b into inventory/firmware/all/bmc_image.
2026-08-14 21:26:29 +02:00
Markus Hilger 586b2f1d77 Report more of a processor than its model
Processor inventory carried a single field, the model, so a platform that
does not give one had a processor in the listing with nothing in it, and
the client, which skips empty values, showed no processor at all.  This
bmc names the manufacturer, the socket and the core and thread counts, and
gives no model.

Carry those, along with the speed, serial and part number where a platform
offers them, and treat a processor as missing only when the bmc says its
state is absent, rather than whenever it does not describe a state.
2026-08-14 21:26:29 +02:00
Markus Hilger 63c5f2dca5 Keep the device available bit out of the firmware version
The top bit of the major revision byte of Get Device ID says the device is
still initialising or taking a firmware update.  It was read as part of
the revision, so immediately after a bmc reset nodefirmware reported "BMC
Version: 131.11" for what is 3.11, and settled down only once the bit
cleared.

Mask it as sdr.py already does for the same byte, so the two agree about
the same field.
2026-08-14 21:26:29 +02:00
Markus Hilger 593dc75145 Do not set an indicator the platform does not have
Reading the identify state says plainly when a platform describes no
indicator, but writing it went ahead and patched IndicatorLED regardless.
This bmc has neither that property nor the boolean that replaced it, and
answered the write with an internal service error, which reached the user
as one and the log as a traceback.

Ask the same question the read asks.  With neither property present there
is nothing to write, so say so in the same words instead of finding out
from the bmc.
2026-08-14 21:26:29 +02:00
Markus Hilger d34cb35f98 Treat a bios link that is not served as no bios link
This bmc advertises a Bios resource on its system and answers 404 for it.
Confluent followed the link and passed the bmc's complaint on as an
unexpected error, so a nodeconfig read printed every bmc setting and then
ended with "The requested resource of type  named 'Bios' was not found",
and the system half of the configuration was a 500 saying the same.

There is already a good answer for a system that offers no bios settings,
and a link that is advertised and not served is the same thing as far as a
caller is concerned, so give it the same one.  The result is checked once
and remembered, including the negative, so this costs one request on the
first ask and nothing after.
2026-08-14 21:26:29 +02:00
Markus Hilger bd5d7ffb94 Address a redfish account by the id the bmc gave it
Redfish identifies an account by a string, and an implementation is free
to use the account name, which this one does.  The handler converted the
last element of the path to an integer, so every per user read, update and
delete answered "invalid literal for int() with base 10: 'root'" as an
unexpected error, with a traceback to match.  Confluent offered the id
itself, listing the account as "root", and then could not accept it back.

Take the element as given.  Everything below already compares ids as
strings, and the input parsing already keeps a non numeric uid, so only
this conversion stood in the way.  nodebmcpassword goes through exactly
this path, reading users/all for the id and then writing to that account,
so it could not work at all on such a bmc.

ipmi users really are numbered slots, so the conversion is right there and
stays, but say so when it fails rather than letting a ValueError surface
as an unexpected error.
2026-08-14 21:26:29 +02:00
Markus Hilger ba3edc2c02 Match sensor categories against modern redfish sensors
A caller asking for fans or energy got nothing from any bmc that serves
the Sensors collection.  Those sensors were filed under their redfish
reading type, Rotational for a fan, while the categories are named after
the ipmi sensor types the rest of the code uses, so nothing matched.
Power appeared to work only by coincidence, Power and Current happening to
be spelled the same in both vocabularies.

Translate the reading type as the sensor is mapped, so a sensor means the
same thing whether it came from the Sensors collection, from the older
Thermal and Power documents, or from ipmi.  On the bmc this was found on,
fans go from nothing to the 24 tachometers, and temperature and power
already agreed with what the same hardware reports over ipmi.

The fan controls stay out, and cannot be brought in.  Their reading type
is Percent, which is also what a battery state of health reports, and this
bmc fills in no PhysicalContext to tell them apart, so there is nothing to
classify them by that would not also drag in unrelated percentages.
2026-08-14 21:26:28 +02:00
Markus Hilger 881e035043 Stop the 6 bit packed name decoder looping forever
The loop decoded the first three bytes of a name and never consumed them,
so any sensor or fru name a bmc encodes as 6 bit packed ascii spins at
full speed, appending the same four characters until the process runs out
of memory.  Measured at about 10 MB a second, so a bmc using an encoding
the spec gives its own worked example of costs a pinned core and, before
long, the daemon.

Consume each group, and decode a trailing group of one or two bytes rather
than dropping it, since those carry a character each and the name would
otherwise come back short.

The arithmetic was already right, it was only never reached a second time.
Verified by encoding names per the packing and reading them back: exact for
every length except those leaving three characters in a three byte group,
where the byte count cannot say whether three or four were meant and a
trailing space is unavoidable.
2026-08-14 21:26:28 +02:00
Markus Hilger ee35fba8bf Read sensor data records from a bmc that has no repository
A bmc may keep its sensor data records on the sensor device instead of in
a repository, and this one does, so it had no sensors, no health and only
a partial inventory over ipmi.

The records themselves are identical, version 0x51 and the same types, so
everything that decodes them is reused as is.  Only the fetching differs:
a command of its own, a reservation of its own, and records held per lun
rather than in one place.  The luns to ask, and a change indicator to
cache on, come from Get Device SDR Info.

The fetch is written out rather than shared with the repository one.  The
loops are alike, but nothing available here has a repository to test
against, and the price of factoring them together is that a mistake would
land on every bmc that works today rather than only on those that do not
work at all.

Records are cached in memory and on disk exactly as repository records
are, keyed on the change indicator, and a device that offers no such
indicator is read afresh each time rather than cached wrongly.

Names are stripped of the nulls that pad a fixed width field.  This bmc
pads every name out to sixteen bytes, and a name carrying them cannot be
matched by a caller asking for a sensor by name.

On the bmc this was written for: 163 sensors and 9 frus, against the 163
the device reports it has.  The 54 temperatures and 36 fans it now reads
match what the same machine reports over redfish, to within the precision
each side gives.
2026-08-14 21:26:28 +02:00
Markus Hilger db22a3e41a Stop asking the bmc who it is on every oem lookup
The oem lookup answers whether it found a handler for the vendor, and that
answer was being stored as whether the lookup had been done at all.  On
anything the map does not name, which is every bmc that is not a Lenovo,
the flag stayed false and each oem_init issued another Get Device ID and
built another handler.

Almost everything goes through oem_init, so this is a round trip added to
almost every operation.  Where those calls are close together it is far
worse than that: reading the sensor data records asks for the event
constants once per record, so a run of 172 records fired 176 Get Device ID
commands back to back, which was enough to make the bmc stop answering and
the read fail with a timeout.  The same sequence now takes 3 commands.

Settling for the generic handler is an answer.  The device id cannot
change within a session, so asking again buys nothing, and the handler it
throws away each time is the one holding the sensor names it had cached.
2026-08-14 21:26:28 +02:00
Markus Hilger 8de6c56998 Say what is missing when sensor records cannot be read
A bmc that keeps its sensor data records on the sensor device rather than
in a repository was answered with a bare NotImplementedError.  With no
text of its own it reached the user as "Unexpected Error:
NotImplementedError" and was logged with a traceback, for a capability the
library had simply never implemented rather than anything having gone
wrong.

Say so instead, in all three places that give up: the two branches for a
bmc without an sdr repository, and the version check that only understands
records of version 0x51, which now names the version it was given.

This makes the inventory usable on such a bmc as a side effect.  The
system fru is gathered before the records are, and an unsupported
operation is tolerated where an unexpected error was not, so nodeinventory
answers with the board, chassis and product data instead of one line of
error.
2026-08-14 21:26:28 +02:00
Markus Hilger 233fc9ec5a Read the event log rather than whatever logs the bmc offers
The redfish event log was taken from every log service the manager
advertises, whatever those turned out to be.  On a bmc that keeps its
systemd journal there, nodeeventlog answered with a thousand lines of
kernel probe failures and daemon chatter, and the log the user asked for
was never read at all, because this implementation keeps it under the
system.  Clearing was worse: of the services it did find, the ones with a
clear action were the dumps, so a clear destroyed diagnostic data, left
the event log untouched, and reported success.

Judge a log service before reading it.  A service whose id or name says
journal, dump, post code, host logger or crash is not an event log, and
both reading and clearing skip it, so a clear can no longer take out
something that was never asked for.

If that leaves the manager with no event log at all, look under the
system, where such an implementation keeps it.  Only then: a bmc that has
one under the manager is served exactly as before, from the same requests,
so this cannot change what an implementation that already worked reports.

The list of services was also being extended in place, and it belongs to
whatever the url cache is holding, so an extra log added by an oem handler
accumulated on every call within the cache window.
2026-08-14 21:26:28 +02:00
Markus Hilger d77e71967a Say plainly when a platform has no alert destinations
Reading the alert destinations of a bmc that has none reported "Unknown
code 0x80 encountered", which is the fallback text for a completion code
the library has no name for.  0x80 on this parameter is not a failure, it
is the platform saying it does not have alert destinations, and the
redfish side of the same resource has said so in words for a while.

The lan parameter fetch already knew how to tell those apart, so build the
alert reads on it rather than on a raw command that raises on any non-zero
code, and raise UnsupportedFunctionality with something to read.  Both the
count and an individual destination are covered, so a platform that offers
one and not the other says the same thing instead of failing differently.

Splitting the completion code handling out of the parameter fetch is what
makes that reuse possible; the interpretation of the payload, and every
answer it gives, is unchanged.

The oem hook for the destination count was passing its byte through ord(),
which raises TypeError on the bytearray it is given.  No handler in tree
implements the hook, so it had never been called; hand it the integer.
2026-08-14 21:26:28 +02:00
Markus Hilger 7b9ad03aaf Read a lan parameter the bmc does not have without crashing
A bmc that does not implement a lan configuration parameter says so in the
completion code and answers with no data at all.  The helper reached
straight into the payload, so such a parameter raised IndexError, and with
it went the whole of nodeconfig over ipmi: the bmc group, the plain, the
detailed, the extended and the advanced reads all ended in "bytearray
index out of range".  It also took the attribute enumeration with it, so
the client then rejected names it should have accepted.

The guard that was there caught an exception carrying the completion code,
but oldraw_command reports the code in its response rather than raising on
it, so nothing was ever caught.  Read the code from the response instead:
parameter not supported and parameter out of range mean the platform does
not have it, and anything else is a real failure that should say what it
was rather than be mistaken for absence.

The address configuration method was looked up in a table with no regard
for whether it had been read at all, so a bmc that does not report it
would have traded the IndexError for a KeyError.  Answer None when it is
absent, as the address above it already does, and name the value when it
is present but unfamiliar.
2026-08-14 21:26:28 +02:00
Markus Hilger 564230cf7e Give an unreachable target an error a user can read
A console whose bmc had gone away reported "Unexpected error - None", and
the api answered 504 with an error of None.  The redfish plugin took the
text for an unreachable target from the strerror of the socket error it
caught, guarded by a hasattr that is always true: every OSError has a
strerror attribute, and it is None on most of the ones a bmc going away
produces, TimeoutError and gaierror among them.  Ask for the text the same
way as everywhere else instead, which also keeps the errno on the errors
that do carry one.

The same applies to an unreachable target raised with no message at all,
so use the same helper there, on both transports.

Underneath that, give the node error messages a default to fall back on
rather than carrying whatever they were handed.  Each subclass already had
one, in an __init__ that an explicit None went straight past; making it a
class attribute the base class applies means it holds however the message
was built, and removes five copies of the same constructor.

Also repair an affluent handler that put a closing parenthesis in the
wrong place, passing its error text to Queue.put_nowait as a second
argument.  Any OSError there other than "no route to host" raised
TypeError from inside the except clause instead of reporting anything.
2026-08-14 21:26:28 +02:00
Markus Hilger 9402df2bdd Stop a websocket console spinning once its peer is gone
The receive loop treated only WSMsgType.CLOSE as the end of a session, but
aiohttp reports a peer that has gone away as CLOSED, and it does so
immediately and for every subsequent call.  Everything that was not CLOSE
fell to an else branch that printed a line and went round again, so a
console whose bmc restarted became a full speed loop writing one line per
iteration: measured at 2.7 million iterations a second, and observed
filling 15 GB of log in a quarter of an hour while the daemon stopped
answering requests.

Treat every message that is not data as the end of the session, clear the
connected flag and report the disconnect once.  A session that ended any
other way than a clean close is recorded in the trace log, unbuffered so
that it survives a daemon that does not, rather than printed.

Both websocket console plugins carried the same loop.  While here, give
the openbmc one the parts tsmsol already had: text frames are data rather
than a surprise, and the client session is closed when the upgrade fails
and when the console does, instead of being leaked.
2026-08-14 21:26:28 +02:00
Markus Hilger 1b2f15c392 Answer a firmware category request over ipmi honestly
nodefirmware <node> disks reported the bmc version, and so did adapters and
misc. The generic handler takes a category and ignores it, and the ipmi plugin
does not filter either, so every category answered with the one entry ipmi can
report.

Apply the rule R13 established for a redfish inventory that does not categorise
itself: the bmc's own firmware is system firmware, so it answers for core and
for nothing else. Filtering in the handler that produces the entry leaves an oem
handler that does categorise its own firmware free to answer for more.
2026-08-14 21:26:28 +02:00
Markus Hilger 58d56426ed Tell a bmc without DCMI apart from a failed request
There is no ipmi command for a bmc hostname, so get_hostname falls back to the
DCMI management controller identifier. A bmc that does not implement the DCMI
group at all rejects that with "invalid command", which was handed to the caller
as if the request had been bad: nodeconfig <node> bmc read seven fields
correctly and then reported "Error: Invalid command", and the api answered 500
Unexpected error.

Route every DCMI request through one helper that turns "invalid command" and
"command disabled or unavailable" into UnsupportedFunctionality, so the
identifier, the asset tag and the hostname all report a platform that cannot do
this rather than a fault. Where the caller asked about a hostname, say that
rather than naming DCMI.

The whole group read still ends in one error line, because an operation a
platform cannot perform and one that failed are the same message to the client.
That is worth separating, but not here.
2026-08-14 21:26:28 +02:00
Markus Hilger d9d9fa3df7 Say which resource a transport does not implement
The unhandled tails of handle_request, handle_configuration and handle_alerts
raised a bare Exception('Not implemented'), so asking for a resource the
transport has no code for was reported as an unexpected error and logged with a
traceback. management_controller/location over ipmi is one such resource: R5
implemented it for redfish only.

Raise UnsupportedFunctionality naming the resource instead, which the plugins
already treat as its own case rather than a fault, and give decode_alert over
redfish the same treatment. Any resource added to the tree without an
implementation on one transport now reports that plainly.
2026-08-14 21:26:28 +02:00
Markus Hilger f6ba3802bc Report an error that carries no message of its own
nodereseat printed "Error: " and nothing else against a bmc that refused the
credentials. The redfish plugin reports it properly, but the message it emits
re-raises TargetEndpointBadCredentials with no arguments when a single node is
addressed, and the enclosure plugin renders that with str(e), which is empty.

Both hardwaremanagement plugins already had a helper for exactly this, one copy
each. Keep one in confluent.exceptions instead, teach it to fall back to the
description a confluent exception carries by class before falling back to the
exception name, and use it in the enclosure plugin too.

get_error_body had the mirror image of the same bug, joining the class
description and the message unconditionally and so answering "Bad Credentials -"
with a separator and nothing after it. The apierrorstr property beside it
already gets this right, so use it.
2026-08-14 21:26:28 +02:00
Markus Hilger 7724a18c43 Do not let a missing pid file break the exit callback
The exit callback opened the pid file unguarded, so when it was already gone the
atexit handler raised FileNotFoundError and python reported an exception ignored
in an atexit callback. The removal of the debug socket immediately above is
guarded, so this was an oversight rather than an intent. Verified by stopping
the service with the pid file deleted first.
2026-08-14 21:26:28 +02:00
Markus Hilger 3563da041b Stop the service without raising through the event loop
terminate() called sys.exit(0) from a signal handler while the asyncio loop was
running. The SystemExit escaped run_forever, and closing the loop afterwards
then failed with "Cannot close a running event loop", so every clean stop wrote
a cascade of tracebacks and left a pending task behind.

Ask the loop to stop instead: deliver the signals through add_signal_handler,
which is the signal safe route, and set an event the main coroutine waits on so
that it returns and asyncio can unwind itself. Flush configuration on the way
out, as the client requested shutdown has always done.

That client requested shutdown went the same way, calling sys.exit from inside a
request coroutine, and it is the route the systemd unit uses to stop the
service. It now asks for the same orderly stop through a hook the running
service registers, keeping the old behaviour when nothing is registered.

Measured on all three routes, with redfish and ipmi sessions to a bmc in flight:
under a third of a second and no tracebacks, where before each one wrote a
cascade.
2026-08-14 21:26:28 +02:00
Markus Hilger 077c169169 Use the csrf header name MegaRAC actually checks
The web session helper sent X-CSRF-Token. MegaRAC checks X-CSRFTOKEN, so the
login succeeded and then every request answered Invalid Authentication, which is
why this helper has never worked. Confirmed both ways against a bmc: the same
request answers 401 with the old spelling and 200 with the new one.
2026-08-14 21:26:28 +02:00
Jarrod Johnson e9e5fc2516 Only add [] if the address is ipv6-like 2026-08-14 14:36:50 -04:00
Markus Hilger 836ab7e896 Read and write the current redfish location indicator property
IndicatorLED is deprecated in redfish in favour of the boolean
LocationIndicatorActive, so a platform that only implements the newer property
reported no identify state and could not be told to light up. Read either one,
preferring the older where both appear, and write the boolean when that is what
the resource offers. The boolean has no way to express blinking, so a request to
blink lights it steadily.

The led resource shares the same reader, so it gains this as well.
2026-08-14 20:04:19 +02:00
Markus Hilger d966b75a87 Report an inventory filter that matched nothing
Asking for the mac addresses of a node whose inventory does not describe any,
which is every node reached over ipmi, printed absolutely nothing and exited
successfully, leaving no way to tell an empty answer from a broken command. Name
what was asked for instead. The exit code stays successful, since an inventory
that does not mention something is a valid answer rather than a failure.
2026-08-14 20:04:19 +02:00
Markus Hilger 58c5827c62 Say when a platform cannot report its ntp state
get_ntp_enabled returns None to mean the platform cannot tell us, and that went
to the caller as the literal text "None", which says nothing at all. Report it as
unsupported instead, which the tooling already renders plainly.
2026-08-14 20:04:19 +02:00
Markus Hilger d41417d326 Stop printing a sensor unit that has no reading
A discrete sensor reports no value, and the unit was appended regardless, so a
watchdog came out as "Watchdog:% " and an event log sensor as "SEL:". The unit
belongs to a reading, so only print it when there is one. On the platform this
was seen on the units field is itself meaningless for such a sensor, carrying a
percent sign and a trailing space from the sdr.
2026-08-14 20:04:19 +02:00
Markus Hilger d2ead735a5 Consult the manager document when picking a redfish oem handler
The lookup fell back to the generic handler whenever it was given the service
root, which is the early call during connection setup, so a bmc that names its
vendor only in the manager document was served by two different handlers on
one connection. Read the manager during that early call too.
2026-08-14 20:04:19 +02:00
Markus Hilger 4ca6ec365d Report failures instead of tracebacks and usage in the client tools
A stray trailing comma made the update detail a one element tuple, so a
firmware error printed as a python tuple. A missing status printed the whole
response dict. A failure that named no node was dropped entirely, which is how
a service data request that the server refused came out as silence and a
success exit code, and nodestorage, nodelicense and nodesupport exited zero
even when they had reported an error.

nodeconsole crashed decoding an absent screenshot, and again on the terminal
calls behind a pipe, where a log replay crashed too; refuse the terminal only
modes cleanly and dump the log when there is no terminal to replay into.
nodedefine raised a ValueError on an argument without an equals sign, and
firmware for a category the target does not describe printed usage as though
the question had been malformed.

On the server side the readability check was applied to the path a download is
saved to, so asking for service data or saved licences at a path that does not
exist yet failed claiming the destination was not readable.
2026-08-14 20:04:19 +02:00
Markus Hilger 6361fd6578 Fix user deletion, media insertion, firmware categories and reseat
Deleting a user gave up after one attempt and reported why the fallback of
blanking the name failed rather than why the delete did. MegaRAC reports a
timeout for a delete that a second ask completes, so retry, check whether the
account went away, and keep the original error.

Attaching media judged a device free by ConnectedVia, which describes how the
device is wired to the host rather than whether anything is in it. Every
device on this bmc reports a fixed value there, so nothing was ever selected
and the attach reported success having done nothing. Judge by whether an image
is loaded, fall back to the properties when an advertised insert action is not
served, and say so when no device would take the image.

The firmware category was passed to the library and ignored, so core,
adapters, disks and misc all returned the same full list. Classify by what
each entry is related to, and answer only for core when a platform says
nothing about where its firmware belongs.

reseat_bay reached for a hardcoded Nvidia action on Chassis_0, so it failed
with a not found for that url instead of saying reseat is unsupported.
2026-08-14 20:04:19 +02:00
Markus Hilger b30ee29ff6 Implement the redfish resources that called missing methods
Five resources called methods that do not exist on the redfish client, so
each answered with an internal error naming the missing attribute: the leds,
the management controller identifier, the domain name, the remote kvm licence
and the alert destinations. Implement the first four from the manager network
protocol, the graphical console and the chassis indicator.

Alert destinations stay unimplemented on purpose. Redfish describes where to
send events with EventService subscriptions, which is a different model from
the numbered PET destinations this resource was built around, so say so and
drop the code that could never run.

The location resource fetched its data and discarded it, so a read produced
no output whatsoever.
2026-08-14 20:04:19 +02:00
Markus Hilger c1545cdf3f Drop the redfish url cache when something is written
Reads are cached for thirty seconds and a write did not invalidate anything,
so setting a boot device or the identify state and then reading it back
reported the value from before the write for up to half a minute. A write can
change documents other than the one written, an action url not being the
resource it acts on, so drop the cache rather than one entry.
2026-08-14 20:04:19 +02:00
Markus Hilger 8649b22605 Give every unsupported operation a message to report
A bare UnsupportedFunctionality() left the user with an error containing no
text at all, or with no output and a success exit code, so asking a platform
for something it does not implement looked like nothing had happened.

Name what is unsupported at each raise, treat it as its own case in the
plugins so it reads as a limitation rather than an unexpected error and does
not log a traceback, and fall back to naming the exception when an exception
still arrives with nothing to say. The generic redfish
get_extended_bmc_configuration was also declared without async while the
caller awaits it.
2026-08-14 20:04:19 +02:00
Markus Hilger b0ae4f201d Make redfish firmware updates work on AMI MegaRAC
Three things stopped a redfish firmware update on MegaRAC. The AMI handler
opened by asking the bmc to preserve fourteen named settings, and a build that
knows a different set rejects the whole request, which aborted the update
before anything was uploaded; send only the keys the bmc advertises. The
multipart push carried the image alone, and the specification has it carry an
UpdateParameters part too, which this firmware enforces. AMI also wants an
OemParameters part naming the kind of image, and nothing was supplying one.

The kind of image is asked for rather than worked out from the file. The
extension is vendor habit rather than format, and the leading bytes answer just
as confidently about an image they have never seen, while being wrong means a
bmc flashed with a bios image. So nodefirmware takes --type, it travels as far
as the handler that wants it, and where the bmc publishes the types it accepts,
an unknown one is refused with the list, as is asking with none. A platform
that reads the kind of firmware out of the image itself refuses the option
rather than dropping it, so nobody aims an update somewhere they did not mean
to. A parameter file still wins, since it can carry more than the image type.

Updating the bmc takes the bmc, and the task being watched, away for minutes.
That is the update working rather than the monitoring failing, so wait a
bounded while for it to answer again instead of reporting a successful flash
as an error.
2026-08-14 20:04:19 +02:00
Jarrod Johnson 8d618b354e dhcpcd is a function in the profile, source the profile to have it available in the script 2026-08-14 13:30:36 -04:00
Jarrod Johnson 357d205ee9 Correct the variable name for errmsg 2026-08-14 10:42:18 -04:00
Jarrod Johnson cf87fd2127 Fix network issue handling issue while managing network configuration. 2026-08-14 10:40:08 -04:00
Jarrod Johnson 969bbe009e Merge pull request #273 from Obihoernchen/ipmi-session
Fix ipmi session sharing
2026-08-14 09:41:15 -04:00
Markus Hilger 0b7f6b1395 Tidy three loose ends around sharing a session
Closing a console gives up its claim on the session, and that talks to the
bmc, so let it fail the same way the console's own deactivate is already
allowed to.

kg was left as the caller passed it in both the register key and the reuse
check, so the mismatch fixed for the name and password still applied to it.

The count for a new socket was taken before binding it and before the io
task was known to be up.  Take it last, once nothing is left that can still
fail.
2026-08-14 14:10:38 +02:00
Markus Hilger 8aecc6959a Do not let a logout come back round into its own notification
Telling a keepalive that the session is gone can end up back in logout,
because reporting it is how a console gives up its claim.  The inner pass
finished by clearing the register of keepalives while the outer was still
walking it, so the next entry was looked up on None.  Two entries is what an
XCC has, the console's and the oem handler's.

Take the callbacks and give up the register before notifying anyone.
2026-08-14 14:10:38 +02:00
Markus Hilger 5aec0e69f5 Recognise a session that is already open to the same bmc
The check for sharing a live session compared the credentials a session
keeps encoded against the strings every caller passes, so it never matched
and each caller built another session beside the one it could not see.
Normalise both sides.  Three commands to one bmc went from three sessions on
three sockets to one, and from five sessions open on the bmc to three.

Not from the port to asyncio: upstream compares the same two things the same
way.
2026-08-14 14:10:38 +02:00
Markus Hilger 9a8fe206a4 Let a session serve several callers without one closing it
One session is now routinely handed to a console and a command at once, and
logout closed it for both, leaving whoever was left holding one that
answered as though it had been lost.  Count the holders and give up a claim
instead, unless the session is no longer usable, which logout is told by
sessionok.

A console had no way to give a claim back: close deactivated its sol payload
and left the session alone, which was right when closing meant closing it
for everybody and is a leak now.  Both of its exits release it.
2026-08-14 14:10:38 +02:00
Markus Hilger ed21c7634a Drop a guard that never ran and would not have worked
__init__ opened by checking for an initialized attribute, meaning it had
been handed a session someone else was establishing.  That attribute is only
ever set further down in the same method, so the check cannot be true, and
the port lost the return that made it work upstream.  Waiting for someone
else's login is done in __new__ now, so this is dead code claiming to
protect something.
2026-08-14 14:10:38 +02:00
Markus Hilger 9408c56639 Give back a socket pool count once, not twice
logout decremented the count of sessions on a socket twice over, and
_mark_broken again for the case logout had not, so it went negative and kept
falling. _assignsocket picks the least used socket and refuses one at
MAX_BMCS_PER_SOCKET, and both of those read that number.
2026-08-14 14:10:38 +02:00
Markus Hilger ef608005cf Register an ipmi session before establishing it, not after
initting_sessions exists so a caller can share a session already on its way,
but the entry was added only once the login had finished, leaving the login
itself uncovered.  Two callers asking at once each built a session, both on
the socket the other had not claimed yet, and replies route by bmc address
and local port, so only the last to transmit was ever answered.  That is the
console session failing about one attempt in three.

Register before the login and remove the entry in a finally.  The two old
removals keyed on the encoded credentials while the register is keyed on the
caller's strings, so they never matched and an entry outlived its session.
A session handed over mid login is now waited for rather than returned as
one that answers as though it had been lost.
2026-08-14 14:10:38 +02:00
Markus Hilger c7b6147e74 Report why an ipmi session could not be established
A session that failed raised with no message at all, so a failed console
read "IpmiException: None".  Record the reason wherever a session is marked
broken.
2026-08-14 14:07:49 +02:00
Jarrod Johnson 3e2c03ed9d Correct imgutil argument handling 2026-08-14 08:02:49 -04:00
Jarrod Johnson 8fe87bf0eb Fix getfetchable hidden argument 2026-08-13 16:34:22 -04:00
Jarrod Johnson f7815bd7cf Improve tab completion of fetch 2026-08-13 16:23:09 -04:00
Jarrod Johnson 2660842b9a Add signature validation when possible to image fetch 2026-08-13 16:08:14 -04:00
Jarrod Johnson 9be1136ba4 Normalize destination directory in osdeploy fetch 2026-08-13 15:27:11 -04:00
Jarrod Johnson b84cf35d79 Show progress with osdeploy fetch 2026-08-13 14:52:03 -04:00
Jarrod Johnson 3aa5222fef Add packaging and some tab completion for osdeploy fetch 2026-08-13 14:41:06 -04:00
Jarrod Johnson a2a1f76ff2 Attempt a new 'osdeploy fetch' to help with downloading some of the ISOs. 2026-08-13 14:25:04 -04:00
Jarrod Johnson 4a5e20735d Normalize scratchdir globally
There remain issues where relative path can screw up the transient mounts.  Normalize it to be consistent using absolute path every time.
2026-08-13 13:30:12 -04:00
Markus Hilger 6e95399528 Repair the asynchronous contracts around ipmi users and extended config
handle_users iterated get_users with async for while the same call is awaited
a few lines below, so listing the users collection, and creating a user,
raised a TypeError about a coroutine having no __aiter__. list_inventory in
the redfish plugin had the mirror of it, awaiting an async generator.

get_extended_bmc_configuration is called with hideadvanced but the ipmi chain
never accepted it, so the extra and extra_advanced resources raised a
TypeError; thread the argument through instead.

A user slot the bmc refuses to describe no longer takes the whole user list
with it: one MegaRAC slot answered Invalid data field for good after an
account was deleted, which broke every user operation.
2026-08-13 18:09:18 +02:00
Markus Hilger 6002b14f52 Send If-Match with the redfish writes that lacked it
MegaRAC refuses a PATCH with no If-Match header, so setting the bmc hostname,
ntp, the bmc network configuration, a location and ejecting media all failed
with a precondition error. set_identify already passed etag=*; do the same at
the call sites that did not, including the firmware push busy flag.
2026-08-13 18:09:08 +02:00
Markus Hilger 6d36e803b1 Fix reading the identify state and the description over redfish
get_identify indexed the sysinfo method object rather than awaiting it, so
reading the identify state raised a KeyError naming a bound method on every
redfish system. Read the indicator from the chassis that owns the physical
led, since some implementations leave the copy on the system stale, and say
plainly when a platform does not report one.

The generic get_description took no fishclient while every other handler and
the caller pass one, so the description resource raised a TypeError on any
non Lenovo bmc.
2026-08-13 18:07:21 +02:00
Jarrod Johnson d2d081d94f Fix SSDP ignoring packets unless rapidly following another
The asyncio rework mistakenly follows up a long waiting acquire with blanking and starting over.

Now use 'True' as a sentinel value to trigger a fetch, otherwise, assume srp is ripe for processing.

Immediately discard srp after unpack and replace with sentinal value.
2026-08-13 09:22:51 -04:00
Jarrod Johnson 8f5d68ba3c Remove extraneous output from the tpm pcr bank identification 2026-08-12 15:48:09 -04:00
Jarrod Johnson 8ec8ce5289 Fix pcr extend
Using the subshell prevented variable from being set.

No longer use a subshell.
2026-08-12 15:36:49 -04:00
Jarrod Johnson d175d06b8d Change extend to change extend by appropriate hash size. 2026-08-12 14:48:57 -04:00
Jarrod Johnson b37af50c3f Use other TPM PCR banks
Some TPMs are configured to use other pcr banks.

For now, prefer the sha256 for compatibility.
2026-08-12 13:51:29 -04:00
Jarrod Johnson ae664717e4 Fix Ubuntu encrypted OS volume setup
Actually install the dependencies needed, and correct path to detect need to re-seal when PCRs requested
2026-08-11 11:46:37 -04:00
Jarrod Johnson e3dbdbd44d Fix missing tpm2-tools
Needed for full TPM based boot
2026-08-11 09:38:41 -04:00
Jarrod Johnson 3c820d7600 Remove redundant print 2026-08-11 09:26:03 -04:00
Jarrod Johnson 7677267ada Change TPM reseal to be more generic
The crypttab in Ubuntu does not have that indication.  Instead just iterate through
crypttab devices and for any tpm2 sealed ones, reseal them.
2026-08-11 09:23:36 -04:00
Jarrod Johnson 0746b843d2 Change Ubuntu to seal to pcrs on firstboot
Consistent with changes for EL, seal PCRs on firstboot to extend the usefulness of some sealing.
2026-08-11 08:43:25 -04:00
Jarrod Johnson 8d519b57a6 Merge pull request #272 from Obihoernchen/ruff-more-checks
Enable more ruff and pyrefly checks
2026-08-11 07:56:41 -04:00
Jarrod Johnson 45155f23ee Merge pull request #271 from Obihoernchen/fix/type-checker-findings
Fix/type checker findings
2026-08-11 07:54:18 -04:00
Markus Hilger b56ff764c2 Rename Pyrefly job 2026-08-11 05:50:55 +02:00
Markus Hilger 47cfa70097 Enable the pyrefly kinds that already report nothing
Fifteen error kinds beyond the two async ones report nothing on this tree
today and have something in it to bite on, so turning them on costs no
findings and keeps it that way.

invalid-syntax is the one that closes a gap rather than covering ground
another check already holds: CI compiles under a modern interpreter, which
accepts syntax the el8 and sles15 interpreters cannot parse. Checking
against python-version rejects it instead, which makes that setting load
bearing for the first time.

Kinds whose subject matter this tree does not contain stay off, among them
everything reached only through typing: the module is never imported, so
TypeVar and namedtuple naming and stale `# type: ignore` have nothing to
find here.
2026-08-11 05:48:17 +02:00
Markus Hilger 5311437d0e Enable the ruff rules that already report nothing
Every rule added here is at zero once the previous commit lands, so it costs
no cleanup: the point is that a future patch cannot introduce one without the
ruff job failing. They are the rest of pyflakes' format-string checks,
flake8-2020, most of bugbear, the pylint warnings that describe bugs rather
than style, four flake8-async rules for blocking calls in coroutines, and
some RUF, LOG, PGH, PIE, ISC and EXE rules in the same spirit. Each was
confirmed to fire on a synthetic violation, so none is silently inert under
the py37 target.

Rules are named by group wherever the group is already clean, and each prefix
stops short of a rule that is not: PLW150 rather than PLW15, which would pull
in PLW1510. Bugbear is listed rule by rule apart from B02 and B03, since
B006, B007 and B018 are all left out on purpose.
2026-08-11 05:47:57 +02:00
Markus Hilger cb5a6fe964 Remove the Windows service entry point
bin/confluentsrv.py is a python2 script that setup.py never lists in
scripts, so no package has ever installed it, and both the systemd unit and
the sysvinit script start bin/confluent instead. It had also drifted out of
step with what it calls: main.run takes the argument vector and this passed
none.

confluentsrv.spec goes with it. It is a PyInstaller spec whose only input
is c:/Python27/Scripts/confluentsrv.py, a path that has never existed in
this tree, left over from the Windows compatibility work.
2026-08-11 05:38:15 +02:00
Markus Hilger 1faa79134f Give the relay web connection a port
WebConnection requires a port, and every other caller passes one, twice in
this very file. Following a relay URL during discovery raised TypeError
instead.
2026-08-11 05:38:00 +02:00
Markus Hilger a22dd62f18 Keep OEM handler signatures in step with their base class
Four calls reached a base method with an argument list it does not accept,
so they raised TypeError about the argument count.

Three of them would have failed either way, since the base only raises
UnsupportedFunctionality. What changes there is that the failure becomes
the intended, catchable one rather than an argument count error the caller
cannot interpret. get_diagnostic_data grew an autosuffix argument
everywhere except the ipmi generic handler, which is the handler used for
unrecognized hardware. The redfish generic handler already had it. The two
storage super() calls dropped the cfgspec they were given.

The fourth is a real fallback rather than a message: the XCC user_delete
dropped the fishclient it receives from redfish/command.py, so deleting a
uid the XCC does not list raised TypeError instead of attempting the
generic Redfish delete.
2026-08-11 05:38:00 +02:00
Markus Hilger c27327dfe0 Name the node when an inventory component is missing
ConfluentTargetNotFound takes the node as its first argument, and every
other caller passes it. The two inventory plugins constructed it with no
arguments at all, so asking for a component that is not in the inventory
map raised TypeError instead of returning the 404 the path was written to
return.

Both now follow the pattern used a few lines further down for volumes and
name the component that was not found.
2026-08-11 05:17:59 +02:00
Markus Hilger 016e08fa63 Catch socket errors, not the socket class, while firmware applies
The retry around the firmware progress poll named socket.socket, which is not
an exception class, so the moment the request it guards actually failed Python
raised "catching classes that do not inherit from BaseException is not
allowed" in place of the error.

socket.error is OSError, which is what a failed poll raises and what the retry
below was written for.
2026-08-11 05:17:05 +02:00
Markus Hilger a325f65076 Pass the address family and type to getaddrinfo by keyword
The loop resolver takes only host and port positionally, so these three calls
raised "BaseEventLoop.getaddrinfo() takes 3 positional arguments but 5 were
given" every time they ran.

get_ipaddr and _find_service have no handler above them, so link local XCC
discovery and a targeted SSDP search both died outright. The snoop copy sits
under an except Exception, which swallowed it and left the MGTIFACE reply
unanswered instead.
2026-08-11 05:17:05 +02:00
Markus Hilger acd6bb228c Clear the last findings in four ruff groups
Each is the only thing keeping its rule group from being selectable whole.
userutil.py imported ctypes with a star; the names it uses are POINTER,
byref, c_char_p, c_int, c_int32, c_uint and cdll. The oem lookup loop had an
else with no break, so the else always ran. The alert parameter table wrapped
int in a lambda that only forwards to it. And the watchdog interval passed 0
where os.environ.get documents a string, which worked because int(0) is 0.
2026-08-11 04:16:29 +02:00
Jarrod Johnson 1a4475f64e Merge pull request #270 from Obihoernchen/fix/fpc-sensor-enumeration
Fix FPC/SMM sensor enumeration
2026-08-10 18:31:20 -04:00
Jarrod Johnson 8cbfaa9662 Merge pull request #269 from Obihoernchen/fix/dangling-asyncio-tasks
Keep spawned asyncio tasks referenced
2026-08-10 18:30:47 -04:00
Jarrod Johnson 05e42a3b09 Merge pull request #268 from Obihoernchen/fix/osdeploy-local-trust-awaits
Await the coroutines in osdeploy local node trust setup
2026-08-10 18:28:38 -04:00
Markus Hilger f5ee86f97e Make the FPC sensor generators coroutines
get_sensor_names and get_sensor_descriptions reach get_psu_count for any
sensor whose table entry carries elementsfun, and get_psu_count is a
coroutine. As plain generators they could not await it, so range() was handed
the coroutine object and enumeration died with "'coroutine' object cannot be
interpreted as an integer".

Every DW612S has such entries, so nodesensors returned nothing for the
enclosure. get_sensor_descriptions was doubly broken: the Lenovo handler
already iterated it with async for, which a plain generator cannot satisfy.

Verified against a DW612S SMM (FPC variant 38). Before, descriptions raised at
the async for and readings raised partway through enumeration; after, both
return all 34 sensors, 19 of which are the PSU entries that never enumerated.
2026-08-11 00:20:49 +02:00
Markus Hilger cb91589363 Enforce that spawned asyncio tasks are kept referenced
RUF006 catches a create_task whose result is discarded. The loop holds only a
weak reference, so such a task can be collected while still pending and the
work disappears without a trace.

Selected last, once the three existing offenders are gone, so the tree stays
clean under it from this commit on.
2026-08-10 23:18:40 +02:00
Markus Hilger 9e41fdc598 Hold the async HTTP handler task until it finishes
run_handler scheduled the coroutine that serves an async HTTP request and
dropped the returned task. The event loop only keeps a weak reference, so the
task could be collected while still pending, leaving the request unanswered
and "Task was destroyed but it is pending!" in the log.

The session already outlives the request in _asyncsessions, so it holds the
task in a set and discards it from a done callback.
2026-08-10 23:18:24 +02:00
Markus Hilger 95d6c96363 Relay console keystrokes from a single ordered consumer
nodeconsole spawned a task per input byte from the stdin reader callback
and kept no reference to it.  Two of those tasks overlap as soon as one
parks in relay_keypresses waiting on the VNC connection, so keystrokes
can reach the node out of order and the escape sequence state (buffer,
inputcontext, modkeys) is mutated by more than one task at a time.  With
the first keystroke relaying slowly, typing abcdef arrives as bcdefa.

Those tasks were also unreferenced, which asyncio documents as
collectable while still pending, so a keypress could be dropped.

Queue the bytes in the reader callback and process them from one
long-lived task instead.  Keep a reference to the watch_input task as
well, since collecting that one takes the whole input handler with it.
2026-08-10 23:17:57 +02:00
Markus Hilger 151fb1efc8 Await the coroutines in osdeploy local node trust setup
local_node_trust_setup() called get_cluster_list() and sign_host_key() without
awaiting them, so "osdeploy initialize -l" aborted with "TypeError: cannot
unpack non-iterable coroutine object" before doing any work.

Both awaits have to land together: sign_host_key() is called in a loop that
unlinks the existing ssh_host_*_key-cert.pub before writing the new one, so
fixing only the unpack would delete every host certificate and then fail.
2026-08-10 23:05:44 +02:00
Jarrod Johnson 486540d24d Correct checking encryptboot pcrs in firstboot 2026-08-10 14:22:12 -04:00
Jarrod Johnson 730f645dc0 Defer PCR sealing to first boot
If someone wants to seal to a PCR
explicitly to prevent booting rescue, the PCR is likely to
extend differently during install.

Leave the volume sealed to the tpm without any PCRs until first boot.

Then wipe the bindings without PCR specified, and seal according to user preferred values.
2026-08-10 13:59:28 -04:00
Jarrod Johnson 120050ae78 Perodically reopen the multicast sockets
It has been observed there are times where an ethernet switch is partially working with MLD snoop/IGMP snoop.  A workaround for the unreliable behavior seems to be to reassert multicast joins ever so often.

Give it a try to restart the SSDP sockets every minute.
2026-08-10 11:24:02 -04:00
Jarrod Johnson 0c18bc1c01 Older ESXi does not support listdetailed, support normal list. 2026-08-10 11:08:32 -04:00
Jarrod Johnson cfcb406dba Merge pull request #267 from Obihoernchen/pyrefly
Add pyrefly for async correctness
2026-08-10 09:10:34 -04:00
Markus Hilger a860211817 Make the pyrefly job blocking
The tree is clean under it now, so the job can fail the run and catch the next
async regression instead of only reporting one.

Nothing is suppressed beyond the two ignores that state their reason at the
line, for an async __new__ and an untyped callback registry, neither of which
pyrefly can model.
2026-08-10 14:36:37 +02:00
Markus Hilger 1c675c5f24 Port the Eaton PDU plugin to asyncio
The plugin was written against the http.client based SecureHTTPConnection, and
when that went away the reference was pointed at the aiohttp WebConnection,
which shares the name and nothing else. Nothing in it could run: the transport
called an async request() without awaiting it and then reached for a
getresponse() the new class does not have, and three PDUClient methods that
were never coroutines were awaited by the entry points.

Two transports now, both local to this plugin. https is aiohttp and stays on
the event loop, since the cert verifier records new fingerprints through
tasks.spawn. http is http.client in a thread, with its own socket so it can
still ask for a smaller segment size before connect: aiohttp only takes a
socket factory from 3.12 on, newer than el9, el10, ubuntu 24.04 or Leap 16
ship. That side has no cert to verify and its credentials arrive already read,
so the thread touches nothing.

connect() establishes and authenticates, wc is just the accessor now, and
logout() no longer sends a session id it never obtained. update() reports an
unsupported element instead of raising NameError.

On the https side cookies follow aiohttp's domain rules and the one POST with
a body goes out as text/plain, where http.client replayed every cookie and
sent no content type. The http side is as before, and neither can be settled
without an Eaton PDU on the bench. Both transports were exercised against a
stand-in: login, outlet read and set, sensors, logout, and a clamped segment
size on the plaintext path.
2026-08-10 14:36:37 +02:00
Markus Hilger a53351a730 Give the virsh console loop a reason to wake
virEventRunDefaultImpl waits for an event that an idle domain need not
produce, so the thread could outlive a deactivation that reported success, and
every later activation was refused while it did. Registering a timeout is what
makes it return: measured, a thread with nothing registered was still running
four seconds after being asked to stop, and with a half second timer it came
out at once.
2026-08-10 14:36:37 +02:00
Markus Hilger 11dc8196b0 Keep hold of a virsh console thread that will not stop
Deactivation dropped its reference once the wait expired, whether or not the
thread had stopped, so the next activation started a second one and revived
the first by setting run_console again. The reference is cleared only when the
thread is really gone, and activation refuses with 0x80 while one is alive.
That makes the wait a courtesy rather than a correctness measure, so it drops
to a second.
2026-08-10 14:36:37 +02:00
Markus Hilger bda403766e Start and stop the virsh console thread with the payload
Activation started an event thread whichever way the base handler had just
answered, so a refusal started one anyway and an already active console got a
second. activated alone cannot tell the two refusals apart, being true
already on the already active path, so the value from before the call decides.

Deactivation joined that thread on the event loop, where it could stall every
other session and its own response. The wait moves off the loop and is
bounded, and the thread is a daemon.
2026-08-10 14:36:37 +02:00
Markus Hilger 42ae44d5bf Finish the server side SOL and cleanup paths
The console awaits its output handler, but both sample BMCs supplied a plain
function, and both dropped the send_data coroutine. virshbmc additionally
receives its stream callback on a libvirt thread, so the send goes through
run_coroutine_threadsafe against the loop captured at activation, called
asyncloop because the class already has a loop method it uses as a thread
target.

ServerConsole asked the session layer to retry, which a ServerSession cannot
do: it never runs Session.__init__, so it has no timeout, and its _timedout
does nothing. IpmiServer.logout was synchronous and an argument short while
_cleanup awaits logout(False). The boot options handler answered, then read an
unbound name and answered again with 0xff.
2026-08-10 14:36:37 +02:00
Markus Hilger 8521c6ad4b Port the IPMI server side to asyncio
bmc.py, serversession.py, fakebmc.py and virshbmc.py were byte identical to
upstream pyghmi: the async port went through the session layer beneath them
and left the server side alone. So every response created a coroutine and
dropped it, and the overrides the now async parent awaits returned None.
Running fakebmc bound no socket, spun a core, and answered nothing.

Everything that sends a response is a coroutine now, and so is the dispatch
that reaches it; the hooks a subclass implements stay ordinary functions, so
an out of tree Bmc is unaffected unless it overrides the payload handlers.
Two things had to leave their constructors, both being coroutines: assigning
the server socket, into bind(), which is why listen() is no longer a
classmethod, and answering the open session request, into
send_open_session_response.

Verified against upstream with ipmitool over nine commands, with identical
output. SOL is not covered: fakebmc reports the payload disabled on both.
2026-08-10 14:36:37 +02:00
Markus Hilger 7ecc2b2818 Replace the housekeeping thread with a loop owned task
Housekeeper dates from when this work was a blocking select loop. Under
asyncio the event loop is already that place, and a thread around it takes a
captured loop reference, a scheduling step before the thread starts, and a
daemon flag, only to sit waiting on a coroutine that never returns with no way
to stop it. A task runs on the right loop by construction and cancels. The
loop is looked up before the coroutine is built, so calling this without one
raises rather than stranding it.

Nothing in this tree used the class. Out of tree callers need
start_housekeeping() instead, and gain the ability to stop it.
2026-08-10 14:36:37 +02:00
Markus Hilger e828ed4ff4 Close the TSM console web session
TsmConsole created an aiohttp ClientSession and never closed it, and leaked it
again when ws_connect failed. Neither was reachable before the connection path
was repaired. It is closed on both paths, and starts as None so that closing
before a connect does not trip over a missing attribute.
2026-08-10 14:36:37 +02:00
Markus Hilger c9d7343e02 Feed console input through a queue
A task per read swallowed failures, could deliver keystrokes out of order, and
at end of input returned with the reader still registered, so a level
triggered selector called it again for the same EOF. The reader queues now,
one consumer sends in order, and it is gathered with the main loop so a
failure reaches the caller.
2026-08-10 14:36:37 +02:00
Markus Hilger 91960527aa Repair the TSM console connection
Three faults in the same few lines. The redfish Command lost its constructor
for an async create, so building one raised TypeError, which the except below
reported as TargetEndpointUnreachable. await_redirect is defined nowhere in
this repository's history, so that call raised too; create performs the
session setup it was meant to trigger. And oem is a coroutine method rather
than an attribute, with its web connection coming from get_wc, which is what
performs the login that sets csrftok.
2026-08-10 14:36:37 +02:00
Markus Hilger 461385522c Accept lastchance in every network manager
The retry pass calls apply_configuration with lastchance=True, which only
NetworkManager accepts. It is unreachable today, since only NetworkManager
returns the 1 that fills the retry list, but it springs the moment either of
the others grows a return, or the retry selection is brought in line with the
first pass. Matching the signatures costs nothing.
2026-08-10 14:36:37 +02:00
Markus Hilger e1071317ed Create shell sessions through the async factory
ConsoleSession grew an async create and lost its constructor, and ShellSession
inherits that. sockapi was updated for the console branch but not the shell
branch immediately below it, so opening a shell session raised TypeError. It
is the only place in the tree that builds one.
2026-08-10 14:36:37 +02:00
Markus Hilger dfde5736e9 Stop the aiohmi event loop from spinning when it has nothing to do
Command.eventloop called wait_for_rsp with no timeout. With nothing waiting or
being kept alive there is no deadline to derive one from, so it returns
without suspending and the loop runs flat out, measured at over 100000
iterations in two tenths of a second.

MAX_IDLE gives it something to wait on, as a ceiling rather than a fixed
delay: real deadlines still shorten it and an arriving packet still ends it
early.
2026-08-10 14:36:37 +02:00
Markus Hilger b996a30a44 Finish porting pyghmicons to async
The Console was never connected, so main_loop ran against a session that had
never been established. Input arrived on a thread that called send_data and
dropped the coroutine; the thread is gone and the loop watches stdin with
add_reader instead. The output handler was a plain function that
Console._print_data awaits.
2026-08-10 14:36:37 +02:00
Markus Hilger ec56e2c67c Rebuild pyghmiutil on the async command API
Command lost its constructor for an async create classmethod, so the utility
raised TypeError before connecting. The onlogon callback it was built around
is gone as well: create establishes the session itself, so both the callback
and the eventloop that waited for it are unnecessary.

docommand awaited nothing, so every operation produced a coroutine that was
printed and dropped, and three of its calls are async generators. Each BMC is
handled in turn now rather than only the last.
2026-08-10 14:36:37 +02:00
Markus Hilger 8428a47d69 Await the console output flushes
ServerConsole._got_sol_payload and Console._got_cons_input both flushed
pending output without awaiting the flush, so nothing was written.
2026-08-10 14:36:37 +02:00
Markus Hilger b96d4b4103 Give the aiohmi command line utilities an event loop
Console.main_loop drives Session.wait_for_rsp, which is a coroutine, so it
spun without ever waiting for a packet. It is a coroutine now, and pyghmicons
runs its main under asyncio.run. pyghmiutil had the same shape around
Command.eventloop.
2026-08-10 14:36:37 +02:00
Markus Hilger 92c9abc74a Await the remaining reachable coroutines
sockapi sent its collective refusal without awaiting tlvdata.send. redfish
handle_sensors returned a coroutine from most branches and None from the short
ones, so the caller's await raised TypeError; it is a coroutine throughout
now. console send_payload waited for a response without awaiting the wait.

Two suppressions are pyrefly limitations rather than bugs: Session defines an
async __new__, and the keepalive registry holds coroutine functions in an
untyped dict.
2026-08-10 14:36:37 +02:00
Markus Hilger 2ea7aed2cc Port the SMM handler, Delta PDU logout and XCC config to async
Taken from fix/asyncio-port-critical, limited to what pyrefly reports.

The SMM handler still used the httplib style connect/request/getresponse that
the async webclient does not have, so nothing was ever sent. It goes through
grab_response_with_status now, with allow_redirects=False to preserve httplib
behaviour, hence the new webclient parameter. That also fixes a login passing
its headers as urlencode's second positional argument.

Delta PDU's logout was a plain function both callers awaited, so the power
paths raised TypeError. XCC's set_system_configuration was the last
synchronous implementation of a method every caller awaits.
2026-08-10 14:36:37 +02:00
Markus Hilger 9148a9e1cf Run the module self tests through asyncio
These __main__ blocks called coroutines as if they were functions, so they did
nothing at all. Single calls go through asyncio.run; sshutil, proxmox and
vcenter needed an _selftest coroutine. Two were invisible to pyrefly because
repr() and list() count as using the result: vcenter's get_vm_serial needs an
await, and proxmox's get_vm_inventory is an async generator.

lldp called _extract_neighbor_data twice, once correctly, so the bare call is
dropped. xcc3.remote_nodecfg is not a self test: every other handler defines
it as a coroutine and selfservice.py awaits it.
2026-08-10 14:36:37 +02:00
Markus Hilger 0bd8a85046 Add a pyrefly job for async correctness
Catches unawaited coroutines and awaits on non-awaitables, which ruff and
compileall cannot see. Only those two kinds are enabled, and in pyrefly.toml
rather than a --only flag so CI, local runs and editors agree; a full check
reports thousands of errors on this largely unannotated tree.

search-path is what lets imports resolve, and without it a large share of the
findings, including everything in the SMM handler, goes unreported. Pyrefly
walks *.py only and a glob does not lift that, so extensionless tools are
named individually. No baseline: it matches by file, kind and column, so a new
mistake at the same indentation as an old one would pass unnoticed.

Advisory until the existing findings are dealt with.
2026-08-10 14:36:37 +02:00
Markus Hilger e6f6b2330f Parse SMM answers as bytes, not as decoded text
lxml refuses a str carrying an encoding declaration, and the SMM declares one
when it answers /data/login, so _webconfigcreds has raised ValueError on the
first thing it does after logging in ever since the switch to lxml. stdlib
ElementTree took the same input, which is why it went unnoticed. fromstring
already means to take either shape, so encode there. Confirmed against a
DW612S.
2026-08-10 14:27:14 +02:00
Jarrod Johnson 4fe1532da9 Merge pull request #266 from Obihoernchen/ruff
Add ruff to CI and fix issues
2026-08-10 08:17:09 -04:00
Markus Hilger 3fe70363e2 Use latest ruff version 2026-08-10 05:32:00 +02:00
Markus Hilger dc141d8a08 Compile the Python files that are not named *.py
compileall only ever compiles *.py. Handed anything else, even by name on the
command line, it skips the file and still exits 0, so the job has been
checking 218 of the 298 Python files in this tree and reporting success for
the rest. Everything without the extension went unchecked: the whole of
confluent_client/bin and confluent_server/bin, the osdeploy deploy scripts,
the loose misc utilities and the setup.py.tmpl templates.

Those files now go to py_compile, which compiles what it is given. The list
is built from python shebangs, read with the shell builtin rather than by
forking head and grep per file, plus four patterns for the files that carry
no shebang at all and so cannot be detected: the setup templates, configbmc,
add_local_repositories and misc/filterpasswd. It comes to the same 298 files
ruff.toml arrives at through extend-include, and wants keeping in step with
it.

The shebang test matches python anywhere in the line rather than after a
slash or space, because several tools use /usr/libexec/platform-python.
2026-08-10 05:32:00 +02:00
Markus Hilger 6a6d559d81 Speedup ShellCheck 2026-08-10 05:32:00 +02:00
Markus Hilger d31dba4029 Enforce the rules the tree is now clean under
Adds the checks whose findings were cleared in the two preceding commits
(F401, F541, E701, E711, E712, E713, PLC0414) plus three that were already
at zero and cost nothing to lock in: E401, B015 and B023.

B905 is deliberately left out even though it also reads as clean: it only
reports on py310+, and satisfying it would mean adding a keyword the oldest
interpreters this tree runs on cannot parse.
2026-08-10 05:32:00 +02:00
Markus Hilger c6c2d3112e Tidy comparisons, statement layout and a redundant alias (E711, E712, E701, PLC0414)
Hand written rather than autofixed, since three of the four need the
surrounding code read to be sure they are equivalent:

- confetty: `powerstate == None` -> `is None`.
- nodeconfig: `setmode != True` / `!= False` -> `not setmode` / `setmode`.
  Safe because setmode only ever holds None, True or False, and the two
  lines above each test normalise None away first.
- pam: split two `if cond: stmt` one-liners.
- imgutil: `from shutil import copytree as copytree`, an alias that renames
  nothing.  Not a re-export marker, this is a script.
2026-08-10 05:32:00 +02:00
Markus Hilger 644843b892 Remove unused imports and pointless f-string prefixes (F401, F541, E713)
Entirely mechanical, produced by `ruff check --fix --select F401,F541,E713`
and reviewed rather than taken on faith: deleting an import is only safe if
nothing imports it for its side effects or re-exports it.  None of the 19
removed names is referenced anywhere in its file, none appears in any string
literal, and none of the touched files uses eval, exec, globals() or
__import__, so there is no dynamic lookup that could reach them.
2026-08-10 05:32:00 +02:00
Markus Hilger 3d79c3535d Add ruff configuration and a CI job
The enforced rule set is deliberately narrow: undefined names, statements
in impossible positions, duplicate definitions, invalid escapes and a
couple of bugbear checks that only fire on genuine defects.  No style
rules, and the tree is clean under it as of the preceding commits.

Discovery needs help.  Ruff only walks *.py, and about a quarter of the
Python here has no extension: every node* CLI tool, the server bin tools,
the osdeploy scripts (some of which carry no shebang either) and the
setup.py templates.  extend-include lists them, and *.sh is excluded so
the shell scripts sharing those directories are not parsed as Python.

The CI job pins both the action and the ruff version, since there is no
pyproject.toml for the action to read a version from and an unpinned
`latest` would let a new ruff release fail an unchanged branch.
2026-08-10 05:32:00 +02:00
Markus Hilger 938070c5b7 Skip pending nodes that have no handler
The guard evaluated the `next` builtin and discarded it, which does
nothing, so a pending node with no handler fell through to
None.NodeHandler(...).  The AttributeError was caught by the enclosing
except and logged as "Unexpected error during discovery", turning a node
that should have been quietly skipped into a spurious error in the log.
2026-08-10 05:32:00 +02:00
Markus Hilger 607845bacf Stop the plugin loader from shadowing the plugin module (F402)
load_plugins() used `plugin` as the loop variable for plugin file names,
which shadows `import confluent.plugin as plugin` for the whole function.
Nothing in the function needed the module, so this was latent rather than
broken, but the next line that does need it would have failed oddly.
2026-08-10 05:32:00 +02:00
Markus Hilger 9984bff909 Fix the dedicated hotspare drive list (B035)
The DedicatedSpareDrives payload was built as a set containing a list
containing a dict comprehension with a constant key, so it collapsed to a
single entry and then raised TypeError on the unhashable list.  Build a
list of drive references, the same shape as the Drives list just above it.
2026-08-10 05:32:00 +02:00
Markus Hilger fcacaca79d Use a raw string for a regex escape (W605)
'\s' is not a recognised string escape.  Python still accepts it today but
warns, and it becomes a syntax error in a future release.
2026-08-10 05:32:00 +02:00
Markus Hilger b119de345b Remove shadowed duplicate definitions (F811)
Three names were defined twice in the same scope, so the first definition
was unreachable:

- lenovo OEM handler: two set_user_access methods, the second silently
  replacing the first.  That made the SMM privilege update dead code.  The
  conditions are mutually exclusive (is_fpc returns None once has_xcc is
  true), so merge both into the surviving method.
- redfish plugin handle_cert_authorities and prepfish
  disable_host_interface: byte identical copies, drop the redundant one.
2026-08-10 05:32:00 +02:00
Markus Hilger 5d9e30de7b Stop loop variables from shadowing what they iterate (B020)
Each of these loops rebinds the name that holds the iterable.  They work
today because the iterable is evaluated once before the loop starts, but
the name is then gone, so any later use reads a loop item instead of the
collection.

- nodeinventory: `for arg in args` / `for arg in arg.split(',')`.
- confignet (common and debian copies): iname holds the comma separated
  interface list and is then reused for each interface in it.
- xcc _get_agentless_firmware: adata holds the adapter query response and
  is then reused for each adapter.

No behaviour change, just distinct names for distinct things.
2026-08-10 05:32:00 +02:00
Markus Hilger 8a3fce85c0 Fix undefined names (F821)
Every one of these raises NameError if its code path is reached:

- nodeapply: run_automation accumulated into an exitcode that only existed
  in run(), so any automation error crashed instead of being reported.  It
  now keeps and returns its own, tracked separately from the exit code of
  the ssh commands: the early exit after the spawn loop tests that one,
  and folding automation failures into it would exit with children already
  running and their pipes abandoned.  Both are reported at the real exits.
- nodeconsole: redraw() reads firstnodename, which was local to
  do_screenshot(); promote it to a module global like the other drawing
  state.
- nodedeploy: the redeploy path appended to a lockednodes list that did not
  exist yet.  The block that follows re-reads the same lock state and acts
  on it, so drop the dead duplicate.
- samples/nodeattrib_from_switch.py, misc/filterpasswd: missing import sys.
- xcc3: fixuuid was never imported.  xcc imports xcc3, so take a local copy
  the way the smm handler does instead of creating an import cycle.
- httpapi: the async session call still passed the WSGI-era env and an
  extra argument to handle_async(), which has taken only querydict since
  the aiohttp port.  Calling it correctly exposed that handle_async()
  registers an AsyncSession before raising on the discontinued long poll
  path, so every request to it would leak a session that is never reaped.
  It now only creates one when there is a websocket handler to yield it to.
- messages: the InputFirmwareUpdate.filename property checked
  self.filebynode[node] with no node in scope.  __init__ already validates
  every expanded path and nodefile() rechecks per node, so drop the checks.
- pam: drop the python2 branches referencing unicode and raw_input.  The
  server has been python3 only since the asyncio port.
- cooltera: the sensor-name listing referenced a nonexistent sensors dict.
  The available sensors depend on the model, which is only known after
  reading the device, so list them from the same status data the readings
  use.
- deltapdu, eatonpdu, geist: the not-implemented response in update() used
  node outside the loop, unlike retrieve() in the same files and unlike
  raritan/enlogic.
- confluentdbgcli: stray self. on a module-level socket connect.
2026-08-10 05:32:00 +02:00
Jarrod Johnson 76aef703ff avoid moving firmware directories if they don't exist 2026-08-07 15:42:10 -04:00
Jarrod Johnson 94c1683663 Add support for specifying tpm2 pcrs in the encryptboot attribute
This allows a user to opt into pcrs if they understand what they are doing.

Some PCRs are sensitive to firmware updates and some are sensitive to boot loader, kernel, boot config, or initramfs.  All of these are an opportunity for an unsuspecting update to remove access to the boot volume.  There are update processes that can be put into place to make this work,
but it is up to the OS update process to address that, and
OS update processes are likely not to address that at this time.
2026-08-06 16:07:19 -04:00
Jarrod Johnson e1839b6c6e Adopt nodes with proxy environment variables set 2026-08-06 15:34:36 -04:00
Jarrod Johnson c063cbff3a Change to using systemd-cryptenroll where available 2026-08-06 15:34:20 -04:00
Jarrod Johnson cc898d5661 Diseregard proxy for various confluent interactions
In some environments, http proxy is set for internet, but does not work internally.

Accommodate by suspending the proxy in confluent contexts.
2026-08-06 12:15:44 -04:00
Jarrod Johnson 91f1010e4c Merge pull request #265 from Obihoernchen/proxydhcp-insecuremode
Honor deployment.useinsecureprotocols for ProxyDHCP boot
2026-08-05 15:25:54 -04:00
Markus Hilger 4570d9f8af Throttle the insecure mode boot refusal log
reply_dhcp4 logs the insecure mode remediation hint on every DHCP
discover it refuses.  A node in this state never receives a reply, so it
retries for as long as it is powered on and the same message repeats
every few seconds.

Rate limit it per hardware address the way the neighbouring boot attempt
messages already do, reusing the ignoremacs window that check_reply uses
for the missing profile hint.
2026-08-05 03:43:14 +02:00
Markus Hilger ad2d021fcc Restore proxyDHCP log throttling
The per-MAC 90 second log throttle in proxydhcp has been inert: the
`skiplogging = True` reset sat in relay_proxydhcp, where it is a dead
local, while the loop in proxydhcp only ever assigns False.  Once the
first packet is handled the flag stays False for the life of the
process, so every retransmitted boot request logs again even though
ignoredisco is updated to suppress it.

Reset the flag at the top of each loop iteration instead, next to the
timestamp check it belongs to, and drop the dead assignment.
2026-08-05 03:43:14 +02:00
Markus Hilger a7b476b3fc Ignore UEFI HTTP boot on ProxyDHCP port 2026-08-05 03:43:14 +02:00
Markus Hilger fea71a0ce4 Honor deployment.useinsecureprotocols for ProxyDHCP boot
reply_dhcp4 declines to answer a PXE boot request unless
deployment.useinsecureprotocols is set to firmware or always, but
proxydhcp had no such check. A node left at the default of never was
therefore still offered a TFTP bootfile and a plain http boot.ipxe URL
whenever the request arrived on port 4011 rather than port 67, so the
attribute silently did nothing in ProxyDHCP deployments alongside an
independent DHCP server.

Apply the same gate, including the UEFI HTTP boot exemption, and log the
same remediation hint. The node attributes are now fetched once and
passed through to get_deployment_profile instead of being looked up
again there.

Requests whose architecture could not be determined are ignored rather
than falling through to the reply. opts_to_dict stops parsing before the
client architecture option whenever the message type is not a request,
and such a packet would otherwise reach the iPXE branch and be handed a
plain http boot.ipxe URL without ever passing the gate.
2026-08-05 03:43:14 +02:00
Jarrod Johnson 5dce6f2b21 Fix identity image deployment of suse 15
The default 'cp' is /lbin/cp, but that fails, use full path to the cp that works.

Perform the hmac registration of api key that was missing.

Remove assumption that the ip will be ipv6, wrapping it only if a : is present in address.
2026-08-04 12:45:52 -04:00
Jarrod Johnson e05707da2e Remove -k from curl invocation 2026-08-04 11:15:53 -04:00
Jarrod Johnson d468f3afb6 Add identity image to the SUSE15 install 2026-08-04 09:21:11 -04:00
Jarrod Johnson b85f15f59e Merge pull request #263 from Obihoernchen/syncfiles2
Apply chown before chmod in syncfileclient permission handling
2026-08-04 08:09:24 -04:00
Jarrod Johnson 4564f51303 Merge pull request #264 from Obihoernchen/syncfiles-preserve-attrs
Preserve attributes in syncfiles without disturbing parent directories
2026-08-04 08:08:12 -04:00
Markus Hilger 271e3b4d93 Preserve attributes in syncfiles without disturbing parent directories
The rsync push carried no preservation flags, so files arrived with their
special permission bits explicitly disabled and a setuid/setgid entry could
only be honored by the permissions= chmod on the client side.

Preservation was turned on once before in e52a9ff70f ("Have syncfiles
attempt to preserve more") and rolled back the same day in c0287e93ed
("Roll back rsync ownership"), because rsync also applied the staging copy's
attributes to the parent directories it merely traversed on the way to the
synced files, clobbering the permissions of system directories such as /etc.
Naming every staged file explicitly through --files-from and adding
--no-implied-dirs confines preservation to the content actually being
synchronized, leaving traversed directories alone and creating missing ones
with default attributes.

Two details follow from the way the staging tree is built. Files are staged as
symlinks, so rsync reads their attributes through to the real file, but
directories are staged as directories and need the source attributes copied
onto them for the otherwise empty ones that have to be named explicitly.
Ownership is mapped from the account the daemon runs as to root, since that
account generally does not exist on the node and would otherwise arrive as a
meaningless numeric id.

--xattrs from that earlier attempt is deliberately left out: with --copy-links
rsync reads xattrs off the symlink rather than its referent, so it transfers
nothing here while adding a failure mode on hosts without xattr support.
2026-08-04 05:59:57 +02:00
Markus Hilger 9788d563aa Apply chown before chmod in syncfileclient permission handling
chown() clears the setuid bit of a file on Linux (and its setgid bit, if
the file is group-executable), even when run by root and even when the
owner/group are unchanged. Since the owner/group chown ran after the
permissions chmod, any syncfiles entry combining owner=/group= with a
setuid/setgid permissions= value silently lost the special bits.
2026-08-04 04:33:04 +02:00
Jarrod Johnson a806466f20 Merge pull request #262 from Obihoernchen/syncfiles
Report syncfiles failures instead of discarding them
2026-08-03 10:21:46 -04:00
Markus Hilger 165d229178 Remove missing old obsolete syncfileclient from consolidation 2026-08-03 15:24:15 +02:00
Markus Hilger c0bc33e494 Report syncfiles failures instead of discarding them
get_syncresult() caught the sync task's exception, logged a repr server
side and returned 200 OK with a null body.  The node then called
.get('options') on that null resulted in:

  c1: 'NoneType' object has no attribute 'get'

and syncfileclient still exited 0 as if syncing had succeeded.

Return the error to the requestor as a 500 with an error payload.  On
the node, unwrap the body that grab_url_with_status raises for a
non-success status, print it once and exit non-zero.  Only a failure the
server deliberately reported for this sync is terminal. Anything else,
such as a dropped connection, is re-raised so the existing retry loop
handles it as before.  The same case now reports

  c1: Error performing syncfiles: Syncing failed due to unreadable files: /etc/dangling.conf
  c1: 'syncfileclient' exited with code 1
2026-08-03 15:17:15 +02:00
Jarrod Johnson e213949bf0 Bring fix in from el8-diskless edition of syncfileclient 2026-08-03 08:49:55 -04:00
Jarrod Johnson c3cf2a402f Remove redundant copies of syncfileclient 2026-08-03 08:48:28 -04:00
Jarrod Johnson 6853fd2833 Move syncfileclient to common
It is largely the samey
2026-08-03 08:47:16 -04:00
Jarrod Johnson a23ab6dd50 Merge pull request #261 from Obihoernchen/attrib_env
Fix -e attribute setting for dotted attribute names
2026-08-03 08:22:00 -04:00
Markus Hilger f9dc92ecb6 Fix -e attribute setting for dotted attribute names
nodeattrib/nodegroupattrib -e replaced '.' with '_' in the attribute name
before handing it to the server, not just when looking up the environment
variable.  Any attribute with a dot in it was therefore rejected, e.g.

$ export info_note=test
$ nodeattrib -e gpu1 info.note
Traceback (most recent call last):
  File "/opt/confluent/bin/nodeattrib", line 97, in <module>
    exitcode=client.updateattrib(session,args,nodetype, noderange, options, argassign)
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/opt/confluent/lib/python/confluent/client.py", line 688, in updateattrib
    key, os.environ[key.upper()])
         ~~~~~~~~~~^^^^^^^^^^^^^
  File "<frozen os>", line 714, in __getitem__
KeyError: 'INFO_NOTE'
$ export INFO_NOTE=test
$ nodeattrib -e gpu1 info.note
Error: Bad Request - info_note attribute on node gpu1 is invalid

Keep the attribute name intact and derive the environment variable name
from it separately.  A missing environment variable now reports which
variable names were looked for instead of raising a bare KeyError
traceback.
2026-08-02 03:14:17 +02:00
Jarrod Johnson 4653d3f959 Move ssh scratch location out of /tmp
/tmp is sometimes locked down, move it into the runtime directory instead.
2026-07-29 15:44:39 -04:00
Jarrod Johnson 5111652e01 Fix spurious log on impossible passkey requests 2026-07-29 09:03:56 -04:00
Jarrod Johnson d7dcb07a3f Implement deployment.storage
This is an attribute for a node to indicate preferences for storage.

For now, 'm2' policy will hit m.2 and mirroring kits.
2026-07-28 15:02:27 -04:00
Jarrod Johnson 3501f70c37 Merge pull request #247 from Obihoernchen/unsquashfs
Use multi-threaded unsquashfs to extract untethered images
2026-07-28 08:59:38 -04:00
Jarrod Johnson 035d6849e8 Merge pull request #257 from Obihoernchen/lenovo-async
Fix the NextScale SMM web path on the asyncio port
2026-07-28 08:48:50 -04:00
Jarrod Johnson 86e4603b82 Merge pull request #258 from Obihoernchen/imgutil-async
imgutil: fix async-port fallout in the image pack/capture path
2026-07-28 08:45:18 -04:00
Jarrod Johnson 46296d4751 Merge pull request #259 from Obihoernchen/versioning
Derive build versions from a tracked VERSION file
2026-07-28 08:37:53 -04:00
Markus Hilger ab13ec9e94 Remove duplicate "dracut_install lsmod ethtool" 2026-07-28 02:04:00 +02:00
Markus Hilger 67e84f15f8 Keep the root filesystem guard reachable when extraction fails
source_remote imageboot.sh is the last thing the diskless cmdline hook
runs, so returning early on a failed extraction ended the hook and left
dracut to time out. Falling through instead reaches the existing
/sysroot/sbin/init guard, which reports the failure and holds the node so
it stays reachable over ssh, as it did before extraction was checked.
2026-07-28 01:49:38 +02:00
Markus Hilger 62465812a3 Degrade gracefully when squashfs-tools is missing
imageboot falls back to cp when unsquashfs is unavailable, but the build
side did not: a bare dracut_install/copy_exec aborts initramfs generation
when the binary is absent, and the capture prerequisite check refused to
capture the image at all.

Mark the initramfs copies optional and report the missing package as an
advisory rather than a hard prerequisite, so such images still build and
capture, just without the faster extraction path. Widen the EL check to
every release past el8 so future ones inherit it.
2026-07-28 01:49:01 +02:00
Markus Hilger d9da502fbe Copy LICENSE into confluent_osdeploy before leaving the directory
The copy ran after the cd to the repo root, so it read ../LICENSE from
outside the checkout and never placed the file in confluent_osdeploy/. The
tarball went out without it and the spec's %install, which does
"cp LICENSE" after %setup cds into the unpacked directory, failed.
imgutil/buildrpm already copies before its cd; do the same here.

The aarch64 spec has its LICENSE lines commented out, so only the x86_64
build broke, but the copy was equally wrong in both scripts.
2026-07-27 20:19:39 +02:00
Markus Hilger a9d7b67929 Derive build versions from a tracked VERSION file
Release tags do not live on master: 3.15.2 through 3.15.6 were tagged on branch
3.15, so git describe reaches only 3.15.1 and dev builds were stamped
3.15.2.dev<n>. Besides being confusing, rpm and dpkg both rank the released
3.15.6 above that, so a dev package will not install over a released one.

Add a top-level VERSION file naming the release the branch is working toward
(4.0.0 on master) and a mkversion helper that stamps packages from it, keeping
the tag-derived value as a floor so a forgotten bump cannot go backwards.
mkversion also replaces the block copy-pasted into seven build scripts, and
makesetup no longer writes a per-package VERSION file, so the stale checked-in
confluent_common/VERSION goes with it.
2026-07-27 20:06:22 +02:00
Markus Hilger 4f112fb78d Skip the profile manifest when the server libraries are absent
confluent_imgutil does not depend on confluent_server, and the yaml
import is optional too, yet both capture and pack dereference osimage
and yaml unconditionally when writing manifest.yaml.  With
confluent_osdeploy present but the server absent that raises rather
than producing a profile.

Guard the manifest on both being importable and say so, since rebase
is what the manifest exists for.  The yaml fallback now binds None
instead of leaving the name undefined.

The two call sites carried the manifest write verbatim in both, so fold
them into one function rather than duplicate the guard as well.
2026-07-27 18:49:48 +02:00
Markus Hilger 9956845009 Do not block the import poll loop with time.sleep
osimport polls import progress from a coroutine, so a blocking sleep
between reads stalls the whole client loop.  It was the only use of time
in the script, so the import goes with it.
2026-07-27 18:49:48 +02:00
Markus Hilger 3e7da14a9a Test the import drain loops for an error before a percentage
Both loops that read the importer's output test for a percentage first,
so an ERROR: line whose text carries a % takes the percentage branch and
float() raises instead of the error being reported.  The import target
name can carry one too, and that one is user supplied.  importmedia runs
as a bare task, so the exception is swallowed and the client polls a
phase that never advances.

Test for ERROR: first and treat an unparsable percentage as no
percentage.  Set percent on the error path of the second loop as well,
as the first already does.
2026-07-27 18:49:48 +02:00
Markus Hilger b081c17b55 Let the import drain loop accumulate a line
The loop that drains the importer's remaining output reads a byte at a
time but clears currline on every iteration, one level out from where
the earlier loop clears it.  currline is therefore never longer than a
single byte, so the percentage and ERROR: branches can never match and
the tail of an import is silently discarded.

Clear it only once a line has been consumed, as the earlier loop does.
2026-07-27 18:49:48 +02:00
Markus Hilger f2c74b0be3 Fingerprint installation media off the event loop
scan_iso walks an entire ISO with blocking libarchive reads, yielding
only once per entry, and the header-sum branch of fingerprint reads the
whole file with no yield at all.  Both run in the daemon, reached from
MediaImporter.init on every fingerprint and importing request.

The scan costs about 8us per entry and is indifferent to media size,
since libarchive seeks past file data rather than reading it: measured
at 80ms for 10k entries whether the image is 0.2 GB or 8.8 GB, and at
310ms for 40k.  The header-sum branch is the one that scales with size,
reading a multi-gigabyte image end to end.

Make the pair plain functions and hand them to a thread instead.
2026-07-27 18:49:48 +02:00
Markus Hilger 547ecf16d4 Release the crypt device if encrypt_image is interrupted
Nothing unwound the loop device and dm-crypt mapping when the copy loop
raised, so interrupting a pack stranded both, still holding the profile's
rootimg.sfs.

Tear them down from a finally.  The retry loop moves with them, so also
honour its tries counter, as unpack_image already does; spinning forever
inside a finally would hang the interrupt it is meant to clean up after.

A bounded retry loop can also give up, and the detach that follows would
then fail with EBUSY and, raising from a finally, replace the exception
that brought us here.  Warn and leave both in place instead.
2026-07-27 17:01:49 +02:00
Markus Hilger 4e9052012f Keep imgutil pack and capture synchronous
Both functions became coroutines solely to await one get_hashes call,
but their bodies are long stretches of blocking work: mksquashfs, the
encrypt_image copy loop, rsync, ssh and osdeploy.

From Python 3.11 on, asyncio.run installs a SIGINT handler that cancels
the main task and returns rather than raising, so an interrupt is only
noticed at the next await.  Interrupting a pack during mksquashfs
surfaced as a CalledProcessError from the dying child instead of a
KeyboardInterrupt, and with a base profile, where nothing is ever
awaited, pack carried on and published the profile before exiting.

Run the loop only around the call that needs it.
2026-07-27 17:01:33 +02:00
Markus Hilger 38be080bec Accept a command list in check_call
check_output unwraps a single list argument, check_call never did, so
callers passing a list hit a TypeError out of create_subprocess_exec.
Two callers do: the genisoimage run behind Windows profile imports,
where an except Exception swallows the failure and the boot.iso is
silently missing, and the nodeconfig run in discovery, which takes out
automatic node configuration on discovery outright.
2026-07-27 17:01:33 +02:00
Markus Hilger aba564914f Limit imgutil manifest hashes to the profile source
capture and pack hash the whole profile directory, which by that point
holds rootimg.sfs, the kernel and the distribution initramfs.  rebase
only ever looks up entries that came from the profile source directory,
so the image blobs cost gigabytes of hashing for nothing.

Pass the source directory as the filter, as generate_stock_profiles
already does.  Older manifests keep working, since rebase reads their
entries with a default.
2026-07-27 17:01:18 +02:00
Markus Hilger 1a9613f22e Hash profile files in larger chunks
The asyncio port added an await between every 2048 byte read, which
roughly doubled the cost of hashing.  imgutil runs entire packed images
through this, and the server pays it on rebase and media import.

Read a megabyte per iteration instead.  That still yields hundreds of
times per gigabyte, so the event loop stays responsive, and sha512 is
independent of the read size, so existing manifests remain valid.
2026-07-27 17:01:18 +02:00
Markus Hilger ec7b96ecfb Decode SMM response bodies before raising them
grab_response_with_status hands back bytes, so every failure path put a
b'<status>error</status>' repr in front of the operator rather than what
the SMM said.  Decode at the raise, replacing rather than failing on a
body that is not valid utf8.  The bodies still reach fromstring() as
bytes, which is what lxml wants when the xml carries an encoding
declaration.
2026-07-27 16:22:46 +02:00
Markus Hilger 4d75c444ca Ride out a transient bad status while firmware applies
The poll loop spends its retry budget on a poll that goes unanswered but
aborted the update on the first non-200, even though an SMM restarting
its web service part way through the apply keeps answering, with
whatever its httpd has to say, before it stops answering at all.  Give a
bad status the same budget as a dead connection.
2026-07-27 16:22:00 +02:00
Markus Hilger f5a90f3e35 Fail set_user_priv on a rejected privilege change
Every other /data call checks the status, this one discarded the
response, so an SMM that refused the user record was reported to the
caller as a successful privilege change.
2026-07-27 16:13:42 +02:00
Markus Hilger 6593863988 Re-establish an SMM web session the chassis has dropped
Staleness is judged by age alone, so a session the SMM ended on its own
reached the operator as a raw error body instead of being retried.
Route the /data calls through a helper that logs back in and retries
once on a 401, which is how the SMM answers once a session is gone.

A hostname or domain write does not end the session, measured on a
DW612S at firmware 1.18, so this covers what the chassis drops by
itself, not a self-inflicted loss.
2026-07-27 16:13:42 +02:00
Markus Hilger c3d6e0ae58 Clear the firmware poll retry budget after a good poll
The counter is there to ride out a few unanswered progress polls, but
nothing ever cleared it, so three failures spread across a long apply
exhausted it and aborted an update that was still making progress.
2026-07-27 16:13:42 +02:00
Markus Hilger d665064dbd Hold the SMM web session across long operations
A firmware update posts on one session for the minutes its apply loop
runs, and an FFDC collection downloads on the session it acquired, but
wc() judges a session by its age alone, so a settings call arriving
thirty seconds in logged that session out from underneath them.  Flag
the long operations the way the IMM and XCC handlers already do and
leave their session in place.
2026-07-27 16:13:42 +02:00
Jarrod Johnson 29ba1d8515 Merge pull request #254 from Obihoernchen/exclude
Add exclude option to confluentdbutil
2026-07-27 09:51:37 -04:00
Jarrod Johnson 27b173f739 Merge pull request #256 from Obihoernchen/client_async
Fix async-port regressions in confluent_client
2026-07-27 08:14:59 -04:00
Markus Hilger f8ea1adec7 Keep the nodediscover CSV import going past a failed assignment
gather propagates the first exception and leaves its siblings running,
so a transport level failure against one node ends the import with a
traceback while the rest of the batch is cancelled at loop shutdown.
The forked children used to contain such a failure to their own node.
assign_macs already reports an error response itself, so this is the
connection dropping rather than the server refusing the assignment.

Collect the exceptions instead, report each one and count it towards
the exit code.  Schedule the assignments as tasks while doing so, since
the plain coroutines are left unawaited if defining a later node raises
before the gather is reached.
2026-07-27 06:15:45 +02:00
Markus Hilger a619b6ed6f Bound the sessions the nodediscover CSV import opens
Replacing the forked children with a gather kept their fan-out: every
row of the import file gets a session of its own and they all start at
once, so a large file opens a local socket and a server side session
task per node simultaneously.

Hold a semaphore for the duration of each node's assignment instead, so
a finished node's session is dropped before the next one starts.  Also
build that session once per node rather than once per MAC, and say why
the caller's session is not reused, which was self evident while this
ran in a forked child.
2026-07-27 06:15:32 +02:00
Markus Hilger 18c24effc6 Initialize the scan total in nodediscover register
register_endpoint primes current but not total, so a first response
without a count field goes straight to

  UnboundLocalError: local variable 'total' referenced before assignment

on the elif.  Start at zero, which skips the progress line until the
server does report a count.
2026-07-27 06:15:08 +02:00
Markus Hilger 1031bad407 Keep the SMM web session across settings operations
Every getter and setter logged out on the way out, which nulled the
cached client and made the session cache inert on exactly the paths it
was meant to serve: a single nodeconfig walk of ntp costs two full
logins for the read and one per server for the write, each of them a
fresh TLS handshake plus, on firmware that omits st2, two extra page
fetches to scrape the tokens.

Leave the session in place and let wc() dispose of it once it expires.
This also stops one coroutine's logout from invalidating the session
another coroutine just fetched and is about to post with.
2026-07-27 04:56:32 +02:00
Markus Hilger 474e2bd975 Drop unreachable web client check in get_diagnostic_data
wc() either returns a client or propagates the exception raised while
logging in; it cannot return None the way connect() could.
2026-07-27 04:56:32 +02:00
Markus Hilger 6e6cbce0a2 Do not report an interrupted firmware update as complete
The retry counter is there to ride out a few unanswered polls, but
exhausting it broke out of the loop with complete still unset and fell
through to the 'complete' return, so an SMM that stopped answering
part way through an apply was reported to the operator as updated.
Raise instead; a genuine finish still leaves the loop on the progress
reaching 100.
2026-07-27 04:56:32 +02:00
Markus Hilger 5135a6cd3d Make the SMM web session cache safe to share
Now that the expiry comparison actually caches a client, the session it
holds is shared, so tearing it down and replacing it needs the same care
the IMM handler already takes:

Dispose of an expired session with a logout instead of dropping the
reference, otherwise every refresh leaves an authenticated session
behind on an SMM that only has a handful of slots.  That logout has to
tolerate a session the SMM has already reaped, hence the except.

Guard the login itself, so two coroutines arriving at an empty or
expired cache do not both log in and orphan one of the two sessions.

Stamp the vintage once the login round trips are done rather than
before, so a slow SMM cannot hand back a client that is already expired.
2026-07-27 04:56:32 +02:00
Markus Hilger ca81907d25 Restore SMM web request semantics lost in the async port
The old WebConnection.request() added a
'Content-Type: application/x-www-form-urlencoded' header to any POST
carrying a body, but grab_response_with_status() only sets a content
type for dict payloads, so the SMM login and every /data form POST now
go out as text/plain.  This is not a fix for an observed failure: an SMM
running FPC variant 38 was measured accepting a text/plain login exactly
as readily as a urlencoded one.  It restores the header the synchronous
code always sent and that the TSM and IMM handlers still set explicitly,
rather than relying on every SMM firmware level being equally lax about
what it will parse.

Also stop hard failing on responses the synchronous code discarded on
purpose.  'set=securityrollback:1' is only understood by newer SMM2
firmware.  And /data/logout answers 401 once the session is gone, as
measured on that same SMM, so raising on a non-200 there turns a
completed hostname, domain or NTP operation into a spurious error.
2026-07-27 04:55:57 +02:00
Markus Hilger 8817ee6deb Await NextScale SMM settings operations
The SMM hostname, domain, and NTP helpers looked synchronous even though their web transport is asynchronous. Removing awaits in the Lenovo OEM handler therefore returned unresolved coroutine work instead of completed settings results.

Convert the SMM settings and logout helpers to the asynchronous web interface, validate HTTP status responses, and await each operation from the OEM handler so callers only observe completed results.
2026-07-27 04:54:44 +02:00
Markus Hilger fbbeda6c86 Fix NextScale asynchronous web client
The NextScale SMM path still used the removed http.client-style interface against the asynchronous WebConnection implementation. Login, configuration, diagnostic, and firmware operations consequently called unavailable methods or left request coroutines unresolved.

Make web-client creation asynchronous, migrate the affected requests to grab_response_with_status(), and await the cached client accessor. Correct the cache expiry comparison so fresh authenticated clients are reused and stale clients are renewed.
2026-07-27 04:54:08 +02:00
Markus Hilger 53f1d4a7c2 Do not block the event loop with time.sleep in the async clients
nodediscover's rescan poll and nodeconsole's screenshot refresh both slept
with time.sleep inside a coroutine.  In nodeconsole --video that stops the
input handler and the VNC streaming tasks for the whole interval.
2026-07-27 01:54:24 +02:00
Markus Hilger 15670f0ab1 Put the local socket in non-blocking mode in the async client
_connect_unix left the socket blocking, while _connect_tls sets a zero
timeout, so every loop.sock_recv and sock_sendall against the local
socket ran the blocking call inline and stalled the whole event loop.
nodeconsole --video showed this most clearly: a power action opens its
own session, so the tiles stopped refreshing and keystrokes went
unhandled until the BMC finished.

asyncio only enforces this in debug mode, where the client failed
outright with ValueError: the socket must be non-blocking.  The
descriptor passing retries in asynctlvdata also assume a non-blocking
socket, since they wait for BlockingIOError.
2026-07-27 01:54:01 +02:00
Markus Hilger 64cfad09af Complete the async port of the nodegroup attribute paths
simple_nodegroups_command awaited the async generators returned by read
and update, which raises

  TypeError: 'async_generator' object can't be awaited

and printgroupattributes was left synchronous, iterating one of those
generators with plain for.  Neither is reachable yet, since nodeattrib
only ever passes a noderange and nodegroupattrib still uses the
traditional client, but they are the paths nodegroupattrib will use once
it is ported.
2026-07-27 01:52:09 +02:00
Markus Hilger 550751d0ff Fix the file descriptor send retry in asynctlvdata
When sendmsg() reports EAGAIN, _sendmsg rescheduled itself with

  loop.add_reader(fd, _sendmsg, loop, fut, sock, fd)

which waits for the socket to become readable rather than writable, and
passes four of the six required arguments, so the callback raised
TypeError once it did fire.  Wait for writability and pass the message
and descriptors through.

Also skip the work in _recvmsg if the future was cancelled while waiting
for data, as _sendmsg already does, so a cancelled read does not end in
InvalidStateError from set_result.

This module is imported by the server as well, so both paths are reached
by the daemon whenever a descriptor is passed over the local socket.
2026-07-27 01:51:53 +02:00
Markus Hilger 7ebc1dc616 Fix nodediscover CSV import in the async port
import_csv was left with several synchronous idioms:

- search_record is a coroutine function, but was called without await.
  The returned coroutine is always truthy, so the rescan on incomplete
  discovery data never happened, and iterating the result raised
  TypeError: 'coroutine' object is not iterable
- the node creation loop iterated an async generator with plain for
- the per-node discovery assignment was forked off with os.fork() while
  the event loop was running, and the child then built a fresh session
  on the inherited selector

Assign discovery entries with asyncio.gather instead of a forked child,
which keeps the assignments concurrent and lets their exit codes
propagate.  The forked child always ended in sys.exit(0), so its
accumulated errorcode was discarded.
2026-07-27 01:51:16 +02:00
Markus Hilger 4197bd9118 Fix nodediscover register and subscribe in the async port
register_endpoint and subscribe_discovery were left as plain functions
iterating the async client generators, so nodediscover register,
subscribe and unsubscribe all failed immediately with

  TypeError: 'async_generator' object is not iterable
2026-07-27 01:50:24 +02:00
Markus Hilger db303ca014 Allocate a fresh node index when merging a backup
A node imported by a merge was assigned a free index and then had it
overwritten by the index carried in the backup, which may already belong
to a node in the target database.  Keep the allocated index instead; a
full restore still honors the dumped index.
2026-07-25 05:17:46 +02:00
Markus Hilger fa3d1ca388 Add exclude option to confluentdbutil
The -x/--exclude option drops matching node and node group attributes
from a dump, restore, or merge, so a backup can leave out dynamic state
such as deployment.state_last_updated or data that should not travel with it.
Patterns use shell-style wildcards, and a bare namespace such as net
excludes every attribute below it.  The node "groups" and "id.index"
attributes and the node group "noderange" attribute are always retained
so that a restore can still reconstruct group membership and node index
assignments.
2026-07-25 05:17:22 +02:00
Jarrod Johnson 3f2ad75b6d Change output of the nodedeploy timestamp
The 'updated' could be confused for OS updates or similar.
2026-07-24 16:23:27 -04:00
Jarrod Johnson bf195bfdeb Fix stray typing in nodedeploy 2026-07-24 15:10:06 -04:00
Jarrod Johnson a65583c325 Clean up some headers missed in the rebase to aio http 2026-07-24 12:38:33 -04:00
Jarrod Johnson 9fa89d712a Format timestamp consistent with nodeveentlog 2026-07-24 12:05:58 -04:00
Jarrod Johnson b3b16c6497 Do not fail on inability to do REUSEPORT 2026-07-24 11:52:41 -04:00
Jarrod Johnson d81ab1d239 Remove microseconds from the last updated timestamp 2026-07-24 09:30:42 -04:00
Jarrod Johnson 43c98552f4 Change to use standard iso format 2026-07-24 09:12:17 -04:00
Jarrod Johnson 2fed3ddb57 Add timestamp to booted information on diskless boot
It can be ambiguous if the node booted recently or not.
2026-07-24 08:46:47 -04:00
Jarrod Johnson 538a51310b Allow confluentdbutil to operate on an alternate directory 2026-07-24 08:12:02 -04:00
Jarrod Johnson 61e0524a56 Some fixup of SELinux contexts for EL10 diskless boot
Unfortunately, the problem of urlmount's selinux context is left open.

urlmount starts before policy load, preventing transition.

However the policy blocks access urlmount needs when loaded.
2026-07-23 15:58:07 -04:00
Jarrod Johnson 57418696f1 Merge remote-tracking branch 'xcat/master' 2026-07-20 17:11:32 -04:00
Jarrod Johnson 5136b95cde Adjust to EL10 grub stub cfg
The syntax changed, make the code more adaptive to a variety of situations.
2026-07-20 17:11:19 -04:00
Jarrod Johnson 933354f454 Merge pull request #251 from Obihoernchen/yaml-fixes
Harden and clean up config dump/restore
2026-07-20 14:29:55 -04:00
Markus Hilger b5f54c9382 Fix tenant enumeration path in dump_db_to_directory
os.path.join(ConfigManager._cfgdir, '/tenants/') discards the cfgdir
because the second component is absolute, so it resolves to /tenants/
rather than <cfgdir>/tenants.
Tenants are not used yet but let's not face this issue in the future.
2026-07-20 19:18:08 +02:00
Jarrod Johnson 9dfb3ea42b Merge pull request #250 from Obihoernchen/stateless-booted-status
Report stateless boot completion via new 'booted' status
2026-07-20 12:55:55 -04:00
Jarrod Johnson 9128f774ea Merge pull request #249 from Obihoernchen/merge
Comment out the MERGE statement in syncfiles
2026-07-20 12:53:41 -04:00
Markus Hilger 93a6c535b2 Clean up YAML dump/restore code
Rename the format parameter to fmt to stop shadowing the builtin,
pass the already parsed key data dict directly to _restore_keys
instead of reserializing it, use yaml.safe_dump for symmetry with
the safe loader, and consolidate the five repeated per-format dump
blocks into one helper.
2026-07-20 18:03:42 +02:00
Markus Hilger 9101b07d54 Report stateless boot completion via new 'booted' status
Diskless profiles had the updatestatus callback in onboot.sh commented
out because no suitable status existed: 'complete' clears
deployment.pendingprofile, which the PXE responder requires to answer
the next network boot of a diskless node.

Add a 'booted' status that records the pending profile as
deployment.profile while leaving pendingprofile armed and skipping
autolock, and enable the onboot.sh callback in all diskless profiles.
nodedeploy now shows 'pending: <profile> (booted)' for a running
stateless node.
2026-07-20 15:59:30 +02:00
Markus Hilger 73c5b9cebf Comment out the MERGE statement in syncfiles
It's confusing for users to have this enabled by default.
This should be opt-in as everything else.
2026-07-18 03:40:48 +02:00
Markus Hilger d415fcd97f Restrict YAML implicit typing on restore
PyYAML implements YAML 1.1 implicit typing, so hand edited values like
'yes', '52:54:00:12:34:56', or '2026-07-17' in a YAML dump would be
restored as bool, sexagesimal int, or date instead of strings (the
date additionally crashing the JSON re-serialization).  Load with a
SafeLoader subclass that only implicitly types scalars the PyYAML
dumper would have quoted when emitting strings, keeping dump/restore
round trips faithful.
2026-07-18 01:26:14 +02:00
Markus Hilger a8b443dc93 Provide clearer error on restore with mismatched dump format
Restoring a YAML dump without --yaml (or vice versa) previously
reported 'Cannot restore without keys, this may be a redacted dump'.
Point at the actual format of the dump instead when the keys file
exists in the other format.
2026-07-17 23:56:05 +02:00
Jarrod Johnson 12ef3fc529 Merge pull request #245 from Obihoernchen/selinux
SELinux label diskless runtime files on EL
2026-07-16 20:10:43 -04:00
Jarrod Johnson 2b902cb67f Merge pull request #246 from Obihoernchen/el10capture
Enable EL10 imgutil capture
2026-07-16 20:09:42 -04:00
Jarrod Johnson 8e5b37e09b Merge pull request #248 from Obihoernchen/add_local_repositories
Skip add_local_repositories if no imgutil build --source is set
2026-07-16 20:07:58 -04:00
Markus Hilger b6fb58b31f Skip add_local_repositories if no imgutil build --source is set
BUILDSRC is only set if imgutil build is run with --source, otherwise
the build host repos are used. If --source is not used, there is no
distribution symlink and add_local_repositores failed with 404.
Check if BUILDSRC is set and skip add_local_repositories if this is the
case.
2026-07-17 00:06:57 +02:00
Jarrod Johnson 5db32996f9 Provide cleaner error on requesting non-existant tenant. 2026-07-16 17:55:22 -04:00
Markus Hilger cfc4490fe1 Use multi-threaded unsquashfs to extract untethered images
`unsquashfs` can use multiple CPU cores during image extraction, significantly reducing boot time.
For example the whole boot time from PXE to shell on a 8-core VM, from approximately 45 seconds to 20 seconds.

This PR adds `squashfs-tools` as a dependency. Since the package is smaller than 1 MB, the additional image size is justified by the performance improvement.

For backward compatibility, the existing `cp`-based extraction method is used when `unsquashfs` is unavailable, such as with images built before this change.

The extraction logic has also been moved into the common functions and is now shared between EL9, EL10, and Ubuntu.

Both untethered `squashfs` images and `confluent_multisquash` images are supported.

Images must be rebuilt to include `unsquashfs` and benefit from the faster extraction path.
2026-07-16 21:24:58 +02:00
Markus Hilger dd120dc5de Allow el10 capture in imgutil 2026-07-16 19:59:43 +02:00
Markus Hilger e1a066bf49 Use dhcpcd for EL10 capture prerequisites
EL10 diskless networking uses dhcpcd and no longer dhclient and ipcalc.
Keep the existing dhclient and ipcalc requirements for older EL capture targets.
2026-07-16 19:58:53 +02:00
Markus Hilger c3c4805f9d SELinux label diskless runtime files on EL
Centralize the SELinux chcon helper and use it for downloaded
systemd units, onboot hooks, and apiclient files across EL7 through EL10.
Include chcon in captured EL initramfs images.

Without this fix the onboot services failed to start on SELinux enabled
captured image.
2026-07-16 19:52:52 +02:00
Jarrod Johnson 8b0fb03c66 Disable implicit tenant creation
If we support more tenants, we will modify that branch.
2026-07-16 12:39:22 -04:00
Jarrod Johnson a0995b63ba Merge remote-tracking branch 'xcat/master' 2026-07-15 14:47:51 -04:00
Jarrod Johnson ea9be4aae0 Merge pull request #239 from Obihoernchen/ci
Add CI pipeline
2026-07-15 13:54:15 -04:00
Jarrod Johnson 2644ad861d Merge pull request #244 from Obihoernchen/caperm
TLS CA permission fixes
2026-07-15 13:51:47 -04:00
Jarrod Johnson b812d2aa21 Merge pull request #243 from Obihoernchen/sysctl
Increase net.core.rmem_max to 4 MB
2026-07-15 13:47:34 -04:00
Jarrod Johnson 164a168694 Add fallback to mac table
If LLDP is uncooperative, maybe the mac was learned.

If no mac apparently learned, then we ping_everywhere in hopes of soliciting traffic, and then rescan the switches.

Then get all mac addresses, try to determine zone from generated mac, and print on success.
2026-07-15 08:24:33 -04:00
Markus Hilger 63a0cd237f Add missing postinst steps to Ubuntu
The following post install steps were missing on Ubuntu builds:

- Permission fixes
- sysctl load
- Service restart
- confluent PAM symlink to /etc/pam.d/sshd. It works without it on
  Ubuntu because it falls back to other which allows login on Ubuntu,
  but the behaviour should be the same on every OS. Furtheremore, an
  admin might implement additional steps to sshd PAM and would like to
  have this in Confluent, too
2026-07-15 06:02:03 +02:00
Markus Hilger 86d90281e0 Speed up find
Spawn just one find process and stop on first hit
2026-07-15 05:55:48 +02:00
Markus Hilger 4b61351373 Create and maintain the TLS CA as the confluent service account
The CA database under /etc/confluent/tls/ca is typically created by a root context such as osdeploy initialize -t, but the confluent service runs as the owner of /etc/confluent, and openssl ca rewrites the database (index, serial) as the invoking user on every issuance. Certificate issuance through the running service (e.g. the /self/tlscert deployment API) then fails on the root-owned database until packaging happens to repair the ownership.

Run the CA creation (full CA and the currently unused simple CA variant) and the openssl ca invocation under normalize_uid, the convention already used when publishing the CA certificate. The issued certificate is staged through a temporary file since the destination may only be writable by the invoking user, e.g. the web server certificate paths.

Existing root-owned CA databases are repaired by packaging or manually via: chown -R --reference=/etc/confluent /etc/confluent/tls
2026-07-15 05:55:48 +02:00
Markus Hilger fb41537555 Increase net.core.rmem_max to 4 MB
This was introduced 8 years ago. Newer OSs have 4194304 as default.
Confluent shouldn't decrease the default.

RHEL 8           212992
RHEL 9          2097152
RHEL 10         2097152
Fedora 44       4194304

Ubuntu 24.04     212992
Ubuntu 26.06    4194304

SLES 15          212992
SLES 16          212992

It was changed to 4MB in upstream kernel, too:
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=a6d4f25888b83b8300aef28d9ee22765c1cc9b34
2026-07-15 03:26:02 +02:00
Markus Hilger fa605160e0 Merge branch 'master' into ci 2026-07-14 17:04:16 +02:00
Jarrod Johnson 07df9a8c94 Ensure prefix is a string 2026-07-14 11:02:15 -04:00
Jarrod Johnson 5cfc3f07af Fix nodemedia attach
The hardening blocked all URL patters.
2026-07-14 11:02:08 -04:00
Markus Hilger 06c7352394 Log exceptions do not swallow it 2026-07-14 17:01:33 +02:00
Jarrod Johnson 3d0b1c1fce Merge remote-tracking branch 'xcat/master' 2026-07-14 08:32:08 -04:00
Jarrod Johnson 53172f760f Have the priority models brought out and more documented 2026-07-14 08:30:34 -04:00
Jarrod Johnson 0263ad0914 Merge pull request #241 from Obihoernchen/ipv6fix
Preserve scoped IPv6 console addresses
2026-07-14 07:55:03 -04:00
Jarrod Johnson 065426b161 Remove derelict devnull open
This was leftover from pre-async days to support old subprocess
mechanism.
2026-07-14 07:51:47 -04:00
Jarrod Johnson 2b1facb2c5 Merge pull request #240 from Obihoernchen/enos
Reuse ENOS health data
2026-07-14 07:50:56 -04:00
Jarrod Johnson 44eb784b8f Merge pull request #242 from Obihoernchen/ruff
ruff auto fixes
2026-07-14 07:50:34 -04:00
Markus Hilger ab6eeb3ced Merge branch 'master' into ruff 2026-07-14 05:28:53 +02:00
Markus Hilger 2af402b13c ruff auto fixes
Apply ruff's safe autofixes.
The changes are mechanical and behaviour-preserving. Issues fixed:

- F401: remove unused imports.
- F841: drop unused local variables and assignments, including discarded
  await/return values, unused "except ... as e" bindings, and unused
  "with ... as name" targets.
- F541: remove the f prefix from f-strings that contain no placeholders.
- E711: compare against None with "is"/"is not" instead of "=="/"!=".
- E712: test truthiness directly instead of comparing to True.
- E713: use "x not in y" instead of "not x in y".
- E714: use "is not" instead of "not ... is".
- E731: convert lambdas bound to a name into def statements.
- W291/W293: trim trailing whitespace on touched lines.
2026-07-14 05:03:58 +02:00
Markus Hilger b8739b1feb Preserve scoped IPv6 console addresses
Before the async port the bmc var was used for a "host=bmc" parameter
which does not exist anymore.

Now add [] around IPv6 addresses with missing brackets but keep the scope zone like %eth0
as this is needed in current code.
2026-07-14 04:00:45 +02:00
Markus Hilger 686fef6730 Reuse ENOS health data 2026-07-14 03:49:55 +02:00
Markus Hilger 8464effada Add CI pipeline
- Shellcheck (errors only)
- Python compileall for most Python versions supported
2026-07-14 02:16:41 +02:00
Markus Hilger dab2f0bb02 Move returns out of finally blocks
Fixes python compile warning:

SyntaxError: 'return' in a 'finally' block
2026-07-14 01:30:16 +02:00
Markus Hilger 69a6714108 Use raw string notation to fix python compile warnings
E.g.: SyntaxError: "\d" is an invalid escape sequence. Did you mean "\\d"? A raw string is also an option.
2026-07-14 01:26:19 +02:00
Jarrod Johnson 090a887ed9 Merge pull request #235 from Obihoernchen/aiohmi
Various asyncio fixes
2026-07-13 13:26:07 -04:00
Jarrod Johnson c208162839 Add a sample to get ip addresses from a switch port 2026-07-13 13:19:32 -04:00
Markus Hilger 44d2534b4b Revert changes flagged by review
Restore pre-PR behavior for three changes flagged by @jjohnson42.
Bigger changes are needed for these. Will be done in a separate PR.
2026-07-13 17:25:58 +02:00
Markus Hilger dc906b9cd0 Await IPMI session challenge callbacks
The IPMI 1.5 session-challenge callback returned coroutine objects from error reporting and session activation instead of expressing an asynchronous callback contract directly. That made completion depend on the dispatch path noticing and awaiting the returned object.

Make the callback asynchronous and explicitly await both onlogon() and _activate_session(), ensuring failure notification and activation finish before callback dispatch continues.
2026-07-13 17:25:22 +02:00
Jarrod Johnson e853db5c36 Support affluent peeraddresses, when available 2026-07-13 11:16:47 -04:00
Jarrod Johnson 81cf17360e arm64 boot assets pick up 2026-07-13 09:59:15 -04:00
Jarrod Johnson 7babc503c3 Merge pull request #233 from Obihoernchen/checkipmac2
Warn on conflicting entries in confluent2hosts and confluent2dnsmasq
2026-07-13 08:58:19 -04:00
Jarrod Johnson 1660fe9c7d Merge pull request #231 from Obihoernchen/ipmioemfix
Improve SEL record type handling in aiohmi
2026-07-13 08:56:28 -04:00
Jarrod Johnson f4a207ed47 Merge pull request #232 from Obihoernchen/timeout
Add connect timeout option to nodeshell
2026-07-13 08:45:46 -04:00
Jarrod Johnson 6bdd44ffcf Merge pull request #234 from Obihoernchen/netsettings
Add net.extra_settings for passthrough network settings
2026-07-13 08:45:23 -04:00
Jarrod Johnson a2c1b2bd44 Merge pull request #238 from Obihoernchen/shellcheck
Fix Shellcheck errors
2026-07-13 07:56:41 -04:00
Markus Hilger 6d606f37f6 Fix Shellcheck errors
Fix SC2045 (error): Iterating over ls output is fragile. Use globs.

Add exception SC2068 exception for confluent_client/confluent_env.sh as this is intended.
SC2068 (error): Double quote array expansions to avoid re-splitting elements.
2026-07-13 05:34:12 +02:00
Markus Hilger 670a11666d Fix remaining hardware async responses 2026-07-13 02:50:12 +02:00
Markus Hilger ca6a54ab05 Fix asynchronous console control dispatch 2026-07-13 02:50:12 +02:00
Markus Hilger 430260becf Await collective address propagation 2026-07-13 02:50:12 +02:00
Markus Hilger 3c6e7d202f Await BMC discovery configuration operations 2026-07-13 02:50:12 +02:00
Markus Hilger b5c5f62789 Fix OEM asynchronous operation dispatch 2026-07-13 02:50:12 +02:00
Markus Hilger 0165fc9935 Fix IPMI coroutine result handling 2026-07-13 02:50:11 +02:00
Markus Hilger 38746b19d5 Import signal for SSH agent cleanup 2026-07-13 02:50:11 +02:00
Markus Hilger b5e0e9f9e4 Fix asynchronous console and shell contracts 2026-07-13 02:50:11 +02:00
Markus Hilger 9c4f9e1935 Fix hardware management async dispatch 2026-07-13 02:50:11 +02:00
Markus Hilger 89d0fa81b9 Fix asynchronous discovery call contracts 2026-07-13 02:50:11 +02:00
Markus Hilger 5fc036a2b7 Await asynchronous configuration mutations 2026-07-13 02:50:11 +02:00
Markus Hilger 91654ea0d1 Fix aiohmi async call contracts 2026-07-13 02:50:11 +02:00
Markus Hilger 2c41841efc Fix additional missing awaits in aiohmi
Await the channel access raw command, the TSMA remote media settings
requests, and the XCC3 volume creation responses. These calls returned
or unpacked coroutine objects, breaking set_channel_access, TSMA
virtual media attach, and RAID volume creation at runtime.
2026-07-13 02:50:11 +02:00
Markus Hilger 2117d120d8 Fix remaining aiohmi async call paths
Await OEM sensor, NTP, retry, and firmware-update operations that otherwise returned or discarded coroutine objects. Return the initialized energy manager for FAPM systems and update stale utility entry points to use asynchronous Command factories.
2026-07-13 02:50:11 +02:00
Markus Hilger 6650a01a25 Fix Redfish OEM handler instantiation
Fallback paths called OEM handler constructors directly, but these handlers are initialized through async create factories. This caused generic, TSMA, and SMM3 selection to fail with 'OEMHandler() takes no arguments'. Use and await the factories consistently.

Also await the asynchronous bmcinfo lookup and forward the TSMA pool argument correctly.
2026-07-13 02:50:11 +02:00
Markus Hilger 2f006e507f Add net.extra_settings for passthrough network settings
Allow arbitrary per-connection network settings, such as static routes
or a firewalld zone, to be specified as semicolon-delimited key=value
pairs on a net.*.extra_settings attribute. The keys are passed through
to the network backend of the deployed OS in its native syntax: nmcli
properties on NetworkManager systems, netplan YAML paths on netplan
systems, and ifcfg variables on wicked systems.
2026-07-11 20:46:46 +02:00
Markus Hilger 60a00c452b Warn on conflicting entries in confluent2hosts and confluent2dnsmasq
Neither tool detected when the attribute database produces conflicting
name/IP data, silently emitting the conflicts.

confluent2hosts now warns when the same hostname is generated for
multiple different addresses within one address family (dual-stack
IPv4+IPv6 pairs stay silent), which happens naturally in -a mode when a
node has several networks without distinct per-net hostnames.

confluent2dnsmasq now warns when generated reservations share a
hostname across different IPs, reserve the same IP more than once
(dnsmasq refuses to start on a duplicate dhcp-host IP), or reuse a MAC.
2026-07-11 04:29:03 +02:00
Markus Hilger 3af93c2a39 Tolerate standard SEL records with malformed bodies
A type 0x02 record whose body cannot be decoded (e.g. a bogus EvM
revision) would raise and abort retrieval of the entire event log.
ipmitool and freeipmi print such entries with whatever fields they can
extract rather than failing; degrade to the same raw passthrough used
for undecodable reserved types instead of raising.
2026-07-10 16:50:25 +02:00
Markus Hilger 2034d522d7 Decode Linux kernel panic SEL records
The Linux kernel ipmi panic logger stores panic strings in SEL records
of type 0xf0, with a chunk sequence number in byte 4 and up to 11
characters of the message in bytes 5-15.  ipmitool and freeipmi both
recognize this convention; do the same rather than presenting such
records as opaque non-timestamped OEM data.
2026-07-10 16:50:25 +02:00
Markus Hilger 9f35965b1b Decode reserved SEL record types as standard events
Some BMCs (e.g. AMI) log events using spec-reserved record types like
0x04 with a standard system event record layout.  Previously only type
0x02 was decoded, leaving such entries with no usable data and tripping
the generic OEM handler.  Follow ipmitool and treat all types below 0xc0
as standard format.  If the body of a reserved type turns out not to
follow the standard layout, fall back to passing it through raw instead
of aborting the whole log retrieval.
2026-07-10 16:50:25 +02:00
Markus Hilger be7a3c753a Fix crash if sel entry is not OEM 2026-07-10 16:50:25 +02:00
Markus Hilger 0f7ba1b70d Add connect timeout option to nodeshell 2026-07-10 04:31:59 +02:00
Jarrod Johnson 0272137e94 Have custom handling for megarac initial password state
Initial password state demands webgui to change password.

So act like the webgui.
2026-07-09 16:42:28 -04:00
Jarrod Johnson 12c35f2b96 Merge pull request #230 from Obihoernchen/crossarch
Add cross-architecture image build support to imgutil
2026-07-09 14:43:02 -04:00
Jarrod Johnson 0bbc75d53e Copy sshd-session helper if present 2026-07-09 14:24:34 -04:00
Jarrod Johnson 263953fc1e Remove nuisance autoncons output when empty
If no serial console detected, don't bother mentioning it.
2026-07-09 14:22:12 -04:00
Jarrod Johnson beab5cd791 Fix for modern python ioctl
Need to actually feed full buffer into modern python ioctl calls.
2026-07-09 14:13:45 -04:00
Jarrod Johnson 5f34fac2bc confluent_nodename variable might not survive to imageboot
Pull it from the confluent.info file.
2026-07-09 14:13:12 -04:00
Jarrod Johnson 149ecad90e Improvements for MegaRAC discovery
Some Megarac fail with Host header looking like link local.

Systems with nVidia architecture have multiple bmcs, select the actual bmc.
2026-07-09 14:12:23 -04:00
Markus Hilger 5945e8f22d Fix imgutil crash without arg 2026-07-09 19:11:09 +02:00
Markus Hilger 7633bac055 Add cross-architecture image build support to imgutil
Allow building EL and Ubuntu diskless images for a foreign architecture (e.g.
aarch64 on an x86_64 host) by leveraging qemu-user-static. The target
architecture is detected automatically from a -s source tree (for EL),
or may be requested explicitly with the new --arch option.

When the target differs from the host, dnf/debootstrap is invoked with
--forcearch/--arch and the presence of an enabled binfmt_misc handler
with the F (fix-binary) flag is verified up front, so emulation keeps working
inside the installroot chroot and a missing setup yields an actionable
error instead of a confusing exec failure mid-build.

The image architecture is recorded in confluentimg.buildinfo so that
pack selects the initramfs addons for the image architecture rather
than the build host, and exec of a foreign-arch root performs the same
binfmt check.
2026-07-09 19:11:09 +02:00
Jarrod Johnson 0a14e019d0 Skip suse16 diskless for now 2026-07-09 10:47:55 -04:00
Jarrod Johnson b724de4230 Merge pull request #223 from Obihoernchen/showsecret
Add server-side confluentdbutil showattrib subcommand
2026-07-08 17:53:12 -04:00
Markus Hilger 6f11dffae8 Add server-side confluentdbutil showattrib subcommand
Adds `confluentdbutil showattrib <noderange> <attribute>...` to print the
node attribute.

In contrast to nodeattrib it can shows secrets and crypted values with -u flag.
It's server-side only: reads the config store and master key directly, never over
the API.
It's read-only and works without confluentd running.
2026-07-08 19:55:18 +02:00
Jarrod Johnson 40f1a85932 Merge pull request #227 from Obihoernchen/autorelease
Auto add releases for new tags
2026-07-08 12:32:33 -04:00
Jarrod Johnson 0d90317d1f Merge remote-tracking branch 'xcat/master' 2026-07-08 11:57:15 -04:00
Jarrod Johnson e41a844aa3 Further tighten routing for "special" cases
Mitigate risk of misdirection through more explicit routing rules.
2026-07-08 11:56:48 -04:00
Jarrod Johnson c9adc7690a http api fixes
Instead of returning a sessionless authdata if webauthn loaded and validation requested, raise a not found indicating missing webauthn module.

Fix str being passod to rsp_write for the 403 return.

Ensure the console and shell session logic triggers only for subordinates of nodes or noderange.

Fix  str being passed to rsp.write for the successful console session
2026-07-08 11:27:26 -04:00
Markus Hilger 1dcaa40970 Auto add releases for new tags 2026-07-08 16:11:47 +02:00
Jarrod Johnson 9e71ea6b6e Merge pull request #225 from Obihoernchen/license
License naming fixes for EPEL
2026-07-08 09:46:41 -04:00
Jarrod Johnson ec0ab527b2 Fix exception name 2026-07-07 16:56:32 -04:00
Jarrod Johnson 25a6fb82d4 Correct exception name in passkey denial 2026-07-07 16:47:50 -04:00
Jarrod Johnson 8e75585f7d Fixes for shell session operation in select paths
The classic console interface is restored.
2026-07-07 16:39:35 -04:00
Jarrod Johnson 9a4653412c Fix webauthn related issues
The block on user modification shorted out webauthn hooks.

Further, be more picky about the prefix before the username in webauthn registered credentials and validation.
2026-07-07 16:37:45 -04:00
Markus Hilger bb7de607c2 Add missing BSD-3-Clause license of tmt.c
confluent_vtbufferd/tmt.c has a BSD-3-Clause license as described in
confluent_vtbufferd/NOTICE, too.
2026-07-07 21:21:56 +02:00
Markus Hilger dba2af71c7 Match Apache-2.0 license name with SPDX expressions
For EPEL the official SPDX license expressions have to be used.
Check:

- https://docs.fedoraproject.org/en-US/packaging-guidelines/LicensingGuidelines/
- https://spdx.org/licenses/
- https://docs.fedoraproject.org/en-US/legal/allowed-licenses/
2026-07-07 21:08:43 +02:00
Jarrod Johnson 1feec98edf Ensure install interface comes up in firstboot 2026-07-02 17:51:56 -04:00
Jarrod Johnson c3b75f0ca1 Remove stale logging output from enlogic 2026-07-02 16:23:32 -04:00
Jarrod Johnson 72dcd9ef2f Merge pull request #222 from Obihoernchen/spelling
Fix typos and small bugs found during a documentation/UI text review
2026-07-02 16:22:45 -04:00
Markus Hilger f4c43394d8 Fix missing format() leaving {0} literal in error message 2026-07-02 22:07:48 +02:00
Markus Hilger 46ada49401 Remove stray debug write to /etc/whatnowhosts 2026-07-02 22:07:48 +02:00
Markus Hilger 89c710f8c3 Fix key typo dropping verified flag in enclosure discovery 2026-07-02 22:07:48 +02:00
Markus Hilger e280651343 Fix discostatus typo hiding records from the unidentified filter 2026-07-02 22:07:38 +02:00
Markus Hilger 7727cd86fc Fix typos in help text, errors, and log messages 2026-07-02 22:07:27 +02:00
Markus Hilger cfc26f1e60 Fix typos in man pages 2026-07-02 21:52:29 +02:00
Jarrod Johnson 77f2094ff5 Merge pull request #219 from Obihoernchen/nodeattrib_doc
Extend nodeattrib net.* documentation
2026-07-02 15:06:24 -04:00
Jarrod Johnson 98190031df Merge pull request #221 from Obihoernchen/defaultdoc
Add more attribute documentation
2026-07-02 15:05:44 -04:00
Jarrod Johnson f6d7a47140 Successfully indicate install_url and TLS setup
While curl and agama download are happy with the CA bundle, zypper was not.  Have pre.sh properly set up the CA certs.

Additionally, indicate the install subdirectory of the repository to agama via it's cmdline conf.
2026-07-02 14:57:26 -04:00
Markus Hilger 0f20c709c0 Add more attribute documentation
- deployment.lock: add missing 'unlocked' (messages.py's
  InputDeploymentLock/DeploymentLock already accept and persist it).
- hardwaremanagement.method: correct stale "ipmi is used if not
  specified" claim. Was changed to null in
  c14165e2bd.
- snmp.privacyprotocol: document that unset is treated as 'des'
  (snmputil.py explicitly groups None with 'des').
2026-07-02 19:50:56 +02:00
Jarrod Johnson a8cd9a24d5 Auto-restart vtbufferd on exit
If vtbuffer is interrupted, then restart it.
2026-07-02 12:13:24 -04:00
Jarrod Johnson 752d04939b Merge pull request #220 from Obihoernchen/pubkeys_addpolicy
Fix pubkeys.addpolicy documentation to match implementation
2026-07-02 10:31:04 -04:00
Jarrod Johnson d24359a86c Add comments clarifying non-voting state with respect to security expectations 2026-07-02 10:27:21 -04:00
Jarrod Johnson a41e20b1ab Place install_url into agama configuration 2026-07-02 10:19:26 -04:00
Markus Hilger 4c0b2e44f4 Fix pubkeys.addpolicy documentation to match implementation
validvalues listed 'automatic'/'manual', but that was outdated.
Commit 454e1b8267 and cc70dcfa2b
implemented unset/'tofu' (trust-on-first-use, the default), 'manual', 'ca-only',
and an implicit 'ca' (any value that isn't otherwise handled falls
through to the standard CA-verification path, keying an already
pinned match without a full CA reverify).
The validvalues fix in ecaa75d967 rejected
these new values. Add new valid values with proper documentation.
2026-07-02 15:55:50 +02:00
Jarrod Johnson ada4cb196d Lock down non-system users to not have open ended access 2026-07-01 21:11:27 -04:00
Jarrod Johnson 0106758ceb Prevent overwrite of existing files when saving licenses 2026-07-01 21:06:32 -04:00
Jarrod Johnson 1934b88b0d Use basename to ensure no path traversal in license filenames 2026-07-01 20:54:36 -04:00
Jarrod Johnson ae290c4419 Ensure the filename cannot have path traversal in XCC2 and older 2026-07-01 20:42:22 -04:00
Jarrod Johnson 57a4c840cb Fix web shell sessions 2026-07-01 14:47:42 -04:00
Jarrod Johnson 6bcf1b73ba Fix stale references to wsgi style env 2026-07-01 13:49:05 -04:00
Markus Hilger 125b5b3ba2 Extend nodeattrib net.* documentation 2026-07-01 18:21:40 +02:00
Jarrod Johnson 3a6887b4b4 Provide nicer message when requested VM does not exist 2026-07-01 09:55:15 -04:00
Jarrod Johnson d761c7e6da Slow down reconnect attempts to powered down Proxmox VMs and better handle closed websockets. 2026-07-01 09:30:28 -04:00
Jarrod Johnson 45b392932d Handle unreachable proxmox host more friendly 2026-07-01 09:14:07 -04:00
Jarrod Johnson 33c67db3c4 Further mitigate potential XML misbehavior
Since it turns out we already incurred lxml dependency, use lxml etree instead of xml and mitigate risky xml features beyond blocking the word '!entity'
2026-07-01 08:28:25 -04:00
Jarrod Johnson fbec09c073 Fix behavior with IPMI bad user/password 2026-06-30 15:17:45 -04:00
Jarrod Johnson 5abd080ba2 Restore some sanity to redfish error handling 2026-06-30 13:56:34 -04:00
Jarrod Johnson 9ec7100042 Merge pull request #216 from Obihoernchen/dnsmasqdhcp
Implement confluent2dnsmasq
2026-06-30 08:25:49 -04:00
Jarrod Johnson 0383115446 Merge pull request #217 from Obihoernchen/hwplugins
Add missing validvalues to attributes.py
2026-06-30 08:13:21 -04:00
Jarrod Johnson 622e8e696a Merge pull request #218 from gosforthcross/eureka-chassis-support
Add MEGWARE Eureka Chassis support + Small change for how IPs/Hostnames are handled in the redfish hardwaremanagement plugin
2026-06-30 08:08:37 -04:00
gosforthcross 6505810833 Improve handling of IPs with colon notation with seperate IPv4 and IPv6 paths, as well as whitespace stripping 2026-06-30 13:05:22 +02:00
gosforthcross fbb79de786 Include EUREKA in necessary files for loading and handling redfish and autodiscovery 2026-06-30 11:30:24 +02:00
gosforthcross 6ab7d573da Add EUREKA discovery handler 2026-06-30 11:29:19 +02:00
gosforthcross 5435acd23e Add redfish OEM implementation for EUREKA Chassis 2026-06-30 11:28:45 +02:00
gosforthcross cacfce214f Add fallback for generic redfish and add specific MEGWARE code path for EUREKA 2026-06-30 11:25:11 +02:00
gosforthcross b4882692ea Add port discovery via colon notation 2026-06-30 11:22:30 +02:00
Markus Hilger aed0bf0bea Add valid_values to hardwaremanagement.method 2026-06-30 04:26:48 +02:00
Markus Hilger ecaa75d967 Fix validvalues typo
valid_values is never checked and is a typo. Use validvalues instead.

Note: This can break existing scripts if invalid values are used.
2026-06-30 04:21:16 +02:00
Markus Hilger 1129089307 Rename confluent2dnsmasqdhcp -> confluent2dnsmasq 2026-06-30 02:59:12 +02:00
Markus Hilger 6a12b6c977 Use ip route for listen-address and detect missing /prefixlen 2026-06-30 02:55:52 +02:00
Markus Hilger 007b374c73 Implement confluent2dnsmasqdhcp
confleunt2dnsmasqdhcp creates static DHCP entries for dnsmasq
for nodes with defined net.*.hwaddr.
2026-06-30 02:55:23 +02:00
Jarrod Johnson 99405aa6c4 Merge pull request #215 from qisback/add-missing-man-pages
doc/man: add man pages for previously undocumented client commands
2026-06-29 12:33:21 -04:00
Jarrod Johnson a2d4285ee9 Merge remote-tracking branch 'xcat/master' 2026-06-29 11:33:23 -04:00
Jarrod Johnson 3ce0988f5a Fix -s on certutil 2026-06-29 11:30:49 -04:00
Markus Hilger 8a655be794 Update download link in README 2026-06-29 01:57:34 +02:00
Markus Hilger 6308703a10 Use new documentation url in README 2026-06-29 01:56:16 +02:00
Laurence 00a772785d doc/man: add man pages for previously undocumented client commands
These commands ship in confluent_client/bin but had no .ronn man page, so
they did not appear in the generated documentation. Add man pages matching
the existing style, with synopsis and options taken from each command's
argument parser:

- confluent2ansible: export node inventory to an Ansible hosts file
- confluent2lxca: export nodes to a Lenovo XClarity Administrator bulk import CSV
- confluent2xcat: export nodes to an xCAT stanza definition (and optional macs.csv)
- dir2img: build a FAT image from a directory for nodemedia upload
- nodecertutil: manage BMC CA certificates and sign BMC certificates
- nodegrouprename: rename a node group
- noderename: rename nodes
2026-06-28 19:22:34 +01:00
Jarrod Johnson 2c669358b3 Restore ability for certutil to run as standalone script
Also make days an argument
2026-06-26 12:12:25 -04:00
Jarrod Johnson 97e0f4d253 Rework autoconsole logic
Match autocons

Skip unless EFI x86_64.

If SPCR, trust it and use that unconditionally.

Otherwise, if only one can respond to TIOCMGET, then use that one.

If multiple can respond, but exactly one shows carrier, use that.
2026-06-25 16:41:14 -04:00
Jarrod Johnson 9462de42ac Fix setboot when network not in bootorder 2026-06-25 15:56:11 -04:00
Jarrod Johnson e748a97eae Only count copernicus replies that have OK status 2026-06-25 15:08:23 -04:00
Jarrod Johnson c1eea55610 When possible, check confluent user access to file
If a confluent user is a system user, do not allow them to
upload paths that their user would not have access to otherwise.

For non-system users, continue with the path based banned behavior.
2026-06-25 12:14:54 -04:00
Jarrod Johnson d59652e0fd Prevent staging of files from indicating path traversal 2026-06-25 10:09:28 -04:00
Jarrod Johnson 46fbc11a93 Other than skipauth type users (unix domain socket root/confluent), no longer allow confluent user addition/manipulation. 2026-06-25 09:55:09 -04:00
Jarrod Johnson 78ffd509c9 Have messages force normalizing the incoming filenames
This avoids downstream code that may expect specific locations from being confused.
2026-06-25 09:48:11 -04:00
Jarrod Johnson 7845376eaf Fix debian deployment on slow network link up
When network link was slow to establish, it would fall right through
the network initilalization code.

Now keep working it until a result is acheived.
2026-06-25 08:36:37 -04:00
Jarrod Johnson 8fdf3a9abe Do not set 0.0.0.0 gateway 2026-06-24 15:58:45 -04:00
Jarrod Johnson 1449eee4f4 Fix to more reliably default to 47 2026-06-23 09:55:29 -04:00
Jarrod Johnson cf7f2f434d Add function for nodes to request a TLS certificate from confluent
Also, make certificate lifetime default configurable as attribute, with 47 as explicit default.
2026-06-23 09:39:16 -04:00
Jarrod Johnson 1068be423a Remove some python2 considerations 2026-06-23 08:30:21 -04:00
Jarrod Johnson d3f3242eea Draft attempt at a tlscert self api 2026-06-22 16:40:45 -04:00
Jarrod Johnson 701a9a7268 Adjustments for Suse 16.1 beta 2026-06-22 09:59:40 -04:00
Jarrod Johnson a117aace05 Fix suse16 firstboot behavior 2026-06-16 16:37:31 -04:00
Jarrod Johnson b89cba1643 Fix non-bonding configuration of ubuntu 2026-06-16 12:41:30 -04:00
Jarrod Johnson 4aef4ac6d3 Bring XCC3 raid config workaround forward 2026-06-16 10:49:31 -04:00
Jarrod Johnson 8f2b044cc9 Make SUSE16 support adaptive to SLES/Leap 2026-06-16 10:49:14 -04:00
Jarrod Johnson 1afba4ecaa Add initprofile for suse profiles 2026-06-15 14:13:09 -04:00
Jarrod Johnson 0b6e5fa63e More work on SUSE16 deployment 2026-06-15 14:04:35 -04:00
Jarrod Johnson fc1a16e77f Fix prepfish for XCC3
XCC3 does not do SSDP over the USB host interface.

Workaround by assuming a 'mac - 1' could work.
2026-06-15 10:23:26 -04:00
Jarrod Johnson af84c6c45b Further implement SUSE16 deployment 2026-06-12 16:51:46 -04:00
Jarrod Johnson 8b5147a321 Actually put the common copy of functions in 2026-06-12 16:17:34 -04:00
Jarrod Johnson 0645788b0d Rework functions to be common file 2026-06-12 16:06:06 -04:00
Jarrod Johnson c33a9730f2 Iterate SUSE16 deployment support 2026-06-12 15:30:12 -04:00
Jarrod Johnson f58cc7d984 Fix sed invocation with slashes in the value 2026-06-12 14:08:51 -04:00
Jarrod Johnson 7ddd18391d Commence work on autoinstall scripts for SUSE16 2026-06-12 13:29:29 -04:00
Jarrod Johnson 1fc1f47e00 Preserve a breadcrumb to indicate autocons results to live env 2026-06-12 08:10:32 -04:00
Jarrod Johnson 1042404a38 Move getinstalldisk for all the Linux into a common place
esxi is the only one with a different version
2026-06-12 07:45:41 -04:00
Jarrod Johnson 033b2e4e3f Make it possible for a common script to be superseded by a profile specific script 2026-06-12 07:43:11 -04:00
Jarrod Johnson f30e2135d5 Store inst.script where Suse16 agama will actually read it
Agama ignores the dracut cmdline.d and uses only /proc/cmdline

Good news is they have their own agama.conf and they append to it rather than rewrite it.
2026-06-12 07:42:10 -04:00
Jarrod Johnson d71d66eb43 Add hook to run custom install logic 2026-06-11 19:17:32 -04:00
Jarrod Johnson ad48b718a1 Fixup aspects of Ubuntu diskless boot
For one, ensure a unique machine-id.  Broadly should be done, but critical for bonds to have unique mac
addressses in event of booting a captured image.

Fix ubuntu slow boot due to waiting forever for a network config. Have the transient network config bake into netplan for
a first pass before confignet comes along to do full configuration.

Clean up spurious error messages about grep true: and device busy on mounting overlay.
2026-06-11 15:42:03 -04:00
Jarrod Johnson d4310cefa1 Remove modprobe errors and influence OpenSUSE to actually start agama 2026-06-11 10:36:01 -04:00
Jarrod Johnson 9ae1925e82 Remove stray . from SUSE16 bootstrap 2026-06-11 09:20:17 -04:00
Jarrod Johnson 9ce0dd0d6c Have SUSE16 root parameter match documentation 2026-06-11 09:09:09 -04:00
Jarrod Johnson ecc001d3c4 Activate networking in initramfs 2026-06-11 09:01:50 -04:00
Jarrod Johnson 178defc0f5 Advance Suse16 implementation 2026-06-11 08:19:46 -04:00
Jarrod Johnson b5c71e46ee SMM3 debug log workaround 2026-06-11 07:50:29 -04:00
Jarrod Johnson 8088c7af94 Amend suse16 bootstrap
While suse resembles EL more closely, it doesn't have python that early.

Switch to clortho/shell/curl instead
2026-06-10 16:20:13 -04:00
Jarrod Johnson 1fb0f59e74 Add suse16 to the os categories in spec file 2026-06-10 13:51:19 -04:00
Jarrod Johnson 48f75cc506 Begin SUE 16 support work 2026-06-10 13:25:17 -04:00
Jarrod Johnson 41851082bc Mask architecture specific libraries in initramfs hook 2026-06-10 08:45:54 -04:00
Jarrod Johnson b3ef8bfc1e Add ubuntu 26.04 diskless 2026-06-10 08:14:20 -04:00
Jarrod Johnson 7de61941e1 Change to pandoc for man rendering 2026-06-10 07:44:50 -04:00
Jarrod Johnson 887e804894 Merge pull request #99 from VersatusHPC/imgutil-build-errors
fix(imgutil): propagate image build failures
2026-06-10 07:25:55 -04:00
Vinícius Ferrão 0757f9b48f fix(imgutil): propagate image build failures
Copy Debian apt sources and keyrings into the target before apt runs. Run apt with DEBIAN_FRONTEND=noninteractive.

Return constrained child status to callers, and make pack fail clearly when no kernel was installed.
2026-06-10 01:25:58 -03:00
Jarrod Johnson 525186ac7f Add empty sensors to VM health 2026-06-09 16:52:59 -04:00
Jarrod Johnson 5bfd44528d Provide unknown health for vcenter and proxmox 2026-06-09 16:35:13 -04:00
Jarrod Johnson 44385388b6 Provide more verbose feedback 2026-06-09 16:17:40 -04:00
Jarrod Johnson c14165e2bd Switch to null by default
Require active choice of ipmi
2026-06-09 16:08:13 -04:00
Jarrod Johnson e5c550f200 Avoird warning on nodeconsole -tv exit with newer python 2026-06-08 14:30:51 -04:00
Jarrod Johnson b91cfa562c Add libdl and libpthread 2026-06-08 11:02:32 -04:00
Jarrod Johnson 2f6a87eccd Fix compatibility with newer ssh-agent
Newer ssh-agent defaults to homedir agent location.

Unfortunately, /var/lib/confluent may be a poor fit, so go back to how openssh used to handle it.
2026-06-08 10:31:29 -04:00
Jarrod Johnson c8c00c8f5f Update to ast.Constant
Python removed ast.Num in 3.14
2026-06-08 09:19:04 -04:00
Jarrod Johnson 96780eb975 Have apiclient and copernicus try port 1900, but accept if they cannot. 2026-06-07 07:45:09 -04:00
Jarrod Johnson 3a09861ef6 Set name/email for debian builds 2026-06-05 14:27:05 -04:00
Jarrod Johnson 0b10c240bd Fx issues with IPMI session management
Do not continue waiting when session is broken.

Do not call _timedout without releasing the lock first.

Properly await on relog with bad rakp4

If an accounting issue pushes logontries too far without touching zero, then still recognize retries were exhausted.

Timeout on missing RAKP2 if retries were already exhausted.
2026-06-05 09:44:10 -04:00
Jarrod Johnson 60ff78ecd6 Fix portions of selfservice api
Must be bytes before returning.
2026-06-04 14:37:09 -04:00
Jarrod Johnson ac9fd075b9 Remove use of deprecated get_event_loop 2026-06-04 13:36:11 -04:00
Jarrod Johnson 86abdc4257 Bring changes forward from pyghmi
HTTP boot enablement and fixes for the firmware parameters.
2026-06-04 08:27:03 -04:00
Jarrod Johnson 1aca69fd14 Allow nodedeploy to request http boot specifically 2026-06-04 08:04:45 -04:00
Jarrod Johnson ea693e1246 Add support for http boot
Some redfish require us to be very specific.
2026-06-04 08:04:39 -04:00
Jarrod Johnson b6629c39db Fix to actually use the new update_firmware interposer 2026-06-01 19:49:58 -04:00
Jarrod Johnson 2a84c68ad3 Add support for passing a parameterfile in updates 2026-06-01 19:49:19 -04:00
Jarrod Johnson 6185917ab8 Port megarac changes over 2026-06-01 19:44:38 -04:00
Jarrod Johnson 3809fb8f84 Merge pull request #214 from Obihoernchen/confignet
confignet: Fix interface type detection for IB VFs
2026-06-01 19:30:13 -04:00
Jarrod Johnson 431d4992e0 Fixes for confignet for Ubuntu
Try to find various layers of network config and normalize.

Ultimately, after post subiquity will do some things and easiest to fix in firstboot instead.
2026-06-01 19:15:52 -04:00
Jarrod Johnson 842144ccab Fix iterating the netplan configuration 2026-06-01 19:15:46 -04:00
Jarrod Johnson 418a52e66d Remove cloud-init netplan if redundant 2026-06-01 19:15:30 -04:00
Markus Hilger 84b678cd78 confignet: Fix interface type detection for IB VFs
IB VFs have the following "ip l" output:

4: ibp129s0: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 2044 qdisc mq state UP mode DEFAULT group default qlen 1000
    link/infiniband 00:00:00:8d:fe:80:00:00:00:00:00:00:60:5e:65:03:00:2c:43:c8 brd 00:ff:ff:ff:ff:12:40:1b:ff:ff:00:00:00:00:00:00:ff:ff:ff:ff
    vf 0     link/infiniband 00:00:00:8d:fe:80:00:00:00:00:00:00:60:5e:65:03:00:2c:43:c8 brd 00:ff:ff:ff:ff:12:40:1b:ff:ff:00:00:00:00:00:00:ff:ff:ff:ff, spoof checking off, NODE_GUID 00:00:00:00:00:00:00:00, PORT_GUID 00:00:00:00:00:00:00:00, link-state enable, trust off, query_rss off
5: eno1: <NO-CARRIER,BROADCAST,MULTICAST,UP> mtu 1500 qdisc mq state DOWN mode DEFAULT group default qlen 1000
    link/ether 30:56:0f:17:c0:b4 brd ff:ff:ff:ff:ff:ff
    altname enp196s0
    altname enx30560f17c0b4

This breaks the detection script because index 0 of the "vf 0 ..." line is not link/<type> anymore.
This commit improves the detection logic to fix this.
2026-06-01 21:00:40 +02:00
Jarrod Johnson 223ae16892 Fix missing mounts for ubuntu cloned install 2026-05-28 17:32:37 -04:00
Jarrod Johnson ac11ff405e Fix regex pattern warnings 2026-05-28 10:55:53 -04:00
Jarrod Johnson 18c539fcbd Take the opportunity to move on from cryptodome
Cryptodome is redundant with cryptography, which we need anyway..
2026-05-28 09:30:17 -04:00
Jarrod Johnson 2534948e59 Update rdma version for 10.2 2026-05-26 11:00:41 -04:00
Jarrod Johnson a9d38d74af Add bonding to netplan management 2026-05-26 10:06:16 -04:00
Jarrod Johnson 36b4de0859 Merge pull request #213 from VersatusHPC/xen-drivers
Include xen-front drivers in initramfs
2026-05-23 14:20:17 -04:00
Vinícius Ferrão e0f309a165 Include xen-front drivers in confluent-curated initramfs 2026-05-23 01:53:23 -03:00
Jarrod Johnson 84ca936753 Add grp module 2026-05-22 18:11:01 -04:00
Jarrod Johnson 00df01d298 Add missing unicodedata to genesis python 2026-05-22 17:58:28 -04:00
Jarrod Johnson ef937b533a Add idna encoding to genesis python 2026-05-22 16:48:37 -04:00
Jarrod Johnson a8f91c36bd Add more missing dependencies 2026-05-22 16:44:38 -04:00
Jarrod Johnson 8ae76ecf32 Add _struct to genesis 2026-05-22 16:31:01 -04:00
Jarrod Johnson 603d5ba42c Add binascii to genesis 2026-05-22 16:26:19 -04:00
Jarrod Johnson 6f12062ef2 Add mokmanager to symlinks 2026-05-22 16:13:54 -04:00
Jarrod Johnson 2d8698d152 Add missing dependencies for python3-numpy, python3-pillow, and python3-matplotlib 2026-05-22 16:08:23 -04:00
Jarrod Johnson 61c14a5ffb Add alien, remove some old stale dependencies 2026-05-22 15:54:23 -04:00
Jarrod Johnson bb205a4e1a Reference the correct ipxe for aarch64 2026-05-22 14:52:07 -04:00
Jarrod Johnson 01a393bf7f Remove python3-dns, sync up the webauthn dependency 2026-05-22 11:37:53 -04:00
Jarrod Johnson 65f2a13755 Add missing dependency for el10 2026-05-22 11:19:59 -04:00
Jarrod Johnson fe6b3ca277 Add MokManager to genesis 2026-05-22 10:53:46 -04:00
Jarrod Johnson 9b9333106e Change to using el9build container for osdeploy el9+ utilities 2026-05-22 10:36:36 -04:00
Jarrod Johnson b63b75e05f Merge remote-tracking branch 'xcat/master' 2026-05-22 10:33:56 -04:00
Jarrod Johnson 1699175460 Have buildscripts fix their own directories 2026-05-22 10:30:29 -04:00
Jarrod Johnson d77d220f6e Merge pull request #212 from Obihoernchen/swraid
Fix software RAID creation with newer mdadm versions

LGTM
2026-05-22 09:41:31 -04:00
Markus Hilger 7d7f001826 Handle hostname-prefixed md device names and clear stale superblocks
Newer mdadm versions may load arrays under names like
`/dev/md/<hostname>:raid` after reboot instead of `/dev/md/raid`. Detect both
naming schemes when waiting for the array device and use the resolved
path consistently when determining the underlying md device name.

Also clear existing md superblocks before wiping signatures to avoid
stale RAID metadata interfering with array creation or assembly.
2026-05-22 02:00:41 +02:00
Markus Hilger 94eee92d36 Fix software RAID creation with newer mdadm versions
Recent mdadm versions introduced an interactive prompt when creating RAID arrays
without an explicit bitmap configuration:

  "To optimize recovery speed, it is recommended to enable write-intent bitmap,
   do you want to enable it now? [y/N]?"

This behavior was introduced by upstream change:
https://github.com/md-raid-utilities/mdadm/commit/e97c4e18c847803016aa60066cb6e57c528d83a6

In non-interactive environments such as Anaconda, this prompt blocks installation
and causes RAID creation to hang.

Fix this by explicitly enabling the internal bitmap when creating RAID arrays.
2026-05-21 23:44:53 +02:00
Jarrod Johnson 33935e93ff Add mokutil for potential secureboot assistance 2026-05-21 11:33:38 -04:00
Jarrod Johnson cc7f0b91d9 Update to EL10 changes 2026-05-21 11:09:17 -04:00
Jarrod Johnson 0f70401ad6 Port pyghmi fixes forward 2026-05-21 10:59:19 -04:00
Jarrod Johnson a3968a6e68 Support more states 2026-05-21 10:02:53 -04:00
Jarrod Johnson e8543269c5 Recognize more storage states 2026-05-20 16:14:46 -04:00
Jarrod Johnson 7787a405bc Merge pull request #210 from VersatusHPC/remove-eventlet
Comment fixup and attempted removal of eventlet workarounds that shouldn't be needed
2026-05-20 15:43:06 -04:00
Jarrod Johnson 6bf534aa56 Fix generic refish boot override handling 2026-05-20 15:29:53 -04:00
Jarrod Johnson 41fe249151 Port diskless enhancements from el9 to ubuntu 2026-05-20 15:29:31 -04:00
Jarrod Johnson 3d5663f9a7 Recognize some arm64 paths for imgutil 2026-05-20 12:39:09 -04:00
Jarrod Johnson 875c5eac2c Add UEFI HTTP boot for arm64 to recognized archs
For now, we serve up the whole image, no need to distinguish arm from x86 here yet
2026-05-20 12:38:15 -04:00
Jarrod Johnson eb21d930cb Various cleanups for urlmount 2026-05-20 11:36:15 -04:00
Jarrod Johnson 7bbd9778b1 Fix off by one in urlmount 2026-05-20 11:27:52 -04:00
Jarrod Johnson bc0177388c Use %onlyarch in ubuntu diskless build 2026-05-19 09:53:35 -04:00
Jarrod Johnson 698b29c3c0 Fix syntax mistake 2026-05-15 08:19:53 -04:00
Jarrod Johnson a7a83f5cf0 Ensure that logoutexpiry is set if it is going to be needed 2026-05-14 17:34:25 -04:00
Jarrod Johnson 1754636c3f Add sample material for nodeconsole automation for MOK manipulation 2026-05-14 17:26:39 -04:00
Jarrod Johnson d6ed984cb5 Fix async signature of xcc3 discovery 2026-05-14 17:24:06 -04:00
Jarrod Johnson fff2fa2101 Add missing dependencies for debs 2026-05-14 16:55:49 -04:00
Jarrod Johnson ae2b86b51f Add mok manager to boot media for imgutil images 2026-05-14 16:55:33 -04:00
Jarrod Johnson 3fe9bc968f Amend nodeconsole completion 2026-05-14 16:48:31 -04:00
Jarrod Johnson 0026c3f62a Fix pixel format to be RGBA actually 2026-05-14 09:16:52 -04:00
Jarrod Johnson 38fe07ea28 Fix name of ssh in various ubuntu scripts 2026-05-13 13:52:29 -04:00
Jarrod Johnson c11fdcc286 Add missing syncfiles examples to ubuntu profiles 2026-05-13 13:51:13 -04:00
Jarrod Johnson a99f3de910 Implement a headless mode
For automation, this can make more sense.
2026-05-12 14:55:45 -04:00
Jarrod Johnson eebc4ed3c7 Add expression support to the nodeconsole automation 2026-05-12 14:55:39 -04:00
Jarrod Johnson a537a7ba5c Provide automation facility for nodeconsole
Allow nodeconsole to walk console according to a script
2026-05-12 14:55:28 -04:00
Jarrod Johnson 0d1cd4b570 Correct path to DEFAULT_TIMEOUT 2026-05-08 14:49:57 -04:00
Jarrod Johnson 5d7e88998a Do not pass None, as some aiohttp vintages don't understand that timeout 2026-05-08 14:21:42 -04:00
Jarrod Johnson b623eea76f Update dependency information for deb/prm 2026-05-08 14:17:28 -04:00
Jarrod Johnson facf8f7cac Remove asyncore dependency 2026-05-08 10:36:48 -04:00
Jarrod Johnson 999a824872 Rework input handling by vnc client
Ensure all input order is preserved as it is processed.

Institute a 10ms delay after a key transmit.  This is to prevent overwhelming QEMU console, and improve it for other consoles.

Tested by pasting a large volume of text and seeing that it was intact.
2026-05-08 09:42:19 -04:00
Jarrod Johnson 3a4c8a4cd8 Resolute build dep 2026-05-07 16:30:05 -04:00
Jarrod Johnson 0a9c4bc6a7 Block more OSes on profile check 2026-05-07 15:37:02 -04:00
Jarrod Johnson 26519ff793 Do not import if we don't have matching profile fodder 2026-05-07 14:10:05 -04:00
Jarrod Johnson 2586393a45 Implement unix socket permission controls 2026-05-07 10:07:08 -04:00
Jarrod Johnson 34f8ad221d Fix shift-enter, and fix SS3 handling in general 2026-05-07 09:28:33 -04:00
Vinícius Ferrão 7315200133 Replace pyghmi.util.webclient with aiohmi.util.webclient 2026-05-06 19:48:04 -03:00
Jarrod Johnson d53113fe86 Update setup.cfg for newer python standards 2026-05-06 16:57:46 -04:00
Jarrod Johnson 7ba816fc25 Add missing osdeploy initialize options to completion 2026-05-06 10:36:00 -04:00
Jarrod Johnson 060bfb0926 Add a '-r' argument to refresh site contents
If an environment manually manages all materials,
provide -r to let
them request packing of those materials
without trying to generate any of the content.
2026-05-06 08:46:09 -04:00
Jarrod Johnson b3c8cf348c Add completion for nodecertutil 2026-05-05 16:44:35 -04:00
Jarrod Johnson 1d5c5028cf Significantly rework '-tv' titlebar behavior 2026-05-05 16:25:23 -04:00
Jarrod Johnson 1c3ff13841 Fix newpolicy assignment 2026-05-05 16:25:06 -04:00
Vinícius Ferrão 5e26f48e10 Restore debugger eventlet backdoor comment per maintainer request 2026-05-05 17:11:22 -03:00
Jarrod Johnson cc70dcfa2b Add ca-only policy
This policy forces CA validation every time.

This also checks things like date validity.
2026-05-05 14:39:42 -04:00
Jarrod Johnson 454e1b8267 Allow a policy that only uses certificate authority
Technically this was possible before by setting
a bad fingerprint, but formalize an addpolicy
2026-05-05 11:12:16 -04:00
Jarrod Johnson 0abe252c24 Implement power commands in nodeconsole -tv 2026-05-05 09:49:35 -04:00
Jarrod Johnson b165977870 Disable 'help' text for now, can't be seen.
Also plant a seed for potential titlebar content add.
2026-05-04 19:54:07 -04:00
Jarrod Johnson 9d474591f9 Make 'titlebars' more prominent 2026-05-04 19:44:05 -04:00
Jarrod Johnson c54eb2919a Have nodeconsole cleanly exit 2026-05-04 18:35:57 -04:00
Jarrod Johnson dc627342e9 Fix handling of special keys
Particularly handle alt-arrows
2026-05-04 13:57:07 -04:00
Jarrod Johnson f911198907 Implement keyboard input and focus changes
Replace input handling with an async, this
permitts screen updates while doing commands.

Implement 'send break' (sysrq) and focus move.

Indicate not-yet-active focus with titlebar color.
2026-05-04 12:20:46 -04:00
Jarrod Johnson e55cf43f7a Begin work to add ctrl-e commands to video nodeconsole 2026-05-03 22:37:58 -04:00
Jarrod Johnson 966cb9a01d Fix streaming video behavior on resize 2026-05-03 13:49:40 -04:00
Jarrod Johnson 19d05bd82e Handle desktop resize
Also, remove dead code.

Potentially improve performance by
having numpy do the alpha channel massage.
2026-05-03 13:34:06 -04:00
Vinícius Ferrão 1964d4a4ca Fix unbound exception variable in CalledProcessError handler 2026-05-03 12:04:37 -03:00
Vinícius Ferrão aafd6967ba Clean up iothread design rationale comment 2026-05-03 12:04:36 -03:00
Vinícius Ferrão 52f2086319 Replace eventlet CalledProcessError workaround with proper catch
The repr() check existed because eventlet broke normal exception
catching. With eventlet removed, catch CalledProcessError directly.
2026-05-03 01:45:43 -03:00
Vinícius Ferrão 0b1c40aa4a Remove stale comments that restated the obvious
Drop NullLock rationale comment (referenced removed library),
consoleserver greenthread spawn comment, and update syncfiles
CalledProcessError workaround comment.
2026-05-03 01:44:12 -03:00
Vinícius Ferrão b195429d6b Remove python3-eventlet from build deps and clean up stale references
Drop python3-eventlet from the Ubuntu Noble build Dockerfile. Clean up
remaining greenthread/greenlet terminology in comments across aiohmi
IPMI modules, consoleserver, macmap, and the IPMI plugin. Remove a
commented-out GreenPool reference in macmap.
2026-05-02 23:07:54 -03:00
Vinícius Ferrão a69d828e69 Remove eventlet dependency, migrate to asyncio/concurrent.futures
Replace eventlet.greenpool with concurrent.futures.ThreadPoolExecutor
in the BMC discovery script, using as_completed() for proper exception
propagation and main-thread result aggregation to avoid race conditions.

Remove dead eventlet socket compatibility code (.fd attribute checks)
from the IPMI session layer, and clean up stale eventlet references
in comments across the codebase.

Closes: xcat2/confluent#197
2026-05-02 22:56:39 -03:00
Jarrod Johnson 98aac78e55 Fix await of vnc client create 2026-05-02 12:41:58 -04:00
Jarrod Johnson b7f6c158ea Switch to homegrown async vnc implementation
The pip ones didn't support tight.

Further, when switching to streaming, they were a bit hiccupy with performance.
2026-05-02 12:37:52 -04:00
Jarrod Johnson 3116416799 First pass at '-v' support 2026-05-02 09:16:35 -04:00
Jarrod Johnson fcb2c3b4f5 Switch to mostly binary image manipulation
This saves a few round trips through base64 and reduce memory footprint.
2026-05-01 15:38:54 -04:00
Jarrod Johnson 490a04f276 Include aarch64 names for key libraries in ubuntu diskless 2026-05-01 14:24:53 -04:00
Jarrod Johnson 08f52b1210 Skip hashing content that didn't come from confluent
For the manifest, only things that *could* be package updated matter.

So add a parameter to let get_hashes skip files that couldn't be related.

This speeds up packimage and rebase dramatically.
2026-05-01 12:30:18 -04:00
Jarrod Johnson f587539c2a Fix missing ubuntu diskless content 2026-05-01 12:13:24 -04:00
Jarrod Johnson d10e49ed0d Bring chrony fixes to other scripts 2026-04-30 11:17:01 -04:00
Jarrod Johnson 98cbd7581a Fix diskless profiles for chrony.conf modification 2026-04-30 10:44:28 -04:00
Jarrod Johnson d03e689660 Fix imgutil async call 2026-04-30 10:26:54 -04:00
Jarrod Johnson 7f604e3e35 Fix async handling of passed file descriptors 2026-04-30 09:25:09 -04:00
Jarrod Johnson bfc27595dc Fold aiohmi into confluent
If someone asks for it independently, we can break it out again.  But for now,
assume it's only for confluent.
2026-04-30 08:48:24 -04:00
Jarrod Johnson ea6ab5dc2a Merge pull request #209 from middelkoopt/tm-el9-aarch64
Fix EL9 aarch support
2026-04-30 08:08:26 -04:00
Timothy Middelkoop a3f40e2982 Fix el8/el9 hook paths corrupted by symlinked el10 in aarch64 spec
In confluent_osdeploy-aarch64.spec.tmpl, el10 was created as a symlink
to el8, so the subsequent `mv el10/initramfs/usr el10/initramfs/var`
inadvertently renamed el8's usr directory, leaving el8 and el9 (also
symlinked to el8) with hooks at var/lib/dracut/hooks/ instead of
usr/lib/dracut/hooks/. Rocky 9 dracut never found the hooks and dropped
to the emergency shell on all aarch64 nodes.

Use `cp -a el8 el10` as the x86_64 spec already does, so the rename
only affects the el10 copy.

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Timothy Middelkoop <tmiddelkoop@internet2.edu>
2026-04-29 16:36:23 -05:00
Jarrod Johnson 75776b77a3 Provide mechanism for client session loss to trigger a teardown 2026-04-29 15:20:39 -04:00
Jarrod Johnson ee8a8bdac7 Include ommitted file from previous commit 2026-04-29 11:57:09 -04:00
Jarrod Johnson d2f23b475f Fix some failure to await 2026-04-29 11:47:55 -04:00
Jarrod Johnson ae5afc526a Fix missing rsp on return 2026-04-29 11:41:33 -04:00
Jarrod Johnson 78e5301ff8 Fix attempt to await non-async get_nowait. 2026-04-29 11:29:21 -04:00
Jarrod Johnson 5064ac80b9 Fix accidental change of iterating data in ipmi 2026-04-29 10:53:18 -04:00
Jarrod Johnson 27b951b7cb Honor 'Done' message to avoid incurring a delay after task is done. 2026-04-29 09:57:36 -04:00
Jarrod Johnson 34bc45aa9e Allow monitor to read attributes by 'all' resource. 2026-04-29 07:51:42 -04:00
Jarrod Johnson 1f969f2962 Fixes for accel redirect and port forwarding 2026-04-28 16:27:12 -04:00
Jarrod Johnson 347c7fdc1e Fix osdeploy list 2026-04-28 16:08:24 -04:00
Jarrod Johnson 16c99efda7 Correct firmware update through http api 2026-04-28 15:33:18 -04:00
Jarrod Johnson 6421097f32 Fix logging when client ip == server ip 2026-04-28 15:02:12 -04:00
Jarrod Johnson 83ac9af196 Fix for staging in async 2026-04-28 15:02:00 -04:00
Jarrod Johnson 069338baf3 Write to stdout as binary
This allows better redirection.

In python3, must write to sys.stdout.buffer.  AttributeError for the unlikely event of a python2 based node being deployed.
2026-04-28 08:16:05 -04:00
Jarrod Johnson 17d3022caf Implement username by passkey 2026-04-24 15:31:16 -04:00
Jarrod Johnson e44145f978 Simplify webauthn by keeping with webauthn defaults 2026-04-24 11:40:43 -04:00
Jarrod Johnson d97eba787d Fix mistake in spec file 2026-04-24 09:29:57 -04:00
Jarrod Johnson 260443c1d6 Add Ubuntu 26.04 2026-04-24 08:35:27 -04:00
Jarrod Johnson 056d690db0 Fully fix webauthn as implemented 2026-04-23 17:46:34 -04:00
Jarrod Johnson ee32b8cefc Merge pull request #208 from forryz/fix-ubuntu-initramfs-routing
Handle confluent= boot arg and IPv4 NIC autodetect
2026-04-23 13:59:03 -04:00
Jarrod Johnson bf6a097083 Simplify webauthn implementation
Stop tracking sign counters (which weren't used).

Remove various management of transient challenges.

Co-authored-by: Copilot <copilot@github.com>
2026-04-23 12:52:47 -04:00
xu_ren_xian f269200004 Handle confluent= boot arg and IPv4 NIC autodetect
Add support for a confluent=<host> kernel argument in init-premount: configure networking, flush interfaces, autodetect the primary NIC (saved to /tmp/autodetectnic), verify TLS connectivity to the provided server, call the whoami endpoint over TLS to obtain the node name, and write results to /custom-installation/confluent/confluent.info (with fallback to copernicus on failure).

Also update casper-bottom logic to handle IPv4 manager addresses: for IPv6 the manager is still bracketed and scoped interface resolved as before; for IPv4 the script now uses the previously detected NIC (/tmp/autodetectnic) or falls back to an `ip route get <mgr>` lookup to determine DEVICE. This ensures routed IPv4 deployments work correctly.
2026-04-23 23:23:26 +08:00
Jarrod Johnson 82744c5d52 Simplify webauthn code in httpapi
Co-authored-by: Copilot <copilot@github.com>
2026-04-22 14:17:39 -04:00
Jarrod Johnson 96d368fda6 Push second part of the webauthn rework
Co-authored-by: Copilot <copilot@github.com>
2026-04-22 10:58:59 -04:00
Jarrod Johnson 3fecec7743 Change webauthn to aiohttp
Co-authored-by: Copilot <copilot@github.com>
2026-04-22 10:55:49 -04:00
Jarrod Johnson 06786f202c Fix deployment/storage handling 2026-04-22 10:18:32 -04:00
Jarrod Johnson b1568ca01e Fix some issues with remote video forwarding 2026-04-21 12:42:20 -04:00
Jarrod Johnson 2c9b2a93f3 Rework async session handling 2026-04-20 17:05:39 -04:00
Jarrod Johnson 6c6dbf9c2b Fix references to headers 2026-04-20 16:53:57 -04:00
Jarrod Johnson 2835e9e804 Fix references to start_response 2026-04-20 16:39:51 -04:00
Jarrod Johnson bc5f9cf1e8 Fix mistakes in the node apoption samples 2026-04-20 09:45:44 -04:00
Jarrod Johnson 2ef0748724 Begin work to make shellserver work async 2026-04-17 16:21:33 -04:00
Jarrod Johnson 2650b11421 Rework consolesessieon to async create function 2026-04-17 16:15:03 -04:00
Jarrod Johnson 2d80647f1c Fix TSM discovery 2026-04-17 14:30:02 -04:00
Jarrod Johnson 650b1ee91f Fix collective rename of nodegroups 2026-04-17 11:56:58 -04:00
Jarrod Johnson aec1d62e44 Fix noderename 2026-04-17 11:53:46 -04:00
Jarrod Johnson fb6c8c2ff3 Convert rpc master calls that write to async 2026-04-17 11:46:42 -04:00
Jarrod Johnson e4b04e4198 Fix mistake in exception handling 2026-04-17 11:40:23 -04:00
Jarrod Johnson 3aeb1389a6 Fix indentation error 2026-04-17 10:45:45 -04:00
Jarrod Johnson f1f5f1b3b8 Fix issues associated with unix domain vs tls 2026-04-17 10:42:40 -04:00
Jarrod Johnson 67860dc7c3 Fix remote client operation with Python 3.12+ 2026-04-17 09:00:39 -04:00
Jarrod Johnson 86534b38eb Correct some collective behavior 2026-04-16 16:17:39 -04:00
Jarrod Johnson e4a00d40cc Fix reseat 2026-04-16 15:37:20 -04:00
Jarrod Johnson a26c1409db Rework several aspects of asyncio in consoleserver 2026-04-15 12:18:11 -04:00
Jarrod Johnson b4e0710a98 Correct arguments to WebConnection when getting SMM neighbors 2026-04-15 12:07:01 -04:00
Jarrod Johnson ede16c6ab0 Rework cert validation
Move a generic callback to the generic function
2026-04-15 11:22:21 -04:00
Jarrod Johnson 2c6acb0212 Fix async cert handling 2026-04-15 10:13:31 -04:00
Jarrod Johnson f4c68032e3 Change noderemove to use sync client for now 2026-04-15 10:03:57 -04:00
Jarrod Johnson 39cd8a3bcb Correct async style in various parts of configmanager and dependent core 2026-04-15 09:58:57 -04:00
Jarrod Johnson 2903e6dc23 Update dependencies for async 2026-04-14 15:30:00 -04:00
Jarrod Johnson ec1efecdae Merge branch 'master' into async 2026-04-14 13:51:36 -04:00
Jarrod Johnson c54ac530e1 Handle some environments where timedatectl does not exist 2026-04-14 13:50:12 -04:00
Jarrod Johnson 8db76b92ee Fix update of pinned cert on CA blessing 2026-04-14 10:53:34 -04:00
Jarrod Johnson 2a32fc85a6 Skip policy setting for now and take defaults. 2026-04-14 10:45:45 -04:00
Jarrod Johnson 2bd13c397d Rework for older python cryptography compatibility 2026-04-14 10:45:03 -04:00
Jarrod Johnson 5250a3a67a Pass subject to the verifier in redfish 2026-04-14 10:24:33 -04:00
Jarrod Johnson 038faaab74 Await clear node attributes 2026-04-14 09:57:33 -04:00
Jarrod Johnson a8f4c437bb Remove duplicate copy of function 2026-04-13 16:12:30 -04:00
Jarrod Johnson 9d17102f60 Await creation of the certificate 2026-04-13 16:05:55 -04:00
Jarrod Johnson f2ce13253f Properly place messages on async queue 2026-04-13 13:17:09 -04:00
Jarrod Johnson 65944a4507 For fixes for sync use of async methods 2026-04-13 13:07:10 -04:00
Jarrod Johnson 7c8cee2480 Certificate list fix 2026-04-13 12:46:04 -04:00
Jarrod Johnson 31a56f9fdc Merge branch 'master' into async 2026-04-08 16:17:23 -04:00
Jarrod Johnson 8990622470 Improve certificate mismatch handling 2026-04-08 15:37:50 -04:00
Jarrod Johnson 93a35d7e77 Improve srlinux error handling 2026-04-08 15:30:43 -04:00
Jarrod Johnson 131fa052e0 Rework merge to be async friendly 2026-04-08 15:13:42 -04:00
Jarrod Johnson f5a1c0d1b4 Merge branch 'master' into async 2026-04-08 15:08:30 -04:00
Jarrod Johnson c49b2fd8ab Update quorum on deletion
If deletion of a node brings quorum, notify followers
of the good news
2026-04-07 14:57:09 -04:00
Jarrod Johnson 3ce2a5bc26 More tightly constrain node profile requests
Normalize paths using abspath and validate the result is within the expected path.
2026-04-06 15:12:44 -04:00
Jarrod Johnson 9dad93fd26 Improve raritan support by persisting info about a pdu
This gets the time for a re-sweep of sensors under a second.

It's come a long way from the 3 minutes it was taking.
2026-04-06 11:40:34 -04:00
Jarrod Johnson f125ff0fb6 Implement session management for raritan
The session token accelerates requests.  Shaves a second off of a sensor pass
2026-04-06 11:02:04 -04:00
Jarrod Johnson ec6cdd6c21 Defer inlet sensor to shave a few more seconds from a sensor sweep. 2026-04-06 10:12:18 -04:00
Jarrod Johnson c53206331c Leverage bulk facility
Study of the web interface showed that bulk requests are a key component.

This takes a sensor sweep from about 3 minutes to about 7 seconds in a test environment.
2026-04-06 10:09:10 -04:00
Jarrod Johnson 7bbc451047 Draft raritan pdu support
Particularly need to replace the sensors logic to provide vaguely credible performance
2026-04-03 11:00:23 -04:00
Jarrod Johnson 5ccbc37aa6 Merge branch 'master' into async 2026-04-03 10:36:25 -04:00
Jarrod Johnson 69d984b9dc Fix syntax mistake in deferred handling in nodeapply 2026-04-03 10:34:20 -04:00
Jarrod Johnson dade4239aa Fix user management with redfish 2026-04-03 08:37:05 -04:00
Jarrod Johnson 6d17d9f0f8 Fix some storage api calls under ipmi 2026-04-02 16:33:35 -04:00
Jarrod Johnson 62e9ee8dac Fix user password manipulation in ipmi 2026-04-02 15:52:54 -04:00
Jarrod Johnson 326b659d28 Merge branch 'master' into async 2026-04-02 15:29:49 -04:00
Jarrod Johnson a123165712 Improve error when unknown user specified in syncfiles 2026-04-02 15:29:31 -04:00
Jarrod Johnson 8c9d55d469 Fix nodesetboot 2026-03-27 17:11:35 -04:00
Jarrod Johnson a1e7cb1e9d Fix nodeinventory for vertiv pdus 2026-03-27 17:00:17 -04:00
Jarrod Johnson 7f4da79679 Fix taskpile to work as intended 2026-03-27 16:35:06 -04:00
Jarrod Johnson fecf766183 Change to returning a list of msgs in enlogic 2026-03-27 13:02:35 -04:00
Jarrod Johnson f842a0a6e3 Fikup TaskPile and it's uses 2026-03-27 12:56:47 -04:00
Jarrod Johnson 2462e07832 Fix async gen invocation of apply license 2026-03-26 16:57:46 -04:00
Jarrod Johnson 939d5c59ea Fix async apply license in ipmi plugin 2026-03-26 16:48:37 -04:00
Jarrod Johnson 36d3cdbe41 Merge branch 'master' into async 2026-03-25 13:00:02 -04:00
Jarrod Johnson b91b10552c EL10 doesn't do setgid keysign
chmod 600 instead
2026-03-25 12:59:40 -04:00
Jarrod Johnson 779b07d2c2 Only try to use ssh_keys if it exists
EL10 changed from using ssh_keys and setgid to just
do setuid root instead.
2026-03-25 12:56:16 -04:00
Jarrod Johnson d44e7f0955 Port some license management to async 2026-03-24 16:36:03 -04:00
Jarrod Johnson cd68225672 Fix nodeconfig and nodefirmware to IPMI targets 2026-03-24 16:23:48 -04:00
Jarrod Johnson f18cc981e2 Fix async discovery of XCC 2026-03-24 15:41:48 -04:00
Jarrod Johnson 8c482d9078 Numerous async modifications for XCC1/2 discovery 2026-03-24 14:48:05 -04:00
Jarrod Johnson 78c708424d async fixes for nodeconfig and netutil 2026-03-24 14:08:54 -04:00
Jarrod Johnson 9b6ead31a7 Consume async generator 2026-03-24 08:29:40 -04:00
Jarrod Johnson 8e508ee858 Properly await async pushes to async queue 2026-03-20 16:24:35 -04:00
Jarrod Johnson 7f80a4d5aa Port enhancements from sync client to async 2026-03-20 16:11:26 -04:00
Jarrod Johnson 315fb0ced8 Port more methodns to await 2026-03-20 15:59:00 -04:00
Jarrod Johnson 6dc57abe28 Numerous async changes
For one, simplify and make more robust normalizing iterating various things.

Add required async/await in varous other places.
2026-03-20 15:09:11 -04:00
Jarrod Johnson 36bf03d65c Fix passing of args in handoff to async 2026-03-20 10:56:18 -04:00
Jarrod Johnson 53c2ace620 Go all in on systemd startup
Become a notify type service.  Reserve the ability to be normal should it come up, but with
the complications around fork, easiest
to punt on daemonize for now.
2026-03-20 10:12:45 -04:00
Jarrod Johnson bec5363877 Merge branch 'master' into async 2026-03-19 18:05:50 -04:00
Jarrod Johnson df73c14475 Support unconfigured good without space
Some platforms try to combine the words
2026-03-19 18:05:38 -04:00
Jarrod Johnson d8c7e1fc2a Add a note why fork might be ok in auth 2026-03-19 17:28:51 -04:00
Jarrod Johnson 7adef74fe9 Delay async until after daemonize
os.fork shouldn't happen after an async event loop if the child might use the async loop
2026-03-19 17:27:24 -04:00
Jarrod Johnson 40da956a06 Fix confluent_selfcheck for asyncio
Most dramatically, rework to avoid os.fork, which
ruins threading and by extension the getaddrinfo behavior.
2026-03-19 17:21:30 -04:00
Jarrod Johnson 07a6eb32ed Merge branch 'master' into async 2026-03-19 12:21:10 -04:00
Jarrod Johnson f78b301143 Update usage text 2026-03-19 09:54:08 -04:00
Jarrod Johnson 57fe186a10 Rework redfish to be more async 2026-03-17 16:43:41 -04:00
Jarrod Johnson b4c4ac0861 Fix async syncfiles 2026-03-17 16:10:21 -04:00
Jarrod Johnson 2ab85bb687 Refactor functions to be a bit more readable 2026-03-17 15:40:01 -04:00
Jarrod Johnson e1a6a1c9bf Merge branch 'master' into async 2026-03-17 13:03:56 -04:00
Jarrod Johnson 9b00fe5521 Don't try to open a file that doesn't exist 2026-03-17 13:03:18 -04:00
Jarrod Johnson 13a6444541 Fix incorrectly matching older versions as 'el10' 2026-03-17 12:58:04 -04:00
Jarrod Johnson a6e7d016ea Restore ansible running in async, complete with recent changes from master 2026-03-13 14:57:59 -04:00
Jarrod Johnson 52db46be93 Fix python detection from ansible with space in shebang 2026-03-13 11:41:16 -04:00
Jarrod Johnson 2c8cce74ab Fix some issues from async ansible running 2026-03-13 11:40:44 -04:00
Jarrod Johnson fed83841bb Merge branch 'master' into async 2026-03-13 09:26:20 -04:00
Jarrod Johnson 550dfbf6a0 Fix reference of inputdata in remoteconfig 2026-03-13 09:26:11 -04:00
Jarrod Johnson 89cc70260f Merge branch 'master' into async 2026-03-13 08:59:15 -04:00
Jarrod Johnson e0951b11a6 Fix filename typo 2026-03-13 08:58:58 -04:00
Jarrod Johnson 794502eb6a Merge branch 'master' into async 2026-03-09 17:30:04 -04:00
Jarrod Johnson 1a87701fee Fix ansible running
Have results available as they happen

change away from stdout, to avoid being stepped on by ansible modules that print to that
2026-03-09 16:48:42 -04:00
Jarrod Johnson e185f2224f Implement ability for user to kick off confluent ansible runs
Add nodeapply -A and associated API.

This permits orchestrating plays without touching the nodes directly by the user.
2026-03-06 16:24:26 -05:00
Jarrod Johnson a4510ae58d Fix ordering of width/height geometry 2026-03-04 16:54:27 -05:00
Jarrod Johnson a7c5b2478c Get screenshots working in redfish asyncio 2026-03-04 16:04:51 -05:00
Jarrod Johnson 14035fce88 Implement fallback for screen geometry
Ideally, we can do TIOCGWINSZ.

Unfortunately, in some cases this breaks, resort to
escape codes.
2026-03-04 16:04:08 -05:00
Jarrod Johnson 42bfde3f86 Restore VNC console handling to async branch 2026-03-04 15:13:30 -05:00
Jarrod Johnson 2c75571a84 Fixes to allow a test deployment to complete under asyncio 2026-03-04 10:47:33 -05:00
Jarrod Johnson a5e7fe93e4 Fix ip address list 2026-03-03 17:12:28 -05:00
Jarrod Johnson bfb41de43b Further asyncio conversion work 2026-03-03 16:06:14 -05:00
Jarrod Johnson c44cdc21ea Fix async definition of decode_alert 2026-03-03 14:53:54 -05:00
Jarrod Johnson c806bf2234 Rework more of ipmi support for async 2026-03-03 14:49:21 -05:00
Jarrod Johnson 61ada4d3d4 Advance async rework of ipmi 2026-03-03 13:05:10 -05:00
Jarrod Johnson f10a173e62 Address numerous reworks of async 2026-03-03 12:50:32 -05:00
Jarrod Johnson 27a3a446fe Fix issues in nodeconfig async 2026-03-03 12:42:29 -05:00
Jarrod Johnson 9abe1f98b7 Numerous await changes for ipmi 2026-03-03 12:42:17 -05:00
Jarrod Johnson 1fd986cc9a async fixes for ipmi 2026-03-02 16:43:12 -05:00
Jarrod Johnson f97c481c62 Normalize various iterables for single node responses 2026-03-02 13:01:28 -05:00
Jarrod Johnson d50fa2b587 Handle some async conversions 2026-03-02 12:17:35 -05:00
Jarrod Johnson a31532d8e4 Fix transmit of notfound errors 2026-03-02 08:41:18 -05:00
Jarrod Johnson c8fcb716fb Fix socket being in blocking mode before async calls. 2026-03-02 08:33:34 -05:00
Jarrod Johnson 289a6a6af8 Adjustments to purge eventlet references in deltapdu 2026-03-02 08:30:12 -05:00
Jarrod Johnson 6379b03051 Purge eventlet from plugins 2026-03-01 13:42:36 -05:00
Jarrod Johnson 580dd283cc Fix some async gaps 2026-02-27 13:37:32 -05:00
Jarrod Johnson 9d85a3c993 Remove eventlet from dependencies 2026-02-27 13:32:08 -05:00
Jarrod Johnson 3636e14628 Rework netutil for async and dependent functions 2026-02-26 17:27:04 -05:00
Jarrod Johnson 84f3614f0a Convert vcenter console to async 2026-02-26 13:16:17 -05:00
Jarrod Johnson 97f1545f97 Initial wave of vcenter async conversion 2026-02-26 12:57:15 -05:00
Jarrod Johnson e6bcf3cf9a Fix proxmox console for async operation 2026-02-26 10:49:33 -05:00
Jarrod Johnson 61c063adf4 Draft of converting tsmsol to asyncio 2026-02-25 16:19:05 -05:00
Jarrod Johnson 36898ba570 Fix openbmc console method with async 2026-02-25 16:14:51 -05:00
Jarrod Johnson 2f2afd970a Merge branch 'master' into async 2026-02-25 10:02:05 -05:00
Jarrod Johnson 69beaad3c9 Induce more versions of openssh to do the proper thing 2026-02-23 15:07:19 -05:00
Jarrod Johnson 74dda48513 Provide helper script for setting up nokia switches 2026-02-23 10:15:55 -05:00
Jarrod Johnson f2de24015b Port mdns to the async architecture similar to ssdp 2026-02-20 10:17:24 -05:00
Jarrod Johnson 9b2eaa90b5 Remove a number of eventlet imports/comments 2026-02-20 10:17:10 -05:00
Jarrod Johnson 16ff57dcfc Fixes for macmap in async 2026-02-20 09:43:34 -05:00
Jarrod Johnson 7adea90169 Fixes for async snmp and SRLinux 2026-02-20 09:29:03 -05:00
Jarrod Johnson f28e6d32f3 Handle NXOS async and advance state of lldp/mac map with async 2026-02-20 08:58:44 -05:00
Jarrod Johnson 09089befc8 SRLinux async handling advancement 2026-02-19 16:34:36 -05:00
Jarrod Johnson 4173356a70 Fixes for srlinux asyncio 2026-02-19 16:25:37 -05:00
Jarrod Johnson ca0c89aa07 Move SRLinux support to asyncio style 2026-02-19 16:08:32 -05:00
Jarrod Johnson ce2487dcb8 Fix import of srlinux 2026-02-19 14:58:31 -05:00
Jarrod Johnson 5b6bd4fbe1 Fix async mistakes from merge 2026-02-19 14:53:28 -05:00
Jarrod Johnson 18d06409d4 Merge branch 'master' into async 2026-02-18 16:54:46 -05:00
Jarrod Johnson 08b2e1d008 Wire up FDB and LLDP for srlinux 2026-02-18 16:53:12 -05:00
Jarrod Johnson 582842aec8 Add mac and lldp retrieval for SRLinux 2026-02-18 16:16:22 -05:00
Jarrod Johnson e0f00d80ed Merge branch 'master' into async 2026-02-17 16:19:51 -05:00
Jarrod Johnson 63307c331e Have nodesensors and nodehealth be more adaptive to partial server data. 2026-02-17 16:19:41 -05:00
Jarrod Johnson 318608cde3 Add draft SRLinux support
Wire up the non-networking facets of Nokia SR Linux support.

Provide stubs for LLDP and FDB
2026-02-17 16:13:43 -05:00
Jarrod Johnson 7efbec7ea2 Merge branch 'master' into async 2026-02-11 11:35:31 -05:00
Jarrod Johnson ef7d2414ad Update nodeconfig usage material 2026-02-11 11:35:09 -05:00
Jarrod Johnson 21867a60d2 Merge branch 'master' into async 2026-02-11 10:55:07 -05:00
Jarrod Johnson 722a0b874a Add notation about certificate and nodemedia 2026-02-11 10:54:30 -05:00
Jarrod Johnson f77c1e9333 Merge branch 'master' into async 2026-02-10 17:10:38 -05:00
Jarrod Johnson 1deb76989e Recognize 1a/2b style enclosure bay in discovery 2026-02-10 17:10:18 -05:00
Jarrod Johnson 9e5c69c286 Merge branch 'master' into async 2026-02-09 13:19:50 -05:00
Jarrod Johnson 480d399f44 Add missing switch member of info with NX switches 2026-02-09 13:17:45 -05:00
Jarrod Johnson 07369667f7 Become incompatible with pysnmp 7.1.16
The EPEL version of pysnmp is broken, block it from dependecies
2026-02-06 15:13:46 -05:00
Jarrod Johnson 25ea00d6d1 Merge branch 'master' into async 2026-02-05 07:58:15 -05:00
Jarrod Johnson e1d4b72f32 Be less picky about megarac url
megarac implementations consistently indicate an .xml file, but wildly vary on what it may be.

Broaden recognition.
2026-02-05 07:57:25 -05:00
Jarrod Johnson 137c3a0688 More async changes for confluent 2026-02-04 16:12:36 -05:00
Jarrod Johnson 00136f61fe More async fixes to remote media related redfish 2026-02-04 15:44:04 -05:00
Jarrod Johnson f9e898a46a Asyncio fixes
Fix ability to receive file descriptions from a unix domain client

Correct invocations to clearbuffer in consoleserver
2026-02-04 15:34:38 -05:00
Jarrod Johnson 06c2299b10 Adjust redfish fetch of diagnostic data 2026-02-04 10:19:13 -05:00
Jarrod Johnson a003ad6e2c Address asyncio changes for consoleserver 2026-02-04 10:02:17 -05:00
Jarrod Johnson d5c85cdff9 Fix async bind handling 2026-02-04 09:51:10 -05:00
Jarrod Johnson cdc668d717 Fix async assumption about list_updates
Turns out that the firmwaremanagemer methods will
generally not be async after all.
2026-02-03 16:43:25 -05:00
Jarrod Johnson 9ea971d9df Begin work to rework firmwaremanager for async
Requires a pool concept to manage concurrent task execution to match
previous expectations.
2026-02-03 16:36:41 -05:00
Jarrod Johnson d2a13f93f3 Fix error handling flow for async in redfish 2026-02-03 11:49:18 -05:00
Jarrod Johnson 850793a73f Merge branch 'master' into async 2026-02-03 07:58:33 -05:00
Jarrod Johnson 86783a2f12 Fix uninitialized privacy_protocol variable 2026-02-03 07:58:07 -05:00
Jarrod Johnson 29ec8e2b54 Merge branch 'master' into async 2026-02-02 10:19:30 -05:00
Jarrod Johnson 99063eb049 Recognize variation in DeviceDescrption.json to see SMM3 2026-02-02 10:17:32 -05:00
Jarrod Johnson c83ac717fe Wire up license management async wise 2026-02-02 10:09:33 -05:00
Jarrod Johnson 69e149f9c5 Adjust ipmi get_licenses for async 2026-01-30 16:58:56 -05:00
Jarrod Johnson b496f2c324 Wire up a number of async style calls 2026-01-29 15:18:46 -05:00
Jarrod Johnson e75a1dc7ad Gracefully accept loop cancellation in async 2026-01-28 16:22:53 -05:00
Jarrod Johnson 5c6fb7f7ef Bring to current asyncio run best practices 2026-01-28 15:42:00 -05:00
Jarrod Johnson 2dcbf76738 More async rework of ipmi 2026-01-28 15:39:00 -05:00
Jarrod Johnson 04d2a5affc More async conversions 2026-01-28 15:32:54 -05:00
Jarrod Johnson 134f339050 Update some ipmi code for async 2026-01-28 15:06:53 -05:00
Jarrod Johnson b4b9a1d1ce Merge branch 'master' into async 2026-01-28 15:05:14 -05:00
Jarrod Johnson 0975bd9e62 Revert "Update some code for async"
This reverts commit 3058dd4141.
2026-01-28 15:04:49 -05:00
Jarrod Johnson 291363c582 Update some code for async 2026-01-28 15:04:32 -05:00
Jarrod Johnson 3058dd4141 Update some code for async 2026-01-28 14:49:58 -05:00
Jarrod Johnson c29494bcf6 Make all the redfish iterators async consistent 2026-01-23 20:48:16 -05:00
Jarrod Johnson 50ec0bbca6 Correct to async for in refish retriev 2026-01-23 20:46:57 -05:00
Jarrod Johnson 52bb240aff Wire up async mechanism in redfish 2026-01-23 20:45:44 -05:00
Jarrod Johnson 667e44983d Fix ordering of confluentbmcname setting 2026-01-23 20:38:29 -05:00
Jarrod Johnson b5771023c3 Fix confetty indentation 2026-01-23 20:33:00 -05:00
Jarrod Johnson 76efea7c44 Use new method of running async in nodeattrib 2026-01-23 20:32:21 -05:00
Jarrod Johnson b0647275df Replace dead references to SecureHTTPConnection 2026-01-23 20:23:23 -05:00
Jarrod Johnson c4616745c4 Remove pyghmi usage across multiple areas 2026-01-23 13:31:09 -05:00
Jarrod Johnson 60c3d5400a Fix up proxmox module for async operation 2026-01-23 10:30:01 -05:00
Jarrod Johnson 6bc9282698 Change to await login 2026-01-22 14:59:31 -05:00
Jarrod Johnson b06ffb293a Asyncify proxmox retrieve function 2026-01-22 14:53:52 -05:00
Jarrod Johnson b548002a8d Fix nodegroup attribute async behavior 2026-01-22 14:50:59 -05:00
Jarrod Johnson 218ecce63f Correct import name 2026-01-22 14:46:34 -05:00
Jarrod Johnson 1c679727ad Correct issues in recent revision 2026-01-22 14:44:36 -05:00
Jarrod Johnson 50e530ebde Replace pyghmi with aiohmi in various plugins, remove some eventlet usage 2026-01-22 14:40:44 -05:00
Jarrod Johnson 7984c02042 Temporarily remove eficompressor dependency 2026-01-22 09:42:05 -05:00
Jarrod Johnson d338f8d586 Temporarily lift some rpm dependencies to work through dev 2026-01-22 09:26:45 -05:00
Jarrod Johnson c0d53ba986 Clean up RPM dependencies for async branch 2026-01-22 09:26:01 -05:00
Jarrod Johnson 68097428a5 Modernize asyncio invocation in main confluent runtime 2026-01-21 16:47:17 -05:00
Jarrod Johnson 7fedbc1810 Replace some pyghmi references and modernize some asyncio invocations 2026-01-21 16:45:42 -05:00
Jarrod Johnson b2f1b8da79 Add tasks management module for async 2026-01-21 16:23:31 -05:00
Jarrod Johnson 21c9158491 Carry forward some dns attributes into a bond 2026-01-21 15:12:23 -05:00
Jarrod Johnson 54735e9857 Carry forward some dns attributes into a bond 2026-01-21 15:11:45 -05:00
Jarrod Johnson 0dabccaec8 Corrections after some mistakes in the merge 2026-01-20 14:55:06 -05:00
Jarrod Johnson d89305ca42 Merge branch 'master' into async
Try to merge in 2025 work into async
2026-01-20 14:24:01 -05:00
Jarrod Johnson e6c19388a2 Add device-manager to container build
Confluent needs device-mapper for imgutil operation
2026-01-16 08:45:12 -05:00
Jarrod Johnson 048780e16d Explicitly mknodes for pack/unpack
In some contexts, udev may be asleep
at the wheel. Explictly have dmsetup
refresh the devnodes.
2026-01-15 15:15:11 -05:00
Jarrod Johnson 61d7a49163 Revert "Fallback to filename for PE format kernels"
This reverts commit a0a5887214.
2026-01-15 14:29:31 -05:00
Jarrod Johnson f8b8ce3847 Fallback to filename for PE format kernels
Some ARM64 kernels ship as EFI executables, but it's
not obvious how to extract version numbers from those properly.
2026-01-15 14:29:23 -05:00
Jarrod Johnson a0a5887214 Fallback to filename for PE format kernels
Some ARM64 kernels ship as EFI executables, but it's
not obvious how to extract version numbers from those properly.
2026-01-15 13:27:21 -05:00
Jarrod Johnson ccaf22f44f Add architecture handling in pkglist
To handle amd64/arm64 profiles, have the pkglist allow for architecture specific qualifiers.

Additionally, soften failure to accomplish selinux changes.
2026-01-15 12:52:07 -05:00
Jarrod Johnson 72c4868073 Update container with more packages, volumes, env, and alma 10 2026-01-15 09:46:23 -05:00
Jarrod Johnson afb6356f9d Change ownership
Container runs as internal 'root' user for now
2026-01-14 16:29:31 -05:00
Jarrod Johnson 6e6ac67b3d Provide some build assets
Provide some dockerfiles for creating build containers
2026-01-13 13:57:37 -05:00
Jarrod Johnson 99d10896e8 Fix parameter count unpack for accelerated switch interrogation 2026-01-08 17:07:39 -05:00
Jarrod Johnson 488f23e3ed Fix spelling of rpmbuild 2026-01-06 15:55:36 -05:00
Jarrod Johnson 6ca62cbb35 Provide optional output directory 2026-01-06 15:54:46 -05:00
Jarrod Johnson 45bc9788b4 Correct mistake in SPECS spelling 2026-01-06 15:51:40 -05:00
Jarrod Johnson 289c31e7ac Ensure in expected directory to start 2026-01-06 15:51:06 -05:00
Jarrod Johnson 1a684f2012 Ensure rpmbuild directory exists before building 2026-01-06 15:49:50 -05:00
Jarrod Johnson a4229fc58d Change name to index in apiclient
confignet was using the index for ipv4
2025-12-12 11:18:33 -05:00
Jarrod Johnson 31c1a865dc Update confignet to match apiclient changes 2025-12-12 09:30:56 -05:00
Jarrod Johnson ff84fcf6e9 Merge branch '3.14' 2025-12-11 13:21:33 -05:00
Jarrod Johnson 56dfb6dc6b Fix spelling issue in man page 2025-12-11 08:46:59 -05:00
Jarrod Johnson d7577a04a7 Fix ESXi compatibility of apiclient
apiclient was using Linux specific network  information.

Change to libc getifaddrs for better cross-platform compatibility.
2025-12-11 08:46:19 -05:00
Jarrod Johnson b72d6c9cfc Fix typo 2025-12-10 14:14:14 -05:00
Jarrod Johnson 523c93dfc3 Tolerate more network circumstances in bluefield deploy
If the networking didn't come up well, the 'functions' routines would not be able to handle.

Switch to using apiclient which is designed specifically to handle less cooperative
initial network conditions.
2025-12-09 08:49:27 -05:00
Jarrod Johnson 04e983a2d3 Handle broader memory information being returned from confluent 2025-12-04 09:52:15 -05:00
Jarrod Johnson 2464e0ff4f Fix location of the apiclient common resource 2025-12-02 14:35:50 -05:00
Jarrod Johnson c196bf9d55 Fix initial startup of a new confluent
The indexes change failed on a brand new install.
2025-12-02 14:31:10 -05:00
Jarrod Johnson 12d886a4f6 Add more imgutil documentation 2025-11-25 13:19:03 -05:00
Jarrod Johnson 6a26ece782 Merge remote-tracking branch 'xcat/master' 2025-11-25 11:59:43 -05:00
Jarrod Johnson 3cbac38d57 Also autoconsole when exactly one serial port is detected at all. 2025-11-25 11:53:50 -05:00
Jarrod Johnson 224f349053 Extend autocons to more use cases
If SPCR comes up blank, see if there is one and exactly one serial with carrier detect

Failing that, give DMI a chance to indicate a preference, for now just SuperMicro, since they have the most
inconsistent carrier detect behavior
but almost always consider ttyS1 to be the answer.
2025-11-25 11:51:07 -05:00
Jarrod Johnson 9d361d376d Merge pull request #203 from Obihoernchen/bond_desc
Add bond alias to team description
2025-11-21 09:47:09 -05:00
Markus Hilger ec39de3df0 Add bond alias to team description 2025-11-21 14:16:07 +01:00
Jarrod Johnson a3b768c70f Draft bluefield deploymeent facilities 2025-11-20 16:44:24 -05:00
Jarrod Johnson 4f75d4942b Modify adoption process:
Restore useinsecureprotocols if set directly on node

Switch from pxe-style to identity-file based node api token for hardened node authentication
2025-11-20 16:05:22 -05:00
Jarrod Johnson 4d2f36917c Restore useinsecureprotocols after adopt 2025-11-20 15:49:51 -05:00
Jarrod Johnson a2a50d34d1 Merge remote-tracking branch 'xcat' 2025-11-19 15:38:01 -05:00
Jarrod Johnson 041008a524 Remove redundant el10 initramfs fixup 2025-11-19 15:37:29 -05:00
Jarrod Johnson 5923feaa18 Merge pull request #202 from Obihoernchen/custom
Add documentation for custom nodeattribs
2025-11-19 07:47:46 -05:00
Jarrod Johnson 73216fc062 Fix architecture name mismatch
Confluent went with aarch64 consistent
with EL naming, but Ubuntu used
debian naming, recognize and just
handle that.
2025-11-18 09:10:30 -05:00
Jarrod Johnson 100944490c Fix potentially uninitialized curridx 2025-11-17 15:07:17 -05:00
Jarrod Johnson 61b07e0af4 Start index at 1 instead of 0 2025-11-17 12:05:03 -05:00
Jarrod Johnson 53760ab5dd Attribute feature enhancement
Add expression functions upper, lower, block_number, and block_offset.

Add an 'id.index' auto-attribute to
yield a number for nodes.
2025-11-17 11:58:04 -05:00
Jarrod Johnson d3e7a49f92 Simplify by recursion
Use _handle_ast_node to process
everything before the function name in an Attribute call
2025-11-15 10:32:11 -05:00
Jarrod Johnson 1f688ead28 Implement .replace() for attribute expressions
Provide an easy to use replace() to allow removing or substiting values
during expression evaluation.
2025-11-14 17:20:06 -05:00
Jarrod Johnson d20c5ac6eb Move handling of the loop directio straight to onboot
There were difficulties in the devfs after
boot, just let the full system handle it.
2025-11-13 15:33:04 -05:00
Jarrod Johnson 4484216198 Fix issues with the tethered memory optimizations 2025-11-13 15:24:26 -05:00
Jarrod Johnson e1efd6a9c5 Implement new 'uncompressed' image method
This allows the FS to just live, uncompressed, in cache.

This is generally a bad idea, however:

- In a hypothetically super-tuned diskless image, the lack of double-cache can offset the lack of compression
- The image will have supreme read performance
- It will have the most deterministic memory behavior
2025-11-13 14:39:53 -05:00
Jarrod Johnson 58d5209595 Port tethered improvments to EL8 2025-11-13 14:35:18 -05:00
Jarrod Johnson 53c918042a Remove double-caching in tethered diskless
By default, the squashfs file was being cached as well as the contents after extraction.

This is superfluous pressure on the cache of the OS.

However, it does help keep the image afloat through 'confignet', so
leave it on until onboot completes, then reclaim cache and disable further caching.
2025-11-13 14:28:25 -05:00
Markus Hilger 9148a841b5 Add documentation for custom nodeattribs 2025-11-13 00:45:53 +01:00
Jarrod Johnson 6ebb6de107 Allow specifiying SNMP privacy protocol
Modern SNMP devices may require AES.

Unfortunately, older ones may refuse AES.

For compatibility, continue to default to DES, but
allow AES to be indicated in attributes.
2025-11-10 10:21:01 -05:00
Jarrod Johnson 20292cdfd0 Do not let diskless.conf persist into EL9 diskless images
It fouls run of kdump building the kdump image.
2025-11-07 13:22:21 -05:00
Jarrod Johnson b07da455c2 Fix SAN generation
The nameconstraint support missed
a branch, fix this.
2025-11-07 11:22:12 -05:00
Jarrod Johnson cc9a81103b Do not autosign if the corresponding cryptography is unavailable
We use cryptography verification, but it's relatively new.

For compatibility, we fall back to fingerprint only.

This is pretty bad when inflicted on
unsuspecting users on autosign,
so skip autosign if cert validation
would break.
2025-11-04 15:51:22 -05:00
Jarrod Johnson 21155d2091 Bring untethered changes to el10 diskless 2025-11-04 11:17:28 -05:00
Jarrod Johnson 6c0d7ea60e Simplify end untethered el9 diskless environment
Rather than treat both as the same, since untethered has everything up front anyway, go ahead and extract the filesystem.

This makes the mount look more straightforward and makes it so deletion of files from
the image also frees ram.
2025-11-04 11:14:52 -05:00
Jarrod Johnson 174d204607 Implement compatibility with newer pysnmp
For now, terminate the async nature
if newer pysnmp is detected.
2025-11-04 09:58:11 -05:00
Jarrod Johnson 2826abb7ab Prune excessive leftover ext config files 2025-11-03 14:21:36 -05:00
Jarrod Johnson 5adb5fa780 Automatically sign XCC certificates on discover
If an XCC doesn't have a 'real' certificate, sign it with the confluent
CA for 47 days.
2025-11-03 14:02:33 -05:00
Jarrod Johnson 5de063212f Prepare for supporting constrained CA
If asked to sign using a name constrained CA,
avoid generating a certificate that
would violate those constraints.
2025-11-03 10:43:34 -05:00
Jarrod Johnson 073f6d1389 Wire up cert signing to nodecertutil 2025-10-31 12:04:27 -04:00
Jarrod Johnson f755ba9f91 Implement method to sign BMC certificates 2025-10-31 10:46:42 -04:00
Jarrod Johnson cf8c01ef13 Merge remote-tracking branch 'lenovo' 2025-10-31 09:48:05 -04:00
Jarrod Johnson 8b12047ae0 Update to handle newer XCC2 firmware 2025-10-31 09:45:59 -04:00
Jarrod Johnson f0a779764d Fix ordering of digest argument
The digest argument was erroneously inserted between startdate and it's
argument, correct this mistake.
2025-10-28 15:39:04 -04:00
Jarrod Johnson 0ad7e99efe Only optionally use cryptography verification
Some supported distributions can't run the newer cryptography.

Make it a feature that only works with newer platforms.
2025-10-27 08:38:14 -04:00
Jarrod Johnson 24a76612ae Use sha284 hash algorithm
Some implementations reject sha256 as inadequate if ecdsa has 384 bit keylength. Bring the digest up to match
the key size for the ECDSA.
2025-10-27 06:41:05 -04:00
Jarrod Johnson 6c9c58f464 Update certutil to prepare for broader usage
For one, apply more rules from CA/B forum. This includes including KU and EKU extensions, marking basicConstraints critical, and
randomized serial numbers.

Also make the backdate and end date configurable, to allow
for the BMC certs to have a more palatable validity interval.
2025-10-26 14:57:26 -04:00
Jarrod Johnson 3125f4171b Begin overhaul of TLS cert management
Begin expanding certutil to sign other certificates from external CSRs more easily.

Have certutil make the CA constraint critical.

Have the fingerprint based validator have a mechanism to check for properly signed certificate in lieu of exact match,
and update the stored fingerprint
on match.

Provide a means to request a custom subject when evaluating a
target.

Change redfish plugin to set that subject in the verifier.
2025-10-24 20:02:51 -04:00
Jarrod Johnson d66df7ee4b Merge branch 'master' into async
Need to rework httpapi further for changes to the firmware staging.
2025-01-08 14:23:52 -05:00
Jarrod Johnson feaa3bb7b4 Rework vinzmanager for async operation 2024-09-10 11:26:18 -04:00
Jarrod Johnson dac383af59 Merge branch 'master' into async 2024-09-10 09:51:28 -04:00
Jarrod Johnson f50db78c8c Merge branch 'master' into async 2024-08-31 07:30:47 -04:00
Jarrod Johnson 8cb34b20bc Fix duplicate lines from merge 2024-08-28 19:21:34 -04:00
Jarrod Johnson 75ae623b70 Merge branch 'master' into async 2024-08-28 19:20:21 -04:00
Jarrod Johnson 55cdfae437 Fix different invocations of check_fish
Particularly nodediscover register can fail.

Those invocations are XCC specific, so the targtype should not matter
in those cases.
2024-08-28 19:18:43 -04:00
Jarrod Johnson 4edc2a6412 Port forward Confluent 3.11 changes 2024-08-28 11:48:12 -04:00
Jarrod Johnson b46aecbeed Fix PXE afterm merge and have rebase work with async 2024-08-25 18:40:26 -04:00
Jarrod Johnson 69afc013f2 Merge branch 'master' into async 2024-08-23 18:26:33 -04:00
Jarrod Johnson 9c3126e9f7 Add client to asyncio pxe 2024-08-23 16:17:50 -04:00
Jarrod Johnson 0b401e8276 Merge branch 'master' into async 2024-08-23 15:51:04 -04:00
Jarrod Johnson b609a0039f Merge branch 'master' into async 2024-08-22 10:34:42 -04:00
Jarrod Johnson fe0a15faf2 Merge branch 'master' into async 2024-08-22 08:42:37 -04:00
Jarrod Johnson 10faac8835 Hook up descriptions to asyncio 2024-08-22 08:40:34 -04:00
Jarrod Johnson 19439463b1 Normalize non-http and http and http async and internal passthrough
Have the core provide normalization and use it across
places that need it.
2024-08-21 14:40:58 -04:00
Jarrod Johnson 1d861e60bb Refactor task management to its own module 2024-08-21 11:38:46 -04:00
Jarrod Johnson 52f172ef57 Merge branch 'master' into async 2024-08-21 09:56:43 -04:00
Jarrod Johnson a0ab71f7bb Fix call to check_fish with wrong number of args 2024-08-20 17:09:39 -04:00
Jarrod Johnson 4ef24351aa Do not await synchronous functions 2024-08-20 17:05:32 -04:00
Jarrod Johnson c9e428bb1b Change browserfs control to async, bring together to single send 2024-08-20 16:24:18 -04:00
Jarrod Johnson 53d0d09ae1 Have browserfs based import work with async 2024-08-20 15:57:56 -04:00
Jarrod Johnson 30b8979e2c Merge branch 'master' into async 2024-08-19 16:55:04 -04:00
Jarrod Johnson 6f776a657c Begin work on selfservice asyncio port
Have a deploycfg call be able to proceed through.
2024-08-16 17:06:49 -04:00
Jarrod Johnson 2f415caead Fix osdeploy updateboot with asyncio 2024-08-16 17:06:16 -04:00
Jarrod Johnson 708170b06a Convert affluent method from eventlet 2024-08-16 15:17:06 -04:00
Jarrod Johnson 5eaf998391 Remove greenlet, and change 'confluent' to asyncio 2024-08-16 14:36:59 -04:00
Jarrod Johnson c43a667299 Remove some debug output 2024-08-16 14:30:50 -04:00
Jarrod Johnson bf56d40fb5 Convert neighutil to asyncio 2024-08-16 14:29:04 -04:00
Jarrod Johnson fab6a5a757 Remove eventlet from log 2024-08-16 14:04:12 -04:00
Jarrod Johnson a076472718 Merge branch 'master' into async 2024-08-16 11:27:14 -04:00
Jarrod Johnson d1659cef97 Merge branch 'master' into async 2024-08-16 09:33:27 -04:00
Jarrod Johnson c0018840c7 Remove stale eventlet import from osimage 2024-08-15 16:40:40 -04:00
Jarrod Johnson 511fdfe6c1 Fix issues in the online debugger 2024-08-15 16:35:55 -04:00
Jarrod Johnson 4945c1f473 Replace 'backdoor' with 'debugger' 2024-08-15 16:03:02 -04:00
Jarrod Johnson 90b893bc28 Bring up ssh asyncio and fix other shell/console async 2024-08-15 14:45:31 -04:00
Jarrod Johnson ac4092ec4b More fixes for asyncio support console usage 2024-08-15 11:38:14 -04:00
Jarrod Johnson 556e40787c Have OpenBMC work with async changes 2024-08-15 11:28:08 -04:00
Jarrod Johnson 45187e0c54 Merge branch 'master' into async 2024-08-15 10:55:11 -04:00
Jarrod Johnson 2cc61a1810 Merge branch 'master' into async 2024-08-14 16:26:55 -04:00
Jarrod Johnson 21f68bb212 Apply formatting changes 2024-08-13 15:29:15 -04:00
Jarrod Johnson 45b17ba855 Get basic redfish running in asyncio 2024-08-13 15:28:52 -04:00
Jarrod Johnson db670b695f Reuse recent_peers
To be consistent, reuse this set rather than creating a new one.
2024-08-09 16:43:27 -04:00
Jarrod Johnson 0d8173cbcb Fix nodediscover clear
Nested async iteration of multiple confluent calls fail, break
it into two distinct queries.
2024-08-09 10:03:15 -04:00
Jarrod Johnson 8b70213c0d Merge branch 'master' into async 2024-08-09 07:56:30 -04:00
Jarrod Johnson e0fa642496 Fix SLP asyncio performance issue
SLP asyncio performance spent too much time tied up in futile
processing,
avoid duplicate deferpeers and simplify the loop iteration.
2024-08-08 17:06:57 -04:00
Jarrod Johnson 42e5a556c1 Make it easier to debug slow callback
Provide a name to create_task to make the
slow callback warning actually usable.
2024-08-08 16:25:04 -04:00
Jarrod Johnson d7d89dd233 Remove spurious import from redfish 2024-08-08 16:07:42 -04:00
Jarrod Johnson 536aa7f212 Fix XCC/XCC2 discovery for async branch 2024-08-08 14:25:47 -04:00
Jarrod Johnson 67a61c5012 Migrate XCC3 discovery to async 2024-08-08 12:45:39 -04:00
Jarrod Johnson 73cd6d52da Remove spurious reintroduction of select to slp 2024-08-08 11:05:51 -04:00
Jarrod Johnson 5be422958b Removed redundant definitions introduced by merge attempt 2024-08-08 10:10:16 -04:00
Jarrod Johnson c754dc2641 Merge branch 'master' into async 2024-08-08 09:45:15 -04:00
Jarrod Johnson 3741db740f Merge branch 'master' into async 2024-07-23 16:20:14 -04:00
Jarrod Johnson e446aa9277 Merge branch 'master' into async 2024-07-09 08:45:43 -04:00
Jarrod Johnson 1edfeba076 Add MegaRAC discovery support for recent MegaRAC
Create a generic redfish discovery and a MegaRAC specific
variant.

This should open the door for more generic common base redfish discovery
for vaguely compatible implementations.  For now, MegaRAC only
overrides the default username and password (which is undefined
in the redfish spec).

Also, have SSDP recognize the variant, and tolerate odd nonsense
like SSDP replies coming from all manner of odd port numbers (no
way to make a sane firewall rule to capture that odd behavior,
but at application level we have a chance).
2024-07-03 14:36:28 -04:00
Jarrod Johnson 2c2fe08d66 Merge branch 'megaracdisco' into async 2024-07-02 15:13:33 -04:00
Jarrod Johnson 362f6ae6d5 Merge branch 'master' into async 2024-07-02 15:13:26 -04:00
Jarrod Johnson 8fbb495ee9 Merge branch 'master' into async 2024-06-24 15:57:55 -04:00
Jarrod Johnson 879fb9c7ab Merge branch 'master' into async 2024-06-14 11:22:05 -04:00
Jarrod Johnson 9b8ec1e493 Merge branch 'master' into async 2024-06-14 11:16:34 -04:00
Jarrod Johnson 9394e83c81 Avoid pam blocking main thread execution
Use processpool to execute pam authentication,
avoiding a hang while waiting for child process.
2024-06-14 10:47:02 -04:00
Jarrod Johnson f42812b836 Fix console over shared websocket
This fixes console behavior in the webui
2024-06-13 16:57:20 -04:00
Jarrod Johnson b6a0250e5c Advance state of asyncio
Add a mechanism to close a session the right way
in tlvdata

Fix confluentdbutil/configmanager to restore/dump db to directory

Move auth to asyncio away from eventlet

Fix some issues with httpapi, enable reading body via aiohttp

Fix health from ipmi plugin

Fix user creation across a collective.
2024-06-13 16:32:02 -04:00
Jarrod Johnson bdb7f064d6 Rework a number of subprecess calls and osdeploy
Some subprocess calls were reworked to use asyncio friendly
variants.

Also, osdeploy initialize was checked, and reworked the ssh and tls
handling.

osdeploy import was also reworked to functional with async only.
2024-05-31 17:22:26 -04:00
Jarrod Johnson 85c8268ad8 Fix proxy console through collective in async 2024-05-30 16:14:39 -04:00
Jarrod Johnson 00eff4a002 Migrate IPMI SOL to asyncio 2024-05-30 15:37:06 -04:00
Jarrod Johnson cbb52739d3 Fix a number of issues with async rework
Have util retain tasks that are 'fire and forget', to avoid
garbage collection trying to delete the background tasks.

Move some utilities explicitly over to asynclient/asynctlvdata that
had previously been reworked.

Implement terminal resize in new asyncssh backend.
2024-05-30 13:59:14 -04:00
Jarrod Johnson 4ba82b7ef4 Merge branch 'master' into async 2024-05-30 09:29:24 -04:00
Jarrod Johnson c5405f832c Advance state of async shellserver
Can successfully run ssh sessions through
confluent with async now
2024-05-29 20:18:07 -04:00
Jarrod Johnson 23d0bbd047 Move nodediscover to async client
The work to convert had already been done, and it may be handy to make
nodediscover do some async tricks in the future.
2024-05-29 12:24:19 -04:00
Jarrod Johnson 4c3f93765f Have async and traditional client
Since a lot of the traditional client did not need async,
make life easier by just having them in parallel for now.

The server must use the async client, but the client applications can
stick with the somewhat more straightforward synchronous client.
2024-05-29 12:23:05 -04:00
Jarrod Johnson 4a2349d9ad Merge branch 'master' into async 2024-05-23 15:15:59 -04:00
Jarrod Johnson 1a9395fc5f Amend EL network bringup
One issue is that there are multiple networkmanager connections,
clean this up, though this seems not to be a functional issue.

However, sometimes the lldpad usage screws up network configuration,
disable the facility by forcibly disabling fcoe sincec that is what triggers lldpad.
wq
2024-05-22 15:44:05 -04:00
Jarrod Johnson b4ae6012c5 Remove eventlet from PXE support 2024-05-20 16:27:11 -04:00
Jarrod Johnson 782991aea3 Switch to asyncio usage of pysnmnp
This requires pysnm 6, the edition that should become the official one,
maintained by lextudio
2024-05-20 11:48:53 -04:00
Jarrod Johnson 6e751c811e Begin rework of macmap.py
Redo offload to asyncio subprocess, and
replace eventlet Events with futures for
messaging.
2024-05-17 17:07:18 -04:00
Jarrod Johnson c03aa728cc Properly detect killed leader
If leader closes connection, then have get_next_msg return None
as it did before.
2024-05-17 16:03:37 -04:00
Jarrod Johnson fbdb35e33d Merge branch 'master' into async 2024-05-16 15:42:22 -04:00
Jarrod Johnson 207cc3471e Fix closing sockets in various contexts
With asyncio, we must close the writer half of a pair

Also rework the get_next_msg to work better.

Still need to allow stop_following to interrupt get_next_msg
2024-05-16 15:40:43 -04:00
Jarrod Johnson 5a9f608451 Fix handling some eatonpdu return values 2024-05-15 12:30:13 -04:00
Jarrod Johnson 100810788c Fix media location search for EL8
EL8 distributions marked the 'OS' as dracut, workaround by trying to use PRETTY_NAME
2024-05-15 12:28:41 -04:00
Jarrod Johnson 90b90ade9c Remove disused iovec
iovec is no longer used due to migration from relevant
recvmsg ctypes call.
2024-05-09 09:49:56 -04:00
Jarrod Johnson f6fc539df9 Remove disused recvmsg ctypes wrapper
Since going to builtin python recvmsg, remove
the ctypes wrapper.
2024-05-09 09:48:11 -04:00
Jarrod Johnson 2e30f7fb86 Prune unneeded ctypes material from pxe
Moving to .recvmsg from python socket eliminates
most of the ctypes requirement. Still using it for sendto.
2024-05-09 09:46:01 -04:00
Jarrod Johnson e1e3244af6 Port PXE to asyncio and re-enable 2024-05-09 09:40:03 -04:00
Jarrod Johnson 5fd0cf2b0b Begin conversion of pxe to asyncio
Also convert to 'natural' recvmsg now that we are requiring
python high enough to have it.
2024-05-08 17:18:07 -04:00
Jarrod Johnson bd2f08d3ad Reactive SSDP in discovery core
Also, fix a getaddrinfo call to be async.
2024-05-08 13:18:38 -04:00
Jarrod Johnson b9a2c9a3ae Convert more XCC handling to asyncio 2024-05-08 13:18:05 -04:00
Jarrod Johnson 42b7cbe421 Implement SSDP asyncio
This covers SSDP devices as well as confluent deployment
discovery.
2024-05-08 13:17:43 -04:00
Jarrod Johnson 96a43013b5 Merge branch 'master' into async 2024-05-08 11:51:16 -04:00
Jarrod Johnson a3506cf0bf Correct misrouting in slp
IPv4 scan responses were lost as
the reader was passed IPv6 socket
no matter what.

Also, remove some needless verbosity.
2024-05-08 11:48:46 -04:00
Jarrod Johnson 25d4d13a96 Finish conversion of slp to asyncio.
Make process_peer async, with socket connection being async,
and dependency.

Have getaddrinfo use the asyncio version.

Rework the snoop to be more effective.

Rework the scan to be less convoluted.
2024-05-08 11:35:33 -04:00
Jarrod Johnson 23658680a5 Have slp mostly work
Advance the SLP discovery code and core discovery
to mostly work.
2024-05-07 17:02:51 -04:00
Jarrod Johnson 2089f5e7e6 Deal with normal generator from a plugin 2024-05-07 17:01:04 -04:00
Jarrod Johnson b3e0117944 Fix getpeername invocation in async 2024-05-07 17:00:43 -04:00
Jarrod Johnson 056a41c985 Fix client async invocations 2024-05-07 17:00:25 -04:00
Jarrod Johnson 6704f23218 Merge branch 'master' into async 2024-05-07 10:07:08 -04:00
Jarrod Johnson 222bdee851 Load firewall before esxi installation begins
Parts of esxi install depend on firewall running.  When
we are done with 'odd' networking, restore firewall
to meet that expectation.
2024-05-07 10:05:50 -04:00
Jarrod Johnson 5e222041bf Merge branch 'master' into async 2024-05-03 10:27:31 -04:00
Jarrod Johnson ee6f869cea Port utilities to asyncio, selfcheck and osdeploy
confluent_selfcheck removes eventlet dependency,

osdeploy reworked to use async methods to work with new client.
2024-04-30 14:30:01 -04:00
Jarrod Johnson b967c552fd Migrate intra-collective requests to asyncio
Update dispatch to be asyncio based, remove eventlet from core

Clean up some overly verbose print statements.
2024-04-30 13:56:00 -04:00
Jarrod Johnson 553916340e Advanced asyncio port progress
Offer a function in core to normalize plugin return.

A plugin might return an async generator, a traditional generator,
or might even return an awaitable wrapping a traditional generator.

Replace eventlet spawn with util spawn in discover core

Have node attribute update await the set_node_attributes appropriately
2024-04-30 10:44:43 -04:00
Jarrod Johnson 0be60b1ce2 Merge branch 'master' into async 2024-04-29 10:55:58 -04:00
Jarrod Johnson a5dc10debf Fix attribute synchronization
Specify a finite read to actually return from the buffer.

Convert some functions to async/await as appropriate.
2024-04-29 10:54:30 -04:00
Jarrod Johnson d2edcb62c6 Begin implementation of asyncio collective
The config synchronization is in progress.
2024-04-26 15:48:14 -04:00
Jarrod Johnson afa0c0df5a Merge branch 'master' into async 2024-04-22 14:36:42 -04:00
Jarrod Johnson 560ec60c12 Merge branch 'master' into async 2024-04-17 15:18:58 -04:00
Jarrod Johnson e890276bf6 Advance state of collective in asyncio
Eventlet is nominally removed from collective manager, however the join process still
needs to be reworked, and a lot more flows need to be adjusted.
2024-04-16 16:53:45 -04:00
Jarrod Johnson c24da59216 Merge branch 'master' into async 2024-04-16 10:39:20 -04:00
Jarrod Johnson e8110551db Port some of the collective management to asyncio 2024-04-15 17:19:27 -04:00
Jarrod Johnson bfe7529d21 Merge branch 'master' into async 2024-04-15 10:04:19 -04:00
Jarrod Johnson c3cafd9bf8 Purge eventlet and greenlet and long-polling support
Rather than try to support long deprecated http api behavior,
purge it for simpler code and remove eventlet/greenlet from the http
stack.
2024-04-11 09:09:02 -04:00
Jarrod Johnson fb8ac158cb Merge branch 'master' into async 2024-04-11 08:14:20 -04:00
Jarrod Johnson 9d828f0998 Merge branch 'master' into async 2024-04-09 13:37:09 -04:00
Jarrod Johnson e8fed28a21 Do not disarm until client notify done
In the unlikely event of a hiccup during the credserver connection,
defer the disarm until the server has transmitted success.
2024-04-09 10:28:21 -04:00
Jarrod Johnson 7b2e32009f Numerous async improvements
Restore 'as available' behavior to noderange over socket

Bring the httpapi to the point where the webui is able to start working,
notably bringing the asynchttp online with the websocket.

Fix a flaw in the async ipmi that would cause hangups.
2024-04-04 17:13:37 -04:00
Jarrod Johnson 587ccd13cc More work toward asyncio
aiohttp now covers a lot of httpapi GET, and some of websocket.
2024-04-03 16:58:40 -04:00
Jarrod Johnson 198ffb8be6 Advance asyncio port
Purge sockapi of remaining eventlet call

Extend asyncio into the credserver to finish out sockapi.

Have client and sockapi complete TLS connection including password checking

Fix confetty ability to 'create'.
2024-04-01 16:38:10 -04:00
Jarrod Johnson c1d680d8d8 Merge branch 'master' into async 2024-04-01 12:15:47 -04:00
Jarrod Johnson 1fbaee6149 Further move toward asyncio and reduce PyOpenSSL dep
Since we are rebasing to at least Python 3.6, and with
some extra ctypes wranging of the ssl context, we can likely
remove PyOpenSSL. Take first steps by removing it from 'sockapi'.

Have confluent executable become the 'top level' for eventlet, to allow
work on 'de-eventleting' on 'main.py'.

Rework tlvdata to deal with either a socket or a reader, writer tuple.
Using TLS with asyncio is easiest with the 'open_connection'
semantics, which force either a Protocol handler (callback based) or
dual streams.  While protocol approach ends with a more socket-like
'transport', the 'protocol' half is a bit unwieldy. So reader and writer
streams instead.
2024-03-29 16:23:45 -04:00
Jarrod Johnson 81428727d3 Merge branch 'master' into async 2024-03-27 14:28:48 -04:00
Jarrod Johnson 668c5af261 Bugfix and rework consoleserver a bit
Fix incorrect syntax in ssh.py, and correct direct asyncio
call of sock_recv when it must be called on the loop.
2024-03-25 15:18:05 -04:00
Jarrod Johnson b1cd7bcd98 Wire up 'configuration/system/all' in async way
This allows the fundamental API call to pass through
2024-03-25 15:16:56 -04:00
Jarrod Johnson 4fe9e1e80b Merge branch 'master' into async 2024-03-25 08:07:33 -04:00
Jarrod Johnson 508adc8d03 Merge master into asyncio 2024-03-22 15:52:04 -04:00
Jarrod Johnson 46edd8a49a Add a stub backdoor replacement 2024-03-15 17:11:12 -04:00
Jarrod Johnson 0570996c36 Merge branch 'master' into async 2024-03-15 15:51:08 -04:00
Jarrod Johnson ce3d4d7256 Merge branch 'master' into async 2024-03-15 13:04:01 -04:00
Jarrod Johnson 94cb1aebc3 Work on nodeconfig async conversion
Refactor nodeconfig to stand a chance at async.
2024-03-15 12:50:46 -04:00
Jarrod Johnson da63543a70 Advance the state of asyncio port 2024-03-15 12:50:04 -04:00
Jarrod Johnson 887207b9fc Merge branch 'master' into async 2024-03-15 12:31:51 -04:00
Jarrod Johnson 142f97c94e Merge branch 'master' into async 2024-03-15 09:58:06 -04:00
Jarrod Johnson 4ca82948ba SSH test by IP, to reflect actual usage and catch issues
One issue is modified ssh_known_hosts wildcard customization
failing to cover IP address.
2024-03-14 11:20:36 -04:00
Jarrod Johnson 399c1467c1 Remove redundant kill on the agent pid
Extraneous kill on the agent pid is removed.
2024-03-14 10:53:13 -04:00
Jarrod Johnson dcb6a1c759 Updates to confluent_selfcheck
Reap ssh-agent to avoid stale agents lying around.

Remove nuisance warnings about virbr0 when present.

Do a full runthrough as the confluent user to ssh to a node when user
requests with '-a', marking known_hosts and automation key issues.
2024-03-14 10:50:01 -04:00
Jarrod Johnson 91dc37d45e Fix nodeapply redoing a single node multiple times 2024-03-12 15:32:44 -04:00
Jarrod Johnson 500a955d79 Fix confetty tab completion with async
async required the async client to be wrapped in sync code.
2024-03-12 13:13:19 -04:00
Jarrod Johnson f4f5fcdb6d Fix lldp when peername is null
Some neighbors result in a null name, handle that.
2024-03-12 09:36:40 -04:00
Jarrod Johnson eed2e74bd0 Have image2disk delay exit on error
Debugging cloning is difficult when system immediately reboots on error.
2024-03-11 17:10:33 -04:00
Jarrod Johnson 4f92e3413a Expose fingerprinting and better error handling to osdeploy
This allows custom name and pre-import checking.
2024-03-11 13:32:45 -04:00
Jarrod Johnson d42e8e0921 Further asyncio port of confluent
Advance state of basic clients to advance testing and soon start doing
deeper activity.
2024-03-06 16:50:34 -05:00
Jarrod Johnson 635ef6073c Fix stray blank line at end of nodelist
Wrong indentation level for nodelist resulting in
spurious line.
2024-03-06 16:28:09 -05:00
Jarrod Johnson 496e7b4ef3 Properly address runansible error relay 2024-03-06 09:27:53 -05:00
Jarrod Johnson 3d33e33ea2 Dump stderr to client if ansible had an utterly disastrous condition 2024-03-06 08:45:23 -05:00
Jarrod Johnson 0a8ec96cdf Further progress toward asyncio
Basic operations can now happen with some async flows.
2024-03-04 16:18:55 -05:00
Jarrod Johnson 25f2698ae6 Opportunisticlly use sshd_config.d when detected 2024-03-04 08:06:01 -05:00
Jarrod Johnson d6bff637db Commence work on async 2024-02-23 11:56:07 -05:00
Jarrod Johnson 91e0aa938c Remove disused bufferlock
We no longer use a lock on buffer communication, eliminate
the stale variable.
2024-02-22 15:07:12 -05:00
Jarrod Johnson ed54bfa11a Change to unix domain for vtbuffer communication
The semaphore arbitrated single channel sharing
was proving to be too slow.  Make the communication
lockless by having dedicated sockets per request.
2024-02-22 15:05:56 -05:00
498 changed files with 53167 additions and 15673 deletions
+127
View File
@@ -0,0 +1,127 @@
name: CI
on:
push:
pull_request:
permissions:
contents: read
jobs:
shellcheck:
name: ShellCheck
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- name: Run ShellCheck (errors only)
# Check every tracked file that has a .sh extension or an sh/bash
# shebang. SC2148 (missing shebang) is excluded because many .sh
# files are sourced fragments or dracut hooks; ShellCheck then
# falls back to checking them as bash.
run: |
shebang_re='^#!.*[/ ](sh|bash|dash|ash|ksh)([[:blank:]]|$)'
{
git ls-files '*.sh'
git ls-files | while IFS= read -r f; do
[ -f "$f" ] || continue
firstline=
# An empty file makes read fail, which under -e would end the run.
IFS= read -r -n 200 firstline < "$f" 2>/dev/null || true
if [[ $firstline =~ $shebang_re ]]; then
printf '%s\n' "$f"
fi
done
} | sort -u > /tmp/shfiles
# A selection that quietly comes up empty would check nothing and
# still pass, so say how many files there are and insist on some.
echo "$(wc -l < /tmp/shfiles) shell files"
[ -s /tmp/shfiles ] || { echo '::error::No shell files found'; exit 1; }
xargs -d '\n' shellcheck --severity=error --exclude=SC2148 < /tmp/shfiles
ruff:
name: Ruff
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
# Pinned to an exact release: unlike actions/checkout, ruff-action
# publishes no moving major tag past v3, so @v4 does not resolve.
- uses: astral-sh/ruff-action@v4.1.0
with:
version: latest
# Rule selection, file discovery (the many extensionless Python
# executables) and exclusions all live in ruff.toml, so the whole
# workspace can be handed over as-is.
args: check --output-format=github
pyrefly:
name: Pyrefly
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: facebook/pyrefly@main
python-compileall:
name: Python compileall
runs-on: ubuntu-latest
env:
# One entry per Python version shipped by the distros confluent
# targets, limited to versions actions/setup-python still provides
# on current runners (sles15/alma8 ship 3.6, which is unavailable).
# Newline-separated so it feeds both setup-python (multiline input)
# and the shell loop below (word-split on whitespace).
PYTHON_VERSIONS: |
3.8
3.9
3.10
3.12
3.13
3.14
steps:
- uses: actions/checkout@v7
- uses: actions/setup-python@v6
with:
python-version: ${{ env.PYTHON_VERSIONS }}
- name: List the Python files without a .py name
# compileall only ever compiles *.py: handed anything else, even by
# name, it skips it and still exits 0. That leaves every extensionless
# CLI tool, deploy script and setup.py.tmpl unchecked, so collect them
# here and feed them to py_compile, which does compile what it is
# given. The first line is read with the shell builtin rather than
# forking head and grep per file.
run: |
shebang_re='^#!.*python'
{
git ls-files | while IFS= read -r f; do
[ -f "$f" ] || continue
case "$f" in *.py) continue ;; esac
firstline=
# An empty file makes read fail, which under -e would end the run.
IFS= read -r -n 200 firstline < "$f" 2>/dev/null || true
if [[ $firstline =~ $shebang_re ]]; then
printf '%s\n' "$f"
fi
done
# Python that carries no shebang at all, so nothing can detect it.
# Kept in step with extend-include in ruff.toml.
git ls-files '*/setup.py.tmpl' '*/scripts/configbmc' \
'*/scripts/add_local_repositories' 'misc/filterpasswd'
} | sort -u > /tmp/pyfiles
# An empty list would leave py_compile with nothing to do and the
# job green, which is the very hole this step exists to close.
echo "$(wc -l < /tmp/pyfiles) files without a .py name"
[ -s /tmp/pyfiles ] || { echo '::error::No such files found'; exit 1; }
- name: Compile all Python files
run: |
rc=0
for v in $PYTHON_VERSIONS; do
echo "::group::Python $v"
ok=0
"python$v" -W error -m compileall -q -x '/\.git/' . || ok=1
xargs -d '\n' "python$v" -W error -m py_compile < /tmp/pyfiles || ok=1
echo "::endgroup::"
if [ "$ok" -ne 0 ]; then
echo "::error::Python $v compileall failed"
rc=1
fi
done
exit "$rc"
+19
View File
@@ -0,0 +1,19 @@
name: Create Release
on:
push:
tags:
- '*'
permissions:
contents: write
jobs:
release:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- name: Create Release
uses: softprops/action-gh-release@v3
with:
generate_release_notes: true
+25 -5
View File
@@ -5,11 +5,10 @@
Confluent is a software package to handle essential bootstrap and operation of scale-out server configurations.
It supports stateful and stateless deployments for various operating systems.
Check [this page](https://hpc.lenovo.com/users/documentation/whatisconfluent.html
) for a more detailed list of features.
Check [this page](https://xcat2.github.io/confluent-docs/) for a more detailed list of features.
Confluent is the modern successor of [xCAT](https://github.com/xcat2/xcat-core).
If you're coming from xCAT, check out [this comparison](https://hpc.lenovo.com/users/documentation/confluentvxcat.html).
If you're coming from xCAT, check out [this comparison](https://xcat2.github.io/confluent-docs/miscellaneous/confluentvxcat/).
# Documentation
@@ -17,9 +16,9 @@ Confluent documentation is hosted on: https://xcat2.github.io/confluent-docs/
# Download
Get the latest version from: https://hpc.lenovo.com/users/downloads/
Get the latest version from: https://xcat2.github.io/confluent-docs/downloads/
Check release notes on: https://hpc.lenovo.com/users/news/
Check release notes on: https://xcat2.github.io/confluent-docs/release_notes/
# Open Source License
@@ -28,3 +27,24 @@ Confluent is made available under the Apache 2.0 license: https://opensource.org
# Developers
Want to help? Submit a [Pull Request](https://github.com/xcat2/confluent/pulls).
## Versioning and releases
The top-level `VERSION` file names the release the current branch is working toward. Build scripts
call `./mkversion`, which turns it into the version stamped on packages: the tag itself on a release
tag (`4.0.1`), otherwise a development version (`4.0.1.dev5+gdeadbee`, or `4.0.1~dev5+gdeadbee` for
the packages that have no `setup.py`).
`mkversion` also derives a version from the newest tag reachable from HEAD and uses whichever is
higher, so forgetting to bump `VERSION` after tagging the current branch cannot walk the version
backwards. Staying ahead of tags on *other* branches is what `VERSION` itself is for: patch releases
are tagged on release branches, which master never sees.
Cutting a new series:
1. Land the release commit on master and tag `X.Y.0` there.
2. Create branch `X.Y` from that tag; it inherits `VERSION=X.Y.0`.
3. Only then bump `VERSION` on master. Doing it before step 2 leaves the release branch on the
wrong series.
Patch releases land on branch `X.Y` and are tagged `X.Y.Z` there; no `VERSION` edit is needed.
+1
View File
@@ -0,0 +1 @@
4.0.0
+8
View File
@@ -0,0 +1,8 @@
FROM almalinux:10
RUN ["yum", "-y","update"]
RUN ["yum", "-y","install","gcc","make","rpm-build","python3-devel","python3-setuptools","createrepo","python3", "perl", "perl-DBI", "perl-JSON", "perl-XML-LibXML", "pinentry-tty", "rpm-sign", "git", "golang"]
ADD rpmmacro /root/.rpmmacros
ADD buildpackages.sh /bin/
#VOLUME ["/rpms", "/srpms"]
CMD ["/bin/bash","/bin/buildpackages.sh"]
+6
View File
@@ -0,0 +1,6 @@
for package in /srpms/*; do
rpmbuild --rebuild $package
done
find ~/rpmbuild/RPMS -type f -exec cp {} /rpms/ \;
+3
View File
@@ -0,0 +1,3 @@
%_gpg_digest_algo sha256
%_gpg_name Lenovo Scalable Infrastructure
+8
View File
@@ -0,0 +1,8 @@
FROM almalinux:8
RUN ["yum", "-y","update"]
RUN ["yum", "-y","install","gcc","make","rpm-build","python3-devel","python3-setuptools","createrepo","python3", "perl", "perl-DBI", "perl-JSON", "perl-Net-DNS", "perl-DB_File", "perl-XML-LibXML", "rpm-sign", "git", "fuse-devel","libcurl-devel"]
ADD rpmmacro /root/.rpmmacros
ADD buildpackages.sh /bin/
#VOLUME ["/rpms", "/srpms"]
CMD ["/bin/bash","/bin/buildpackages.sh"]
+6
View File
@@ -0,0 +1,6 @@
#!/bin/bash
for package in /srpms/*; do
rpmbuild --rebuild $package
done
find ~/rpmbuild/RPMS -type f -exec cp {} /rpms/ \;
+2
View File
@@ -0,0 +1,2 @@
%_gpg_digest_algo sha256
%_gpg_name Lenovo Scalable Infrastructure
+10
View File
@@ -0,0 +1,10 @@
FROM almalinux:9
RUN ["yum", "-y","update"]
RUN ["yum", "-y","install","gcc","make","rpm-build","python3-devel","python3-setuptools","createrepo","python3", "perl", "perl-DBI", "perl-JSON", "perl-Net-DNS", "perl-DB_File", "perl-XML-LibXML", "pinentry-tty", "rpm-sign", "epel-release", "git"]
RUN ["crb", "enable"]
RUN ["yum", "-y","install","fuse-devel","libcurl-devel"]
ADD rpmmacro /root/.rpmmacros
ADD buildpackages.sh /bin/
#VOLUME ["/rpms", "/srpms"]
CMD ["/bin/bash","/bin/buildpackages.sh"]
+6
View File
@@ -0,0 +1,6 @@
#!/bin/bash
for package in /srpms/*; do
rpmbuild --rebuild $package
done
find ~/rpmbuild/RPMS -type f -exec cp {} /rpms/ \;
+2
View File
@@ -0,0 +1,2 @@
%_gpg_digest_algo sha256
%_gpg_name Lenovo Scalable Infrastructure
+12
View File
@@ -0,0 +1,12 @@
FROM ubuntu:noble
ADD stdeb.patch /tmp/
ADD buildapt.sh /bin/
ADD distributions.tmpl /bin/
RUN ["apt-get", "update"]
RUN ["apt-get", "install", "-y", "reprepro", "python3-stdeb", "gnupg-agent", "devscripts", "debhelper", "libsoap-lite-perl", "libdbi-perl", "quilt", "git", "python3-pyparsing", "python3-netifaces", "dh-python", "libjson-perl", "pandoc", "alien", "gcc", "make", "alien"]
RUN ["mkdir", "-p", "/sources/git/"]
RUN ["mkdir", "-p", "/debs/"]
RUN ["mkdir", "-p", "/apt/"]
RUN ["bash", "-c", "patch -p1 < /tmp/stdeb.patch"]
CMD ["/bin/bash", "/bin/buildapt.sh"]
+21
View File
@@ -0,0 +1,21 @@
#cp -a /sources/git /tmp
for builder in $(find /sources/git -name builddeb); do
cd $(dirname $builder)
./builddeb /debs/
done
cp /prebuilt/* /debs/
cp /osd/*.deb /debs/
mkdir -p /apt/conf/
CODENAME=$(grep VERSION_CODENAME= /etc/os-release | sed -e 's/.*=//')
if [ -z "$CODENAME" ]; then
CODENAME=$(grep VERSION= /etc/os-release | sed -e 's/.*(//' -e 's/).*//')
fi
if ! grep $CODENAME /apt/conf/distributions; then
sed -e s/#CODENAME#/$CODENAME/ /bin/distributions.tmpl >> /apt/conf/distributions
fi
cd /apt/
reprepro includedeb $CODENAME /debs/*.deb
for dsc in /debs/*.dsc; do
reprepro includedsc $CODENAME $dsc
done
+7
View File
@@ -0,0 +1,7 @@
Origin: Lenovo HPC Packages
Label: Lenovo HPC Packages
Codename: #CODENAME#
Architectures: amd64 source
Components: main
Description: Lenovo HPC Packages
+34
View File
@@ -0,0 +1,34 @@
diff -urN t/usr/lib/python3/dist-packages/stdeb/cli_runner.py t.patch/usr/lib/python3/dist-packages/stdeb/cli_runner.py
--- t/usr/lib/python3/dist-packages/stdeb/cli_runner.py 2024-06-11 18:30:13.930328999 +0000
+++ t.patch/usr/lib/python3/dist-packages/stdeb/cli_runner.py 2024-06-11 18:32:05.392731405 +0000
@@ -8,7 +8,7 @@
from ConfigParser import SafeConfigParser # noqa: F401
except ImportError:
# python 3.x
- from configparser import SafeConfigParser # noqa: F401
+ from configparser import ConfigParser # noqa: F401
from distutils.util import strtobool
from distutils.fancy_getopt import FancyGetopt, translate_longopt
from stdeb.util import stdeb_cmdline_opts, stdeb_cmd_bool_opts
diff -urN t/usr/lib/python3/dist-packages/stdeb/util.py t.patch/usr/lib/python3/dist-packages/stdeb/util.py
--- t/usr/lib/python3/dist-packages/stdeb/util.py 2024-06-11 18:32:53.864776149 +0000
+++ t.patch/usr/lib/python3/dist-packages/stdeb/util.py 2024-06-11 18:33:02.063952870 +0000
@@ -730,7 +730,7 @@
example.
"""
- cfg = ConfigParser.SafeConfigParser()
+ cfg = ConfigParser.ConfigParser()
cfg.read(cfg_files)
if cfg.has_section(module_name):
section_items = cfg.items(module_name)
@@ -801,7 +801,7 @@
if len(cfg_files):
check_cfg_files(cfg_files, module_name)
- cfg = ConfigParser.SafeConfigParser(cfg_defaults)
+ cfg = ConfigParser.ConfigParser(cfg_defaults)
for cfg_file in cfg_files:
with codecs.open(cfg_file, mode='r', encoding='utf-8') as fd:
cfg.readfp(fd)
+11
View File
@@ -0,0 +1,11 @@
FROM ubuntu:resolute
ADD stdeb.patch /tmp/
ADD buildapt.sh /bin/
ADD distributions.tmpl /bin/
RUN ["apt-get", "update"]
RUN ["apt-get", "install", "-y", "reprepro", "python3-stdeb", "gnupg-agent", "devscripts", "debhelper", "libsoap-lite-perl", "libdbi-perl", "quilt", "git", "python3-pyparsing", "python3-dnspython", "python3-netifaces", "python3-asyncssh", "dh-python", "libjson-perl", "pandoc", "alien", "gcc", "make"]
RUN ["mkdir", "-p", "/sources/git/"]
RUN ["mkdir", "-p", "/debs/"]
RUN ["mkdir", "-p", "/apt/"]
CMD ["/bin/bash", "/bin/buildapt.sh"]
+21
View File
@@ -0,0 +1,21 @@
#cp -a /sources/git /tmp
for builder in $(find /sources/git -name builddeb); do
cd $(dirname $builder)
./builddeb /debs/
done
cp /prebuilt/* /debs/
cp /osd/*.deb /debs/
mkdir -p /apt/conf/
CODENAME=$(grep VERSION_CODENAME= /etc/os-release | sed -e 's/.*=//')
if [ -z "$CODENAME" ]; then
CODENAME=$(grep VERSION= /etc/os-release | sed -e 's/.*(//' -e 's/).*//')
fi
if ! grep $CODENAME /apt/conf/distributions; then
sed -e s/#CODENAME#/$CODENAME/ /bin/distributions.tmpl >> /apt/conf/distributions
fi
cd /apt/
reprepro includedeb $CODENAME /debs/*.deb
for dsc in /debs/*.dsc; do
reprepro includedsc $CODENAME $dsc
done
+7
View File
@@ -0,0 +1,7 @@
Origin: Lenovo HPC Packages
Label: Lenovo HPC Packages
Codename: #CODENAME#
Architectures: amd64 source
Components: main
Description: Lenovo HPC Packages
+34
View File
@@ -0,0 +1,34 @@
diff -urN t/usr/lib/python3/dist-packages/stdeb/cli_runner.py t.patch/usr/lib/python3/dist-packages/stdeb/cli_runner.py
--- t/usr/lib/python3/dist-packages/stdeb/cli_runner.py 2024-06-11 18:30:13.930328999 +0000
+++ t.patch/usr/lib/python3/dist-packages/stdeb/cli_runner.py 2024-06-11 18:32:05.392731405 +0000
@@ -8,7 +8,7 @@
from ConfigParser import SafeConfigParser # noqa: F401
except ImportError:
# python 3.x
- from configparser import SafeConfigParser # noqa: F401
+ from configparser import ConfigParser # noqa: F401
from distutils.util import strtobool
from distutils.fancy_getopt import FancyGetopt, translate_longopt
from stdeb.util import stdeb_cmdline_opts, stdeb_cmd_bool_opts
diff -urN t/usr/lib/python3/dist-packages/stdeb/util.py t.patch/usr/lib/python3/dist-packages/stdeb/util.py
--- t/usr/lib/python3/dist-packages/stdeb/util.py 2024-06-11 18:32:53.864776149 +0000
+++ t.patch/usr/lib/python3/dist-packages/stdeb/util.py 2024-06-11 18:33:02.063952870 +0000
@@ -730,7 +730,7 @@
example.
"""
- cfg = ConfigParser.SafeConfigParser()
+ cfg = ConfigParser.ConfigParser()
cfg.read(cfg_files)
if cfg.has_section(module_name):
section_items = cfg.items(module_name)
@@ -801,7 +801,7 @@
if len(cfg_files):
check_cfg_files(cfg_files, module_name)
- cfg = ConfigParser.SafeConfigParser(cfg_defaults)
+ cfg = ConfigParser.ConfigParser(cfg_defaults)
for cfg_file in cfg_files:
with codecs.open(cfg_file, mode='r', encoding='utf-8') as fd:
cfg.readfp(fd)
+9
View File
@@ -0,0 +1,9 @@
cd ~/confluent
git pull
rm ~/rpmbuild/RPMS/noarch/*osdeploy*
rm ~/rpmbuild/SRPMS/*osdeploy*
sh confluent_osdeploy/buildrpm-aarch64
mkdir -p $HOME/el9/
mkdir -p $HOME/el10/
podman run --rm -it -v $HOME:/build el9build bash /build/confluent/confluent_vtbufferd/buildrpm /build/el9/
+22 -2
View File
@@ -24,6 +24,7 @@ import os
import re
import select
import sys
import csv
path = os.path.dirname(os.path.realpath(__file__))
path = os.path.realpath(os.path.join(path, '..', 'lib', 'python'))
@@ -51,6 +52,8 @@ argparser.add_option('-s', '--skipcommon', action='store_true',
'groups, useful when combined with -d')
argparser.add_option('-c', '--count', action='store_true',
help='Also display count of nodes in a given group')
argparser.add_option('-t', '--tabular', action='store_true',
help='Output in CSV format')
argparser.add_option('-r', '--reverse', action='store_true',
help='Reverse sort order to show biggest output group '
'last')
@@ -68,13 +71,30 @@ if options.abbreviate:
else:
grouped = tg.GroupedData()
def print_current():
def print_current(tabular=False):
if options.diff:
grouped.print_deviants(skipmodal=options.skipcommon, count=options.count,
reverse=options.reverse, basenode=options.base)
elif options.groupcount:
grouped.generate_byoutput()
print(len(grouped.byoutput))
elif tabular:
databynode = {}
fields = ['node']
for n in grouped.bynode:
for ln in grouped.bynode[n]:
label, value = ln.split(':', 1)
label = label.strip()
value = value.strip()
databynode.setdefault(n, {})[label] = value
if label not in fields:
fields.append(label)
writer = csv.DictWriter(sys.stdout, fieldnames=fields)
writer.writeheader()
for n in databynode:
row = dict(databynode[n])
row['node'] = n
writer.writerow(row)
else:
grouped.print_all(skipmodal=options.skipcommon,
count=options.count,
@@ -125,5 +145,5 @@ while fullline:
if printpending:
if clearpending:
sys.stdout.write('\x1b[2J\x1b[;H') # clear screen
print_current()
print_current(options.tabular)
+196 -27
View File
@@ -41,7 +41,6 @@
# esc-( would interfere with normal esc use too much
# ~ I will not use for now...
import math
import getpass
import optparse
import os
@@ -108,8 +107,8 @@ class BailOut(Exception):
def print_help():
print("confetty provides a filesystem like interface to confluent. "
"Navigation is done using the same commands as would be used in a "
"filesystem. Tab completion is supported to aid in navigation,"
"as is up arrow to recall previous commands and control-r to search"
"filesystem. Tab completion is supported to aid in navigation, "
"as is up arrow to recall previous commands and control-r to search "
"previous command history, similar to using bash\n\n"
"The supported commands are:\n"
"cd [location] - Set the current command context, similar to a "
@@ -131,8 +130,19 @@ def print_help():
#common with the api document
def writeout(data):
while True:
try:
select.select((), (sys.stdout,), ())
sys.stdout.write(data)
break
except BlockingIOError:
continue
def updatestatus(stateinfo={}):
global powerstate, powertime, clearpowermessage
if opts.headless:
return
status = consolename
info = []
for statekey in stateinfo:
@@ -150,10 +160,10 @@ def updatestatus(stateinfo={}):
if 'state' in stateinfo: # currently only read power means anything
newpowerstate = stateinfo['state']['value']
if newpowerstate != powerstate and newpowerstate == 'off':
sys.stdout.write("\x1b[2J\x1b[;H[powered off]\r\n")
writeout("\x1b[2J\x1b[;H[powered off]\r\n")
clearpowermessage = True
if newpowerstate == 'on' and clearpowermessage:
sys.stdout.write("\x1b[2J\x1b[;H")
writeout("\x1b[2J\x1b[;H")
clearpowermessage = False
powerstate = newpowerstate
if 'clientcount' in laststate and laststate['clientcount'] != 1:
@@ -171,7 +181,7 @@ def updatestatus(stateinfo={}):
if info:
status += ' [' + ','.join(info) + ']'
if os.environ.get('TERM', '') not in ('linux'):
sys.stdout.write('\x1b]0;console: %s\x07' % status)
writeout('\x1b]0;console: %s\x07' % status)
sys.stdout.flush()
@@ -451,12 +461,14 @@ def do_command(command, server):
print_result(res)
elif argv[0] == 'start':
targpath = fullpath_target(argv[1])
nodename = targpath.split('/')[-3]
nodename = targpath.split('/')[2]
currconsole = targpath
startrequest = {'operation': 'start', 'path': targpath,
'parameters': {}}
height, width = struct.unpack(
'hh', fcntl.ioctl(sys.stdout, termios.TIOCGWINSZ, b'....'))[:2]
height, width = 31, 100
if not opts.headless:
height, width = struct.unpack(
'hh', fcntl.ioctl(sys.stdout, termios.TIOCGWINSZ, b'........')[:4])[:2]
startrequest['parameters']['width'] = width
startrequest['parameters']['height'] = height
for param in argv[2:]:
@@ -612,7 +624,7 @@ def do_resize(a, b):
if not inconsole:
return
height, width = struct.unpack(
'hh', fcntl.ioctl(sys.stdout, termios.TIOCGWINSZ, b'....'))[:2]
'hh', fcntl.ioctl(sys.stdout, termios.TIOCGWINSZ, b'........')[:4])[:2]
tlvdata.send(session.connection, {'operation': 'resize', 'width': width,
'height': height})
@@ -624,17 +636,19 @@ def startconsole(nodename):
signal.signal(signal.SIGWINCH, do_resize)
didconsole = True
consolename = nodename
tty.setraw(sys.stdin.fileno())
currfl = fcntl.fcntl(sys.stdin.fileno(), fcntl.F_GETFL)
fcntl.fcntl(sys.stdin.fileno(), fcntl.F_SETFL, currfl | os.O_NONBLOCK)
if not opts.headless:
tty.setraw(sys.stdin.fileno())
currfl = fcntl.fcntl(sys.stdin.fileno(), fcntl.F_GETFL)
fcntl.fcntl(sys.stdin.fileno(), fcntl.F_SETFL, currfl | os.O_NONBLOCK)
inconsole = True
check_automation('') # give any leading 'sends' a chance
def quitconfetty(code=0, fullexit=False, fixterm=True):
global inconsole
global currconsole
global didconsole
if fixterm or didconsole:
if (fixterm or didconsole) and not opts.headless:
currfl = fcntl.fcntl(sys.stdin.fileno(), fcntl.F_GETFL)
fcntl.fcntl(sys.stdin.fileno(), fcntl.F_SETFL, currfl & ~os.O_NONBLOCK)
if oldtcattr is not None:
@@ -659,9 +673,10 @@ def get_session_node(shellargs):
return targ, shellargs[0]
if len(shellargs) == 2 and shellargs[0] == 'start':
args = [s for s in shellargs[1].split('/') if s]
if len(args) == 4 and args[0] == 'nodes' and args[2] == 'console' and \
args[3] == 'session':
return shellargs[1], args[1]
if len(args) == 4 and args[0] == 'nodes':
if args[2] == 'console' and \
args[3] == 'session':
return shellargs[1], args[1]
if len(args) == 5 and args[0] == 'nodes' and args[2] == 'shell' and \
args[3] == 'sessions':
return shellargs[1], args[1]
@@ -844,6 +859,20 @@ def check_escape_seq(currinput, filehandle):
currinput += filehandle.read()
return currinput
automation_directives = []
current_automation_directive = None
automation_map = {
'<up>': '\x1b[A',
'<down>': '\x1b[B',
'<right>': '\x1b[C',
'<left>': '\x1b[D',
'<enter>': '\r',
'<esc>': '\x1b',
'<tab>': '\t',
}
parser = optparse.OptionParser()
parser.add_option("-s", "--server", dest="netserver",
help="Confluent instance to connect to",
@@ -851,12 +880,64 @@ parser.add_option("-s", "--server", dest="netserver",
parser.add_option("-c", "--control", dest="controlpath",
help="Path to offer terminal control",
metavar="PATH")
parser.add_option('-a', '--automation', type='string', default=None,
help='Specify an automation script to run', metavar='SCRIPT')
parser.add_option('-e', '--headless', action='store_true', default=False,
help='Run in headless mode, which is designed for use with '
'automation scripts and disables interactive features')
parser.add_option(
'-m', '--mintime', default=0,
help='Minimum time to run or else pause for input (used to keep a '
'terminal from closing quickly on error)')
opts, shellargs = parser.parse_args()
def parse_automation_script(script, session_node):
global current_automation_directive
for line in script.splitlines():
line = line.strip()
if not line or line.startswith('#'):
continue
if line.startswith('exit'):
automation_directives.append(('exit', None))
continue
if ' ' not in line:
sys.stderr.write("Invalid line in automation script: %s\n" % line)
continue
directive, arg = line.split(' ', 1)
directive = directive.strip().lower()
if directive not in ('expect', 'send', 'forget'):
sys.stderr.write("Unknown directive in automation script: %s\n" % directive)
continue
arg = arg.strip()
origarg = arg
if arg[0] not in ('"', "'"):
arg = '"' + arg + '"'
if arg[0] == "'" and arg[-1] == "'":
# do not process '<>' sequences in single quotes
arg = arg[1:-1]
arg = bytes(arg, "utf-8").decode("unicode_escape")
elif arg[0] == '"' and arg[-1] == '"':
arg = bytes(arg[1:-1], "utf-8").decode("unicode_escape")
for key, value in automation_map.items():
arg = arg.replace(key, value)
arg = re.sub(r'<env:(\w+)>', lambda m: os.environ[m.group(1)], arg)
if '{' in arg: # support confluent expressions
for res in session.create('/nodes/{0}/attributes/expression'.format(session_node),
{'expression': arg}):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
sys.exit(1)
if 'value' in res:
arg = res['value']
automation_directives.append((directive, arg, origarg))
if automation_directives:
current_automation_directive = automation_directives.pop(0)
if opts.headless and current_automation_directive[0] == 'expect':
sys.stdout.write(f'Expecting {repr(current_automation_directive[2])}\r\n')
sys.stdout.flush()
username = None
passphrase = None
def server_connect():
@@ -909,6 +990,7 @@ def main():
# sys.stdout.write('\x1b[H\x1b[J')
# sys.stdout.flush()
global powerstate, powertime, clearpowermessage
if sys.stdout.isatty():
@@ -925,6 +1007,9 @@ def main():
targ, session_node = get_session_node(shellargs)
if session_node is not None:
consoleonly = True
if opts.automation:
with open(opts.automation) as f:
parse_automation_script(f.read(), session_node)
do_command("start %s" % targ, netserver)
doexit = True
elif shellargs:
@@ -936,9 +1021,13 @@ def main():
while inconsole or not doexit:
if inconsole:
if opts.headless:
handles = [session.connection]
else:
handles = (sys.stdin, session.connection)
try:
rdylist, _, _ = select.select(
(sys.stdin, session.connection), (), (), 10)
handles, (), (), 10)
except select.error:
rdylist = ()
for fh in rdylist:
@@ -953,7 +1042,7 @@ def main():
except IOError:
pass
if powerstate is None or powertime < time.time() - 10: # Check powerstate every 10 seconds
if powerstate == None:
if powerstate is None:
powerstate = True
powertime = time.time()
check_power_state()
@@ -976,6 +1065,72 @@ fgcolor = None
bgcolor = None
fgshifted = False
pendseq = ''
automation_check = ''
def check_automation(data):
global automation_check
global current_automation_directive
if type(data) != str:
data = data.decode('utf-8', errors='ignore')
while current_automation_directive:
if current_automation_directive[0] == 'forget' and data:
current_automation_directive = None
automation_check = ''
if automation_directives:
current_automation_directive = automation_directives.pop(0)
if opts.headless and current_automation_directive[0] == 'expect':
sys.stdout.write(f'Expecting {repr(current_automation_directive[2])}\n')
sys.stdout.flush()
return
if current_automation_directive[0] == 'expect':
expected = current_automation_directive[1]
combined = automation_check + data
if expected and expected in combined:
data = data[combined.rindex(expected) + len(expected):]
if opts.headless:
sys.stdout.write(f'Detected {repr(current_automation_directive[2])}\r\n')
sys.stdout.flush()
current_automation_directive = None
automation_check = ''
if automation_directives:
current_automation_directive = automation_directives.pop(0)
if opts.headless and current_automation_directive[0] == 'expect':
sys.stdout.write(f'Expecting {repr(current_automation_directive[2])}\r\n')
sys.stdout.flush()
else:
# Check if there's potential start of expected data in the incoming data
combined = automation_check + data
automation_check = ''
for i in range(1, min(len(expected), len(combined)) + 1):
if expected.startswith(combined[-i:]):
automation_check = combined[-i:]
return # wait for next check
elif current_automation_directive[0] == 'send':
data = ''
automation_check = ''
if opts.headless:
sys.stdout.write(f'Sending {repr(current_automation_directive[2])}\r\n')
sys.stdout.flush()
if current_automation_directive[1]:
tlvdata.send(session.connection, current_automation_directive[1])
current_automation_directive = None
if automation_directives:
current_automation_directive = automation_directives.pop(0)
if opts.headless and current_automation_directive[0] == 'expect':
sys.stdout.write(f'Expecting {repr(current_automation_directive[2])}\r\n')
sys.stdout.flush()
elif current_automation_directive[0] == 'exit':
if opts.headless:
sys.stdout.write('Automation completed\r\n')
sys.stdout.flush()
return True
if opts.headless and not current_automation_directive:
sys.stdout.write('Automation completed\r\n')
sys.stdout.flush()
return True
return False
def consume_termdata(fh, bufferonly=False):
global clearpowermessage
global fgcolor, bgcolor, fgshifted, pendseq
@@ -987,6 +1142,11 @@ def consume_termdata(fh, bufferonly=False):
updatestatus(data)
return ''
if data is not None:
shouldexit = check_automation(data)
if opts.headless:
if shouldexit:
quitconfetty(fullexit=True)
return ''
indata = pendseq + client.stringify(data)
pendseq = ''
data = ''
@@ -1045,14 +1205,21 @@ def consume_termdata(fh, bufferonly=False):
clearpowermessage = False
if bufferonly:
return data
try:
sys.stdout.write(data)
except UnicodeEncodeError:
sys.stdout.buffer.write(data.encode('utf8'))
except IOError: # Some times circumstances are bad
# resort to byte at a time...
for d in data:
sys.stdout.write(d)
data = data.encode('utf8')
written = False
while not written:
select.select((), (sys.stdout,), ())
try:
sys.stdout.buffer.write(data)
written = True
except BlockingIOError:
continue
except IOError: # Some times circumstances are bad
# resort to byte at a time...
raise
for d in data:
sys.stdout.write(d)
written = True
now = time.time()
if ('showtime' not in laststate or
(now // 60) != laststate['showtime'] // 60):
@@ -1066,6 +1233,8 @@ def consume_termdata(fh, bufferonly=False):
# this scenario comfortable that it
# will come out soon enough
pass
if shouldexit:
quitconfetty(fullexit=True)
else:
deadline = 5
connected = False
+567
View File
@@ -0,0 +1,567 @@
#!/usr/bin/python3
# Generate a dnsmasq static DHCP configuration (a file under /etc/dnsmasq.d)
# from the confluent attribute database. One "dhcp-host=" reservation is
# emitted for every network of every node that has a hardware (MAC) address
# defined (net.hwaddr, net.<iface>.hwaddr, net.bmc.hwaddr, ...).
import argparse
import ipaddress
import json
import os
import re
import signal
import subprocess
import sys
import time
try:
signal.signal(signal.SIGPIPE, signal.SIG_DFL)
except AttributeError:
pass
path = os.path.dirname(os.path.realpath(__file__))
path = os.path.realpath(os.path.join(path, '..', 'lib', 'python'))
if path.startswith('/opt'):
sys.path.append(path)
import confluent.client as client
TARGET = '/etc/dnsmasq.d/confluent-dhcp.conf'
TAGPREFIX = 'cdhcp'
# Invocation flags that do not affect the generated content; excluded from the
# recorded "Regenerate:" line and from the settings-changed comparison.
NONCONFIG_ARGS = {'-y', '--yes', '-n', '--stdout'}
def split_list(val):
"""Split a confluent comma/whitespace delimited attribute value."""
if not val:
return []
return [v for v in re.split(r'[\s,]+', val.strip()) if v]
def config_args(argv):
"""Join the invocation arguments that affect the generated content."""
return ' '.join(a for a in argv if a not in NONCONFIG_ARGS)
def parse_net_attrib(attrib):
"""Split a net.* attribute into (network_name, field).
The field is always the final dotted component; everything between
'net.' and the field is the network name, which may itself span
several dotted components.
"""
parts = attrib.split('.')
field = parts[-1]
if len(parts) == 2:
currnet = None
else:
currnet = '.'.join(parts[1:-1])
return currnet, field
def subnet_sort_key(netaddr):
try:
return tuple(int(o) for o in netaddr.split('.'))
except ValueError:
return (0, 0, 0, 0)
def parse_routes():
"""Parse `ip --json route` into a list of (network, route_dict) tuples for
routes with a concrete destination prefix (default routes and the like are
skipped). Returns [] if iproute2 or its JSON output is unavailable."""
try:
out = subprocess.check_output(['ip', '--json', 'route'],
stderr=subprocess.DEVNULL)
except (OSError, subprocess.SubprocessError):
return []
try:
entries = json.loads(out.decode('utf-8', 'replace'))
except ValueError:
return []
routes = []
for route in entries:
dst = route.get('dst')
if not dst or dst == 'default':
continue
if '/' not in dst:
dst += '/32'
try:
net = ipaddress.ip_network(dst, strict=False)
except ValueError:
continue
routes.append((net, route))
return routes
def connected_network_for_ip(routes, ipstr):
"""Most specific directly-attached (prefsrc-bearing) route network that
contains `ipstr`, or None. Used to fill in a missing prefix length."""
try:
ip = ipaddress.ip_address(ipstr)
except ValueError:
return None
best = None
for net, route in routes:
if 'prefsrc' not in route:
continue
if ip in net and (best is None or net.prefixlen > best.prefixlen):
best = net
return best
def prefsrc_for_network(routes, network):
"""This host's preferred source address on `network` per the routing table
(the address dnsmasq should listen on for that subnet), or None."""
fallback = None
for net, route in routes:
prefsrc = route.get('prefsrc')
if not prefsrc:
continue
if net == network:
return prefsrc
if network.network_address in net:
fallback = prefsrc
return fallback
def existing_regen_args(targetpath):
"""Return the content arguments recorded in an existing config's
'Regenerate:' header line, or None if not present/readable."""
try:
with open(targetpath) as cfg:
for line in cfg:
if line.startswith('# Regenerate:'):
rest = line[len('# Regenerate:'):].strip()
marker = 'confluent2dnsmasq'
if rest.startswith(marker):
return rest[len(marker):].strip()
return rest
if not line.startswith('#'):
break
except OSError:
pass
return None
def warn_conflicts(hostentries):
"""Warn about reservations that conflict with one another: the same
hostname on different IPs, one IP reserved more than once (dnsmasq
refuses to start on a duplicate dhcp-host IP), or one MAC reserved
more than once. Diagnostics only; every entry is still emitted."""
byname = {}
byip = {}
bymac = {}
for e in hostentries:
if e['hostname']:
prev = byname.setdefault(e['hostname'], e)
if prev is not e and prev['ip'] != e['ip']:
sys.stderr.write(
"WARNING: hostname '{0}' maps to both {1} ({2}) and "
'{3} ({4})\n'.format(
e['hostname'], prev['ip'], prev['node'],
e['ip'], e['node']))
prev = byip.setdefault(e['ip'], e)
if prev is not e:
sys.stderr.write(
'WARNING: address {0} is reserved for both {1} ({2}) and '
'{3} ({4}); dnsmasq refuses to start on a duplicate '
'dhcp-host IP address\n'.format(
e['ip'], ','.join(prev['macs']), prev['node'],
','.join(e['macs']), e['node']))
for mac in e['macs']:
prev = bymac.setdefault(mac.lower(), e)
if prev is not e:
sys.stderr.write(
'WARNING: MAC {0} is reserved for both {1} ({2}) and '
'{3} ({4})\n'.format(
mac, prev['ip'], prev['node'],
e['ip'], e['node']))
def main():
ap = argparse.ArgumentParser(
description='Create /etc/dnsmasq.d static DHCP reservations for a '
'noderange from the confluent database')
ap.add_argument('noderange',
help='Noderange to generate DHCP reservations for')
ap.add_argument('--target', default=TARGET, metavar='PATH',
help='Output configuration file (default: %(default)s).')
ap.add_argument('--tag-prefix', default=TAGPREFIX, metavar='PREFIX',
help='Prefix for generated dnsmasq tag names '
'(default: %(default)s).')
ap.add_argument('--no-bind-dynamic', dest='emit_bind_dynamic',
action='store_false',
help="Do not emit the 'bind-dynamic' line "
'(it is emitted by default).')
ap.add_argument('--no-listen-address', dest='detect_listen',
action='store_false',
help='Do not autodetect and emit a listen-address line for '
"this host's address on each managed subnet. By "
'default they are emitted together with localhost '
'(127.0.0.1 and ::1).')
ap.add_argument('--no-range', dest='emit_range', action='store_false',
help='Do not emit dhcp-range lines (manage ranges '
'elsewhere). Note: dnsmasq will not answer DHCP on a '
'subnet that has no covering dhcp-range.')
ap.add_argument('--lease', default='24h', metavar='TIME',
help='Lease time for emitted dhcp-range lines '
'(default: 24h; e.g. 1h, 24h, infinite). '
'Pass an empty string to omit it.')
ap.add_argument('-n', '--stdout', action='store_true',
help='Write the configuration to stdout instead of '
'writing the target file.')
ap.add_argument('-y', '--yes', dest='assume_yes', action='store_true',
help='Do not prompt for confirmation when an existing '
'config was generated with different settings.')
ap.add_argument('--allow-empty', action='store_true',
help='Write the file even if no reservations were '
'produced (default: refuse, to avoid clobbering an '
'existing config on an empty/failed read).')
# Optional extra data injected into the DHCP reply, sourced from the
# confluent database. All are OFF by default and scoped per-subnet.
grp = ap.add_argument_group(
'optional DHCP reply data (sourced from confluent, off by default)')
grp.add_argument('--all-options', action='store_true',
help='Enable every optional reply datum listed below.')
grp.add_argument('--gateway', '--router', dest='opt_gateway',
action='store_true',
help='router, option 3, from net.<net>.ipv4_gateway')
grp.add_argument('--dns', dest='opt_dns', action='store_true',
help='dns-server, option 6, from dns.servers')
grp.add_argument('--domain', dest='opt_domain', action='store_true',
help='domain-name, option 15, from dns.domain')
grp.add_argument('--ntp', dest='opt_ntp', action='store_true',
help='ntp-server, option 42, from ntp.servers')
grp.add_argument('--mtu', dest='opt_mtu', action='store_true',
help='interface-mtu, option 26, from net.<net>.mtu')
args = ap.parse_args()
if args.all_options:
args.opt_gateway = True
args.opt_dns = True
args.opt_domain = True
args.opt_ntp = True
args.opt_mtu = True
c = client.Command()
# Kernel routing table (ip --json route), used to fill in a missing prefix
# length on a node address and to pick this host's listen-address (prefsrc)
# for each managed subnet.
routes = parse_routes()
# Node level attributes
domainbynode = {}
dnsbynode = {}
ntpbynode = {}
# Per (node, net) network attributes: nets[node][net] = {field: value}
nets = {}
read_error = False
for ent in c.read('/noderange/{0}/attributes/current'.format(
args.noderange)):
if 'error' in ent:
sys.stderr.write(ent['error'] + '\n')
read_error = True
continue
ent = ent.get('databynode', {})
for node in ent:
for attrib in ent[node]:
val = ent[node][attrib]
if not isinstance(val, dict):
continue
val = val.get('value', None)
if val in (None, ''):
continue
if attrib == 'dns.domain':
domainbynode[node] = val
elif attrib == 'dns.servers':
dnsbynode[node] = val
elif attrib == 'ntp.servers':
ntpbynode[node] = val
elif attrib.startswith('net.'):
currnet, field = parse_net_attrib(attrib)
if field not in ('hwaddr', 'ipv4_address',
'ipv4_gateway', 'hostname', 'mtu'):
continue
nets.setdefault(node, {}).setdefault(currnet, {})
nets[node][currnet][field] = val
# Build reservation entries and aggregate per-subnet information.
hostentries = [] # {subnetkey, macs, ip, hostname}
subnets = {} # subnetkey (netaddr, prefixlen) -> aggregate dict
for node in sorted(nets):
netnames = sorted(nets[node], key=lambda n: (n is not None, n or ''))
for currnet in netnames:
slot = nets[node][currnet]
hwaddr = slot.get('hwaddr')
if not hwaddr:
continue
label = currnet if currnet else '(primary)'
netprefix = (currnet + '.') if currnet else ''
addr = slot.get('ipv4_address')
if not addr:
sys.stderr.write(
'{0}: net {1} has a hwaddr but no net.{2}ipv4_address; '
'skipping\n'.format(node, label, netprefix))
continue
# If confluent has no prefix length on the address, borrow it from
# the matching directly-attached route.
if '/' not in addr:
cnet = connected_network_for_ip(routes, addr)
if cnet is not None:
addr = '{0}/{1}'.format(addr, cnet.prefixlen)
ip = addr.split('/', 1)[0]
# Reservation hostname: first token of the per-net hostname,
# else the node name for the primary (unnamed) network, else
# omit it (named net with no hostname).
hostname = None
if slot.get('hostname'):
hostname = split_list(slot['hostname'])[0]
elif currnet is None:
hostname = node
# Derive the subnet (needed for dhcp-range and for scoping
# the optional dhcp-option tags).
subnetkey = None
if '/' in addr:
try:
iface = ipaddress.ip_interface(addr)
except ValueError as e:
sys.stderr.write(
'{0}: net {1} has invalid address {2} ({3}); '
'skipping\n'.format(node, label, addr, e))
continue
network = iface.network
subnetkey = (str(network.network_address), network.prefixlen)
sub = subnets.get(subnetkey)
if sub is None:
tag = '{0}_{1}_{2}'.format(
args.tag_prefix,
str(network.network_address).replace('.', '_'),
network.prefixlen)
sub = {'netmask': str(network.netmask), 'tag': tag,
'prefsrc': prefsrc_for_network(routes, network),
'gateway': set(), 'dns': set(), 'domain': set(),
'ntp': set(), 'mtu': set()}
subnets[subnetkey] = sub
if slot.get('ipv4_gateway'):
sub['gateway'].add(slot['ipv4_gateway'])
if node in dnsbynode:
sub['dns'].add(dnsbynode[node])
if node in domainbynode:
sub['domain'].add(domainbynode[node])
if node in ntpbynode:
sub['ntp'].add(ntpbynode[node])
if slot.get('mtu'):
sub['mtu'].add(slot['mtu'])
elif args.emit_range:
sys.stderr.write(
'{0}: net {1} address {2} has no /prefix; cannot derive a '
'subnet for dhcp-range (add a prefix or use --no-range). '
'Emitting the dhcp-host anyway.\n'.format(
node, label, addr))
hostentries.append({'subnetkey': subnetkey,
'macs': split_list(hwaddr),
'ip': ip, 'hostname': hostname,
'node': node})
warn_conflicts(hostentries)
# Resolve optional per-subnet options from the aggregated values.
def resolve(values, what, subnetkey):
if not values:
return None
if len(values) > 1:
chosen = sorted(values)[0]
sys.stderr.write(
'subnet {0}/{1}: inconsistent {2} across nodes {3}; using '
'{4}\n'.format(subnetkey[0], subnetkey[1], what,
sorted(values), chosen))
return chosen
return next(iter(values))
optionlines = {} # subnetkey -> [ 'option:...,value', ... ]
for subnetkey, sub in subnets.items():
lines = []
if args.opt_gateway:
gw = resolve(sub['gateway'], 'net.ipv4_gateway', subnetkey)
if gw:
lines.append('option:router,{0}'.format(gw))
if args.opt_dns:
dv = resolve(sub['dns'], 'dns.servers', subnetkey)
if dv:
lines.append('option:dns-server,{0}'.format(
','.join(split_list(dv))))
if args.opt_domain:
dom = resolve(sub['domain'], 'dns.domain', subnetkey)
if dom:
lines.append('option:domain-name,{0}'.format(dom))
if args.opt_ntp:
nv = resolve(sub['ntp'], 'ntp.servers', subnetkey)
if nv:
lines.append('option:ntp-server,{0}'.format(
','.join(split_list(nv))))
if args.opt_mtu:
mv = resolve(sub['mtu'], 'net.mtu', subnetkey)
if mv:
lines.append('option:mtu,{0}'.format(mv))
if lines:
optionlines[subnetkey] = lines
orderedsubnets = sorted(subnets,
key=lambda k: subnet_sort_key(k[0]) + (k[1],))
# Render.
out = ['# Managed by confluent2dnsmasq -- DO NOT EDIT BY HAND.',
'# Regenerate: confluent2dnsmasq {0}'.format(
config_args(sys.argv[1:])),
'# Generated {0}'.format(time.strftime('%Y-%m-%d %H:%M:%S %z')),
'']
if args.emit_bind_dynamic:
out.append('# Allow confluent and dnsmasq to share the same network '
'for DHCP')
out.append('bind-dynamic')
out.append('')
detected = []
if args.detect_listen:
for subnetkey in orderedsubnets:
prefsrc = subnets[subnetkey].get('prefsrc')
if prefsrc and prefsrc not in detected:
detected.append(prefsrc)
out.append('# Listen on localhost plus this host\'s address on each '
'managed subnet,')
out.append('# so dnsmasq serves these networks alongside confluent '
'and bind-dynamic')
out.append('# (a listen-address line also overrides local-service in '
'dnsmasq.conf).')
out.append('listen-address=127.0.0.1')
out.append('listen-address=::1')
for addr in detected:
out.append('listen-address={0}'.format(addr))
out.append('')
hosts_by_subnet = {}
nosubnet_hosts = []
for e in hostentries:
fields = list(e['macs'])
if e['subnetkey'] in optionlines:
fields.append('set:{0}'.format(subnets[e['subnetkey']]['tag']))
fields.append(e['ip'])
if e['hostname']:
fields.append(e['hostname'])
line = 'dhcp-host={0}'.format(','.join(fields))
if e['subnetkey'] is None:
nosubnet_hosts.append(line)
else:
hosts_by_subnet.setdefault(e['subnetkey'], []).append(line)
for subnetkey in orderedsubnets:
sub = subnets[subnetkey]
netaddr, prefixlen = subnetkey
out.append('# subnet {0}/{1}'.format(netaddr, prefixlen))
if args.emit_range:
rangefields = [netaddr, 'static', sub['netmask']]
if args.lease:
rangefields.append(args.lease)
out.append('dhcp-range={0}'.format(','.join(rangefields)))
for optline in optionlines.get(subnetkey, []):
out.append('dhcp-option=tag:{0},{1}'.format(sub['tag'], optline))
for line in sorted(hosts_by_subnet.get(subnetkey, [])):
out.append(line)
out.append('')
if nosubnet_hosts:
out.append('# reservations with no derivable subnet (address had no '
'/prefix)')
out.extend(sorted(nosubnet_hosts))
out.append('')
text = '\n'.join(out).rstrip('\n') + '\n'
# Warnings about configuration choices that affect coexistence with
# confluent's own DHCP service.
if args.detect_listen and subnets and not detected:
sys.stderr.write(
'WARNING: no local address was found on any managed subnet, so '
'dnsmasq will only listen on localhost. Run this on the dnsmasq '
'host, or set interface/listen-address manually.\n')
if not args.detect_listen:
sys.stderr.write(
'WARNING: listen-address autodetection disabled. For dnsmasq to '
'serve these networks together with confluent and bind-dynamic, '
'set interface, except-interface or listen-address manually, or '
'disable local-service=host in dnsmasq.conf.\n')
if not args.emit_bind_dynamic:
sys.stderr.write(
'WARNING: bind-dynamic not emitted; this can conflict with '
"confluent's DHCP unless bind-dynamic is set elsewhere in the "
'dnsmasq configuration.\n')
if args.stdout:
sys.stdout.write(text)
return
if not hostentries and not args.allow_empty:
sys.stderr.write(
'No DHCP reservations were generated (no networks with both a '
'hwaddr and an ipv4_address were found){0}. Refusing to overwrite '
'{1}; pass --allow-empty to override.\n'.format(
' and there were read errors' if read_error else '',
args.target))
sys.exit(1)
# If a config already exists that was generated with different
# content-affecting settings, confirm before overwriting it.
if os.path.exists(args.target):
prev = existing_regen_args(args.target)
cur = config_args(sys.argv[1:])
if prev is not None and prev != cur:
sys.stderr.write(
'Existing {0} was generated with different settings:\n'
' old: confluent2dnsmasq {1}\n'
' new: confluent2dnsmasq {2}\n'.format(
args.target, prev or '(no arguments)',
cur or '(no arguments)'))
if not args.assume_yes:
if not sys.stdin.isatty():
sys.stderr.write(
'Refusing to regenerate with changed settings '
'non-interactively; pass -y/--yes to confirm.\n')
sys.exit(1)
try:
resp = input('Are you sure you want to regenerate using the new settings? [y/N] ')
except EOFError:
resp = ''
if resp.strip().lower() not in ('y', 'yes'):
sys.stderr.write(
'Aborted; {0} left unchanged.\n'.format(args.target))
sys.exit(1)
targetdir = os.path.dirname(args.target)
targetname = os.path.basename(args.target)
if targetdir and not os.path.isdir(targetdir):
os.makedirs(targetdir, exist_ok=True)
# Write via a hidden temp file in the same directory, then rename: the
# rename is atomic, and the dot-prefix keeps a dnsmasq conf-dir from
# loading the temp file if it scans mid-write.
tmp = os.path.join(targetdir, '.' + targetname + '.tmp')
with open(tmp, 'w') as cfg:
cfg.write(text)
os.replace(tmp, args.target)
sys.stderr.write('Wrote {0} reservation(s) across {1} subnet(s) to {2}\n'
.format(len(hostentries), len(subnets), args.target))
sys.stderr.write('Be sure to restart dnsmasq for the changes to take effect\n')
if __name__ == '__main__':
main()
+11 -3
View File
@@ -3,6 +3,7 @@ import argparse
import os
import re
import signal
import shutil
import sys
try:
signal.signal(signal.SIGPIPE, signal.SIG_DFL)
@@ -14,7 +15,6 @@ path = os.path.realpath(os.path.join(path, '..', 'lib', 'python'))
if path.startswith('/opt'):
sys.path.append(path)
import confluent.client as client
import confluent.sortutil as sortutil
def partitionhostsline(line):
comment = ''
@@ -33,6 +33,8 @@ class HostMerger(object):
self.byip = {}
self.byname = {}
self.byname6 = {}
self.ipbyname = {}
self.ipbyname6 = {}
self.sourcelines = []
self.targlines = []
@@ -54,14 +56,21 @@ class HostMerger(object):
def add_entry(self, ip, names):
targ = self.byname
ipbyname = self.ipbyname
if ':' in ip:
targ = self.byname6
ipbyname = self.ipbyname6
line = '{:<39} {}'.format(ip, names)
x = len(self.sourcelines)
self.sourcelines.append(line)
for name in names.split():
if not name:
continue
if ipbyname.setdefault(name, ip) != ip:
sys.stderr.write(
'WARNING: {0} is mapped to both {1} and {2}; all entries '
'will be written to /etc/hosts\n'.format(
name, ipbyname[name], ip))
targ[name] = x
self.byip[ip] = x
@@ -183,7 +192,6 @@ def main():
names.append(fqdn)
names = ' '.join(names)
merger.add_entry(ipdb[node][currnet], names)
merger.write_out('/etc/whatnowhosts')
else:
namesbynode = {}
ipbynode = {}
@@ -198,7 +206,7 @@ def main():
merger.add_entry(ipbynode[node], namesbynode[node])
if os.path.exists('/etc/hosts'):
merger.read_target('/etc/hosts')
os.rename('/etc/hosts', '/etc/hosts.confluentbkup')
shutil.copy2('/etc/hosts', '/etc/hosts.confluentbkup')
merger.write_out('/etc/hosts')
+1 -1
View File
@@ -57,7 +57,7 @@ def create_image(directory, image, label=None, esize=0, totalsize=None):
if __name__ == '__main__':
if len(sys.argv) < 3:
sys.stderr.write("Usage: {0} <directory> <imagefile>".format(
sys.stderr.write("Usage: {0} <directory> <imagefile>\n".format(
sys.argv[0]))
sys.exit(1)
label = None
+53 -10
View File
@@ -36,6 +36,38 @@ import confluent.client as client
import confluent.sortutil as sortutil
devnull = None
def run_automation(noderange, category, c):
exitcode = 0
automationbynode = {}
for res in c.update('/noderange/{0}/deployment/remote_config/run'.format(noderange), {
'category': category,
}):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode |= res.get('errorcode', 1)
if 'created' in res:
nodename = res['created'].split('/')[2]
automationbynode[nodename] = res['created']
while automationbynode:
for node in list(automationbynode):
for res in c.read(automationbynode[node]):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode |= res.get('errorcode', 1)
for result in res.get('results', []):
sys.stdout.write('{0}: Task [{1}] {2}\n'.format(
node, result['task_name'], result['state']))
for warning in result.get('warnings', []):
sys.stderr.write('{0}: [WARNING] {1}\n'.format(node, warning))
if 'errorinfo' in result:
for errorline in result['errorinfo'].splitlines():
sys.stderr.write('{0}: [ERROR] {1}\n'.format(node, errorline))
if res.get('complete', False):
del automationbynode[node]
sys.stdout.write('{0}: Automation complete\n'.format(node))
return exitcode
def run():
global devnull
devnull = open(os.devnull, 'rb')
@@ -51,6 +83,8 @@ def run():
help='Run the syncfiles associated with the currently completed OS profile on the noderange')
argparser.add_option('-P', '--scripts',
help='Re-run specified scripts, with full path under scripts, e.g. post.d/first,firstboot.d/second')
argparser.add_option('-A', '--automation',
help='Run the automation scripts associated with the current OS profile on the noderange, specifying category (onboot.d/firstboot.d/post.d)')
argparser.add_option('-m', '--maxnodes', type='int',
help='Specify a maximum number of '
'nodes to run remote ssh command to, '
@@ -72,18 +106,18 @@ def run():
pipedesc = {}
pendingexecs = deque()
exitcode = 0
# Kept apart from exitcode: a failed automation run must not trip the
# early exit below, which would abandon ssh children already spawned.
autoexitcode = 0
c.stop_if_noderange_over(args[0], options.maxnodes)
if options.automation:
autoexitcode = run_automation(args[0], options.automation, c)
nodemap = {}
cmdparms = []
nodes = []
for res in c.read('/noderange/{0}/nodes/'.format(args[0])):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode |= res.get('errorcode', 1)
break
node = res['item']['href'][:-1]
nodes.append(node)
cmdstorun = []
if options.security:
@@ -94,8 +128,17 @@ def run():
for script in options.scripts.split(','):
cmdstorun.append(['run_remote', script])
if not cmdstorun:
if options.automation:
sys.exit(autoexitcode)
argparser.print_help()
sys.exit(1)
for res in c.read('/noderange/{0}/nodes/'.format(args[0])):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode |= res.get('errorcode', 1)
break
node = res['item']['href'][:-1]
nodes.append(node)
idxbynode = {}
cmdvbase = ['bash', '/etc/confluent/functions']
for sshnode in nodes:
@@ -107,7 +150,7 @@ def run():
else:
pendingexecs.append((sshnode, cmdv))
if not all or exitcode:
sys.exit(exitcode)
sys.exit(exitcode | autoexitcode)
rdy = poller.poll(10)
while all:
pernodeout = {}
@@ -145,7 +188,7 @@ def run():
run_cmdv(node, cmdv, all, poller, pipedesc)
elif pendingexecs:
node, cmdv = pendingexecs.popleft()
run_cmdv(node, cmdv, all, poller. pipedesc)
run_cmdv(node, cmdv, all, poller, pipedesc)
singlepoller.close()
for node in sortutil.natural_sort(pernodeout):
for line in pernodeout[node]:
@@ -155,7 +198,7 @@ def run():
sys.stdout.flush()
if all:
rdy = poller.poll(10)
sys.exit(exitcode)
sys.exit(exitcode | autoexitcode)
def run_cmdv(node, cmdv, all, poller, pipedesc):
+118 -113
View File
@@ -1,4 +1,4 @@
#!/usr/bin/python2
#!/usr/bin/python3
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2017 Lenovo
@@ -17,6 +17,7 @@
__author__ = 'alin37'
import asyncio
from getpass import getpass
import optparse
import os
@@ -34,126 +35,130 @@ path = os.path.realpath(os.path.join(path, '..', 'lib', 'python'))
if path.startswith('/opt'):
sys.path.append(path)
import confluent.client as client
import confluent.asynclient as client
argparser = optparse.OptionParser(
usage='''\n %prog [-b] noderange [list of attributes or 'all'] \
\n %prog -c noderange <list of attributes> \
\n %prog -e noderange <attribute names to set> \
\n %prog noderange attribute1=value1 attribute2=value,...
\n ''')
argparser.add_option('-b', '--blame', action='store_true',
help='Show information about how attributes inherited')
argparser.add_option('-e', '--environment', action='store_true',
help='Set attributes, but from environment variable of '
'same name')
argparser.add_option('-c', '--clear', action='store_true',
help='Clear attributes')
argparser.add_option('-p', '--prompt', action='store_true',
help='Prompt for attribute values interactively')
argparser.add_option('-m', '--maxnodes', type='int',
help='Prompt if trying to set attributes on more '
'than specified number of nodes')
argparser.add_option('-s', '--set', dest='set', metavar='settings.batch',
default=False, help='set attributes using a batch file')
(options, args) = argparser.parse_args()
async def main():
argparser = optparse.OptionParser(
usage='''\n %prog [-b] noderange [list of attributes or 'all'] \
\n %prog -c noderange <list of attributes> \
\n %prog -e noderange <attribute names to set> \
\n %prog noderange attribute1=value1 attribute2=value,...
\n ''')
argparser.add_option('-b', '--blame', action='store_true',
help='Show information about how attributes are inherited')
argparser.add_option('-e', '--environment', action='store_true',
help='Set attributes, but from environment variable of '
'same name')
argparser.add_option('-c', '--clear', action='store_true',
help='Clear attributes')
argparser.add_option('-p', '--prompt', action='store_true',
help='Prompt for attribute values interactively')
argparser.add_option('-m', '--maxnodes', type='int',
help='Prompt if trying to set attributes on more '
'than specified number of nodes')
argparser.add_option('-s', '--set', dest='set', metavar='settings.batch',
default=False, help='set attributes using a batch file')
(options, args) = argparser.parse_args()
#setting minimal output to only output current information
showtype = 'current'
requestargs=None
try:
noderange = args[0]
nodelist = '/noderange/{0}/nodes/'.format(noderange)
except IndexError:
argparser.print_help()
sys.exit(1)
client.check_globbing(noderange)
session = client.Command()
exitcode = 0
#Sets attributes
nodetype="noderange"
if len(args) > 1:
if "=" in args[1] or options.clear or options.environment or options.prompt:
if "=" in args[1] and options.clear:
print("Can not clear and set at the same time!")
argparser.print_help()
sys.exit(1)
argassign = None
if options.prompt:
argassign = {}
for arg in args[1:]:
oneval = 1
twoval = 2
while oneval != twoval:
oneval = getpass('Enter value for {0}: '.format(arg))
twoval = getpass('Confirm value for {0}: '.format(arg))
if oneval != twoval:
print('Values did not match.')
argassign[arg] = twoval
session.stop_if_noderange_over(noderange, options.maxnodes)
exitcode=client.updateattrib(session,args,nodetype, noderange, options, argassign)
#setting minimal output to only output current information
showtype = 'current'
requestargs=None
try:
# setting user output to what the user inputs
if args[1] == 'all':
showtype = 'all'
requestargs=args[2:]
elif args[1] == 'current':
showtype = 'current'
requestargs=args[2:]
else:
showtype = 'all'
requestargs=args[1:]
except:
pass
elif options.clear or options.environment or options.prompt:
sys.stderr.write('Attribute names required with specified options\n')
argparser.print_help()
exitcode = 400
noderange = args[0]
nodelist = '/noderange/{0}/nodes/'.format(noderange)
except IndexError:
argparser.print_help()
sys.exit(1)
client.check_globbing(noderange)
session = client.Command()
exitcode = 0
elif options.set:
arglist = [noderange]
showtype='current'
argfile = open(options.set, 'r')
argset = argfile.readline()
while argset:
#Sets attributes
nodetype="noderange"
if len(args) > 1:
if "=" in args[1] or options.clear or options.environment or options.prompt:
if "=" in args[1] and options.clear:
print("Can not clear and set at the same time!")
argparser.print_help()
sys.exit(1)
argassign = None
if options.prompt:
argassign = {}
for arg in args[1:]:
oneval = 1
twoval = 2
while oneval != twoval:
oneval = getpass('Enter value for {0}: '.format(arg))
twoval = getpass('Confirm value for {0}: '.format(arg))
if oneval != twoval:
print('Values did not match.')
argassign[arg] = twoval
await session.stop_if_noderange_over(noderange, options.maxnodes)
exitcode = await client.updateattrib(session,args,nodetype, noderange, options, argassign)
try:
argset = argset[:argset.index('#')]
except ValueError:
# setting user output to what the user inputs
if args[1] == 'all':
showtype = 'all'
requestargs=args[2:]
elif args[1] == 'current':
showtype = 'current'
requestargs=args[2:]
else:
showtype = 'all'
requestargs=args[1:]
except:
pass
argset = argset.strip()
if argset:
arglist += shlex.split(argset)
elif options.clear or options.environment or options.prompt:
sys.stderr.write('Attribute names required with specified options\n')
argparser.print_help()
exitcode = 400
elif options.set:
arglist = [noderange]
showtype='current'
argfile = open(options.set, 'r')
argset = argfile.readline()
session.stop_if_noderange_over(noderange, options.maxnodes)
exitcode=client.updateattrib(session,arglist,nodetype, noderange, options, None)
if exitcode != 0:
while argset:
try:
argset = argset[:argset.index('#')]
except ValueError:
pass
argset = argset.strip()
if argset:
arglist += shlex.split(argset)
argset = argfile.readline()
await session.stop_if_noderange_over(noderange, options.maxnodes)
exitcode= await client.updateattrib(session,arglist,nodetype, noderange, options, None)
if exitcode != 0:
sys.exit(exitcode)
# Lists all attributes
if len(args) > 0:
# setting output to all so it can search since if we do have something to search, we want to show all outputs even if it is blank.
if requestargs is None:
showtype = 'current'
elif requestargs == []:
#showtype already set
pass
else:
try:
requestargs.remove('all')
requestargs.remove('current')
except ValueError:
pass
exitcode = await client.printattributes(session, requestargs, showtype,nodetype, noderange, options)
else:
for res in session.read(nodelist):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode = 1
else:
print(res['item']['href'].replace('/', ''))
sys.exit(exitcode)
# Lists all attributes
if __name__ == '__main__':
asyncio.run(main())
if len(args) > 0:
# setting output to all so it can search since if we do have something to search, we want to show all outputs even if it is blank.
if requestargs is None:
showtype = 'current'
elif requestargs == []:
#showtype already set
pass
else:
try:
requestargs.remove('all')
requestargs.remove('current')
except ValueError:
pass
exitcode = client.printattributes(session, requestargs, showtype,nodetype, noderange, options)
else:
for res in session.read(nodelist):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode = 1
else:
print(res['item']['href'].replace('/', ''))
sys.exit(exitcode)
+2 -2
View File
@@ -40,7 +40,7 @@ argparser.add_option('-m', '--maxnodes', type='int',
argparser.add_option('-p', '--prompt', action='store_true',
help='Prompt for password values interactively')
argparser.add_option('-e', '--environment', action='store_true',
help='Set passwod, but from environment variable of '
help='Set password, but from environment variable of '
'same name')
@@ -92,7 +92,7 @@ for rsp in session.read('/noderange/{0}/configuration/management_controller/user
continue
for user in rsp['databynode'][node]['users']:
if user['username'] == username:
if not user['uid'] in uid_dict:
if user['uid'] not in uid_dict:
uid_dict[user['uid']] = node
continue
uid_dict[user['uid']] = uid_dict[user['uid']] + ',{}'.format(node)
+15
View File
@@ -76,6 +76,10 @@ if __name__ == '__main__':
list_parser = subparsers.add_parser('listbmccacerts', help='List BMC CA certificates')
sign_bmc_parser = subparsers.add_parser('signbmccert', help='Sign BMC certificate')
sign_bmc_parser.add_argument('--days', type=int, help='Number of days the certificate is valid for')
sign_bmc_parser.add_argument('--added-names', type=str, help='Additional names to include in the certificate')
args = parser.parse_args()
c = client.Command()
if args.command == 'installbmccacert':
@@ -84,6 +88,17 @@ if __name__ == '__main__':
removebmccacert(args.noderange, args.id, c)
elif args.command == 'listbmccacerts':
listbmccacerts(args.noderange, c)
elif args.command == 'signbmccert':
payload = {}
if args.days is not None:
payload['days'] = args.days
else:
print("Error: --days is required for signbmccert", file=sys.stderr)
sys.exit(1)
if args.added_names:
payload['added_names'] = args.added_names
for res in c.update(f'/noderange/{args.noderange}/configuration/management_controller/certificate/sign', payload):
print(repr(res))
else:
parser.print_help()
sys.exit(1)
+116 -108
View File
@@ -15,7 +15,7 @@
# See the License for the specific language governing permissions and
# limitations under the License.
import asyncio
import os
import signal
import optparse
@@ -31,7 +31,7 @@ path = os.path.realpath(os.path.join(path, '..', 'lib', 'python'))
if path.startswith('/opt'):
sys.path.append(path)
import confluent.client as client
import confluent.asynclient as client
class NullOpt(object):
blame = None
@@ -62,7 +62,7 @@ argparser.add_option('-e', '--extra', dest='extra',
'to be extra configuration')
argparser.add_option('-x', '--exclude', dest='exclude',
action='store_true', default=False,
help='Treat positional arguments as items to not '
help='Treat named settings as items to not '
'examine, compare, or restore default')
argparser.add_option('-a', '--advanced', dest='advanced',
action='store_true', default=False,
@@ -73,11 +73,14 @@ argparser.add_option('-r', '--restoredefault', default=False,
dest='restoredefault', metavar="COMPONENT",
help='Restore the configuration of the node '
'to factory default for given component. '
'Currently only uefi is supported')
'Currently "uefi" and "bmc" are supported components')
argparser.add_option('-m', '--maxnodes', type='int',
help='Specify a maximum number of '
'nodes to configure, '
'prompting if over the threshold')
argparser.add_option('-p', '--pending', dest='showpending',
action='store_true', default=False,
help='Show pending changes for attributes')
(options, args) = argparser.parse_args()
cfgpaths = {
@@ -170,7 +173,7 @@ def parse_config_line(arguments, single=False):
if '=' in param or param[-1] == ':' or forceset:
if setmode is None:
setmode = True
if setmode != True:
if not setmode:
bailout('Cannot do set and query in same command: Query detected but "{0}" appears to be set'.format(param))
if '=' in param:
key, _, value = param.partition('=')
@@ -182,7 +185,7 @@ def parse_config_line(arguments, single=False):
else:
if setmode is None:
setmode = False
if setmode != False:
if setmode:
bailout('Cannot do set and query in same command: Set mode detected but "{0}" appears to be a query'.format(param))
if '.' not in param:
if param == 'bmc':
@@ -222,109 +225,114 @@ def parse_config_line(arguments, single=False):
queryparms[path] = {}
queryparms[path][attrib] = param
if options.batch:
printsys = []
argfile = open(options.batch, 'r')
argset = argfile.readline()
while argset:
try:
argset = argset[:argset.index('#')]
except ValueError:
pass
argset = argset.strip()
if argset:
parse_config_line(shlex.split(argset), single=True)
async def main():
global printsys
if options.batch:
printsys = []
argfile = open(options.batch, 'r')
argset = argfile.readline()
else:
parse_config_line(args[1:])
session = client.Command()
rcode = 0
if options.restoredefault:
session.stop_if_noderange_over(noderange, options.maxnodes)
if options.restoredefault.lower() in (
'sys', 'system', 'uefi', 'bios'):
for fr in session.update(
'/noderange/{0}/configuration/system/clear'.format(noderange),
{'clear': True}):
rcode |= client.printerror(fr)
sys.exit(rcode)
elif options.restoredefault.lower() in (
'bmc', 'imm', 'xcc'):
for fr in session.update(
'/noderange/{0}/configuration/management_controller/clear'.format(noderange),
{'clear': True}):
rcode |= client.printerror(fr)
sys.exit(rcode)
while argset:
try:
argset = argset[:argset.index('#')]
except ValueError:
pass
argset = argset.strip()
if argset:
parse_config_line(shlex.split(argset), single=True)
argset = argfile.readline()
else:
sys.stderr.write(
'Unrecognized component to restore defaults: {0}\n'.format(
options.restoredefault))
sys.exit(1)
if setmode:
session.stop_if_noderange_over(noderange, options.maxnodes)
if options.exclude:
sys.stderr.write('Cannot use exclude and assign at the same time\n')
sys.exit(1)
updatebypath = {}
attrnamebypath = {}
for key in assignment:
if key not in cfgpaths:
if key.startswith('bmc.'):
path = 'configuration/management_controller/extended/all'
attrib = key.replace('bmc.', '')
else:
path = 'configuration/system/all'
attrib = key
parse_config_line(args[1:])
session = client.Command()
rcode = 0
if options.restoredefault:
await session.stop_if_noderange_over(noderange, options.maxnodes)
if options.restoredefault.lower() in (
'sys', 'system', 'uefi', 'bios'):
async for fr in session.update(
'/noderange/{0}/configuration/system/clear'.format(noderange),
{'clear': True}):
rcode |= client.printerror(fr)
sys.exit(rcode)
elif options.restoredefault.lower() in (
'bmc', 'imm', 'xcc'):
async for fr in session.update(
'/noderange/{0}/configuration/management_controller/clear'.format(noderange),
{'clear': True}):
rcode |= client.printerror(fr)
sys.exit(rcode)
else:
path, attrib = cfgpaths[key]
if path not in updatebypath:
updatebypath[path] = {}
attrnamebypath[path] = {}
updatebypath[path][attrib] = assignment[key]
attrnamebypath[path][attrib] = key
# well, we want to expand things..
# check ipv4, if requested change method to static
for path in updatebypath:
for fr in session.update('/noderange/{0}/{1}'.format(noderange, path),
updatebypath[path]):
rcode |= client.printerror(fr)
for node in fr.get('databynode', []):
r = fr['databynode'][node]
if 'value' not in r:
continue
keyval = r['value']
key, val = keyval.split('=')
if key in attrnamebypath[path]:
key = attrnamebypath[path][key]
print('{0}: {1}: {2}'.format(node, key, val))
else:
for path in queryparms:
if options.comparedefault:
continue
rcode |= client.print_attrib_path(path, session, list(queryparms[path]),
NullOpt(), queryparms[path])
if printsys == 'all' or printextbmc or printbmc or printallbmc:
if printbmc or not printextbmc:
rcode |= client.print_attrib_path(
'/noderange/{0}/configuration/management_controller/extended/all'.format(noderange),
session, printbmc, options, attrprefix='bmc.')
if options.extra:
if options.advanced:
rcode |= client.print_attrib_path(
'/noderange/{0}/configuration/management_controller/extended/extra_advanced'.format(noderange),
session, printextbmc, options)
sys.stderr.write(
'Unrecognized component to restore defaults: {0}\n'.format(
options.restoredefault))
sys.exit(1)
if setmode:
await session.stop_if_noderange_over(noderange, options.maxnodes)
if options.exclude:
sys.stderr.write('Cannot use exclude and assign at the same time\n')
sys.exit(1)
updatebypath = {}
attrnamebypath = {}
for key in assignment:
if key not in cfgpaths:
if key.startswith('bmc.'):
path = 'configuration/management_controller/extended/all'
attrib = key.replace('bmc.', '')
else:
path = 'configuration/system/all'
attrib = key
else:
rcode |= client.print_attrib_path(
'/noderange/{0}/configuration/management_controller/extended/extra'.format(noderange),
session, printextbmc, options)
if printsys or options.exclude:
if printsys == 'all':
printsys = []
if (options.comparedefault or printsys == []) and not options.advanced:
path = '/noderange/{0}/configuration/system/all'.format(noderange)
else:
path = '/noderange/{0}/configuration/system/advanced'.format(
noderange)
rcode = client.print_attrib_path(path, session, printsys,
options)
sys.exit(rcode)
path, attrib = cfgpaths[key]
if path not in updatebypath:
updatebypath[path] = {}
attrnamebypath[path] = {}
updatebypath[path][attrib] = assignment[key]
attrnamebypath[path][attrib] = key
# well, we want to expand things..
# check ipv4, if requested change method to static
for path in updatebypath:
async for fr in session.update('/noderange/{0}/{1}'.format(noderange, path),
updatebypath[path]):
rcode |= client.printerror(fr)
for node in fr.get('databynode', []):
r = fr['databynode'][node]
if 'value' not in r:
continue
keyval = r['value']
key, val = keyval.split('=')
if key in attrnamebypath[path]:
key = attrnamebypath[path][key]
print('{0}: {1}: {2}'.format(node, key, val))
else:
for path in queryparms:
if options.comparedefault:
continue
rcode |= await client.print_attrib_path(path, session, list(queryparms[path]),
NullOpt(), queryparms[path], showpending=options.showpending)
if printsys == 'all' or printextbmc or printbmc or printallbmc:
if printbmc or not printextbmc:
rcode |= await client.print_attrib_path(
'/noderange/{0}/configuration/management_controller/extended/all'.format(noderange),
session, printbmc, options, attrprefix='bmc.', showpending=options.showpending)
if options.extra:
if options.advanced:
rcode |= await client.print_attrib_path(
'/noderange/{0}/configuration/management_controller/extended/extra_advanced'.format(noderange),
session, printextbmc, options, showpending=options.showpending)
else:
rcode |= await client.print_attrib_path(
'/noderange/{0}/configuration/management_controller/extended/extra'.format(noderange),
session, printextbmc, options, showpending=options.showpending)
if printsys or options.exclude:
if printsys == 'all':
printsys = []
if (options.comparedefault or printsys == []) and not options.advanced:
path = '/noderange/{0}/configuration/system/all'.format(noderange)
else:
path = '/noderange/{0}/configuration/system/advanced'.format(
noderange)
rcode |= await client.print_attrib_path(path, session, printsys,
options, showpending=options.showpending)
sys.exit(rcode)
if __name__ == '__main__':
asyncio.run(main())
File diff suppressed because it is too large Load Diff
+36 -26
View File
@@ -1,4 +1,4 @@
#!/usr/bin/python2
#!/usr/bin/python3
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2017 Lenovo
@@ -15,6 +15,7 @@
# See the License for the specific language governing permissions and
# limitations under the License.
import asyncio
import optparse
import os
import signal
@@ -30,29 +31,38 @@ path = os.path.realpath(os.path.join(path, '..', 'lib', 'python'))
if path.startswith('/opt'):
sys.path.append(path)
import confluent.client as client
import confluent.asynclient as client
argparser = optparse.OptionParser(
usage='''\n %prog noderange attribute1=value1 attribute2=value,...
\n ''')
(options, args) = argparser.parse_args()
requestargs=None
try:
noderange = args[0]
except IndexError:
argparser.print_help()
sys.exit(1)
client.check_globbing(noderange)
session = client.Command()
exitcode = 0
attribs = {'name': noderange}
for arg in args[1:]:
key, val = arg.split('=', 1)
attribs[key] = val
for r in session.create('/noderange/', attribs):
if 'error' in r:
sys.stderr.write(r['error'] + '\n')
exitcode |= 1
if 'created' in r:
print('{0}: created'.format(r['created']))
sys.exit(exitcode)
async def main():
argparser = optparse.OptionParser(
usage='''\n %prog noderange attribute1=value1 attribute2=value,...
\n ''')
(options, args) = argparser.parse_args()
requestargs = None
try:
noderange = args[0]
except IndexError:
argparser.print_help()
sys.exit(1)
client.check_globbing(noderange)
session = client.Command()
exitcode = 0
attribs = {'name': noderange}
for arg in args[1:]:
if '=' not in arg:
sys.stderr.write(
'Attributes must be given as attribute=value, got "{0}"\n'.format(
arg))
sys.exit(1)
key, val = arg.split('=', 1)
attribs[key] = val
async for r in session.create('/noderange/', attribs):
if 'error' in r:
sys.stderr.write(r['error'] + '\n')
exitcode |= 1
if 'created' in r:
print('{0}: created'.format(r['created']))
sys.exit(exitcode)
if __name__ == '__main__':
asyncio.run(main())
+17 -10
View File
@@ -2,6 +2,7 @@
import argparse
import datetime
import os
import sys
path = os.path.dirname(os.path.realpath(__file__))
@@ -52,7 +53,7 @@ def setpending(nr, profile, profilebynodes, cli):
if profilebynodes:
for node in sortutil.natural_sort(profilebynodes):
prof = profilebynodes[node]
args = {'deployment.pendingprofile': prof, 'deployment.state': '', 'deployment.state_detail': ''}
args = {'deployment.pendingprofile': prof, 'deployment.state': '', 'deployment.state_detail': '', 'deployment.state_last_updated': ''}
if not prof.startswith('genesis-'):
args['deployment.stagedprofile'] = ''
args['deployment.profile'] = ''
@@ -60,7 +61,7 @@ def setpending(nr, profile, profilebynodes, cli):
args):
pass
return
args = {'deployment.pendingprofile': profile, 'deployment.state': '', 'deployment.state_detail': ''}
args = {'deployment.pendingprofile': profile, 'deployment.state': '', 'deployment.state_detail': '', 'deployment.state_last_updated': ''}
if not profile.startswith('genesis-'):
args['deployment.stagedprofile'] = ''
args['deployment.profile'] = ''
@@ -78,8 +79,9 @@ def main(args):
ap = argparse.ArgumentParser(description='Deploy OS to nodes')
ap.add_argument('-c', '--clear', help='Clear any pending deployment action', action='store_true')
ap.add_argument('-n', '--network', help='Initiate deployment over PXE/HTTP', action='store_true')
ap.add_argument('-b', '--bootmethod', help='Specify network boot method (e.g., network, http)', default='network')
ap.add_argument('-p', '--prepareonly', help='Prepare only, skip any interaction with a BMC associated with this deployment action', action='store_true')
ap.add_argument('-m', '--maxnodes', help='Specifiy a maximum nodes to be deployed')
ap.add_argument('-m', '--maxnodes', help='Specify a maximum nodes to be deployed')
ap.add_argument('-r', '--redeploy', help='Redeploy nodes with the current or pending profile', action='store_true')
ap.add_argument('noderange', help='Set of nodes to deploy')
ap.add_argument('profile', nargs='?', help='Profile name to deploy')
@@ -131,11 +133,6 @@ def main(args):
curr = nodeinfo[attr].get('value', '')
if curr and node not in profilebynode:
profilebynode[node] = curr
for lockinfo in c.read('/noderange/{0}/deployment/lock'.format(args.noderange)):
for node in lockinfo.get('databynode', {}):
lockstate = lockinfo['databynode'][node]['lock']['value']
if lockstate == 'locked':
lockednodes.append(node)
if args.profile and profilebynode:
sys.stderr.write('The -r/--redeploy option cannot be used with a profile, it redeploys the current or pending profile\n')
return 1
@@ -178,7 +175,7 @@ def main(args):
if node not in databynode:
databynode[node] = {}
for attr in dbn[node]:
if attr in ('deployment.pendingprofile', 'deployment.apiarmed', 'deployment.stagedprofile', 'deployment.profile', 'deployment.state', 'deployment.state_detail'):
if attr in ('deployment.pendingprofile', 'deployment.apiarmed', 'deployment.stagedprofile', 'deployment.profile', 'deployment.state', 'deployment.state_detail', 'deployment.state_last_updated'):
databynode[node][attr] = dbn[node][attr].get('value', '')
for node in sortutil.natural_sort(databynode):
profile = databynode[node].get('deployment.pendingprofile', '')
@@ -192,6 +189,8 @@ def main(args):
profile = databynode[node].get('deployment.profile', '')
if profile:
profile = 'completed: {}'.format(profile)
elif args.clear:
profile = 'No profile applied, any pending profile has been cleared'
else:
profile= 'No profile pending or applied'
armed = databynode[node].get('deployment.apiarmed', '')
@@ -202,18 +201,26 @@ def main(args):
stateinfo = ''
deploymentstate = databynode[node].get('deployment.state', '')
if deploymentstate:
deploymentdate = databynode[node].get('deployment.state_last_updated', '')
if deploymentdate:
try:
deploymentdate = datetime.datetime.fromisoformat(deploymentdate).strftime('%m/%d/%Y %H:%M:%S')
except ValueError:
pass
statedetails = databynode[node].get('deployment.state_detail', '')
if statedetails:
stateinfo = '{}: {}'.format(deploymentstate, statedetails)
else:
stateinfo = deploymentstate
if deploymentdate:
stateinfo += ' - {}'.format(deploymentdate)
if stateinfo:
print('{0}: {1} ({2})'.format(node, profile, stateinfo))
else:
print('{0}: {1}{2}'.format(node, profile, armed))
sys.exit(0)
if not args.clear and args.network and not args.prepareonly:
rc = c.simple_noderange_command(args.noderange, '/boot/nextdevice', 'network',
rc = c.simple_noderange_command(args.noderange, '/boot/nextdevice', args.bootmethod,
bootmode='uefi',
persistent=False,
errnodes=errnodes)
+92 -76
View File
@@ -1,4 +1,4 @@
#!/usr/bin/python2
#!/usr/bin/python3
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2017 Lenovo
@@ -15,19 +15,18 @@
# See the License for the specific language governing permissions and
# limitations under the License.
import asyncio
import csv
import optparse
import os
import sys
import time
path = os.path.dirname(os.path.realpath(__file__))
path = os.path.realpath(os.path.join(path, '..', 'lib', 'python'))
if path.startswith('/opt'):
sys.path.append(path)
import confluent.client as client
import confluent.sortutil as sortutil
import confluent.asynclient as client
defcolumns = ['Node', 'Model', 'Serial', 'UUID', 'Mac Address', 'Type',
'Current IP Addresses']
@@ -50,10 +49,11 @@ columnmapping = {
}
#TODO: add chassis uuid
def register_endpoint(options, session, addr):
async def register_endpoint(options, session, addr):
neednewline = False
current = 0
for rsp in session.update('/discovery/register', {'addresses': addr}):
total = 0
async for rsp in session.update('/discovery/register', {'addresses': addr}):
if 'count' in rsp:
total = rsp['count']
elif total > 1:
@@ -69,29 +69,28 @@ def register_endpoint(options, session, addr):
if neednewline:
print('')
def subscribe_discovery(options, session, subscribe, targ):
async def subscribe_discovery(options, session, subscribe, targ):
keyn = 'subscribe' if subscribe else 'unsubscribe'
payload = {keyn: targ}
if subscribe:
for rsp in session.update('/discovery/subscriptions/{0}'.format(targ), payload):
async for rsp in session.update('/discovery/subscriptions/{0}'.format(targ), payload):
if 'status' in rsp:
print(rsp['status'])
else:
for rsp in session.delete('/discovery/subscriptions/{0}'.format(targ)):
async for rsp in session.delete('/discovery/subscriptions/{0}'.format(targ)):
if 'status' in rsp:
print(rsp['status'])
def print_disco(options, session, currmac, outhandler, columns):
async def print_disco(options, session, currid, outhandler, columns):
procinfo = {}
for tmpinfo in session.read('/discovery/by-mac/{0}'.format(currmac)):
async for tmpinfo in session.read('/discovery/by-id/{0}'.format(currid)):
procinfo.update(tmpinfo)
if 'Switch' in columns or 'Port' in columns:
if 'switch' in procinfo:
procinfo['port'] = procinfo['switchport']
else:
for tmpinfo in session.read(
'/networking/macs/by-mac/{0}'.format(currmac)):
elif ':' in currid:
async for tmpinfo in session.read(
'/networking/macs/by-mac/{0}'.format(currid)):
if 'ports' in tmpinfo:
# The api sorts so that the most specific available value
# is last
@@ -160,12 +159,12 @@ def datum_complete(datum):
searchkeys = set(['mac', 'serial', 'uuid'])
def search_record(datum, options, session):
async def search_record(datum, options, session):
for searchkey in searchkeys:
options.__dict__[searchkey] = None
for searchkey in searchkeys & set(datum):
options.__dict__[searchkey] = datum[searchkey]
return list(list_matching_macs(options, session))
return [x async for x in list_matching_ents(options, session)]
def datum_to_attrib(datum):
@@ -180,7 +179,10 @@ def datum_to_attrib(datum):
unique_fields = frozenset(['serial', 'mac', 'uuid'])
def import_csv(options, session):
# Cap how many nodes hold a discovery session at once while importing
maxconcurrentassign = 128
async def import_csv(options, session):
nodedata = []
unique_data = {}
exitcode = 0
@@ -211,12 +213,12 @@ def import_csv(options, session):
alldata.append(nodedatum)
allthere = True
for nodedatum in alldata:
if not search_record(nodedatum, options, session) and not broken:
if not await search_record(nodedatum, options, session) and not broken:
allthere = False
blocking_scan(session)
await blocking_scan(session)
break
for nodedatum in alldata:
if not allthere and not search_record(nodedatum, options, session):
if not allthere and not await search_record(nodedatum, options, session):
sys.stderr.write(
"Could not match the following data: " +
repr(nodedatum) + '\n')
@@ -224,11 +226,13 @@ def import_csv(options, session):
nodedata.append(nodedatum)
if broken:
sys.exit(1)
assignments = []
assignlimit = asyncio.Semaphore(maxconcurrentassign)
for datum in nodedata:
maclist = search_record(datum, options, session)
maclist = await search_record(datum, options, session)
datum = datum_to_attrib(datum)
nodename = datum['name']
for res in session.create('/nodes/', datum):
async for res in session.create('/nodes/', datum):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode |= res.get('errorcode', 1)
@@ -237,13 +241,26 @@ def import_csv(options, session):
print('Defined ' + res['created'])
else:
print(repr(res))
child = os.fork()
if child:
continue
assignments.append(
asyncio.create_task(assign_macs(maclist, nodename, assignlimit)))
for rcode in await asyncio.gather(*assignments, return_exceptions=True):
if isinstance(rcode, BaseException):
sys.stderr.write('Error assigning discovery data: {0}\n'.format(rcode))
rcode = 1
exitcode |= rcode
if exitcode:
sys.exit(exitcode)
async def assign_macs(maclist, nodename, assignlimit):
exitcode = 0
async with assignlimit:
# A session of our own, since the connection carries one request at a
# time and the caller's is busy defining the remaining nodes
mysess = client.Command()
for mac in maclist:
mysess = client.Command()
for res in mysess.update('/discovery/by-mac/{0}'.format(mac),
{'node': nodename}):
async for res in mysess.update('/discovery/by-id/{0}'.format(mac),
{'node': nodename}):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode |= res.get('errorcode', 1)
@@ -252,17 +269,10 @@ def import_csv(options, session):
print('Discovered ' + res['assigned'])
else:
print(repr(res))
sys.exit(0)
while True:
try:
os.wait()
except ChildProcessError:
break
if exitcode:
sys.exit(exitcode)
return exitcode
def list_discovery(options, session):
async def list_discovery(options, session):
orderby = None
if options.fields:
columns = []
@@ -278,23 +288,24 @@ def list_discovery(options, session):
if options.order.lower() == field.lower():
orderby = field
outhandler = client.Tabulator(columns)
for mac in list_matching_macs(options, session):
print_disco(options, session, mac, outhandler, columns)
for infoid in [x async for x in list_matching_ents(options, session)]:
await print_disco(options, session, infoid, outhandler, columns)
if options.csv:
outhandler.write_csv(sys.stdout, orderby)
else:
for row in outhandler.get_table(orderby):
print(row)
def clear_discovery(options, session):
for mac in list_matching_macs(options, session):
for res in session.delete('/discovery/by-mac/{0}'.format(mac)):
async def clear_discovery(options, session):
allidentifiers = [x async for x in list_matching_ents(options, session)]
for infoid in allidentifiers:
async for res in session.delete('/discovery/by-id/{0}'.format(infoid)):
if 'deleted' in res:
print('Cleared info for {0}'.format(res['deleted']))
else:
print(repr(res))
def list_matching_macs(options, session, node=None, checknode=True):
async def list_matching_ents(options, session, node=None, checknode=True):
path = '/discovery/'
if node:
path += 'by-node/{0}/'.format(node)
@@ -313,23 +324,22 @@ def list_matching_macs(options, session, node=None, checknode=True):
options.state = 'unidentified'
path += 'by-state/{0}/'.format(options.state).lower()
if options.mac:
path += 'by-mac/{0}'.format(options.mac)
result = list(session.read(path))[0]
path += 'by-id/{0}'.format(options.mac)
result = list([x async for x in session.read(path)])[0]
if 'error' in result:
return []
return [options.mac.replace(':', '-')]
return
yield options.mac.replace(':', '-')
return
else:
path += 'by-mac/'
ret = []
for x in session.read(path):
path += 'by-id/'
async for x in session.read(path):
if 'item' in x and 'href' in x['item']:
ret.append(x['item']['href'])
return ret
yield x['item']['href']
def assign_discovery(options, session, needid=True):
async def assign_discovery(options, session, needid=True):
abort = False
if options.importfile:
return import_csv(options, session)
return await import_csv(options, session)
if not options.node:
sys.stderr.write("Node (-n) must be specified for assignment\n")
abort = True
@@ -340,16 +350,19 @@ def assign_discovery(options, session, needid=True):
abort = True
if abort:
sys.exit(1)
matches = list_matching_macs(options, session, None if needid else options.node, False)
matches = [x async for x in list_matching_ents(options, session, None if needid else options.node, False)]
if not matches:
# Do a rescan to catch missing requested data
blocking_scan(session)
matches = list_matching_macs(options, session, None if needid else options.node, False)
if needid and options.mac and '.' in options.mac:
await register_endpoint(options, session, options.mac)
else:
await blocking_scan(session)
matches = [x async for x in list_matching_ents(options, session, None if needid else options.node, False)]
if not matches:
sys.stderr.write("No matching discovery candidates found\n")
sys.exit(1)
exitcode = 0
for res in session.update('/discovery/by-mac/{0}'.format(matches[0]),
async for res in session.update('/discovery/by-id/{0}'.format(matches[0]),
{'node': options.node}):
if 'assigned' in res:
print('Assigned: {0}'.format(res['assigned']))
@@ -361,14 +374,14 @@ def assign_discovery(options, session, needid=True):
if exitcode:
sys.exit(exitcode)
def blocking_scan(session):
list(session.update('/discovery/rescan', {'rescan': 'start'}))
while(list(session.read('/discovery/rescan'))[0].get('scanning', False)):
time.sleep(0.5)
list(session.update('/networking/macs/rescan', {'rescan': 'start'}))
async def blocking_scan(session, aggressive=False):
list([x async for x in session.update('/discovery/rescan', {'rescan': 'aggressive' if aggressive else 'start'})])
while(list([x async for x in session.read('/discovery/rescan')])[0].get('scanning', False)):
await asyncio.sleep(0.5)
list([x async for x in session.update('/networking/macs/rescan', {'rescan': 'start'})])
def main():
async def main():
parser = optparse.OptionParser(
usage='Usage: %prog [list|assign|rescan|clear|subscribe|unsubscribe|register] [options]')
# -a for 'address' maybe?
@@ -378,6 +391,9 @@ def main():
# flush to clear old data out? (e.g. no good way to age pxe data)
# also delete discovery datum... more targeted
# defect: -t lenovo-imm returns all
parser.add_option('-a', '--aggressive', dest='aggressive',
help='Use aggressive scanning and fingerprinting that is slower and more intrusive',
action='store_true')
parser.add_option('-m', '--model', dest='model',
help='Operate with nodes matching the specified model '
'number', metavar='MODEL')
@@ -392,8 +408,8 @@ def main():
'UUID', metavar='UUID')
parser.add_option('-n', '--node', help='Operate with the given nodename')
parser.add_option('-e', '--ethaddr', dest='mac',
help='Operate against the system with the specified MAC '
'address', metavar='MAC')
help='Operate against the system with the specified MAC or IP '
'address if mac not available', metavar='MAC')
parser.add_option('-t', '--type', dest='type',
help='Operate against the system of the specified type',
metavar='TYPE')
@@ -419,23 +435,23 @@ def main():
sys.exit(1)
session = client.Command()
if args[0] == 'list':
list_discovery(options, session)
await list_discovery(options, session)
if args[0] == 'clear':
clear_discovery(options, session)
await clear_discovery(options, session)
if args[0] == 'assign':
assign_discovery(options, session)
await assign_discovery(options, session)
if args[0] == 'reassign':
assign_discovery(options, session, False)
await assign_discovery(options, session, False)
if args[0] == 'register':
register_endpoint(options, session, args[1])
await register_endpoint(options, session, args[1])
if args[0] == 'subscribe':
subscribe_discovery(options, session, True, args[1])
await subscribe_discovery(options, session, True, args[1])
if args[0] == 'unsubscribe':
subscribe_discovery(options, session, False, args[1])
await subscribe_discovery(options, session, False, args[1])
if args[0] == 'rescan':
blocking_scan(session)
await blocking_scan(session, aggressive=options.aggressive)
print("Rescan complete")
if __name__ == '__main__':
main()
asyncio.run(main())
+62 -2
View File
@@ -53,7 +53,13 @@ argparser.add_option('-t', '--timeframe', type='string',
'entries from the last hours or days. '
'1h would be one hour, 4d would be four days. '
'format <num>h or <num>d'
)
)
argparser.add_option('-s', '--source', type='string',
help='only show entries from the named logs, comma '
'delimited. "-s list" names the logs this node '
'offers instead of showing entries')
(options, args) = argparser.parse_args()
try:
noderange = args[0]
@@ -72,6 +78,19 @@ if len(args) == 2:
argparser.print_help()
sys.exit(1)
listmode = bool(options.source) and options.source.lower() == 'list'
wantsources = None
if options.source and not listmode:
wantsources = set(x.strip().lower() for x in options.source.split(',')
if x.strip())
if options.source and deletemode:
# Clearing a subset is not something the platforms offer, and clearing more
# than was asked for is not something to do quietly
sys.stderr.write(
'Clearing a selection of logs is not supported, so --source cannot be '
'combined with clear\n')
sys.exit(1)
session = client.Command()
exitcode = 0
@@ -125,6 +144,7 @@ if options.timeframe:
sys.exit(1)
timeframe = dt.now() - tdelta
sources_seen = {}
event_dict = {}
nodes = []
for res in session.read('/noderange/{0}/nodes/'.format(args[0])):
@@ -145,6 +165,17 @@ for rsp in func('/noderange/{0}/events/hardware/log'.format(noderange)):
exitcode |= 1
if 'events' in thisdata:
evtdata = thisdata['events']
for evt in evtdata:
if evt.get('log_id', None):
sources_seen.setdefault(node, set()).add(evt['log_id'])
if listmode:
continue
if wantsources is not None:
# Filter before any tail is taken, or asking for the last
# few entries of one log would answer with fewer
evtdata = [x for x in evtdata if (x.get('log_id', '')
or '').lower()
in wantsources]
if options.lines:
event_dict[node].extend(evtdata)
else:
@@ -158,7 +189,34 @@ for rsp in func('/noderange/{0}/events/hardware/log'.format(noderange)):
else:
print('{0}: {1}'.format(node, format_event(evt)))
if options.lines:
if listmode:
for node in nodes:
found = sorted(sources_seen.get(node, ()))
if found:
print('{0}: {1}'.format(node, ','.join(found)))
else:
sys.stderr.write(
'No event log sources reported for "{0}"\n'.format(node))
elif wantsources is not None:
for node in nodes:
found = set(x.lower() for x in sources_seen.get(node, ()))
missing = wantsources - found
if not missing:
continue
if found:
sys.stderr.write('{0}: no such log source: {1}, this node has {2}\n'
.format(node, ','.join(sorted(missing)),
','.join(sorted(sources_seen[node]))))
else:
# A log that holds no entries names no source, so an empty answer
# here is as often a quiet platform as a wrong name. Either way
# saying nothing would look like the log simply had nothing in it.
sys.stderr.write(
'{0}: no entries from any log source, so nothing could match '
'{1}\n'.format(node, ','.join(sorted(missing))))
exitcode |= 1
if options.lines and not listmode:
for node in nodes:
evtdata_list = event_dict[node]
if len(evtdata_list) != 0:
@@ -172,3 +230,5 @@ if options.lines:
print('{0}: {1}'.format(node, format_event(evt)))
else:
print('{0}: {1}'.format(node, format_event(evt)))
sys.exit(exitcode)
+37 -5
View File
@@ -56,9 +56,14 @@ components = ['all']
argparser = optparse.OptionParser(
usage="Usage: "
"%prog <noderange> [list][updatestatus][update [--backup <file>]]|[<components>]")
"%prog <noderange> [list][updatestatus][updatetypes][update [--backup] "
"[--parameterfile <file>] <file>]|[<components>]")
argparser.add_option('-b', '--backup', action='store_true',
help='Target a backup bank rather than primary')
argparser.add_option('-p', '--parameterfile', type='string',
help='When updating, use the specified parameter file; see '
'nodefirmware(8), some platforms need it to say what '
'kind of firmware the image holds')
argparser.add_option('-m', '--maxnodes', type='int',
help='When updating, prompt if more than the specified '
'number of servers will be affected')
@@ -66,6 +71,7 @@ argparser.add_option('-m', '--maxnodes', type='int',
(options, args) = argparser.parse_args()
upfile = None
querystatus = False
querytypes = False
try:
noderange = args[0]
if len(args) > 1:
@@ -77,6 +83,8 @@ try:
comps = args[2:]
elif args[1] == 'updatestatus':
querystatus = True
elif args[1] == 'updatetypes':
querytypes = True
else:
comps = args[1:]
components = []
@@ -94,7 +102,7 @@ def get_update_progress(session, url):
for res in session.read(url):
status = res.get('phase', 'error')
percent = res.get('progress', None)
detail = res.get('detail', repr(res)),
detail = res.get('detail', repr(res))
if status == 'error':
text = 'error!'
else:
@@ -112,6 +120,10 @@ def update_firmware(session, filename):
upargs = {'filename': filename}
if options.backup:
upargs['bank'] = 'backup'
if options.parameterfile:
with open(options.parameterfile, 'rb') as pf:
pfdata = pf.read()
upargs['parameterdata'] = pfdata
noderrs = {}
if session.unixdomain:
filesbynode = {}
@@ -134,6 +146,11 @@ def update_firmware(session, filename):
pass
for res in session.create(resource, upargs):
if 'created' not in res:
if not res.get('databynode', None):
# A failure that is not attributed to any node still has to be
# reported, or the command looks like it quietly did nothing
exitcode |= client.printerror(res)
continue
for nodename in res.get('databynode', ()):
output.set_output(nodename, 'error!')
noderrs[nodename] = res['databynode'][nodename].get(
@@ -187,7 +204,21 @@ def show_firmware(session):
if not nodes_matched:
sys.stderr.write('No matching nodes for noderange "{0}"\n'.format(noderange))
elif not firmware_shown and not exitcode:
argparser.print_help()
# Asking about firmware the target does not describe is a legitimate
# question with an empty answer, not a mistake in how it was asked
sys.stderr.write('No firmware reported for "{0}"\n'.format(
','.join(components)))
def show_update_types(session):
global exitcode
for res in session.read(
'/noderange/{0}/inventory/firmware/updatetypes'.format(noderange)):
exitcode |= client.printerror(res)
for node in res.get('databynode', {}):
types = res['databynode'][node].get('types', None)
if types:
print('{0}: {1}'.format(node, ','.join(types)))
try:
@@ -195,12 +226,13 @@ try:
if querystatus:
for res in session.read(
'/noderange/{0}/inventory/firmware/updatestatus'.format(noderange)):
exitcode |= client.printerror(res)
for node in res.get('databynode', {}):
currstat = res['databynode'][node].get('status', None)
if currstat:
print('{}: {}'.format(node, currstat))
else:
print(repr(res))
elif querytypes:
show_update_types(session)
elif upfile is None:
show_firmware(session)
else:
+1 -1
View File
@@ -42,7 +42,7 @@ argparser = optparse.OptionParser(
\n %prog [options] nodegroup nodes=value1,value2
\n ''')
argparser.add_option('-b', '--blame', action='store_true',
help='Show information about how attributes inherited')
help='Show information about how attributes are inherited')
argparser.add_option('-e', '--environment', action='store_true',
help='Set attributes, but from environment variable of '
'same name')
+4 -4
View File
@@ -84,13 +84,13 @@ def main():
healthexplanations[node] = []
for sensor in health[node]['sensors']:
explanation = sensor['name'] + ':'
if sensor['value'] is not None:
if sensor.get('value', None) is not None:
explanation += str(sensor['value'])
if sensor['units'] is not None:
if sensor.get('units', None) is not None:
explanation += sensor['units']
if sensor['states']:
if sensor.get('states', None):
explanation += ','
if sensor['states']:
if sensor.get('states', None):
explanation += ','.join(sensor['states'])
healthexplanations[node].append(explanation)
if node in healthbynode and node in healthexplanations:
+12 -2
View File
@@ -38,6 +38,8 @@ if sys.version_info[0] < 3:
sys.stdout = codecs.getwriter('utf8')(sys.stdout)
filters = []
wanted = []
matched = False
def pretty(text):
@@ -121,8 +123,10 @@ if len(args) > 1:
os.execlp('nodefirmware', 'nodefirmware', noderange)
else:
url = '/noderange/{0}/inventory/hardware/all/system'
for arg in args:
for arg in arg.split(','):
for rawarg in args:
for arg in rawarg.split(','):
if arg in ('serial', 'model', 'uuid', 'mac'):
wanted.append(arg)
if arg == 'serial':
filters.append(re.compile('serial number'))
elif arg == 'model':
@@ -203,6 +207,7 @@ try:
continue
if info[datum] is None:
continue
matched = True
if options.json:
if node not in databynode:
databynode[node] = {}
@@ -212,6 +217,11 @@ try:
print(u'{0}: {1} {2}: {3}'.format(node, prefix,
pretty(datum),
info[datum]))
if filters and not matched and not options.store:
# An inventory that does not describe what was asked for is a valid
# answer, but saying nothing at all leaves the caller guessing
sys.stderr.write('No {0} reported for "{1}"\n'.format(
'/'.join(wanted) if wanted else 'matching inventory', noderange))
if options.json:
print(json.dumps(databynode, sort_keys=True, indent=4,
separators=(',', ': ')))
+1 -1
View File
@@ -162,7 +162,7 @@ for end_node in end_nodeslist:
if end_node:
end_switches = host_to_switch(end_node, eface)
if not end_switches:
print('Error: net.{0}.switch attribute is not valid')
print('Error: net.{0}.switch attribute is not valid'.format(eface))
continue
path = path_between_nodes(start_switches, end_switches)
print(f'{start_node} to {end_node}: {path}')
+1 -1
View File
@@ -19,7 +19,6 @@ import optparse
import os
import signal
import sys
import time
try:
signal.signal(signal.SIGPIPE, signal.SIG_DFL)
@@ -120,6 +119,7 @@ def show_licenses(session):
for res in session.read(
'/noderange/{0}/configuration/management_controller/licenses/'
'all'.format(noderange)):
exitcode |= client.printerror(res)
for node in res.get('databynode', {}):
for license in res['databynode'][node].get('License', []):
msg = '{0}: {1}'.format(node, license.get('feature',
+8 -7
View File
@@ -1,4 +1,4 @@
#!/usr/libexec/platform-python
#!/usr/bin/python3
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2015-2017 Lenovo
@@ -21,6 +21,7 @@ import optparse
import os
import signal
import sys
import asyncio
@@ -33,14 +34,14 @@ path = os.path.realpath(os.path.join(path, '..', 'lib', 'python'))
if path.startswith('/opt'):
sys.path.append(path)
import confluent.client as client
import confluent.asynclient as client
def main():
async def main():
argparser = optparse.OptionParser(
usage="Usage: %prog noderange\n"
" or: %prog [options] noderange <nodeattribute>...")
argparser.add_option('-b', '--blame', action='store_true',
help='Show information about how attributes inherited')
help='Show information about how attributes are inherited')
argparser.add_option('-d', '--delim', metavar="STRING", default = "\n",
help='Delimiter separating the values')
(options, args) = argparser.parse_args()
@@ -59,9 +60,9 @@ def main():
requestargs=args[1:]
nodetype='noderange'
if len(args) > 1:
exitcode=client.printattributes(session,requestargs,showtype,nodetype,noderange,options)
exitcode=await client.printattributes(session,requestargs,showtype,nodetype,noderange,options)
else:
for res in session.read(nodelist):
async for res in session.read(nodelist):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode = 1
@@ -73,4 +74,4 @@ def main():
sys.exit(exitcode)
if __name__ == '__main__':
main()
asyncio.run(main())
+1 -1
View File
@@ -1,4 +1,4 @@
#!/usr/bin/python2
#!/usr/bin/python3
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2015-2017 Lenovo
+30 -23
View File
@@ -1,4 +1,4 @@
#!/usr/bin/python2
#!/usr/bin/python3
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2017 Lenovo
@@ -15,6 +15,7 @@
# See the License for the specific language governing permissions and
# limitations under the License.
import asyncio
import optparse
import os
import signal
@@ -32,26 +33,32 @@ if path.startswith('/opt'):
import confluent.client as client
argparser = optparse.OptionParser(
usage='''\n %prog <noderange>
async def main():
argparser = optparse.OptionParser(
usage='''\n %prog <noderange>
\n ''')
argparser.add_option('-m', '--maxnodes', type='int',
help='Specify a maximum number of '
'nodes to delete, '
'prompting if over the threshold')
(options, args) = argparser.parse_args()
if len(args) != 1:
argparser.print_help()
sys.exit(1)
noderange = args[0]
client.check_globbing(noderange)
session = client.Command()
exitcode = 0
session.stop_if_noderange_over(noderange, options.maxnodes)
for r in session.delete('/noderange/{0}'.format(noderange)):
if 'error' in r:
sys.stderr.write(r['error'] + '\n')
exitcode |= 1
if 'deleted' in r:
print('{0}: deleted'.format(r['deleted']))
sys.exit(exitcode)
argparser.add_option('-m', '--maxnodes', type='int',
help='Specify a maximum number of '
'nodes to delete, '
'prompting if over the threshold')
(options, args) = argparser.parse_args()
if len(args) != 1:
argparser.print_help()
sys.exit(1)
noderange = args[0]
client.check_globbing(noderange)
session = client.Command()
exitcode = 0
session.stop_if_noderange_over(noderange, options.maxnodes)
for r in session.delete('/noderange/{0}'.format(noderange)):
if 'error' in r:
sys.stderr.write(r['error'] + '\n')
exitcode |= 1
if 'deleted' in r:
print('{0}: deleted'.format(r['deleted']))
sys.exit(exitcode)
if __name__ == '__main__':
asyncio.run(main())
-1
View File
@@ -35,7 +35,6 @@ if path.startswith('/opt'):
import confluent.client as client
import confluent.screensqueeze as sq
import confluent.sortutil as sortutil
def run():
+16 -6
View File
@@ -107,7 +107,7 @@ def sensorpass(showout=True, appendtime=False):
if 'error' in reading:
sys.stderr.write('Error: {0}\n'.format(reading['error']))
if 'errorcode' in reading:
exitcode |= exitcode
exitcode |= reading['errorcode']
else:
exitcode |= 1
if 'databynode' not in reading:
@@ -120,10 +120,14 @@ def sensorpass(showout=True, appendtime=False):
sys.stderr.write(
'{0}: Error: {1}\n'.format(node,
reading[node]['error']))
if 'errorcode' in reading[node]:
exitcode |= reading[node]['errorcode']
else:
exitcode |= 1
if 'sensors' not in reading[node]:
continue
for sensedata in reading[node]['sensors']:
if sensedata['value'] is None and options.skipnumberless:
if sensedata.get('value', None) is None and options.skipnumberless:
continue
for redundant_state in ('Non-Critical', 'Critical'):
try:
@@ -134,17 +138,20 @@ def sensorpass(showout=True, appendtime=False):
resultdata[node][sensedata['name']] = sensedata
sensorname = sensedata['name']
sensorheaders[sensorname] = sensorname
if sensedata['units'] not in (None, u''):
if sensedata.get('units', None) not in (None, u''):
sensorheaders[sensorname] += u' ({0})'.format(
sensedata['units'])
if showout:
if sensedata['value'] is None:
if sensedata.get('value', None) is None:
showval = ''
elif isinstance(sensedata['value'], float):
showval = u' {0} '.format(floatformat(sensedata['value']))
else:
showval = u' {0} '.format(sensedata['value'])
if sensedata['units'] not in (None, u''):
showval = u' {0} '.format(sensedata.get('value', ''))
# A discrete sensor has no reading for a unit to apply
# to, and printing the unit on its own says nothing
if (showval
and sensedata.get('units', None) not in (None, u'')):
showval += sensedata['units']
if sensedata.get('health', 'ok') != 'ok':
datadescription = [sensedata['health']]
@@ -258,3 +265,6 @@ try:
except KeyboardInterrupt:
print('')
sys.exit(0)
# Falling off the end exited 0 whatever sensorpass recorded, so only the
# --interval path ever reported a failed read.
sys.exit(exitcode)
+1 -1
View File
@@ -32,7 +32,7 @@ if path.startswith('/opt'):
import confluent.client as client
argparser = optparse.OptionParser(
usage='Usage: %prog [options] <noderange> [default|cd|network|setup|hd|usb|floppy]')
usage='Usage: %prog [options] <noderange> [default|cd|network|http|setup|hd|usb|floppy]')
argparser.add_option('-b', '--bios', dest='biosmode',
action='store_true', default=False,
help='Request BIOS style boot (rather than UEFI)')
+4
View File
@@ -53,6 +53,8 @@ def run():
help='Specify a custom port for ssh')
argparser.add_option('-s', '--substitutename',
help='Use a different name other than the nodename for ssh')
argparser.add_option('-t', '--timeout', type='int', default=0,
help='Timeout in seconds for each node ssh connection')
argparser.add_option('-x', '--noexpression', action='store_true',
help='Suppress expression expansion of command')
argparser.add_option('-m', '--maxnodes', type='int',
@@ -119,6 +121,8 @@ def run():
cmdv += ['-p', '{0}'.format(options.port)]
if options.loginname:
cmdv += ['-l', options.loginname]
if options.timeout:
cmdv += ['-o', 'ConnectTimeout={0}'.format(options.timeout)]
cmdv += [sshnode, cmd]
if currprocs < concurrentprocs:
currprocs += 1
+2
View File
@@ -244,6 +244,8 @@ def main():
sys.stdout.write('Aborting\n')
sys.exit(1)
handler(noderange, options, args[2:])
# The handlers record failures in exitcode, so let a caller see them
sys.exit(exitcode)
if __name__ == '__main__':
+9
View File
@@ -76,6 +76,11 @@ def download_servicedata(noderange, media, options):
session.stop_if_noderange_over(noderange, options.maxnodes)
for res in session.create(resource, upargs):
if 'created' not in res:
if not res.get('databynode', None):
# A failure that is not attributed to any node still has to be
# reported, or the command looks like it quietly did nothing
printerror(res)
continue
for nodename in res.get('databynode', ()):
output.set_output(nodename, 'error!')
noderrs[nodename] = res['databynode'][nodename].get(
@@ -148,5 +153,9 @@ def main():
argparser.print_help()
sys.exit(1)
handler(noderange, media, options)
# The handlers record failures in exitcode, so let a caller see them
sys.exit(exitcode)
if __name__ == '__main__':
main()
+860
View File
@@ -0,0 +1,860 @@
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2014 IBM Corporation
# Copyright 2015-2019 Lenovo
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import asyncio
import ctypes
import ctypes.util
import dbm
import csv
import errno
import fnmatch
import hashlib
import os
import shlex
import socket
import ssl
import sys
import confluent.asynctlvdata as tlvdata
import confluent.sortutil as sortutil
libssl = ctypes.CDLL(ctypes.util.find_library('ssl'))
libssl.SSL_CTX_set_cert_verify_callback.argtypes = [
ctypes.c_void_p, ctypes.c_void_p, ctypes.c_void_p]
SO_PASSCRED = 16
_attraliases = {
'bmc': 'hardwaremanagement.manager',
'bmcuser': 'secret.hardwaremanagementuser',
'switchuser': 'secret.hardwaremanagementuser',
'bmcpass': 'secret.hardwaremanagementpassword',
'switchpass': 'secret.hardwaremanagementpassword',
}
try:
getinput = raw_input
except NameError:
getinput = input
class PyObject_HEAD(ctypes.Structure):
_fields_ = [
("ob_refcnt", ctypes.c_ssize_t),
("ob_type", ctypes.c_void_p),
]
# see main/Modules/_ssl.c, only caring about the SSL_CTX pointer
class PySSLContext(ctypes.Structure):
_fields_ = [
("ob_base", PyObject_HEAD),
("ctx", ctypes.c_void_p),
]
@ctypes.CFUNCTYPE(ctypes.c_int, ctypes.c_void_p, ctypes.c_void_p)
def verify_stub(store, misc):
return 1
class NestedDict(dict):
def __missing__(self, key):
value = self[key] = type(self)()
return value
def stringify(instr):
# Normalize unicode and bytes to 'str', correcting for
# current python version
if isinstance(instr, bytes) and not isinstance(instr, str):
return instr.decode('utf-8', errors='replace')
elif not isinstance(instr, bytes) and not isinstance(instr, str):
return instr.encode('utf-8')
return instr
class Tabulator(object):
def __init__(self, headers):
self.headers = headers
self.rows = []
def add_row(self, row):
self.rows.append(row)
def get_table(self, order=None):
i = 0
fmtstr = ''
separator = []
for head in self.headers:
if order and order == head:
order = i
neededlen = len(head)
for row in self.rows:
if len(row[i]) > neededlen:
neededlen = len(row[i])
separator.append('-' * (neededlen + 1))
fmtstr += '{{{0}:>{1}}}|'.format(i, neededlen + 1)
i = i + 1
fmtstr = fmtstr[:-1]
yield fmtstr.format(*self.headers)
yield fmtstr.format(*separator)
if order is not None:
for row in sorted(
self.rows,
key=lambda x: sortutil.naturalize_string(x[order])):
yield fmtstr.format(*row)
else:
for row in self.rows:
yield fmtstr.format(*row)
def write_csv(self, output, order=None):
output = csv.writer(output)
output.writerow(self.headers)
i = 0
for head in self.headers:
if order and order == head:
order = i
i = i + 1
if order is not None:
for row in sorted(
self.rows,
key=lambda x: sortutil.naturalize_string(x[order])):
output.writerow(row)
else:
for row in self.rows:
output.writerow(row)
def printerror(res, node=None):
exitcode = 0
if 'errorcode' in res:
exitcode = res['errorcode']
for node in res.get('databynode', {}):
exitcode = res['databynode'][node].get('errorcode', exitcode)
if 'error' in res['databynode'][node]:
sys.stderr.write(
'{0}: {1}\n'.format(node, res['databynode'][node]['error']))
if exitcode == 0:
exitcode = 1
if 'error' in res:
if node:
sys.stderr.write('{0}: {1}\n'.format(node, res['error']))
else:
sys.stderr.write('{0}\n'.format(res['error']))
if 'errorcode' not in res:
exitcode = 1
return exitcode
def cprint(txt):
try:
print(txt)
except UnicodeEncodeError:
print(txt.encode('utf8'))
sys.stdout.flush()
def _parseserver(string):
if ']:' in string:
server, port = string[1:].split(']:')
elif string[0] == '[':
server = string[1:-1]
port = '13001'
elif ':' in string:
server, port = string.split(':')
else:
server = string
port = '13001'
return server, port
class Command(object):
def __init__(self, server=None):
self._prevdict = None
self._prevkeyname = None
self.connection = None
self._currnoderange = None
self.unixdomain = False
if server is None:
if 'CONFLUENT_HOST' in os.environ:
self.serverloc = os.environ['CONFLUENT_HOST']
else:
self.serverloc = '/var/run/confluent/api.sock'
else:
self.serverloc = server
self.connected = False
async def ensure_connected(self):
if self.connected:
return True
if os.path.isabs(self.serverloc) and os.path.exists(self.serverloc):
self._connect_unix()
self.unixdomain = True
elif self.serverloc == '/var/run/confluent/api.sock':
raise Exception('Confluent service is not available')
else:
await self._connect_tls()
self.protversion = int((await tlvdata.recv(self.connection)).split(
b'--')[1].strip()[1:])
authdata = await tlvdata.recv(self.connection)
if authdata['authpassed'] == 1:
self.authenticated = True
else:
self.authenticated = False
if not self.authenticated and 'CONFLUENT_USER' in os.environ:
username = os.environ['CONFLUENT_USER']
passphrase = os.environ['CONFLUENT_PASSPHRASE']
await self.authenticate(username, passphrase)
self.connected = True
async def add_file(self, name, handle, mode):
await self.ensure_connected()
if self.protversion < 3:
raise Exception('Not supported with connected confluent server')
if not self.unixdomain:
raise Exception('Can only add a file to a unix domain connection')
await tlvdata.send(self.connection, {'filename': name, 'mode': mode}, handle)
async def authenticate(self, username, password):
await tlvdata.send(self.connection,
{'username': username, 'password': password})
authdata = await tlvdata.recv(self.connection)
if authdata['authpassed'] == 1:
self.authenticated = True
def add_precede_key(self, keyname):
self._prevkeyname = keyname
def add_precede_dict(self, dict):
self._prevdict = dict
def handle_results(self, ikey, rc, res, errnodes=None, outhandler=None):
if 'error' in res:
if errnodes is not None:
errnodes.add(self._currnoderange)
sys.stderr.write('Error: {0}\n'.format(res['error']))
if 'errorcode' in res:
return res['errorcode']
else:
return 1
if 'databynode' not in res:
return 0
res = res['databynode']
for node in res:
if 'error' in res[node]:
if errnodes is not None:
errnodes.add(node)
sys.stderr.write('{0}: Error: {1}\n'.format(
node, res[node]['error']))
if 'errorcode' in res[node]:
rc |= res[node]['errorcode']
else:
rc |= 1
elif ikey in res[node]:
if 'value' in res[node][ikey]:
val = res[node][ikey]['value']
elif 'isset' in res[node][ikey]:
val = '********' if res[node][ikey] else ''
else:
val = repr(res[node][ikey])
if self._prevkeyname and self._prevkeyname in res[node]:
cprint('{0}: {2}->{1}'.format(
node, val, res[node][self._prevkeyname]['value']))
elif self._prevdict and node in self._prevdict:
cprint('{0}: {2}->{1}'.format(
node, val, self._prevdict[node]))
else:
cprint('{0}: {1}'.format(node, val))
elif outhandler:
outhandler(node, res)
return rc
async def simple_noderange_command(self, noderange, resource, input=None,
key=None, errnodes=None, promptover=None, outhandler=None, **kwargs):
try:
self._currnoderange = noderange
rc = 0
if resource[0] == '/':
resource = resource[1:]
# The implicit key is the resource basename
if key is None:
ikey = resource.rpartition('/')[-1]
else:
ikey = key
if input is None:
async for res in self.read('/noderange/{0}/{1}'.format(
noderange, resource)):
rc = self.handle_results(ikey, rc, res, errnodes, outhandler)
else:
await self.stop_if_noderange_over(noderange, promptover)
kwargs[ikey] = input
async for res in self.update('/noderange/{0}/{1}'.format(
noderange, resource), kwargs):
rc = self.handle_results(ikey, rc, res, errnodes, outhandler)
self._currnoderange = None
return rc
except KeyboardInterrupt:
cprint('')
return 0
async def stop_if_noderange_over(self, noderange, maxnodes):
if maxnodes is None:
return
nsize = await self.get_noderange_size(noderange)
if nsize > maxnodes:
if nsize == 1:
nodename = [x async for x in self.read(
'/noderange/{0}/nodes/'.format(noderange))][0].get('item', {}).get('href', None)
nodename = nodename[:-1]
p = getinput('Command is about to affect node {0}, continue (y/n)? '.format(nodename))
else:
p = getinput('Command is about to affect {0} nodes, continue (y/n)? '.format(nsize))
if p.lower() != 'y':
sys.stderr.write('Aborting at user request\n')
sys.exit(1)
raise Exception("Aborting at user request")
async def get_noderange_size(self, noderange):
numnodes = 0
async for node in self.read('/noderange/{0}/nodes/'.format(noderange)):
if node.get('item', {}).get('href', None):
numnodes += 1
else:
raise Exception("Error trying to size noderange {0}".format(noderange))
return numnodes
async def simple_nodegroups_command(self, noderange, resource, input=None, key=None, **kwargs):
try:
rc = 0
if resource[0] == '/':
resource = resource[1:]
# The implicit key is the resource basename
if key is None:
ikey = resource.rpartition('/')[-1]
else:
ikey = key
if input is None:
async for res in self.read('/nodegroups/{0}/{1}'.format(
noderange, resource)):
rc = self.handle_results(ikey, rc, res)
else:
kwargs[ikey] = input
async for res in self.update('/nodegroups/{0}/{1}'.format(
noderange, resource), kwargs):
rc = self.handle_results(ikey, rc, res)
return rc
except KeyboardInterrupt:
cprint('')
return 0
async def read(self, path, parameters=None):
await self.ensure_connected()
if not self.authenticated:
raise Exception('Unauthenticated')
async for rsp in send_request(
'retrieve', path, self.connection, parameters):
yield rsp
async def update(self, path, parameters=None):
await self.ensure_connected()
if not self.authenticated:
raise Exception('Unauthenticated')
async for rsp in send_request(
'update', path, self.connection, parameters):
yield rsp
async def create(self, path, parameters=None):
await self.ensure_connected()
if not self.authenticated:
raise Exception('Unauthenticated')
async for rsp in send_request(
'create', path, self.connection, parameters):
yield rsp
async def delete(self, path, parameters=None):
await self.ensure_connected()
if not self.authenticated:
raise Exception('Unauthenticated')
async for rsp in send_request(
'delete', path, self.connection, parameters):
yield rsp
def _connect_unix(self):
self.connection = socket.socket(socket.AF_UNIX, socket.SOCK_STREAM)
self.connection.setsockopt(socket.SOL_SOCKET, SO_PASSCRED, 1)
self.connection.connect(self.serverloc)
self.connection.setblocking(False)
async def _connect_tls(self):
server, port = _parseserver(self.serverloc)
for res in await asyncio.get_running_loop().getaddrinfo(
server, port, socket.AF_UNSPEC, socket.SOCK_STREAM):
af, socktype, proto, canonname, sa = res
try:
self.connection = socket.socket(af, socktype, proto)
self.connection.setsockopt(
socket.IPPROTO_TCP, socket.TCP_NODELAY, 1)
except:
self.connection = None
continue
try:
self.connection.settimeout(5)
self.connection.connect(sa)
self.connection.settimeout(0)
except:
raise
self.connection.close()
self.connection = None
continue
break
if self.connection is None:
raise Exception("Failed to connect to %s" % self.serverloc)
#TODO(jbjohnso): server certificate validation
clientcfgdir = os.path.join(os.path.expanduser("~"), ".confluent")
try:
os.makedirs(clientcfgdir)
except OSError as exc:
if not (exc.errno == errno.EEXIST and os.path.isdir(clientcfgdir)):
raise
cacert = os.path.join(clientcfgdir, "ca.pem")
certreqs = ssl.CERT_REQUIRED
knownhosts = False
if not os.path.exists(cacert):
cacert = None
certreqs = ssl.CERT_NONE
knownhosts = True
ctx = ssl.SSLContext(ssl.PROTOCOL_TLSv1_2)
ssl_ctx = PySSLContext.from_address(id(ctx)).ctx
libssl.SSL_CTX_set_cert_verify_callback(ssl_ctx, verify_stub, 0)
sreader = asyncio.StreamReader()
sreaderprot = asyncio.StreamReaderProtocol(sreader)
cloop = asyncio.get_running_loop()
tport, _ = await cloop.create_connection(
lambda: sreaderprot, sock=self.connection, ssl=ctx, server_hostname='x')
swriter = asyncio.StreamWriter(tport, sreaderprot, sreader, cloop)
self.connection = (sreader, swriter)
#self.connection = ssl.wrap_socket(self.connection, ca_certs=cacert,
# cert_reqs=certreqs)
if knownhosts:
certdata = tport.get_extra_info('ssl_object').getpeercert(binary_form=True)
# certdata = self.connection.getpeercert(binary_form=True)
fingerprint = 'sha512$' + hashlib.sha512(certdata).hexdigest()
fingerprint = fingerprint.encode('utf-8')
hostid = '@'.join((port, server))
khf = dbm.open(os.path.join(clientcfgdir, "knownhosts"), 'c', 384)
if hostid in khf:
if fingerprint == khf[hostid]:
return
else:
replace = getinput(
"MISMATCHED CERTIFICATE DATA, ACCEPT NEW? (y/n):")
if replace not in ('y', 'Y'):
raise Exception("BAD CERTIFICATE")
cprint('Adding new key for %s:%s' % (server, port))
khf[hostid] = fingerprint
async def send_request(operation, path, server, parameters=None):
"""This function iterates over all the responses
received from the server.
:param operation: The operation to request, retrieve, update, delete,
create, start, stop
:param path: The URI path to the resource to operate on
:param server: The socket to send data over
:param parameters: Parameters if any to send along with the request
"""
payload = {'operation': operation, 'path': path}
if parameters is not None:
payload['parameters'] = parameters
await tlvdata.send(server, payload)
result = await tlvdata.recv(server)
while '_requestdone' not in result:
try:
yield result
except GeneratorExit:
while '_requestdone' not in result:
result = await tlvdata.recv(server)
raise
result = await tlvdata.recv(server)
def attrrequested(attr, attrlist, seenattributes, node=None):
for candidate in attrlist:
truename = candidate
if candidate.startswith('hm'):
candidate = candidate.replace('hm', 'hardwaremanagement', 1)
if candidate in _attraliases:
candidate = _attraliases[candidate]
if fnmatch.fnmatch(attr.lower(), candidate.lower()):
if node is None:
seenattributes.add(truename)
else:
seenattributes[node][truename] = True
return True
elif attr.lower().startswith(candidate.lower() + '.'):
if node is None:
seenattributes.add(truename)
else:
seenattributes[node][truename] = 1
return True
return False
async def printattributes(session, requestargs, showtype, nodetype, noderange, options):
path = '/{0}/{1}/attributes/{2}'.format(nodetype, noderange, showtype)
return await print_attrib_path(path, session, requestargs, options)
def _sort_attrib(k):
if isinstance(k[1], dict) and k[1].get('sortid', None) is not None:
return sortutil.naturalize_string('{}'.format(k[1]['sortid']))
return sortutil.naturalize_string(k[0])
def _is_nondefault(currattr):
if currattr.get('default', None) is None:
return False # no default specified cannot tell
if 'value' not in currattr: # can't compare if no value specified
return False
if currattr['value'] != currattr['default']:
return True
# ok, the potentially pending value is the same, but check active
if 'active' not in currattr:
return False
if currattr['active'] != currattr['default']:
return True
return False
async def print_attrib_path(path, session, requestargs, options, rename=None, attrprefix=None, showpending=False):
exitcode = 0
seenattributes = NestedDict()
allnodes = set([])
async for res in session.read(path):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode = 1
continue
for node in sorted(res['databynode']):
allnodes.add(node)
for attr, val in sorted(res['databynode'][node].items(), key=_sort_attrib):
if attr == 'error':
sys.stderr.write('{0}: Error: {1}\n'.format(node, val))
continue
if attr == 'errorcode':
exitcode |= val
continue
seenattributes[node][attr] = True
if rename:
printattr = rename.get(attr, attr)
else:
printattr = attr
if attrprefix:
printattr = attrprefix + printattr
currattr = res['databynode'][node][attr]
if show_attr(attr, requestargs, seenattributes, options, node):
if 'value' in currattr:
currval = currattr.get('value', None)
if isinstance(currval, list):
currval = ','.join(currval)
activeval = currattr.get('active', currval)
if isinstance(activeval, list):
activeval = ','.join(activeval)
if activeval != currval:
if currval is None:
currval = ''
if activeval is None:
activeval = ''
if showpending:
attrout = f'{node}: {printattr}: {activeval} -> {currval}'
else:
attrout = f'{node}: {printattr}: {currval} (Pending)'
elif showpending:
continue
elif currval is not None:
attrout = '{0}: {1}: {2}'.format(
node, printattr, currval).strip()
else:
attrout = '{0}: {1}:'.format(node, printattr)
elif 'isset' in currattr:
if currattr['isset']:
attrout = '{0}: {1}: ********'.format(node,
printattr)
else:
attrout = '{0}: {1}:'.format(node, printattr)
elif isinstance(currattr, dict) and 'broken' in currattr:
attrout = '{0}: {1}: *ERROR* BROKEN EXPRESSION: ' \
'{2}'.format(node, printattr,
currattr['broken'])
elif isinstance(currattr, list) or isinstance(currattr, tuple):
attrout = '{0}: {1}: {2}'.format(node, attr, ','.join(map(str, currattr)))
elif isinstance(currattr, dict):
dictout = []
for k, v in currattr.items:
dictout.append("{0}={1}".format(k, v))
attrout = '{0}: {1}: {2}'.format(node, printattr, ','.join(map(str, dictout)))
else:
cprint("CODE ERROR" + repr(attr))
try:
blame = options.blame
except AttributeError:
blame = False
if blame or (isinstance(currattr, dict) and 'broken' in currattr):
blamedata = []
if 'inheritedfrom' in currattr:
blamedata.append('inherited from group {0}'.format(
currattr['inheritedfrom']
))
if 'expression' in currattr:
blamedata.append(
'derived from expression "{0}"'.format(
currattr['expression']))
if blamedata:
attrout += ' (' + ', '.join(blamedata) + ')'
try:
comparedefault = options.comparedefault
except AttributeError:
comparedefault = False
if comparedefault:
try:
exclude = options.exclude
except AttributeError:
exclude = False
if ((requestargs and not exclude) or
_is_nondefault(currattr)):
currval = ','.join(currattr['value']) if isinstance(
currattr['value'], list) else currattr['value']
activeval = ','.join(currattr['active']) if isinstance(
currattr.get('active'), list) else currattr.get('active')
dval = ','.join(currattr['default']) if isinstance(
currattr['default'], list) else currattr['default']
outmsg = '{0}: {1}: {2} (Default: {3}'.format(
node, printattr, currval, dval)
if 'active' in currattr:
outmsg += ', Pending change from: {0}'.format(activeval)
outmsg += ')'
cprint(outmsg)
else:
try:
details = options.detail
except AttributeError:
details = False
if details:
if currattr.get('help', None):
attrout += u' (Help: {0})'.format(
currattr['help'])
if currattr.get('possible', None):
try:
attrout += u' (Choices: {0})'.format(
','.join(currattr['possible']))
except TypeError:
pass
cprint(attrout)
somematched = set([])
printmissing = set([])
badnodes = NestedDict()
if not exitcode:
if requestargs:
for attr in requestargs:
for node in allnodes:
if attr in seenattributes[node]:
somematched.add(attr)
else:
badnodes[node][attr] = True
exitcode = 1
for node in sortutil.natural_sort(badnodes):
for attr in badnodes[node]:
if attr in somematched:
sys.stderr.write(
'Error: {0} matches no valid value for {1}\n'.format(
attr, node))
else:
printmissing.add(attr)
for missing in printmissing:
sys.stderr.write('Error: {0} not a valid attribute\n'.format(missing))
return exitcode
def show_attr(attr, requestargs, seenattributes, options, node):
try:
reverse = options.exclude
except AttributeError:
reverse = False
if requestargs is None or requestargs == []:
return True
processattr = attrrequested(attr, requestargs, seenattributes, node)
if reverse:
processattr = not processattr
return processattr
async def printgroupattributes(session, requestargs, showtype, nodetype, noderange, options):
exitcode = 0
seenattributes = set([])
async for res in session.read('/{0}/{1}/attributes/{2}'.format(nodetype, noderange, showtype)):
if 'error' in res:
sys.stderr.write(res['error'] + '\n')
exitcode = 1
continue
for attr in res:
seenattributes.add(attr)
currattr = res[attr]
if (requestargs is None or requestargs == [] or attrrequested(attr, requestargs, seenattributes)):
if 'value' in currattr:
if currattr['value'] is not None:
attrout = '{0}: {1}: {2}'.format(
noderange, attr, currattr['value'])
else:
attrout = '{0}: {1}:'.format(noderange, attr)
elif 'isset' in currattr:
if currattr['isset']:
attrout = '{0}: {1}: ********'.format(noderange, attr)
else:
attrout = '{0}: {1}:'.format(noderange, attr)
elif isinstance(currattr, dict) and 'broken' in currattr:
attrout = '{0}: {1}: *ERROR* BROKEN EXPRESSION: ' \
'{2}'.format(noderange, attr,
currattr['broken'])
elif 'expression' in currattr:
attrout = '{0}: {1}: (will derive from expression {2})'.format(noderange, attr, currattr['expression'])
elif isinstance(currattr, list) or isinstance(currattr, tuple):
attrout = '{0}: {1}: {2}'.format(noderange, attr, ','.join(map(str, currattr)))
elif isinstance(currattr, dict):
dictout = []
for k, v in currattr.items:
dictout.append("{0}={1}".format(k, v))
attrout = '{0}: {1}: {2}'.format(noderange, attr, ','.join(map(str, dictout)))
else:
cprint("CODE ERROR" + repr(attr))
cprint(attrout)
if not exitcode:
if requestargs:
for attr in requestargs:
if attr not in seenattributes:
sys.stderr.write('Error: {0} not a valid attribute\n'.format(attr))
exitcode = 1
return exitcode
async def updateattrib(session, updateargs, nodetype, noderange, options, dictassign=None):
# update attribute
exitcode = 0
if options.clear:
targpath = '/{0}/{1}/attributes/all'.format(nodetype, noderange)
keydata = {}
for attrib in updateargs[1:]:
keydata[attrib] = None
async for res in session.update(targpath, keydata):
for node in res.get('databynode', {}):
for warnmsg in res['databynode'][node].get('_warnings', []):
sys.stderr.write('Warning: ' + warnmsg + '\n')
if 'error' in res:
if 'errorcode' in res:
exitcode = res['errorcode']
sys.stderr.write('Error: ' + res['error'] + '\n')
sys.exit(exitcode)
elif hasattr(options, 'environment') and options.environment:
for key in updateargs[1:]:
key = key.replace('.', '_')
value = os.environ.get(
key, os.environ[key.upper()])
# Let's do one pass to make sure that there's not a usage problem
for key in updateargs[1:]:
key = key.replace('.', '_')
value = os.environ.get(
key, os.environ[key.upper()])
if (nodetype == "nodegroups"):
exitcode = await session.simple_nodegroups_command(noderange,
'attributes/all',
value, key)
else:
exitcode = await session.simple_noderange_command(noderange,
'attributes/all',
value, key)
sys.exit(exitcode)
elif dictassign:
for key in dictassign:
if nodetype == 'nodegroups':
exitcode = await session.simple_nodegroups_command(
noderange, 'attributes/all', dictassign[key], key)
else:
exitcode = await session.simple_noderange_command(
noderange, 'attributes/all', dictassign[key], key)
else:
if "=" in updateargs[1]:
update_ready = True
for arg in updateargs[1:]:
if '=' not in arg:
update_ready = False
exitcode = 1
if not update_ready:
sys.stderr.write('Error: {0} Can not set and read at the same time!\n'.format(str(updateargs[1:])))
sys.exit(exitcode)
try:
for val in updateargs[1:]:
val = val.split('=', 1)
if val[0][-1] in (',', '-', '^'):
key = val[0][:-1]
if val[0][-1] == ',':
value = {'prepend': val[1]}
elif val[0][-1] in ('-', '^'):
value = {'remove': val[1]}
else:
key = val[0]
value = val[1]
if (nodetype == "nodegroups"):
exitcode = await session.simple_nodegroups_command(noderange, 'attributes/all',
value, key)
else:
exitcode = await session.simple_noderange_command(noderange, 'attributes/all',
value, key)
except Exception:
sys.stderr.write('Error: {0} not a valid expression\n'.format(str(updateargs[1:])))
exitcode = 1
sys.exit(exitcode)
return exitcode
# So we try to prevent bad things from happening when globbing
# We tried to head this off at the shell, but the various solutions would end
# up breaking the shell in various ways (breaking pipe capability if using
# DEBUG, breaking globbing if in pipe, etc)
# Then we tried to parse the original commandline instead, however shlex isn't
# going to parse full bourne language (e.g. knowing that '|' and '>' and
# a world of other things would not be in our command line
# so finally, just make sure the noderange appears verbatim in the command line
# if we glob to something, then bash will change noderange and this should
# detect it and save the user from tragedy
def check_globbing(noderange):
if not os.path.exists(noderange):
return True
rawargs = os.environ.get('CURRENT_CMDLINE', None)
if rawargs:
rawargs = shlex.split(rawargs)
for arg in rawargs:
if arg.startswith('$'):
arg = arg[1:]
if arg.endswith(';'):
arg = arg[:-1]
arg = os.environ.get(arg, '$' + arg)
if arg.startswith(noderange):
break
else:
sys.stderr.write(
'Shell glob conflict detected, specified target "{0}" '
'not in command line, but is a file. You can use "set -f" in '
'bash or change directories such that there is no filename '
'that would conflict.'
'\n'.format(noderange))
sys.exit(1)
+313
View File
@@ -0,0 +1,313 @@
# vim: tabstop=4 shiftwidth=4 softtabstop=4
# Copyright 2014 IBM Corporation
# Copyright 2015 Lenovo
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import array
import asyncio
import ctypes
import ctypes.util
import confluent.tlv as tlv
import socket
from datetime import datetime
import json
import os
import struct
try:
unicode
except NameError:
unicode = str
try:
range = xrange
except NameError:
pass
class iovec(ctypes.Structure): # from uio.h
_fields_ = [('iov_base', ctypes.c_void_p),
('iov_len', ctypes.c_size_t)]
iovec_ptr = ctypes.POINTER(iovec)
class cmsghdr(ctypes.Structure): # also from bits/socket.h
_fields_ = [('cmsg_len', ctypes.c_size_t),
('cmsg_level', ctypes.c_int),
('cmsg_type', ctypes.c_int)]
@classmethod
def init_data(cls, cmsg_len, cmsg_level, cmsg_type, cmsg_data):
Data = ctypes.c_ubyte * ctypes.sizeof(cmsg_data)
class _flexhdr(ctypes.Structure):
_fields_ = cls._fields_ + [('cmsg_data', Data)]
datab = Data(*bytearray(cmsg_data))
return _flexhdr(cmsg_len=cmsg_len, cmsg_level=cmsg_level,
cmsg_type=cmsg_type, cmsg_data=datab)
def CMSG_LEN(length):
sizeof_cmshdr = ctypes.sizeof(cmsghdr)
return ctypes.c_size_t(CMSG_ALIGN(sizeof_cmshdr).value + length)
SCM_RIGHTS = 1
class msghdr(ctypes.Structure): # from bits/socket.h
_fields_ = [('msg_name', ctypes.c_void_p),
('msg_namelen', ctypes.c_uint),
('msg_iov', ctypes.POINTER(iovec)),
('msg_iovlen', ctypes.c_size_t),
('msg_control', ctypes.c_void_p),
('msg_controllen', ctypes.c_size_t),
('msg_flags', ctypes.c_int)]
def CMSG_ALIGN(length): # bits/socket.h
ret = (length + ctypes.sizeof(ctypes.c_size_t) - 1
& ~(ctypes.sizeof(ctypes.c_size_t) - 1))
return ctypes.c_size_t(ret)
def CMSG_SPACE(length): # bits/socket.h
ret = CMSG_ALIGN(length).value + CMSG_ALIGN(ctypes.sizeof(cmsghdr)).value
return ctypes.c_size_t(ret)
class ClientFile(object):
def __init__(self, name, mode, fd):
self.fileobject = os.fdopen(fd, mode)
self.filename = name
def _sendmsg(loop, fut, sock, msg, fds, wfd):
if wfd is not None:
loop.remove_writer(wfd)
if fut.cancelled():
return
try:
retdata = sock.sendmsg(
[msg],
[(socket.SOL_SOCKET, socket.SCM_RIGHTS, array.array("i", fds))])
except (BlockingIOError, InterruptedError):
fd = sock.fileno()
loop.add_writer(fd, _sendmsg, loop, fut, sock, msg, fds, fd)
except Exception as exc:
fut.set_exception(exc)
else:
fut.set_result(retdata)
def send_fds(sock, msg, fds):
cloop = asyncio.get_running_loop()
fut = cloop.create_future()
_sendmsg(cloop, fut, sock, msg, fds, None)
return fut
def _recvmsg(loop, fut, sock, msglen, maxfds, rfd):
if rfd is not None:
loop.remove_reader(rfd)
if fut.cancelled():
return
fds = array.array("i") # Array of ints
try:
msg, ancdata, flags, addr = sock.recvmsg(
msglen, socket.CMSG_LEN(maxfds * fds.itemsize))
except (BlockingIOError, InterruptedError):
fd = sock.fileno()
loop.add_reader(fd, _recvmsg, loop, fut, sock, msglen, maxfds, fd)
except Exception as exc:
fut.set_exception(exc)
else:
for cmsg_level, cmsg_type, cmsg_data in ancdata:
if (cmsg_level == socket.SOL_SOCKET
and cmsg_type == socket.SCM_RIGHTS):
# Append data, ignoring any truncated integers at the end.
fds.frombytes(
cmsg_data[
:len(cmsg_data) - (len(cmsg_data) % fds.itemsize)])
fut.set_result((msg, list(fds)))
def recv_fds(sock, msglen, maxfds):
cloop = asyncio.get_running_loop()
fut = cloop.create_future()
_recvmsg(cloop, fut, sock, msglen, maxfds, None)
return fut
def decodestr(value):
ret = None
try:
ret = value.decode('utf-8')
except UnicodeDecodeError:
try:
ret = value.decode('cp437')
except UnicodeDecodeError:
ret = value
except AttributeError:
return value
return ret
def unicode_dictvalues(dictdata):
for key in dictdata:
if isinstance(dictdata[key], bytes):
dictdata[key] = decodestr(dictdata[key])
elif isinstance(dictdata[key], datetime):
dictdata[key] = dictdata[key].strftime('%Y-%m-%dT%H:%M:%S')
elif isinstance(dictdata[key], list):
_unicode_list(dictdata[key])
elif isinstance(dictdata[key], dict):
unicode_dictvalues(dictdata[key])
def _unicode_list(currlist):
for i in range(len(currlist)):
if isinstance(currlist[i], str):
currlist[i] = decodestr(currlist[i])
elif isinstance(currlist[i], dict):
unicode_dictvalues(currlist[i])
elif isinstance(currlist[i], list):
_unicode_list(currlist[i])
async def sendall(handle, data):
if isinstance(handle, tuple):
handle[1].write(data)
return await handle[1].drain()
else:
cloop = asyncio.get_running_loop()
return await cloop.sock_sendall(handle, data)
def get_socket(handle):
if isinstance(handle, tuple):
return handle[1].transport.get_extra_info('socket')
else:
return handle
async def close(handle):
if isinstance(handle, tuple):
handle[1].close()
try:
await handle[1].wait_closed()
except Exception:
pass
else:
handle.close()
async def send(handle, data, filehandle=None):
cloop = asyncio.get_running_loop()
if isinstance(data, unicode):
try:
data = data.encode('utf-8')
except AttributeError:
pass
if isinstance(data, bytes) or isinstance(data, unicode):
# plain text, e.g. console data
tl = len(data)
if tl == 0:
# if you don't have anything to say, don't say anything at all
return
if tl < 16777216:
# type for string is '0', so we don't need
# to xor anything in
await sendall(handle, struct.pack("!I", tl))
else:
raise Exception("String data length exceeds protocol")
await sendall(handle, data)
elif isinstance(data, dict): # JSON currently only goes to 4 bytes
# Some structured message, like what would be seen in http responses
unicode_dictvalues(data) # make everything unicode, assuming UTF-8
sdata = json.dumps(data, ensure_ascii=False, separators=(',', ':'))
sdata = sdata.encode('utf-8')
tl = len(sdata)
if tl > 16777215:
raise Exception("JSON data exceeds protocol limits")
# xor in the type (0b1 << 24)
if filehandle is None:
tl |= 16777216
await sendall(handle, struct.pack("!I", tl))
await sendall(handle, sdata)
elif isinstance(handle, tuple):
raise Exception("Cannot send filehandle over network socket")
else:
tl |= (2 << 24)
await cloop.sock_sendall(handle, struct.pack("!I", tl))
await send_fds(handle, sdata, [filehandle])
async def _grabhdl(handle, size):
if isinstance(handle, tuple):
return await handle[0].read(size)
else:
cloop = asyncio.get_running_loop()
return await cloop.sock_recv(handle, size)
async def recvall(handle, size):
rd = await _grabhdl(handle, size)
while len(rd) < size:
nd = await _grabhdl(handle, size - len(rd))
if not nd:
raise Exception("Error reading data")
rd += nd
return rd
async def recv(handle):
tl = await _grabhdl(handle, 4)
if not tl:
return None
while len(tl) < 4:
ndata = await _grabhdl(handle, 4 - len(tl))
if not ndata:
raise Exception("Error reading data")
tl += ndata
if len(tl) == 0:
return None
tl = struct.unpack("!I", tl)[0]
if tl & 0b10000000000000000000000000000000:
raise Exception("Protocol Violation, reserved bit set")
# 4 byte tlv
dlen = tl & 16777215 # grab lower 24 bits
datatype = (tl & 2130706432) >> 24 # grab 7 bits from near beginning
if dlen == 0:
return None
if datatype == tlv.Types.filehandle:
if isinstance(handle, tuple):
raise Exception('Filehandle not supported over TLS socket')
msg, filehandles = await recv_fds(handle, dlen, 4)
data = json.loads(bytes(msg))
return ClientFile(data['filename'], data['mode'], filehandles[0])
else:
data = await _grabhdl(handle, dlen)
while len(data) < dlen:
ndata = await _grabhdl(handle, dlen - len(data))
if not ndata:
raise Exception("Error reading data")
data += ndata
if datatype == tlv.Types.text:
return data
elif datatype == tlv.Types.json:
return json.loads(data)
+71 -25
View File
@@ -280,6 +280,7 @@ class Command(object):
def stop_if_noderange_over(self, noderange, maxnodes):
if maxnodes is None:
return
maxnodes = int(maxnodes)
nsize = self.get_noderange_size(noderange)
if nsize > maxnodes:
if nsize == 1:
@@ -391,8 +392,13 @@ class Command(object):
cacert = None
certreqs = ssl.CERT_NONE
knownhosts = True
self.connection = ssl.wrap_socket(self.connection, ca_certs=cacert,
cert_reqs=certreqs)
tlsctx = ssl.create_default_context()
if certreqs == ssl.CERT_NONE:
tlsctx.check_hostname = False
tlsctx.verify_mode = certreqs
if cacert:
tlsctx.load_verify_locations(cacert)
self.connection = tlsctx.wrap_socket(self.connection, server_hostname=server)
if knownhosts:
certdata = self.connection.getpeercert(binary_form=True)
fingerprint = 'sha512$' + hashlib.sha512(certdata).hexdigest()
@@ -467,7 +473,21 @@ def _sort_attrib(k):
return sortutil.naturalize_string('{}'.format(k[1]['sortid']))
return sortutil.naturalize_string(k[0])
def print_attrib_path(path, session, requestargs, options, rename=None, attrprefix=None):
def _is_nondefault(currattr):
if currattr.get('default', None) is None:
return False # no default specified cannot tell
if 'value' not in currattr: # can't compare if no value specified
return False
if currattr['value'] != currattr['default']:
return True
# ok, the potentially pending value is the same, but check active
if 'active' not in currattr:
return False
if currattr['active'] != currattr['default']:
return True
return False
def print_attrib_path(path, session, requestargs, options, rename=None, attrprefix=None, showpending=False):
exitcode = 0
seenattributes = NestedDict()
allnodes = set([])
@@ -495,12 +515,26 @@ def print_attrib_path(path, session, requestargs, options, rename=None, attrpref
currattr = res['databynode'][node][attr]
if show_attr(attr, requestargs, seenattributes, options, node):
if 'value' in currattr:
if currattr['value'] is not None:
val = currattr['value']
if isinstance(val, list):
val = ','.join(val)
currval = currattr.get('value', None)
if isinstance(currval, list):
currval = ','.join(currval)
activeval = currattr.get('active', currval)
if isinstance(activeval, list):
activeval = ','.join(activeval)
if activeval != currval:
if currval is None:
currval = ''
if activeval is None:
activeval = ''
if showpending:
attrout = f'{node}: {printattr}: {activeval} -> {currval}'
else:
attrout = f'{node}: {printattr}: {currval} (Pending)'
elif showpending:
continue
elif currval is not None:
attrout = '{0}: {1}: {2}'.format(
node, printattr, val).strip()
node, printattr, currval).strip()
else:
attrout = '{0}: {1}:'.format(node, printattr)
elif 'isset' in currattr:
@@ -548,15 +582,19 @@ def print_attrib_path(path, session, requestargs, options, rename=None, attrpref
except AttributeError:
exclude = False
if ((requestargs and not exclude) or
(currattr.get('default', None) is not None and
currattr.get('value', None) is not None and
currattr['value'] != currattr['default'])):
_is_nondefault(currattr)):
cval = ','.join(currattr['value']) if isinstance(
currattr['value'], list) else currattr['value']
activeval = ','.join(currattr['active']) if isinstance(
currattr.get('active'), list) else currattr.get('active')
dval = ','.join(currattr['default']) if isinstance(
currattr['default'], list) else currattr['default']
cprint('{0}: {1}: {2} (Default: {3})'.format(
node, printattr, cval, dval))
outmsg = '{0}: {1}: {2} (Default: {3}'.format(
node, printattr, cval, dval)
if 'active' in currattr:
outmsg += ', Pending change from: {0}'.format(activeval)
outmsg += ')'
cprint(outmsg)
else:
try:
@@ -677,23 +715,31 @@ def updateattrib(session, updateargs, nodetype, noderange, options, dictassign=N
sys.stderr.write('Error: ' + res['error'] + '\n')
sys.exit(exitcode)
elif hasattr(options, 'environment') and options.environment:
for key in updateargs[1:]:
key = key.replace('.', '_')
value = os.environ.get(
key, os.environ[key.upper()])
# Let's do one pass to make sure that there's not a usage problem
for key in updateargs[1:]:
key = key.replace('.', '_')
value = os.environ.get(
key, os.environ[key.upper()])
# The environment variable name substitutes '_' for '.'
attrvalues = {}
# Let's do one pass to make sure that there's not a usage problem
for attrib in updateargs[1:]:
envkey = attrib.replace('.', '_')
if envkey in os.environ:
attrvalues[attrib] = os.environ[envkey]
elif envkey.upper() in os.environ:
attrvalues[attrib] = os.environ[envkey.upper()]
else:
sys.stderr.write(
'Error: {0} requested, but neither {1} nor {2} is set in '
'the environment\n'.format(attrib, envkey, envkey.upper()))
sys.exit(1)
for attrib in updateargs[1:]:
if (nodetype == "nodegroups"):
exitcode = session.simple_nodegroups_command(noderange,
'attributes/all',
value, key)
attrvalues[attrib],
attrib)
else:
exitcode = session.simple_noderange_command(noderange,
'attributes/all',
value, key)
attrvalues[attrib],
attrib)
sys.exit(exitcode)
elif dictassign:
for key in dictassign:
@@ -707,7 +753,7 @@ def updateattrib(session, updateargs, nodetype, noderange, options, dictassign=N
if "=" in updateargs[1]:
update_ready = True
for arg in updateargs[1:]:
if not '=' in arg:
if '=' not in arg:
update_ready = False
exitcode = 1
if not update_ready:
+4
View File
@@ -175,6 +175,10 @@ class LogReplay(object):
def _replay_to_console(txtfile, binfile):
if not sys.stdin.isatty():
# Interactive replay needs a terminal to put in raw mode and to take
# navigation keys from, so without one just write the log out
return dump_to_console(txtfile)
replay = LogReplay(txtfile, binfile)
oldtcattr = termios.tcgetattr(sys.stdin.fileno())
tty.setraw(sys.stdin.fileno())
+47 -3
View File
@@ -16,11 +16,55 @@ import fcntl
import sys
import struct
import termios
import select
def get_screengeom():
def get_screengeom(escfallback=False):
# returns height in cells, width in cells, width in pixels, height in pixels
return struct.unpack('hhhh', fcntl.ioctl(sys.stdout, termios.TIOCGWINSZ,
b'........'))
geom = list(struct.unpack('hhhh', fcntl.ioctl(sys.stdout, termios.TIOCGWINSZ,
b'........')))
if escfallback and (geom[0] == 0 or geom[1] == 0):
cellgeom = get_cell_geometry_using_esc_18t()
geom[0], geom[1] = cellgeom
if escfallback and (geom[2] == 0 or geom[3] == 0):
pixgeom = get_pixel_geometry_using_esc_14t()
geom[2], geom[3] = pixgeom
return geom
def get_pixel_geometry_using_esc_14t():
sys.stdout.write('\x1b[14t')
sys.stdout.flush()
rlist, _, _ = select.select([sys.stdin], [], [], 1)
if not rlist:
return 0, 0
response = ''
while True:
c = sys.stdin.read(1)
if c == 't':
break
response += c
if not response.startswith('\x1b[4;'):
return 0, 0
pixgeom = response[4:].split(';')
return int(pixgeom[1]), int(pixgeom[0])
def get_cell_geometry_using_esc_18t():
sys.stdout.write('\x1b[18t')
sys.stdout.flush()
rlist, _, _ = select.select([sys.stdin], [], [], 1)
if not rlist:
return 0, 0
response = ''
while True:
c = sys.stdin.read(1)
if c == 't':
break
response += c
if not response.startswith('\x1b[8;'):
return 0, 0
cellgeom = response[4:].split(';')
return int(cellgeom[0]), int(cellgeom[1])
class ScreenPrinter(object):
def __init__(self, noderange, client, textlen=4):
-2
View File
@@ -138,8 +138,6 @@ class GroupedData(object):
def print_all(self, output=sys.stdout, skipmodal=False, reverse=False,
count=False):
self.generate_byoutput()
modaloutput = None
ismodal = True
if reverse:
outdatalist = sorted(
+2 -6
View File
@@ -19,12 +19,8 @@ import array
import ctypes
import ctypes.util
import confluent.tlv as tlv
try:
import eventlet.green.socket as socket
import eventlet.green.select as select
except ImportError:
import socket
import select
import socket
import select
from datetime import datetime
import json
import os
+346
View File
@@ -0,0 +1,346 @@
import asyncio
from PIL import Image
import io
import numpy as np
import queue
import threading
import time
import zlib
# This results in an RGBA organization of pixels
MYPIXFORMAT = bytearray([
32, # bits per pixel
24, # depth
0, # big endian
1, # true color
0, 255, # red max
0, 255, # green max
0, 255, # blue max
0, 8, 16, # red shift, green shift, blue shift
0, 0, 0 # padding
])
class ByteStream:
def __init__(self):
self.buffer = b''
def add_number(self, number, num_bytes):
data = number.to_bytes(num_bytes, byteorder='big', signed=True)
self.buffer += data
def extend(self, data):
self.buffer += data
def get_bytes(self):
return self.buffer
def clear(self):
self.buffer = b''
def flush(self, writer):
writer.write(self.buffer)
self.clear()
class VNCClient:
async def __aenter__(self):
return self
async def __aexit__(self, exc_type, exc_val, exc_tb):
await self.close()
return False
@classmethod
async def create(cls, url, outputfile=None, fps=15):
self = cls()
self.outputfile = outputfile
self.fps = fps
self.video_writer = None
self._video_size = None
self._last_frame = None
self._last_frame_time = None
self._video_queue = None
self._video_thread = None
self._cv2 = None
if outputfile:
try:
import cv2
except ImportError:
raise ImportError("OpenCV is required for video output but is not installed.")
self._cv2 = cv2
self._video_queue = queue.Queue()
self._video_thread = threading.Thread(
target=self._video_worker, daemon=True)
self._video_thread.start()
if url.startswith('unix://'):
url = url.replace('unix://', '')
if url.startswith('/'):
self.reader, self.writer = await asyncio.open_unix_connection(url)
elif url.startswith('@'):
url = '\0' + url[1:]
self.reader, self.writer = await asyncio.open_unix_connection(url)
elif url.startswith('tcp://'):
url = url.replace('tcp://', '')
host, port = url.split(':')
self.reader, self.writer = await asyncio.open_connection(host, int(port))
else:
raise ValueError('Unsupported URL: {}'.format(url))
self.receiver = None
self.framebuffer = None
self.copytext = None
self._updating = True
self.decompressor = zlib.decompressobj()
self._input_queue = asyncio.Queue()
self._input_task = asyncio.create_task(self._input_worker())
await self._do_vnc_handshake()
return self
async def _input_worker(self):
while True:
keys, modifierkeys = await self._input_queue.get()
payload = ByteStream()
for modkey in (modifierkeys or []):
payload.add_number(4, 1) # Key event
payload.add_number(1, 1) # Down
payload.add_number(0, 2) # Padding
payload.add_number(modkey.value, 4)
payload.flush(self.writer)
await self.writer.drain()
for key in keys:
keynumber = key.value if hasattr(key, 'value') else key
payload.add_number(4, 1) # Key event
payload.add_number(1, 1) # Down
payload.add_number(0, 2) # Padding
payload.add_number(keynumber, 4)
payload.add_number(4, 1) # Key event
payload.add_number(0, 1) # Up
payload.add_number(0, 2) # Padding
payload.add_number(keynumber, 4)
payload.flush(self.writer)
await self.writer.drain()
for modkey in (modifierkeys or []):
payload.add_number(4, 1) # Key event
payload.add_number(0, 1) # Up
payload.add_number(0, 2) # Padding
payload.add_number(modkey.value, 4)
payload.flush(self.writer)
await self.writer.drain()
await asyncio.sleep(0.01) # Have to slow down keypresses for some servers
# Still shouldn't be noticable interactively, but does slow down paste to a fast typist...
self._input_queue.task_done()
async def send_keypresses(self, keys, modifierkeys=None):
await self._input_queue.put((keys, modifierkeys))
async def _read_number(self, num_bytes):
data = await self.reader.readexactly(num_bytes)
return int.from_bytes(data, byteorder='big', signed=True)
def _write_number(self, number, num_bytes):
data = number.to_bytes(num_bytes, byteorder='big', signed=True)
self.writer.write(data)
return data
async def get_screenshot(self):
while self._updating:
await asyncio.sleep(0.1)
await asyncio.sleep(0)
if self.framebuffer is None:
raise Exception('No framebuffer data available')
self._updating = True
return self.framebuffer.copy()
async def _do_vnc_handshake(self):
rfbver = await self.reader.readline()
if not rfbver.startswith(b'RFB 003.008'):
self.writer.close()
await self.writer.wait_closed()
raise Exception('Unsupported RFB version')
self.writer.write(b'RFB 003.008\n')
numsectypes = await self._read_number(1)
if not numsectypes:
self.writer.close()
await self.writer.wait_closed()
raise Exception('No security types supported by the server')
sectypes = await self.reader.readexactly(numsectypes)
sectypes = bytearray(sectypes)
secresult = 1
if 1 in sectypes:
self.writer.write(b'\x01')
await self.writer.drain()
secresult = await self._read_number(4) # Security result
if secresult != 0:
self.writer.close()
await self.writer.wait_closed()
raise Exception('VNC authentication failed')
self.writer.write(b'\x01') # Share display
self.width = await self._read_number(2)
self.height = await self._read_number(2)
pixformat = await self.reader.readexactly(16)
name_length = await self._read_number(4)
self.name = await self.reader.readexactly(name_length)
payload = ByteStream()
if pixformat != MYPIXFORMAT:
payload.add_number(0, 1) # Set pixel format
payload.add_number(0, 3) # Padding
payload.extend(MYPIXFORMAT)
payload.flush(self.writer)
self.receiver = asyncio.create_task(self._receive_loop())
payload.add_number(2, 1) # Set encodings
payload.add_number(0, 1) # Padding
payload.add_number(4, 2) # Number of encodings
payload.add_number(6, 4) # zlib
payload.add_number(7, 4) # tight
payload.add_number(-223, 4) # desktopsize
payload.add_number(-308, 4) # extended desktopsize
payload.flush(self.writer)
self._request_screen_update(incremental=False)
def _request_screen_update(self, incremental=True):
incremental = 1 if incremental else 0
payload = ByteStream()
payload.add_number(3, 1) # Framebuffer update request
payload.add_number(incremental, 1) # Incremental
payload.add_number(0, 2) # x position
payload.add_number(0, 2) # y position
payload.add_number(self.width, 2) # width
payload.add_number(self.height, 2) # height
payload.flush(self.writer)
async def _receive_loop(self):
while True:
try:
message_type = await self._read_number(1)
if message_type == 0: # Framebuffer update
await self._handle_framebuffer_update()
elif message_type == 1: # Set color map entries
raise NotImplementedError('Set color map entries not implemented')
elif message_type == 2: # Bell
pass
elif message_type == 3: # Server cut text
padding = await self._read_number(3)
length = await self._read_number(4)
self.copytext = await self.reader.readexactly(length)
else:
raise Exception(f'Unknown message type: {message_type}')
except Exception as e:
print(f"Error in receive loop: {e}")
break
async def _handle_framebuffer_update(self):
_ = await self._read_number(1) # Padding
num_rects = await self._read_number(2)
self._updating = True
for _ in range(num_rects):
await self._handle_rectangle()
self._updating = False
self._write_video_frame()
self._request_screen_update(incremental=True)
def _write_video_frame(self):
if not self._cv2 or self.framebuffer is None:
return
# Snapshot the framebuffer now and hand it to the writer thread. Frames
# captured while a write is in progress simply queue up behind it.
frame = np.ascontiguousarray(
np.array(self.framebuffer.convert('RGB'))[:, :, ::-1])
self._video_queue.put((frame, time.monotonic()))
def _video_worker(self):
cv2 = self._cv2
while True:
frame, now = self._video_queue.get()
if frame is None:
# Flush the final frame for the time it stayed on screen
if self.video_writer is not None and self._last_frame is not None:
nframes = max(1, round(
(now - self._last_frame_time) * self.fps))
for _ in range(nframes):
self.video_writer.write(self._last_frame)
if self.video_writer is not None:
self.video_writer.release()
self.video_writer = None
return
if self.video_writer is None:
self._video_size = (frame.shape[1], frame.shape[0])
fourcc = cv2.VideoWriter_fourcc(*'mp4v')
self.video_writer = cv2.VideoWriter(
self.outputfile, fourcc, self.fps, self._video_size)
if (frame.shape[1], frame.shape[0]) != self._video_size:
frame = cv2.resize(frame, self._video_size)
if self._last_frame is None:
self._last_frame = frame
self._last_frame_time = now
continue
# Hold the previous frame for the real time it was displayed
nframes = max(1, round((now - self._last_frame_time) * self.fps))
for _ in range(nframes):
self.video_writer.write(self._last_frame)
self._last_frame = frame
self._last_frame_time = now
async def _handle_rectangle(self):
if self.framebuffer is None:
self.framebuffer = Image.new('RGBA', (self.width, self.height))
x = await self._read_number(2)
y = await self._read_number(2)
width = await self._read_number(2)
height = await self._read_number(2)
encoding_type = await self._read_number(4)
pixel_data = None
if encoding_type == 6:
compressed_data_length = await self._read_number(4)
compressed_data = await self.reader.readexactly(compressed_data_length)
# Decompress the data using zlib and store it in the framebuffer
pixel_data = self.decompressor.decompress(compressed_data)
elif encoding_type == 0:
pixel_data = await self.reader.readexactly(width * height * 4) # Assuming 32 bits per pixel
if encoding_type in (-223, -308): # desktopsize
self.width = width
self.height = height
self.framebuffer = Image.new('RGBA', (self.width, self.height))
if encoding_type == -308:
nscreens = await self._read_number(1)
_ = await self._read_number(3) # padding
for _ in range(nscreens):
_ = await self.reader.readexactly(16) # screen info
elif pixel_data:
pixel_data = np.frombuffer(pixel_data, dtype=np.uint8).reshape((height, width, 4)).copy()
pixel_data[:, :, 3] = 0xff
img = Image.fromarray(pixel_data, 'RGBA')
self.framebuffer.paste(img, (x, y))
elif encoding_type == 7: # tight
# Best document I could see was:
# https://github.com/TurboVNC/tightvnc/blob/main/vnc_winsrc/rfb/rfbproto.h
tightheader = await self._read_number(1)
streamid = tightheader & 0x0F
if streamid:
raise NotImplementedError('tight encoding with streamid not implemented')
comptype = (tightheader >> 4) & 0x0F
if comptype != 9:
raise NotImplementedError(f'tight encoding with comptype {comptype} not implemented')
compressed_data_length = await self._read_tight_length()
compressed_data = await self.reader.readexactly(compressed_data_length)
with io.BytesIO(compressed_data) as jpgimg:
img = Image.open(jpgimg)
img.load()
self.framebuffer.paste(img, (x, y))
else:
raise Exception(f'Unsupported encoding type: {encoding_type}')
async def _read_tight_length(self):
length = 0
for i in range(3):
byte = await self._read_number(1)
length |= ((byte & 0x7F) << (i * 7))
if not (byte & 0x80):
break
return length
async def close(self):
if self._video_thread is not None:
# Signal the writer thread to flush and finalize the file
self._video_queue.put((None, time.monotonic()))
await asyncio.to_thread(self._video_thread.join)
self._video_thread = None
self.writer.close()
await self.writer.wait_closed()
+14 -1
View File
@@ -11,7 +11,7 @@ Name: %{name}
Version: %{version}
Release: %{release}
Source0: %{name}-%{fversion}.tar.gz
License: Apache2
License: Apache-2.0
Group: Development/Libraries
BuildRoot: %{_tmppath}/%{name}-%{version}-%{release}-buildroot
Prefix: %{_prefix}
@@ -19,6 +19,13 @@ BuildArch: noarch
Vendor: Lenovo
Url: http://github.com/lenovo/confluent
Obsoletes: confluent_common
# SUSE keeps Pillow's upstream capitalisation and rpm capability names are
# case sensitive, so the lowercase EL spelling resolves to nothing there.
%if 0%{?suse_version}
Requires: python3-numpy python3-Pillow python3-matplotlib
%else
Requires: python3-numpy python3-pillow python3-matplotlib
%endif
%description
This package enables python development and command line access to
@@ -41,6 +48,12 @@ python2 setup.py install --single-version-externally-managed -O1 --root=$RPM_BUI
python3 setup.py install --single-version-externally-managed -O1 --root=$RPM_BUILD_ROOT --record=INSTALLED_FILES --install-scripts=/opt/confluent/bin --install-purelib=/opt/confluent/lib/python
%endif
%if 0%{?suse_version}
# openSUSE patches setuptools to write a literal "#!python" placeholder into
# installed scripts and expects the packager to rewrite it. The stock
# %%python3_fix_shebang only looks at %%{_bindir}, so do it for our prefix.
sed -i '1s|^#!python|#!/usr/bin/python3|' $RPM_BUILD_ROOT/opt/confluent/bin/*
%endif
%clean
rm -rf $RPM_BUILD_ROOT
+55 -3
View File
@@ -19,6 +19,7 @@ alias nodebmcreset='CURRENT_CMDLINE=$(HISTTIMEFORMAT= builtin history 1); export
alias nodeboot='CURRENT_CMDLINE=$(HISTTIMEFORMAT= builtin history 1); export CURRENT_CMDLINE; nodeboot'
alias nodeconfig='CURRENT_CMDLINE=$(HISTTIMEFORMAT= builtin history 1); export CURRENT_CMDLINE; nodeconfig'
alias nodeconsole='CURRENT_CMDLINE=$(HISTTIMEFORMAT= builtin history 1); export CURRENT_CMDLINE; nodeconsole'
alias nodecertutil='CURRENT_CMDLINE=$(HISTTIMEFORMAT= builtin history 1); export CURRENT_CMDLINE; nodecertutil'
alias nodedeploy='CURRENT_CMDLINE=$(HISTTIMEFORMAT= builtin history 1); export CURRENT_CMDLINE; nodedeploy'
alias nodedefine='CURRENT_CMDLINE=$(HISTTIMEFORMAT= builtin history 1); export CURRENT_CMDLINE; nodedefine'
alias nodeeventlog='CURRENT_CMDLINE=$(HISTTIMEFORMAT= builtin history 1); export CURRENT_CMDLINE; nodeeventlog'
@@ -52,6 +53,8 @@ _confluent_get_args()
CMPARGS+=("")
fi
GENNED=""
# Whitespace separates candidate groups, while commas group synonyms.
# shellcheck disable=SC2068
for CAND in ${COMP_CANDIDATES[@]}; do
candarray=(${CAND//,/ })
matched=0
@@ -153,14 +156,22 @@ _confluent_osimage_completion()
{
_confluent_get_args
if [ $NUMARGS == 2 ]; then
COMPREPLY=($(compgen -W "initialize import importcheck updateboot rebase" -- ${COMP_WORDS[COMP_CWORD]}))
COMPREPLY=($(compgen -W "initialize import importcheck updateboot rebase fetch" -- ${COMP_WORDS[COMP_CWORD]}))
return
elif [ ${CMPARGS[1]} == 'initialize' ]; then
COMPREPLY=($(compgen -W "-h -u -s -t -i" -- ${COMP_WORDS[COMP_CWORD]}))
COMPREPLY=($(compgen -W "-h -a -g -u -s -k -t -p -i -l -r" -- ${COMP_WORDS[COMP_CWORD]}))
elif [ ${CMPARGS[1]} == 'import' ] || [ ${CMPARGS[1]} == 'importcheck' ]; then
compopt -o default
COMPREPLY=()
return
elif [ ${CMPARGS[1]} == 'fetch' ]; then
if [ $NUMARGS == 3 ]; then
COMPREPLY=($(compgen -W "$(osdeploy getfetchable)" -- "${COMP_WORDS[COMP_CWORD]}"))
return
fi
compopt -o dirnames
COMPREPLY=()
return
elif [ ${CMPARGS[1]} == 'updateboot' -o ${CMPARGS[1]} == 'rebase' ]; then
COMPREPLY=($(compgen -W "-n $(confetty show /deployment/profiles|sed -e 's/\///')" -- "${COMP_WORDS[COMP_CWORD]}"))
return
@@ -293,6 +304,30 @@ _confluent_nodegroupattrib_completion()
COMP_CANDIDATES=$(nodegroupattrib 'everything' all | awk '{print $2}'|sed -e 's/://')
_confluent_generic_ng_completion
}
_confluent_nodecertutil_completion()
{
_confluent_get_args
if [ $NUMARGS -lt 3 ]; then
_confluent_nr_completion
return
fi
if [ $NUMARGS == 3 ]; then
COMPREPLY=($(compgen -W "installbmccacert removebmccacert listbmccacerts signbmccert" -- ${COMP_WORDS[COMP_CWORD]}))
return
fi
if [ ${CMPARGS[2]} == 'installbmccacert' ]; then
compopt -o default
COMPREPLY=()
return
fi
if [ ${CMPARGS[2]} == 'signbmccert' ]; then
COMP_CANDIDATES=("--days --added-names")
_confluent_get_args
COMPREPLY=($(compgen -W "$GENNED" -- ${COMP_WORDS[COMP_CWORD]}))
fi
}
_confluent_nn_completion()
{
_confluent_get_args
@@ -308,6 +343,22 @@ _confluent_nn_completion()
COMPREPLY=($(compgen -W "$(nodelist | sed -e s/^/$PREFIX/)" -- "${COMP_WORDS[COMP_CWORD]}"))
}
_confluent_nodeconsole_completion()
{
_confluent_get_args
if [ ${CMPARGS[-2]} == '-a' ] || [ ${CMPARGS[-2]} == '--automation' ]; then
compopt -o default
COMPREPLY=()
return
fi
if [[ ${COMP_WORDS[COMP_CWORD]} == -* ]]; then
COMPREPLY=($(compgen -W "-a --automation -e --headless -t --tile -l --log -T --Timestamp -s --screenshot -i --interval -v --video -w --windowed" -- "${COMP_WORDS[COMP_CWORD]}"))
return
fi
_confluent_nn_completion
}
_confluent_nr_completion()
{
CMPARGS=($COMP_LINE)
@@ -362,7 +413,7 @@ complete -F _confluent_nodeattrib_completion nodeattrib
complete -F _confluent_nr_completion nodebmcreset
complete -F _confluent_nodesetboot_completion nodeboot
complete -F _confluent_nr_completion nodeconfig
complete -F _confluent_nn_completion nodeconsole
complete -F _confluent_nodeconsole_completion nodeconsole
complete -F _confluent_define_completion nodedefine
complete -F _confluent_define_completion nodegroupdefine
complete -F _confluent_nr_completion nodeeventlog
@@ -375,6 +426,7 @@ complete -F _confluent_ng_completion nodegroupremove
complete -F _confluent_nr_completion nodehealth
complete -F _confluent_nodeidentify_completion nodeidentify
complete -F _confluent_nr_completion nodeinventory
complete -F _confluent_nodecertutil_completion nodecertutil
complete -F _confluent_nodeattrib_completion nodelist
complete -F _confluent_nodemedia_completion nodemedia
complete -F _confluent_nodepower_completion nodepower
+1 -1
View File
@@ -42,7 +42,7 @@ commands.
* `show` **ELEMENT**, `cat` **ELEMENT**:
Display the result of reading a specific element (by full or relative path)
* `unset` **ELEMENT** **ATTRIBUTE**
For an element with attributes, request to clear the value of the attribue
For an element with attributes, request to clear the value of the attribute
* `set` **ELEMENT** **ATTRIBUTE**=**VALUE**
Set the specified attribute to the given value
* `start` **ELEMENT**
@@ -0,0 +1,35 @@
confluent2ansible(8) -- Export confluent node inventory to an Ansible hosts file
================================================================================
## SYNOPSIS
`confluent2ansible <noderange> -o <ansible.hosts>` [`-a`]
## DESCRIPTION
`confluent2ansible` reads the nodes matched by `<noderange>` from confluent and
writes a corresponding Ansible inventory (hosts) file, allowing an existing
confluent inventory to be used directly by Ansible.
## OPTIONS
* `-o FILE`, `--output=FILE`:
Write the Ansible hosts file to FILE.
* `-a`, `--append`:
Append to an existing hosts file rather than overwriting it.
* `-h`, `--help`:
Show help message and exit
## EXAMPLES
* Write an Ansible hosts file for the compute group:
`# confluent2ansible compute -o /etc/ansible/hosts`
* Append a rack to an existing hosts file:
`# confluent2ansible rack3 -o /etc/ansible/hosts -a`
## SEE ALSO
confluent2xcat(8), confluent2lxca(8), nodelist(8)
@@ -0,0 +1,243 @@
confluent2dnsmasq(8) -- Generate dnsmasq static DHCP reservations for nodes
===============================================================================
## SYNOPSIS
`confluent2dnsmasq [options] <noderange>`
`confluent2dnsmasq <noderange>`
`confluent2dnsmasq --all-options <noderange>`
`confluent2dnsmasq -n <noderange>`
## DESCRIPTION
`confluent2dnsmasq` generates a dnsmasq configuration fragment of static
DHCP reservations for a noderange, using the confluent database as the source
of truth. Each `net.<network>.<attribute>` group is pulled together and one
`dhcp-host=` reservation is emitted for **every network that has a hardware
address** (`net.hwaddr`, `net.bmc.hwaddr`, `net.eth.mynetwork.hwaddr`, and so
on; the network component may itself contain several dotted parts). Networks
without a `hwaddr` (for example InfiniBand interfaces identified only by a
GUID) are skipped.
For each reservation, `net.<network>.ipv4_address` supplies the reserved
address. The `/prefixlen` suffix on that address is used to compute the
enclosing subnet; if the address carries no `/prefixlen`, the prefix length is
filled in from the matching directly-attached route in the host's routing
table. One `dhcp-range=<network>,static,<netmask>,<lease>` line is emitted per
distinct subnet. dnsmasq will not answer DHCP requests on a subnet that has no
covering `dhcp-range`, even when every address on it is statically reserved, so
these ranges are emitted by default (see `--no-range`). An address that has
neither an explicit prefix nor a matching local route still yields a
`dhcp-host` reservation, but no subnet or `dhcp-range` is derived for it and a
warning is printed.
The reservation name is the first token of `net.<network>.hostname`, falling
back to the node name itself for the primary (unnamed) network, and omitted for
a named network that has no hostname. Only IPv4 node addresses are used for
reservations and ranges.
By default the fragment begins with a `bind-dynamic` directive so that dnsmasq
can bind DHCP to networks confluent is using, too (see `--no-bind-dynamic`).
Also by default, a `listen-address` line is emitted for `127.0.0.1`, for
`::1`, and for this host's own address on each managed subnet (see
`--no-listen-address`). The per-subnet addresses are read from the routing
table of the host the command runs on, so run it on the dnsmasq host; only
subnets the host is directly attached to are added. Besides telling dnsmasq
where to listen, the presence of any `listen-address` line disables the default
`local-service` (or `local-service=host`) restriction in `dnsmasq.conf` that
would otherwise limit dnsmasq to localhost.
Additional data can optionally be added to the DHCP reply from the confluent
database (see OPTIONS; all are off by default). Each such option is scoped to
the subnet it belongs to using a dnsmasq tag: the `dhcp-host` lines for that
subnet are given a `set:<tag>` and the corresponding `dhcp-option` lines a
matching `tag:<tag>`. Values are aggregated per subnet; if two nodes on the
same subnet disagree (for example two different gateways), a warning is printed
and the first value is used.
The configuration is written to `/etc/dnsmasq.d/confluent-dhcp.conf` (override
with `--target`) and regenerated on each run, replacing the file atomically.
If the existing file was generated with different content-affecting arguments
(the preview and confirmation flags `-n` and `-y` are ignored for this
comparison), you are asked to confirm before it is overwritten; the prompt is
skipped with `-y` and is refused when there is no controlling terminal. If a
run produces no reservations at all (for example an empty noderange or a failed
read), `confluent2dnsmasq` refuses to overwrite the existing file unless
`--allow-empty` is given.
## OPTIONS
* `--target PATH`:
Path of the configuration file to write. Defaults to
`/etc/dnsmasq.d/confluent-dhcp.conf`.
* `--tag-prefix PREFIX`:
Prefix used when generating dnsmasq tag names for per-subnet options.
Defaults to `cdhcp`, yielding tags such as `cdhcp_10_28_104_0_21`.
* `--no-bind-dynamic`:
Do not emit the `bind-dynamic` line. It is emitted by default. Note that
`bind-dynamic` is a global dnsmasq directive and must not be combined with
a `bind-interfaces` directive set elsewhere. Without it, dnsmasq may conflict
with confluent's own DHCP unless `bind-dynamic` is set elsewhere in the
dnsmasq configuration; a warning is printed.
* `--no-listen-address`:
Do not autodetect or emit `listen-address` lines. By default a
`listen-address` line is emitted for `127.0.0.1`, for `::1`, and for this
host's address on each managed subnet. With this flag you must set
`interface`, `except-interface`, or `listen-address` yourself (or disable
`local-service`/`local-service=host` in `dnsmasq.conf`) for dnsmasq to serve
the cluster networks alongside confluent; a warning is printed as a reminder.
* `--no-range`:
Do not emit `dhcp-range` lines, leaving range and subnet declarations to be
managed elsewhere. Remember that dnsmasq will not serve DHCP on a subnet
that has no covering `dhcp-range`. Per-subnet options are still scoped via
host tags, so they keep working with a range you declare yourself.
* `--lease TIME`:
Lease time applied to the emitted `dhcp-range` lines. Defaults to `24h`.
Accepts any dnsmasq lease syntax (for example `1h`, `24h`, `infinite`).
Pass an empty string to omit the lease field entirely.
* `-n`, `--stdout`:
Write the configuration to standard output instead of the target file, which
is left untouched. Useful for previewing.
* `-y`, `--yes`:
Do not prompt for confirmation when the existing configuration was generated
with different settings; regenerate anyway. Without it a settings change is
confirmed interactively, and refused when there is no controlling terminal.
* `--allow-empty`:
Write the file even when no reservations were produced. By default an empty
result does not overwrite the existing file, to avoid discarding a good
configuration after an empty or failed read.
* `--all-options`:
Enable all of the optional reply-data options below at once.
* `--gateway`, `--router`:
Add a router (DHCP option 3) per subnet, taken from
`net.<network>.ipv4_gateway`.
* `--dns`:
Add DNS servers (DHCP option 6) per subnet, taken from `dns.servers`.
* `--domain`:
Add a domain name (DHCP option 15) per subnet, taken from `dns.domain`.
* `--ntp`:
Add NTP servers (DHCP option 42) per subnet, taken from `ntp.servers`.
* `--mtu`:
Add an interface MTU (DHCP option 26) per subnet, taken from
`net.<network>.mtu`.
* `-h`, `--help`:
Show a help message and exit.
## EXAMPLES
* Generate reservations for a noderange and write the default file:
`# confluent2dnsmasq everything`
* Preview the configuration for one node without writing anything:
`# confluent2dnsmasq -n node01`
* Include all optional reply data (gateway, DNS, domain, NTP, MTU):
`# confluent2dnsmasq --all-options everything`
* Include only the gateway, and use an eight hour lease:
`# confluent2dnsmasq --gateway --lease 8h everything`
* Regenerate non-interactively (for example from cron) after changing options, without the confirmation prompt:
`# confluent2dnsmasq --all-options -y everything`
* Emit only reservations (no ranges) to a custom file kept beside your own range and listen interface definitions:
`# confluent2dnsmasq --no-listen-address --no-range --target /etc/dnsmasq.d/confluent-reservations.conf everything`
## GENERATED CONFIGURATION
Given the following confluent attributes for node `node01`:
node01: net.hwaddr: 10:ff:e0:af:af:f5
node01: net.ipv4_address: 10.28.90.1/21
node01: net.ipv4_gateway: 10.28.88.1
node01: net.bmc.hwaddr: 10:ff:e0:a4:cf:b6
node01: net.bmc.hostname: node01-bmc
node01: net.bmc.ipv4_address: 10.28.106.1/21
node01: net.bmc.ipv4_gateway: 10.28.104.1
node01: net.ib0.hostname: node01-ib0
node01: net.ib0.ipv4_address: 10.28.101.1/21
`confluent2dnsmasq node01`, run on a dnsmasq host whose own addresses on the
two managed subnets are `10.28.88.250` and `10.28.104.250`, writes the file
below. The `ib0` network is skipped because it has no `hwaddr`; the two
remaining networks each yield a subnet (with its `dhcp-range`) and a
reservation, and a `listen-address` line is emitted for localhost and for this
host's address on each subnet:
# Managed by confluent2dnsmasq -- DO NOT EDIT BY HAND.
# Regenerate: confluent2dnsmasq node01
# Generated 2026-06-28 12:00:00 +0000
# Allow confluent and dnsmasq to share the same network for DHCP
bind-dynamic
# Listen on localhost plus this host's address on each managed subnet,
# so dnsmasq serves these networks alongside confluent and bind-dynamic
# (a listen-address line also overrides local-service in dnsmasq.conf).
listen-address=127.0.0.1
listen-address=::1
listen-address=10.28.88.250
listen-address=10.28.104.250
# subnet 10.28.88.0/21
dhcp-range=10.28.88.0,static,255.255.248.0,24h
dhcp-host=10:ff:e0:af:af:f5,10.28.90.1,node01
# subnet 10.28.104.0/21
dhcp-range=10.28.104.0,static,255.255.248.0,24h
dhcp-host=10:ff:e0:a4:cf:b6,10.28.106.1,node01-bmc
Adding `--gateway` scopes a router option to each subnet through a tag, and the
matching `dhcp-host` lines gain a `set:` tag so the option reaches them:
# subnet 10.28.88.0/21
dhcp-range=10.28.88.0,static,255.255.248.0,24h
dhcp-option=tag:cdhcp_10_28_88_0_21,option:router,10.28.88.1
dhcp-host=10:ff:e0:af:af:f5,set:cdhcp_10_28_88_0_21,10.28.90.1,node01
# subnet 10.28.104.0/21
dhcp-range=10.28.104.0,static,255.255.248.0,24h
dhcp-option=tag:cdhcp_10_28_104_0_21,option:router,10.28.104.1
dhcp-host=10:ff:e0:a4:cf:b6,set:cdhcp_10_28_104_0_21,10.28.106.1,node01-bmc
## NOTES
`bind-dynamic` is a global dnsmasq option; if it is set here it must not be
contradicted by a `bind-interfaces` directive elsewhere in the configuration.
The autodetected `listen-address` lines double as the mechanism that relaxes
dnsmasq's default `local-service`/`local-service=host` restriction (any
`interface`, `except-interface`, or `listen-address` line does), which is what
lets dnsmasq answer on the cluster networks instead of only localhost.
Detection only covers subnets the host running the command is directly attached
to, so run it there; with `--no-listen-address` you must supply `interface`,
`except-interface`, or `listen-address` yourself, or disable `local-service`.
After (re)generating the file, validate it with `dnsmasq --test` and restart
dnsmasq (for example `systemctl restart dnsmasq`); a reload (SIGHUP) does not
re-read `listen-address` or interface bindings.
## FILES
* `/etc/dnsmasq.d/confluent-dhcp.conf`:
Default output file (override with `--target`).
## SEE ALSO
nodeattrib(8), noderange(5), dnsmasq(8)
@@ -0,0 +1,30 @@
confluent2lxca(8) -- Export confluent nodes to a Lenovo XClarity Administrator bulk import file
==============================================================================================
## SYNOPSIS
`confluent2lxca <noderange> -o <bulkimport.csv>`
## DESCRIPTION
`confluent2lxca` reads the nodes matched by `<noderange>` from confluent and
writes a CSV file suitable for bulk import into Lenovo XClarity Administrator
(LXCA), so that hardware already known to confluent can be brought under LXCA
management.
## OPTIONS
* `-o FILE`, `--output=FILE`:
Write the XClarity Administrator bulk import CSV to FILE.
* `-h`, `--help`:
Show help message and exit
## EXAMPLES
* Generate a bulk import file for all nodes:
`# confluent2lxca everything -o bulkimport.csv`
## SEE ALSO
confluent2xcat(8), confluent2ansible(8), nodelist(8)
@@ -0,0 +1,36 @@
confluent2xcat(8) -- Export confluent nodes to an xCAT stanza definition
========================================================================
## SYNOPSIS
`confluent2xcat <noderange> -o <xcatnodes.def>` [`-m <macs.csv>`]
## DESCRIPTION
`confluent2xcat` reads the nodes matched by `<noderange>` from confluent and
writes an xCAT object definition stanza file, easing interoperation with or
migration to/from an xCAT environment. Optionally it can also write an xCAT
`macs.csv` file capturing the MAC addresses of the nodes.
## OPTIONS
* `-o FILE`, `--output=FILE`:
Write the xCAT stanza definition to FILE.
* `-m FILE`, `--macs=FILE`:
Write an xCAT macs.csv file to FILE.
* `-h`, `--help`:
Show help message and exit
## EXAMPLES
* Export all nodes to an xCAT stanza file:
`# confluent2xcat everything -o xcatnodes.def`
* Also capture MAC addresses:
`# confluent2xcat everything -o xcatnodes.def -m macs.csv`
## SEE ALSO
confluent2ansible(8), confluent2lxca(8), nodelist(8)
+57 -9
View File
@@ -3,16 +3,26 @@ confluentdbutil(8) -- Backup or restore confluent database
## SYNOPSIS
`confluentdbutil [options] [dump|restore] <path>`
`confluentdbutil [options] [dump|restore|merge] <path>`
`confluentdbutil [-u] [-v] showattrib <noderange> <attribute>...`
## DESCRIPTION
**confluentdbutil** is a utility to export/import the confluent attributes
to/from json files. The path is a directory that holds the json version.
In order to perform restore, the confluent service must not be running. It
is required to indicate how to treat the usernames/passwords are treated in
is required to indicate how the usernames/passwords are treated in
the json files (password protected, removed from the files, or unprotected).
The `showattrib` subcommand prints the stored value of any attribute for the
given noderange, reading directly from the configuration store. Because it
does not talk to the confluent daemon, it can be used to read attribute values
even while confluent is stopped. It must be run directly on the confluent
server as root. By default, `secret.*` and `crypted.*` values are masked as
`********`, just as nodeattrib(8) shows them. With `-u`, they are revealed:
`secret.*` values decrypted to plaintext and `crypted.*` values as their
stored one-way hashes.
## OPTIONS
* `-p PASSWORD`, `--password=PASSWORD`:
@@ -24,10 +34,13 @@ the json files (password protected, removed from the files, or unprotected).
* `-r`, `--redact`:
Indicates to replace usernames and passwords with a dummy string rather
than included.
than including them.
* `-u`, `--unprotected`:
The keys.json file will include the encryption keys without any protection.
With `dump`, the keys.json file will include the encryption keys without
any protection. With `showattrib`, show `secret.*` values decrypted to
plaintext and `crypted.*` values as their stored hashes rather than
masking them as `********`.
* `-s`, `--skipkeys`:
This specifies to dump the encrypted data without
@@ -35,11 +48,46 @@ the json files (password protected, removed from the files, or unprotected).
suitable for an automated incremental backup, where an
earlier password protected dump has a protected
keys.json file, and only the protected data is needed.
keys do not change and as such they do not require
Keys do not change and as such they do not require
incremental backup.
* `-y`, `--yaml
* `-x ATTRIBUTE`, `--exclude=ATTRIBUTE`:
Exclude matching node and node group attributes from `dump`, `restore`,
or `merge`.
The option may be specified multiple times. Attribute names may use
shell-style wildcards such as `net.*`.
A bare namespace such as `net` excludes all attributes below that namespace.
An exclusion omits the attribute from the data being written, it does not
preserve the value already in the database. A `restore` replaces the
database outright, so attributes excluded there are missing from the
restored configuration entirely; use `merge` to leave existing objects
untouched.
Node `groups`, node `id.index`, and node-group `noderange` are retained
so that restore can reconstruct node-group membership and preserve
node index assignments.
During merge, existing nodes and node groups are skipped as whole objects;
exclusions affect only new objects imported from the backup.
* `-y`, `--yaml`:
Use YAML instead of JSON as file format
* `-h`, `--help`:
* `-v`, `--value-only`:
With `showattrib`, print only the values, without node and attribute names
(useful for scripting).
* `-h`, `--help`:
Show help message and exit
## EXAMPLES
* Show the decrypted BMC password confluent uses for a node (run on the
server as root):
`# confluentdbutil -u showattrib n1 secret.hardwaremanagementpassword`
`n1: secret.hardwaremanagementpassword: mypassword123`
* Get just the value, for use in a script:
`# confluentdbutil -u -v showattrib n1 secret.hardwaremanagementpassword`
`mypassword123`
* Dump configuration without (dynamic) deployment state:
`# confluentdbutil -u -x 'deployment.state*' dump /root/confluent-backup`
+41
View File
@@ -0,0 +1,41 @@
dir2img(8) -- Create a disk image from a directory for nodemedia upload
=======================================================================
## SYNOPSIS
`dir2img` [`-e <extra_mb>`] [`-s <total_mb>`] `<directory> <imagefile>` [`<label>`]
## DESCRIPTION
`dir2img` packages the contents of `<directory>` into a FAT-formatted "big
floppy" disk image written to `<imagefile>`. The resulting image is suitable
for upload with nodemedia(8) so that its files are presented to a node as
removable media. An optional volume `<label>` may be supplied.
By default the image is sized to fit the directory contents. The `-e` and `-s`
options allow reserving additional free space.
This command relies on `mtools` (specifically `mformat` and `mcopy`) being
installed.
## OPTIONS
* `-e EXTRA_MB`:
Reserve EXTRA_MB megabytes of free space in the image in addition to the
space required for the directory contents.
* `-s TOTAL_MB`:
Make the image TOTAL_MB megabytes in total, rather than sizing it to the
directory contents.
## EXAMPLES
* Create an image from a directory:
`# dir2img /var/tmp/drivers drivers.img`
* Create a 64 MB labelled image:
`# dir2img -s 64 /var/tmp/drivers drivers.img DRIVERS`
## SEE ALSO
nodemedia(8)
+19
View File
@@ -43,6 +43,16 @@
* `-s`, `--source` <directory>:
Directory to pull installation from, typically a subdirectory of `/var/lib/confluent/distributions`. By default, the repositories for the build system are used. For Ubuntu, this is not supported; the build system repositories are always used.
* `--arch` <architecture>:
Target architecture to build for (`x86_64` or `aarch64`). For Ubuntu, the
Debian-style names `amd64` and `arm64` are accepted as aliases. Building
for a foreign architecture is supported on EL and Ubuntu build hosts and
requires a statically linked qemu-user emulator registered in binfmt_misc
with the F (fix-binary) flag; on Ubuntu this is provided by the
qemu-user-static package (qemu-user-binfmt as of Ubuntu 26.04). For EL,
the architecture is also detected automatically when building from a `-s`
source tree.
* `-y`, `--non-interactive`:
Avoid prompting for confirmation.
@@ -99,6 +109,15 @@ Build a diskless image from a distribution:
imgutil build -s alma-9.6-x86_64 /tmp/myimage
Build an aarch64 EL diskless image on an x86_64 host (the architecture is
detected from the source tree):
imgutil build -s alma-9.6-aarch64 /tmp/myimage
Build an aarch64 Ubuntu diskless image on an amd64 Ubuntu host:
imgutil build --arch aarch64 /tmp/myimage
Execute a shell in an unpacked image:
imgutil exec /tmp/myimage
+1 -1
View File
@@ -9,7 +9,7 @@ nodeapply(8) -- Execute command on many nodes in a noderange through ssh
Provides shortcut access to a number of common operations against deployed
nodes. These operations include refreshing ssh certificates and configuration,
rerunning syncflies, and executing specified postscripts.
rerunning syncfiles, and executing specified postscripts.
## OPTIONS
+37 -3
View File
@@ -24,6 +24,13 @@ For a full list of attributes, run `nodeattrib <node> all` against a node.
If `-c` is specified, this will set the nodeattribute to a null value.
This is different from setting the value to an empty string.
Arbitrary custom attributes can also be created with the `custom.` prefix.
Network attributes (`net.*`) may similarly be qualified with an arbitrary name, or dotted chain of names, between
`net` and the attribute, for example `net.compute.ipv4_address` or `net.eth.compute.ipv4_address` and
`net.eth.management.ipv4_address`. Any such `net.<name*>.*` attributes together describe one logical network
interface, so custom networks may be defined freely by choosing your own name(s).
Attributes may be specified by wildcard, for example `net.*switch` will report
all attributes that begin with `net.` and end with `switch`.
@@ -31,11 +38,15 @@ If the word all is specified, then all available attributes are given.
Omitting any attribute name or the word 'all' will display only attributes
that are currently set.
The values of `secret.*` and `crypted.*` attributes are never shown here; they
can only be retrieved on the confluent server with
`confluentdbutil -u showattrib`.
For the `groups` attribute, it is possible to add a group by doing
`groups,=<newgroup>` and to remove by doing `groups^=<oldgroup>`
Note that `nodeattrib <group>` will likely not provide the expected behavior.
See nodegroupattrib(8) command on how to manage attributes on a group level. Running
See the nodegroupattrib(8) command for how to manage attributes on a group level. Running
nodeattrib on a group will simply set node-specific attributes on each individual
member of the group.
@@ -61,8 +72,8 @@ to a blank value will allow masking a group defined attribute with an empty valu
or environment variables.
* `-s`, `--set`:
Set attributes using a batch file rather than the command line. The attributes in the batch file
can be specified as one line of key=value pairs simmilar to command line or each attribute can
Set attributes using a batch file rather than the command line. The attributes in the batch file
can be specified as one line of key=value pairs similar to command line or each attribute can
be in its own line. Lines that start with # sign will be read as a comment. See EXAMPLES for batch
file syntax.
@@ -111,6 +122,29 @@ to a blank value will allow masking a group defined attribute with an empty valu
`n2: console.method: serial`
`n2: hardwaremanagement.manager: 172.30.3.2`
* Setting up a named network interface as a bond (e.g. a two-port LACP bond used for compute traffic), using an
expression to calculate the IP address:
`# nodeattrib node12 net.compute.team_mode=lacp net.compute.connection_name=bond0 net.compute.interface_names=enp33s0f0np0,enp33s0f1np1 net.compute.ipv4_address='172.17.0.{n1}/16'`
`node12: net.compute.connection_name: bond0`
`node12: net.compute.interface_names: enp33s0f0np0,enp33s0f1np1`
`node12: net.compute.ipv4_address: 172.17.0.12/16`
`node12: net.compute.team_mode: lacp`
* Passing additional settings to the network backend of the deployed OS with `net.extra_settings`
(semicolon-delimited key=value pairs, keys in the native syntax of the respective backend).
On a NetworkManager based OS (e.g. Enterprise Linux), keys are nmcli properties:
`# nodeattrib node12 net.mgmt.extra_settings='connection.zone=internal;ipv4.routes=10.0.0.0/8 192.168.1.254, 172.16.0.0/12 192.168.1.254;ipv4.route-metric=200'`
`node12: net.mgmt.extra_settings: connection.zone=internal;ipv4.routes=10.0.0.0/8 192.168.1.254, 172.16.0.0/12 192.168.1.254;ipv4.route-metric=200`
* On a netplan based OS (e.g. Ubuntu), keys are netplan YAML paths (nested keys dotted, values in YAML flow syntax).
Since Confluent uses braces for attribute expressions, literal braces must be escaped as `{{` and `}}` when setting the attribute:
`# nodeattrib node13 net.mgmt.extra_settings='routes=[{{to: 10.0.0.0/8, via: 192.168.1.254}}];nameservers.search=[lab.example.com]'`
`node13: net.mgmt.extra_settings: routes=[{to: 10.0.0.0/8, via: 192.168.1.254}];nameservers.search=[lab.example.com]`
* On a wicked based OS (e.g. SUSE), keys are ifcfg variables (routes are not supported through this mechanism on wicked):
`# nodeattrib node14 net.mgmt.extra_settings='ZONE=internal;ETHTOOL_OPTIONS=-K iface tso off'`
`node14: net.mgmt.extra_settings: ZONE=internal;ETHTOOL_OPTIONS=-K iface tso off`
* Clear attribute on nodes of a simple noderange, if you want to retain the variable set the attribute to "":
`# nodeattrib n1-n2 -c console.method`
`# nodeattrib n1-n2 console.method`
@@ -44,7 +44,7 @@ It is sometimes the case that the number must be formatted a different way,
either specifying 0 padding or converting to hexadecimal. This can be done by a
number of operators at the end to indicate formatting changes.
`{n1:02x} - Zero pad to two decimal places, and convert to hexadecimal, as mightbe used for generating MAC addresses`
`{n1:02x} - Zero pad to two decimal places, and convert to hexadecimal, as might be used for generating MAC addresses`
`{n1:x} - Hexadecimal without padding, as may be used in a generated IPv6 address`
`{n1:X} - Uppercase hexadecimal`
`{n1:02d} - Zero pad a normal numeric representation of the number.`
@@ -20,9 +20,9 @@ nodebmcpassword(8) -- Change management controller password for a specified user
## EXAMPLES:
* Reset the management controller for nodes n1 through n4:
`# nodebmcreset n1-n4`
`n1: Password Change Successful`
`n2: Password Change Successful`
`n3: Password Change Successful`
`n4: Password Change Successful`
* Change the management controller password for user USERID on nodes n1 through n4:
`# nodebmcpassword n1-n4 USERID newp4ssw0rd`
`n1: Password Change Successful`
`n2: Password Change Successful`
`n3: Password Change Successful`
`n4: Password Change Successful`
@@ -0,0 +1,50 @@
nodecertutil(8) -- Manage BMC CA certificates and sign BMC certificates
=======================================================================
## SYNOPSIS
`nodecertutil <noderange> <command>` [<args>]
## DESCRIPTION
`nodecertutil` manages the certificate authority (CA) certificates trusted by
the baseboard management controllers (BMCs) of the nodes matched by
`<noderange>`, and can sign a BMC certificate. One of the subcommands below
must be supplied.
## COMMANDS
* `installbmccacert <filename>`:
Install the CA certificate contained in `<filename>` onto the BMCs.
* `removebmccacert <id>`:
Remove the CA certificate identified by `<id>` from the BMCs. Use
`listbmccacerts` to discover the identifiers.
* `listbmccacerts`:
List the CA certificates currently installed on the BMCs.
* `signbmccert` [`--days <days>`] [`--added-names <names>`]:
Sign a BMC certificate. `--days` is required and sets the number of days
the certificate is valid for; `--added-names` adds additional subject
alternative names to the certificate.
## OPTIONS
* `-h`, `--help`:
Show help message and exit
## EXAMPLES
* List the CA certificates installed on a node's BMC:
`# nodecertutil n1 listbmccacerts`
* Install a CA certificate on a range of nodes:
`# nodecertutil n1-n4 installbmccacert /etc/confluent/myca.pem`
* Sign a BMC certificate valid for one year:
`# nodecertutil n1 signbmccert --days 365`
## SEE ALSO
nodeconfig(8), nodeattrib(8)
+6 -3
View File
@@ -49,8 +49,11 @@ actually be in effect until a reboot.
* `-r COMPONENT`, `--restoredefault=COMPONENT`:
Request that the specified component of the targeted nodes will have its
configuration reset to default. Currently the only component implemented
is uefi.
configuration reset to default. The component may be "uefi" or "bmc".
* `-p, --pending`:
When possible, indicate only the settings that are awaiting a reboot to take effect. Note that not all nodes will differentiate between
pending and current settings, and for such platforms this will come up empty.
* `-m MAXNODES`, `--maxnodes=MAXNODES`:
Specify a maximum number of nodes to configure, prompting if over
@@ -69,7 +72,7 @@ actually be in effect until a reboot.
`s4: bmc.ipv4_method: DHCP`
`s4: bmc.ipv4_gateway: 172.30.0.6`
* Changing nodes `s3` and `s4` to have the ip addressess 10.1.2.3 and 10.1.2.4 with a 16 bit subnet mask:
* Changing nodes `s3` and `s4` to have the IP addresses 10.1.2.3 and 10.1.2.4 with a 16 bit subnet mask:
`# nodeconfig s3,s4 bmc.ipv4_address=10.1.2.{n1}/16`
## SEE ALSO
+9 -9
View File
@@ -37,7 +37,7 @@ console process which will result in the console window closing.
manager at this time.
* `-T`, `--Timestamp`:
Dump the log with Timpstamps on the current, local log in /var/log/confluent/consoles.
Dump the log with Timestamps on the current, local log in /var/log/confluent/consoles.
If in collective mode, this only makes sense to use on the current collective
manager at this time.
@@ -65,9 +65,9 @@ console process which will result in the console window closing.
environment, to open a set of consoles for a range of
nodes in separate Windows Terminal windows, with the
title set for each node, set **NODECONSOLE_WINDOWED_COMMAND**
to `wt.exe wsl.exe -d AlmaLinux-8 --shell-type login. If the
to `wt.exe wsl.exe -d AlmaLinux-8 --shell-type login`. If the
NODECONSOLE_WINDOWED_COMMAND environment variable isn't set,
xterm will be used bydefault.
xterm will be used by default.
## ESCAPE SEQUENCE COMMANDS
@@ -102,7 +102,7 @@ keystroke will be interpreted as a command. The following commands are availabl
Request immediate boot to network
* `r`:
[send Resize]
This queries the current terminal and sends stty commands to advertise the user termineal
This queries the current terminal and sends stty commands to advertise the user terminal
size to the remote console
* `?`:
Get a list of supported commands
@@ -112,10 +112,10 @@ keystroke will be interpreted as a command. The following commands are availabl
## PASSTHROUGH OPTIONS
While opening a windowed console with xterm or any other console of choice. The
nodeconsole command gives capality to specify passthrough options targeted at
the console. All options after the -- will be parsed the console program. For
example, opening a windowed console using xterm with a black background.
`nodeconconsole -w n1 -- -bg black`
While opening a windowed console with xterm or any other console of choice. The
nodeconsole command gives capability to specify passthrough options targeted at
the console. All options after the -- will be passed to the console program. For
example, opening a windowed console using xterm with a black background.
`nodeconsole -w n1 -- -bg black`
+2 -2
View File
@@ -3,7 +3,7 @@ nodedefine(8) -- Define new confluent nodes
## SYNOPSIS
`nodedefine <noderange> [nodeattribute1=value1> <nodeattribute2=value2> ...]`
`nodedefine <noderange> [<nodeattribute1=value1> <nodeattribute2=value2> ...]`
## DESCRIPTION
@@ -26,4 +26,4 @@ that `nodeattrib(8)` will error if a node does not exist.
## SEE ALSO
noderange(5), nodeattribexpressions(8)
noderange(5), nodeattribexpressions(5)
+3 -3
View File
@@ -27,14 +27,14 @@ deployment status.
Prepare the network services for deployment, but do not interact with BMCs. This is intended for scenarios where
the boot device control and server restart will be handled outside of confluent.
* `-m MAXNODES`, `--maxnodes=MAXNODES`:
Specifiy a maximum nodes to be deployed.
* `-m MAXNODES`, `--maxnodes=MAXNODES`:
Specify a maximum number of nodes to be deployed.
* `-h`, `--help`:
Show help message and exit
## EXAMPLES
* Begin the instalalation of a profile of CentOS 8.2:
* Begin the installation of a profile of CentOS 8.2:
`# nodedeploy d4 -n centos-8.2-x86_64-default`
`d4: network`
`d4: reset`
+5 -2
View File
@@ -35,8 +35,8 @@ devices. Generally every effort is made to passively detect devices as they
become available (as they boot or are plugged in), however sometimes an active
scan is the best approach to catch something that appears to be missing.
**nodedsicover clear** requests the server forget about a selection of
detected device. It takes the same arguments as **nodediscover list**.
**nodediscover clear** requests the server to forget about a selection of
detected devices. It takes the same arguments as **nodediscover list**.
**nodediscover subscribe** and **unsubscribe** instructs confluent to subscribe to or
unsubscribe from the designated switch running affluent with system discovery support.
@@ -75,6 +75,9 @@ the nodes.
## OPTIONS
* `-a`, `--aggressive`:
Perform more aggressive scanning and fingerprinting. This will probe as many addresses as it can find with ssh and https connections. By default more targeted measures are used
that will only tend to probe things that are easily detectable as devices implementing explicit mechanisms for detection and identification.
* `-m MODEL`, `--model=MODEL`:
Operate with nodes matching the specified model number
* `-s SERIAL`, `--serial=SERIAL`:
+35 -4
View File
@@ -3,13 +3,28 @@ nodeeventlog(8) -- Pull eventlog from confluent nodes
## SYNOPSIS
`nodeeventlog [options] <noderange> [clear]`
`nodeeventlog [options] [-s <sources>] <noderange> [clear]`
## DESCRIPTION
`nodeeventlog` pulls and optionally clears the event log from the requested
noderange.
A platform may keep its events in more than one log, and some of those are
noisier than others. Reading gathers every log the platform describes as an
event log, wherever it keeps them, and `-s` narrows the output to the ones
asked for. `-s list` names the logs a node offers rather than showing entries.
A source is named by the entries that come from it, so a log that holds no
entries is not among them, and a node whose logs are all empty names none. An
ipmi managed node keeps its events in the one log the spec gives it, which is
reported as `SEL`.
Clearing is deliberately narrower than reading: it only touches the logs the
platform's own manager publishes, so a log that only a read reaches is never
destroyed by one. A selection cannot be cleared, so `-s` and `clear` are
refused together.
## OPTIONS
* `-m MAXNODES`, `--maxnodes=MAXNODES`:
@@ -20,15 +35,31 @@ noderange.
return the last <n> entries for each node in the eventlog.
* `-t TIMEFRAME`, `--timeframe=TIMEFRAME`:
return entries within a specified timeframe for each node's event log.
This will return entries from the last <n> hours or days. 1h would be
entries from with the last one hour.
return entries within a specified timeframe for each node's event log.
This will return entries from the last <n> hours or days. 1h would be
entries from within the last one hour.
* `-s SOURCES`, `--source=SOURCES`:
only show entries from the named logs, comma delimited and matched without
regard to case. `-s list` names the logs the node offers instead of showing
entries. A platform that named a log "list" would be shadowed by that.
Asking for a log the node named nothing from is reported on stderr and
exits non-zero, rather than looking like an empty log.
* `-h`, `--help`:
Show help message and exit
## EXAMPLES
* Ask which logs a node keeps:
`# nodeeventlog n2 -s list`
`n2: AuditLog,EventLog,Logs,SEL`
* Pull only the sel and the chassis log, leaving the audit noise out:
`# nodeeventlog n2 -s SEL,Logs`
`n2: 07/23/2026 13:03:12 SEL: Processor - Thermal Trip`
`n2: 08/05/2026 04:30:26 Logs: PowerUnit - Power Off`
* Pull the event log from n2 and n3:
`# nodeeventlog n2,n3`
`n2: 05/03/2017 11:44:25 Event Log Disabled - SEL Fullness - Log clear`
+37 -2
View File
@@ -3,7 +3,7 @@ nodefirmware(8) -- Report firmware information on confluent nodes
## SYNOPSIS
`nodefirmware <noderange> [list][updatestatus][update [--backup <file>]]|[<components>]`
`nodefirmware <noderange> [list][updatestatus][updatetypes][update [--backup] [--parameterfile <file>] <file>]|[<components>]`
## DESCRIPTION
@@ -20,6 +20,10 @@ the FPGA version where applicable).
The updatestatus argument will describe the state of firmware updates on the
nodes.
The updatetypes argument will name the kinds of firmware image the platform has
to be told about, for use in a parameter file. Platforms that read the kind of
firmware from the image itself have none to name and say so.
In the update form, it accepts a single file and attempts to update it using
the out of band facilities. Firmware updates can end in one of three states:
@@ -35,12 +39,43 @@ the out of band facilities. Firmware updates can end in one of three states:
* `-m MAXNODES`, `--maxnodes=MAXNODES`:
When updating, prompt if more than the specified number of servers will
be affected
* `-p PARAMETERFILE`, `--parameterfile=PARAMETERFILE`:
For updating, a parameter file to provide along with the update payload. See
PARAMETER FILE below
* `-h`, `--help`:
Show help message and exit
## PARAMETER FILE
The parameter file given with `-p` holds JSON that is sent alongside the image.
Its keys are the parts a Redfish multipart update accepts, so `UpdateParameters`
for the standard fields and `OemParameters` for whatever a vendor asks for on top
of them. Without one, an update names no targets, which means "whatever this
image is for", and that is what most platforms want.
Some platforms will not take an image without being told what kind of firmware it
holds, and refuse the update rather than guess. On an AMI MegaRAC that is an
`ImageType` under `OemParameters`:
{"OemParameters": {"ImageType": "BIOS"}}
`nodefirmware <noderange> updatetypes` names the accepted types without
attempting an update. The value is also checked against the same list before
anything is uploaded, so a name the platform does not accept is refused rather
than sent.
## EXAMPLES
* Ask what kinds of firmware image a node has to be told about:
`# nodefirmware n1 updatetypes`
`n1: BMC,BIOS,MB_CPLD,SCM_CPLD,BPB_CPLD,HPM_BMC,HPM_BIOS,HPM_SCP,HPM_BIOS2`
* Update the BIOS on a platform that has to be told:
`# echo '{"OemParameters": {"ImageType": "BIOS"}}' > /tmp/bios.json`
`# nodefirmware n1 update -p /tmp/bios.json /tmp/image.bin`
* Pull firmware from a node:
`# nodefirmware r1`
`r1: IMM: 3.70 (TCOO26H 2016-11-29T05:09:51)`
@@ -11,7 +11,7 @@ nodegroupattrib(8) -- List or change confluent nodegroup attributes
## DESCRIPTION
`nodegroupattrip` queries the confluent server to get information about nodes.
`nodegroupattrib` queries the confluent server to get information about nodes.
In the simplest form, it simply takes the given group and lists the attributes of that group.
Contrasted with nodeattrib(8), settings managed by nodegroupattrib will be added
@@ -23,6 +23,10 @@ node after using the nodeattrib(8) command will not have attributes change autom
It's easiest to see by using the `nodeattrib <noderange> -b` to understand how
the attributes are set on the node versus a group to which a node belongs.
The values of `secret.*` and `crypted.*` attributes are never shown here; they
can only be retrieved on the confluent server with
`confluentdbutil -u showattrib`.
## OPTIONS
* `-b`, `--blame`:
@@ -3,7 +3,7 @@ nodegroupdefine(8) -- Define new confluent node group
## SYNOPSIS
`nodegroupdefine <groupname> [nodeattribute1=value1> <nodeattribute2=value2> ...]`
`nodegroupdefine <groupname> [<nodeattribute1=value1> <nodeattribute2=value2> ...]`
## DESCRIPTION
@@ -21,4 +21,4 @@ that `nodegroupattrib(8)` will error if a node group does not exist.
## SEE ALSO
nodeattribexpressions(8), nodegroupattrib(8), nodegroupremove(8)
nodeattribexpressions(5), nodegroupattrib(8), nodegroupremove(8)
@@ -0,0 +1,27 @@
nodegrouprename(8) -- Rename a node group in the confluent management service
=============================================================================
## SYNOPSIS
`nodegrouprename <group> <newname>`
## DESCRIPTION
`nodegrouprename` changes the name of an existing node group from `<group>` to
`<newname>`. Group membership and the attributes defined on the group are
preserved under the new name.
## OPTIONS
* `-h`, `--help`:
Show help message and exit
## EXAMPLES
* Rename a group:
`# nodegrouprename compute computenodes`
`compute: computenodes`
## SEE ALSO
nodegroupdefine(8), nodegroupremove(8), nodegroupattrib(8), nodegrouplist(8)
+3 -3
View File
@@ -3,11 +3,11 @@ nodeidentify(8) -- Control the identify LED of confluent nodes
## SYNOPSIS
`nodidentify <noderange> [on|off|blink]`
`nodeidentify <noderange> [on|off|blink]`
## DESCRIPTION
`nodeidentify` allows you to turn on or off the location LED of conflueunt nodes,
`nodeidentify` allows you to turn on or off the location LED of confluent nodes,
making it easier to determine the physical location of the nodes. The following
options are supported:
@@ -24,7 +24,7 @@ options are supported:
`n3: on`
`n4: on`
* Turn off the identify LED on nodes n1 thorugh n4:
* Turn off the identify LED on nodes n1 through n4:
`# nodeidentify n1-n4 off`
`n1: off`
`n2: off`
+1 -1
View File
@@ -9,7 +9,7 @@ nodeinventory(8) -- Get hardware inventory of confluent node
`nodeinventory` pulls information about hardware of a node. This includes
information such as adapters, serial numbers, processors, and memory modules,
as supported by the platforms hardware management implementation. It accepts
as supported by the platform's hardware management implementation. It accepts
arguments such as serial or model or others as listed above to filter
output to specific data.
@@ -1,10 +1,10 @@
nodel2traceroute(8) -- returns the layer 2 route through an Ethernet network managed by confluent given 2 end points.
nodel2traceroute(8) -- Returns the layer 2 route through an Ethernet network managed by confluent given 2 end points.
==============================
## SYNOPSIS
`nodel2traceroute [options] <start_node> <end_noderange>`
## DESCRIPTION
**nodel2traceroute** is a command that returns the layer 2 route for the configered interfaces in nodeattrib.
**nodel2traceroute** is a command that returns the layer 2 route for the configured interfaces in nodeattrib.
It can also be used with the -i and -e options to check against specific interfaces on the endpoints. If the
--interface or --eface option are not used then the command will check for routes against all the defined
interfaces in nodeattrib (net.*.switch) for the nodes.
@@ -17,7 +17,7 @@ interfaces in nodeattrib (net.*.switch) for the nodes.
## OPTIONS
* ` -e` EFACE, --eface=INTERFACE
interface to check against for the second end point or end points if using checking against multiple nodes
interface to check against for the second end point or end points if checking against multiple nodes
* ` -i` INTERFACE, --interface=INTERFACE
interface to check against for the first end point
* ` -c` CUMULUS, --cumulus=CUMULUS
+1 -1
View File
@@ -8,7 +8,7 @@ nodelicense(8) -- Manage license keys on BMC
## DESCRIPTION
`nodelicense` manages license keys on supported BMCs. Without an argument, the command
lists currently installed license. Using `delete` will remove the specified license name
lists currently installed licenses. Using `delete` will remove the specified license name
from the BMC. The `save` subcommand will take the passed directory (which may be in the form
of /path/to/{node}/ to have the node name substituted for each node) and back up installed licenses
to that directory. The `install` command will take the specified filename and install. The filename
+3 -3
View File
@@ -15,7 +15,7 @@ matching nodes, one line at a time.
If a list of node attribute names are given, the value of those are also
displayed. If `-b` is specified, it will also display information on
how inherited and expression based attributes are defined. There is more
information on node attributes in nodeattributes(5) man page.
information on node attributes in the nodeattrib(8) man page.
Attributes may be specified by wildcard, for example `net.*switch` will report
all attributes that begin with `net.` and end with `switch`.
@@ -25,7 +25,7 @@ all attributes that begin with `net.` and end with `switch`.
* `-b`, `--blame`:
Annotate inherited and expression based attributes to show their base value.
* `-d`, `--delim`:
Choose a delimiter to separat the values. Default - ENTER.
Choose a delimiter to separate the values. Default - ENTER.
## EXAMPLES
* Listing matching nodes of a simple noderange:
`# nodelist n1-n4`
@@ -40,7 +40,7 @@ all attributes that begin with `net.` and end with `switch`.
`n2: hardwaremanagement.manager: 172.30.3.2`
* Getting a group of attributes while determining what group defines them:
`# nodelist n1,n2 hardwaremanegement --blame`
`# nodelist n1,n2 hardwaremanagement --blame`
`n1: hardwaremanagement.manager: 172.30.3.1`
`n1: hardwaremanagement.method: ipmi (inherited from group everything)`
`n1: hardwaremanagement.switch: r8e1`
+7
View File
@@ -32,6 +32,13 @@ BMCs map a virtual USB device to that url. Content is loaded on demand, and
as such that URL is referenced potentially once for every IO operation that
the host platform attempts.
## NOTES
When doing an attach of an https:// url, you may hit an error if you have not enrolled your certificate authority.
In a general confluent environment, you can usually address it by:
`# for cert in /var/lib/confluent/public/site/tls/*.pem; do nodecertutil s1-s4 installbmccacert $cert; done`
## OPTIONS
* `-h`, `--help`:
+2 -2
View File
@@ -14,8 +14,8 @@ It can also be used with the `-s` flag to change the ping location to something
* `-h`, `--help`:
Show help message and exit
* `-s` SUBSTITUTENAME, --substitutename=SUBSTITUTENAME
Use a different name other than the nodename for ping. This may be a
expression, such as {bmc} or, if no { character is present, it is treated as a suffix. -s -eth1 would make n1 become n1-eth1, for example.
Use a different name other than the nodename for ping. This may be an
expression, such as {bmc} or, if no { character is present, it is treated as a suffix. -s -eth1 would make n1 become n1-eth1, for example.
## EXAMPLES
+34
View File
@@ -0,0 +1,34 @@
noderename(8) -- Rename nodes in the confluent management service
=================================================================
## SYNOPSIS
`noderename <noderange> <newname>`
## DESCRIPTION
`noderename` changes the name of the node or nodes matched by the given
noderange to `<newname>`. As with other attribute changes, the rename is
applied through the confluent datastore and does not by itself alter the
operating system hostname of a running node.
When renaming more than one node, `<newname>` is normally an expression so that
each matched node receives a distinct name (see noderange(5) and the attribute
expression syntax).
## OPTIONS
* `-h`, `--help`:
Show help message and exit
## EXAMPLES
* Rename a single node:
`# noderename n1 n2`
* Rename a range of nodes using an expression:
`# noderename n1-n4 'compute{n1}'`
## SEE ALSO
nodedefine(8), noderemove(8), nodeattrib(8), noderange(5)
+1 -1
View File
@@ -24,7 +24,7 @@ interval of 1 second is used.
the results, which may have a different ordering than non-CSV usage of nodesensors.
* `-i`, `--interval`=**SECONDS**:
Repeat data gathering waiting, waiting the specified time between samples. Unless `-n` is
Repeat data gathering, waiting the specified time between samples. Unless `-n` is
specified, indefinite retrieval is assumed.
* `-n`, `--numreadings`=**SAMPLES**:
+1 -1
View File
@@ -10,7 +10,7 @@ nodesetboot(8) -- Check or set next boot device for noderange
Requests that the next boot occur from the specified device. Unless otherwise
specified, this is a one time boot option, and does not change the normal boot
behavior of the system. This is useful for taking a system that normally boots
to the hard drive and startking a network install, or to go into the firmware
to the hard drive and starting a network install, or to go into the firmware
setup menu without having to hit a keystroke at the correct time on the console.
Generally, it's a bit more convenient and direct to use the nodeboot(8) command,
+5 -2
View File
@@ -30,13 +30,16 @@ as stderr, unlike psh which combines all stdout and stderr into stdout.
* `-p PORT`, `--port=PORT`
Specify a custom port for ssh
* `-s SUBSTITUTION`, `--substitutename=SUBSTITITUTION`
* `-s SUBSTITUTION`, `--substitutename=SUBSTITUTION`
Specify a substitution name instead of the nodename. If no {} are in the substitution,
it is considered to be an append. For example, '-s -ib' would produce 'node1-ib' from 'node1'.
Full expression syntax is supported, in which case the substitution is considered to be the entire
new name. {node}-ib would be equivalent to -ib. For example, nodeshell -s {bmc} node1
new name. {node}-ib would be equivalent to -ib. For example, nodeshell -s {bmc} node1
would ssh to the BMC instead of the node.
* `-t TIMEOUT`, `--timeout=TIMEOUT`
Timeout in seconds for each node ssh connection attempt
## EXAMPLES
* Running `echo hi` on for nodes:
+3 -3
View File
@@ -3,7 +3,7 @@ stats(8) -- Common basic statistics on typical numeric data in output
## SYNOPSIS
`<other command> | stats [-c N] [-d D] [-x|-g|-t|-o image.png] [-s N] [-v] [-b N]
`<other command> | stats [-c N] [-d D] [-x|-g|-t|-o image.png] [-s N] [-v] [-b N]`
## DESCRIPTION
@@ -11,11 +11,11 @@ The **stats** command helps analyze common numerical data such as performance nu
or temperatures or any other numerical value.
By default it looks for the last numerical output on the first line to identify the numerical column
and analyze that number. This can be overriden by **-c COLUMN** to indicate a column. By default,
and analyze that number. This can be overridden by **-c COLUMN** to indicate a column. By default,
whitespace and commas are treated to delimit columns, and **-d DELIMITER** can be used to override.
By default it outputs basic statistics, but a histogram is available either text or through X11 output
or sixel or output to an image file depending on whether **-x**, **-g**, **-t*, or **-o image.png** is
or sixel or output to an image file depending on whether **-x**, **-g**, **-t**, or **-o image.png** is
used.
## OPTIONS
+14 -1
View File
@@ -5,7 +5,20 @@ cd `dirname $0`/doc/man
mkdir -p ../../man/man1
mkdir -p ../../man/man5
mkdir -p ../../man/man8
ronn -r *.ronn
for ronn in *.ronn; do
# The first heading line is of the form "name(section) -- description"
header=`grep -m1 -E '\([0-9]+\) *--+ ' "$ronn"`
name=`echo "$header" | sed -E 's/^#* *([A-Za-z0-9._-]+)\(([0-9]+)\).*/\1/'`
section=`echo "$header" | sed -E 's/^#* *([A-Za-z0-9._-]+)\(([0-9]+)\).*/\2/'`
# Strip the ronn title heading (and any setext underline) and promote the
# remaining section headings so pandoc emits them as man .SH sections.
sed -E '/\([0-9]+\) *--+ /d; /^=+$/d' "$ronn" | \
pandoc --standalone --from markdown --to man \
--shift-heading-level-by=-1 \
--metadata title="$name" \
--metadata section="$section" \
-o "$name.$section"
done
mv *.1 ../../man/man1/
mv *.5 ../../man/man5/
mv *.8 ../../man/man8/
@@ -13,6 +13,7 @@
import confluent.client as cl
import socket
import struct
import sys
c = cl.Command()
macs = []
interface = sys.argv[1]
-1
View File
@@ -1 +0,0 @@
1.0.1

Some files were not shown because too many files have changed in this diff Show More