2
0
mirror of https://github.com/xcat2/confluent.git synced 2026-10-02 08:51:48 +00:00
Commit Graph

176 Commits

Author SHA1 Message Date
Jarrod Johnson dbe194d6fb Supersede generic behavior for SMM3
SMM3 puts everything on chassis, expose that.
2026-09-25 17:03:14 -04:00
Jarrod Johnson ae1afcabf8 Replace generic unsupported error with more specific
The generic error was unclear for a fairly common attempt to check power state.
2026-09-25 16:50:49 -04:00
Jarrod Johnson 2759404f4c Merge remote-tracking branch 'xcat/master' 2026-09-25 16:14:34 -04:00
Jarrod Johnson a31d1a9efa Stub out useless 'system' for SMMv3
SMMv3 is really a BMC without a 'system', reflect that by stubbing out the system information.
2026-09-25 16:13:40 -04:00
Markus Hilger 1bf8c01867 Report node health and unreachable nodes from EUREKA health
get_health looked only at each node's State, so an Enabled node whose
Health was Warning or Critical left the chassis reported as ok. It also
skipped a node whose resource could not be read, counting it as
healthy while an unreadable chassis was flagged. Report the node's
Health when it is not OK, and a node that cannot be read as
Unreachable.
2026-09-25 19:22:38 +02:00
Markus Hilger 38e94c73d6 Read EUREKA sensors with $expand when the firmware supports it
The sensor collection has 668 members, and reading each one on its own
takes about 92 seconds and rebuilds on every nodesensors call. Opt in
to $expand=. for that collection, checking once whether the firmware
really inlines the members, so firmware that ignores $expand keeps the
per-member reads.
2026-09-25 19:02:39 +02:00
Markus Hilger 859a8db873 Reseat EUREKA nodes with the Reseat reset type
ForceRestart only restarts the host, so the node BMC kept running and
a reseat did not recover a hung one. The EUREKA firmware provides a
Reseat reset type that removes all power from the slot.
2026-09-25 19:02:39 +02:00
Markus Hilger 7cc8972170 Skip CPU temperatures from unbooted EUREKA node BMCs
A node whose BMC is not reporting returns Reading 0 from its
temperature sensors, dragging averages down with meaningless
values. The ComputerSystem Oem data flags this via HasBMCMetrics;
skip such nodes, and keep collecting when the flag is absent.
2026-09-25 19:02:39 +02:00
Markus Hilger e59230c63a Pass rootinfo through to EUREKA OEM handler
The service root was already fetched by the caller; dropping it
forced OEMHandler.create to request /redfish/v1/ again.
2026-09-25 19:02:39 +02:00
Markus Hilger 4d952892b1 Report detail and severity from EUREKA health
get_health collected an issues list but returned it nowhere, and
flattened every problem to Warning. Return the findings as
badreadings using SensorReading, honor verbose, and map chassis
health through _healthmap so Critical is no longer downgraded.
2026-09-25 19:02:39 +02:00
Markus Hilger 5399ba25a0 Fix EUREKA CPU temperature readings
The sensor URL was built as BMC{N}CpuCPU{X}Temp instead of
BMC{N}CPU{X}Temp, so every request returned 404. The entries also
referenced const.SensorUnits, which does not exist, and used a
dict shape the only consumer, get_average_processor_temperature,
cannot read - it expects thermal-style dicts with ReadingCelsius.
2026-09-25 19:02:39 +02:00
Markus Hilger 40b1bf5e3b Report a missing redfish manager instead of raising TypeError
ManagedBy is optional, and when a system does not link a manager
get_bmcurl() answers None. Most callers passed that straight into a
web request, which failed with "Constructor parameter should be str"
from the url library. Route those callers through bmcinfo(), which now
raises UnsupportedFunctionality so confluent reports it plainly. The
event log falls through to its system and chassis fallback, and
list_media lists nothing, since neither needs a manager.
2026-09-25 17:55:05 +02:00
Markus Hilger 5050580bb8 Stop redfish sensor health from raising on a missing Health
SensorReading computed health and states with .get() and then
overwrote both with a strict lookup, so a sensor whose Status lacks
Health, or carries a value outside the health map, raised KeyError
and aborted the whole sensor listing. Keep the tolerant lookup, and
drop the states of a reading that is OK, as the copy in command.py
already does.
2026-09-24 21:47:41 +02:00
Jarrod Johnson 57f6c18ebd Fix early handling in Lenovo OEM handling
If mgrinfo is needed, have logic local, to avoid
the mess during early initialization.
2026-09-21 10:16:48 -04:00
Jarrod Johnson 3dad193926 Fix erroneous cache retention
Do not refresh cache vintage an every access.

Also, give callers finer grained control over cache.
2026-09-18 13:58:05 -04:00
Jarrod Johnson 046ace6dba Improve redfish sensor reading performance
Move to the oem handler and leverage the expansion facility
to speed up supported redfish BMCs.
2026-09-18 13:35:15 -04:00
Jarrod Johnson e6aefd39b8 Remove pyrefly directive
Evidently it goes from needing this to needing it not to be ther.
2026-09-11 14:38:27 -04:00
Jarrod Johnson c7d2e17545 Avoid failure on receiving malformed SOL payload packets
Some devices can emit invalid IPMI packets with missing payloads.  Do not be overly burdened by such packets.
2026-09-10 15:09:32 -04:00
Markus Hilger 497a7abffb Do not cache an SDR that failed to build
init_sdr assigned self._sdr before initialize() ran, so a failure left the
half built object in the cache. The next call saw a non-None _sdr and handed
back that partial repository rather than trying again.

The visible symptom is a first call raising and the second appearing to
succeed. The real cost is on a bmc where the read fails once: the client
keeps the incomplete sdr for the life of the session and every later sensor
lookup answers from it without complaint.
2026-09-03 02:15:22 +02:00
Jarrod Johnson 7afbbca691 Merge pull request #291 from Obihoernchen/fix/multi-nic-refusal
Say which interfaces a multi-homed BMC has, rather than crash
2026-09-02 10:29:29 -04:00
Jarrod Johnson d4c5fd850a Merge pull request #290 from Obihoernchen/fix/report-failures-not-crashes
Let a failed request report itself
2026-09-02 10:16:16 -04:00
Jarrod Johnson 26c69e6cc0 Merge pull request #289 from Obihoernchen/fix/health-optional-collections
Handle Redfish services that omit optional collections
2026-09-02 10:13:26 -04:00
Markus Hilger 9012888cc0 Report a refused XCC web login instead of returning None
get_webclient falls off its end when /api/login answers anything but 200, so
it returned None. wc() passes that back, and thirty of the thirty-four call
sites use it unchecked, so a refused login arrived as "'NoneType' object has
no attribute 'grab_json_response'" from wherever it landed.

Raised where the failure is known, and with the status: 404 is firmware with
no web api, or a Redfish-only capture of one, while 401 is credentials it
will not take. Neither was distinguishable before.

The other four call sites are the inventory reads, and they always did check.
They answer partially when the web interface is out of reach, which is why
nodeinventory still says something useful. They ask through wc_if_available
now, so that tolerance is stated rather than resting on a None.
2026-09-01 23:52:49 +02:00
Markus Hilger 7a7bd758a4 Let a failed request report itself
Two places where the code that exists to explain a failure fails instead, and
the caller is shown the second failure rather than the first.

_do_web_request builds its message from the response body when that body is
not the JSON error document the spec asks for. An html 404 page is exactly
that, the body is bytes, and str + bytes raised TypeError. The status and the
body never reached anyone.

LenovoFirmwareConfig raised a bare Exception when python-lxml and
python-eficompressor are absent. Confluent has no handler for one, so a
missing dependency showed as "Unexpected Error" and hid a message that
already said what to install.

Neither changes what fails, only what the caller is told.
2026-09-01 23:52:49 +02:00
Markus Hilger 1d8dcae6b0 Refuse plainly when a BMC lists no interfaces at all
Same function, one line above. EthernetInterfaces is optional, and when it is
absent the None went into a request and raised TypeError from the url library
rather than saying what was missing.

Kept beside the ambiguous case because the two are one question asked twice.
2026-09-01 23:52:25 +02:00
Markus Hilger 6909d1b5df Say which interfaces a multi-homed BMC has, rather than crash
Dedicated plus shared is how BMCs are built, so several NICs is ordinary.
_get_bmc_nic_url only reaches the count when the address the session came in
on matches none of them, which is what a tunnel or a NAT does, and it then
raised the bare PyghmiException base class.

Confluent has no handler for the base class, so it fell through to the
generic one and showed "Unexpected Error" plus a traceback, for a machine
doing nothing wrong. UnsupportedFunctionality now, which the redfish plugin
already reports plainly and which stays inside PyghmiException.

The message said "does not have exactly one interface" without saying how
many, which ones, or what to do. Every caller takes a name, so it lists the
candidates, and the empty case reads differently from the ambiguous one.
2026-09-01 23:52:25 +02:00
Markus Hilger 26151c506f Read power and boot from a service that has no system
A Redfish service may publish no Systems collection, and power and cooling
equipment does exactly that. sysurl is then None, and get_power and
get_bootdev handed it straight to a request, raising TypeError from inside
the url library.

sysinfo already guarded the same field. Both now ask through _system_url and
get a refusal naming what is missing. DMTF publish three services of this
shape, which is why they sit commented out in inventory-dmtf.yaml.
2026-09-01 23:51:32 +02:00
Markus Hilger 5d5ad821f4 Read health from a service that publishes no PCIe
PCIeDevices and PCIeFunctions are optional, and get_health indexed both
without checking. _get_adp_urls in the same file already spells it
.get('PCIeDevices', []), so these two were the outliers.

Power and cooling equipment has no PCIe at all. The KeyError escaped the
health read and reached the user as "Unexpected Error" with a traceback
behind it. Reproduces offline against DMTF's public-rackmount1 mockup.
2026-09-01 23:51:32 +02:00
Markus Hilger 2e7ce63b51 Read a fractional second as a fraction
parse_time read the digits after the decimal point as whole milliseconds, so
'.5' became 5ms instead of 500ms. Only a three digit fraction came out right,
and a BMC may write either.

'.5', '.50' and '.500' are all 500ms now, '.125' is 125ms. Every other format
parse_time accepts is untouched.
2026-09-01 23:51:00 +02:00
Markus Hilger ef9afa7d72 Fix advanced settings
Disabled this by accident when doing the async fixes.
Upstream pyghmi passess this as well.
2026-08-20 00:41:34 +02:00
Jarrod Johnson 2bd931658c Defer creating locks until loop is running
In python 3.9, this pattern was causing issues.
2026-08-19 13:13:20 -04:00
Jarrod Johnson a0d99a15ce Fix use of positional arguments in getadrrinfo
The new getaddrinfo needs this as a keyvalue pair.
2026-08-19 11:05:41 -04:00
Jarrod Johnson ec64a6d46c Change to asyncio based name lookup in various places 2026-08-19 10:28:55 -04:00
Jarrod Johnson 31647a52e1 Remove disused RLock and simplify code. Fix name lookup stalls on connection. 2026-08-19 09:56:28 -04:00
Markus Hilger 9cdcbe4046 Look elsewhere when the manager publishes no log services
The early return sat before the fallback, so the one layout it was
written for was the one it could not reach.
2026-08-19 15:10:55 +02:00
Markus Hilger caa857ead0 Gather every event log a platform keeps, wherever it keeps it
A read only ever looked at the manager's log services, and fell back to the
system's when the manager published none.  A platform that keeps an event log
in both places had the second one invisible: an AMI MegaRAC keeps power unit
and thermal events in a chassis log that nothing read, 103 records that no
command could reach.

Clearing deliberately does not follow.  It stays where it was, so a log that
only a read reaches is never destroyed by one, and clearing a platform that
keeps its only event log on the system still works.

The name test now ignores spacing, since a build that calls its post code log
"BIOS POST Code Log" was read as an event log and merged 2719 post codes in.
2026-08-19 15:10:55 +02:00
Jarrod Johnson 3fa1236d34 Merge pull request #278 from Obihoernchen/fix/configurable-endpoints
configurable sockapi path and ipmi port
2026-08-19 08:33:06 -04:00
Jarrod Johnson c87bdc5a40 Merge pull request #274 from Obihoernchen/openbmc-support
Various fixes and features for AMI MegaRAC and OpenBMC BMCs
2026-08-19 08:22:26 -04:00
Markus Hilger 9e811dd81c Leave it to OpenBMC to say which of its logs are not events
Reading an event log skipped every log service whose id or name said journal,
dump, post code, host logger or crash, on every implementation.  Those names
are bmcweb's: the AMI and Lenovo bmcs call theirs SEL, EventLog, AuditLog and
PlatformLog.  bmcweb does need the distinction, keeping its event log under the
system while a clear would destroy its dumps, so it gets a handler that names
the words and generic names none, reading whatever a platform publishes.
2026-08-18 22:58:48 +02:00
Markus Hilger 811d48ed42 Ask only MegaRAC for the parameters part it insists on
Generic added an UpdateParameters part to every multipart firmware push,
because the specification has one carried.  Only the AMI firmware was seen to
insist on it, so name it in that handler and let generic send what the caller
passed.
2026-08-18 22:58:48 +02:00
Markus Hilger eaa1a8bfc7 Let a console work through a forwarded ipmi port
A bmc behind a forward answers Activate Payload with the port it listens on
itself, and the advertised-port check refused that, so a console failed where
command traffic worked. The advertised port is never sent to, so the check now
applies only on the default port. A bmc on another port advertising a third one
is no longer refused outright, which nothing here could have served anyway.
2026-08-16 21:37:50 +02:00
Markus Hilger c5f831e589 Start the bmc reset grace when the bmc actually goes
The deadline was set when monitoring began, so a flash that kept
answering for longer than the grace period had already spent it by the
time the reboot it covers arrived, and reported a working update as a
failure.
2026-08-15 13:46:18 +02:00
Markus Hilger 3ace6af075 Reserve the sensor records of the lun they are read from
A device holds its sensor records per lun and scopes the reservation the
same way, so a token taken on lun 0 can be refused for lun 1 with 0xc5.
The retry then took the same wrong token again without end.
2026-08-15 13:45:49 +02:00
Markus Hilger 274b1b310d Record why identify writes the system and not the chassis
An SD665-N V3 refuses every IndicatorLED value on its chassis and takes
all of them on the system, so making the write follow the read breaks it.
2026-08-15 13:44:49 +02:00
Markus Hilger c9804368f2 Only read a user slot as empty when the bmc says it is
Any completion code counted as an absent slot, so a bmc that was busy
or still starting up quietly shortened the user list.
2026-08-15 13:44:49 +02:00
Markus Hilger 803c4f2ceb Skip a shared enclosure when reading the leds
get_identify already ignores a chassis several systems share, so that
one node does not report the enclosure's indicator as its own.
2026-08-15 13:44:49 +02:00
Markus Hilger be1304e560 Name the firmware categories beyond core, adapters and disks
Firmware for a supply or a fan matched no fragment and so was called
core, and nothing could ever answer for misc.

Only the collections that mean one thing are matched by url.  A Storage
resource is the subsystem, so its firmware is the controller rather than
a drive, and a Processor is a cpu as readily as an accelerator, so that
one is decided by asking the processor what it is.
2026-08-15 13:44:37 +02:00
Markus Hilger 8c884e2ece Keep unrelated firmware in core rather than nowhere
An entry naming no RelatedItem was dropped from every category as soon
as any other entry named one, so core lost the bmc and uefi versions.
2026-08-15 12:58:47 +02:00
Markus Hilger e69f82a6a2 Do not let a failed lookup pass for a failed delete
The check for whether the account went sat outside the try, so an error
reading it escaped instead of falling back to blanking the account.
2026-08-15 12:58:32 +02:00
Markus Hilger 24e8cd7e00 Check for a deleted account without the cache
The delete that failed left the account collection cached as it was, so
asking whether the account is gone could only ever answer no.
2026-08-15 12:58:16 +02:00