mirror of
https://github.com/xcat2/xcat-core.git
synced 2026-09-05 04:27:55 +00:00
96c81d7646
The respawn of the install monitor was paced by a retry budget that, once spent, made the daemon stop trying for good. That put xcatiport back in the state the respawn was added to fix: with no monitor alive there is nothing left to reset the counter, so the port stays dead until the whole daemon is restarted, and a port that frees up a minute later is never picked back up. It only reached that state more slowly than before. Pacing itself is needed. do_installm_service dies when it cannot bind the port, so an unguarded re-fork spins as fast as fork allows while something else holds it, and keeps re-entering that function's USR2 socket-takeover handshake. Replace the budget with an exponential backoff that has a ceiling but no end: the delay doubles from XCATD_MON_RESPAWN_MIN_INTERVAL (default 5s) to XCATD_MON_RESPAWN_MAX_INTERVAL (default 300s) and stays there. A monitor that cannot start therefore costs one fork per five minutes for as long as that lasts, and is back within five minutes of the port becoming free, with no restart and no operator action. A monitor that ran for XCATD_MON_RESPAWN_HEALTHY seconds (default 60) plainly got the socket and served, so its eventual death resets the delay: an isolated death is retried at once and the backoff only builds up during a real streak of failures to start. The ceiling is reported once per streak rather than on every attempt, and says that xcatd is still retrying instead of that it has stopped. The pacing lives in a marked mon-respawn-policy region, free of forking and of daemon state, so xCAT-test/unit/xcatd_monitor_respawn.t drives the real code rather than a copy of it. Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>