mirror of
https://github.com/xcat2/xcat-core.git
synced 2026-09-05 12:37:54 +00:00
5970c4efd1
The respawn added for the install monitor is paced by a retry budget, and once that budget is spent the daemon stops trying. That reintroduces the failure the respawn exists to remove: with no monitor alive there is nothing left to reset the counter, so xcatiport stays dead until the whole daemon is restarted, and a port that becomes free a minute later is never picked back up. Pacing the retries is necessary -- an unguarded re-fork spins as fast as fork allows while the port is held, and keeps re-entering do_installm_service's USR2 socket-takeover handshake -- but pacing must not decay into giving up. The property that matters is therefore behavioural, not structural: the monitor comes back on its own, at a bounded rate, no matter how long it has been failing. Assert it by driving xcatd's real pacing code rather than grepping for it -- extract the marked mon-respawn-policy region from the script verbatim, the way build_ubunturepo_lock.t drives build-ubunturepo's real lock, and run it. Over a virtual clock, check that the delay backs off to a ceiling and holds there, that the daemon is still forking monitors three hours into a failure, and that a monitor which stayed up long enough to serve resets the pacing when it later dies. Then do it for real against a genuinely held TCP port: fail several times, release the port, and require that a respawned monitor binds it and stays up without the daemon being restarted. Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>