mirror of
https://github.com/xcat2/xcat-core.git
synced 2026-09-05 12:37:54 +00:00
75d9a3a6d5
Three defects found by running the respawn against a live xcatd on an MN rather than only against its unit tests. A monitor whose child dies between xfork() returning and the assignment to $pid_MON is lost for good. ssl_reaper matches $CHILDPID against $pid_MON, so a child reaped in that window is compared against a stale value and missed, and $pid_MON is then left naming a pid that no longer exists. The service loop reads !$pid_MON to decide whether to respawn, so it never respawns again -- the same permanently dead xcatiport this whole change exists to prevent, reached by a different route. Block SIGCHLD across the fork and the assignment at both fork sites; the child unblocks on the same line, since it needs to reap its own children. Reproduced with a widened window before the fix and confirmed closed after. Recovery took 30 seconds on an idle daemon. The respawn only gets a turn when the service loop comes round, and the loop parks in $bothwatcher->can_read(30) when there is nothing to serve, so the full select timeout was being added to the respawn delay. Wait in 5s hops while the monitor is down and at the usual 30s otherwise, so an idle daemon pays a few extra wakeups only while xcatiport is actually dead. Measured on the MN afterwards: a killed monitor returns in 5s, then 10s, then 21s across three kills in a row -- the backoff, visible in wall-clock time -- reclaiming the port each time, with the SSL listener holding the same pid throughout. The tunables are read from %ENV and were compared before being validated, so an empty or misspelt XCATD_MON_RESPAWN_* put "Argument isn't numeric" in the daemon log at every start. Anything that is not a plain non-negative integer is now treated as unset. Signed-off-by: Daniel Hilst <392820+dhilst@users.noreply.github.com>