public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
* [pve-devel] [PATCH ha-manager v2 0/5] watchdog: sync log to disk before and after expiring
@ 2025-06-25 13:23 Maximiliano Sandoval
  2025-06-25 13:23 ` [pve-devel] [PATCH ha-manager v2 1/5] watchdog-mux: Use #define for 60s timeout Maximiliano Sandoval
                   ` (4 more replies)
  0 siblings, 5 replies; 6+ messages in thread
From: Maximiliano Sandoval @ 2025-06-25 13:23 UTC (permalink / raw)
  To: pve-devel

Without a clear-cut message in the log, it is very hard to provide a definitive
answer to whether a host fenced or not. In some cases the journal on the disk
can be missing up to 2 minutes since its last logged entry and the time where
another node detects the corosync link is down, with such a gap, the fenced node
would not even record that it lost conenction and it is not possible to
fully-determine if the node was fenced or not.

This series:
 - adds a second warning 10 seconds before the watchdog expires
 - syncs the journal to disk after the warning was issued
 - syncs the journal to disk after the watchdog expires

Differences from v1:
 - Define the warning cuttoff based on the 60 second timeout
 - Change log messages and constant names
 - When not immediately fencing, run journal sync in double fork

Maximiliano Sandoval (5):
  watchdog-mux: Use #define for 60s timeout
  watchdog-mux: split if block in two if blocks
  watchdog-mux: warn when about to expire
  watchdog-mux: sync journal after logging expiration message
  watchdog-mux: sync journal right after fencing warning

 src/watchdog-mux.c | 52 +++++++++++++++++++++++++++++++++++++++++-----
 1 file changed, 47 insertions(+), 5 deletions(-)

-- 
2.39.5



_______________________________________________
pve-devel mailing list
pve-devel@lists.proxmox.com
https://lists.proxmox.com/cgi-bin/mailman/listinfo/pve-devel


^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2025-06-25 13:24 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2025-06-25 13:23 [pve-devel] [PATCH ha-manager v2 0/5] watchdog: sync log to disk before and after expiring Maximiliano Sandoval
2025-06-25 13:23 ` [pve-devel] [PATCH ha-manager v2 1/5] watchdog-mux: Use #define for 60s timeout Maximiliano Sandoval
2025-06-25 13:23 ` [pve-devel] [PATCH ha-manager v2 2/5] watchdog-mux: split if block in two if blocks Maximiliano Sandoval
2025-06-25 13:23 ` [pve-devel] [PATCH ha-manager v2 3/5] watchdog-mux: warn when about to expire Maximiliano Sandoval
2025-06-25 13:23 ` [pve-devel] [PATCH ha-manager v2 4/5] watchdog-mux: sync journal after logging expiration message Maximiliano Sandoval
2025-06-25 13:23 ` [pve-devel] [PATCH ha-manager v2 5/5] watchdog-mux: sync journal right after fencing warning Maximiliano Sandoval

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal