[ClusterLabs] Can stalled libqb SHM I/O delay Pacemaker's operation timer?

Brian Cosgrove cosgroveb at gmail.com
Mon Aug 10 20:03:20 UTC 2026


Could Pacemaker’s controller block on a stalled block-device read while
faulting in a swapped-
out libqb SHM page, before it receives the executor reply and starts the
operation timer?

Would that behavior be expected?

I'm investigating a lockup where the PostgreSQL primary and Pacemaker DC
were on the same node when a degraded RAID array caused reads to appear
unbounded without returning errors. Pacemaker did not stop the PostgreSQL
resource or promote a standby. crm_mon showed the DC as standby (with
active resources).

Testing that simulates the incident reproduces that production Pacemaker
behavior. In that setup a libqb SHM page is swapped out and controller
blocks in shmem_fault() before the executor reply and operation timer. The
surviving production logs (such as the crm_mon output) match. Corosync
membership remained intact in both the incident and the test.


-- 
Brian Cosgrove
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <https://lists.clusterlabs.org/pipermail/users/attachments/20260810/dc4ea071/attachment.htm>


More information about the Users mailing list