From: Alwin Antreich <alwin@antreich.com>
To: pve-user@lists.proxmox.com, Marco Gaiarin <gaio@lilliput.linux.it>
Subject: Re: Two nodes asymmetrical cluster, backup, stall of one node.
Date: Tue, 01 Sep 2026 10:55:22 +0000 [thread overview]
Message-ID: <20260901105522.Horde.vgrn9fzQ416tczxJi0dVq0Y@drive.antreich.com> (raw)
In-Reply-To: <snrfmm-hso.ln1@leia.lilliput.linux.it>
Hi Marco,
"Marco Gaiarin" gaio@lilliput.linux.it – 31. August 2026 um 23:10
>
> I've hit a trouble i cannot understand very well; we have some asymmetrical
> clusters, composed of one main server and a 'backup', a little underpowered, server
> that act as backup, exporting their fs via NFS.
>
> At every backup run (weekly), the backup node got some pressure (this is
> expected and common to other clusters), kernel start to complain as:
>
> 2026-08-30T00:18:43.125643+02:00 pdpve2 pvescheduler[1474088]: jobs: cfs-lock 'file-jobs_cfg' error: got lock request timeout
> 2026-08-30T00:18:43.723075+02:00 pdpve2 pvescheduler[1474087]: replication: cfs-lock 'file-replication_cfg' error: got lock request timeout
> [...]
> 2026-08-30T00:49:59.269199+02:00 pdpve2 kernel: [481202.580590] INFO: task pvescheduler:1478618 blocked for more than 122 seconds.
>
> after that, backup node get 'offline' (from main node), pveproxy respond in
> backup node but i canot login (local PAM account or domain ones).
>
> If i SSH on the backup node, all seems normal and working, no service
> stopped, a little test VM running on it working as expected.
>
> Only a note, there what seems a stalled backup job:
>
> root@pdpve2:~# ps aux | grep pves[r]
> root 1497581 0.0 0.6 308788 110928 ? Ds Aug30 0:00 /usr/bin/perl -T /usr/bin/pvesr prepare-local-job 100-0 --scan local-zfs local-zfs:vm-100-disk-0 --last_sync 1788035400
>
> i've trid to kill it, even with '-9', but whitout success.
Processes in D state can't be killed, only a reboot helps.
The pvesr is doing a snapshot sync, which is not involving NFS. Are you running two different things, one replication and a vzdump backup at the same time?
>
>
> If i reboot the node, all come back as normal.
>
>
> Someone have some clue?
>
>
> If i have a set ov VM on node A and a set of VM on node B, and i run a
> backup of all VMs, there's some way to 'serialize' it, eg, do backup of VMs
> of node A and after that backup of VM on node B?
>
> I suppose that part of the trouble came fron the fact that backup start in
> parrallels on both nodes, but node B is also the destination of backup, so
> it get 3 time the IO (write from node A, read from node B, write to node B).
>
> I can surely setup TWO backup job, but... thanks.
That's the solution. Plus add a bandwidth limit to reduce the IO pressure.
Cheers,
Alwin
prev parent reply other threads:[~2026-09-01 10:55 UTC|newest]
Thread overview: 2+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-08-31 8:48 Two nodes asymmetrical cluster, backup, stall of one node Marco Gaiarin
2026-09-01 10:55 ` Alwin Antreich [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260901105522.Horde.vgrn9fzQ416tczxJi0dVq0Y@drive.antreich.com \
--to=alwin@antreich.com \
--cc=gaio@lilliput.linux.it \
--cc=pve-user@lists.proxmox.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox