From: "Elias Huhsovitz" <e.huhsovitz@proxmox.com>
To: "Dominik Csapak" <d.csapak@proxmox.com>, <pve-devel@lists.proxmox.com>
Subject: Re: [RFC qemu-server v2 0/4] fix #7282: allow (NUMA aware) vCPU pinning
Date: Fri, 25 Sep 2026 11:31:58 +0200 [thread overview]
Message-ID: <DLOAKASNY1LL.2XWVQZKJCRHH2@proxmox.com> (raw)
In-Reply-To: <faaf878f-e44d-4d38-a947-441b11c7629b@proxmox.com>
Thanks for the feedback!
Sorry for not running make tidy, I always forget... I also had some
kind of auto-formatter that was applied on save via my editor.
See further comments/questions below.
On Thu Sep 24, 2026 at 2:25 PM CEST, Dominik Csapak wrote:
> I'm in the progress of looking through the code, but want
> to give some high level feedback 'early'
>
> I think the approach goes to far for what was intended:
>
> The pinning setting now also influences memory behavior
> but we already have the numaX options for that, the intention
> was that this is just related to the virtual cores.
> I'd prefer for it to stay that way.
This is because i noticed major performance inconsistencies when I
didn't explicity define the numaX in this shape to bind the memory:
numaX: cpus=0-1,hostnodes=0,memory=2048,policy=bind
> If we want to modify the memory handling, we'll have to do
> that in the memory/numa/etc. options (or a new separate
> one) but don't conflate it with the CPU pinning.
I will include this in a later patch. My reasoning was that I didn't
want to shift the responsibility to the user, but I guess that should
be handled in another layer of abstraction.
> Also, AFAIU this can also change the guest hardware
> depending on the host topology no?
> If yes, this is a no-go since this breaks live migration,
> potentially suspend/resume and backup/restore.
To my understanding binding the numa nodes to specific memory
locations would definitely break live migration. (But i am not sure how
to even fix this)
> It also can mess up the guest licensing when this is fixed
> for some sockets for example.
I didn't think about this, good point.
> In general such an options should never overwrite the
> explicit values in the config.
I belive I only added values and didn't override them, but I guess that
is also additional config that the user did not ask for.
> Iff we want such a thing, this must be reflected in the
> sockets/cores/etc config (e.g. by setting 'auto' or
> some thing) and must be live migration blocker.
>
> In a first applied version I'd really like for it to focus
> on pinning vcpus on specific cores (or sets of cores).
>
> This should be opt-in and align with already existing options
> (such as numa/memory/sockets/etc) and not break
> live migration (e.g. it is ok to pin to different
> cpus/hostnodes on another node, but the hardware must
> stay the same)
>
> Making memory hostnode selection automatic and inferring guest
> topology can also make sense, but should be a separate
> patch/series and isn't really what the bug was about.
>
> Let's keep the focus narrow, this also will make the patches
> smaller and more easily digestible ;)
>
> I'll answer on the individual patches for style/code things too
I prose the following
order of patches:
-----------------
1. Read/Store CPU Topology. i.e. have some kind of internal
datastrucutre that can be queried for CPU Topology information.
2. Introduce pinning=one-to-one|numa settings. (SMP aware via CPU
topology).
3. Automatic host node selection and memory binding in NUMA mode.
(Opt-in)
4. I/O Thread control, as described in [0]. The format could mirror the
affinity setting, e.g. a user could configure: iothreadpin: 0-5.
5. CPU Isolation: using a combination of
- isolcpus
- cpuset
- qemu.slice
(not sure what is preffered at the moment. Some users seem to
swear by simply setting `isolcpus`.)
6. Expose configurations via UI.
Optional Patches:
-----------------
a. Make CPU Topology via API (Perhaps useful for future UI features, or
making scripting easier).
b. Introduce pinning=balanced mode.
c. pre-Huge Page memory compaction, as mentioned in [0].
Does this sound ok, or do you have a better sequence in mind?
Additional questions:
---------------------
Q1: Is it not ok to introduce new perl modules? I didn't really find any
existing perl modules where was happy to put the code into from a
semantic standpoint.
Q2: What is your opinion on the pinning=balanced mode (better name TBD)?
Q3: Since we are planning a to move the PCI code to rust: Should the CPU
Topology code ideally also live in rust?
References
----------
[0] https://forum.proxmox.com/threads/hey-proxmox-community-lets-talk-about-resources-isolation.124256/
prev parent reply other threads:[~2026-09-25 9:32 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-21 9:53 [RFC qemu-server v2 0/4] fix #7282: allow (NUMA aware) vCPU pinning Elias Huhsovitz
2026-09-21 9:54 ` [RFC qemu-server v2 1/4] pinning: add topology discovery and config parsing Elias Huhsovitz
2026-09-24 12:35 ` Dominik Csapak
2026-09-21 9:54 ` [RFC qemu-server v2 2/4] pinning: add NUMA allocator and reservation tracking Elias Huhsovitz
2026-09-21 9:54 ` [RFC qemu-server v2 3/4] memory: integrate pinning-aware NUMA memory binding Elias Huhsovitz
2026-09-21 9:54 ` [RFC qemu-server v2 4/4] pinning: integrate cpu pinning into vm lifecycle Elias Huhsovitz
2026-09-24 12:25 ` [RFC qemu-server v2 0/4] fix #7282: allow (NUMA aware) vCPU pinning Dominik Csapak
2026-09-25 9:31 ` Elias Huhsovitz [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DLOAKASNY1LL.2XWVQZKJCRHH2@proxmox.com \
--to=e.huhsovitz@proxmox.com \
--cc=d.csapak@proxmox.com \
--cc=pve-devel@lists.proxmox.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox