From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id 1D5E91FF0AF for ; Thu, 08 Oct 2026 12:25:57 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 071732137D; Thu, 08 Oct 2026 12:25:55 +0200 (CEST) Mime-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset=UTF-8 Date: Thu, 08 Oct 2026 12:25:49 +0200 Message-Id: Subject: Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes From: "Elias Huhsovitz" To: "Filip Schauer" , X-Mailer: aerc 0.20.0 References: <20261007083726.26067-1-e.huhsovitz@proxmox.com> In-Reply-To: X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1791455150059 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.532 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: TDSCOCCOO3P273EYRJ6JE2BLKC3AXREX X-Message-ID-Hash: TDSCOCCOO3P273EYRJ6JE2BLKC3AXREX X-MailFrom: e.huhsovitz@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote: > On 07/10/2026 10:37, Elias Huhsovitz wrote: >> The current Problem >> ------------------- >> Console session processes can outlive their container when it stops or >> is destroyed. The cleanup currently runs in pvestatd every 10 seconds. >> It calls vmstatus() and kills stale processes by PID. >>=20 >> The killing of the process can result in a race condition: >> If the console gets closed by something else, just before the kill >> command is executed, a different process can reuse the PID and we kill >> a completely unrelated process. >>=20 >> Also checking this every 10s is too much unnecessary overhead IMO. >>=20 >> My understanding of a "better" architecture >> ------------------------------------------- >> IMO console deletion should be tied to the state of the underlying >> container. Also I was unable to find a "convenient" time/event for a >> cleanup job to run. >>=20 >> Proposed implementation >> ----------------------- >> Wrap console commands in transient systemd scopes and bind them to the >> container service via the PartOf=3D property. When the container stops, >> systemd stops the scope, terminating the console session and the dtach >> session. This allows removing the pvestatd periodic cleanup entirely. > > Getting away from the polling in pvestatd would be nice. > However, I don't think `PartOf=3D` can be relied on here. > > As a test I created a scope running `sleep 1000` with > `PartOf=3Dpve-container@103.service`. Stopping the container left the > scope and its process running. It seems like `lxc-stop` ends the service > without a systemd stop job, and `PartOf=3D` only propagates explicit stop > jobs. The same goes for a `poweroff` inside the container. As far as I > can tell, the scope would only be stopped by `systemctl stop` or > `systemctl restart`. So most stop paths are not covered. Thats interesting. I was not able to re-produce this behaviour when calling poweroff from inside the container Did you create the console using `pct`, or via the API=20 (e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`. But nevertheless, you are correct. My assumptions about the `PartOf=3D` were wrong. > > But do we even need to kill the lxc-console process in the first place? > lxc-console already exits automatically when the container stops. Or am > I missing something? Is there any case where the process lingers? There are a few scenarios where lxc-console remains: 1. rm-rf .conf=20 while the process is running. Killing the container via lxc-stop leaves the console running. 2. pid=3D$(pgrep -f "lxc-console.*") kill -STOP $pid lxc-stop also results in a remaining console. In my testing yesterday, calling=20 pct detory --force while the container is still running, resulted in an orphaned console, which prompted me to create this patch in the first place But I am unable to re-produce this currently. So probably the issue was something else... Bottom Line ----------- The scenarios that leave an orphaned are not standard lifecycle events. What about: Moving the `remove_stale_lxc_consoles` logic (with some improvements) to pve-container. Introduce a new pct command `pct console cleanup` (name TBD), so that administrators can cleanup the orphaned consoles in the edge cases where an lxc-console might get left behind. I would love to hear your thoughts on this matter!