From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id 986111FF0AF for ; Thu, 08 Oct 2026 15:15:38 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id E1126214BF; Thu, 08 Oct 2026 15:15:36 +0200 (CEST) Message-ID: <82fceb8e-f52d-4e56-9dcf-5cd61393fa26@proxmox.com> Date: Thu, 8 Oct 2026 15:15:28 +0200 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes To: Elias Huhsovitz , pve-devel@lists.proxmox.com References: <20261007083726.26067-1-e.huhsovitz@proxmox.com> Content-Language: en-US From: Filip Schauer In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1791465332401 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.525 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: 3O5KBSIMQ5HWUU4XIB74GJ47AJUJTPYN X-Message-ID-Hash: 3O5KBSIMQ5HWUU4XIB74GJ47AJUJTPYN X-MailFrom: f.schauer@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On 08/10/2026 12:25, Elias Huhsovitz wrote: > On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote: >> On 07/10/2026 10:37, Elias Huhsovitz wrote: >>> The current Problem >>> ------------------- >>> Console session processes can outlive their container when it stops or >>> is destroyed. The cleanup currently runs in pvestatd every 10 seconds. >>> It calls vmstatus() and kills stale processes by PID. >>> >>> The killing of the process can result in a race condition: >>> If the console gets closed by something else, just before the kill >>> command is executed, a different process can reuse the PID and we kill >>> a completely unrelated process. >>> >>> Also checking this every 10s is too much unnecessary overhead IMO. >>> >>> My understanding of a "better" architecture >>> ------------------------------------------- >>> IMO console deletion should be tied to the state of the underlying >>> container. Also I was unable to find a "convenient" time/event for a >>> cleanup job to run. >>> >>> Proposed implementation >>> ----------------------- >>> Wrap console commands in transient systemd scopes and bind them to the >>> container service via the PartOf= property. When the container stops, >>> systemd stops the scope, terminating the console session and the dtach >>> session. This allows removing the pvestatd periodic cleanup entirely. >> >> Getting away from the polling in pvestatd would be nice. >> However, I don't think `PartOf=` can be relied on here. >> >> As a test I created a scope running `sleep 1000` with >> `PartOf=pve-container@103.service`. Stopping the container left the >> scope and its process running. It seems like `lxc-stop` ends the service >> without a systemd stop job, and `PartOf=` only propagates explicit stop >> jobs. The same goes for a `poweroff` inside the container. As far as I >> can tell, the scope would only be stopped by `systemctl stop` or >> `systemctl restart`. So most stop paths are not covered. > > Thats interesting. I was not able to re-produce this behaviour when > calling poweroff from inside the container > > Did you create the console using `pct`, or via the API > (e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`. 1. I started container 109. 2. Then I ran `systemd-run --scope --unit=pve-lxc-console-109 --property "PartOf=pve-container@109.service" sleep 1000` 3. I watched the sleep process with `watch -n 0.1 'COLUMNS= ps aux | grep sleep'` 4. Then I logged into the container via xterm.js in the web UI and ran the `poweroff` command. 5. The container stopped, but the sleep process kept running. 6. Only once explicitly calling `systemctl stop pve-container@109.service` did the sleep process terminate. > > But nevertheless, you are correct. My assumptions about the `PartOf=` > were wrong. > >> >> But do we even need to kill the lxc-console process in the first place? >> lxc-console already exits automatically when the container stops. Or am >> I missing something? Is there any case where the process lingers? > > There are a few scenarios where lxc-console remains: > > 1. rm-rf .conf > > while the process is running. Killing the container via lxc-stop > leaves the console running. This one I could not reproduce. 1. I started a container. 2. I opened its console in the web UI. 3. I deleted the container config at /etc/pve/lxc/.conf. 4. I stopped the container with `lxc-stop `. (also tried with `lxc-stop --kill `) 5. The lxc-console process was no longer running after this. > > 2. pid=$(pgrep -f "lxc-console.*") > kill -STOP $pid > lxc-stop > > also results in a remaining console. This one I was indeed able to reproduce. > > In my testing yesterday, calling > > pct detory --force > while the container is still running, resulted in an orphaned console, > which prompted me to create this patch in the first place > > But I am unable to re-produce this currently. So probably the issue was > something else... I am not able to reproduce this one either. When I run `pct destroy --force` it first stops the container: ``` forced to stop CT before destroying! ``` > > Bottom Line > ----------- > The scenarios that leave an orphaned are not standard lifecycle events. > > What about: > > Moving the `remove_stale_lxc_consoles` logic (with some > improvements) to pve-container. > > Introduce a new pct command `pct console cleanup` (name TBD), so that > administrators can cleanup the orphaned consoles in the edge cases > where an lxc-console might get left behind. > > I would love to hear your thoughts on this matter! Hmm... I am not very keen on adding a new command just for cleaning up what, as far as my testing goes, looks like an artificial situation. If `lxc-console` processes don't linger under realistic conditions, we should evaluate whether we even need the cleanup, or if `remove_stale_lxc_consoles` was simply leftover legacy code. And even if we want to handle this, I think it should remain automatic. Maybe we can bridge the gap to something I tried here: "add container console scrollback buffer" https://lore.proxmox.com/pve-devel/20260121112335.84473-1-f.schauer@proxmox.com/ As it is right now, my v1 is not ready, but my point is that we could maybe have `lxc-console` + `dtach` running automatically at all times alongside `lxc-start`, by also having `lxc-console` managed by `pve-container@.service`. This way, we could manage both processes under the same systemd unit. Not sure if that's the direction we want to take, but it's an idea.