From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [45.144.208.40]) by lore.proxmox.com (Postfix) with ESMTPS id ABF911FF0B2 for ; Fri, 09 Oct 2026 11:44:30 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 2731C2145C; Fri, 09 Oct 2026 11:44:28 +0200 (CEST) Mime-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset=UTF-8 Date: Fri, 09 Oct 2026 11:44:21 +0200 Message-Id: Subject: Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes From: "Elias Huhsovitz" To: "Filip Schauer" , X-Mailer: aerc 0.20.0 References: <20261007083726.26067-1-e.huhsovitz@proxmox.com> <82fceb8e-f52d-4e56-9dcf-5cd61393fa26@proxmox.com> In-Reply-To: <82fceb8e-f52d-4e56-9dcf-5cd61393fa26@proxmox.com> X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1791539062071 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.525 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: 4DDDTWLJN6SM56TZYEHLUJ7O7UWPXYGW X-Message-ID-Hash: 4DDDTWLJN6SM56TZYEHLUJ7O7UWPXYGW X-MailFrom: e.huhsovitz@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On Thu Oct 8, 2026 at 3:15 PM CEST, Filip Schauer wrote: > On 08/10/2026 12:25, Elias Huhsovitz wrote: >> On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote: >>> On 07/10/2026 10:37, Elias Huhsovitz wrote: >>>> The current Problem >>>> ------------------- >>>> Console session processes can outlive their container when it stops or >>>> is destroyed. The cleanup currently runs in pvestatd every 10 seconds. >>>> It calls vmstatus() and kills stale processes by PID. >>>> >>>> The killing of the process can result in a race condition: >>>> If the console gets closed by something else, just before the kill >>>> command is executed, a different process can reuse the PID and we kill >>>> a completely unrelated process. >>>> >>>> Also checking this every 10s is too much unnecessary overhead IMO. >>>> >>>> My understanding of a "better" architecture >>>> ------------------------------------------- >>>> IMO console deletion should be tied to the state of the underlying >>>> container. Also I was unable to find a "convenient" time/event for a >>>> cleanup job to run. >>>> >>>> Proposed implementation >>>> ----------------------- >>>> Wrap console commands in transient systemd scopes and bind them to the >>>> container service via the PartOf=3D property. When the container stops= , >>>> systemd stops the scope, terminating the console session and the dtach >>>> session. This allows removing the pvestatd periodic cleanup entirely. >>> >>> Getting away from the polling in pvestatd would be nice. >>> However, I don't think `PartOf=3D` can be relied on here. >>> >>> As a test I created a scope running `sleep 1000` with >>> `PartOf=3Dpve-container@103.service`. Stopping the container left the >>> scope and its process running. It seems like `lxc-stop` ends the servic= e >>> without a systemd stop job, and `PartOf=3D` only propagates explicit st= op >>> jobs. The same goes for a `poweroff` inside the container. As far as I >>> can tell, the scope would only be stopped by `systemctl stop` or >>> `systemctl restart`. So most stop paths are not covered. >>=20 >> Thats interesting. I was not able to re-produce this behaviour when >> calling poweroff from inside the container >>=20 >> Did you create the console using `pct`, or via the API >> (e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`. > > 1. I started container 109. > 2. Then I ran `systemd-run --scope --unit=3Dpve-lxc-console-109 --propert= y "PartOf=3Dpve-container@109.service" sleep 1000` > 3. I watched the sleep process with `watch -n 0.1 'COLUMNS=3D ps aux | gr= ep sleep'` > 4. Then I logged into the container via xterm.js in the web UI and ran > the `poweroff` command. > 5. The container stopped, but the sleep process kept running. > 6. Only once explicitly calling > `systemctl stop pve-container@109.service` > did the sleep process terminate. > > >>=20 >> But nevertheless, you are correct. My assumptions about the `PartOf=3D` >> were wrong. >>=20 >>> >>> But do we even need to kill the lxc-console process in the first place? >>> lxc-console already exits automatically when the container stops. Or am >>> I missing something? Is there any case where the process lingers? >>=20 >> There are a few scenarios where lxc-console remains: >>=20 >> 1. rm-rf .conf >>=20 >> while the process is running. Killing the container via lxc-stop >> leaves the console running. > > This one I could not reproduce. > 1. I started a container. > 2. I opened its console in the web UI. > 3. I deleted the container config at /etc/pve/lxc/.conf. > 4. I stopped the container with `lxc-stop `. > (also tried with `lxc-stop --kill `) > 5. The lxc-console process was no longer running after this. I did the following: 1. Remove the subroutine `remove_stale_lxc_consoles` from pvestatd.pm 2. systemctl restart pvestatd 3. pct start 202 4. open the xterm web console for container 202 pgrep -af 'console|dtach|spiceterm|vncterm|termproxy': 15088 /usr/bin/termproxy 5900 --path /vms/202 --perm VM.Console --vncticket-endpoint --verify-port --ticket-fd 6 -- /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1 15093 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-cons= ole -n 202 -e -1 15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-cons= ole -n 202 -e -1 15095 lxc-console -n 202 -e -1 5. rm -rf /etc/pve/lxc/202.conf 6. I close the xterm web console pgrep -af 'console|dtach|spiceterm|vncterm|termproxy': 15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-cons= ole -n 202 -e -1 15095 lxc-console -n 202 -e -1 7. lxc-stop 202 lxc-console process is gone (I think i need to better standardzie my testing process :/) > >>=20 >> 2. pid=3D$(pgrep -f "lxc-console.*") >> kill -STOP $pid >> lxc-stop >>=20 >> also results in a remaining console. > > This one I was indeed able to reproduce. > > >>=20 >> In my testing yesterday, calling >>=20 >> pct detory --force >> while the container is still running, resulted in an orphaned console, >> which prompted me to create this patch in the first place >>=20 >> But I am unable to re-produce this currently. So probably the issue was >> something else... > > I am not able to reproduce this one either. When I run > `pct destroy --force` it first stops the container: > ``` > forced to stop CT before destroying! > ``` Yeah, I think I made some mistake in my original testing here. Sorry about that. > >>=20 >> Bottom Line >> ----------- >> The scenarios that leave an orphaned are not standard lifecycle events. >>=20 >> What about: >>=20 >> Moving the `remove_stale_lxc_consoles` logic (with some >> improvements) to pve-container. >>=20 >> Introduce a new pct command `pct console cleanup` (name TBD), so that >> administrators can cleanup the orphaned consoles in the edge cases >> where an lxc-console might get left behind. >>=20 >> I would love to hear your thoughts on this matter! > > Hmm... I am not very keen on adding a new command just for cleaning up > what, as far as my testing goes, looks like an artificial situation. > > If `lxc-console` processes don't linger under realistic conditions, we > should evaluate whether we even need the cleanup, or if > `remove_stale_lxc_consoles` was simply leftover legacy code. True, I am starting to belive now as well that this might just be leftover legacy code. The current `remove_stale_lxc_consoles` in pvestatd.pm, only removes lxc-consoles if the underlying container is not returned by=20 `PVE::LXC::vmstatus()`, i.e., if the container does not exist anymore. To my knowledge this can only happen if the container config is broken or missing.=20 I just figured that the existance of this code, might indicate some important edge case, but perhaps we just should consider just removing it. > And even if we want to handle this, I think it should remain automatic. > > Maybe we can bridge the gap to something I tried here: > "add container console scrollback buffer" > https://lore.proxmox.com/pve-devel/20260121112335.84473-1-f.schauer@proxm= ox.com/ > As it is right now, my v1 is not ready, but my point is that we could > maybe have `lxc-console` + `dtach` running automatically at all times > alongside `lxc-start`, by also having `lxc-console` managed by > `pve-container@.service`. This way, we could manage both processes under > the same systemd unit. Not sure if that's the direction we want to take, > but it's an idea. I think this is a great idea. Since we start the `pve-container@.service` for each container anyway, I belive utilizing it to the related processes/resources is correct direction. My understanding ---------------- Maybe my idea is a bit too basic, but what about this: Currently the pve-container@.service contains ExecStop=3D/usr/share/lxc/pve-container-stop-wrapper %i To my knowledge this is only executed in a non-failure case. We could extend the .service file to contain ExecStopPost=3D implement a new `pve-container-stop-post-wrapper` Add it to the .service config: ExecStopPost=3D/usr/share/lxc/pve-container-stop-post-wrapper %i according to [1], we can use this field when "the service exited unexpectedly". Inside the new wrapper we call something like: lxc-stop # other cleanup ? Your patch ---------- Let me know some of the details on how you plan for for pve-container@.service to manage the lxc-console and the improved dtach process. I think this idea is really interesting and I would like to support/review your patches on this matter. References ---------- [1] https://www.freedesktop.org/software/systemd/man/latest/systemd.service= .html