From: "Elias Huhsovitz" <e.huhsovitz@proxmox.com>
To: "Filip Schauer" <f.schauer@proxmox.com>, <pve-devel@lists.proxmox.com>
Subject: Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
Date: Fri, 09 Oct 2026 11:44:21 +0200 [thread overview]
Message-ID: <DM07LF3LD9RU.2VK0W1IC6WVKE@proxmox.com> (raw)
In-Reply-To: <82fceb8e-f52d-4e56-9dcf-5cd61393fa26@proxmox.com>
On Thu Oct 8, 2026 at 3:15 PM CEST, Filip Schauer wrote:
> On 08/10/2026 12:25, Elias Huhsovitz wrote:
>> On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote:
>>> On 07/10/2026 10:37, Elias Huhsovitz wrote:
>>>> The current Problem
>>>> -------------------
>>>> Console session processes can outlive their container when it stops or
>>>> is destroyed. The cleanup currently runs in pvestatd every 10 seconds.
>>>> It calls vmstatus() and kills stale processes by PID.
>>>>
>>>> The killing of the process can result in a race condition:
>>>> If the console gets closed by something else, just before the kill
>>>> command is executed, a different process can reuse the PID and we kill
>>>> a completely unrelated process.
>>>>
>>>> Also checking this every 10s is too much unnecessary overhead IMO.
>>>>
>>>> My understanding of a "better" architecture
>>>> -------------------------------------------
>>>> IMO console deletion should be tied to the state of the underlying
>>>> container. Also I was unable to find a "convenient" time/event for a
>>>> cleanup job to run.
>>>>
>>>> Proposed implementation
>>>> -----------------------
>>>> Wrap console commands in transient systemd scopes and bind them to the
>>>> container service via the PartOf= property. When the container stops,
>>>> systemd stops the scope, terminating the console session and the dtach
>>>> session. This allows removing the pvestatd periodic cleanup entirely.
>>>
>>> Getting away from the polling in pvestatd would be nice.
>>> However, I don't think `PartOf=` can be relied on here.
>>>
>>> As a test I created a scope running `sleep 1000` with
>>> `PartOf=pve-container@103.service`. Stopping the container left the
>>> scope and its process running. It seems like `lxc-stop` ends the service
>>> without a systemd stop job, and `PartOf=` only propagates explicit stop
>>> jobs. The same goes for a `poweroff` inside the container. As far as I
>>> can tell, the scope would only be stopped by `systemctl stop` or
>>> `systemctl restart`. So most stop paths are not covered.
>>
>> Thats interesting. I was not able to re-produce this behaviour when
>> calling poweroff from inside the container
>>
>> Did you create the console using `pct`, or via the API
>> (e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`.
>
> 1. I started container 109.
> 2. Then I ran `systemd-run --scope --unit=pve-lxc-console-109 --property "PartOf=pve-container@109.service" sleep 1000`
> 3. I watched the sleep process with `watch -n 0.1 'COLUMNS= ps aux | grep sleep'`
> 4. Then I logged into the container via xterm.js in the web UI and ran
> the `poweroff` command.
> 5. The container stopped, but the sleep process kept running.
> 6. Only once explicitly calling
> `systemctl stop pve-container@109.service`
> did the sleep process terminate.
>
>
>>
>> But nevertheless, you are correct. My assumptions about the `PartOf=`
>> were wrong.
>>
>>>
>>> But do we even need to kill the lxc-console process in the first place?
>>> lxc-console already exits automatically when the container stops. Or am
>>> I missing something? Is there any case where the process lingers?
>>
>> There are a few scenarios where lxc-console remains:
>>
>> 1. rm-rf <id>.conf
>>
>> while the process is running. Killing the container via lxc-stop <id>
>> leaves the console running.
>
> This one I could not reproduce.
> 1. I started a container.
> 2. I opened its console in the web UI.
> 3. I deleted the container config at /etc/pve/lxc/<id>.conf.
> 4. I stopped the container with `lxc-stop <id>`.
> (also tried with `lxc-stop --kill <id>`)
> 5. The lxc-console process was no longer running after this.
I did the following:
1. Remove the subroutine `remove_stale_lxc_consoles` from pvestatd.pm
2. systemctl restart pvestatd
3. pct start 202
4. open the xterm web console for container 202
pgrep -af 'console|dtach|spiceterm|vncterm|termproxy':
15088 /usr/bin/termproxy 5900 --path /vms/202 --perm VM.Console
--vncticket-endpoint --verify-port --ticket-fd 6 -- /usr/bin/dtach
-A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15093 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15095 lxc-console -n 202 -e -1
5. rm -rf /etc/pve/lxc/202.conf
6. I close the xterm web console
pgrep -af 'console|dtach|spiceterm|vncterm|termproxy':
15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15095 lxc-console -n 202 -e -1
7. lxc-stop 202
lxc-console process is gone (I think i need to better standardzie my
testing process :/)
>
>>
>> 2. pid=$(pgrep -f "lxc-console.*<id>")
>> kill -STOP $pid
>> lxc-stop <id>
>>
>> also results in a remaining console.
>
> This one I was indeed able to reproduce.
>
>
>>
>> In my testing yesterday, calling
>>
>> pct detory <id> --force
>> while the container is still running, resulted in an orphaned console,
>> which prompted me to create this patch in the first place
>>
>> But I am unable to re-produce this currently. So probably the issue was
>> something else...
>
> I am not able to reproduce this one either. When I run
> `pct destroy <id> --force` it first stops the container:
> ```
> forced to stop CT <id> before destroying!
> ```
Yeah, I think I made some mistake in my original testing here.
Sorry about that.
>
>>
>> Bottom Line
>> -----------
>> The scenarios that leave an orphaned are not standard lifecycle events.
>>
>> What about:
>>
>> Moving the `remove_stale_lxc_consoles` logic (with some
>> improvements) to pve-container.
>>
>> Introduce a new pct command `pct console cleanup` (name TBD), so that
>> administrators can cleanup the orphaned consoles in the edge cases
>> where an lxc-console might get left behind.
>>
>> I would love to hear your thoughts on this matter!
>
> Hmm... I am not very keen on adding a new command just for cleaning up
> what, as far as my testing goes, looks like an artificial situation.
>
> If `lxc-console` processes don't linger under realistic conditions, we
> should evaluate whether we even need the cleanup, or if
> `remove_stale_lxc_consoles` was simply leftover legacy code.
True, I am starting to belive now as well that this might just be
leftover legacy code.
The current `remove_stale_lxc_consoles` in pvestatd.pm, only removes
lxc-consoles if the underlying container is not returned by
`PVE::LXC::vmstatus()`, i.e., if the container does not exist anymore.
To my knowledge this can only happen if the container config is broken
or missing.
I just figured that the existance of this code, might indicate some
important edge case, but perhaps we just should consider just removing it.
> And even if we want to handle this, I think it should remain automatic.
>
> Maybe we can bridge the gap to something I tried here:
> "add container console scrollback buffer"
> https://lore.proxmox.com/pve-devel/20260121112335.84473-1-f.schauer@proxmox.com/
> As it is right now, my v1 is not ready, but my point is that we could
> maybe have `lxc-console` + `dtach` running automatically at all times
> alongside `lxc-start`, by also having `lxc-console` managed by
> `pve-container@.service`. This way, we could manage both processes under
> the same systemd unit. Not sure if that's the direction we want to take,
> but it's an idea.
I think this is a great idea. Since we start the `pve-container@.service`
for each container anyway, I belive utilizing it to the related
processes/resources is correct direction.
My understanding
----------------
Maybe my idea is a bit too basic, but what about this:
Currently the pve-container@<id>.service contains
ExecStop=/usr/share/lxc/pve-container-stop-wrapper %i
To my knowledge this is only executed in a non-failure case.
We could extend the .service file to contain ExecStopPost=
implement a new `pve-container-stop-post-wrapper`
Add it to the .service config:
ExecStopPost=/usr/share/lxc/pve-container-stop-post-wrapper %i
according to [1], we can use this field when "the service exited
unexpectedly".
Inside the new wrapper we call something like:
lxc-stop <id>
# other cleanup ?
Your patch
----------
Let me know some of the details on how you plan for for
pve-container@<id>.service to manage the lxc-console and the improved
dtach process.
I think this idea is really interesting and I would like to
support/review your patches on this matter.
References
----------
[1] https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html
prev parent reply other threads:[~2026-10-09 9:44 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 8:37 [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07 8:37 ` [PATCH manager 1/2] fix #8093: pvestatd: remove cleanup of stale lxc consoles Elias Huhsovitz
2026-10-07 8:37 ` [PATCH container 2/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07 13:38 ` [PATCH container/manager 0/2] " Filip Schauer
2026-10-08 10:25 ` Elias Huhsovitz
2026-10-08 13:15 ` Filip Schauer
2026-10-09 9:44 ` Elias Huhsovitz [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=DM07LF3LD9RU.2VK0W1IC6WVKE@proxmox.com \
--to=e.huhsovitz@proxmox.com \
--cc=f.schauer@proxmox.com \
--cc=pve-devel@lists.proxmox.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox