all lists on lists.proxmox.com
 help / color / mirror / Atom feed
From: "Elias Huhsovitz" <e.huhsovitz@proxmox.com>
To: "Filip Schauer" <f.schauer@proxmox.com>, <pve-devel@lists.proxmox.com>
Subject: Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
Date: Fri, 09 Oct 2026 11:44:21 +0200	[thread overview]
Message-ID: <DM07LF3LD9RU.2VK0W1IC6WVKE@proxmox.com> (raw)
In-Reply-To: <82fceb8e-f52d-4e56-9dcf-5cd61393fa26@proxmox.com>

On Thu Oct 8, 2026 at 3:15 PM CEST, Filip Schauer wrote:
> On 08/10/2026 12:25, Elias Huhsovitz wrote:
>> On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote:
>>> On 07/10/2026 10:37, Elias Huhsovitz wrote:
>>>> The current Problem
>>>> -------------------
>>>> Console session processes can outlive their container when it stops or
>>>> is destroyed. The cleanup currently runs in pvestatd every 10 seconds.
>>>> It calls vmstatus() and kills stale processes by PID.
>>>>
>>>> The killing of the process can result in a race condition:
>>>> If the console gets closed by something else, just before the kill
>>>> command is executed, a different process can reuse the PID and we kill
>>>> a completely unrelated process.
>>>>
>>>> Also checking this every 10s is too much unnecessary overhead IMO.
>>>>
>>>> My understanding of a "better" architecture
>>>> -------------------------------------------
>>>> IMO console deletion should be tied to the state of the underlying
>>>> container. Also I was unable to find a "convenient" time/event for a
>>>> cleanup job to run.
>>>>
>>>> Proposed implementation
>>>> -----------------------
>>>> Wrap console commands in transient systemd scopes and bind them to the
>>>> container service via the PartOf= property. When the container stops,
>>>> systemd stops the scope, terminating the console session and the dtach
>>>> session. This allows removing the pvestatd periodic cleanup entirely.
>>>
>>> Getting away from the polling in pvestatd would be nice.
>>> However, I don't think `PartOf=` can be relied on here.
>>>
>>> As a test I created a scope running `sleep 1000` with
>>> `PartOf=pve-container@103.service`. Stopping the container left the
>>> scope and its process running. It seems like `lxc-stop` ends the service
>>> without a systemd stop job, and `PartOf=` only propagates explicit stop
>>> jobs. The same goes for a `poweroff` inside the container. As far as I
>>> can tell, the scope would only be stopped by `systemctl stop` or
>>> `systemctl restart`. So most stop paths are not covered.
>> 
>> Thats interesting. I was not able to re-produce this behaviour when
>> calling poweroff from inside the container
>> 
>> Did you create the console using `pct`, or via the API
>> (e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`.
>
> 1. I started container 109.
> 2. Then I ran `systemd-run --scope --unit=pve-lxc-console-109 --property "PartOf=pve-container@109.service" sleep 1000`
> 3. I watched the sleep process with `watch -n 0.1 'COLUMNS= ps aux | grep sleep'`
> 4. Then I logged into the container via xterm.js in the web UI and ran
>     the `poweroff` command.
> 5. The container stopped, but the sleep process kept running.
> 6. Only once explicitly calling
>     `systemctl stop pve-container@109.service`
>     did the sleep process terminate.
>
>
>> 
>> But nevertheless, you are correct. My assumptions about the `PartOf=`
>> were wrong.
>> 
>>>
>>> But do we even need to kill the lxc-console process in the first place?
>>> lxc-console already exits automatically when the container stops. Or am
>>> I missing something? Is there any case where the process lingers?
>> 
>> There are a few scenarios where lxc-console remains:
>> 
>> 1. rm-rf <id>.conf
>> 
>> while the process is running. Killing the container via lxc-stop <id>
>> leaves the console running.
>
> This one I could not reproduce.
> 1. I started a container.
> 2. I opened its console in the web UI.
> 3. I deleted the container config at /etc/pve/lxc/<id>.conf.
> 4. I stopped the container with `lxc-stop <id>`.
>     (also tried with `lxc-stop --kill <id>`)
> 5. The lxc-console process was no longer running after this.

I did the following:
1. Remove the subroutine `remove_stale_lxc_consoles` from pvestatd.pm
2. systemctl restart pvestatd
3. pct start 202
4. open the xterm web console for container 202

pgrep -af 'console|dtach|spiceterm|vncterm|termproxy':

15088 /usr/bin/termproxy 5900 --path /vms/202 --perm VM.Console
--vncticket-endpoint --verify-port --ticket-fd 6 -- /usr/bin/dtach
-A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15093 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15095 lxc-console -n 202 -e -1

5. rm -rf /etc/pve/lxc/202.conf
6. I close the xterm web console

pgrep -af 'console|dtach|spiceterm|vncterm|termproxy':
15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15095 lxc-console -n 202 -e -1

7. lxc-stop 202
lxc-console process is gone (I think i need to better standardzie my
testing process :/)

>
>> 
>> 2. pid=$(pgrep -f "lxc-console.*<id>")
>> kill -STOP $pid
>> lxc-stop <id>
>> 
>> also results in a remaining console.
>
> This one I was indeed able to reproduce.
>
>
>> 
>> In my testing yesterday, calling
>> 
>> pct detory <id> --force
>> while the container is still running, resulted in an orphaned console,
>> which prompted me to create this patch in the first place
>> 
>> But I am unable to re-produce this currently. So probably the issue was
>> something else...
>
> I am not able to reproduce this one either. When I run
> `pct destroy <id> --force` it first stops the container:
> ```
> forced to stop CT <id> before destroying!
> ```

Yeah, I think I made some mistake in my original testing here.
Sorry about that.

>
>> 
>> Bottom Line
>> -----------
>> The scenarios that leave an orphaned are not standard lifecycle events.
>> 
>> What about:
>> 
>> Moving the `remove_stale_lxc_consoles` logic (with some
>> improvements) to pve-container.
>> 
>> Introduce a new pct command `pct console cleanup` (name TBD), so that
>> administrators can cleanup the orphaned consoles in the edge cases
>> where an lxc-console might get left behind.
>> 
>> I would love to hear your thoughts on this matter!
>
> Hmm... I am not very keen on adding a new command just for cleaning up
> what, as far as my testing goes, looks like an artificial situation.
>
> If `lxc-console` processes don't linger under realistic conditions, we
> should evaluate whether we even need the cleanup, or if
> `remove_stale_lxc_consoles` was simply leftover legacy code.

True, I am starting to belive now as well that this might just be
leftover legacy code.

The current `remove_stale_lxc_consoles` in pvestatd.pm, only removes
lxc-consoles if the underlying container is not returned by 
`PVE::LXC::vmstatus()`, i.e., if the container does not exist anymore.

To my knowledge this can only happen if the container config is broken
or missing. 

I just figured that the existance of this code, might indicate some
important edge case, but perhaps we just should consider just removing it.

> And even if we want to handle this, I think it should remain automatic.
>
> Maybe we can bridge the gap to something I tried here:
> "add container console scrollback buffer"
> https://lore.proxmox.com/pve-devel/20260121112335.84473-1-f.schauer@proxmox.com/
> As it is right now, my v1 is not ready, but my point is that we could
> maybe have `lxc-console` + `dtach` running automatically at all times
> alongside `lxc-start`, by also having `lxc-console` managed by
> `pve-container@.service`. This way, we could manage both processes under
> the same systemd unit. Not sure if that's the direction we want to take,
> but it's an idea.

I think this is a great idea. Since we start the `pve-container@.service`
for each container anyway, I belive utilizing it to the related
processes/resources is correct direction.

My understanding
----------------
Maybe my idea is a bit too basic, but what about this:

Currently the pve-container@<id>.service contains

ExecStop=/usr/share/lxc/pve-container-stop-wrapper %i

To my knowledge this is only executed in a non-failure case.

We could extend the .service file to contain ExecStopPost=

implement a new `pve-container-stop-post-wrapper`

Add it to the .service config:
ExecStopPost=/usr/share/lxc/pve-container-stop-post-wrapper %i

according to [1], we can use this field when "the service exited
unexpectedly".

Inside the new wrapper we call something like:
lxc-stop <id>
# other cleanup ?

Your patch
----------
Let me know some of the details on how you plan for for
pve-container@<id>.service to manage the lxc-console and the improved
dtach process.

I think this idea is really interesting and I would like to
support/review your patches on this matter.

References
----------
[1] https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html




      reply	other threads:[~2026-10-09  9:44 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07  8:37 [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07  8:37 ` [PATCH manager 1/2] fix #8093: pvestatd: remove cleanup of stale lxc consoles Elias Huhsovitz
2026-10-07  8:37 ` [PATCH container 2/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07 13:38 ` [PATCH container/manager 0/2] " Filip Schauer
2026-10-08 10:25   ` Elias Huhsovitz
2026-10-08 13:15     ` Filip Schauer
2026-10-09  9:44       ` Elias Huhsovitz [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DM07LF3LD9RU.2VK0W1IC6WVKE@proxmox.com \
    --to=e.huhsovitz@proxmox.com \
    --cc=f.schauer@proxmox.com \
    --cc=pve-devel@lists.proxmox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal