public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
From: "Elias Huhsovitz" <e.huhsovitz@proxmox.com>
To: "Filip Schauer" <f.schauer@proxmox.com>, <pve-devel@lists.proxmox.com>
Subject: Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
Date: Fri, 09 Oct 2026 11:44:21 +0200	[thread overview]
Message-ID: <DM07LF3LD9RU.2VK0W1IC6WVKE@proxmox.com> (raw)
In-Reply-To: <82fceb8e-f52d-4e56-9dcf-5cd61393fa26@proxmox.com>

On Thu Oct 8, 2026 at 3:15 PM CEST, Filip Schauer wrote:
> On 08/10/2026 12:25, Elias Huhsovitz wrote:
>> On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote:
>>> On 07/10/2026 10:37, Elias Huhsovitz wrote:
>>>> The current Problem
>>>> -------------------
>>>> Console session processes can outlive their container when it stops or
>>>> is destroyed. The cleanup currently runs in pvestatd every 10 seconds.
>>>> It calls vmstatus() and kills stale processes by PID.
>>>>
>>>> The killing of the process can result in a race condition:
>>>> If the console gets closed by something else, just before the kill
>>>> command is executed, a different process can reuse the PID and we kill
>>>> a completely unrelated process.
>>>>
>>>> Also checking this every 10s is too much unnecessary overhead IMO.
>>>>
>>>> My understanding of a "better" architecture
>>>> -------------------------------------------
>>>> IMO console deletion should be tied to the state of the underlying
>>>> container. Also I was unable to find a "convenient" time/event for a
>>>> cleanup job to run.
>>>>
>>>> Proposed implementation
>>>> -----------------------
>>>> Wrap console commands in transient systemd scopes and bind them to the
>>>> container service via the PartOf= property. When the container stops,
>>>> systemd stops the scope, terminating the console session and the dtach
>>>> session. This allows removing the pvestatd periodic cleanup entirely.
>>>
>>> Getting away from the polling in pvestatd would be nice.
>>> However, I don't think `PartOf=` can be relied on here.
>>>
>>> As a test I created a scope running `sleep 1000` with
>>> `PartOf=pve-container@103.service`. Stopping the container left the
>>> scope and its process running. It seems like `lxc-stop` ends the service
>>> without a systemd stop job, and `PartOf=` only propagates explicit stop
>>> jobs. The same goes for a `poweroff` inside the container. As far as I
>>> can tell, the scope would only be stopped by `systemctl stop` or
>>> `systemctl restart`. So most stop paths are not covered.
>> 
>> Thats interesting. I was not able to re-produce this behaviour when
>> calling poweroff from inside the container
>> 
>> Did you create the console using `pct`, or via the API
>> (e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`.
>
> 1. I started container 109.
> 2. Then I ran `systemd-run --scope --unit=pve-lxc-console-109 --property "PartOf=pve-container@109.service" sleep 1000`
> 3. I watched the sleep process with `watch -n 0.1 'COLUMNS= ps aux | grep sleep'`
> 4. Then I logged into the container via xterm.js in the web UI and ran
>     the `poweroff` command.
> 5. The container stopped, but the sleep process kept running.
> 6. Only once explicitly calling
>     `systemctl stop pve-container@109.service`
>     did the sleep process terminate.
>
>
>> 
>> But nevertheless, you are correct. My assumptions about the `PartOf=`
>> were wrong.
>> 
>>>
>>> But do we even need to kill the lxc-console process in the first place?
>>> lxc-console already exits automatically when the container stops. Or am
>>> I missing something? Is there any case where the process lingers?
>> 
>> There are a few scenarios where lxc-console remains:
>> 
>> 1. rm-rf <id>.conf
>> 
>> while the process is running. Killing the container via lxc-stop <id>
>> leaves the console running.
>
> This one I could not reproduce.
> 1. I started a container.
> 2. I opened its console in the web UI.
> 3. I deleted the container config at /etc/pve/lxc/<id>.conf.
> 4. I stopped the container with `lxc-stop <id>`.
>     (also tried with `lxc-stop --kill <id>`)
> 5. The lxc-console process was no longer running after this.

I did the following:
1. Remove the subroutine `remove_stale_lxc_consoles` from pvestatd.pm
2. systemctl restart pvestatd
3. pct start 202
4. open the xterm web console for container 202

pgrep -af 'console|dtach|spiceterm|vncterm|termproxy':

15088 /usr/bin/termproxy 5900 --path /vms/202 --perm VM.Console
--vncticket-endpoint --verify-port --ticket-fd 6 -- /usr/bin/dtach
-A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15093 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15095 lxc-console -n 202 -e -1

5. rm -rf /etc/pve/lxc/202.conf
6. I close the xterm web console

pgrep -af 'console|dtach|spiceterm|vncterm|termproxy':
15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15095 lxc-console -n 202 -e -1

7. lxc-stop 202
lxc-console process is gone (I think i need to better standardzie my
testing process :/)

>
>> 
>> 2. pid=$(pgrep -f "lxc-console.*<id>")
>> kill -STOP $pid
>> lxc-stop <id>
>> 
>> also results in a remaining console.
>
> This one I was indeed able to reproduce.
>
>
>> 
>> In my testing yesterday, calling
>> 
>> pct detory <id> --force
>> while the container is still running, resulted in an orphaned console,
>> which prompted me to create this patch in the first place
>> 
>> But I am unable to re-produce this currently. So probably the issue was
>> something else...
>
> I am not able to reproduce this one either. When I run
> `pct destroy <id> --force` it first stops the container:
> ```
> forced to stop CT <id> before destroying!
> ```

Yeah, I think I made some mistake in my original testing here.
Sorry about that.

>
>> 
>> Bottom Line
>> -----------
>> The scenarios that leave an orphaned are not standard lifecycle events.
>> 
>> What about:
>> 
>> Moving the `remove_stale_lxc_consoles` logic (with some
>> improvements) to pve-container.
>> 
>> Introduce a new pct command `pct console cleanup` (name TBD), so that
>> administrators can cleanup the orphaned consoles in the edge cases
>> where an lxc-console might get left behind.
>> 
>> I would love to hear your thoughts on this matter!
>
> Hmm... I am not very keen on adding a new command just for cleaning up
> what, as far as my testing goes, looks like an artificial situation.
>
> If `lxc-console` processes don't linger under realistic conditions, we
> should evaluate whether we even need the cleanup, or if
> `remove_stale_lxc_consoles` was simply leftover legacy code.

True, I am starting to belive now as well that this might just be
leftover legacy code.

The current `remove_stale_lxc_consoles` in pvestatd.pm, only removes
lxc-consoles if the underlying container is not returned by 
`PVE::LXC::vmstatus()`, i.e., if the container does not exist anymore.

To my knowledge this can only happen if the container config is broken
or missing. 

I just figured that the existance of this code, might indicate some
important edge case, but perhaps we just should consider just removing it.

> And even if we want to handle this, I think it should remain automatic.
>
> Maybe we can bridge the gap to something I tried here:
> "add container console scrollback buffer"
> https://lore.proxmox.com/pve-devel/20260121112335.84473-1-f.schauer@proxmox.com/
> As it is right now, my v1 is not ready, but my point is that we could
> maybe have `lxc-console` + `dtach` running automatically at all times
> alongside `lxc-start`, by also having `lxc-console` managed by
> `pve-container@.service`. This way, we could manage both processes under
> the same systemd unit. Not sure if that's the direction we want to take,
> but it's an idea.

I think this is a great idea. Since we start the `pve-container@.service`
for each container anyway, I belive utilizing it to the related
processes/resources is correct direction.

My understanding
----------------
Maybe my idea is a bit too basic, but what about this:

Currently the pve-container@<id>.service contains

ExecStop=/usr/share/lxc/pve-container-stop-wrapper %i

To my knowledge this is only executed in a non-failure case.

We could extend the .service file to contain ExecStopPost=

implement a new `pve-container-stop-post-wrapper`

Add it to the .service config:
ExecStopPost=/usr/share/lxc/pve-container-stop-post-wrapper %i

according to [1], we can use this field when "the service exited
unexpectedly".

Inside the new wrapper we call something like:
lxc-stop <id>
# other cleanup ?

Your patch
----------
Let me know some of the details on how you plan for for
pve-container@<id>.service to manage the lxc-console and the improved
dtach process.

I think this idea is really interesting and I would like to
support/review your patches on this matter.

References
----------
[1] https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html




      reply	other threads:[~2026-10-09  9:44 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07  8:37 [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07  8:37 ` [PATCH manager 1/2] fix #8093: pvestatd: remove cleanup of stale lxc consoles Elias Huhsovitz
2026-10-07  8:37 ` [PATCH container 2/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07 13:38 ` [PATCH container/manager 0/2] " Filip Schauer
2026-10-08 10:25   ` Elias Huhsovitz
2026-10-08 13:15     ` Filip Schauer
2026-10-09  9:44       ` Elias Huhsovitz [this message]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=DM07LF3LD9RU.2VK0W1IC6WVKE@proxmox.com \
    --to=e.huhsovitz@proxmox.com \
    --cc=f.schauer@proxmox.com \
    --cc=pve-devel@lists.proxmox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal