public inbox for pve-devel@lists.proxmox.com
 help / color / mirror / Atom feed
From: Filip Schauer <f.schauer@proxmox.com>
To: Elias Huhsovitz <e.huhsovitz@proxmox.com>, pve-devel@lists.proxmox.com
Subject: Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
Date: Thu, 8 Oct 2026 15:15:28 +0200	[thread overview]
Message-ID: <82fceb8e-f52d-4e56-9dcf-5cd61393fa26@proxmox.com> (raw)
In-Reply-To: <DLZDUMGURYB7.1MGPBENZ88IGS@proxmox.com>

On 08/10/2026 12:25, Elias Huhsovitz wrote:
> On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote:
>> On 07/10/2026 10:37, Elias Huhsovitz wrote:
>>> The current Problem
>>> -------------------
>>> Console session processes can outlive their container when it stops or
>>> is destroyed. The cleanup currently runs in pvestatd every 10 seconds.
>>> It calls vmstatus() and kills stale processes by PID.
>>>
>>> The killing of the process can result in a race condition:
>>> If the console gets closed by something else, just before the kill
>>> command is executed, a different process can reuse the PID and we kill
>>> a completely unrelated process.
>>>
>>> Also checking this every 10s is too much unnecessary overhead IMO.
>>>
>>> My understanding of a "better" architecture
>>> -------------------------------------------
>>> IMO console deletion should be tied to the state of the underlying
>>> container. Also I was unable to find a "convenient" time/event for a
>>> cleanup job to run.
>>>
>>> Proposed implementation
>>> -----------------------
>>> Wrap console commands in transient systemd scopes and bind them to the
>>> container service via the PartOf= property. When the container stops,
>>> systemd stops the scope, terminating the console session and the dtach
>>> session. This allows removing the pvestatd periodic cleanup entirely.
>>
>> Getting away from the polling in pvestatd would be nice.
>> However, I don't think `PartOf=` can be relied on here.
>>
>> As a test I created a scope running `sleep 1000` with
>> `PartOf=pve-container@103.service`. Stopping the container left the
>> scope and its process running. It seems like `lxc-stop` ends the service
>> without a systemd stop job, and `PartOf=` only propagates explicit stop
>> jobs. The same goes for a `poweroff` inside the container. As far as I
>> can tell, the scope would only be stopped by `systemctl stop` or
>> `systemctl restart`. So most stop paths are not covered.
> 
> Thats interesting. I was not able to re-produce this behaviour when
> calling poweroff from inside the container
> 
> Did you create the console using `pct`, or via the API
> (e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`.

1. I started container 109.
2. Then I ran `systemd-run --scope --unit=pve-lxc-console-109 --property "PartOf=pve-container@109.service" sleep 1000`
3. I watched the sleep process with `watch -n 0.1 'COLUMNS= ps aux | grep sleep'`
4. Then I logged into the container via xterm.js in the web UI and ran
    the `poweroff` command.
5. The container stopped, but the sleep process kept running.
6. Only once explicitly calling
    `systemctl stop pve-container@109.service`
    did the sleep process terminate.


> 
> But nevertheless, you are correct. My assumptions about the `PartOf=`
> were wrong.
> 
>>
>> But do we even need to kill the lxc-console process in the first place?
>> lxc-console already exits automatically when the container stops. Or am
>> I missing something? Is there any case where the process lingers?
> 
> There are a few scenarios where lxc-console remains:
> 
> 1. rm-rf <id>.conf
> 
> while the process is running. Killing the container via lxc-stop <id>
> leaves the console running.

This one I could not reproduce.
1. I started a container.
2. I opened its console in the web UI.
3. I deleted the container config at /etc/pve/lxc/<id>.conf.
4. I stopped the container with `lxc-stop <id>`.
    (also tried with `lxc-stop --kill <id>`)
5. The lxc-console process was no longer running after this.


> 
> 2. pid=$(pgrep -f "lxc-console.*<id>")
> kill -STOP $pid
> lxc-stop <id>
> 
> also results in a remaining console.

This one I was indeed able to reproduce.


> 
> In my testing yesterday, calling
> 
> pct detory <id> --force
> while the container is still running, resulted in an orphaned console,
> which prompted me to create this patch in the first place
> 
> But I am unable to re-produce this currently. So probably the issue was
> something else...

I am not able to reproduce this one either. When I run
`pct destroy <id> --force` it first stops the container:
```
forced to stop CT <id> before destroying!
```


> 
> Bottom Line
> -----------
> The scenarios that leave an orphaned are not standard lifecycle events.
> 
> What about:
> 
> Moving the `remove_stale_lxc_consoles` logic (with some
> improvements) to pve-container.
> 
> Introduce a new pct command `pct console cleanup` (name TBD), so that
> administrators can cleanup the orphaned consoles in the edge cases
> where an lxc-console might get left behind.
> 
> I would love to hear your thoughts on this matter!

Hmm... I am not very keen on adding a new command just for cleaning up
what, as far as my testing goes, looks like an artificial situation.

If `lxc-console` processes don't linger under realistic conditions, we
should evaluate whether we even need the cleanup, or if
`remove_stale_lxc_consoles` was simply leftover legacy code.

And even if we want to handle this, I think it should remain automatic.

Maybe we can bridge the gap to something I tried here:
"add container console scrollback buffer"
https://lore.proxmox.com/pve-devel/20260121112335.84473-1-f.schauer@proxmox.com/
As it is right now, my v1 is not ready, but my point is that we could
maybe have `lxc-console` + `dtach` running automatically at all times
alongside `lxc-start`, by also having `lxc-console` managed by
`pve-container@.service`. This way, we could manage both processes under
the same systemd unit. Not sure if that's the direction we want to take,
but it's an idea.





  reply	other threads:[~2026-10-08 13:15 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-07  8:37 [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07  8:37 ` [PATCH manager 1/2] fix #8093: pvestatd: remove cleanup of stale lxc consoles Elias Huhsovitz
2026-10-07  8:37 ` [PATCH container 2/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07 13:38 ` [PATCH container/manager 0/2] " Filip Schauer
2026-10-08 10:25   ` Elias Huhsovitz
2026-10-08 13:15     ` Filip Schauer [this message]
2026-10-09  9:44       ` Elias Huhsovitz

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=82fceb8e-f52d-4e56-9dcf-5cd61393fa26@proxmox.com \
    --to=f.schauer@proxmox.com \
    --cc=e.huhsovitz@proxmox.com \
    --cc=pve-devel@lists.proxmox.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal