all lists on lists.proxmox.com
 help / color / mirror / Atom feed
* [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
@ 2026-10-07  8:37 Elias Huhsovitz
  2026-10-07  8:37 ` [PATCH manager 1/2] fix #8093: pvestatd: remove cleanup of stale lxc consoles Elias Huhsovitz
                   ` (2 more replies)
  0 siblings, 3 replies; 7+ messages in thread
From: Elias Huhsovitz @ 2026-10-07  8:37 UTC (permalink / raw)
  To: pve-devel; +Cc: Elias Huhsovitz

The current Problem
-------------------
Console session processes can outlive their container when it stops or
is destroyed. The cleanup currently runs in pvestatd every 10 seconds.
It calls vmstatus() and kills stale processes by PID.

The killing of the process can result in a race condition:
If the console gets closed by something else, just before the kill
command is executed, a different process can reuse the PID and we kill
a completely unrelated process.

Also checking this every 10s is too much unnecessary overhead IMO.

My understanding of a "better" architecture
-------------------------------------------
IMO console deletion should be tied to the state of the underlying
container. Also I was unable to find a "convenient" time/event for a
cleanup job to run.

Proposed implementation
-----------------------
Wrap console commands in transient systemd scopes and bind them to the
container service via the PartOf= property. When the container stops,
systemd stops the scope, terminating the console session and the dtach
session. This allows removing the pvestatd periodic cleanup entirely.

Summary of Changes
------------------

pve-manager:

Elias Huhsovitz (1):
  fix #8093: pvestatd: remove cleanup of stale lxc consoles

 PVE/Service/pvestatd.pm | 18 ------------------
 1 file changed, 18 deletions(-)


pve-container:

Elias Huhsovitz (1):
  fix #8093: console: bind console sessions to container lifetime via
    systemd scopes

 src/PVE/API2/LXC.pm |  6 +++---
 src/PVE/CLI/pct.pm  | 10 +++++++---
 src/PVE/LXC.pm      | 41 +++++++++++++++++++++++++++++++++++++++++
 3 files changed, 51 insertions(+), 6 deletions(-)


Summary over all repositories:
  4 files changed, 51 insertions(+), 24 deletions(-)

-- 
Generated by murpp 0.12.0




^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH manager 1/2] fix #8093: pvestatd: remove cleanup of stale lxc consoles
  2026-10-07  8:37 [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
@ 2026-10-07  8:37 ` Elias Huhsovitz
  2026-10-07  8:37 ` [PATCH container 2/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
  2026-10-07 13:38 ` [PATCH container/manager 0/2] " Filip Schauer
  2 siblings, 0 replies; 7+ messages in thread
From: Elias Huhsovitz @ 2026-10-07  8:37 UTC (permalink / raw)
  To: pve-devel; +Cc: Elias Huhsovitz

Responsibility is moved to pve-container.

Signed-off-by: Elias Huhsovitz <e.huhsovitz@proxmox.com>
---
 PVE/Service/pvestatd.pm | 18 ------------------
 1 file changed, 18 deletions(-)

diff --git a/PVE/Service/pvestatd.pm b/PVE/Service/pvestatd.pm
index 34f9da71..cec66e00 100755
--- a/PVE/Service/pvestatd.pm
+++ b/PVE/Service/pvestatd.pm
@@ -444,20 +444,6 @@ sub update_qemu_status {
     PVE::PullMetric::update($pull_txn, 'qemu', $vmstatus, $ctime);
 }
 
-sub remove_stale_lxc_consoles {
-
-    my $vmstatus = PVE::LXC::vmstatus();
-    my $pidhash = PVE::LXC::find_lxc_console_pids();
-
-    foreach my $vmid (keys %$pidhash) {
-        next if defined($vmstatus->{$vmid});
-        syslog('info', "remove stale lxc-console for CT $vmid");
-        foreach my $pid (@{ $pidhash->{$vmid} }) {
-            kill(9, $pid);
-        }
-    }
-}
-
 my $rebalance_error_count = {};
 
 my $NO_REBALANCE;
@@ -848,10 +834,6 @@ sub update_status {
     $err = $@;
     syslog('err', "storage status update error: $err") if $err;
 
-    eval { remove_stale_lxc_consoles(); };
-    $err = $@;
-    syslog('err', "lxc console cleanup error: $err") if $err;
-
     eval { rotate_authkeys(); };
     $err = $@;
     syslog('err', "authkey rotation error: $err") if $err;
-- 
2.47.3





^ permalink raw reply related	[flat|nested] 7+ messages in thread

* [PATCH container 2/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
  2026-10-07  8:37 [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
  2026-10-07  8:37 ` [PATCH manager 1/2] fix #8093: pvestatd: remove cleanup of stale lxc consoles Elias Huhsovitz
@ 2026-10-07  8:37 ` Elias Huhsovitz
  2026-10-07 13:38 ` [PATCH container/manager 0/2] " Filip Schauer
  2 siblings, 0 replies; 7+ messages in thread
From: Elias Huhsovitz @ 2026-10-07  8:37 UTC (permalink / raw)
  To: pve-devel; +Cc: Elias Huhsovitz

Console processes can outlive their useful session when a container
stops or disappears while a console is attached. The previous cleanup
ran in pvestatd, which polled every 10s: forked lxc-info for every
running container, and killed orphaned processes by their PID.

Manage console lifetime in the container lifecycle instead. Add
wrap_in_console_scope() to start console commands in a systemd scope.
Bind each scope with PartOf= to the respective pve-container service.
Use the wrapper for pct console/enter/exec, and the API console
endpoints.

When the container service stops, systemd stops the related console
scopes.

Signed-off-by: Elias Huhsovitz <e.huhsovitz@proxmox.com>
---
 src/PVE/API2/LXC.pm |  6 +++---
 src/PVE/CLI/pct.pm  | 10 +++++++---
 src/PVE/LXC.pm      | 41 +++++++++++++++++++++++++++++++++++++++++
 3 files changed, 51 insertions(+), 6 deletions(-)

diff --git a/src/PVE/API2/LXC.pm b/src/PVE/API2/LXC.pm
index 5f94d5a..3fd2bf8 100644
--- a/src/PVE/API2/LXC.pm
+++ b/src/PVE/API2/LXC.pm
@@ -1015,7 +1015,7 @@ __PACKAGE__->register_method({
         my $ticket = PVE::AccessControl::assemble_vnc_ticket($authuser, $authpath, $port);
 
         my $conf = PVE::LXC::Config->load_config($vmid, $node);
-        my $concmd = PVE::LXC::get_console_command($vmid, $conf, -1);
+        my $concmd = PVE::LXC::get_console_command_scoped($vmid, $conf, -1);
 
         my $shcmd = [
             '/usr/bin/dtach',
@@ -1146,7 +1146,7 @@ __PACKAGE__->register_method({
         my $ticket = PVE::AccessControl::assemble_vnc_ticket($authuser, $authpath, $port);
 
         my $conf = PVE::LXC::Config->load_config($vmid, $node);
-        my $concmd = PVE::LXC::get_console_command($vmid, $conf, -1);
+        my $concmd = PVE::LXC::get_console_command_scoped($vmid, $conf, -1);
 
         my $shcmd = [
             '/usr/bin/dtach',
@@ -1288,7 +1288,7 @@ __PACKAGE__->register_method({
 
         die "CT $vmid not running\n" if !PVE::LXC::check_running($vmid);
 
-        my $concmd = PVE::LXC::get_console_command($vmid, $conf);
+        my $concmd = PVE::LXC::get_console_command_scoped($vmid, $conf);
 
         my $shcmd = [
             '/usr/bin/dtach',
diff --git a/src/PVE/CLI/pct.pm b/src/PVE/CLI/pct.pm
index ceb0bea..11e85bc 100755
--- a/src/PVE/CLI/pct.pm
+++ b/src/PVE/CLI/pct.pm
@@ -153,7 +153,7 @@ __PACKAGE__->register_method({
         # test if container exists on this node
         my $conf = PVE::LXC::Config->load_config($param->{vmid});
 
-        my $cmd = PVE::LXC::get_console_command($param->{vmid}, $conf, $param->{escape});
+        my $cmd = PVE::LXC::get_console_command_scoped($param->{vmid}, $conf, $param->{escape});
         exec(@$cmd);
     },
 });
@@ -207,7 +207,9 @@ __PACKAGE__->register_method({
 
         my @lxc_attach_cmd = ('lxc-attach', '-n', $vmid);
         push @lxc_attach_cmd, $keep_env ? '--keep-env' : '--clear-env';
-        exec(@lxc_attach_cmd);
+
+        my $scoped_cmd = PVE::LXC::wrap_in_console_scope($vmid, \@lxc_attach_cmd);
+        exec(@$scoped_cmd);
     },
 });
 
@@ -250,7 +252,9 @@ __PACKAGE__->register_method({
         my @lxc_attach_cmd = ('lxc-attach', '-n', $vmid);
         push @lxc_attach_cmd, $keep_env ? '--keep-env' : '--clear-env';
         push @lxc_attach_cmd, '--', @{ $param->{'extra-args'} };
-        exec(@lxc_attach_cmd);
+
+        my $scoped_cmd = PVE::LXC::wrap_in_console_scope($vmid, \@lxc_attach_cmd);
+        exec(@$scoped_cmd);
     },
 });
 
diff --git a/src/PVE/LXC.pm b/src/PVE/LXC.pm
index dee073d..14bc6e4 100644
--- a/src/PVE/LXC.pm
+++ b/src/PVE/LXC.pm
@@ -954,6 +954,39 @@ sub verify_searchdomain_list {
     return join(' ', @list);
 }
 
+=head2 wrap_in_console_scope
+
+Wraps a command array in a transient systemd scope and binds it to the
+container's main systemd service. This ensures the console process is
+automatically terminated by systemd when the container stops, preventing
+orphaned processes without requiring manual cleanup scripts.
+
+=cut
+
+sub wrap_in_console_scope {
+    my ($vmid, $cmd) = @_;
+
+    my $session_id = "$$-" . time();
+    my $unit_name = "pve-lxc-console-$vmid-$session_id";
+
+    my $service = "pve-container\@$vmid.service";
+    # service name appends "-debug" when container is started in debug mode
+    my $debug_service = "pve-container-debug\@$vmid.service";
+
+    my @scope_cmd = (
+        'systemd-run',
+        '--scope',
+        '--unit', $unit_name,
+        '--description', "PVE LXC Console for $vmid",
+        '--property', "PartOf=$service $debug_service",
+        '--quiet',
+    );
+
+    push @scope_cmd, @$cmd;
+
+    return \@scope_cmd;
+}
+
 sub get_console_command {
     my ($vmid, $conf, $escapechar) = @_;
 
@@ -978,6 +1011,14 @@ sub get_console_command {
     return $cmd;
 }
 
+sub get_console_command_scoped {
+    my ($vmid, $conf, $escapechar) = @_;
+
+    my $cmd = get_console_command($vmid, $conf, $escapechar);
+
+    return wrap_in_console_scope($vmid, $cmd)
+}
+
 sub get_primary_ips {
     my ($conf) = @_;
 
-- 
2.47.3





^ permalink raw reply related	[flat|nested] 7+ messages in thread

* Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
  2026-10-07  8:37 [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
  2026-10-07  8:37 ` [PATCH manager 1/2] fix #8093: pvestatd: remove cleanup of stale lxc consoles Elias Huhsovitz
  2026-10-07  8:37 ` [PATCH container 2/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
@ 2026-10-07 13:38 ` Filip Schauer
  2026-10-08 10:25   ` Elias Huhsovitz
  2 siblings, 1 reply; 7+ messages in thread
From: Filip Schauer @ 2026-10-07 13:38 UTC (permalink / raw)
  To: Elias Huhsovitz, pve-devel

On 07/10/2026 10:37, Elias Huhsovitz wrote:
> The current Problem
> -------------------
> Console session processes can outlive their container when it stops or
> is destroyed. The cleanup currently runs in pvestatd every 10 seconds.
> It calls vmstatus() and kills stale processes by PID.
> 
> The killing of the process can result in a race condition:
> If the console gets closed by something else, just before the kill
> command is executed, a different process can reuse the PID and we kill
> a completely unrelated process.
> 
> Also checking this every 10s is too much unnecessary overhead IMO.
> 
> My understanding of a "better" architecture
> -------------------------------------------
> IMO console deletion should be tied to the state of the underlying
> container. Also I was unable to find a "convenient" time/event for a
> cleanup job to run.
> 
> Proposed implementation
> -----------------------
> Wrap console commands in transient systemd scopes and bind them to the
> container service via the PartOf= property. When the container stops,
> systemd stops the scope, terminating the console session and the dtach
> session. This allows removing the pvestatd periodic cleanup entirely.

Getting away from the polling in pvestatd would be nice.
However, I don't think `PartOf=` can be relied on here.

As a test I created a scope running `sleep 1000` with
`PartOf=pve-container@103.service`. Stopping the container left the
scope and its process running. It seems like `lxc-stop` ends the service
without a systemd stop job, and `PartOf=` only propagates explicit stop
jobs. The same goes for a `poweroff` inside the container. As far as I
can tell, the scope would only be stopped by `systemctl stop` or
`systemctl restart`. So most stop paths are not covered.

But do we even need to kill the lxc-console process in the first place?
lxc-console already exits automatically when the container stops. Or am
I missing something? Is there any case where the process lingers?




^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
  2026-10-07 13:38 ` [PATCH container/manager 0/2] " Filip Schauer
@ 2026-10-08 10:25   ` Elias Huhsovitz
  2026-10-08 13:15     ` Filip Schauer
  0 siblings, 1 reply; 7+ messages in thread
From: Elias Huhsovitz @ 2026-10-08 10:25 UTC (permalink / raw)
  To: Filip Schauer, pve-devel

On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote:
> On 07/10/2026 10:37, Elias Huhsovitz wrote:
>> The current Problem
>> -------------------
>> Console session processes can outlive their container when it stops or
>> is destroyed. The cleanup currently runs in pvestatd every 10 seconds.
>> It calls vmstatus() and kills stale processes by PID.
>> 
>> The killing of the process can result in a race condition:
>> If the console gets closed by something else, just before the kill
>> command is executed, a different process can reuse the PID and we kill
>> a completely unrelated process.
>> 
>> Also checking this every 10s is too much unnecessary overhead IMO.
>> 
>> My understanding of a "better" architecture
>> -------------------------------------------
>> IMO console deletion should be tied to the state of the underlying
>> container. Also I was unable to find a "convenient" time/event for a
>> cleanup job to run.
>> 
>> Proposed implementation
>> -----------------------
>> Wrap console commands in transient systemd scopes and bind them to the
>> container service via the PartOf= property. When the container stops,
>> systemd stops the scope, terminating the console session and the dtach
>> session. This allows removing the pvestatd periodic cleanup entirely.
>
> Getting away from the polling in pvestatd would be nice.
> However, I don't think `PartOf=` can be relied on here.
>
> As a test I created a scope running `sleep 1000` with
> `PartOf=pve-container@103.service`. Stopping the container left the
> scope and its process running. It seems like `lxc-stop` ends the service
> without a systemd stop job, and `PartOf=` only propagates explicit stop
> jobs. The same goes for a `poweroff` inside the container. As far as I
> can tell, the scope would only be stopped by `systemctl stop` or
> `systemctl restart`. So most stop paths are not covered.

Thats interesting. I was not able to re-produce this behaviour when
calling poweroff from inside the container

Did you create the console using `pct`, or via the API 
(e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`.

But nevertheless, you are correct. My assumptions about the `PartOf=`
were wrong.

>
> But do we even need to kill the lxc-console process in the first place?
> lxc-console already exits automatically when the container stops. Or am
> I missing something? Is there any case where the process lingers?

There are a few scenarios where lxc-console remains:

1. rm-rf <id>.conf 

while the process is running. Killing the container via lxc-stop <id>
leaves the console running.

2. pid=$(pgrep -f "lxc-console.*<id>")
kill -STOP $pid
lxc-stop <id>

also results in a remaining console.

In my testing yesterday, calling 

pct detory <id> --force
while the container is still running, resulted in an orphaned console,
which prompted me to create this patch in the first place

But I am unable to re-produce this currently. So probably the issue was
something else...

Bottom Line
-----------
The scenarios that leave an orphaned are not standard lifecycle events.

What about:

Moving the `remove_stale_lxc_consoles` logic (with some
improvements) to pve-container.

Introduce a new pct command `pct console cleanup` (name TBD), so that
administrators can cleanup the orphaned consoles in the edge cases
where an lxc-console might get left behind.

I would love to hear your thoughts on this matter!






^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
  2026-10-08 10:25   ` Elias Huhsovitz
@ 2026-10-08 13:15     ` Filip Schauer
  2026-10-09  9:44       ` Elias Huhsovitz
  0 siblings, 1 reply; 7+ messages in thread
From: Filip Schauer @ 2026-10-08 13:15 UTC (permalink / raw)
  To: Elias Huhsovitz, pve-devel

On 08/10/2026 12:25, Elias Huhsovitz wrote:
> On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote:
>> On 07/10/2026 10:37, Elias Huhsovitz wrote:
>>> The current Problem
>>> -------------------
>>> Console session processes can outlive their container when it stops or
>>> is destroyed. The cleanup currently runs in pvestatd every 10 seconds.
>>> It calls vmstatus() and kills stale processes by PID.
>>>
>>> The killing of the process can result in a race condition:
>>> If the console gets closed by something else, just before the kill
>>> command is executed, a different process can reuse the PID and we kill
>>> a completely unrelated process.
>>>
>>> Also checking this every 10s is too much unnecessary overhead IMO.
>>>
>>> My understanding of a "better" architecture
>>> -------------------------------------------
>>> IMO console deletion should be tied to the state of the underlying
>>> container. Also I was unable to find a "convenient" time/event for a
>>> cleanup job to run.
>>>
>>> Proposed implementation
>>> -----------------------
>>> Wrap console commands in transient systemd scopes and bind them to the
>>> container service via the PartOf= property. When the container stops,
>>> systemd stops the scope, terminating the console session and the dtach
>>> session. This allows removing the pvestatd periodic cleanup entirely.
>>
>> Getting away from the polling in pvestatd would be nice.
>> However, I don't think `PartOf=` can be relied on here.
>>
>> As a test I created a scope running `sleep 1000` with
>> `PartOf=pve-container@103.service`. Stopping the container left the
>> scope and its process running. It seems like `lxc-stop` ends the service
>> without a systemd stop job, and `PartOf=` only propagates explicit stop
>> jobs. The same goes for a `poweroff` inside the container. As far as I
>> can tell, the scope would only be stopped by `systemctl stop` or
>> `systemctl restart`. So most stop paths are not covered.
> 
> Thats interesting. I was not able to re-produce this behaviour when
> calling poweroff from inside the container
> 
> Did you create the console using `pct`, or via the API
> (e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`.

1. I started container 109.
2. Then I ran `systemd-run --scope --unit=pve-lxc-console-109 --property "PartOf=pve-container@109.service" sleep 1000`
3. I watched the sleep process with `watch -n 0.1 'COLUMNS= ps aux | grep sleep'`
4. Then I logged into the container via xterm.js in the web UI and ran
    the `poweroff` command.
5. The container stopped, but the sleep process kept running.
6. Only once explicitly calling
    `systemctl stop pve-container@109.service`
    did the sleep process terminate.


> 
> But nevertheless, you are correct. My assumptions about the `PartOf=`
> were wrong.
> 
>>
>> But do we even need to kill the lxc-console process in the first place?
>> lxc-console already exits automatically when the container stops. Or am
>> I missing something? Is there any case where the process lingers?
> 
> There are a few scenarios where lxc-console remains:
> 
> 1. rm-rf <id>.conf
> 
> while the process is running. Killing the container via lxc-stop <id>
> leaves the console running.

This one I could not reproduce.
1. I started a container.
2. I opened its console in the web UI.
3. I deleted the container config at /etc/pve/lxc/<id>.conf.
4. I stopped the container with `lxc-stop <id>`.
    (also tried with `lxc-stop --kill <id>`)
5. The lxc-console process was no longer running after this.


> 
> 2. pid=$(pgrep -f "lxc-console.*<id>")
> kill -STOP $pid
> lxc-stop <id>
> 
> also results in a remaining console.

This one I was indeed able to reproduce.


> 
> In my testing yesterday, calling
> 
> pct detory <id> --force
> while the container is still running, resulted in an orphaned console,
> which prompted me to create this patch in the first place
> 
> But I am unable to re-produce this currently. So probably the issue was
> something else...

I am not able to reproduce this one either. When I run
`pct destroy <id> --force` it first stops the container:
```
forced to stop CT <id> before destroying!
```


> 
> Bottom Line
> -----------
> The scenarios that leave an orphaned are not standard lifecycle events.
> 
> What about:
> 
> Moving the `remove_stale_lxc_consoles` logic (with some
> improvements) to pve-container.
> 
> Introduce a new pct command `pct console cleanup` (name TBD), so that
> administrators can cleanup the orphaned consoles in the edge cases
> where an lxc-console might get left behind.
> 
> I would love to hear your thoughts on this matter!

Hmm... I am not very keen on adding a new command just for cleaning up
what, as far as my testing goes, looks like an artificial situation.

If `lxc-console` processes don't linger under realistic conditions, we
should evaluate whether we even need the cleanup, or if
`remove_stale_lxc_consoles` was simply leftover legacy code.

And even if we want to handle this, I think it should remain automatic.

Maybe we can bridge the gap to something I tried here:
"add container console scrollback buffer"
https://lore.proxmox.com/pve-devel/20260121112335.84473-1-f.schauer@proxmox.com/
As it is right now, my v1 is not ready, but my point is that we could
maybe have `lxc-console` + `dtach` running automatically at all times
alongside `lxc-start`, by also having `lxc-console` managed by
`pve-container@.service`. This way, we could manage both processes under
the same systemd unit. Not sure if that's the direction we want to take,
but it's an idea.





^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes
  2026-10-08 13:15     ` Filip Schauer
@ 2026-10-09  9:44       ` Elias Huhsovitz
  0 siblings, 0 replies; 7+ messages in thread
From: Elias Huhsovitz @ 2026-10-09  9:44 UTC (permalink / raw)
  To: Filip Schauer, pve-devel

On Thu Oct 8, 2026 at 3:15 PM CEST, Filip Schauer wrote:
> On 08/10/2026 12:25, Elias Huhsovitz wrote:
>> On Wed Oct 7, 2026 at 3:38 PM CEST, Filip Schauer wrote:
>>> On 07/10/2026 10:37, Elias Huhsovitz wrote:
>>>> The current Problem
>>>> -------------------
>>>> Console session processes can outlive their container when it stops or
>>>> is destroyed. The cleanup currently runs in pvestatd every 10 seconds.
>>>> It calls vmstatus() and kills stale processes by PID.
>>>>
>>>> The killing of the process can result in a race condition:
>>>> If the console gets closed by something else, just before the kill
>>>> command is executed, a different process can reuse the PID and we kill
>>>> a completely unrelated process.
>>>>
>>>> Also checking this every 10s is too much unnecessary overhead IMO.
>>>>
>>>> My understanding of a "better" architecture
>>>> -------------------------------------------
>>>> IMO console deletion should be tied to the state of the underlying
>>>> container. Also I was unable to find a "convenient" time/event for a
>>>> cleanup job to run.
>>>>
>>>> Proposed implementation
>>>> -----------------------
>>>> Wrap console commands in transient systemd scopes and bind them to the
>>>> container service via the PartOf= property. When the container stops,
>>>> systemd stops the scope, terminating the console session and the dtach
>>>> session. This allows removing the pvestatd periodic cleanup entirely.
>>>
>>> Getting away from the polling in pvestatd would be nice.
>>> However, I don't think `PartOf=` can be relied on here.
>>>
>>> As a test I created a scope running `sleep 1000` with
>>> `PartOf=pve-container@103.service`. Stopping the container left the
>>> scope and its process running. It seems like `lxc-stop` ends the service
>>> without a systemd stop job, and `PartOf=` only propagates explicit stop
>>> jobs. The same goes for a `poweroff` inside the container. As far as I
>>> can tell, the scope would only be stopped by `systemctl stop` or
>>> `systemctl restart`. So most stop paths are not covered.
>> 
>> Thats interesting. I was not able to re-produce this behaviour when
>> calling poweroff from inside the container
>> 
>> Did you create the console using `pct`, or via the API
>> (e.g. `/nodes/{node}/lxc/{vmid}/vncproxy)`.
>
> 1. I started container 109.
> 2. Then I ran `systemd-run --scope --unit=pve-lxc-console-109 --property "PartOf=pve-container@109.service" sleep 1000`
> 3. I watched the sleep process with `watch -n 0.1 'COLUMNS= ps aux | grep sleep'`
> 4. Then I logged into the container via xterm.js in the web UI and ran
>     the `poweroff` command.
> 5. The container stopped, but the sleep process kept running.
> 6. Only once explicitly calling
>     `systemctl stop pve-container@109.service`
>     did the sleep process terminate.
>
>
>> 
>> But nevertheless, you are correct. My assumptions about the `PartOf=`
>> were wrong.
>> 
>>>
>>> But do we even need to kill the lxc-console process in the first place?
>>> lxc-console already exits automatically when the container stops. Or am
>>> I missing something? Is there any case where the process lingers?
>> 
>> There are a few scenarios where lxc-console remains:
>> 
>> 1. rm-rf <id>.conf
>> 
>> while the process is running. Killing the container via lxc-stop <id>
>> leaves the console running.
>
> This one I could not reproduce.
> 1. I started a container.
> 2. I opened its console in the web UI.
> 3. I deleted the container config at /etc/pve/lxc/<id>.conf.
> 4. I stopped the container with `lxc-stop <id>`.
>     (also tried with `lxc-stop --kill <id>`)
> 5. The lxc-console process was no longer running after this.

I did the following:
1. Remove the subroutine `remove_stale_lxc_consoles` from pvestatd.pm
2. systemctl restart pvestatd
3. pct start 202
4. open the xterm web console for container 202

pgrep -af 'console|dtach|spiceterm|vncterm|termproxy':

15088 /usr/bin/termproxy 5900 --path /vms/202 --perm VM.Console
--vncticket-endpoint --verify-port --ticket-fd 6 -- /usr/bin/dtach
-A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15093 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15095 lxc-console -n 202 -e -1

5. rm -rf /etc/pve/lxc/202.conf
6. I close the xterm web console

pgrep -af 'console|dtach|spiceterm|vncterm|termproxy':
15094 /usr/bin/dtach -A /var/run/dtach/vzctlconsole202 -r winch -z lxc-console -n 202 -e -1
15095 lxc-console -n 202 -e -1

7. lxc-stop 202
lxc-console process is gone (I think i need to better standardzie my
testing process :/)

>
>> 
>> 2. pid=$(pgrep -f "lxc-console.*<id>")
>> kill -STOP $pid
>> lxc-stop <id>
>> 
>> also results in a remaining console.
>
> This one I was indeed able to reproduce.
>
>
>> 
>> In my testing yesterday, calling
>> 
>> pct detory <id> --force
>> while the container is still running, resulted in an orphaned console,
>> which prompted me to create this patch in the first place
>> 
>> But I am unable to re-produce this currently. So probably the issue was
>> something else...
>
> I am not able to reproduce this one either. When I run
> `pct destroy <id> --force` it first stops the container:
> ```
> forced to stop CT <id> before destroying!
> ```

Yeah, I think I made some mistake in my original testing here.
Sorry about that.

>
>> 
>> Bottom Line
>> -----------
>> The scenarios that leave an orphaned are not standard lifecycle events.
>> 
>> What about:
>> 
>> Moving the `remove_stale_lxc_consoles` logic (with some
>> improvements) to pve-container.
>> 
>> Introduce a new pct command `pct console cleanup` (name TBD), so that
>> administrators can cleanup the orphaned consoles in the edge cases
>> where an lxc-console might get left behind.
>> 
>> I would love to hear your thoughts on this matter!
>
> Hmm... I am not very keen on adding a new command just for cleaning up
> what, as far as my testing goes, looks like an artificial situation.
>
> If `lxc-console` processes don't linger under realistic conditions, we
> should evaluate whether we even need the cleanup, or if
> `remove_stale_lxc_consoles` was simply leftover legacy code.

True, I am starting to belive now as well that this might just be
leftover legacy code.

The current `remove_stale_lxc_consoles` in pvestatd.pm, only removes
lxc-consoles if the underlying container is not returned by 
`PVE::LXC::vmstatus()`, i.e., if the container does not exist anymore.

To my knowledge this can only happen if the container config is broken
or missing. 

I just figured that the existance of this code, might indicate some
important edge case, but perhaps we just should consider just removing it.

> And even if we want to handle this, I think it should remain automatic.
>
> Maybe we can bridge the gap to something I tried here:
> "add container console scrollback buffer"
> https://lore.proxmox.com/pve-devel/20260121112335.84473-1-f.schauer@proxmox.com/
> As it is right now, my v1 is not ready, but my point is that we could
> maybe have `lxc-console` + `dtach` running automatically at all times
> alongside `lxc-start`, by also having `lxc-console` managed by
> `pve-container@.service`. This way, we could manage both processes under
> the same systemd unit. Not sure if that's the direction we want to take,
> but it's an idea.

I think this is a great idea. Since we start the `pve-container@.service`
for each container anyway, I belive utilizing it to the related
processes/resources is correct direction.

My understanding
----------------
Maybe my idea is a bit too basic, but what about this:

Currently the pve-container@<id>.service contains

ExecStop=/usr/share/lxc/pve-container-stop-wrapper %i

To my knowledge this is only executed in a non-failure case.

We could extend the .service file to contain ExecStopPost=

implement a new `pve-container-stop-post-wrapper`

Add it to the .service config:
ExecStopPost=/usr/share/lxc/pve-container-stop-post-wrapper %i

according to [1], we can use this field when "the service exited
unexpectedly".

Inside the new wrapper we call something like:
lxc-stop <id>
# other cleanup ?

Your patch
----------
Let me know some of the details on how you plan for for
pve-container@<id>.service to manage the lxc-console and the improved
dtach process.

I think this idea is really interesting and I would like to
support/review your patches on this matter.

References
----------
[1] https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html




^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-10-09  9:44 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz follow: Atom feed
-- links below jump to the message on this page --
2026-10-07  8:37 [PATCH container/manager 0/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07  8:37 ` [PATCH manager 1/2] fix #8093: pvestatd: remove cleanup of stale lxc consoles Elias Huhsovitz
2026-10-07  8:37 ` [PATCH container 2/2] fix #8093: console: bind console sessions to container lifetime via systemd scopes Elias Huhsovitz
2026-10-07 13:38 ` [PATCH container/manager 0/2] " Filip Schauer
2026-10-08 10:25   ` Elias Huhsovitz
2026-10-08 13:15     ` Filip Schauer
2026-10-09  9:44       ` Elias Huhsovitz

This is an external index of several public inboxes,
see mirroring instructions on how to clone and mirror
all data and code used by this external index.
Service provided by Proxmox Server Solutions GmbH | Privacy | Legal