From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id 205321FF0B1 for ; Fri, 09 Oct 2026 04:04:20 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 53CDC21320; Fri, 09 Oct 2026 04:04:19 +0200 (CEST) Message-ID: <1f05d528-4fb9-49c2-bd73-11a3759665cd@proxmox.com> Date: Fri, 9 Oct 2026 04:04:11 +0200 MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Beta Subject: Re: [PATCH proxmox{,-backup} 00/28] append-only sync jobs and snapshot retention timespan To: Christian Ebner , pbs-devel@lists.proxmox.com References: <20260813171002.809441-1-c.ebner@proxmox.com> Content-Language: en-US, de-DE From: Thomas Lamprecht In-Reply-To: <20260813171002.809441-1-c.ebner@proxmox.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-Bm-Milter-Handled: 55990f41-d878-4baa-be0a-ee34c49e34d2 X-Bm-Transport-Timestamp: 1791511452811 X-SPAM-LEVEL: Spam detection results: 0 AWL 0.596 Adjusted score from AWL reputation of From: address DMARC_MISSING 0.1 Missing DMARC policy KAM_DMARC_STATUS 0.01 Test Rule for DKIM or SPF Failure with Strict Alignment (newer systems) RCVD_IN_DNSWL_MED -2.3 Sender listed at https://www.dnswl.org/, medium trust SPF_HELO_NONE 0.001 SPF: HELO does not publish an SPF Record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: MPWFEUAOIMFJFU6EIJHKLLQT54T2FT5R X-Message-ID-Hash: MPWFEUAOIMFJFU6EIJHKLLQT54T2FT5R X-MailFrom: t.lamprecht@proxmox.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox Backup Server development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On 13/08/2026 19:10, Christian Ebner wrote: > Currently sync jobs cannot be configured to be fully append-only > since namespace creation requires datastore modify privileges to do > so. Further, permissions would also allow to restore or modify owned > content. > > This patch series therefore extends the current permissions > and roles to allow for append only sync jobs, by only allowing > the minimally required permissions and roles. > > In particular, for push the sync jobs local user on the source > requires RemoteSyncAppendOperator as well as DatastoreReader on the > source datastore, with DatastoreAppend and DatastoreAudit (latter for > listing privs of pre-existing contents without restore) permissions > for the user on the remote instance used for connection. > > For pull, the user on the target must be able to append to the > datastore via DatastoreAppend and able to read from the remote source > by the respective RemoteSyncOperator permissions on the remote and > by either DatastoreBackup or DatastoreReader permissions to access the > contents. > > Further, sync jobs are extended to allow setting a retention timespan > for which synced snapshots cannot be pruned, neither by the sync job, > nor by prune jobs. Only root@pam is allowed to change the retention > period. After the retention period, snapshots behave like regular > snapshots again and can be pruned. > > To protect from sync jobs setting unintended retention timespans, > it is now also possible to configure a maximum allowed reteniton time > on the datastore. > > Sending this as RFC for some initial feedback on the overall > implementation approach, plan to further have a look into object > locking and retention on s3 object stores [0] and changes required > for immutable storage [1]. > > [0] https://bugzilla.proxmox.com/show_bug.cgi?id=6780 > [1] https://bugzilla.proxmox.com/show_bug.cgi?id=4293 Thanks for the series, I like the overall direction. The append-only privileges combined with an absolute per-snapshot retain-until is a good base. It shouldn't block adding support for S3 Object Lock later, but only if we pin the semantics down now, so that a storage backend can later enforce exactly what PBS enforces today. So what I'd potentially do for a v2 - albeit you're probably much better here w.r.t. checking if these proposals have any merit: 1. Keep retain-until an absolute timestamp. Robert's idea of renewing the lock on every sync is good, but it should extend the stored timestamp rather than count from the file's mtime. A lock relative to mtime is nothing a storage backend can enforce. S3 Object Lock stores exactly that too, i.e. a retain-until date per object, set on upload or later via `PutObjectRetention`, so an absolute retain-until could be passed through as is. 2. Only allow extending a lock. The root endpoint from patch 27 can currently shorten or clear it. That makes it weaker than what S3 compliance mode would later give us, so I'd start with extend-only. 3. Always bound the lock. Without a configured max, any user with backup privileges can lock data for decades, and any Datastore.Modify user can remove the max again. I'd only accept retain-until once a max is configured, and enforce it on every path a value can come in. Today that misses pull without a timespan (which copies the remote's value as is), tape restore and s3-refresh. 4. Respect the lock on every delete path. Today only BackupDir::destroy checks it. These still remove retained snapshots, FWICT: - destroying a datastore with its data; - deleting a namespace recursively, where the S3 bulk delete runs before any per-group check; - push remove-vanished against an S3 remote; - resync-corrupt, which replaces the manifest and with it the retain-until. 5. Skip retained snapshots instead of failing. Prune and group deletion should treat them like protected snapshots. For prune, I'd compute the keep marks as now and show retained ones as "kept (retained)", so the lock doesn't use up keep-* slots, would avoid quite a bit of noise, e.g. as a PVE jobs can prune after every backup. With 1. to 3. we should be future proof for S3, 4. and 5. is not required for that per se, but should also fit well into how S3 works anyway. Potential issues in current implementation here: - Pull seems to rewrite the whole manifest. That would explain losing the file list that Robert noticed, and it probably also invalidates the signature. Only the unprotected part should change. - Repeated pulls look like they stop the whole group. From reading the code (not tested), it looks like from the second run on, the newest snapshot, which pull always re-syncs, trips the new check, and the `result?` aborts the rest of the group until the lock expires. Some other points to check out, even if relevant they still might be follow-ups though: - A server-side default lock per datastore, maybe per namespace, applied to every new snapshot when the backup finishes, independent of what the client sends. PVE doesn't pass a retain-until today, and a compromised PVE host wouldn't anyway, so this is what would actually protect PVE backups. A client-provided value could then only extend the lock, up to the max. If PVE later probably should get a per-job setting, to trigger an extend-only call after the backup, like vzdump already does for e.g. setting the notes and the protected flag. That would avoid threading a new parameter through QEMU and proxmox-backup-qemu lib, and works for VMs and containers alike. - Splitting Datastore.Modify into finer privileges, as Robert suggested, for example for creating namespaces, creating snapshots, and modifying or removing content. That would let us express append-only roles directly instead of special-casing the permission checks. It needs a backward-compatible mapping for existing ACLs and roles, so a separate series seems better suited (but IMO worth it, Datastore.Modify is as is rather too powerful anyway). We could also keep Datastore.Modify and just add newer split up privs and roles and allow either or, that should allow us to introduce it backward compatible without any old-to-new mapping code. - download_previous serves any file of the previous snapshot by name, so Append-users can also read blobs like the guest config or the client log. Incremental backups only need the index files and the manifest, so either restrict Append users to those, or document reading the rest as intended. With these addressed, I'd be fine with dropping the RFC tag for v2 (as always, nice work!)