From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: from gate001.proxmox.com (gate001.proxmox.com [IPv6:2a0f:8001:1:32::40]) by lore.proxmox.com (Postfix) with ESMTPS id 036B51FF0AD for ; Sun, 20 Sep 2026 18:23:29 +0200 (CEST) Received: from gate001.proxmox.com (localhost.localdomain [127.0.0.1]) by gate001.proxmox.com (Proxmox) with ESMTP id 966E2214CC; Sun, 20 Sep 2026 18:23:24 +0200 (CEST) ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=rn3jLhizFTt/p/W0gkzmDOCkGrDvYxpi4iAYsEweFnnEmDPck+Qs52e1Ijm2wUJCKJC690Psudst3N85IHwOs6wALM0dcVV5oHBkZ36qDpbbL9ayvTJdk+Sr7xPf8JVYg/xhjwo6BCkLeR0p5wuVmhAU8SKurGMa+6/5vLbe1Q5G2j3f2Dv/+zSyltaWZXPzSiLWhSi8KW4KDQI7ekBtzP6ZNb2EVJu9/5iHhXo97zjC6c35oDZPAuljBVwha+hcloFDvqCe4LERDQ0DhrB+SI9ZfsF8ngDDnH6GF3mQybXf1hPQ/KRaDUPQEjJxHyIaZZaSmf0HDJvZmPHWvD8nrw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=TdxjNNs6U7HQIbKgKOcblWVtLSQDJlH0fNM0HBp6JtQ=; b=kMBRNCqZSUknaYtjYEmdVnXuzneHFaZCfcovqpxOgjDeLV0PEgsMZ0ptzWVzAIb5aVljMLxihvdmGJgFComXv1h5lEm5i4Oa9Fn2R3BfHqhNmZcaztkssMFq1odN7YJqgHAoYIFfe5PVlEJrAiyw4Pd/fexSqdSh0K8IFAxZGdLZHnG70KXtS/GdGLRPE6nm/0ijw5izhBUZ4BEOzRyQyJXTahnjz94HD1mlvspjiMmYAsHvVnM3R6oZpx9ZaX1Xq9wAkOPe3IsJ3jAFxxRLCzX0FMlyHiXsKb9frEBhIr1pzKk0uFTT5iPCpApp/zjfhPIo92DzbOku5Gll74KCqw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=ryomherold.dk; dmarc=pass action=none header.from=ryomherold.dk; dkim=pass header.d=ryomherold.dk; arc=none DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=RyomHerold.dk; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=TdxjNNs6U7HQIbKgKOcblWVtLSQDJlH0fNM0HBp6JtQ=; b=AvIh9R3BcRV1Cd8c/28mnCghEGRbTBUDUnnK6LGG3rCWxf2vEOtcURZWw55gsQPaSQu6OsZivdCN9VnqwfKpy2/wiclYYGy4snW4FGH4ZDmx2HVazo4DJ4LZd05kLLrSkOxgCevskplyE65T1U14M1/lJy5lUOOO25WHoTx8i2kw1FvPhBgePU5TYNjsIOJUaIAJ6SgpSI2qoaXfOWX9LyIHd6lwNAtq5T0whZO6k57dO/p3CulyFRI/fwhJshbWYTnHZL5LaGS3XXTIZHU3A1OfNbRUQXOyGp/oqu4ovvYPMK6nrLWTaYFRIhTCRBwPaTd0uVTSPxcTQR+kqctsww== Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=RyomHerold.dk; From: Michael Ryom To: pve-devel@lists.proxmox.com Subject: [PATCH ha-manager 6/7] usage: dynamic: smooth the unaccounted node load Date: Sun, 20 Sep 2026 18:22:13 +0200 Message-ID: <20260920162220.574802-7-Michael@RyomHerold.dk> X-Mailer: git-send-email 2.43.0 In-Reply-To: <20260920162220.574802-1-Michael@RyomHerold.dk> References: <20260920162220.574802-1-Michael@RyomHerold.dk> Content-Transfer-Encoding: 8bit Content-Type: text/plain X-ClientProxiedBy: TL2P290CA0008.ISRP290.PROD.OUTLOOK.COM (2603:1096:950:2::11) To AS8PR10MB7231.EURPRD10.PROD.OUTLOOK.COM (2603:10a6:20b:619::17) MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: AS8PR10MB7231:EE_|AM0PR10MB883391:EE_ X-MS-Office365-Filtering-Correlation-Id: bd2fb272-de10-43b8-ccc0-08df173375f3 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|23010399003|376014|1800799024|366016|10067099003|6133799003|22082099003|18002099003|56012099006; X-Microsoft-Antispam-Message-Info: emqeoJ4JD4KYGegDfQQD59j6dd90dJHFB9bXjU6jtxlWymE4JpfmPA4XqIEixcQ7deBacx1BoqoLMzn3+rZIhVKkK5IrIhmZ66wCvnUsUPn/n75hTIcQG8CHGOfYRpdXQp1ddHq+pQiZKi1BNMyPzw4cROQ8b9eldMTbFGVBeli8lGT9tmC0KWOz9TyHqtHwqx2gupaY5cL6jKRSDGzmCDjdwj0kKu9JTSQalcHHm4X6qXRsjs5weAw9onb46HBCtzBFCD4mKmhFvppXsjLoY9r5T0y7mOs2306kwNicAvVAeBHzMidfTZKGrLxHmT7T6EbRmF6hCw7XhFoi+2YqqXlP7j9MpBfUYv02pgR0QHFFwLrvvctstyB7520nP1Tp2EF7T6BsU4XdUC2RyT52ty9mZ6NHFkjetPLIh+38R5utJWrrwosBFoUJeW27biyH17Q4uVFFLV/0SB4i+hkk51YpX3m9WjL3ThZSJk4JhC4/OcXn16ULYIUPrbZ+u58CZZ3gIssOxRpG2/Slayd4Mb5ZhU401KjLuIo/4q9lgXLX1BBsc/A/8dGkTp9z+hDcIj5Raq0rLuyXV+ryF8+kxz8q0z2CFs7hdrb6pVMDQk4fB8KzrIp2nuQ7eNbf71bLLIBxCpTrF+obesa2rDhabDds74xhrQfT6YCZ0QZidmc= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:AS8PR10MB7231.EURPRD10.PROD.OUTLOOK.COM;PTR:;CAT:NONE;SFS:(13230040)(23010399003)(376014)(1800799024)(366016)(10067099003)(6133799003)(22082099003)(18002099003)(56012099006);DIR:OUT;SFP:1102; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?us-ascii?Q?YXI2ls729GbPD1sfy7A2QbSbYdR/FvZ7e/ahntQrX/wLE+5rNX3pw2dxx337?= =?us-ascii?Q?dAFswSE3rxu51gmjQJuA0ck7wqU2UvZ8pVrLyJ0OJapZKjD6ovLNh30KumiL?= =?us-ascii?Q?AOxvAv03KF4SCffr/dHfNEy966keuycVv8PeesriD22c+nPZe3lo8FS7zWgU?= =?us-ascii?Q?gYcsLdy0D3vW1fpieayvTmYYJLQJsU4KeBkMkoe0SRv0SjsfjjbbALEkXvhb?= =?us-ascii?Q?OjcgiJrrl1yZKOYSEfWBVt7UiOvbIqkmOEj0BhGyJDLk0fbSM9+qc6Ipgxch?= =?us-ascii?Q?5ZRp6rj57sjZuExyUWnYVxPniGpx1VQvQEq2DmANSF13Loomiek3Su/kLFzO?= =?us-ascii?Q?OQmeY6RjE90b3uzChVh5nLSiyJBTfoW6Uko5C6OuBuUmwRxj90sV04szrJBW?= =?us-ascii?Q?qPSYxW8WxOtfgCvST8vkiKTZJ0gir+gADMbs6PWuFOflgmLAgt9iUpN8oaj/?= =?us-ascii?Q?SasetBBey+DaMtey8uHgV2CWB6+T5sQPBMgnWg2MG9x9b/90H745HyYyoLwd?= =?us-ascii?Q?m0Oyh9oxD8wKYfQHsqnn4qQR3KsRh7c2z4YkExl501G79Xo1aeX6ZszgkFOY?= =?us-ascii?Q?37p1nHN7dd+PKW0/neXGg+H6b20zbDaLkLWjVqTChhoM6FB8nCJbwgV4kcRG?= =?us-ascii?Q?CmtpRBvlnNGtmEfTy0O724MrYkC6iTupdEdi//rHAM3EMh49gehxPgWjOuUn?= =?us-ascii?Q?P9tccx/V5iN83pC33VFwUtZLbwcnr2ch7Tnxx3gZU0Z/p/1IhunvIaextZEq?= =?us-ascii?Q?lGDbrXCzkyG9mrSIQ94JE2E5FqfTbNnePx6KEHyWA0BvMxupbLnkU8hEx4n7?= =?us-ascii?Q?lSVmksC5JDmJfctZF7RSY71Rm6xijpQVhBU2qM1SR5nuE0ft5hMhCyTq4DHo?= =?us-ascii?Q?nuO4cIpNcvzCysyUafQdU822NjPoP8sc8lCFbfBBg3VMtApGf+7A3p8oLIto?= =?us-ascii?Q?Etanum3dxvp39vsF5w+DWiqlgo1Tyg26R0NViIFLgG6HFDj7GYOLrDpT/9rp?= =?us-ascii?Q?Vo8snjXIIZ+bQnDauUzEwOeov9bC+f+AblFG6wChHXUGPOvSwyw9bNIturOY?= =?us-ascii?Q?zQ3/TEn3CgMcHuRWW/w0OS9ZD+MwDOm8Ekpv9ioP4ElG/xpEvz66NcOBjBXK?= =?us-ascii?Q?bHF+FijDWYEbgFOZuaXAPoPlB73q0ASA8IjGilnTFI6p/zSakdV+aiVaNoT3?= =?us-ascii?Q?SmkGjlRP2QZn7oioYVUJiL6Ob0rPYsqvct5h8afM+ySJcPscfxiZKiXQSNma?= =?us-ascii?Q?7JoKkyj3Y2FBdP2GgPoZn+NkJNTFpjmExinRqby+bNQzRhTb4I/XuW4a6yQE?= =?us-ascii?Q?mp5mvjFlwVJ1PVhHiHFzjEam9NgiEAwA3a92/09xxxZTKEoaJf8nrS6NuZ+h?= =?us-ascii?Q?ppRFrXpArFB3l34YZiCTxZnTgzc8R94Olo8uHzlOUA4FxjvbQBE6IxAubHOx?= =?us-ascii?Q?NyCcdrsXkgtU5mCAQlHSErMOyduDBMNtaI+e/iTMk8m0oA/DIqhfAdwGzHcT?= =?us-ascii?Q?1HQiZLv2n6Avs+vpNoSf7kbdEVmaqrX6HzYAAxhI7oEFlaYZ1c/qiMaiF0qC?= =?us-ascii?Q?brkeaglLdk/rT1Si/1uc3q9mz1D4QumgHQxJ7IKM4CRIEhzn4OzOK9S5aaTa?= =?us-ascii?Q?viJFRTNmMSUuerzYVyj1I/8y2BvEJZoW5WBc9rNBCajSaDK4oBZp/mnQ5Zp5?= =?us-ascii?Q?7QuCh9WQkK+QRG9Z80EorYIEHzUOZKzvjX1ujf3VWSpjxSQAPJZ3KlJn/gs3?= =?us-ascii?Q?rwQZTCNy8w=3D=3D?= X-OriginatorOrg: RyomHerold.dk X-MS-Exchange-CrossTenant-Network-Message-Id: bd2fb272-de10-43b8-ccc0-08df173375f3 X-MS-Exchange-CrossTenant-AuthSource: AS8PR10MB7231.EURPRD10.PROD.OUTLOOK.COM X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 20 Sep 2026 16:23:10.1491 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: 26a2499e-4646-4888-a695-c64736a56807 X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: Ch9BZrIK4YPPiP4WoBwHcgCZVlRAeqj6pHjHHdHDOyi1KETjr54QfMlMjo0LYH5LTJvdSKbLhOY9Za1ZVPMjzg== X-MS-Exchange-Transport-CrossTenantHeadersStamped: AM0PR10MB883391 X-SPAM-LEVEL: Spam detection results: 0 DKIM_SIGNED 0.1 Message has a DKIM or DK signature, not necessarily valid DKIM_VALID -0.1 Message has at least one valid DKIM or DK signature DKIM_VALID_AU -0.1 Message has a valid DKIM or DK signature from author's domain DKIM_VALID_EF -0.1 Message has a valid DKIM or DK signature from envelope-from domain DMARC_PASS -0.1 DMARC pass policy RCVD_IN_DNSWL_NONE -0.0001 Sender listed at https://www.dnswl.org/, no trust SPF_HELO_PASS -0.001 SPF: HELO matches SPF record SPF_PASS -0.001 SPF: sender matches SPF record Message-ID-Hash: IUW6A3BSMRQFJMP5TQKCQ63ZZUGAAG6O X-Message-ID-Hash: IUW6A3BSMRQFJMP5TQKCQ63ZZUGAAG6O X-MailFrom: Michael@RyomHerold.dk X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header CC: Michael Ryom X-Mailman-Version: 3.3.10 Precedence: list List-Id: Proxmox VE development discussion List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: The dynamic scheduler acts on point-in-time samples of the node and guest usage stats. The part of the node load that is not caused by any HA-managed service (non-HA guests, the ZFS ARC, Ceph daemons, the IO of a rebalance migration itself, sampling skew between node and guest stats) is noisy, and acting on a single sample makes the automatic load balancer prone to oscillation: a single dominant HA resource on a small cluster qualifies for a migration (and afterwards, for its reverse) as soon as the node base loads fluctuate around each other by a fraction of the resource's own load. Decompose each node's load into the usage of the HA-managed services actually running on it plus a residual, and smooth only the residual with an exponentially moving average that the manager keeps across scheduling rounds (not persisted for a CRM failover, like the other balancer state). The service usages themselves are deliberately not smoothed: they are fully modeled by the scheduler, so known structural changes (a service was migrated, started or stopped, or its measured load changed) take effect immediately, transient service load spikes remain handled by the hold duration, and the expected log of no existing regression test changes. The residual, on the other hand, cannot be attributed to anything the balancer could act on, and its expected reaction to a rebalance motion is zero, so a lagging average is safe to act on. Add a regression test where the node base loads of a two-node cluster fluctuate around each other by half a percentage point under a dominant HA resource. Without the smoothing, every fluctuation qualifies a migration of the dominant resource to the other node, moving it back and forth indefinitely (with the per-resource cooldown, once its cooldown has expired); with the smoothing, the effective base load difference stays well below the qualification point and no motion is issued. Signed-off-by: Michael Ryom --- src/PVE/HA/Manager.pm | 10 ++- src/PVE/HA/Usage/Dynamic.pm | 73 ++++++++++++++++++- .../test-crs-dynamic-auto-rebalance8/README | 19 +++++ .../test-crs-dynamic-auto-rebalance8/cmdlist | 19 +++++ .../datacenter.cfg | 6 ++ .../dynamic_service_stats | 3 + .../hardware_status | 4 + .../log.expect | 28 +++++++ .../manager_status | 1 + .../service_config | 3 + .../static_service_stats | 3 + 11 files changed, 167 insertions(+), 2 deletions(-) create mode 100644 src/test/test-crs-dynamic-auto-rebalance8/README create mode 100644 src/test/test-crs-dynamic-auto-rebalance8/cmdlist create mode 100644 src/test/test-crs-dynamic-auto-rebalance8/datacenter.cfg create mode 100644 src/test/test-crs-dynamic-auto-rebalance8/dynamic_service_stats create mode 100644 src/test/test-crs-dynamic-auto-rebalance8/hardware_status create mode 100644 src/test/test-crs-dynamic-auto-rebalance8/log.expect create mode 100644 src/test/test-crs-dynamic-auto-rebalance8/manager_status create mode 100644 src/test/test-crs-dynamic-auto-rebalance8/service_config create mode 100644 src/test/test-crs-dynamic-auto-rebalance8/static_service_stats diff --git a/src/PVE/HA/Manager.pm b/src/PVE/HA/Manager.pm index 2c4c273..6e1827f 100644 --- a/src/PVE/HA/Manager.pm +++ b/src/PVE/HA/Manager.pm @@ -81,6 +81,10 @@ sub new { failures => {}, # "$sid:$target" => { count => $count, until => $time } cooldown => {}, # $sid => $time }, + # exponentially moving average of the unaccounted node load, kept + # across rounds by PVE::HA::Usage::Dynamic; also not persisted for a + # CRM failover + dynamic_stats_smoothing_state => {}, group_migration_round => 3, # wait a little bit }, $class; @@ -588,7 +592,11 @@ sub recompute_online_node_usage { if ($have_dynamic_scheduling) { $online_node_usage = eval { $service_stats = $haenv->get_dynamic_service_stats(); - my $scheduler = PVE::HA::Usage::Dynamic->new($haenv, $service_stats); + my $scheduler = PVE::HA::Usage::Dynamic->new( + $haenv, + $service_stats, + $self->{dynamic_stats_smoothing_state}, + ); $scheduler->add_node($_) for $online_nodes->@*; return $scheduler; }; diff --git a/src/PVE/HA/Usage/Dynamic.pm b/src/PVE/HA/Usage/Dynamic.pm index 76d0fea..fc88ddf 100644 --- a/src/PVE/HA/Usage/Dynamic.pm +++ b/src/PVE/HA/Usage/Dynamic.pm @@ -8,12 +8,83 @@ use PVE::RS::ResourceScheduling::Dynamic; use base qw(PVE::HA::Usage); +# weight of the newest sample in the exponentially moving average of the +# unaccounted node load; with one sample per CRM scheduling round (~10s), +# 0.25 averages over roughly the last minute +my $usage_smoothing_alpha = 0.25; + +# Smooths the *unaccounted* part of each node's load with an exponentially +# moving average kept in $state across invocations. +# +# The node load is decomposed into the usage of the HA-managed services +# running on it plus a residual, and only the residual is smoothed. The +# residual captures the load that cannot be attributed to any HA-managed +# service (non-HA guests, ZFS ARC, Ceph daemons, the IO of a rebalance +# migration itself, sampling skew between the node and guest stats). Those +# point-in-time samples are noisy, and acting on a single sample makes the +# load balancer prone to oscillation, while their expected reaction to a +# rebalance motion is zero, so a lagging average is safe to act on. +# +# The service usages themselves are deliberately not smoothed: they are +# fully modeled by the scheduler, i.e. known structural changes (a service +# was migrated, started or stopped, or its load changed) take effect +# immediately, and transient service load spikes are already handled by the +# hold duration of the load balancer. +my $smooth_unaccounted_node_load = sub { + my ($state, $node_stats, $service_stats) = @_; + + # usage per node of the HA-managed services actually running on it + my $accounted = {}; + for my $sid (keys %$service_stats) { + my ($node, $running, $usage) = $service_stats->{$sid}->@{qw(node running usage)}; + next if !$running || !defined($node) || !$usage; + next if !defined($usage->{cpu}) || !defined($usage->{mem}); + + $accounted->{$node}->{cpu} += $usage->{cpu}; + $accounted->{$node}->{mem} += $usage->{mem}; + } + + for my $node (sort keys %$node_stats) { + my $stats = $node_stats->{$node}; + next if !defined($stats->{cpu}) || !defined($stats->{mem}); + + # the residual may become negative if the samples disagree, e.g. + # right after a service stopped; only the reconstructed node load is + # clamped, so that the residual average is not biased upwards + my $residual_sample = { + cpu => $stats->{cpu} - ($accounted->{$node}->{cpu} // 0.0), + mem => $stats->{mem} - ($accounted->{$node}->{mem} // 0), + }; + + my $residual = $state->{$node}; + if (!defined($residual)) { + $residual = $state->{$node} = $residual_sample; + } else { + for my $field (qw(cpu mem)) { + $residual->{$field} = $usage_smoothing_alpha * $residual_sample->{$field} + + (1 - $usage_smoothing_alpha) * $residual->{$field}; + } + } + + my $cpu = $residual->{cpu} + ($accounted->{$node}->{cpu} // 0.0); + my $mem = $residual->{mem} + ($accounted->{$node}->{mem} // 0); + $stats->{cpu} = $cpu > 0.0 ? $cpu : 0.0; + $stats->{mem} = $mem > 0 ? int($mem) : 0; + } + + # drop state of meanwhile removed nodes + delete $state->{$_} for grep { !$node_stats->{$_} } keys %$state; +}; + sub new { - my ($class, $haenv, $service_stats) = @_; + my ($class, $haenv, $service_stats, $smoothing_state) = @_; my $node_stats = eval { $haenv->get_dynamic_node_stats() }; die "did not get dynamic node usage information - $@" if $@; + $smooth_unaccounted_node_load->($smoothing_state, $node_stats, $service_stats) + if defined($smoothing_state); + my $scheduler = eval { PVE::RS::ResourceScheduling::Dynamic->new() }; die "unable to initialize dynamic scheduling - $@" if $@; diff --git a/src/test/test-crs-dynamic-auto-rebalance8/README b/src/test/test-crs-dynamic-auto-rebalance8/README new file mode 100644 index 0000000..0264b58 --- /dev/null +++ b/src/test/test-crs-dynamic-auto-rebalance8/README @@ -0,0 +1,19 @@ +Test that smoothing the unaccounted node load prevents a single dominant HA +resource from oscillating between two nodes whose base loads (i.e. load not +caused by any HA-managed service, here set via the node base load of the +simulated hardware) fluctuate around each other. + +The cluster has two nodes with an equal base load and one dominant movable +HA resource vm:100 on node1. The imbalance is permanently above the trigger +threshold, but no motion can improve it, so the load balancer stays armed +without acting. Then the node base loads are repeatedly flipped by about half +a percentage point of node load in alternating directions. + +Each raw sample after a flip would qualify a migration of vm:100 to the less +loaded node (and after the next flip, back again): the expected relative +imbalance improvement of about 10.4% exceeds the 10% margin and the absolute +improvement of about 5.5 percentage points exceeds the required minimum of 5. +Acting on the raw samples would move the dominant resource back and forth +indefinitely. With the exponentially moving average over the unaccounted +node load, the effective base load difference stays well below the point +where a motion qualifies, so no motion may be issued at any point. diff --git a/src/test/test-crs-dynamic-auto-rebalance8/cmdlist b/src/test/test-crs-dynamic-auto-rebalance8/cmdlist new file mode 100644 index 0000000..eeee26c --- /dev/null +++ b/src/test/test-crs-dynamic-auto-rebalance8/cmdlist @@ -0,0 +1,19 @@ +[ + [ "power node1 on", "power node2 on" ], + [ + "node node1 set-dynamic-stats cpu 5.064 mem 0", + "node node2 set-dynamic-stats cpu 4.536 mem 0" + ], + [ + "node node1 set-dynamic-stats cpu 4.536 mem 0", + "node node2 set-dynamic-stats cpu 5.064 mem 0" + ], + [ + "node node1 set-dynamic-stats cpu 5.064 mem 0", + "node node2 set-dynamic-stats cpu 4.536 mem 0" + ], + [ + "node node1 set-dynamic-stats cpu 4.536 mem 0", + "node node2 set-dynamic-stats cpu 5.064 mem 0" + ] +] diff --git a/src/test/test-crs-dynamic-auto-rebalance8/datacenter.cfg b/src/test/test-crs-dynamic-auto-rebalance8/datacenter.cfg new file mode 100644 index 0000000..01c8114 --- /dev/null +++ b/src/test/test-crs-dynamic-auto-rebalance8/datacenter.cfg @@ -0,0 +1,6 @@ +{ + "crs": { + "ha": "dynamic", + "ha-auto-rebalance": 1 + } +} diff --git a/src/test/test-crs-dynamic-auto-rebalance8/dynamic_service_stats b/src/test/test-crs-dynamic-auto-rebalance8/dynamic_service_stats new file mode 100644 index 0000000..ef14918 --- /dev/null +++ b/src/test/test-crs-dynamic-auto-rebalance8/dynamic_service_stats @@ -0,0 +1,3 @@ +{ + "vm:100": { "cpu": 9.6, "mem": 0 } +} diff --git a/src/test/test-crs-dynamic-auto-rebalance8/hardware_status b/src/test/test-crs-dynamic-auto-rebalance8/hardware_status new file mode 100644 index 0000000..bb0cf81 --- /dev/null +++ b/src/test/test-crs-dynamic-auto-rebalance8/hardware_status @@ -0,0 +1,4 @@ +{ + "node1": { "power": "off", "network": "off", "maxcpu": 24, "maxmem": 34359738368, "cpu": 4.8, "mem": 0 }, + "node2": { "power": "off", "network": "off", "maxcpu": 24, "maxmem": 34359738368, "cpu": 4.8, "mem": 0 } +} diff --git a/src/test/test-crs-dynamic-auto-rebalance8/log.expect b/src/test/test-crs-dynamic-auto-rebalance8/log.expect new file mode 100644 index 0000000..73e381a --- /dev/null +++ b/src/test/test-crs-dynamic-auto-rebalance8/log.expect @@ -0,0 +1,28 @@ +info 0 hardware: starting simulation +info 20 cmdlist: execute power node1 on +info 20 node1/crm: status change startup => wait_for_quorum +info 20 node1/lrm: status change startup => wait_for_agent_lock +info 20 cmdlist: execute power node2 on +info 20 node2/crm: status change startup => wait_for_quorum +info 20 node2/lrm: status change startup => wait_for_agent_lock +info 20 node1/crm: got lock 'ha_manager_lock' +info 20 node1/crm: status change wait_for_quorum => master +info 20 node1/crm: using scheduler mode 'dynamic' +info 20 node1/crm: node 'node1': state changed from 'unknown' => 'online' +info 20 node1/crm: node 'node2': state changed from 'unknown' => 'online' +info 20 node1/crm: adding new service 'vm:100' on node 'node1' +info 20 node1/crm: service 'vm:100': state changed from 'request_start' to 'started' (node = node1) +info 21 node1/lrm: got lock 'ha_agent_node1_lock' +info 21 node1/lrm: status change wait_for_agent_lock => active +info 21 node1/lrm: starting service vm:100 +info 21 node1/lrm: service status vm:100 started +info 22 node2/crm: status change wait_for_quorum => slave +info 120 cmdlist: execute node node1 set-dynamic-stats cpu 5.064 mem 0 +info 120 cmdlist: execute node node2 set-dynamic-stats cpu 4.536 mem 0 +info 220 cmdlist: execute node node1 set-dynamic-stats cpu 4.536 mem 0 +info 220 cmdlist: execute node node2 set-dynamic-stats cpu 5.064 mem 0 +info 320 cmdlist: execute node node1 set-dynamic-stats cpu 5.064 mem 0 +info 320 cmdlist: execute node node2 set-dynamic-stats cpu 4.536 mem 0 +info 420 cmdlist: execute node node1 set-dynamic-stats cpu 4.536 mem 0 +info 420 cmdlist: execute node node2 set-dynamic-stats cpu 5.064 mem 0 +info 1020 hardware: exit simulation - done diff --git a/src/test/test-crs-dynamic-auto-rebalance8/manager_status b/src/test/test-crs-dynamic-auto-rebalance8/manager_status new file mode 100644 index 0000000..0967ef4 --- /dev/null +++ b/src/test/test-crs-dynamic-auto-rebalance8/manager_status @@ -0,0 +1 @@ +{} diff --git a/src/test/test-crs-dynamic-auto-rebalance8/service_config b/src/test/test-crs-dynamic-auto-rebalance8/service_config new file mode 100644 index 0000000..60688cf --- /dev/null +++ b/src/test/test-crs-dynamic-auto-rebalance8/service_config @@ -0,0 +1,3 @@ +{ + "vm:100": { "node": "node1", "state": "started" } +} diff --git a/src/test/test-crs-dynamic-auto-rebalance8/static_service_stats b/src/test/test-crs-dynamic-auto-rebalance8/static_service_stats new file mode 100644 index 0000000..7eee778 --- /dev/null +++ b/src/test/test-crs-dynamic-auto-rebalance8/static_service_stats @@ -0,0 +1,3 @@ +{ + "vm:100": { "maxcpu": 12.0, "maxmem": 8589934592 } +} -- 2.47.3