---
title: "1,000 Clusters, One Engine, No Tuning - Calinora Blog"
description: "We pointed one proposal engine at 1,000 randomly generated Kafka clusters - three brokers to a thousand, a thousand partitions to a million, every rack and replication shape we could think of. No per-cluster tuning. Here is how little the answer moves when you move everything else."
url: "https://blog.calinora.io/posts/proposal-engine-at-scale/"
date: "2026-05-23"
---

May 23, 2026

# 1,000 Clusters, One Engine, No Tuning

Most rebalancer benchmarks pick one friendly cluster, hit run, and screenshot the result. We generated a thousand. Three-broker dev boxes and thousand-broker fleets. A thousand partitions and a million. Single-rack and five-rack, replication factor two through four, half-empty clusters waiting to be filled, clusters where the busiest 5% of partitions carry most of the load. Then we pointed the same proposal engine at every one of them. No per-cluster tuning. No constants nudged between runs. The same seven-metric objective, a thousand times.

This post is about one thing: how little the answer moves when you move everything else.

## Three things the sweep shows

- **It does the same thing at every scale.** From a three-broker box to a thousand-broker fleet, from a thousand partitions to a million, the engine runs the same pipeline with the same constants. Nothing is swapped in for the big clusters or the small ones.
- **The result barely moves when the input does.** Hand the engine a cluster sitting at 4pp of imbalance or one at 96pp; it lands both in the same narrow band near the target. Across the sweep the median cluster goes from 53pp of imbalance to 4.85pp, about a tenfold reduction.
- **It tells you when a cluster cannot be balanced.** A handful of brokers spread across three racks has a floor no rebalancer can cross. The engine pushes everything it can, then names the metric the topology is holding back rather than pretending it is done.

## The one picture

Each dot is one cluster. Left to right is how unbalanced it started. Bottom to top is where the engine left it. The point of the chart is the shape: wide spread sideways, tight band along the bottom.

Image: Pre-plan vs post-plan average variance across the full sweep of a thousand clusters. Starting imbalance ranges from 4 to 128 percentage points; post-plan variance collapses into a narrow band near the balance target regardless of where each cluster started.

Image text: 1,000 clusters: where each one started vs. where the engine left it; 5pp default target; 0; 5; 10; 15; 20; post-plan avg variance (pp); 0; 25; 50; 75; 100; 125; pre-plan avg variance (pp); average reached the cluster’s target; topology-limited

Each dot is one cluster. Starting imbalance spans 4pp to 128pp across the x-axis. Post-plan variance lands in a tight band along the bottom. Green dots above the 5pp line are clusters that deliberately ran a looser target, for example 10pp on a six-broker box, and met it.

The median cluster starts at 53pp of imbalance and ends at 4.85pp. Three quarters of them land their average under their own balance target. The handful sitting higher all have the same explanation, a structural floor the topology imposes, and the engine names the metric in each case. Move the input freely along the x-axis; the engine still puts you near the bottom of the y-axis.

## What we measure

The engine balances seven metrics at once on every cluster, grouped into three families.

**Placement:** leader count, follower count, and disk bytes per broker. **Producer load:** producer message rate and producer byte rate per broker. **Consumer load:** consumer message rate and consumer byte rate per broker.

Each metric is reported as per-broker variance normalized to the cluster mean, in percentage points. A value of 5pp means the worst-loaded broker sits about 5% off the cluster average for that dimension. The default balance target is 5pp: the engine keeps moving until every metric is below it, or it hits a structural limit and says so.

Cluster shape is real Kafka shape: brokers across racks, rack-aware placement, partition replica sets, per-partition disk size drawn from a log-normal distribution, traffic synthesized through six configurable profiles. The data is synthetic; the constraints are real.

## Stability under every axis

The full sweep randomizes everything at once. To see what each variable does on its own, we also swept one axis at a time, holding the rest at a 100-broker, 50,000-partition baseline. The story is the same in every direction: change the input, the output stays put.

Skew is the clearest case. As we crank the input skew, the starting imbalance climbs from 4pp to 96pp, a 24-fold spread. The plan lands in the same place every time.

Image: As input skew rises from 0 to 0.7, starting imbalance climbs from 4 to 96 percentage points while post-plan variance stays flat near 4 percentage points. Same 100-broker cluster, same engine, same constants.

Image text: Input imbalance climbs 24x. The result does not move.; 5pp target; 0; 25; 50; 75; 100; avg variance (pp); 0; 0.05; 0.1; 0.15; 0.2; 0.3; 0.4; 0.5; 0.6; 0.7; input skew parameter; pre-plan (96pp); post-plan (4.9pp)

The dashed line is each cluster’s starting imbalance; the solid green line is where the engine leaves it. At skew 0 the cluster is already balanced and the engine correctly does nothing, zero moves. From there the input rises 24x and the result holds flat at the target.

Every other axis behaves the same way:

| Axis swept | Range | Pre-plan imbalance | Post-plan band |
| - | - | - | - |
| Partition count | 1,000 -> 200,000 | 41-47pp | 3.9-4.8pp |
| Broker count | 12 -> 250 | 38-42pp | 3.8-6.8pp |
| Rack count | 1 -> 5 | 42pp | 3.8-4.0pp |
| Replication factor | 2 -> 4 | 42pp | 3.9-4.0pp |
| Topic count | 250 -> 16,666 | 42pp | 4.0pp |
| Empty brokers | 0 -> 50 | 42-100pp | 4.0-5.0pp |
| Traffic profile | 6 shapes | 41-43pp | 4.0-4.8pp |
| Skew | 0 -> 0.7 | 4-96pp | 3.5-4.9pp |

Partition count is worth sitting with: from a thousand partitions to two hundred thousand, a 200-fold range, the post-plan variance stays inside a single percentage point. Plan-generation time scales close to linearly with partition count, so the engine stays predictable as the cluster grows rather than falling off a cliff. The broker axis is the one with a wide band, and that is the structural floor at very small broker counts, covered below; from twelve brokers up it sits in the same 4-to-7pp range as everything else.

## It is deterministic

We ran the baseline cluster, 100 brokers and 50,000 partitions at replication factor 3, ten times from a cold start. All ten plans came back identical: the same 27,802 moves, the same 4.22 TB scheduled to the byte, the same 3.9928pp final variance. Same cluster in, same plan out, every time. You can diff two runs and get nothing back.

Determinism is not a nicety. It is what lets you review a plan, sleep on it, regenerate it the next morning, and know you are applying the thing you approved.

And the plan does not exaggerate. On every cluster in the sweep, the post-plan variance the proposal projects matches the variance you actually get after the moves, to the digit. The engine is not selling you a number it cannot deliver.

## Traffic shape is not the hard part

Rebalancer folklore says traffic skew is the thing that breaks you. We ran the same 50,000-partition cluster six different ways and got the same balance every time.

Image: Six traffic profiles run against the same 100-broker cluster. All six start near 42pp imbalance and land between 4.0 and 4.8pp, every metric under the 5pp target.

Image text: 0pp; 100pp variance; 5pp target; uniform; 41.7 -> 3.99pp; hot-skew; 41.0 -> 4.62pp; produce-heavy; 41.6 -> 4.16pp; consume-heavy; 41.6 -> 4.22pp; mostly-idle; 43.1 -> 4.78pp; mixed; 42.1 -> 4.55pp

Uniform, hot-skew, produce-heavy, consume-heavy, mostly-idle, mixed. Same cluster, same engine. All six start near 42pp and land under the 5pp target.

Structure does the heavy lifting: broker count, racks, replication factor, partition count. The traffic profile rides along. If you have been carrying a mental model that says “we cannot rebalance because our traffic is too skewed,” this is the evidence to retire it. The engine also respects `min.insync.replicas` on every move, and that constraint costs nothing in proposal quality.

## One cluster up close: half a million partitions, load piled on a sliver

The most strained cluster in the sweep: 150 brokers, 500,000 partitions, replication factor 3, and savagely skewed traffic, the shape that streams clickstream or ad-impression data, where the busiest 5% of partitions carry roughly 85% of the load. It starts at 69pp of average imbalance.

Image: 500,000-partition hot-skew cluster: all seven metrics fall from about 69pp to at or under the 5pp target.

Image text: 0pp; 100pp variance; 5pp balance target; Leader; 70.0 -> 4.94pp; Follower; 68.7 -> 4.57pp; Disk; 67.5 -> 5.00pp; Producer msg rate; 69.7 -> 4.30pp; Producer byte rate; 69.8 -> 4.56pp; Consumer msg rate; 69.6 -> 4.52pp; Consumer byte rate; 69.7 -> 4.63pp; AVERAGE; 69.3 -> 4.65pp

Faint gray bar: the pre-plan variance. Green bar: where the plan lands. All seven metrics fall from roughly 69pp to at or under the 5pp target.

The plan: 192,580 moves, 61 TB to relocate, and every one of the seven metrics at or under 5pp. Half a million partitions, the hardest traffic shape, and the result is the same flat band as the dev box.

## Capacity expansion does the right thing

Add empty brokers to a cluster and a cautious rebalancer either ignores them or churns the whole fleet for marginal gain. Across the sweep we added 5 to 50 empty brokers to a 100-broker cluster. Each time the starting imbalance climbs, fifty empty brokers reads as 100pp because half the fleet carries nothing, and each time the plan lands back near 4-to-5pp with every broker filled and leaders rotated across the loaded ones.

One of the largest in the sweep, 600,000 partitions across 200 brokers, goes from 41.9pp down to 4.26pp: 276,186 moves, 46 TB scheduled, the same flat landing as everything else.

## When the topology sets the limit

Not every cluster can be balanced, and the honest answer matters more than a clean screenshot. A five-broker cluster spread across two or three racks has a hard floor: with so few brokers, leader and follower placement is largely forced, and no sequence of moves gets every metric under 5pp.

One such cluster in the sweep starts at 99pp of imbalance. The engine pulls it down to 15pp, with producer and consumer rates landing around 10 to 12pp, and then reports follower and disk as topology-limited rather than spending moves it knows will not help. That is the difference between a rebalancer you can leave running and one you have to babysit.

Across the whole sweep, where a metric stays above target it is almost always a small cluster or a deliberately tightened target, and the typical miss is under 2pp over the line. The engine never blows up; it tells you exactly how close it got and what stopped it.

## What holds true at every scale

- One engine, one set of constants, a thousand configurations. No per-cluster tuning anywhere.
- Deterministic: same cluster in, same plan out, byte for byte.
- Rack-aware placement, replication factor, and replica co-location respected on every run. Zero rack violations.
- The plan matches reality: projected and achieved post-plan variance agree on every config we validated.
- Cluster expansion does the right thing. Add empty brokers; you get a plan that fills them.
- “Already balanced” means the engine genuinely cannot do better, and when that conclusion is structural it names the metric and why.

## The bottom line

A thousand clusters, one engine, no per-cluster tuning, and the answer barely moved. Whether a cluster walked in at 4pp of imbalance or 96pp, ran on six brokers or a thousand, carried even traffic or load piled onto a sliver of partitions, the engine left it in the same narrow band near the target. Where a topology set a limit, it named the metric and stopped, instead of churning the cluster for nothing.

These are synthetic clusters, built to push the engine into corners most real fleets never reach. The one that produced every plan above is the same engine that runs in production, and what it hands you carries exactly what you have seen here: a per-metric variance breakdown and a topology-limit note where one applies. Point it at your cluster and you get the same answer, every time.

## Try it on your own cluster

Pilot’s proposal generation, what-if simulation, and the large-scale tests above are all license-free. The engine is the same one that runs in production.

- [Quick start](https://docs.calinora.io/getting-started/quickstart): point it at your bootstrap servers, no broker changes required
- [Generate a proposal](https://docs.calinora.io/features/proposals): one API call, no commitment to apply
- [The seven-metric objective explained](https://docs.calinora.io/features/proposals#how-it-works): why these seven, why this threshold

The proposal arrives with the same fields you see in this post: a per-metric variance breakdown and a topology-limit annotation where one applies. What you see here is what comes back from the API.
