WebsiteDocs
← Back to Blog

Replacing a Kafka Broker Without Under-Replicated Partitions

A lot of broker maintenance doesn’t need a drain. If you have RF=3, healthy ISR, and the broker comes back quickly with its data intact - a kernel patch, a service restart, a short OS update - you can just shut it down. Kafka shifts leadership to followers, the cluster keeps serving, and the broker catches up when it returns. Pilot’s proposal engine will smooth out any leadership skew afterwards. Don’t over-engineer the easy case.

Draining is for the maintenance cases where shutting down isn’t safe:

In these cases the broker comes back empty, or doesn’t come back at all. The replacement booting empty isn’t what hurts - that’s true either way. What hurts is the window between. Skip the drain and the moment broker-5 disappears, its replicas drop to RF-1 and stay there for hours until the new broker boots and catches up: the whole maintenance window runs under-replicated, one bad event away from crossing min.ISR. Drain first and those replicas already live on the remaining brokers before broker-5 goes offline. The cluster sits at full replication factor through the entire window. Once the new broker is online, Pilot’s proposal engine takes over: it rebalances the cluster onto the empty broker in a coordinated way, picking moves that improve balance without disrupting traffic. No manual planning, no rush.

Pilot turns the drain into one API call. The maintenance engine plans the move, applies it, and runs the reassignments in the background.

The Scenario

We’re migrating broker-5 to a new instance type, so the existing broker won’t come back - a fresh one will take its place with empty disks. Before broker-5 goes offline, it has to be empty: no leaders, no follower replicas, nothing. The cluster has five brokers across three racks. We want the drain to keep rack diversity, avoid leadership churn, and finish in a predictable amount of time.

Click through to see how Pilot handles it. Each step is one API call.

pilot - drain broker-5
1 / 5
Step 1

Preview what would happen

Ask Pilot to plan the drain without applying it. `preview: true` returns the full move plan so you can sanity-check it first.

POSThttp://pilot:8080/api/v1/brokers/5/maintenance
-d {"preview": true}
HTTP 200 OK
success:true
preview:true
brokerId:5
totalPartitions:47
affectedTopics:payments-events, orders, audit-events, user-events

The drain itself is one POST. Sending preview: true first is optional - same endpoint, just shows you the plan without applying it. A polling GET confirms when broker-5 is empty, and a DELETE brings it back into rotation after the maintenance is done. No planning by hand, no waiting on a rebuild after the broker rejoins, no guessing whether broker-5 is really empty.

What Just Happened?

Preview first, then commit

POST /api/v1/brokers/{id}/maintenance accepts a small body:

{ "preview": true }

With preview: true, Pilot computes the full move plan and returns it without touching anything. You see how many partitions would move, which topics are affected, and whether the plan keeps rack diversity. If something looks off (a topic that shouldn’t move, a target broker that’s also scheduled for maintenance), you stop and adjust.

When the plan looks right, send the same request with preview: false. Pilot marks broker-5 as in maintenance, registers the new replica assignments with Kafka, and starts the reassignments in the background. For every replica that lived on broker-5, the engine picks a new home on one of the remaining brokers - replication factor stays the same. The response comes back immediately; the actual data movement runs without you having to babysit it.

By default Pilot uses rack-aware placement, which is what you want for production drains: every partition’s replicas land in different racks, so a rack-level failure later doesn’t cost you a partition. There’s a strategy field if you ever need to override it, but you almost never do.

Leadership moves are unavoidable during a drain - any partition that broker-5 currently leads has to elect a new leader before broker-5 can go offline. The maintenance engine handles this carefully: for those partitions, it picks the new leader from the existing ISR, so promotion is instant and producers see at most a brief reconnection. Partitions where broker-5 is only a follower keep their current leader untouched. There’s no flag to set for any of this - it’s how the engine plans the drain.

Poll for “broker is empty”

Once the drain is running, the question “is broker-5 safe to shut down yet?” has a clean answer: ask Pilot for partition counts per broker.

GET /api/v1/cluster/logdirs?summary=true returns one row per broker with partitionCount and totalSize. Watch broker-5’s row drop. When it reads 0 partitions, 0 GB, the broker holds no data and no leadership - you can shut it down.

This polling pattern works for any drain. You don’t need to track reassignment IDs, parse progress percentages, or trust an ETA. The state is observable directly: how many partitions are still on the broker. When it’s zero, you’re done.

Bring it back when ready

Once the new broker-5 is online and registered with the cluster, one call removes it from maintenance:

curl -sX DELETE $PILOT/api/v1/brokers/5/maintenance

Broker-5 is now eligible to receive new partitions again. Pilot’s continuous proposal engine will naturally start placing partitions there over the next hours as the cluster rebalances. If you want to speed that up, you can apply a fresh proposal explicitly - but for a routine drain it’s usually fine to let the cluster settle on its own.

The API at a Glance

The whole workflow:

# 1. Preview the drain curl -sX POST $PILOT/api/v1/brokers/5/maintenance \ -d '{"preview": true}' # 2. Apply it curl -sX POST $PILOT/api/v1/brokers/5/maintenance \ -d '{"preview": false}' # 3. Watch broker-5 empty out curl -s "$PILOT/api/v1/cluster/logdirs?summary=true" # 4. After the maintenance: bring it back curl -sX DELETE $PILOT/api/v1/brokers/5/maintenance

Three endpoints:

Full OpenAPI reference at docs.calinora.io .

Why This Matters

Broker maintenance shouldn’t require a runbook. The work that needs to happen on a drain - picking target replicas, respecting rack diversity, keeping leaders stable, tracking progress - is mechanical. There’s no judgment call hidden in it. Pilot does the mechanical part so the operator can focus on the things that actually matter: what to do once the broker is down, how to communicate the change, and what to verify after it comes back.

Every action is audit-logged to __pilot_audit_log with the user identity, so the security team gets a clean trail of who drained which broker when. The reassignments themselves go through Pilot’s normal executor with its own audit events. Nothing happens silently.

Try It

Pilot ships as a single binary and a Docker image. The API is the same surface the UI and the MCP server both use, so anything you can click you can curl. Check out the documentation  for the full API reference, or visit calinora.io  to learn more.