WebsiteDocs
← Back to Blog

From Alert to Fix in 60 Seconds: AI-Powered Kafka Operations

It’s Wednesday afternoon. The payments team pings your on-call channel: “orders are processing slowly.” You open a terminal, SSH into a jumpbox, run kafka-consumer-groups.sh --describe, and stare at a wall of partition offsets trying to figure out which partitions are lagging and why.

Twenty minutes later, you’ve found the consumer group has lag. But is it a slow consumer? A hot broker? A topic partition imbalance? You need to cross-reference consumer lag with partition assignments, broker disk usage, and activity metrics. That’s another three CLI commands, a spreadsheet, and a lot of context-switching.

There’s a better way.

The Scenario

Pilot ships with a built-in AI agent that connects directly to your Kafka cluster through the Model Context Protocol (MCP) . Instead of memorizing CLI flags and cross-referencing outputs, you describe what you’re seeing in natural language. The agent has access to over 60 specialized tools - from consumer group inspection to broker health checks to partition rebalancing.

Here’s a real interaction. Press play to watch.

pilot - agent session
ready
⬡Press play to watch the demo
0/25 steps

Five messages from the operator. Seven tool calls by the agent. One approval gate. The entire investigation-to-fix loop completed in under a minute - a workflow that traditionally takes an hour or more of manual CLI work, JSON wrangling, and spreadsheet analysis.

What Just Happened?

Let’s break down the agent’s investigation.

1. Consumer Group Inspection

The operator reported a symptom - slow order processing. The agent immediately called two tools in parallel: describe_consumer_group for group state and membership, and get_consumer_group_lag for per-partition lag numbers. Within seconds, it identified that lag was concentrated on two specific partitions (P3 and P7), not spread evenly - a pattern that points to an infrastructure issue rather than a slow consumer.

2. Root Cause Analysis

Without being asked, the agent followed the thread: if two partitions are lagging while others aren’t, what do those partitions have in common? It called get_partition_activity and get_cluster_health in parallel, and found the answer. Both partitions were led by broker-4, which was at 87% disk utilization while the cluster average was 54%. High disk pressure causes I/O contention that directly impacts consumer fetch latency.

This kind of multi-hop reasoning - from consumer lag to partition assignment to broker health - is exactly what makes Kafka troubleshooting time-consuming. The agent connected the dots in seconds.

3. Options, Not Orders

When the operator asked “what can we do?”, the agent didn’t jump straight to a solution. It presented two options: a quick fix (leader elections only - instant relief but temporary) and a lasting fix (full rebalance proposal that also addresses the disk imbalance). It explained the trade-offs and recommended the full proposal, but left the decision to the operator.

Once asked, the generate_proposal tool triggered Pilot’s multi-phase rebalancing pipeline. This isn’t a cluster-wide shuffle. The engine generated a targeted plan: 5 leader elections (instant, zero data transfer) to immediately redistribute read load, plus 3 replica moves (14.2 GB) to prevent the problem from recurring. Broker-4’s disk usage drops from 87% to 71%.

The distinction between leader elections and replica moves matters. Leader elections change which broker serves reads for a partition - they complete in milliseconds with zero data movement. Replica moves physically transfer partition data between brokers - they take time but fix the underlying imbalance. The engine maximizes elections and minimizes transfers.

4. Approval Gate

When the agent called apply_proposal, Pilot didn’t just execute it. The tool returned pending_approval and the agent explained what would happen: which moves are instant, how much data will transfer, and at what rate. Every mutation - whether it’s a partition reassignment, a config change, or an ACL update - goes through this gate. The AI can diagnose and recommend, but it cannot act without your consent.

60+ Tools, One Natural Language Interface

The MCP integration exposes Pilot’s full operational surface as structured tools that the AI can compose dynamically:

Diagnostics - get_cluster_overview, get_cluster_health, get_partition_activity, get_broker_racks, search_topics, list_consumer_groups, describe_consumer_group, get_consumer_group_lag

Simulation - what_if_simulate (broker failure, rack failure, traffic spike), blast_radius_analyze (single-failure impact analysis), generate_proposal (multi-metric optimization)

Operations - apply_proposal, redistribute_topic, redistribute_partition, preferred_leader_election, update_topic_config, reset_consumer_group_offsets, create_acls, update_quota

Audit - search_audit_log, revert_audit_event (undo previous changes)

Read-only tools execute automatically. Mutation tools always require approval. This isn’t a configuration option - it’s a hard architectural constraint.

Why This Matters

Kafka operations have a knowledge problem. The platform is powerful but the operational surface is vast: broker configs, topic configs, partition assignments, consumer groups, quotas, ACLs, rack awareness, replication factors, ISR management. Most teams have one or two people who understand it deeply, and everyone else is searching Stack Overflow during incidents.

An MCP-powered agent doesn’t replace expertise - it makes expertise accessible. A junior engineer can report “orders are slow” and get the same root cause analysis that a senior SRE would produce. The approval gate ensures they can’t accidentally break anything.

The combination of natural language understanding, structured tool access, and human-in-the-loop safety creates an operations model where:

Try It

Pilot is available as a Docker image. The AI assistant works with Anthropic (Claude), OpenAI, and local models - bring your own API key or run fully offline. Check out the documentation  for setup instructions, or visit calinora.io  to learn more.