Kafka Rebalance Duration Estimator

How long a rebalance takes, how much of it is just noticing a dead consumer, what a rolling restart costs, and what static membership removes.

Group

A member that returns inside session.timeout.ms keeps its own partitions and no rebalance happens at all.

Timeouts
What is happening

A clean close is noticed at once. SIGKILL, an OOM kill or a lost node costs the full session timeout.

kafka-rebalance.txt

updates as you type

    Common mistakes

    These are the ones that fail silently. The config is accepted, nothing raises an error, and the consequence arrives later.

    1. Restarting brokers faster than the cluster recovers

      Each restart leaves partitions under-replicated until the replica catches up. Restarting the next broker before then can drop below min.insync.replicas and stop writes.

      Instead:Wait for UnderReplicatedPartitions to return to zero between brokers.

    2. Ignoring controlled shutdown

      Without it, leadership fails over reactively and clients see errors. With it, leadership moves before the broker stops.

      Instead:Leave controlled.shutdown.enable on and allow time for it.

    3. Estimating from broker count alone

      The cost is dominated by the data each broker must re-replicate, which depends on partition count and retention, not on how many brokers there are.

      Instead:Estimate from bytes to move and the replication throughput available.

    What a rebalance actually costs

    Detection, a barrier that waits for the slowest member, and then assignment. Most of the time is in the first two, and both are configurable.

    A clean close costs nothing to detect, a kill costs the session timeout

    A consumer that calls close() sends LeaveGroup and the coordinator reacts immediately. A consumer that is killed sends nothing, so it is only noticed when its session expires, which at the 45 second default is roughly a hundred times longer than the rebalance itself. Most reports of slow rebalances are really reports of consumers being SIGKILLed, and the usual cause is a container termination grace period shorter than the shutdown hook needs. That is a deployment setting, not a Kafka one.

    The barrier is the slowest member, not the average

    Assignment cannot happen until every member has sent JoinGroup, and a member only notices the rebalance when it next polls. So the whole group waits for whichever consumer is deepest into processing its current batch. One member with a slow downstream call sets the rebalance time for everyone. Reducing max.poll.records reduces this directly, and it is more effective than any timeout change.

    Static membership removes rebalances rather than shortening them

    With group.instance.id set, a member that returns within session.timeout.ms takes its own partitions back and no rebalance happens at all. A rolling restart of twenty consumers goes from twenty rebalances to zero, provided each restart completes inside the session timeout. This is the single largest available win on consumer availability and it costs one config line, with one condition: the deploy has to be faster than the session timeout, so a slow image pull turns the win back into a rebalance.

    Eager stops everything, cooperative stops what moves

    The eager protocol revokes every partition from every member, including the partitions that are about to be assigned straight back to the same consumer. Cooperative revokes only what actually moves and pays for it with a second round. The difference is not the rebalance duration so much as what stops consuming during it. Switching needs a rolling restart with both assignors listed before the old one is removed, which is a two-deploy operation rather than a config change.

    Exceeding max.poll.interval.ms is a loop, not a warning

    If processing between polls outlasts max.poll.interval.ms, the coordinator evicts the consumer and rebalances. The evicted consumer then reprocesses the same records, takes just as long, and is evicted again. That is the group that rebalances continuously and never commits an offset, and raising the timeout usually postpones it rather than fixing it. Smaller batches fix it.

    What this cannot see

    It does not know your assignor's actual computation time, your network, or how long your consumers take to warm a cache after being assigned a new partition, which for a stateful consumer can dwarf everything modelled here. The SyncGroup figure is a small linear estimate rather than a measurement. Treat the outputs as the shape of the cost and where it sits, not as a prediction to the millisecond.

    More kafka tools

    Kafka Confluent Wire Format Decoder The five junk bytes in front of your payload Kafka Key to Partition Mapper Which partition does this key land on? Kafka Topic Name Validator Legal, risky, or 249 characters too long? Kafka Replication Safety Checker How many brokers can you lose Kafka Producer Config Linter Will it start, and will it lose a record? Kafka Message Payload Decoder The first five bytes are usually not data Kafka Connect Source Connector Generator tasks.max is a ceiling, not a count Kafka Connect Sink Connector Generator A dead letter queue with no context headers is a pile of records Kafka Connect SMT Chain Builder The order is the transforms list Kafka MirrorMaker 2 Config Generator It renames every topic by default Kafka Partition Reassignment Generator The throttle is not optional Strimzi Kafka Resource Generator Without the cluster label, nothing happens Kafka mTLS Config Generator The certificate is the identity Kafka Schema Registry Config Generator The compatibility direction is your deployment order Kafka Exactly-Once Config Generator Half of it is worse than none Kafka Broker and KRaft Config Generator The internal topics that break a one-broker cluster Kafka Quota Generator Byte rates are per broker, not per cluster Kafka Streams Config Generator application.id is four things at once Kafka Connect Worker Config Generator Security three times, or the tasks fail Kafka Retention and Unit Converter log.retention.hours does not take milliseconds Kafka Timestamp Converter Two sentinels and two meanings Kafka .properties to YAML Converter Dotted keys stay flat Kafka Streams Internal Topic Predictor Create them before Streams does Kafka ACL Generator The grant you forgot is on another resource type Kafka Topic Config Generator min.insync.replicas is the one that matters Kafka client.properties Generator The file every CLI tool asks for Kafka Producer Config Generator No password field, on purpose Kafka Consumer Config Generator The commit mode decides the semantics Kafka Disk and Retention Calculator retention.bytes is per partition Kafka Partition Count Calculator The number you can never reduce Kafka Cluster Sizing Calculator The traffic no client metric shows Kafka Consumer Lag Catch-Up Calculator Whether it ever clears, not just when Kafka Producer Batching Calculator linger.ms=0 still batches Kafka Segment and Index Sizing Why retention.ms is a lower bound Kafka Cost Estimator Your rates, so nothing goes stale Kafka Config Explorer by Version The answer depends on the release Kafka Default Config Reference What moved under a config you never edited Kafka OAuth Bearer Token Decoder Will Kafka accept it, and can it refresh Kafka Record Header Viewer Headers are a list, not a map Kafka Topic Regex Subscription Tester Kafka matches the whole name Kafka ACL Permission Matrix Viewer DENY beats every ALLOW Kafka Connect Config Validator The mistakes that raise no error Kafka Consumer Group Id Validator Which broker coordinates the group Kafka Partition Assignment Visualizer Leadership is the load, not replicas Kafka Consumer Assignment Visualizer The three assignors disagree Kafka ZooKeeper to KRaft Config Converter The authorizer class nobody changes Kafka Config to Strimzi Half of it belongs elsewhere Kafka Docker Compose Generator (KRaft) Reachable from inside and outside Kafka JAAS Config Decoder The line that stops SASL working Kafka CRC32C Calculator Which CRC, over which bytes Kafka Config Upgrade Checker What breaks when you upgrade Kafka Kafka Config Diff Which change actually changed something Kafka Consumer Config Linter Why the group rebalances, and where the records went Kafka Avro Schema Validator The defaults Avro accepts and rejects Kafka Schema Compatibility Checker What the registry will say, before you ask it Kafka Avro Schema Diff Which direction each change breaks Kafka Compression Comparison Measured on your bytes Kafka Delivery Semantics Exactly-once has a consumer half Kafka ksqlDB Query Builder It looks like SQL and the rules are not Kafka Connect SMT Predicate Tester negate reads backwards Kafka Streams Topology Viewer Count the repartitions Kafka Connect Pipeline Visualizer The order things really run in Kafka Protobuf Binary Decoder Works without the .proto Kafka Protobuf JSON Converter Why your JSON does not round-trip Kafka Protobuf to Avro Schema What does not survive the conversion Kafka Avro Binary Decoder Wrong schema, no error Kafka Avro JSON Converter Why the console producer rejects your line Kafka Avro Sample Data Generator Records that actually serialize Kafka JSON to Avro Schema What JSON cannot tell you Kafka JSON Schema to Avro What does not survive the conversion Kafka SASL JAAS Generator One login module, four syntaxes Kafka CLI Command Builder kcat is librdkafka, not Kafka

    Elsewhere on the site