Kafka Segment and Index Sizing
When your segments actually roll, how many exist per partition, what the index files cost, and how far past retention.ms your data really lives.
kafka-segments.txt
updates as you type Common mistakes
These are the ones that fail silently. The config is accepted, nothing raises an error, and the consequence arrives later.
Expecting retention.ms to delete on time
Retention applies to CLOSED segments. An open segment is never deleted, so data older than the retention stays until the segment rolls, which segment.bytes and segment.ms control.
Instead:Set segment.ms if timely deletion matters. retention.ms alone is a lower bound.
Using very small segments to delete sooner
Each segment costs file handles and index memory, and a topic with thousands of tiny segments slows broker startup and recovery noticeably.
Instead:Balance the two. Segments in the hundreds of megabytes are typical.
Forgetting compacted topics keep the active segment forever
Compaction never touches the active segment, so the latest value per key can sit uncompacted indefinitely on a low-traffic topic.
Instead:Set segment.ms on compacted topics too.
Why retention.ms is a lower bound rather than a promise
Kafka deletes whole segments and never touches the active one, so how long data really lives depends on when segments roll.
The active segment cannot be deleted
Retention works by deleting closed segments. The segment currently being written to is not a candidate, whatever its age, so the oldest record in it survives for its retention plus however long that segment stays open. If segment.ms is longer than retention.ms and the write rate is too low to fill segment.bytes, data stays for the segment span regardless of what retention says. That is a disk planning problem when retention was set for space, and a compliance problem when it was set to satisfy a deletion promise, because the topic quietly keeps records past the date you told someone it would not.
Three things roll a segment, and the third is a surprise
segment.bytes and segment.ms are the two people know about. The third is the offset index filling up: one index entry is written every index.interval.bytes, each entry is 8 bytes, and the index file is capped at segment.index.bytes. That makes index.interval.bytes an indirect ceiling on segment size. With the defaults it works out to terabytes and never binds, but lower index.interval.bytes to make lookups finer and you can find segments rolling at a few megabytes with segment.bytes set to a gigabyte and having no effect at all.
Index files are preallocated in full
Both the offset index and the time index for the active segment are created at segment.index.bytes and only truncated when the segment rolls. At the 10 MiB default that is 20 MiB of disk per active partition per replica, occupied before a single record is written. On a broker with a few thousand partitions this is real space that df reports and no Kafka metric explains, and it is why a freshly created topic appears to consume disk immediately.
Segment count is a restart time question
Every segment is three open file handles: the log, the offset index and the time index. That sets an ulimit requirement, and the failure mode is Too many open files during recovery rather than during steady state. It also sets broker startup time, because startup scans segments to rebuild state. A broker that restarted in a minute a year ago and takes twenty now usually has not changed configuration; it has accumulated segments.
Smaller segments make retention accurate and files numerous
Deletion granularity is one segment, so a topic with two segments over its retention window swings between holding roughly the window and roughly twice it. Smaller segments track retention more closely at the cost of more files, more handles and slower startup. There is no correct answer, only the trade, and this page shows both ends of it so the choice is visible.