Kafka Compression Comparison
Paste one representative record and get the real ratio for gzip, snappy and lz4 on your own bytes. Measured on a full batch rather than a single record, because that is what Kafka actually compresses, and a single record understates every codec.
Project the saving
Set records per second to zero to skip the projection. The saving applies to the produce, to replication factor - 1 replication transfers, and to every consumer group, which is why it is worth more than the ratio alone suggests.
Or drop a file anywhere on this panel. Nothing is uploaded: the analysis runs in this tab.
The answer appears here
Paste on the left and press Measure. Nothing leaves this tab.
Formatted
Results
Nothing else to flag.
No formatting problems, and nothing the rules object to. Worth remembering what that covers: this reads the file you pasted, not the account or cluster it will be applied to.
No finding matches that filter.
Common mistakes
These are the ones that fail silently. The config is accepted, nothing raises an error, and the consequence arrives later.
Choosing a codec from a general benchmark
Compression ratio depends almost entirely on the data. A benchmark on text says nothing about your Avro or your JSON.
Instead:Measure on your own records, which is what this page does.
Compressing at the producer and again at the broker
If the topic's compression.type differs from the producer's, the broker decompresses and recompresses every batch, which costs broker CPU for no benefit.
Instead:Set the topic to producer so the batch is stored as sent.
Measuring a single record
Kafka compresses a batch, and a codec that looks poor on one 200-byte record performs completely differently across a full batch.
Instead:Measure a realistic batch. This tool builds one for that reason.
Every published comparison is measured on somebody else's data
Compression ratio depends almost entirely on the bytes. JSON with repeated keys compresses five or six times; an already-compressed image gets slightly larger. A table of typical numbers is therefore the one thing that cannot answer which codec you should use, which is why this page compresses what you paste.
Kafka compresses a batch, not a record
The producer's batch.size and linger.ms decide how much data the codec sees at once, and a codec given one 200-byte record has almost nothing to work with. This is why measuring a single record understates every codec badly, and why this page repeats your sample to Kafka's default 16 kB batch before measuring. If your producer sends with linger.ms=0, the real ratios are worse than these and the fix is batching rather than a different codec: setting linger.ms to 20 usually beats any codec choice.
batch.size=16384
linger.ms=20 compression.type=producer on the topic is what avoids recompression
The topic-level default is producer, which means the broker stores the batch exactly as it arrived. Setting any codec name there instead makes the broker decompress and recompress every batch it receives, on the request-handling path, for every partition. It is the commonest accidental cause of broker CPU saturation and it is completely invisible unless you know to look. Choose the codec on the producer and leave the topic alone.
The saving is paid three times over
A compressed batch is smaller on the network to the broker, smaller on disk on every replica, and smaller again on the network to every consumer group. With a replication factor of 3 and two consumer groups, a 60% saving applies to four network transfers and three disk copies. That multiplier is why compression usually pays even when the ratio on its own looks unexciting, and it is what the projection below computes.
zstd is usually the right answer and is not measured here
Every browser zstd compressor is a WebAssembly build of the reference library, the smallest around 900 kB, and loading one needs a network fetch from a shipped chunk, which this site does not do. Guessing a number would be worse than omitting one, because it would look like a measurement. The command below gets the real figure on the same bytes in about a second. Since 2.1, zstd at its default level typically compresses better than gzip for less CPU, which is why it is the usual recommendation.
zstd -3 -c sample.json | wc -c
gzip -6 -c sample.json | wc -c What this cannot see
Your producer's real CPU cost. The timings here come from JavaScript implementations in your browser and are indicative of the ordering only: the relative cost of the codecs is roughly right and the absolute numbers have nothing to do with a JVM producer. It also cannot see your batch composition, since a real batch holds different records and this one holds copies of yours, which makes every codec look slightly better than production.