Kafka Protobuf Binary Decoder

Decode a protobuf payload, with a .proto or without one. The wire format carries field numbers and wire types, so the structure is always readable. Where a field could be a string, bytes, an embedded message or a packed array, every reading is listed rather than one being guessed.

Leave the schema empty for a structural decode, which is the useful case when you have bytes off a topic and cannot find the .proto. The payload accepts base64, hex or a decimal byte array.

Paste below, or drop a file anywhere on this panel

Or drop a file anywhere on this panel. Nothing is uploaded: the analysis runs in this tab.

The answer appears here

Paste on the left and press Decode. Nothing leaves this tab.

Examples

Real input you can load into the tool above. Each one shows a different thing going wrong, because that is what the tool is for.

Schemaless decode

Every reading the wire format admits for these bytes, because a field tag alone does not fix the type

089601120548656c6c6f

A nested message

A length-delimited field that could be a string, bytes or a sub-message

0a0508961612 03616263

Common mistakes

These are the ones that fail silently. The config is accepted, nothing raises an error, and the consequence arrives later.

  1. Assuming a field tag identifies the type

    The wire type distinguishes varint from length-delimited and nothing more. A length-delimited field could be a string, bytes, a packed array or a nested message.

    Instead:Decode with the .proto, or read every possible interpretation, which is what this page shows.

  2. Expecting an unknown field to be an error

    Protobuf ignores unknown fields by design, which is what makes it forward compatible. A typo in a field number is silently dropped.

    Instead:Check the field numbers against the schema.

  3. Reading a negative int32 as a small number

    A negative value encoded as int32 uses ten bytes, because it is sign-extended. sint32 uses zigzag and is much smaller.

    Instead:Use sint32 or sint64 for values that are often negative.

Protobuf bytes are partially self-describing, which is why this works with no schema

Every field is prefixed with a tag holding its field number and one of five wire types, so a payload can be walked with no .proto at all. What the tag does not say is the declared type, and that is where the honesty has to come in.

One wire type, four possible readings, and the bytes do not choose

Wire type 2 is shared by string, bytes, every embedded message and every packed repeated field. The three bytes 08 96 01 are both the message {1: 150} and a perfectly ordinary byte string, and nothing in the encoding distinguishes them. This page lists every reading a field admits instead of picking one, because picking one is what produces the confident nonsense that other decoders are known for. Wire type 0 has the same problem more mildly: int32, int64, uint32, uint64, sint32, sint64, bool and every enum share it, so the value is shown read every way at once.

The five bytes in front, and the array after them

A value from Confluent's serialiser starts with a zero magic byte and a four-byte big-endian schema id. For protobuf, and only for protobuf, a message-index array follows that, so the payload begins at byte 5 plus the array rather than at byte 5. Reading from byte 5 leaves the array in front of the payload, which is exactly the decode that comes out garbage by a few bytes. This page detects both and reports the id so the registry can be asked for the schema.

curl -s http://registry:8081\
  /schemas/ids/42 | jq -r .schema

Sign extension is where hand-decoding goes wrong

An int32 holding -1 is widened to 64 bits before encoding, so it occupies ten bytes rather than the five somebody counting bits expects. A buffer truncated inside one of those runs is the usual cause of a last field that reads as nonsense. Values are decoded as arbitrary-precision integers rather than as JavaScript numbers, because a uint64 runs to 18446744073709551615 and a double stops being exact at 2^53.

Stopping early is the diagnosis, not a crash

When the walk cannot read a byte it reports what it read, where it stopped and how much was left, because that is usually the answer: a length that runs past the end of the buffer means either the payload is truncated or the offset it was read from is wrong. Fields decoded before the fault are kept rather than discarded.

What this cannot see

Which message these bytes are. Protobuf payloads are not self-describing at the message level, so with several messages in a .proto nothing in the buffer says which one applies. It also cannot tell an absent field from one set to its default: proto3 does not write defaults, so they are the same bytes. There is no checksum here either, because a protobuf message has none.