Ambiguous encoding
Bytes that could be raw Avro or Confluent wire format, and why the tool refuses to guess
0c48656c6c6f
Decode an Avro payload against its schema, byte by byte with an offset per field. The wrong schema does not fail, so the signals that do exist are the verdict: bytes left over, a union index that cannot exist, a length that runs past the end.
The payload accepts base64, hex or a decimal byte array. A short all-hex string is genuinely valid as both hex and base64, and the two give different bytes, so the reading that produces a plausible first byte is taken and the other one is reported.
Or drop a file anywhere on this panel. Nothing is uploaded: the analysis runs in this tab.
The answer appears here
Paste on the left and press Decode. Nothing leaves this tab.
Nothing else to flag.
No formatting problems, and nothing the rules object to. Worth remembering what that covers: this reads the file you pasted, not the account or cluster it will be applied to.
No finding matches that filter.
Real input you can load into the tool above. Each one shows a different thing going wrong, because that is what the tool is for.
Bytes that could be raw Avro or Confluent wire format, and why the tool refuses to guess
0c48656c6c6f
An int and a string decoded against the schema, field by field
9601 0a 68 65 6c 6c 6f
These are the ones that fail silently. The config is accepted, nothing raises an error, and the consequence arrives later.
The first five bytes are a magic byte and a schema id. Decoding from byte zero either fails or produces nonsense that looks like corruption.
Instead:Strip five bytes, or let the tool detect the framing.
It is not. The bytes carry no field names and no types; without the writer schema they are unreadable.
Instead:Keep the schema with the data, or use a registry.
Avro encodes integers as zigzag varints, so a value's byte length depends on its magnitude.
Instead:Decode with the schema rather than by offset.
A record is its fields back to back, with no names, no lengths and no tags. That is what makes it small, and it is why decoding with the wrong schema does not fail: it returns plausible garbage.
A correct schema consumes the buffer exactly. Anything remaining means the writer had fields this schema does not, or that more than one record was pasted. Since the values themselves will look entirely reasonable either way, the leftover count is the verdict on this page rather than a note underneath one. It is evidence and not proof in the other direction too: two schemas of the same shape decode the same bytes into different field names, and nothing in the encoding distinguishes them.
A value produced by the Confluent Avro serializer starts with a zero magic byte and a four-byte big-endian schema id, and the Avro datum begins at byte 5. Decoding from byte 0 reads the magic byte as the first field and everything after it is misaligned. This page detects the header and removes it, and reports the id so you can fetch the exact schema it was written with.
curl -s http://registry:8081\
/schemas/ids/42 | jq -r .schema A random byte read as an array count becomes a request for billions of items, and an array of nulls costs nothing per item, so a naive decoder loops rather than failing. A block count past a hundred thousand is refused here instead of attempted. Same for a variable-length integer past ten bytes, which cannot encode a 64-bit value and is always a misaligned read rather than a large number.
Avro's long is 64 bits and a JavaScript number holds integers exactly only to 2^53. A timestamp in microseconds is already past 2^50. Anything larger is decoded as an arbitrary-precision integer rather than rounded, because the point at which a decoder starts quietly lying is not a good thing to find out from a reconciliation report.
Which schema the payload was written with. If it came from the Confluent serializer the id is in the header and the registry has the answer; otherwise there is nothing in the bytes to go on. It also does not verify a checksum, because an Avro datum has none: the CRC lives on the record batch, one level up, and the CRC tool on this site reads that.