Inferred from samples
A number with no decimal part is inferred as int, which overflows above two billion
{"id":1,"name":"a"}
{"id":2,"name":null}Infer an Avro schema from sample records. Every guess is listed with its path, because JSON cannot tell an int from a long, an absent key from a null one, or an empty array's element type from nothing at all.
Paste one record, a JSON array of them, or one JSON document per line. More samples give a better schema, because an optional field is only visible when it is missing from one of them.
Or drop a file anywhere on this panel. Nothing is uploaded: the analysis runs in this tab.
The answer appears here
Paste on the left and press Infer schema. Nothing leaves this tab.
Nothing else to flag.
No formatting problems, and nothing the rules object to. Worth remembering what that covers: this reads the file you pasted, not the account or cluster it will be applied to.
No finding matches that filter.
Real input you can load into the tool above. Each one shows a different thing going wrong, because that is what the tool is for.
A number with no decimal part is inferred as int, which overflows above two billion
{"id":1,"name":"a"}
{"id":2,"name":null}Fields absent in any sample become nullable, because the inference cannot assume
{"id":1,"email":"a@b.com"}
{"id":2}These are the ones that fail silently. The config is accepted, nothing raises an error, and the consequence arrives later.
A field absent in the sample is absent from the schema, and a null is inferred as null-only.
Instead:Give several samples covering the real variation.
A number with no decimal part infers as int, which overflows at 2,147,483,647. Ids reach that.
Instead:Widen to long for anything that counts up.
A field present in every sample is inferred required, which breaks the first time it is absent.
Instead:Mark optional fields explicitly rather than relying on the samples.
Inferring a schema from data is guessing, and the guesses are the interesting part. Every one of them is listed here with its path, and the ones that cannot be guessed are marked rather than filled in.
A field is optional when it is absent from some records and present in others. From a single sample every field looks present and required, which produces a schema that cannot be evolved and rejects half the topic. Paste several records, one JSON document per line: a dump of a few hundred is better than a hand-written example, because it contains the cases nobody thought of.
kcat -C -t orders -e -c 500 \
-f '%s\n' > samples.jsonl A value of 3 could be an int, a long, a float or a fixed-point amount, and nothing in the document says which. This picks the narrowest type that holds every sample and tells you it did. Widening turns every integer into a long instead, which is worth doing for anything that will grow: an int that overflows fails at serialization rather than truncating, and widening a field from int to long later is a backward compatible change while the reverse is not. For money, neither is right: use bytes with a decimal logical type and a fixed scale.
This is the one thing here that more data of the same kind cannot fix. An array that is empty in every sample says nothing about what goes in it, so a placeholder is emitted and the finding is loud. Nothing else in this tool invents a type it has no evidence for.
Names are letters, digits and underscores, and must not start with a digit. There is no escaping and no alias that helps, because an alias has to be a legal name too. So a JSON key with a hyphen, a dot or a space is renamed here and reported as critical: the schema you get does not describe the JSON you pasted, and the key has to change in whatever produces the records. A Connect ReplaceField transformation does it in a pipeline.
There is nothing in the data to infer a default from, so every field comes out without one. That is fine for a first version and it means none of these fields can be added to a schema that is already registered. Adding defaults before you register is a minute's work and saves the argument later. The schema validator on this site says the same thing about any schema you paste.
Semantics. It cannot tell a timestamp from a number, a currency amount from a float, or an id from any other string, so no logical type is ever guessed. Everything runs in your browser, which matters more than usual here: sample records are real records.