Kafka Avro Sample Data Generator

Generate sample records from an Avro schema, one JSON document per line, in the form a producer actually takes. Every record is serialized against the schema before it is shown.

Optional fields are null by default, which is what most records look like. Filling them exercises the union branches instead, which is what you want when testing a consumer.

Paste below, or drop a file anywhere on this panel

Or drop a file anywhere on this panel. Nothing is uploaded: the analysis runs in this tab.

The answer appears here

Paste on the left and press Generate. Nothing leaves this tab.

Examples

Real input you can load into the tool above. Each one shows a different thing going wrong, because that is what the tool is for.

Sample records

Data shaped by the schema, for a producer test that does not need real data

{"type":"record","name":"User","fields":[{"name":"id","type":"int"},{"name":"email","type":["null","string"],"default":null}]}

A schema with a union

How an optional field appears in Avro JSON, wrapped in its type

{"type":"record","name":"E","fields":[{"name":"ts","type":"long"},{"name":"tag","type":["null","string"],"default":null}]}

Common mistakes

These are the ones that fail silently. The config is accepted, nothing raises an error, and the consequence arrives later.

  1. Using generated data as a fixture

    It is shaped by the schema and means nothing. Assertions written against it test the generator.

    Instead:Use it to exercise a pipeline, not to verify behaviour.

  2. Forgetting the union wrapping

    Sample data in Avro JSON wraps union values in their type. Feeding it to a plain JSON consumer fails.

    Instead:Pick the output format that matches what will read it.

  3. Assuming defaults fill optional fields

    Avro JSON requires every field present. A default applies to schema resolution, not encoding.

    Instead:Fill optional fields when the consumer expects them.

Records that serialize, which is the only claim worth making

Sample data is easy to generate and easy to generate wrongly. Every record here is serialized against the schema it came from before it is shown, so anything the schema cannot represent is reported instead of printed.

The Avro JSON encoding is what a producer takes

kafka-avro-console-producer, the REST proxy and every Avro JSON parser expect the wrapped form, where each non-null union value is a one-key object named for the branch's type and a record, enum or fixed branch uses its fullname. The plain form is easier to read and will be rejected. Both are available here and the plain one carries a warning, because that mismatch is the usual reason a first attempt fails.

kafka-avro-console-producer \
  --bootstrap-server broker:9093 \
  --topic orders \
  --property value.schema="$(cat schema.avsc)" \
  < samples.jsonl

It is deterministic on purpose

Values come from the schema and from a hash of each field's path, never from a random source. The same schema always produces the same records, so a change in the output means a change in the schema and nothing else. That makes it usable in a fixture, and it means two people looking at this page see the same thing.

Recursion terminates by taking the null branch

A recursive schema has no natural end, so a generator that follows one produces an infinite record. Nested records stop at a fixed depth and a nullable union takes null once it gets deep enough, which is why a linked-list schema gives a short chain rather than hanging. It is also why filling optional fields does not fill every one of them.

Logical types get plausible values

A timestamp-millis is a real epoch millisecond, a date is a day count, and a uuid is a well-formed v4-shaped string. They are not realistic data and they are the right type, which is what a serializer, a sink connector or a schema migration needs from them.

What this is not for

Anything that depends on the shape of real data. Sizes, cardinalities, distributions and compression ratios from these records mean nothing. Use it to exercise a consumer, a topic, a converter configuration or a Connect pipeline, and take the numbers from production.