Convert CSV to Avro
Real OCF binary,
in your browser.
Free CSV to Apache Avro converter. Creates a genuine Avro Object Container File using pure JavaScript — zig-zag varint encoding, auto-generated schema, nullable union types. Compatible with fastavro, Apache Spark, Kafka and Confluent Schema Registry. No upload, no libraries.
How Avro encodes data
The binary primitives used to build the .avro file — all implemented in JavaScript.
-1 → zigzag(1) → [0x01]
"foo" → [0x02] + encStr("foo")
How the Avro OCF binary is built
Schema generation + OCF header
CSVShift infers an Avro record schema with nullable union fields (["null", "string"], ["null", "long"] etc.) from the CSV columns. The schema JSON is embedded in the Avro file metadata map, preceded by the 4-byte magic Obj\x01 and followed by a 16-byte random sync marker.
Record serialisation with varint
Each CSV row is serialised as an Avro record. For each field, the union branch index is written as a zig-zag varint (0 for null, 1 for value). Then the value itself — string as length + UTF-8 bytes, long as zig-zag varint, double as 8-byte LE float64, boolean as single byte.
Data block + EOF
All serialised records are wrapped in an Avro data block: record count (varint) + byte size (varint) + record bytes + sync marker (16 bytes). A final empty block (varint 0) marks the end of file. The output is a standard Avro OCF readable by any Avro-compatible tool.
Read the Avro file in Python
When do you need CSV to Avro?
Apache Kafka data pipelines
Kafka producers and consumers in data pipelines use Avro with Confluent Schema Registry for schema enforcement. Converting CSV data to Avro produces messages compatible with Kafka's Avro serializer without writing a producer script.
Hadoop and data lake ingestion
HDFS, S3 and Azure Data Lake Store data lakes often use Avro for raw data storage — especially in streaming ingestion pipelines from Kafka or Flume. Converting CSV exports to Avro produces files compatible with Hive external tables and Spark jobs.
Schema evolution
Unlike CSV, Avro stores the schema with the data. When downstream consumers need to handle schema changes (adding or removing columns), Avro's schema evolution rules handle backward and forward compatibility — making it more robust than CSV for long-lived pipelines.
fastavro workflows
Python data engineers using fastavro for reading Avro files from S3 or HDFS can use CSVShift to create test fixtures and seed data in Avro format — matching the schema of production files without setting up a full Avro producer.
Related CSV tools
Avro OCF written
in your browser. Zero libraries.
CSVShift writes the Apache Avro Object Container File format directly in JavaScript — zig-zag varint encoding, Avro map encoding for file metadata, nullable union types, random sync marker. No fastavro, no avro-js, no WebAssembly. The output is byte-compatible with files produced by the official Apache Avro SDK.
Standard Avro OCF format
Nullable union types
Schema in result preview
Free, no conditions
How do I convert CSV to Avro?
fastavro.reader(open('f.avro','rb')) in Python.What is Apache Avro and why is it used?
How do I read the Avro file in Python?
import fastavro; records = list(fastavro.reader(open('f.avro','rb'))). Requires pip install fastavro. The schema is read automatically from the file — no need to provide it separately. Convert to pandas: import pandas as pd; df = pd.DataFrame(records).