Serialisation Formats
Date: 2026-08-17
Turning structured data into bytes that can be stored or sent, and back again. Each format encodes a different assumption about who reads it — and choosing wrong is usually discovered years later, when the schema needs to change.
Serialisation converts in-memory data into a linear sequence of bytes. Deserialisation reverses it. The format decides what survives the trip: types, structure, ordering, and whether the reader needs to know the shape in advance.
The same data, four ways
JSON — self-describing, text
{"id":42,"name":"Serum","price":24.99,
"tags":["new"]}CSV — tabular, no nesting
id,name,price,tags
42,Serum,24.99,newXML — verbose, with attributes, namespaces and schemas
<product id="42">
<name>Serum</name>
</product>Protobuf — binary, and unreadable without the .proto schema
[0x08 0x2a 0x12 0x05 ...]
What each assumes about the reader
| Reader needs schema | Human-readable | Types | Nesting | |
|---|---|---|---|---|
| JSON | No | Yes | Weak | Yes |
| CSV | Sort of | Yes | None | No |
| XML | Optional | Barely | Via schema | Yes |
| Protobuf / Avro | Yes | No | Strong | Yes |
| MessagePack | No | No | Weak | Yes |
That first column is the real decision. Self-describing formats let an unknown reader parse them; schema-required formats are smaller and stricter but useless to anyone without the schema.
JSON’s gaps, which you will hit
It’s the default for good reasons — ubiquitous, readable, native to JavaScript — and it has specific holes:
NO DATE TYPE
"2026-08-17T10:00:00Z" is a string
everyone agrees to parse
NUMBERS ARE DOUBLES
integers above 2^53 lose precision
→ send IDs as strings
NO COMMENTS
configuration files suffer
NO BINARY
base64, +33% size
NO TRAILING COMMAS
a real source of hand-editing errors
The large-integer problem is the one that draws blood. A 64-bit database ID silently rounds:
9007199254740993 sent
9007199254740992 received
No error, no warning, wrong record. Send identifiers as strings — Character Encoding.
CSV is not a format
It’s a family of conventions that disagree. There is no universal standard for quoting, escaping, line endings, or encoding, and every tool has its own dialect.
name,note
"Smith, John","said ""hello"""
↑ comma inside quotes
↑ escaped quote
Excel may write UTF-8 with a BOM,
or CP-1252, depending on locale
Use a real parser, never split(','). And for money, remember CSV has no types — 007 becomes 7, and long numbers become scientific notation the moment a spreadsheet opens the file.
Schema formats, and why they exist
Protobuf, Avro and Thrift require a schema shared by writer and reader. That buys three things a self-describing format can’t:
- Size. Field names aren’t in the payload — a tag number is
- Speed. No parsing of text, no type guessing
- Evolution rules. The schema defines what a compatible change is, so old readers and new writers are a solved problem rather than a hope — Backwards Compatibility
The cost is coupling: the reader must have the schema, and distributing schemas is its own infrastructure. This is the trade behind every event pipeline — Event Taxonomy Design.
Choosing
public API, unknown clients JSON
config a human edits YAML/TOML
tabular export to a spreadsheet CSV
high-volume internal traffic Protobuf
analytics events at scale Avro/Parquet
browser ↔ server JSON
Default to JSON and be deliberate about leaving it. The reasons to leave are volume, strict typing, or schema evolution you can’t manage by convention.
The failure that matters most
Not choosing wrong — not planning for the schema to change. Every format needs an answer to “what happens when a field is added, renamed or removed while old and new code both run”. JSON’s answer is convention and discipline; Protobuf’s is enforced by the schema. Having no answer is the actual mistake — Versioning.