Tags: web-dev concept

Serialisation Formats

Date: 2026-08-17


Turning structured data into bytes that can be stored or sent, and back again. Each format encodes a different assumption about who reads it — and choosing wrong is usually discovered years later, when the schema needs to change.


Serialisation converts in-memory data into a linear sequence of bytes. Deserialisation reverses it. The format decides what survives the trip: types, structure, ordering, and whether the reader needs to know the shape in advance.

The same data, four ways

JSON — self-describing, text

{"id":42,"name":"Serum","price":24.99,
 "tags":["new"]}

CSV — tabular, no nesting

id,name,price,tags
42,Serum,24.99,new

XML — verbose, with attributes, namespaces and schemas

<product id="42">
  <name>Serum</name>
</product>

Protobuf — binary, and unreadable without the .proto schema

[0x08 0x2a 0x12 0x05 ...]

What each assumes about the reader

Reader needs schemaHuman-readableTypesNesting
JSONNoYesWeakYes
CSVSort ofYesNoneNo
XMLOptionalBarelyVia schemaYes
Protobuf / AvroYesNoStrongYes
MessagePackNoNoWeakYes

That first column is the real decision. Self-describing formats let an unknown reader parse them; schema-required formats are smaller and stricter but useless to anyone without the schema.

JSON’s gaps, which you will hit

It’s the default for good reasons — ubiquitous, readable, native to JavaScript — and it has specific holes:

NO DATE TYPE
  "2026-08-17T10:00:00Z" is a string
  everyone agrees to parse

NUMBERS ARE DOUBLES
  integers above 2^53 lose precision
  → send IDs as strings

NO COMMENTS
  configuration files suffer

NO BINARY
  base64, +33% size

NO TRAILING COMMAS
  a real source of hand-editing errors

The large-integer problem is the one that draws blood. A 64-bit database ID silently rounds:

9007199254740993   sent
9007199254740992   received

No error, no warning, wrong record. Send identifiers as strings — Character Encoding.

CSV is not a format

It’s a family of conventions that disagree. There is no universal standard for quoting, escaping, line endings, or encoding, and every tool has its own dialect.

name,note
"Smith, John","said ""hello"""
  ↑ comma inside quotes
                ↑ escaped quote

Excel may write UTF-8 with a BOM,
or CP-1252, depending on locale

Use a real parser, never split(','). And for money, remember CSV has no types — 007 becomes 7, and long numbers become scientific notation the moment a spreadsheet opens the file.

Schema formats, and why they exist

Protobuf, Avro and Thrift require a schema shared by writer and reader. That buys three things a self-describing format can’t:

  • Size. Field names aren’t in the payload — a tag number is
  • Speed. No parsing of text, no type guessing
  • Evolution rules. The schema defines what a compatible change is, so old readers and new writers are a solved problem rather than a hope — Backwards Compatibility

The cost is coupling: the reader must have the schema, and distributing schemas is its own infrastructure. This is the trade behind every event pipeline — Event Taxonomy Design.

Choosing

public API, unknown clients    JSON
config a human edits           YAML/TOML
tabular export to a spreadsheet CSV
high-volume internal traffic   Protobuf
analytics events at scale      Avro/Parquet
browser ↔ server               JSON

Default to JSON and be deliberate about leaving it. The reasons to leave are volume, strict typing, or schema evolution you can’t manage by convention.

The failure that matters most

Not choosing wrong — not planning for the schema to change. Every format needs an answer to “what happens when a field is added, renamed or removed while old and new code both run”. JSON’s answer is convention and discipline; Protobuf’s is enforced by the schema. Having no answer is the actual mistake — Versioning.