Tags: web-dev concept

SQL vs NoSQL

Date: 2026-08-17


A badly-named debate. “NoSQL” covers four unrelated families of database, and the decision is never about the query language — it’s about whether you know your access patterns in advance, and whether you need transactions across records.


“SQL” conventionally means a relational database with a schema and transactions. “NoSQL” means everything else, which is not a category so much as an absence of one.

The four families under one label

DOCUMENT        MongoDB, DynamoDB, Firestore
  JSON-ish documents, nested
  flexible shape, no joins

KEY-VALUE       Redis, Memcached
  get and set by key, nothing else
  extremely fast, no querying

WIDE-COLUMN     Cassandra, HBase, Bigtable
  rows with dynamic columns
  built for write throughput and scale

GRAPH           Neo4j, Neptune
  nodes and edges as first-class
  traversal is native and cheap

These have nothing in common except not being relational. Comparing “SQL vs NoSQL” is comparing one thing to four, which is why the debate never resolves.

The actual question

Do you know your access patterns in advance?

RELATIONAL
  model the DATA correctly
  → query it however you like later
  → the planner works out how
  → new questions cost nothing to ask

DOCUMENT / WIDE-COLUMN
  model the QUERIES
  → store data shaped for how it's read
  → the anticipated query is very fast
  → an UNANTICIPATED query is expensive
    or impossible without restructuring

That’s the trade, and everything else follows from it. A relational schema is a bet that you’ll be asked questions you haven’t thought of. A document schema is a bet that you won’t.

For a business where analysts, reporting and new features keep asking new questions — which is nearly every commercial business — that bet usually favours relational — The Relational Model.

Where the myths are wrong

“NoSQL is schemaless.” There is always a schema; the only question is whether the database enforces it or your application code does. Schema-on-read moves the problem to every reader, and old documents in old shapes remain forever — Serialisation Formats.

“NoSQL scales, SQL doesn’t.” Postgres and MySQL handle very large workloads. Horizontal write-scaling is genuinely harder relationally, but almost nobody hits that ceiling — and most who claim to have a query problem, not a scale one.

“NoSQL has no transactions.” Most now offer them, often with restrictions on scope. The older blanket claim is out of date.

“JSON in a relational database is a hack.” Postgres jsonb is indexable and queryable, and covers most document use cases without giving up joins and constraints. This is frequently the right answer and it’s routinely overlooked — you don’t have to choose.

Where non-relational genuinely wins

  • Caching and sessions — Redis, and it isn’t a close call
  • Genuinely huge write volume with simple access — telemetry, event streams
  • Graph traversal of unknown depth — recommendation and social graphs. Recursive SQL exists — Common Table Expressions — and is unpleasant
  • Full-text search — Elasticsearch or a dedicated engine, not LIKE '%term%'
  • Documents with genuinely variable shape, where each record legitimately differs
  • Analytics at scale — columnar warehouses, which are relational but a different implementation entirely — Warehouse-First Analytics

Polyglot persistence

The realistic answer for most systems:

Postgres      orders, customers, products
              ← the system of record
Redis         sessions, cart, rate limits
Elasticsearch product search
BigQuery      analytics

One system of record, several derived stores. The rule that keeps it sane: the derived stores can be rebuilt from the system of record. If losing Redis loses data you can’t reconstruct, it has quietly become a system of record without the guarantees of one.

Choosing

Do you need transactions across records?
  yes → relational

Are access patterns fixed and known?
  no  → relational

Will you outgrow one large server?
  probably not → relational

Is it a cache, a queue, or a search index?
  → a purpose-built tool

Otherwise → relational, with jsonb where
            the shape is genuinely variable

The default is relational, and the burden of proof is on leaving it. Not because it’s always best, but because reimplementing joins, constraints and transactions in application code is the most common and most expensive way this decision goes wrong.