Hashing
Date: 2026-08-17
A one-way function turning any input into a fixed-size fingerprint. Hashing is not encryption — nothing decrypts a hash — and the hash function suitable for a lookup table is the wrong one for a password, for opposite reasons.
A hash function maps input of any size to a fixed-size output, deterministically, such that the input cannot be recovered from the output.
"hello" → 2cf24dba5fb0a30e...
"hello!" → 334d016f755cd6dc...
↑ one character changed,
completely different output
DETERMINISTIC same input → same output,
always
FIXED SIZE any input length → same
output length
ONE-WAY cannot reverse
AVALANCHE tiny input change →
unrecognisable output
Hashing is not encryption
The distinction that matters most:
HASHING one-way. No key.
Nothing decrypts it.
→ verify, don't recover
ENCRYPTION two-way, with a key.
→ recover the original
“We hash passwords so we can decrypt them if you forget” is nonsense. You verify a login by hashing the attempt and comparing hashes — the original is never recoverable, which is the entire point — Encryption Basics.
Three jobs, three different functions
1 INTEGRITY / FINGERPRINTING
"is this the same file?"
→ SHA-256
→ fast is GOOD
2 PASSWORD STORAGE
"is this the right password?"
→ Argon2, bcrypt, scrypt
→ fast is CATASTROPHIC
3 HASH TABLES / SHARDING
"which bucket does this go in?"
→ non-cryptographic, xxHash, MurmurHash
→ speed is everything, security
irrelevant
Using the wrong one for job 2 is the most consequential mistake in this note.
Why passwords need slow hashes
A fast hash is a liability the moment a database leaks:
MODERN GPU RIG, SHA-256
billions of guesses per second
→ an 8-character password falls in
hours
ARGON2 / BCRYPT, TUNED
deliberately slow and memory-hard
→ the same attack takes years
Password hashes are designed to be slow, with a tunable work factor you raise as hardware improves. That’s the whole design goal, and it’s why a general-purpose hash — however cryptographically strong — is wrong here.
Salt and pepper
WITHOUT SALT
same password → same hash
→ one precomputed table cracks every
user at once
→ identical hashes reveal identical
passwords
WITH SALT
a unique random value per user,
stored alongside the hash
→ every user must be attacked
separately
→ precomputation is useless
Salting is not optional and is not secret — it’s stored in plaintext next to the hash. Modern password libraries generate and embed it automatically, which is a strong argument for using one rather than assembling this yourself.
Pepper is an additional secret held outside the database, so a database-only leak isn’t enough. Useful, and secondary to getting salt and a slow function right.
Collisions
Two inputs producing the same output. Unavoidable in principle — infinite inputs, finite outputs — and the question is whether one can be found.
MD5 broken. Collisions on demand.
SHA-1 broken. Do not use.
SHA-256 no practical collision known
SHA-3 the newer standard
BLAKE3 fast, modern
MD5 and SHA-1 are fine for non-security fingerprinting — cache keys, deduplication, change detection — and unacceptable anywhere an adversary benefits from a collision.
Comparing hashes safely
A naive comparison leaks information through timing — it returns faster when the first bytes differ:
INSECURE if (hash === expected)
SECURE crypto.timingSafeEqual(a, b)
Matters for tokens, signatures and webhook verification. Any comparison of a secret should be constant-time.
Where you’ll actually meet it
- Webhook signatures — HMAC over the payload with a shared secret, so you can verify the sender. Compare in constant time — Shopify
- Content-addressed filenames —
app.a3f9c.js, and the image pipeline in this vault - ETags — a hash of the response body, for conditional requests — HTTP Semantics
- Git commits — every object is addressed by its hash
- Identity hashing in analytics — hashing an email to join datasets without sharing it. Note the limit: a hashed email is still personal data, because the input space is small enough to enumerate — Identity Stitching, Consent Management