Tags: web-dev concept

AI Coding Tools

Date: 2026-08-17


Assistants that generate, complete and modify code. They’re genuinely useful for a specific and predictable set of tasks and confidently wrong on another — and the failure mode is that wrong output looks exactly like right output, which shifts the whole cost onto review.


AI coding tools generate or transform code from natural-language instructions or surrounding context — completion in the editor, chat-based generation, and agents that edit across files and run commands.

[CHECK: capabilities in this area change quickly. Verify current tool behaviour rather than relying on any general claim here.]

Where they’re reliably good

The pattern: high-volume, well-specified, verifiable work.

BOILERPLATE      a component skeleton, a
                 config file, a test
                 harness

MECHANICAL       renaming across files,
TRANSFORMS       converting a format,
                 migrating an API's call
                 sites

TESTS FROM       given a function, the
EXISTING CODE    edge cases you'd have
                 written

UNFAMILIAR       "what does this regex do"
SYNTAX           "how do I do X in bash"

FIRST DRAFTS     of documentation, commit
                 messages, error messages

EXPLAINING       a codebase you've just
                 inherited

The common property is that the output is cheap to verify. You can read the generated test and know whether it’s right.

Where they’re reliably poor

NOVEL LOGIC      anything where the answer
                 isn't in the training
                 distribution

YOUR CONVENTIONS unless given them
                 explicitly

WHOLE-SYSTEM     it sees the context
DESIGN           window, not the system

CURRENT VERSIONS confident output using an
                 API that changed
                 — Versioning

SECURITY-        plausible auth code with
SENSITIVE CODE   a subtle hole
                 — Common Vulnerabilities

ANYTHING WHERE   the confident wrong answer
BEING WRONG      is indistinguishable from
IS EXPENSIVE     the right one

See: Versioning · Common Vulnerabilities

That last row is the whole risk. A human who doesn’t know says so; a model produces fluent, well-formatted, plausible code with the same confidence either way.

The review problem

GENERATION      fast
VERIFICATION    the same as always

The bottleneck moves to review, and the volume of code needing review goes up. That’s a bad trade if review capacity is already the constraint — which it usually is — Code Review.

The specific hazard: generated code is easy to skim. It’s well-formatted, plausibly named and reads fluently, all of which suppress the scrutiny that a colleague’s slightly-odd code would attract.

Rules that hold

  • You own what you commit. “The AI wrote it” is not a defence in review, in an incident, or in a security assessment
  • Never accept code you can’t explain. If you can’t say why it works, you can’t debug it at 3am
  • Verify against the actual documentation for anything version-dependent. This is the single most common failure — plausible code using an API that changed — Dependency Management
  • Never paste secrets or customer data into a tool. Check where the data goes and what’s retained — Secrets Management, PII in Analytics
  • Treat suggested dependencies with suspicion. Models occasionally suggest packages that don’t exist, and typosquatters have registered such names — Supply Chain Risk
  • Give it your conventions explicitly. A rules file, or a linked example — otherwise it writes generic code that passes review and doesn’t match anything

Where it fits in the existing pipeline

Nothing about these tools changes what verification is for. Generated code goes through exactly the same gates:

types            — Type Checking in CI
lint             — Linting
tests            — Testing Strategy
review           — Code Review

See: Type Checking in CI · Linting · Testing Strategy · Code Review

If anything, the gates matter more, because the volume is higher and the author’s understanding is shallower. A codebase with weak automated verification gets worse faster with these tools, not better.

The realistic assessment

Substantial productivity gain on the mechanical half of the job; no substitute for understanding the system.

The developers who benefit most are the ones who already know what good looks like — they can spot the wrong answer immediately and accept the right one without reading it twice. The risk concentrates on people who can’t yet make that judgement, which is a real problem for how the skill gets learned.

Use them. Verify everything. Don’t ship what you don’t understand.