Tags: commerce experimentation landscape
2026-09-28 - Landscape - AI and Conversion
Captured: 2026-09-28
Confidence: mixed — stated per section. Protocol and product facts mostly from primary sources; nearly every traffic and outcome figure is vendor data. Researched on the day; OpenAI’s own pages were unreachable, so OpenAI facts rest on reporting of spokesperson statements.
Watch for: Google’s agentic checkout (UCP) reaching the UK — announced “later”, no date · the FCA’s response on agentic payments · Adobe’s monthly AI-traffic growth rate settling
Where AI meets conversion work, September 2026. AI has moved into the top of the funnel and into the tools — not into the checkout. Shoppers increasingly discover products through AI assistants and arrive on merchant sites later in the decision and converting better; agents that buy on the shopper’s behalf are live at the card-network level and barely anywhere else. Testing tools build experiments faster, and there is still no independent evidence they build better ones.
Previously on this topic: 2026-08-16 - Landscape - Industry, sections Experimentation tooling, Crawlers and discoverability and AI in the development workflow. Checked against it:
- Its experimentation table was already out of date when captured. It had Statsig “still operating independently” after OpenAI’s acquisition; Amplitude had taken over Statsig’s platform and customers in May 2026, three months earlier. Eppo, described as “folded into product analytics”, relaunched as Datadog Experiments in April 2026. VWO and AB Tasty, “merging”, launched as Wingify in September 2026
- Its expectation 4 — “AI agent adoption is the number most likely to have moved” — partly holds. Agents moved in payments infrastructure; measured agent usage is still small (see Agents in the measurement)
- Its crawler claim — AI crawlers don’t render JavaScript — needs qualifying: user-directed agents are a different class, and ChatGPT’s agent runs JavaScript as a normal browser
Agentic checkout — the protocols
Confidence: high on what exists, low on adoption. Specs and announcements verified against primary sources; usage isn’t published by anyone.
Agentic commerce means an AI agent carrying out some or all of a purchase for the shopper — finding the product, filling the basket, paying. Four layers of standard are competing to carry it:
| Protocol | From | What it does | Status, Sept 2026 |
|---|---|---|---|
| ACP — Agentic Commerce Protocol | OpenAI and Stripe, Sept 2025 | product feed, a checkout API the merchant implements, delegated payment token to the merchant’s payment provider | beta; latest snapshot April 2026 adds cart, orders, auth and MCP (Model Context Protocol — the standard way to connect an AI to a tool) |
| AP2 — Agent Payments Protocol | Google, Sept 2025, 60+ partners | signed “mandates” proving the shopper authorised this specific cart and payment | announced; later version and a reported move to the FIDO Alliance are secondary-sourced only |
| UCP — Universal Commerce Protocol | Google with Shopify, Etsy, Wayfair, Target, Jan 2026 | the whole journey: discovery, cart, checkout, account linking, order updates, payment-token exchange; carries AP2 for payment | live on Google surfaces in the US, Canada and Australia; Amazon, Meta, Microsoft, Salesforce and Stripe joined its council in April 2026 |
| Card-network layer | Visa Trusted Agent Protocol (Oct 2025); Mastercard Agent Pay (Apr 2025) | tell verified agents from bots; tokenised credentials an agent can use | live in Europe including the UK — see below |
The significant event of the period: checkout inside the chat retreated. OpenAI launched Instant Checkout in ChatGPT with ACP in September 2025 and by March 2026 was “allowing merchants to use their own checkout experiences while we focus our efforts on product discovery”. ChatGPT now hands the shopper to the merchant’s site, or to a merchant-built ChatGPT app running on the merchant’s own checkout. Reports of how few merchants went live and how it converted are secondary-sourced and not repeated here.
The pattern that has emerged: the AI finds the product; the merchant’s site closes the sale. Google’s UCP is the exception — checkout on Google’s surface with the retailer as merchant of record — and it isn’t live in the UK.
Other moves worth knowing:
- Shopify — Agentic Storefronts (products syndicated into AI assistants), its Catalog opened to non-Shopify brands, and checkout embedded in Microsoft Copilot (US), all announced January 2026. Reports that Agentic Storefronts were switched on for all stores by default are secondary-sourced
- Salesforce — shopper and merchant agents generally available June 2026; ChatGPT and Google integrations announced, not confirmed shipped
- Adobe Commerce — committed to UCP and ACP (Feb 2026); no availability date found
- PayPal — Store Sync (catalogue into AI platforms) and instant purchase inside Perplexity, late 2025
Amazon blocks agents that don’t identify themselves — 47 AI bots in robots.txt from November 2025, and Meta’s shopping agent at checkout in September 2026 — while sitting on the UCP council and running its own agents (Buy for Me; Alexa for Shopping in the US). Its injunction against Perplexity’s Comet browser was granted in March 2026 and overturned on appeal in August, the appeal court treating an agent acting on a user’s instruction as the user. The case continues. The signal: agent access on the platform’s terms, not no access.
Feeds: Checkout Instrumentation Constraints · Headless Architecture · Shopify
The UK position
Confidence: high. Regulator and network announcements are primary.
- Live: card-network agent payments. Visa went live across Europe in July 2026 with UK issuers Barclays, HSBC UK, Lloyds, Nationwide and NatWest, and merchants including lastminute.com and Frasers. Mastercard reports its European issuers enabled. No volumes disclosed by either
- Not live: Google’s UCP checkout (“later”, no date); ChatGPT shopping in the UK is unconfirmed
- The CMA (Competition and Markets Authority) published Agentic AI and consumers and business guidance in March 2026: a business is responsible for what its AI agent does as it would be for an employee, and the DMCC Act (Digital Markets, Competition and Consumers Act 2024) regime allows fines up to 10% of worldwide turnover. No new legislation proposed
- The FCA (Financial Conduct Authority) said in March 2026 it would consider whether payments regulation needs to change for agents; its July 2026 review of AI in retail finance recommends frameworks for agent authorisation, identity and liability. No rules yet
For a UK merchant, the practical question this year is discovery — being found and represented correctly in AI assistants — not agent checkout.
Feeds: Deceptive Design · Testing and Compliance
AI-referred traffic
Confidence: medium. Consistent across sources, but all vendor data, and measured by referrer — which undercounts.
AI-referred traffic is visits arriving from a link in an AI assistant (ChatGPT, Gemini, Copilot, Perplexity).
Adobe’s figures for US retail — Adobe Analytics data across the top 2,000 US retailers:
AI-referred visits, conversion vs
year on year non-AI traffic
Mar 2025 — 38% worse
Nov–Dec 2025 +693% better (crossover ~Sept 2025)
Q1 2026 +393% 42% better (March)
May 2026 +138% 54% better
Jul 2026 +62% 60% better · revenue/visit +53% · bounce −33%
Two things to notice. Growth is decelerating fast — mostly a base effect, as each month now compares with a year in which AI traffic had already grown. And the conversion gap reversed inside a year, from 38% worse to 60% better.
UK (Adobe, reported September 2026): AI referrals +149% year on year in August 2026, converting 20% better than other traffic; May 2026 was the first month AI-referred traffic out-converted the rest — roughly eight months behind the US. The same release gives two inconsistent cumulative-growth figures; neither is repeated here.
What nobody publishes: AI’s share of total retail traffic. Adobe gives growth rates only; Similarweb (September 2026) calls it “a relatively small traffic volume compared to search”, and reports 89% of shoppers who research with AI also use search engines.
The reading for CRO — interpretation, not a finding: AI-referred visitors behave like people arriving late in a decision — fewer bounces, more add-to-baskets, higher conversion. The assistant has done the comparison work the category page and product page used to do. Adobe reports the behaviour, not the cause.
In plain terms: the traffic is small but growing, and the people it sends have mostly made up their minds.
Feeds: Answer Engines · Channel Taxonomy · Multi-Touch Attribution
AI answers and organic click-through
Confidence: medium-high on direction; the size varies by study and query type.
- Pew Research (US, March 2025, browsing trackers on 900+ adults): users clicked a traditional result on 8% of searches showing an AI summary, against 15% without. Clicking a link inside the summary: 1% of visits
- Ahrefs (Dec 2023 vs Dec 2025, 300,000 keywords, desktop, correlational, informational-heavy): position-one click-through rate 58% lower where an AI Overview appears
- Seer Interactive (53 brands, Jan 2025–Feb 2026): organic click-through on queries with an AI Overview fell to 1.31% in December 2025 and recovered to 2.36% by February 2026, against 3.82% without an Overview. Being cited in the Overview roughly doubled clicks per impression against not being cited. Paid click-through held at about 13–17%
- Google’s AI Mode has been in the UK since July 2025
The figure to stop quoting: “AI Overviews cut click-through by 61%” — from Seer’s September 2025 study, superseded by its own update showing partial recovery.
Feeds: Search Engine Optimisation · Technical SEO · Answer Engines
Agents in the measurement
Confidence: medium. Mechanisms verified; agent volume figures are vendor and partly unverified.
The measurement problem is that the most capable agents look like people. ChatGPT’s agent sends an ordinary Chrome user-agent and runs JavaScript, so client-side tags fire and its visits enter session counts and conversion-rate denominators — as almost-never-converting sessions.
What exists for telling them apart:
- Signed requests. OpenAI signs agent requests using HTTP Message Signatures with a header identifying ChatGPT. Verifiable at the server or CDN (content delivery network) edge — not in client-side analytics, which never sees request headers
- An IETF standard in progress — the Internet Engineering Task Force’s web bot authentication working group has a first working-group draft (September 2026). A draft, not a standard
- Cloudflare uses the same signatures for verified bots, and from September 2026 blocks mixed-use AI crawlers by default on ad-supported pages for new and free-tier sites
- GA4 added an “AI Assistant” default channel in May 2026, classified by referrer — ChatGPT, Gemini, Copilot and others. It excludes Google’s own AI Overviews and AI Mode, which remain in Organic, and it doesn’t identify agents at all
- Adobe has you build AI-traffic classification yourself from user-agent, referrer and campaign parameters, and warns that some agents run limited JavaScript, so client-side data will be incomplete
Scale: HUMAN Security (vendor) put agents at 1.7% of AI-driven traffic in December 2025, with training crawlers at 74%. Small — but concentrated on exactly the pages where purchase decisions happen.
What this does to the numbers, each a known failure mode with a note in the vault:
- Agent sessions inflate the denominator, pulling conversion rate down with no change in human behaviour — the same shape as any bot problem — Bot and Internal Traffic · Sessionisation · Metric Decomposition
- Humans arriving from AI apps that strip the referrer land in Direct, hiding AI’s contribution and inflating Direct — Direct Traffic and Lost Referrers
- AI Overview clicks sit in Organic, so “AI traffic” in any report is an undercount
- Last-click attribution under-credits AI when most AI-researchers also search before buying — Attribution Models · Self-Reported Attribution
Feeds: Guide - Diagnosing a Metric Movement · Fingerprinting
AI in the testing tools
Confidence: high on what shipped; low on whether it works. Release notes are primary. Every outcome figure is from a vendor.
What has shipped, generally available unless marked:
| Vendor | What’s shipped |
|---|---|
| Optimizely (Opal) | agents that generate test ideas from past results, build variations in the visual editor, score the backlog, estimate annualised value, and check a new test against running ones before launch |
| Adobe | Experimentation Agent (Sept 2025) over Target and Journey Optimizer: summarises results, analyses impact, suggests opportunities — separately licensed. Experience Cloud rebranded “CX Enterprise” (April 2026), moving to credit-based pricing. Generative features in Target’s visual composer: status unconfirmed |
| Wingify (VWO + AB Tasty) | unified brand launched Sept 2026 with an embedded AI engine; the two products still run separately. VWO Copilot (ideation, variants, replay analysis); AB Tasty’s Evi agent (hypothesis to analysis) |
| Kameleoon | prompt-based experimentation — describe the change, it builds variant, targeting and metrics |
| GrowthBook, Convert | MCP servers, so a general AI assistant can draft, check and summarise experiments in the platform |
| LaunchDarkly | runtime switching of prompts and models, with experiments comparing them |
| Datadog Experiments (was Eppo) | relaunched April 2026, tying tests to observability data |
| Statsig | platform and customers to Amplitude (May 2026); engineers stayed at OpenAI. Its announced evaluation features aren’t shipped |
The pattern: AI is being applied to the cheap parts of experimentation — ideation, building, summarising — which is where time went, but not where tests fail. Nothing shipped addresses underpowered tests, peeking or post-hoc segmentation, and faster building makes the multiple-comparisons exposure larger, not smaller — The Multiple Comparisons Problem · Experimentation Velocity.
The evidence on outcomes:
- Optimizely’s analysis of 173,000 experiments (2026) reports customers using its AI saw +20.6% average test impact and +70% velocity against their own earlier baseline — and states itself that it is “observational, not causal”. Its agentic benchmark draws on its heaviest users: textbook Selection Bias
- Kameleoon and AB Tasty report build-time and velocity gains; neither reports win rates or effect sizes
- The one independent-ish study (arXiv, August 2026, pre-registered, 330 real A/B tests): a language model shown screenshots predicted the winner detectably above chance, but not reliably enough to act on. The authors sell synthetic testing, so not neutral. The same paper found 44% of “winners” in a leading CRO agency’s case library came from non-significant tests — a finding about the industry’s evidence base more than about AI
No independent evidence that AI-generated variants win more often or by more. Velocity is the only well-supported gain, and velocity without power is more false positives, faster — Win Rate and Expected Value.
The reverse direction — testing AI features — has converged on one pattern: offline evaluations screen candidate prompts or models; an online experiment, randomised by user, measures the business effect, with more variance to allow for because the outputs are non-deterministic — Evaluating Non-Deterministic Systems · Building with Language Models · Feature Flags.
Feeds: Test Prioritisation · Optimizely · GrowthBook · Adobe Target
What it means for the practice
Confidence: low. Interpretation. No credible survey of how CRO roles are changing was found.
- Discovery is becoming a conversion surface. If the assistant does the comparison, the inputs it reads — product data, structured data, reviews, clear specifications — decide whether you’re in the shortlist. Adobe reported in April 2026 that about a third of product pages couldn’t be read properly by AI. This looks like technical SEO and merchandising more than like landing-page optimisation — Technical SEO · Merchandising
- The on-site job shifts toward closing. Visitors arriving pre-decided need the price, delivery, returns and trust answered fast, not persuasion from the top — Trust Signals · Risk Reversal · Checkout Design
- Measurement literacy gets more valuable. Separating agents from people, and AI-influenced journeys from Direct and Organic, is now part of reading any conversion rate
- Statistical judgement gets more valuable as building gets cheaper. When a test costs minutes to build, the scarce skills are choosing what’s worth testing and refusing to believe a noisy winner
- Checkout moving off-site isn’t this year’s problem in the UK — but a CRO programme whose metrics assume every purchase happens on the site will eventually be measuring a shrinking share of them
Poorly sourced — don’t repeat
- “AI is 1% (or 2%) of retail traffic” — no primary source for any share figure
- Figures on how ChatGPT Instant Checkout converted, how many merchants went live, and OpenAI’s fee — secondary only
- “AI Overviews cut clicks 61%” — superseded by the same firm’s update
- Ahrefs’ “58% fewer clicks” quoted without “desktop, informational-heavy, correlational”
- Any AI-variant win-rate claim — none independent exists
- Salesforce’s “AI influenced 20% of holiday sales” — “influenced” is undefined
What I expect to be wrong
Recorded so the next snapshot on this topic can check:
- “Checkout stays on the merchant’s site.” This is the claim most exposed to one large launch — UCP checkout arriving in the UK, or a revived in-chat checkout — overturning it
- The conversion premium of AI-referred traffic. I expect it to narrow as the traffic broadens from early adopters to everyone
- Agent share of traffic. Small now; the number most likely to have grown, and the one analytics tools are least equipped to show
- “No independent evidence on AI-generated variants.” Someone will publish a controlled comparison; I expect a modest or null effect on win rate
- Wingify — I expect the two products to have merged into one, or one to be sunset
- UK regulation — I expect FCA consultation on agent payment authorisation and liability, and no CMA action beyond guidance
Feeds, durable concepts this touches: Answer Engines · Attribution Models · Bot and Internal Traffic · The Multiple Comparisons Problem · Experimentation Velocity · Checkout Instrumentation Constraints