Case study
When the Shopper Is Software: Making a Storefront Agent-Ready
How a direct-to-consumer brand stops being dropped from AI shortlists: price, stock, and terms published as facts a machine can read, one source of truth behind the page and the feed, and a person keeping the authority over what any agent is allowed to transact.
No client engagement behind this piece: this is how we would transform this category of product, with benchmark-sourced targets.
Who this is for
A typical direct-to-consumer brand selling 300 to 2,000 SKUs
Industry
E-commerce and retail, 5-50 staff
Legacy stack
Engagement
Concept
Contents
A shopper asks an assistant to find them a trail running shoe under two hundred dollars that ships this week. The assistant reads a handful of product pages, builds a shortlist of three, and hands back two names. This brand makes a better shoe than either of them, priced in the middle, in stock, with a sixty-day return policy that beats both. It was not on the list. Its price lives inside the hero image, its stock is signalled by a greyed-out button, and its returns policy is a PDF, so the assistant recorded three unknowns and moved on.
This is the transformation we would run for that brand. It is presented as a concept, built on the same method we use in client engagements, so the before, the after, and the path between them are all visible.
What being unreadable costs a storefront
The infrastructure for this arrived quickly. Stripe and OpenAI published the Agentic Commerce Protocol in September 2025, an open Apache-licensed standard for checkout between buyers, agents, and businesses, launched with ChatGPT Instant Checkout. Google announced the Universal Commerce Protocol on 11 January 2026, co-developed with Shopify, Etsy, Wayfair, Target, and Walmart.
The buying behaviour is on a slower clock, and pretending otherwise would make this study useless. Bain & Company’s forecast of 17 December 2025 puts US agentic commerce at 300 to 500 billion dollars by 2030, or 15 to 25 percent of e-commerce, and observes that most consumers are still uncomfortable letting AI complete a purchase end to end. What is already routine is the step before: Bain puts 30 to 45 percent of US consumers using generative AI for product research today.
That is the part worth acting on now. Long before an agent buys anything on your behalf, it is building the shortlist, and a brand whose price cannot be resolved never enters the comparison.
Be careful about how much is claimed here, because the elimination happens without a click and therefore never appears in an analytics report. That cuts both ways: a brand cannot blame a soft quarter on this, any more than it can prove the loss. The only honest response is to measure the channel directly instead of inferring it from revenue, with the extraction check described below and the probe method for AI visibility run alongside it.
Before and after: what an agent can resolve
Drag the handle. The before is the recreated extraction from a normal product page; the after is the transactable surface designed in its place.
Transactable surface · one source of truth
Nothing was clicked, so nothing was recorded.
This comparison never appears in your analytics.
The page a person sees barely changes. What changes is that the facts behind it move out of images, PDFs, and script-rendered elements and into text and markup, and that one set of numbers now feeds the page, the feed, and any checkout an agent calls.
What we built
The concept above is not a mockup exercise. It is the output of the same engagement steps we run on real operators, including the parts we argue against:
- We start with the eight fields that decide a shortlist, not with the catalog: what it is, price, availability, delivery, returns, how to buy, who it suits, and how to reach a person. One category first, so the pattern is proven before it is repeated across 2,000 SKUs.
- We build one read model behind them, so the product page, the structured feed, and the checkout all resolve from the same source. The alternative, syncing three systems that each hold their own copy, is the arrangement that produces a page saying one price and a feed saying another, and it is the most common way this work goes wrong.
- We publish stock with a deliberate buffer rather than a live count. Under-promising to an agent costs one sale; overselling costs a refund, a support thread, and a public review. A more aggressive engineer would expose the live number, and that is a defensible position we disagree with at this stage of the technology.
- We leave pricing authority un-automated on purpose. Nothing in this system writes a price. The rules layer publishes a floor and a discount limit that a person sets, and the checkout enforces the floor rather than trusting whatever the page last rendered.
- We wire the reconciliation before opening anything: a daily job replays the previous day’s agent orders against the price and stock truth and flags divergence for a human. Until that exists and has run clean, the checkout stays closed to agents and the work stops at publishing facts.
- We refuse the volume play. No generated per-agent landing pages, no thin SKU-variant pages written to feed a crawler. That is the content-farm tactic in new clothes and it degrades the storefront for both audiences.
Week one is smaller than most people expect: audit one category’s product pages with scripts disabled, agree the eight fields, and publish them as text and structured data on about twenty SKUs. The metric for that slice is the extraction check, which scores how many of the eight a machine can resolve without running anything.
Two things are worth saying plainly before anyone budgets for this. Most mainstream platforms already emit a product feed, so the pipe usually exists; the work is almost never plumbing, it is that the facts going through the pipe are missing, stale, or disagree with the page. And the cost to the brand is not only invoiced hours: it needs one person who genuinely owns product data for roughly half a day a week in the first month, because deciding which of three recorded lead times is the true one requires someone with the authority to decide, and that cannot be outsourced to us.
How it works
Three layers sit under the storefront. Catalog facts carry price, stock, and specifications. Policy facts carry returns, shipping, and warranty. Authority rules carry what may be sold, at what floor, and to which agents. An agent checkout reads all three, and the order it produces carries the same price, the same stock truth, and the same policy a person would have received, because there is only one set of numbers to read.
The layer that matters most is the last one. An agent never sets its own terms; it reads the ones the brand published. That is not a convention we invented, it is how the protocols are specified: the ACP specification lets a merchant accept or decline on a per-agent, per-transaction basis, and decide which products can be sold at all. The design question is not whether to keep control, it is whether the brand has written its rules down clearly enough to enforce them.
Nothing gets ripped out. The platform stays, the design stays, the team keeps its tools. What is added is a layer of stated facts and a rules file that a person owns.
Why this pays back
The return does not depend on the forecast being right, and that is the whole reason to start now. Every fix in this work also helps a human in a hurry, on a phone, on a poor connection, or using a screen reader: a price you can read before checkout, a return window stated in a sentence, a spec you do not have to open a tab to see. If agent-completed purchases stay marginal for another three years, the brand is left with a clearer, faster, more accessible storefront and a product catalog that finally agrees with itself. That is a project worth doing on its own merits, wearing a trend’s clothes.
The honest limit is sharper than the usual one. This makes the brand comparable, not preferable. Clarity gets you into the comparison and then shows the buyer exactly where you stand, so a brand whose price and terms are genuinely worse will lose faster and more visibly than it does today, when confusion was hiding it. That is the correct outcome and it is not always a welcome one. If the numbers do not survive daylight, fix the offer before you fix the markup.
The outcomes
US agentic commerce by 2030
$300-500B1
Industry benchmarkUS consumers already using generative AI for product research
30-45%2
Industry benchmarkProduct facts an agent can resolve without running scripts
4 of 8 fields 8 of 8 fields
Unknown to stated3
Design targetTime to the first shippable slice
4-6 wk4
Design target1 Bain & Company, 17 December 2025, 2030 Forecast: How Agentic AI Will Reshape US Retail: 300 to 500 billion dollars of US agentic commerce by 2030, or 15 to 25 percent of e-commerce, counting purchases initiated, influenced, or completed by AI agents and excluding journeys that only use AI-assisted search or discovery.
2 Bain & Company, 17 December 2025, same forecast. Bain also notes most consumers remain uncomfortable with end-to-end AI transactions today, which is why the near-term value sits in the research and shortlisting stage rather than in checkout.
3 Design target for the transactable surface, measured with a scripts-disabled extraction check against the recreated product page shown in the comparison above.
4 Standard first-slice scope: one workflow, one metric, evals and an approval gate included. Here that is one product category, its eight fields, and the extraction check that scores them.
Frequently asked questions
Is agentic commerce real enough to spend money on in 2026?
The plumbing is real and recent: Stripe and OpenAI published the Agentic Commerce Protocol in September 2025 alongside ChatGPT Instant Checkout, and Google announced the Universal Commerce Protocol on 11 January 2026 with Shopify, Etsy, Wayfair, Target and Walmart. The buying behaviour is slower: Bain's December 2025 forecast reaches 15 to 25 percent of US e-commerce only by 2030. So the answer is yes for the data work, which pays back on its own, and no for a platform migration, which does not.
What actually gets built here, in plain terms?
One source of truth for the facts a buyer needs, published in three places that can never disagree: the product page as readable text, the structured feed, and the checkout an agent can call. Plus a rules layer that says what may be sold, at what floor price, and which agents are allowed to transact. Roughly four to six weeks for the first product category, then the same pattern repeated.
What stops an agent selling at the wrong price or overselling stock?
Three things, deliberately layered. A price floor enforced at checkout rather than trusted from the page, so a stale cached price cannot become a stale sale. A stock buffer, because under-promising to an agent costs a sale while overselling costs a refund and a review. And a daily reconciliation that replays the previous day's agent orders against the price and stock truth and flags any divergence for a person. If those three are not in place, the checkout stays closed to agents.
Do we have to open our checkout to AI agents at all?
No, and most brands should not start there. The first phase publishes facts, which affects whether you appear in a shortlist at all, and is reversible in an afternoon. Opening a checkout is a separate decision with its own risk, and the protocols are built for that caution: the ACP specification lets a merchant accept or decline on a per-agent and per-transaction basis. The authority stays with the merchant by design.
How much of our own time does this take?
More than most vendors admit. One person who genuinely owns product data, roughly half a day a week for the first month, because the hard part is not the engineering: it is settling which of the three conflicting versions of a fee, a lead time, or a return window is the true one. That work has to be done by someone with the authority to decide, and it cannot be outsourced to us.
Go deeper on the method
The Agent-Ready Website: What Breaks When Your Visitor Is Software
A second kind of visitor is reading your site: one that does not scroll, does not wait for your scripts, and never emails to ask what something costs. It takes a record and leaves. Here is what it can and cannot see, why most sites fail at the same step, and the fixes that pay off whether or not the agents ever arrive.
Read articleAnswer Engine Optimization: How Customers Find You When AI Answers First
Your next customer is asking ChatGPT, not scrolling Google. The answer cites two or three sources, and either you are one of them or you are invisible. A scorecard shows whether an answer engine would cite you today, and the fixes are more honest than any SEO trick.
Read articleHow to Measure AI Search Visibility When There Is No Rank Report
Buyers ask an assistant now, and no rank tracker can see inside that conversation. The honest instrument for AI search visibility is a sample: ten buyer questions, asked every month, scored on one axis. Build your probe set here and score your first run in about fifteen minutes.
Read articleMore transformations
AI Claim Triage From Photos: First Notice of Loss to a Reviewed Assessment
How a property claims operation turns a folder of unsorted phone photos and a 62-page policy into a structured damage assessment: every line citing the photo it came from, the coverage clause quoted, uncertain lines flagged, and an adjuster approving before anything moves.
32.4 days
Average property claim, filing to finished repairs
AI Quote Desk for a Wholesale Distributor: RFQ Inbox to Priced Quote in Minutes
How a parts distributor replaces the morning RFQ pile and the three-system price hunt with a copilot that extracts every line item, prices it from the company's own item master and contracts, checks stock, and drafts the reply, with a rep approving every quote before it leaves.
60-70%
Share of routine admin work generative AI can absorb
Get the next transformation in your inbox.
When we publish something worth your time, you will be first to know. No spam, unsubscribe anytime.