Tony Xhufi, founder of WEM, argues that the race to let AI agents shop has moved faster than the evidence layer underneath the prices they rely on. Ask an AI assistant what a pair of headphones costs and you will have an answer in about two seconds. A model, a shop, a price. £229, at a retailer you have heard of.
It reads like a fact. It is presented like a fact. Most of the time, it probably is one. But consider what you cannot see.
You cannot see when that £229 was last checked an hour ago, or last Tuesday. You cannot see what evidence sits underneath it: whether somebody checked the retailer’s own page, or whether the number came from product information distributed elsewhere. You cannot see whether availability was actually checked.
And you cannot always see whether the headphones being priced are exactly the product you asked about, rather than a different model, specification or market variant. Today, if the number is wrong, you click through and find out. Mildly annoying, and usually that is the end of it.
That is changing.
When The Assistant Stops Advising And Starts Buying
The race is no longer just to recommend products. It is to let software complete the purchase. OpenAI, Google, Stripe, Shopify and others are building the discovery, checkout and payment infrastructure that makes AI-native commerce possible. That makes the quality of the information immediately before a transaction more important, not less.
An agent that buys changes the consequence of a wrong commerce fact. It stops being merely a wasted click and starts raising harder questions around authorisation, evidence and accountability. If something goes wrong, the questions become much more concrete. What did the agent see?
What did the user authorise? What did the merchant commit to? And when did each of those things happen?
Those questions will matter increasingly as assistants move from recommending products to acting on behalf of consumers.
What Happened When We Tested Ourselves
We recently published the rules WEM has to satisfy before it can describe a retail offer as verified. Not a mission statement, specific conditions, some of which deliberately make it harder for us to show a result.
Then we converted those rules into tests against our own system.
They exposed inconsistencies between what we intended the system to do and what it actually did in production, including how missing availability information was handled and a case where public behaviour did not match what we had documented. We corrected them.
That experience reinforced the point.
A rule that lives in someone’s head is an intention. A rule that someone else can test is an accountability mechanism.
We did not discover those problems simply by promising to be careful. We found them because we published something we were capable of failing.
Three Rules In Plain English
- Say nothing rather than something plausible – If an offer does not have sufficient evidence to qualify as verified, it should not be presented as verified. That sounds obvious, but every instinct in comparison pushes the other way. A missing result can look like a broken product, while a plausible estimate makes the experience appear complete. The problem is that once estimates and observations look identical, the shopper has no way of knowing which is which. An offer can still be useful for discovery without being verified. Those are different things.
- “We don’t know” is not the same as “yes” – Availability has three states, not two: reported available, reported unavailable and not reported. If availability was not checked, the answer should remain unknown. This sounds like a technicality, but small assumptions made by software are capable of becoming large claims when they are repeated at scale.
- Every price needs a time attached to it – A price is not a permanent fact about the world. It is something observed at a particular moment. “Checked four hours ago” is weaker than “live”. It is also more honest. Retailers can change prices at any moment, and the retailer remains authoritative at checkout.
More from Artificial Intelligence
- Why Is OpenAI Pre-Announcing Astra’s Safety Limits Instead Of Its Capabilities?
- How Does Machine Learning Work? The Five Algorithms You See Every Day
- Stop Losing Money By Ignoring AI: Advice From $1B Company Founder Shane Morand
- Lunar Cyber Launches Token Exposure Monitoring As Infostealers Target Developer And AI Credentials
- Malware Is Now Stealing Claude Sessions To Drain Paid AI Usage – How Does That Work?
- 5 Of The Most Promising Humanoid Robotics Companies Changing The Industry
- What Does It Mean To Be AI Agnostic?
- AI Is Hungry For Power, Is Nuclear The Answer?
The Feed And The Observation
There are two kinds of price commonly used in online commerce, and the word “price” hides an important difference.
An advertised price is supplied for distribution, for example through an affiliate, comparison or promotional product feed. Those feeds are genuinely useful. They help discovery systems understand what retailers say they sell.
An observed price comes from a separate observation of the retailer’s customer-facing or transactional surface at a stated moment. A feed is evidence of what a merchant said. An observation is evidence of what a shopper could actually see. Both ultimately originate with the retailer.
The distinction is not that one is retailer-controlled and the other is not. The distinction is that they are separate evidence paths. Distributed product information and what a shopper actually sees can drift apart. If all you possess is the first representation, you cannot independently detect that difference.
Verification requires a second look
We Get Paid When You Buy
WEM (https://wem3.ai/shop) earns affiliate commission when somebody buys through certain retailer links. We disclose that because neutrality is not the same thing as having no business model. Equally, saying that a service does not take affiliate commission does not, by itself, establish neutrality.
Could The Money Change The Answer?
At WEM, commercial terms are separated from verification and offer-ranking decisions. Commission does not determine whether an offer qualifies as verified or where it appears. The principle is straightforward: commercial incentives can exist without being permitted to determine the evidence.
The Piece That Is Missing
None of this is an argument that the industry has built nothing. It has built a great deal, very quickly.
Schema.org has long given the web a common vocabulary for describing offers. Google’s Universal Commerce Protocol and the Agentic Commerce Protocol developed by OpenAI and Stripe are helping build interoperability across AI-native discovery, commerce and checkout. Model Context Protocol, introduced by Anthropic, gives AI systems a standard way to connect to external tools and data. The Agent Payments Protocol goes further on authorisation and accountability, cryptographically binding important parts of an agentic transaction.
All of that is meaningful progress. It also answers a different question.
These mechanisms can help establish what was selected, what was authorised and what transaction was ultimately executed. They do not, by themselves, establish that the discovery-stage commerce fact which caused an agent to favour one product or merchant over another had first been independently observed.
You can therefore have a robust record of an authorised transaction while still asking a more basic question about the information that led to it.
Was The Commerce Fact Supported By Evidence When The Agent Relied On It?
That is the layer I think is missing. Not how to buy.
Not how to prove that somebody authorised the purchase. But what evidence should exist underneath a commerce fact before software states it and eventually acts on it.
Go And Check
We have published our own answer as the WEM Verified Offer Standard.
The standard sets out the minimum evidence WEM requires before describing a retail offer as verified. It is versioned, openly accessible and designed so that the resulting claims can be challenged rather than simply trusted.
We are not claiming to certify anybody else. We are not claiming that WEM sees every price on the internet.
And we are not claiming that a verified observation guarantees what the retailer will charge when somebody eventually checks out. The point is smaller, and more stubborn.
If an assistant is going to state a price as fact and increasingly spend a consumer’s money on the strength of that fact, there should be evidence underneath it that somebody other than the assistant can inspect.
Agentic commerce does not only need better ways to buy. It needs better ways to know what is true.
Tony Xhufi is the founder of WEM, a UK AI-powered commerce discovery and price-intelligence platform building verification infrastructure for assistant-led commerce. WEM is funded by affiliate commission, disclosed on pages carrying retailer links.
