Why is it difficult for AI shopping agents to become widely adopted?

By Anderl

Translated by Chopper, Foresight News

At the moment, there’s a common narrative circulating in both the AI and crypto industries: equip AI agents with wallets so they can fully handle shopping on your behalf—this is said to be AI’s most core, practical use case. The pitch sounds neat and flawless, as if it were the future. But at its core, the logic doesn’t hold up. It muddles the hard parts of shopping with the simple ones, placing the most basic payment step at the center.

Let’s set payment aside first and return to shopping itself.

Two core behaviors in shopping

At its essence, shopping consists of two categories of actions that are now tightly coupled: information retrieval and value judgment. Retrieval (gathering, filtering, comparing, and performing initial ranking) has standardized characteristics and can almost be fully delegated to machine agents. Value judgment (whether the item is good, whether it fits you, whether the merchant is trustworthy) is the part deeply bound to human subjective emotion.

Data has already shown that information retrieval is rapidly shifting toward AI. Adobe Analytics data shows that from July 2024 to July 2025, traffic routed to U.S. retail websites by generative AI surged by about 4700%. But the “AI wallet” narrative assumes that intelligent agents can simultaneously handle both retrieval and value judgment—this is the key conceptual sleight of hand. Retrieval can be fully handed over to machines, while value judgment can only be partially delivered under specific conditions; in most scenarios, it cannot be delegated at all.

Value judgment is also divided into two layers

More importantly, value judgment itself is not a single dimension—it is divided into two parts. One part is evaluation: testing options against a utility function. The other part is demand definition—meaning the utility function is set first: which dimensions matter, what the weights are, which values impose constraints, and what “good” ultimately means.

Demand definition is not something completed once at the very beginning; it runs through the entire shopping process. Compliance standards for goods are decided by you; whether damaged zippers mean you abandon the item depends on your judgment criteria; choosing which merchant depends on the value orientation you care about. Each filtering layer follows the logic of “human subjective criteria × AI machine evaluation.” Automation can only replace the evaluation step; the sovereignty of defining demand always remains with humans.

Many people wrongly think humans only need to write a list of standards once, and then they can fully let go—this severely underestimates human decision-making logic. A large body of research in decision science proves that people’s preferences are not fixed. Psychologist Paul Slovic proposed the constructive preference theory: there isn’t a ready-made set of fixed likes inside us waiting to be retrieved; preferences form gradually during the process of making choices. The “preference reversal” experiment confirms this: when using two different but equivalent methods—one based on choice and the other on pricing—people produce completely different rankings of the same goods, violating the foundational axioms of rational choice.

Ariely, Loewenstein, and Prelec’s 2003 “coherent arbitrariness” theory: even random numbers with no relation (like the last few digits of your Social Security) can anchor people’s psychological bids for ordinary products; moreover, this anchoring effect does not disappear with consumption experience or market transactions. So-called “stable preferences” are merely an illusion of order constructed by humans.

Therefore, human-AI interaction can’t be a one-time filling-out of a standards checklist; it needs continuous iterative communication. At each filtering node, an AI agent asks targeted questions like: “You previously valued durability—once the premium reaches what price level does durability stop being worth it to you?” Then humans define that boundary in real time.

The real dividing line: is procurement a chore, or is it enjoyment itself?

The industry often separates scenarios using “standardized products / personalized products.” The disagreement isn’t whether standardization can be quantified. The real dividing line is whether the act of making a choice itself carries experiential value.

For items like printer paper, batteries, or goods that need regular restocking, the selection process has no experiential value. Nobody wants to spend effort comparing two ink cartridges that are essentially identical. These products naturally fit a model where AI agents handle everything completely; automatic repurchasing won’t sacrifice any experience.

For enjoyment-based consumption, the situation is the opposite. Wine, furniture, coats, books—the act of picking them is itself part of the fun of consumption. If you hand judgment to a machine, you save time costs, but you directly strip away the core pleasure of consumption. Even if AI offers free Q&A end to end, people still won’t fully delegate.

For enjoyment-based goods, AI agents should not make autonomous full decisions; instead, they should switch to being information collectors: complete retrieval, do initial filtering, match parameters, verify merchant credentials, extract common product issues from massive reviews, reduce 200 options to 5, and then leave the final choice to humans.

A threefold dilemma: AI questioning, historical simulation, and autonomous decision-making

Someone might say: “Why not just have AI ask directly about my evaluation criteria?” But that is exactly the inefficient low-effect pattern most ordinary people dislike in daily life. More dangerously, repeatedly interrogating people about their criteria distorts their final choices.

In 1991, Wilson and Schooler conducted a jam-tasting experiment. In the group where participants were asked in advance to organize the reasons behind their preferences, the final rankings deviated more from professional tasting standards. Follow-up experiments also showed that forcing people to list the reasons for each choice results in even lower post-choice satisfaction for decorative paintings. Language can only describe easily expressible surface features; it can’t capture true internal preferences. Sensory experiences like taste and aesthetics can only be perceived—not defined precisely in words.

This creates a threefold dilemma that can’t all be optimized at once:

Directly asking the AI: closely matches real preferences at the moment, but creates very high interaction friction, even distorting choices;

Simulating based on historical behavior: smooth operation, but it locks people into past preferences and kills new consumption exploration;

Humans making autonomous judgments end to end: fully preserves choice rights, but costs a lot of time and effort.

A fourth compromise can avoid the above problems: identification-based interaction, not one-by-one questioning. AI directly presents 3 options; humans just choose directly. This method is both convenient and accurate because it doesn’t require abstractly naming standards that you often can’t even name.

The classic choice overload experiments can back this up. In 2000, Iyengar and Lepper ran a jam tasting stall experiment: when 6 jam options were presented, conversion rates were far higher than when 24 jams were presented. However, the theory is controversial; multiple follow-up analysis experiments overturned this conclusion. After aggregating many experiments, Scheibehenne et al. in 2010 found no overall universal choice overload effect. Analytical coverage by Chernev, Böckenholt, and Goodman in 2015 across 99 studies showed that the overload effect only appears when product complexity is high, decision difficulty is high, and a person’s own preferences are vague. The original authors’ later retrospective also noted that when facing 24 options, consumers lacked enough time to organize their own preferences. True autonomy exists between “delegating the agent to apply my standards” and “having the agent fabricate standards out of thin air based on my history.”

Back to “AI wallets”: payment is only the least important part

Once these logic chains are clarified, the narrative gap in “equipping AI with wallets” becomes obvious. This story mixes up three completely independent things: the decision-making party, the execution party, and the funds-holding party. “Equipping AI with a wallet” only solves the funds-holding problem. Wallet custody only matters when the AI also has decision-making power.

There are three cases. In the first case, humans decide and pay themselves. In this situation, the agent doesn’t pay—it acts as a scout. In the second case, humans make the decision and delegate execution to the agent (“Yes, buy that”). Then the agent handles checkout, but it doesn’t need to custody funds; it only needs a limited, revocable authorization for this already approved purchase. Only in the third case—when the agent decides and pays autonomously, with no human on-site required for checkout—does the wallet itself take responsibility for payment.

Interestingly, the payment industry has already implemented layered authorization schemes globally in 2025, clearly separating “authorization” from “funds custody”:

OpenAI teamed up with Stripe to launch a smart agent commercial agreement, issuing shared payment tokens bound to a single merchant, a fixed amount, and a time-limited one-time use; the AI cannot obtain the full card number;

Mastercard released Agent Pay in April 2025, generating specialized tokens limited to a defined agent, specified merchant, and user-authorized rules;

Google launched the AP2 smart payment protocol in September 2025, clearly splitting “user demand authorization credentials” and “AI procurement list credentials.” Both are verifiable cryptographic credentials, perfectly matching the layered logic of “demand definition / machine evaluation” above;

Visa launched a trusted agent protocol in October 2025, based on the same approach as the above.

Major payment providers are, in their own ways, demonstrating the core viewpoint of this article: you don’t need to hand funds to AI—just grant limited operational permissions.

So, where are the real applicable scenarios for an AI independently custodial wallet? Consumer retail scenarios almost don’t require AI wallet custody. The real space for this solution is in standardized bulk commodities and automated settlement between machines. Coinbase and Cloudflare jointly introduced the x402 protocol, filling gaps in traditional card payment rails: it supports automated inter-agent payments with no human intervention, 24/7 operation, and API-call-based billing. Within months of launch, transaction volume exceeded 100 million. Mastercard simultaneously rolled out Agent Pay for machines, serving high-frequency, low-latency small machine settlements. This is the underlying infrastructure for machine economies, not for individual shopping.

This narrative itself has no factual errors, but its importance is completely inverted: an independently custodial AI wallet has meaningful value only in scenarios where goods are highly homogeneous and per-transaction amounts are relatively low.

Where wallets should truly go

This also explains the shift in emphasis in the crypto space over the past two years: it no longer mainly promotes a narrative of “freeing consumers,” and instead focuses on institutional underlying infrastructure—stablecoin clearing, tokenization of assets, and enterprise-grade services. This isn’t abandoning the AI shopping lane; it’s returning to the area where wallet custody is actually a fit: the enterprise side. Corporate procurement departments are themselves an institutional embodiment of “procurement as pure chores.” Procurement processes strip away subjective aesthetics and personal feelings, relying entirely on specifications, price, performance/fulfillment, and contract terms for decision-making—essentially an employee version of intelligent agent procurement. An AI wallet only automates workflows that enterprises have already matured; it doesn’t overturn behavioral patterns.

Today, many enterprises outsource office supplies and low-value consumables. Platforms like Mercateo and Amazon Business compete not primarily on low prices, but on reducing process costs: unified procurement catalogs and consolidated billing. Enterprises are willing to accept small premiums on individual items in exchange for a dramatic reduction in procurement labor costs; for low-value consumables, process costs are often higher than the value of the goods themselves. AI agents can reduce the ordering labor cost to nearly zero while also searching fragmented massive supplier networks. Procurement platforms only keep compliance verification functions: vendor onboarding reviews, unified reconciliation, and anti-counterfeiting checks—yet verification is exactly the real bottleneck in the end-to-end process.

So the deployment strategy can’t be summarized simply as “wallet services for enterprises, not individuals.” The precise formulation should be: tools matched to scenarios.

Self-custody wallets (AI autonomous decision-making, holding funds, completing payments): suitable mainly for enterprise procurement of standardized goods and automated settlement between machines, focusing on high-frequency B2B/M2M repeat purchase scenarios;

Personal consumption scenarios: no need to custody funds; instead, grant one-time, limited-scope payment tokens, and only after humans confirm the order can the AI complete settlement.

“Equipping AI with a wallet” as a C-end promotional headline targets only a tiny slice of scenarios. The truly core deployment market for this approach is the enterprise back-end automation system.

Additional note: enterprise procurement isn’t all standardized goods. Some procurement decisions also can’t be handled by AI autonomous custody and decision-making. Choosing law firms, acquisition targets, and core suppliers are strategic decisions—subjective consequences are significant. Just like choosing a bottle of wine personally, these must be decided by humans, and self-custody wallets have no room to help in such scenarios.

This rule applies everywhere: independently custodial AI wallets create value only at the standardized goods layer, while standardized business concentrates inside enterprises—where the biggest scale and the densest demand are found.

The real bottleneck has never been payments

The technology for moving funds is already mature; there is no bottleneck in the payment step. The real bottlenecks for AI shopping deployment lie in the other two areas.

First, the lack of trustworthy data sources. The premise for automated judgment is that data must be truly reliable; once underlying information is distorted, autonomous-decision AI will amplify the negative effects of false information at machine speed. The proliferation of fake reviews is already a public industry problem. In 2024, the U.S. Federal Trade Commission issued new rules (effective October 21) explicitly prohibiting fake product reviews. It also clearly lowers the threshold for generating large batches of fake reviews using generative AI. For violations, fines per single instance can be up to $51,744, and by the end of 2025, regulators had issued multiple warning letters.

The problem of counterfeit physical goods is also severe. According to 2025 data statistics from the OECD and the EU Intellectual Property Office, in 2021 the global counterfeit trade volume was about $467 billion, accounting for 2.3% of total global trade; in the EU, counterfeit imports accounted for 4.7% of imported volume. Apparel, shoes/bags, and luxury goods are the worst-hit categories—precisely the kinds of experience-based consumption products. Product-level traceability documents, verifiable genuine reviews, independent third-party authentication, and product-level transfer/record storage (the EU anti-counterfeiting directives and the U.S. drug supply chain law have already been forcing such documentation on medicine packaging) are prerequisites for safe AI judgment.

Second, human demand-definition authority cannot be automated. As long as the demand criteria are defined by humans, subsequent filtering, comparison, and settlement can be automated. But defining your own demand can never be handed to a machine. If AI generates the demand standards for you, the resulting preferences do not belong to you.

Equipping AI agents with wallets only solves the simplest funds-transfer step. The direction worth deep investment is safely and controllably automating filtering and evaluation, while returning the two core rights to humans: defining evaluation criteria and enjoying the pleasure of making the final choice.

Summary

Finally, it’s important to emphasize that for experience-based consumer goods, once procurement channels become “commoditized,” the product is everywhere. At that point, your product isn’t the product anymore—choice is. Platforms need to optimize their own information standardization level so AI can retrieve effectively: complete, verifiable traceability; clean, structured product data; and guidance that funnels users to the platform. On that basis, preserve the core advantage that automation cannot replace, broaden users’ aesthetic boundaries for entirely new consumption exploration, and ultimately deliver the satisfying experience of selecting the right product. Let AI agents handle information collection, and focus the platform’s core competitiveness on creating a better “choice experience.”

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned