It Turns Out AI Agents Have Unpredictable Shopping Habits
Increasingly, consumers and businesses are looking to AI agents to assist them in their purchases – go out and find the best prices and product fits. New research, however, suggests that what agents come up with after their shopping runs is unpredictable and inconsistent.
AI agents represent a new way to shop. At least 42% of millennials, for example, would trust an AI tool to make purchasing decisions for them for values up to $250, a study by RTB House finds. Overall, across all ages, leading AI tools, including Google AI Overviews and ChatGPT, “have overtaken social media and traditional media for shopper trust,” the study shows.
Similarly, 53% of shoppers trust AI tools, including an AI shopping assistant, as much as brand websites, according to a separate survey by Rithum and Retail Dive. In addition, 64% of shoppers – Gen Z in this case – say they’re likely to purchase based on an AI recommendation without verifying it anywhere else.
However, just how reliable is AI for decisioning for shopping choices? A new study led by researchers out of the University of Pennsylvania casts doubts on the consistency and predictability of agent-assisted shopping or purchasing.
The study shows that even minor changes to the search process can influence AI-generated product recommendations, “making agentic shopping decisions less consistent and harder to predict,” the study’s authors, which included Ethan Mollick, professor at the University of Pennsylvania, reports. "When LLMs saw no sources or additional context, each model had a strong preference for a single product."
The researchers conducted 200 runs of multiple frontier AI models, including Claude Haiku 4.5 and Opus 4.8, GPT-5 Mini and GPT-5.5, as well as Gemini 3.1 Flash Lite and Gemini 3.5 Flash.
With additions to the shopping prompts, and use of different models, agents’ recommendations changed, Mollick and his co-authors observed. “Two users issuing the same request, or the same user on a different day or a different model, may receive different products without any visible explanation." Inexplicably, one product may have been favored over another that was superior on price, rating, and review count.
In the course of e-commerce, “agents will likely encounter reviews, recommendations, search results, user memories, and other prior information before making a purchase decision," the researchers state. However, changing context may complicate the process.
For those companies or individuals selling online, this is a challenge, as they likely have little control over this process. This presents a challenge deeper than traditional SEO, they observe. It means "limited control rather than new leverage.”
That’s because “prior content clearly moves agent choices, but which way it moves them depends on the model, the combination of sources encountered, the order in which they appear, and even how the results are packaged into tool calls — none of which a seller observes or controls. A seller cannot know which model is shopping, what else it may have already read, or how its harness retrieves information.”
For buyers, “it means that LLM recommendations can shift without obvious reasons, making it harder to get consistent results,” the researchers add.
Another problem is agents designed for discovery and purchasing agents “are likely to underestimate real-world variability, since deployed agents will inevitably encounter a messier search process with prior content before making a purchase.,” the researchers state.
The net result: “as the number of moving parts in the search process increases, product recommendations become less consistent and harder to predict, with specific model and path dependencies. The overall search and recommendation process is unstable and hard to predict.”
Companies deploying AI shopping agents need to consider these inconsistencies in their design processes. “Unlike human shoppers who can be made conscious of promotional content, AI agents have no obvious mechanism to detect or discount strategically planted prior content,” the researchers state. They urge companies to “develop robustness testing that includes adversarial prior content and consider architectural interventions like transparent flagging of potentially influential user memory statements.”
Loading article...