New Research: AIs are highly inconsistent when recommending brands or products

Paper · Source
Knowledge After the Web

Source: SparkToro / Gumshoe.ai · 2026-01-28

The Problem: For the last few years, companies have been investing inordinate sums into AI tracking and AI visibility for their brands and products. $100M+/yr is already estimated to be spent on this new version of search analytics and yet, I could find absolutely no research showing whether AI tools are consistent enough, when prompted for lists of brand/product recommendations, to produce valid visibility metrics. There are loads of studies about AI accuracy (in fact, we used Carnegie Mellon’s Estimating LLM Consistency: A User Baseline vs Surrogate Metrics as a model for this work), but when it came to answering the question of whether ChatGPT, Claude, or Google AI produce consistent enough lists of results to be usefully tracked... Nothing.

Our Experiment’s Hypothesis: AI tools produce such randomized lists of recommendations, and user prompts are so highly varied that attempting to track rankings or visibility of a brand/product for a given topic space or user intent is pointless (and brands with too much money should just pay ChatGPT, et al. for the impressions data in their upcoming ads products).

Step One: Ask a bunch of folks to run the same AI prompts over and over, recording their results.

We chose the three most popular AI tools in the US: ChatGPT, Claude, and Google search’s AI Overview (and AI Mode, if/when AI Overviews didn’t show). 600 volunteers ran 12 different prompts through each of the 3 tools a combined 2,961 times. They copied and pasted the AI tool responses into survey forms, which Patrick (and my new Chief of Staff, Kristy Morrison) normalized into ordered product/brand results.

What you’re seeing with the green (ChatGPT), orange (Claude), and blue (Google AI) bars is the number of unique brands, products, and entities recommended by the AI response to the prompts. The range appears to be all over the place, but is in fact closely tied to how many entities frequently appear in documents in the AI corpora around the topic (e.g. there are less than a dozen Volvo dealerships in Los Angeles, but thousands of recently-published Science Fiction novels). Those pink scatterplot-style dots, meanwhile (graphed on the secondary axis) show the average number of responses the AI tools gave (another factor complexifying the ranking/visibility problem).

If you ask an AI tool for brand/product recommendations a hundred times nearly every response will be unique in three ways:

To get mathematical about it, there’s a <1 in 100 chance that ChatGPT or Google’s AI, if asked 100X, will give you the same list of brands in any two responses. Claude is just slightly more likely to give you the same list twice in a hundred runs, but even less likely to do so in the same order.

In fact, when it comes to ordering, AI tool responses are so random that it’s more like 1 in 1,000 runs before you’d see two lists in the same order. And we didn’t even try to collect data on how the AIs described each brand or how positive/negative sentiment was around the recommendation.

The bottom line is: AIs do not give consistent lists of brand or product recommendations. If you don’t like an answer, or your brand doesn’t show up where you want it to, just ask a few more times.

Even if rankings are random to the point of near-uselessness, appearances across dozens or hundreds of runs of the same prompt indicates a set of brands that the AI’s system generally associates more (or less) as a good answer for the prompt intent. Measuring that percent visibility is (probably) a reasonable way to know how prominent or invisible your entity is within the AI’s consideration set.

Another example — men’s fashion influencer Adam Gallagher only came up 36X in the 73 responses Google’s AI gave when asked for recommendations in this space.

Let’s briefly return to that example of West Coast cancer care hospitals. In ChatGPT’s results, City of Hope hospital in Los Angeles showed up in 69/71 answers: a 97% visibility rate.

Across 142 responses, there were barely two prompts that, if you squinted, looked similar to me at all. This taught me that AI prompting is nothing like searching Google. People don’t reduce their search intent to the fewest, most logical set of 2-5 keywords, they get creative, and weird, and highly specific.

Overall, semantic similarity of these prompts was 0.081. Or, in the language of recipes, it’s like the average pair of prompts were Kung Pao Chicken and Peanut Butter — key ingredients had overlap, but other than being foods with peanuts in them, they’re not especially close.

Running 142 human-crafted prompts about choosing the best headphones for a traveling family member multiple times resulted in nearly a thousand responses: 994 to be precise. And across those 994 AI responses, headphones like Bose, Sony, Sennheiser, and Apple showed up 55-77% of the time; remarkably similar to what we saw with top 3 brands in the tighter spaces from our survey-takers’ results in topics like L.A. Volvo dealerships, cloud computing providers for SaaS startups, and cancer care hospitals on the West Coast.

When Gumshoe ran their analysis showing the visibility percentages and rankings for the headphones prompts, the top brands often had 90-100% visibility, while the brand design agencies prompt produced high numbers in the 30%s-40%s. Note that they also make far prettier, easier-to-read charts to synthesize this data than I did above (apologies, I ran out of time to do more sophisticated data visualization).

All of this reinforces the notion that visibility percent, across loads of prompts, whether written by humans or generated synthetically, are likely to be decent proxies for how brands actually show up in real AI answers.

These tools are probability engines: they’re designed to generate unique answers every time. Thinking of them as sources of truth or consistency is provably nonsensical.

Users almost never craft similar prompts, even when they have the same intent. The variation of brands/recs in AI answers around a space in the messy wilds of AI prompting is likely much higher than what our controlled experiments revealed here.

Measuring your brand’s presence in AI answers with precision is a fool’s errand. You can, with enough prompts run enough times, get a dartboard-pattern-like answer comparing you with others. I’ve been swayed from my initial position and now believe visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric.

But, any tool that gives a “ranking position in AI” is full of baloney.

If you do nothing else after reading this report, please, please, marketers, analysts, and execs: stop throwing money at AI tracking products that don’t provide stats-backed, publicly-reviewable research. Before you spend a dime tracking AI visibility, make sure your provider answers the questions we’ve surfaced here and shows their math.

Lines of inquiry this paper opens 18

Research framings built by reading the notes related to this paper — the questions it feeds into.

Are AI-generated articles systematically disadvantaged in search ranking and user engagement? How do educators verify student capability when AI can produce indistinguishable work? How can we reduce inherent biases in LLM-based evaluation judges? How do real-world evaluations reveal AI capabilities that benchmarks hide? How does awareness of evaluation context influence model behavior? How do network effects and self-selection distort aggregated rating accuracy? Can AI systems evade safety evaluations through reasoning manipulation? What human oversight must AI research systems have? What explains the gap between benchmark scores and true reasoning capability? Does AI-assisted research sacrifice exploration breadth for productivity gains? What gaps exist between benchmark performance and real deployment outcomes? How do AI hiring systems affect authenticity, fairness, and candidate preferences?