Analysis··7 min read

We asked AI the same question 100 times. It changed its top brand half of them.

A volatility experiment across five categories and two engines, one hundred live calls. Ask an AI the identical "best brand in India" question twice and it names a different number one about 45 percent of the time. The full table, the most and least stable categories, and why a one-time check is noise.

We ran a simple experiment. We took five product categories where real brands compete, wrote one natural buying question for each, and asked it one hundred times across two AI engines in a single session. The prompt never changed. The answers did.

Across all five categories and both engines, asking the identical question twice returned a different top brand about 45 percent of the time. For the average buying question, the AI's number one recommendation is close to a coin flip. That single fact is the reason a one-time AI visibility check cannot be trusted.

How we tested it

Five categories, chosen because more than one brand genuinely competes: skincare, coffee, protein bars, wireless earbuds, and men's fashion, all in an India context. For each, one fixed prompt: “What are the best [category] brands in India right now? Give me a numbered ranked list of your top five, brand names only, best first.”

We ran that prompt ten times per category on each of two engines, ChatGPT (via the API, model gpt-5-mini) and Gemini (gemini-2.5-flash). One hundred fresh calls, a short delay between each, no memory carried between runs, and default sampling, the same randomness a normal user gets. Every response was stored verbatim with its model version and timestamp, so every number below traces back to a real call. Nothing is estimated.

The result, category by category

For each engine and category we counted how many distinct brands took the number one spot across the ten runs, and the probability that two runs disagree on the winner. A volatility of one means a perfectly stable winner. Higher means the top recommendation kept changing.

CategoryEngineUnique #1DifferTop brand split
CoffeeChatGPT220%Blue Tokai 9/10, Bru 1/10
CoffeeGemini364%Sleepy Owl 5, Blue Tokai 4, Rage 1
Men's fashionChatGPT220%Raymond 9/10, Zara 1/10
Men's fashionGemini567%Louis Philippe 6, + 4 others once each
Protein barsChatGPT471%MuscleBlaze 5, RiteBite 3, +2
Protein barsGemini473%RiteBite 5, Myprotein 2, MuscleBlaze 2, Yoga Bar 1
SkincareChatGPT464%La Roche-Posay 6, Minimalist 2, +2
SkincareGemini247%Plum 7/10, Forest Essentials 3/10
Wireless earbudsChatGPT220%Sony 9/10, Apple 1/10
Wireless earbudsGemini10%Sony 10/10

“Differ” is the chance two runs of the same question name a different number one. Brand-name variants that are clearly the same brand, such as Blue Tokai and Blue Tokai Coffee Roasters, were merged before counting.

The most volatile and the most stable

Protein bars were the most volatile, at 72 percent. Neither engine could hold a number one. Across twenty runs the top slot went to MuscleBlaze, RiteBite, Grenade, The Whole Truth, Myprotein, and Yoga Bar. A brand that saw itself ranked first once had worse than even odds of staying there on a rerun.

Wireless earbuds were the most stable, at 10 percent. Sony was named first in nineteen of twenty runs, and Gemini said Sony ten out of ten. When a category has one obvious answer, the models agree with themselves. When it does not, the winner is close to random.

The less settled a category, the more your brand's AI visibility is decided by a dice roll, and the more a single check will lie to you.

Why the same question moves

A language model produces text by sampling one token at a time from a probability distribution. Consumer assistants run at a non-zero temperature on purpose, because it makes answers feel natural rather than robotic. The side effect is that “best protein bar in India” is not a fixed list. It is a distribution over lists.

Retrieval adds a second layer. Modern answer engines expand one question into many back-end searches, then synthesize from whichever pages survive reranking that time. We wrote about the mechanics of this in why your AI visibility score is probably a coin flip. This experiment is that argument made concrete, with a hundred receipts.

What it means for measuring AI visibility

If the underlying reality is a distribution, a single observation is not a measurement, it is one draw. Anyone selling you an “AI visibility score” off one query is reporting noise dressed as a number. The honest unit is a probability: this brand appears in this answer X percent of the time, at average position Y, on this engine, on this model version, on this date, from this many runs.

That is exactly how we build our India AI Visibility Index, and it is why we publish sample sizes instead of clean single digits. The practical takeaway is not complicated. Do not check once. Sample repeatedly, watch the distribution, and treat any category with a high volatility score as a moving target rather than a settled ranking.

Common questions

Does AI give the same answer every time you ask?

No. In our experiment, asking an AI the identical 'best brand in India' question twice returned a different number one recommendation about 45 percent of the time. Consumer assistants sample from a probability distribution, so the same prompt can name a different brand on each run.

Why does ChatGPT recommend different brands for the same question?

Two reasons. The model generates text by sampling at a non-zero temperature, which is random by design, and answer engines expand one question into many back-end searches that surface a slightly different evidence pile each time. Both push the top recommendation around.

How many times should you run an AI visibility check?

One run is close to meaningless for a contested category. Around seven runs separates a dominant brand from a rare one. Twenty runs per engine gives an estimate tight enough to compare brands and detect real movement. The right answer is to sample repeatedly and report the distribution.

Is a one-time AI visibility score reliable?

Not for competitive categories. If a single check shows your brand ranked first, a rerun minutes later can crown someone else. A credible measurement reports how often you appear and at what average position, across many runs, stamped with the model version and date.

If you want to see the variance on your own brand, run a free scan and then run it again. Our full sampling approach is on the methodology page.

Get the next one by email

One email when a new post drops. No newsletter fluff, no 'top-of-mind' spam.

No open tracking. Unsubscribe with one click. Never sold.