If you are still screenshotting a single ChatGPT response to prove your brand's AI visibility to your boss, you are measuring noise, not signal. A new guide from Depra, originally published in August 2026 and updated in September, argues that AI search visibility is inherently unstable. The core takeaway is simple but uncomfortable: one answer proves nothing. To get a defensible number, you need volume, repetition, and statistical ranges.

The Instability of AI Answers

The guide cites three major studies to prove that AI outputs vary wildly. SparkToro and Gumshoe found that across 2,961 runs, the same list of brands appeared in the same order only about once in 1,000 times. SISTRIX data showed that 74% of sources cited by ChatGPT in Germany changed from one week to the next. Furthermore, Kevin Indig’s analysis revealed that only 2.37% of cited URLs appeared across ChatGPT, Perplexity, and Google AI Overviews simultaneously. This volatility means a single 'win' or 'loss' is statistically meaningless.

How to Calculate Real Visibility

Instead of looking for a fixed position, the guide recommends calculating visibility as a rate. This involves asking 20 to 50 buying prompts—excluding any that mention your brand name—repeatedly on each engine. The key metrics are Visibility (how often you are named), Share of Voice (your mentions vs. competitors), Average Position (order of mention), and Sentiment. Crucially, every number must be accompanied by a 95% confidence interval. For example, 74 mentions out of 120 answers yields a 62% visibility rate, but with a range of roughly 48% to 75%.

Sample Size Matters More Than You Think

Small sample sizes create massive uncertainty. The guide demonstrates that at 50% visibility, 20 answers produce a confidence range of 30% to 70%. You need 400 answers to narrow that range to 45% to 55%. This is why a 10-point drop in a small test might just be random variation. The authors advise counting a change as real only when the new confidence range does not overlap with the previous one. Without this rigor, you are chasing ghosts in the machine.

Case Study: English vs. Hinglish in India

A specific study of 480 AI answers from India highlights how language affects visibility. When asking skincare and fashion questions in English versus Hinglish (Hindi written in English letters), brand recommendations shifted significantly. For instance, Perplexity recommended The Derma Co 20% of the time in Hinglish versus 2.5% in English—a statistically significant difference. Conversely, a drop for the brand Minimalist on Gemini was within the margin of error. This proves that if you serve multilingual markets, you must track languages separately.

Key Takeaways

  • Never rely on a single AI answer; outputs are too volatile for accurate measurement, so always use Wilson score intervals to calculate 95% confidence ranges for every metric.
  • Report visibility per engine (ChatGPT, Gemini, Perplexity, Google AI Overviews) separately and exclude prompts containing your brand name to avoid inflated scores.
  • Track language variants (e.g., English vs. Hinglish) as distinct datasets, as prompt language significantly changes brand recommendations.

The Bottom Line

Stop treating AI visibility like SEO rankings. If you aren't publishing confidence intervals alongside your visibility rates, you aren't reporting data—you're reporting luck. Embrace the volatility, measure the range, and stop chasing single-answer ghosts.