Skip to main content
Provider Trends Published Jul 18, 2026

Same question, different answer: almost one in four Gemini answer pairs shares not a single brand

Achtung.app asks every AI platform every tracked question several times, usually in varying wordings. In a share of those cases the exact same wording comes up twice, under identical conditions and about one second apart. Those pairs act as a control measurement, and they are the only ones evaluated here. As a side effect, they answer a question that has rarely been backed by data: how stable are the answers of ChatGPT, Gemini and Perplexity in the first place?

The result across 3,313 answer pairs from 1 June to 18 July 2026: on average, ChatGPT names 69 percent of the same brands as in its own answer from one second earlier. Perplexity reaches 49 percent, Gemini 41 percent.

Average overlap between two answers to the identical question

Jaccard overlap of the brands named in both answers, asked about one second apart. 1 June to 18 July 2026, 3,313 answer pairs.

ChatGPT
69.0 %
Perplexity
48.5 %
Gemini
40.7 %

The extremes make it even clearer. Only ChatGPT regularly returns exactly the same brand list twice, in 37 percent of pairs. Gemini tips the other way: almost one in four answer pairs shares not a single brand. Ask Gemini the same question twice and in 24 percent of cases the two brand lists share nothing at all; in over half of those cases one of the two answers names no brand whatsoever.

The extremes: identical answer vs. no overlap at all

Share of answer pairs per platform, 1 June to 18 July 2026.

ChatGPT
Identical
37.0
Disjoint
10.0
Perplexity
Identical
9.2
Disjoint
8.4
Gemini
Identical
9.7
Disjoint
24.3

In percent of answer pairs. Identical: both answers name exactly the same brands. Disjoint: the two answers share no brand at all.

  • ChatGPT is the most stable of the three platforms and has recently become more stable still: since late June, overlap sits at 70 to 74 percent, and the share of identical answers rose from roughly a third to over 40 percent.
  • Perplexity has been remarkably consistent since early May, but at a low level: a stable core of brands persists while the rest of the list changes almost every time.
  • Gemini oscillates between 35 and 53 percent overlap with no discernible direction. Identical answers are the exception here (10 percent).

The weekly view shows this is not a fluke: the order ChatGPT ahead of Perplexity ahead of Gemini has held in every single measurement week since mid-May. The chart below is updated weekly with the latest measurements.

Answer stability by week

ChatGPT 73.3 Gemini 46.5 Perplexity 57.0
0 25 50 75 100 % W21 W24 W27 W30 W33

Average overlap (Jaccard) between two answers to the identical question, asked about one second apart. Weekly means per platform, updated weekly. Latest data point: W33/2026.

Answer stability by week
Week ChatGPT Gemini Perplexity
W21 66.3 % 52.5 % 57.9 %
W22 64.8 % 42.9 % 46.6 %
W23 65.9 % 40.2 % 48.7 %
W24 67.6 % 35.2 % 49.5 %
W25 66.8 % 37.4 % 46.9 %
W26 68.0 % 45.3 % 47.1 %
W27 72.8 % 43.9 % 49.0 %
W28 71.1 % 43.0 % 48.3 %
W29 73.0 % 41.6 % 50.1 %
W30 74.5 % 44.4 % 51.4 %
W31 70.6 % 37.3 % 56.0 %
W32 71.5 % 36.8 % 55.8 %
W33 73.3 % 46.5 % 57.0 %

Important context: the measurement runs at temperature 0 on all three platforms. None of the providers offers a usable seed on the search-grounded path, and temperature 0 alone does not make a model deterministic. Part of the measured deviation therefore comes from the model itself rather than from the web search. That alone does not explain the size of the differences: the platforms run a live web search per query and assemble the answer from scratch every time. An AI answer is a snapshot, not a database lookup.

For brands, this means: whether a brand appears in a single AI answer is decided anew with every query. AI visibility only becomes meaningful across repeated measurements. That is exactly why Achtung.app asks every tracked question multiple times and averages over weeks rather than over single answers.

Sample: 3,313 queries Window: 48 days Achtung.app executes every visibility query twice: two runs with an identical prompt, identical model and temperature 0, about one second apart. The brands named in the two answers are compared via the Jaccard coefficient (intersection divided by union of the two brand lists). Core window: 1 June to 18 July 2026, 3,313 answer pairs across 124 tracked queries in 151 wordings from multiple industries; weekly series from 4 May 2026, since when all three platforms have run on the same model throughout. Models measured via the respective APIs: ChatGPT (gpt-4o-mini), Gemini (gemini-2.5-flash), Perplexity (sonar); the consumer apps may behave differently. Answer pairs in which neither answer named a brand are excluded (Gemini: 279 of 1,201 pairs, ChatGPT: 8, Perplexity: 5). Claude is not included: weekly rather than daily measurement and a much smaller sample do not allow a fair comparison. The sampling parameters are identical across all three platforms (temperature 0); none of the providers offers a usable seed on the search-grounded path. The measured variance therefore contains both the platforms' live web search and a residual model component, and these data cannot cleanly separate the two.

Want insights specific to your brand?

Start with a free AI visibility scan and see how your brand performs across AI assistants.

Start Free Scan Get in Touch