Skip to main content

Diese Seite gibt es auch auf Deutsch. Auf Deutsch lesen

Provider Trends Published Jul 18, 2026

Same question, different answer: almost one in four Gemini answer pairs shares not a single brand

Achtung.app asks every AI platform every tracked question several times, usually in varying wordings. In a share of those cases the exact same wording comes up twice, under identical conditions and in the same daily batch. Those pairs act as a control measurement, and they are the only ones evaluated here. As a side effect, they answer a question that has rarely been backed by data: how stable are the answers of ChatGPT, Gemini and Perplexity in the first place?

The result across 3,313 answer pairs from 1 June to 18 July 2026: on average, ChatGPT names 69 percent of the same brands as in the second run of the same day. Perplexity reaches 49 percent, Gemini 41 percent.

Average overlap between two answers to the identical question

Jaccard overlap of the brands named in both answers, where both runs used identical question text. 1 June to 18 July 2026, 3,313 answer pairs.

ChatGPT
69.0 %
Perplexity
48.5 %
Gemini
40.7 %

The extremes make it even clearer. Only ChatGPT regularly returns exactly the same brand list twice, in 37 percent of pairs. Gemini tips the other way: almost one in four answer pairs shares not a single brand. Ask Gemini the same question twice and in 24 percent of cases the two brand lists share nothing at all; in over half of those cases one of the two answers names no brand whatsoever.

The extremes: identical answer vs. no overlap at all

Share of answer pairs per platform, 1 June to 18 July 2026.

ChatGPT
Identical
37.0
Disjoint
10.0
Perplexity
Identical
9.2
Disjoint
8.4
Gemini
Identical
9.7
Disjoint
24.3

In percent of answer pairs. Identical: both answers name exactly the same brands. Disjoint: the two answers share no brand at all.

  • ChatGPT is the most stable of the three platforms and has recently become more stable still: since late June, overlap sits at 70 to 75 percent, and the share of identical answers rose from roughly a third to over 40 percent.
  • Perplexity held steady at just under 50 percent from May through July and has become noticeably more stable since late July, at 56 to 57 percent. Identical answers remain the exception (8 percent), a stable core of brands persists while the rest of the list changes.
  • Gemini oscillates between 35 and 53 percent overlap with no discernible direction. Identical answers are the exception here (10 percent).

The weekly view shows this is not a fluke: the order ChatGPT ahead of Perplexity ahead of Gemini has held in every single measurement week since mid-May. The chart below is updated weekly with the latest measurements.

Answer stability by week

ChatGPT 65.0 Gemini 44.5 Perplexity 63.7
0 25 50 75 100 % W24 W27 W30 W33 W36

Average overlap (Jaccard) between two answers to the identical question text, both collected in the same daily batch. Weekly means per platform, updated weekly. Latest data point: W36/2026.

Answer stability by week
Week ChatGPT Gemini Perplexity
W24 67.2 % 34.7 % 47.8 %
W25 66.8 % 37.4 % 46.9 %
W26 68.0 % 45.3 % 47.1 %
W27 72.8 % 43.9 % 49.0 %
W28 71.1 % 43.0 % 48.3 %
W29 73.0 % 41.6 % 50.1 %
W30 74.5 % 44.4 % 51.4 %
W31 70.6 % 37.3 % 56.0 %
W32 71.5 % 36.8 % 55.8 %
W33 73.3 % 46.5 % 57.0 %
W34 65.3 % 41.7 % 55.2 %
W35 67.5 % 45.4 % 60.6 %
W36 65.0 % 44.5 % 63.7 %

As of 18 August 2026: exactly one number has moved since publication. Across the three weeks since late July, Perplexity sits at 56 to 57 percent overlap instead of the 49 percent of the core window, and answer pairs without a single shared brand have fallen from 8.4 to 3.7 percent. It is not a different set of questions: restricted to the queries already measured in June, Perplexity still comes out at 57 percent. ChatGPT sits at 72 percent over the same weeks, Gemini is unchanged at 41 percent, and the order of the three platforms has not flipped in any measurement week.

Since 5 August, Achtung.app additionally discards Gemini answers that were produced without a web search. That puts the obvious objection to the Gemini figures to the test: the share of pairs in which at least one of the two answers names no brand fell from 34 to 26 percent, overlap did not move (42 percent instead of 41), and the share of disjoint pairs stayed at 25 percent. Gemini's volatility does not come from answers given without a web search.

Important context: the measurement runs at temperature 0 on all three platforms. None of the providers offers a usable seed on the search-grounded path, and temperature 0 alone does not make a model deterministic. Part of the measured deviation therefore comes from the model itself rather than from the web search. That alone does not explain the size of the differences: the platforms run a live web search per query and assemble the answer from scratch every time. An AI answer is a snapshot, not a database lookup.

For brands, this means: whether a brand appears in a single AI answer is decided anew with every query. AI visibility only becomes meaningful across repeated measurements. That is exactly why Achtung.app asks every tracked question multiple times and averages over weeks rather than over single answers.

Sample: 3,313 queries Window: 48 days Achtung.app runs every visibility query several times, in varying wordings. Only the pairs where both runs used identical question text are evaluated, at identical model and temperature 0; both runs come from the same daily batch. The brands named in the two answers are compared via the Jaccard coefficient (intersection divided by union of the two brand lists). Core window: 1 June to 18 July 2026, 3,313 answer pairs across 124 tracked queries in 151 wordings from multiple industries; weekly series from 4 May 2026, since when all three platforms have run on the same model throughout. Models measured via the respective APIs: ChatGPT (gpt-4o-mini), Gemini (gemini-2.5-flash), Perplexity (sonar); the consumer apps may behave differently. Answer pairs in which neither answer named a brand are excluded (Gemini: 279 of 1,201 pairs, ChatGPT: 8, Perplexity: 5). Claude is not included: weekly rather than daily measurement and a much smaller sample do not allow a fair comparison. The sampling parameters are identical across all three platforms (temperature 0); none of the providers offers a usable seed on the search-grounded path. The measured variance therefore contains both the platforms' live web search and a residual model component, and these data cannot cleanly separate the two. Since 5 August 2026 the measurement discards Gemini answers in which the model ran no web search (one retry, then no data point); earlier weeks still contain those answers. The core window is unchanged; the current figures quoted in the text come from the measurement weeks from 27 July 2026 onwards (ChatGPT 504, Perplexity 534, Gemini 400 answer pairs).

Want insights specific to your brand?

Start with a free AI visibility scan and see how your brand performs across AI assistants.

Start Free Scan Get in Touch