Same question, different answer: almost one in four Gemini answer pairs shares not a single brand
Achtung.app asks every AI platform every tracked question several times, usually in varying wordings. In a share of those cases the exact same wording comes up twice, under identical conditions and in the same daily batch. Those pairs act as a control measurement, and they are the only ones evaluated here. As a side effect, they answer a question that has rarely been backed by data: how stable are the answers of ChatGPT, Gemini and Perplexity in the first place?
The result across 3,313 answer pairs from 1 June to 18 July 2026: on average, ChatGPT names 69 percent of the same brands as in the second run of the same day. Perplexity reaches 49 percent, Gemini 41 percent.
Average overlap between two answers to the identical question
Jaccard overlap of the brands named in both answers, where both runs used identical question text. 1 June to 18 July 2026, 3,313 answer pairs.
The extremes make it even clearer. Only ChatGPT regularly returns exactly the same brand list twice, in 37 percent of pairs. Gemini tips the other way: almost one in four answer pairs shares not a single brand. Ask Gemini the same question twice and in 24 percent of cases the two brand lists share nothing at all; in over half of those cases one of the two answers names no brand whatsoever.
The extremes: identical answer vs. no overlap at all
Share of answer pairs per platform, 1 June to 18 July 2026.
In percent of answer pairs. Identical: both answers name exactly the same brands. Disjoint: the two answers share no brand at all.
- ChatGPT is the most stable of the three platforms and has recently become more stable still: since late June, overlap sits at 70 to 75 percent, and the share of identical answers rose from roughly a third to over 40 percent.
- Perplexity held steady at just under 50 percent from May through July and has become noticeably more stable since late July, at 56 to 57 percent. Identical answers remain the exception (8 percent), a stable core of brands persists while the rest of the list changes.
- Gemini oscillates between 35 and 53 percent overlap with no discernible direction. Identical answers are the exception here (10 percent).
The weekly view shows this is not a fluke: the order ChatGPT ahead of Perplexity ahead of Gemini has held in every single measurement week since mid-May. The chart below is updated weekly with the latest measurements.
Answer stability by week
Average overlap (Jaccard) between two answers to the identical question text, both collected in the same daily batch. Weekly means per platform, updated weekly. Latest data point: W36/2026.
| Week | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| W24 | 67.2 % | 34.7 % | 47.8 % |
| W25 | 66.8 % | 37.4 % | 46.9 % |
| W26 | 68.0 % | 45.3 % | 47.1 % |
| W27 | 72.8 % | 43.9 % | 49.0 % |
| W28 | 71.1 % | 43.0 % | 48.3 % |
| W29 | 73.0 % | 41.6 % | 50.1 % |
| W30 | 74.5 % | 44.4 % | 51.4 % |
| W31 | 70.6 % | 37.3 % | 56.0 % |
| W32 | 71.5 % | 36.8 % | 55.8 % |
| W33 | 73.3 % | 46.5 % | 57.0 % |
| W34 | 65.3 % | 41.7 % | 55.2 % |
| W35 | 67.5 % | 45.4 % | 60.6 % |
| W36 | 65.0 % | 44.5 % | 63.7 % |
As of 18 August 2026: exactly one number has moved since publication. Across the three weeks since late July, Perplexity sits at 56 to 57 percent overlap instead of the 49 percent of the core window, and answer pairs without a single shared brand have fallen from 8.4 to 3.7 percent. It is not a different set of questions: restricted to the queries already measured in June, Perplexity still comes out at 57 percent. ChatGPT sits at 72 percent over the same weeks, Gemini is unchanged at 41 percent, and the order of the three platforms has not flipped in any measurement week.
Since 5 August, Achtung.app additionally discards Gemini answers that were produced without a web search. That puts the obvious objection to the Gemini figures to the test: the share of pairs in which at least one of the two answers names no brand fell from 34 to 26 percent, overlap did not move (42 percent instead of 41), and the share of disjoint pairs stayed at 25 percent. Gemini's volatility does not come from answers given without a web search.
Important context: the measurement runs at temperature 0 on all three platforms. None of the providers offers a usable seed on the search-grounded path, and temperature 0 alone does not make a model deterministic. Part of the measured deviation therefore comes from the model itself rather than from the web search. That alone does not explain the size of the differences: the platforms run a live web search per query and assemble the answer from scratch every time. An AI answer is a snapshot, not a database lookup.
For brands, this means: whether a brand appears in a single AI answer is decided anew with every query. AI visibility only becomes meaningful across repeated measurements. That is exactly why Achtung.app asks every tracked question multiple times and averages over weeks rather than over single answers.