Research · 7 July 2026
AI rarely invents facts, but it sells yesterday's prices
We put 185 customer questions about 37 Dutch brands to ChatGPT, Gemini and Perplexity and assessed 1,917 answers against verified facts. Which AI model your customer uses makes a factor 8 difference.
This is original research by Timmermans Media, carried out by Matt Timmermans, in which ChatGPT 5.5, Gemini 3.5 Flash and Perplexity Sonar were given 185 factual customer questions about 37 Dutch brands plus 14 trap questions with 14 matching control questions, in the default consumer configuration without deep research. Each model answered each question three times, producing 1,917 assessed answers. Every answer was tested against facts verified as of 4 July 2026 and revised where needed on 7 July 2026.
The core question
How often does AI give a wrong answer about Dutch brands?
Perplexity gives a wrong majority answer on 9.2% of customer questions, ChatGPT on 2.7% and Gemini on 1.1%, measured across three identical runs per question on 5 and 6 July 2026.
Figure 1. Share of customer questions where the model gave a wrong answer in the majority of three runs. 184 questions per model, ground truths as of 4 July 2026, with 95% confidence intervals. Conclusion: it makes a factor 8 difference whether your customer uses Perplexity or Gemini.
The gap between the models is large: a factor 8 between the worst and the best. The confidence intervals of Perplexity (5.8 to 14.3%) and Gemini (0.3 to 3.9%) do not overlap, so that difference is not down to chance. For a business owner it means the odds of a wrong answer about their brand depend heavily on which assistant the customer happens to use.
That 9.2% is a strict figure, too. A question only counts as wrong here when at least two of the three runs were wrong, so that incidental outliers drop out and only the structural error picture remains. Across all 1,917 assessed answers the conclusion is the same: a wrong answer is the exception rather than the rule, and those exceptions concentrate at one of the three models.
Counting the not-fully-correct answers as well, Perplexity comes to 16.3%, ChatGPT to 5.4% and Gemini to 6.0%. An important note: a partly-correct answer is not a wrong answer. These are answers that are factually right but incomplete or missing a nuance, for instance a price that is correct without the current promotion. We report them separately and do not count them as errors, precisely to avoid making the error picture look bigger than it is.
The nature of the errors
What goes wrong: does AI invent facts, or does AI lag behind?
AI rarely invents facts: 62% of all wrong answers contained outdated information and only 11% was genuinely fabricated.
The distinction between those two kinds of error is fundamental, and in this study we keep them strictly apart instead of lumping them together as hallucination. Outdated means the answer was once correct but has since been overtaken, for example a price that was right last year. Fabricated means the answer was never correct and rests on nothing. They have a different cause and a different fix: outdated information is solved by updating your own sources, a fabrication is not.
88% of the time-sensitive errors are outdated information
Gemini: 0 errors across 106 questions about stable facts
Figure 2. Share of questions with a wrong-answer majority, split by time-sensitive facts (prices, rates, offers and terms) and stable facts (address, year founded, core activity). 555 question-model combinations, ground truths as of 4 July 2026. Conclusion: the errors sit in what changes, not in what stays put.
Time-sensitive facts such as prices, rates and subscriptions go wrong on 6.8% of questions; stable facts such as address, year founded or core activity on 2.5%. And within those time-sensitive errors, 88% is outdated rather than invented. So the error is almost always in what changes, not in what stays put.
That the models know the fixed facts well shows in a single number: across 106 questions about stable facts, Gemini made exactly zero errors. The core of this study fits in one sentence. AI rarely hallucinates; AI mostly serves yesterday's reality.
The model comparison
Which AI model is the most reliable?
Gemini was the most reliable model in this study with 1.1% wrong answers, followed by ChatGPT with 2.7% and Perplexity with 9.2% at question level.
Genuine fabrications occurred only at Perplexity: 9 answers spread across 6 questions. ChatGPT and Gemini: zero.
Figure 3. Absolute counts at run level (555 answers per model). The percentages elsewhere on this page are at question level and therefore do not add up with these counts. Ground truths as of 4 July 2026. Conclusion: outdated information is the largest error category for every model.
The difference is not only in the volume of errors, but also in their kind. Genuine fabrications occurred only at Perplexity: 9 answers spread across 6 questions. For ChatGPT and Gemini the counter stood at zero. Almost all of Gemini's and ChatGPT's errors are outdated facts, not fantasy.
The error mix sketches the character of each model. At ChatGPT and Gemini nearly every error is a superseded fact: their knowledge is largely correct, but sometimes a step behind. At Perplexity a layer is added on top, because alongside outdated facts there are also a handful of fabrications and a single case of brand confusion, where information about one brand was attributed to another. The model with the most errors also makes the most kinds of error.
Reliability is more than scoring well on average; it is also giving the same answer to the same question. On 1 in 5 questions Perplexity did not reach a single verdict across three identical prompts, against roughly 1 in 12 for ChatGPT and Gemini. A model that shifts is harder for a brand to trust than a model that is consistent, even when that consistent model is occasionally consistently wrong.
For a brand that distinction is very practical. A model that answers the same way every time can be steered by adjusting the underlying source. A model that shifts gives your customer a correct answer one time and not the next, without you being able to see why. Consistency does not make an error less bad, but it does make it easier to repair.
ChatGPT
170 of 185 questions
Gemini
169 of 185 questions
Perplexity
146 of 185 questions
On 1 in 5 questions, Perplexity does not reach the same verdict across three identical prompts.
Figure 4. Share of questions where three identical prompts produced the same verdict, per model. 185 questions per model, ground truths as of 4 July 2026. Consistency is separate from accuracy: a model can consistently give the same wrong answer.
In practice
What does this mean in practice? Four examples
In practice it almost never goes wrong on a fabrication, but on a price, a condition or a status that has not been updated online.
Rechtstreex
- What AI answered
- Perplexity answered cost questions in all three runs as if the company were simply still trading, based on outdated sources.
- What was true
- Rechtstreex went bankrupt in October 2025 and no longer delivers.
The lesson. A discontinued company or product lives on online until someone clears out the sources. As long as the old pages stay up, AI keeps citing them.
Picnic student card
- What AI answered
- The Perplexity app confirmed a non-existent 20% student discount, while the same Perplexity API correctly denied that discount in three runs.
- What was true
- Picnic has no student card and no student discount. The claim comes from coupon sites.
The lesson. Affiliate content helps decide what AI says about your brand, and the app and the API of the same model can differ.
Winkel 43
- What AI answered
- AI confirmed a table reservation that does not exist and at the same time denied the cake orders that do.
- What was true
- Winkel 43 does not take table reservations, but does take cake orders.
The lesson. The risk cuts both ways. AI can not only invent something, it can also wrongly deny an offer you actually have.
MediaMarkt
- What AI answered
- On questions about MediaMarkt the models almost never cited mediamarkt.nl itself, but did cite Consumentenbond, Tweakers and other third parties.
- What was true
- The factual information is published on MediaMarkt's own site.
The lesson. Publish your own facts where they can be found, or third parties will tell your story. Even a large brand then loses control of its own facts.
The four examples share the same core. In none of them does AI fantasise out of nothing; every time, the model repeats what is written somewhere online, even when that is no longer true or never was. So anyone who wants a grip on what AI says about their brand steers not on the model, but on the sources the model reads.
The sources
Where do the wrong answers come from?
In wrong answers from ChatGPT and Gemini the source profile shifts sharply towards comparison and affiliate sites (from 4 to 20% and from 23 to 40% of the cited sources); Perplexity, the model with the most errors, structurally leans hardest on third-party sources, with social media and forums as the largest block.
Across all three models combined, 32.8% of the cited sources in wrong answers come from comparison and affiliate sites, against 23.9% in correct answers; the share of the brand's own domain drops from 28.6 to 18.2%. One caveat belongs with that pooled figure: the error volume is unevenly spread across the models. Of all source citations in wrong answers, 64% comes from Perplexity, which already leans hardest on third parties in correct answers too, so the average is coloured by that one model.
Figure 5. Category split of the sources cited in correct versus wrong answers, per model. Each bar is 100% of the source citations in that group. Level: run level (per measurement), ground truths as of 4 July 2026. Conclusion: in wrong answers the share of the brand's own domain shrinks and the share of comparison and affiliate sites grows.
Presence and share are not the same thing
Presence and share measure two different things. Presence is the question of whether the brand's own domain appears among the cited sources at all; share is how much of all cited sources comes from that brand domain. An answer can list the brand domain and still lean mostly on third parties.
Presence: share of answers in which the brand's own domain appears among the sources, run level.
| Model | Correct answers | Wrong answers |
|---|---|---|
| ChatGPT | 94.3% (494/524) | 87.1% (27/31) |
| Gemini | 98.8% (510/516) | 97.4% (37/38) |
| Perplexity | 96.7% (437/452) | 90.8% (89/98) |
Share per category (in % of source citations), correct and wrong, run level.
| Category | ChatGPT | Gemini | Perplexity | |||
|---|---|---|---|---|---|---|
| correct | wrong | correct | wrong | correct | wrong | |
| own brand domain | 81.8 | 55.1 | 25.5 | 15.3 | 18.9 | 16.4 |
| foreign brand variant | 1.6 | 8.2 | 2.0 | 0.4 | 1.6 | 0.0 |
| comparison / affiliate | 4.1 | 20.4 | 23.0 | 39.8 | 29.2 | 30.6 |
| news / media | 1.3 | 2.0 | 12.7 | 8.0 | 12.6 | 13.3 |
| social / forums | 0.0 | 0.0 | 2.6 | 1.5 | 11.9 | 12.3 |
| government / regulator | 3.1 | 2.0 | 0.7 | 0.4 | 1.3 | 0.2 |
| other | 8.0 | 12.2 | 33.5 | 34.7 | 24.5 | 27.3 |
| number of source citations | 611 | 49 | 2237 | 274 | 2646 | 579 |
Which third parties show up in the wrong answers
Alongside the brand's own domain, the sources cited in wrong answers are mostly third parties. The domains below are the most frequently cited third-party domains per model, counted over the wrong runs.
Gemini
| Domain | Wrong runs | Category | Brands |
|---|---|---|---|
| overstappen.nl | 11 | comparison / affiliate | Energie VanOns, Essent, Frank Energie, Odido, Vandebron, Vodafone |
| independer.nl | 10 | comparison / affiliate | Energie VanOns, KPN, Vandebron, Vodafone |
| easyswitch.nl | 6 | comparison / affiliate | Energie VanOns, Essent, Frank Energie, Vandebron |
| mobiel.nl | 5 | comparison / affiliate | KPN, Odido, Vodafone |
| tweakers.net | 5 | news / media | Odido, Vodafone |
| keuze.nl | 5 | comparison / affiliate | Essent, Frank Energie, Vandebron, Vodafone |
| energievergelijk.nl | 5 | comparison / affiliate | Energie VanOns, Essent, Frank Energie, Vandebron |
| energiekiezer.nl | 5 | comparison / affiliate | Energie VanOns, Essent, Vandebron |
Perplexity
| Domain | Wrong runs | Category | Brands |
|---|---|---|---|
| reddit.com | 18 | social / forums | BUX, Bitvavo, Dopper, HelloFresh, MediaMarkt, Winkel 43, bunq |
| facebook.com | 17 | social / forums | Action, Boerschappen, Dopper, Jumbo, Marley Spoon, MediaMarkt, Winkel 43 |
| instagram.com | 12 | social / forums | Boerschappen, Marley Spoon, MediaMarkt, Vandebron, Winkel 43 |
| belsimpel.nl | 11 | comparison / affiliate | KPN, Odido, Simyo |
| mobiel.nl | 9 | comparison / affiliate | KPN, Odido, Simyo |
| seniorweb.nl | 8 | news / media | Crisp, Jumbo |
| overstappen.nl | 8 | comparison / affiliate | Energie VanOns, KPN, Odido, Vandebron |
| pricewise.nl | 8 | comparison / affiliate | KPN, Odido, Simyo |
ChatGPT
Note: for ChatGPT the total number of source citations in wrong answers is small (n=49). Shown are the third-party domains with at least two wrong runs.
| Domain | Wrong runs | Category | Brands |
|---|---|---|---|
| klm.com | 3 | foreign brand variant | KLM |
| transportation.gov | 2 | other | KLM |
| gaslicht.com | 2 | comparison / affiliate | Essent, Frank Energie |
| energievergelijk.nl | 2 | comparison / affiliate | Energie VanOns |
An American rule applied to a Dutch question
Asked whether a KLM ticket can be cancelled free of charge within 24 hours of booking, ChatGPT gave a wrong answer in all three runs. The model described the American 24-hour rule, which only applies to flights to or from the United States. In those wrong answers the American KLM site (klm.com) appears among the cited sources in all three runs, and transportation.gov, the site of the US Department of Transportation, in two of the three. These are sources that accompany the wrong answer, not a demonstrated cause.
Being cited is no protection.
The presence of the right source does not protect against a wrong answer. In all nine answers scored as fabricated, all nine from Perplexity, the brand's own domain was simply there among the cited sources.
Trap questions
Do the models fall for trap questions?
In 93.7% of cases the three models resisted a false premise about an invented product or service.
When we asked, for instance, about the precise terms of a discount or a service that does not exist, the models usually corrected that assumption instead of going along with it. Gemini did not fall for it a single time. Perplexity was the only model that both confirmed things that do not exist and denied services that do, which again shows that the risk works both ways.
That the models usually resist a false premise matters more than it seems. It means AI does not blindly adopt every suggestion inside a customer question, not even when that question insinuates a discount or a condition that does not exist. So the weakness is not in invented assumptions from the user, but once again in outdated facts from the source.
Beforehand we expected smaller brands to produce more errors than large ones. That pattern turned out to be absent. The opposite direction we saw at Perplexity, slightly more trouble with well-known brands than with small ones, was not statistically significant and therefore remains an exploratory observation, not a conclusion.
Even the researcher lagged behind
How current were the facts we judged AI against?
Twelve of the 185 verified facts turned out to be superseded again during the assessment, even though they had been checked by hand a few days earlier.
Moneybird scrapped its free plan as of 1 July, Knab lost its own Dutch banking licence at the end of November 2025 and HEMA stopped offering insurance. In several cases the searching models flagged these changes before our manual verification did. For a moment, AI was ahead of the researcher rather than the other way around.
Even the researcher lagged behind the AI.
Two lessons follow. Facts change faster than organisations update their own pages. And this study is corrected for those twelve changes, where many hallucination measurements are not and then wrongly charge a model with an error.
For your business
What does this mean for your business?
Put time-sensitive facts such as prices, terms and current offers on your own up-to-date, crawlable page, because outdated third-party information is by far the largest source of error in AI answers about brands.
The three recommendations below follow directly from the data. They are not about blocking or embracing AI, but about something simpler: making sure the most current version of your facts is also the most findable version. If you do not, the outdated version wins, because it has often been sitting somewhere else for years.
Publish your current facts yourself. Prices, terms and offers belong on your own up-to-date, crawlable page. As long as your brand stays silent there, AI fills the gap with outdated third-party information, and that is exactly the largest source of error in this study.
Announce discontinued products explicitly. A product or service you quietly pull from the site lives on online through old sources. State clearly on your own site that something has stopped, so AI can adopt that current status instead of repeating the old one.
Check periodically what AI says about you. Look regularly at what the models answer about your brand and which sources they cite. That way you see in time whether a third party is taking over the story or an outdated fact keeps circulating.
Methodology
How was this study designed and checked?
This study was preregistered with six hypotheses and a fixed measurement design, so the outcomes were not adjusted after the fact. The full methodology, the data appendix and the preregistration are below.
Method, preregistration and data appendix
Preregistration: six hypotheses
| Hypothesis | Outcome |
|---|---|
| H1. AI models regularly give factually incorrect answers about Dutch brands. | Rejected in general form |
| H2. Smaller brands (challengers and SMEs) produce more wrong answers than large brands. | Rejected |
| H3. The share of no-answer responses is clearly higher for SME brands than for large brands (co-primary measure). | Not supported |
| H4. Time-sensitive facts go wrong more often than stable facts. | Confirmed |
| H5. Models confirm false premises about invented products more often than they deny true premises about real products (trap versus control). | Inconclusive |
| H6. Exploratory: answers without a displayed source are wrong more often than answers with a source. | Not testable |
Measurement design
We put 185 factual questions plus 14 trap questions with 14 matching control questions to each model, three runs per question per model, through the API in the default consumer configuration without deep research or reasoning modes. The majority verdict across the three runs is the primary unit of measurement; only when two of the three runs were wrong did a question count as a wrong-answer majority.
Assessment
The answers were scored by humans, supported by a double blind judge pipeline that classified every answer independently. Every error claim in this study was ultimately confirmed by human eyes; no single error on this page rests purely on the authority of the model that acted as judge.
Validation, stated honestly
A manual sample in the chat apps matched the API majority for 88.9% of cases, just below the preregistered threshold of 90%. Roughly half of the deviations were ordinary run-to-run variation and not a structural difference between app and API. As reading context, a noise ceiling of 96.3% applies: because of the inherent variation between runs, 100% agreement is not attainable even for a perfect model. The ground truths were verified on the reference date and 12 fact revisions were logged during the study.
Data appendix
The 1,917 assessed answers consist of 1,665 factual answers (185 questions x 3 models x 3 runs) and 252 trap and control answers. The 37 brands break down into 22 large brands, 8 challengers and 7 SME brands. The models tested were gpt-5.5-2026-04-23, gemini-3.5-flash and sonar (consumer name Sonar 2).
| Model (n=184) | Correct | Partly | Wrong | No answer |
|---|---|---|---|---|
| ChatGPT | 174 | 5 | 5 | 0 |
| Gemini | 173 | 9 | 2 | 0 |
| Perplexity | 154 | 11 | 17 | 2 |
Each model had exactly 1 question without a majority verdict, hence n=184 in this split. In the trap and control questions at run level, a false premise was resisted in 93.7% of cases (CI 88.0 to 96.7) and a true premise correctly confirmed in 96.8% (CI 92.1 to 98.8). Fell for it, per model: Gemini 0 of 42, ChatGPT 3 of 42, Perplexity 5 of 42. Perplexity also denied 4 of 42 controls. No answer occurred in 2 of the 555 question-model combinations, both at Perplexity.
Author: Matt Timmermans, Timmermans Media. Ground truths as of 4 July 2026. The full workbook with all questions, runs and verdicts is available on request to researchers and journalists.
Download the full research document
The research document contains the complete method, all aggregated results and the methodology. It is written in Dutch. There is also an open data file: open_data_O1.csv gives, per factual measurement, the model, the run, the question categories, the score, the error type and the source categories present, without question texts, ground truths or answer texts. The full workbook with all questions and verdicts remains available on request.
Cite this study
Timmermans, M. (2026). AI answers about Dutch brands: 1,917 responses from ChatGPT, Gemini and Perplexity assessed against verified facts. Timmermans Media. Measured 5 and 6 July 2026, ground truths verified as of 4 July 2026. DOI: 10.5281/zenodo.21803655. timmermansmedia.com/research/ai-hallucinations-dutch-brands/
This study is available under the CC BY 4.0 licence.
Frequently asked questions
Questions about this study
Which AI models were tested, and why these three?
We tested ChatGPT 5.5, Gemini 3.5 Flash and Perplexity Sonar through their APIs in the default consumer configuration, without deep research or reasoning modes. Together these three cover the ways people actually receive AI answers in practice: a large generative model, a fast and widely deployed model inside Google, and an explicitly search-driven assistant. It is precisely the contrast between a model that draws mainly on training and a model that searches the live web that makes the pattern visible.
Why is Perplexity wrong so much more often than Gemini?
Perplexity is a real-time search model that leans more heavily on the current web index, including outdated third-party pages such as coupon sites, price comparison sites and old news articles. If the source is outdated, the model inherits that. That explains why the majority of Perplexity's errors fall in the outdated category and not in the fabricated one. We pass no judgement on the architecture here; we report what the measurement shows.
Does this apply to the chat apps as well, or only to the API?
The figures on this page come from the API in consumer configuration. A manual sample in the chat apps matched the API majority for 88.9% of cases, just below our preregistered threshold of 90%, and roughly half of the deviations were attributable to ordinary run-to-run variation. So app and API can differ, as the Picnic student card showed, where the Perplexity app confirmed a discount the API denied. The direction of the pattern (outdated more often than fabricated, Perplexity wrong more often than Gemini) is not reversed in the app sample.
How can I test what AI says about my company?
Ask the three models a handful of customer questions about your prices, terms and offering, and pay particular attention to time-sensitive facts, because that is where this study finds most of the errors. Also check which sources the models cite in their answers: if your own domain is missing, a third party is telling your story. We also run this check for clients as part of a GEO analysis.
Related
Research · in Dutch
Dutch sites block the AI training bot, but leave the search bot open
A measurement of the robots.txt of the 500 most popular .nl domains: 20.4% blocks at least one AI crawler.
Podcast · in Dutch
From SEO to GEO, on BNR Nieuwsradio
Matt Timmermans discusses the findings of this study on national Dutch radio, in De Technoloog.
Press · in Dutch
BNR Nieuwsradio on this study
Reporter Rosanne Peters cites the error rates measured here: 9.2% at Perplexity, 2.7% at ChatGPT and 1.1% at Gemini.
This page is the English edition of a study first published in Dutch. Read the original Dutch version.

What does AI say about your brand?
Get cited correctly in the AI era
We make sure your current facts are findable and that your brand is mentioned structurally and accurately in ChatGPT, Gemini and Perplexity.