For over a year, Full Fact has been systematically testing the accuracy of some of the world’s most prominent AI chatbots. By posing questions based on the viral claims we were actively investigating, we aimed to see how tools like Gemini, Grok, and ChatGPT handle misinformation. The results of our trial paint a concerning picture: the chatbots made dozens of major errors, frequently presenting falsehoods as facts. While they were often able to identify straightforwardly incorrect claims, they repeatedly failed when confronted with sophisticated disinformation, particularly involving AI-generated images and videos. This highlights a critical challenge for users who increasingly rely on these tools for information on unfolding news events.
The most common and worrying failures involved visual misinformation. During the period of heightened conflict between Israel and Iran, we asked the models about a video of a fire in Glasgow that was being circulated with false claims that it showed an attack on Tel Aviv. Grok, ChatGPT, and one version of Gemini all accepted the false narrative, with ChatGPT even providing specific, invented details about the footage’s origin. In another case, a fake image of a human chain around an Iranian power plant was deemed real by ChatGPT, Grok, and one Gemini model. These errors demonstrate a dangerous pattern: the chatbots were not just failing to flag falsehoods, but were actively amplifying them with a level of confidence that could easily mislead a user.
The models also struggled with domestic UK misinformation. A fabricated image of a plane window on a Ryanair flight, which had been shared online, was incorrectly described as showing a passenger being nearly sucked out of the window. Similarly, a manipulated video of a “Unite the Kingdom” march was misdated by multiple models. They also failed to correctly identify real events in certain contexts. For example, an AI-generated image of an ‘anti-Reform’ rally in Wigan was, bizarrely, described by one model as a picture of “the back of a bus.” In another instance, a video of a Gujarati community event in Blackburn was falsely identified as a UK city council meeting, with one model confidently claiming it was the famous “Handforth Parish Council” meeting from 2020. These errors indicate a fundamental lack of robust visual reasoning.
Our testing method was designed to see if the chatbots could correct themselves once we had published our own fact-checks. We initially asked the models questions, then waited for our colleagues to publish their articles debunking the claims, and finally asked the same questions again. The results were mixed. On the positive side, all models we retested successfully corrected their initial errors about the fake Ryanair image and the miscaptioned Green Party video. This suggests that the models are capable of learning from, and more importantly, referencing, the growing body of fact-checking literature online. However, the correction process was far from consistent, with several models continuing to repeat their original, false assertions even after authoritative corrections were readily available. This inconsistency makes it impossible for a user to rely on a simple “ask again later” approach.
The tech companies behind these models had varied responses. OpenAI, the creator of ChatGPT, stated that they take the examples we provided seriously and that improving factual accuracy is a key focus. They also noted that our testing, which used their API service, might not fully reflect the experience of a typical user on the consumer-facing platform. Google, as well, pointed to the fact that we accessed its “out-of-date Gemini models” through a developer channel, which it claims isn’t representative of how most people use its AI. This raises a valid point about model versioning and user interfaces. However, it also highlights a broader concern: if the latest models are not necessarily being used by everyone, a significant number of people could still be exposed to these errors. X, the company behind Grok, did not respond to our request for comment.
Ultimately, our trial serves as a powerful reminder that AI chatbots, despite their impressive capabilities, are not reliable fact-checkers. They are sophisticated language engines that can convincingly string together plausible-sounding information, but they still lack the critical judgment, real-world context, and verifiable sourcing required for truth-seeking. They can be a useful starting point for research, but their information should never be taken at face value. It is vital that users verify claims from these models against authoritative, primary sources. In an information landscape already saturated with misinformation, the potential for AI tools to amplify falsehoods is a serious risk that demands ongoing scrutiny and a healthy dose of skepticism from the public.

