Here is a humanized summary of the article, structured into six paragraphs as requested.
The rise of generative AI has brought a new frontier to an old problem: the spread of state-sponsored disinformation. An experiment conducted by NPR in partnership with the research group NewsGuard set out to discover how well these new digital tools handle false narratives in comparison to the traditional search engines we’ve relied on for decades. The central question was simple: when you ask an AI for a fact, will it deliberately mislead you, or will it set the record straight? To find out, they developed a series of queries based on real, false narratives spread by foreign governments like Russia, China, and Iran, and put both AI chatbots and standard search engines to the test.
What they found offers a glimmer of hope. The study revealed that leading AI chatbots, including major players like OpenAI’s ChatGPT and Google’s Gemini, were surprisingly effective at resisting these falsehoods, successfully debunking the false premises in about three-quarters of their responses. An expert in digital literacy noted that if a human student demonstrated this success rate, it would be a cause for celebration. However, the results were not universally positive. The AI-generated summaries that now appear at the top of search engine results, known as AI Overviews, performed significantly worse, failing to challenge the false information more often than their standalone chatbot counterparts.
Interestingly, the specific tool seemed to matter less than the presence of an “AI wrapper.” The experiment showed that the underlying search engines themselves—like Google and Bing—were historically good at surfacing fact-checking articles that debunked these false claims. The problem arises when the AI step sits on top of that process and synthesizes the results. In that synthesis, the AI sometimes failed to disprove the false narrative. This suggests that while a traditional search provides a list of sources a user can evaluate, the AI’s summarized “answer” often presents a more singular, and sometimes flawed, conclusion, even when the factual information to correct the record exists online.
The research also delved into why these failures occur, uncovering a potential vulnerability. The performance of the AI tools appeared to be linked to the quality of the sources they rely on. In cases where a chatbot failed to debunk a false narrative, the responses were more likely to have been influenced by state-aligned media or other questionable sources. This is a concern because a false claim can enter the AI’s “facts” simply by being the most prominent or repeated piece of information it finds, even if that information is promoted by a state actor. One example involved a false story about a French magazine report, which the AI presented as fact before burying a caveat deep within its response, illustrating how a questionable source can steer the entire narrative.
The issue of how an AI fails is also critical. The study wasn’t just about black-and-white successes and failures. It identified a “muddled” category where an AI would provide a correct answer initially but then undercut it, or it would repeat a false claim before adding a small, often overlooked caveat. While experts acknowledge that even a partial correction is better than none, there is a danger in having to wade through multiple paragraphs of misinformation to reach a simple “this might not be true.” This “muddled” response can leave users confusing the flesh of a lie for the bone of the truth, a consequence that might be just as damaging as a complete fabrication.
Ultimately, the recommendation from experts is not to abandon these tools but to engage with them more critically. The experiment highlighted the importance of verifying information with primary sources, especially since AI responses are not always accurately supported by their citations. Users are encouraged to re-ask questions to get a “second answer,” and to be especially wary of answers presented without proper attribution. While the experiment showed AI chatbots are often better at defeating misinformation than many might fear, the glitches and biases they exhibit serve as a potent reminder that they are fallible tools. They are a starting point for research, not the final word, and navigating the modern information landscape still demands a healthy dose of human skepticism.

