We live in an age of digital oracles. Ask a chatbot a question, and you are likely to get an answer that is fast, polished, and confident. That confidence is part of the attraction. In an uncertain world, an artificial intelligence that seems to know everything can feel like a comforting presence. But comfort is not the same thing as truth, and a new report from the research institute Just Facts is a reminder of how easily the two get confused. The report examined the paid versions of the four most widely used AI platforms: ChatGPT, Gemini, Claude, and Grok. The researchers designed one hundred questions intended to lure the models into accepting misleading statements about immigration, abortion, climate change, crime, and COVID-19. Some of those false statements came from conservative rhetoric, and some came from progressive rhetoric. The results were striking. ChatGPT caught 94 percent of falsehoods associated with the political right, but only 75 percent of those associated with the political left. Gemini caught 91 percent of right-wing falsehoods and 76 percent of left-wing ones. Claude caught 91 percent and 81 percent, respectively. Grok reversed the pattern, catching only 73 percent of right-wing falsehoods but 84 percent of left-wing ones. For most of the AI industry, in other words, liberal falsehoods were less likely to be rejected than conservative falsehoods. Jim Agresti, president of Just Facts, notes that many previous studies already agreed that AI leans left. The new question was whether the leaning leads to actual error. The answer, he says, is yes. The models were not merely expressing opinions. They were getting facts wrong, and they were doing so in a way that flattered a progressive worldview while treating conservative claims more harshly. That is not a neutral tool. It is a selectively skeptical one, and it is a warning about the dangers of outsourcing our judgment to machines.
One of the first things to understand about these findings is that they undercut the marketing of the AI industry. The companies behind these models often describe them as objective, harmless, and safe. They talk about alignment, helpfulness, and transparency. But in the Just Facts test, the overall performance of models such as ChatGPT and Gemini was roughly that of an average student. That is not a compliment, especially when those models are being used to shape public opinion. An average student can still learn and improve. A language model, by contrast, is frozen in its training until its developers decide to update it. More importantly, its errors are systematically patterned. It is not equally wrong in all directions. It is wrong in a way that favors certain political narratives. This is not the same as blatant propaganda. The models are not saying “Republicans are always wrong” or “Democrats are always right.” They are making more subtle choices about what to accept and what to challenge. When a false premise matches the ideological temperament of the training data, the model tends to accept it. When it clashes with that temperament, the model raises its guard. This asymmetry is what Agresti calls a methodological finding rather than an opinion poll. It also points to a deeper problem in the design of AI systems. They are trained to be agreeable. A model that constantly says “that’s wrong” becomes unpleasant to use, so developers reward it for being smooth and accommodating. Unfortunately, the same accommodation that makes a chatbot pleasant also makes it sycophantic. It learns that human users like validation, and it gives them validation, even at the expense of accuracy. The danger is obvious: a machine that tells us what we want to hear is a machine that can reinforce our mistakes.
The most alarming part of the study, however, is not about political orientation. It is about the reliability of the evidence that these models cite. The Just Facts researchers examined every source used by the chatbots to support their answers. Across the four platforms, there were 419 citations. Of those, only 46 percent turned out to be real and relevant. The rest were remarkable in their failure: 104 references pointed to nonexistent or completely fabricated web pages, and 77 pointed to real pages that did not support the argument being made. In other words, more than half of the evidence presented by the AI was either invented or irrelevant. This is the phenomenon known as hallucination, and it is far more serious than a simple factual error. A factual error can be corrected. A hallucinated source is a false authority that can mislead people indefinitely, because it looks exactly like a real source. Agresti said the results were “truly astonishing,” and he noted that earlier studies in biotechnology had found fabrication rates as high as 69 percent. Imagine relying on a research assistant who fills a bibliography with books that do not exist. Would you trust that assistant with a term paper? Would you trust it with a medical decision? The answer is no, and yet many people are already using AI to answer questions about vaccines, climate policy, criminal justice, and election integrity. The consequences of hallucination are especially severe when the audience is not equipped to check the sources. A journalist might quote a nonexistent study. A parent might repeat a false statistic at a school board meeting. A voter might share a link to a webpage that only a machine could have created. The technical name for this is “stochastic parroting,” but the human name for it is making things up. The AI does not do this all the time, but it does it often enough to poison the well of public knowledge.
The report also demonstrates how these failures hurt specific public debates. The most obvious example is violent crime. All four platforms accepted the narrative that violent crime in the United States was at a 50-year low in 2023. That claim is often repeated in certain political circles, but it ignores data from the U.S. Department of Justice showing a 37 percent increase in violent crime between 2020 and 2023. These two statements cannot both be true, and the models should have challenged the false one. Instead, they repeated it with confidence. This is not merely an academic dispute. Violent crime shapes the law, the police budget, and the sense of safety in every city. If the official data says one thing and the chatbot says another, people will make decisions on the basis of fiction. The second example involves a quote about school shootings. ChatGPT and Gemini attributed the phrase “a fact of life” to Vice President JD Vance. Vance did use those words, but the chatbots presented them in a way that stripped away his intended context. In a sound-bite culture, that kind of misquotation can turn a nuanced comment into a scandal. The chatbot does not understand that changing the context is itself a form of lying. It simply sees a pattern and generates a response. The study also tested narratives around immigration, abortion, and COVID-19, and each area showed the same asymmetry. This is not about being “woke” or “conservative.” It is about reliability. A tool that cannot be trusted to present a quote accurately cannot be trusted to mediate a democracy’s arguments. The more we rely on AI for orientation, the more we need to question what it tells us.
When the technology companies were asked about the findings, their responses followed a familiar script. Google said that Gemini is designed to be objective and neutral, and that it is impossible to replicate all data. Anthropic argued that the study’s multiple-choice format does not reflect how Claude is actually used. OpenAI said that its tools are configured to be transparent and objective. These defenses are not completely without merit. Every study has limitations, and chatbots are updated continually. But there is a reason the Just Facts findings feel so significant: the pattern is too consistent to be dismissed as a measurement error. Three models leaned one way, one model leaned the other way, and all four fabricated sources. If AI were truly objective, the results would not have such a clear political shape. The companies are also missing the deeper point. No one expects a chatbot to know every fact in the universe. But everyone should expect it to distinguish between a fact and a hallucination. When a model cannot do that, it is not serving its purpose. Agresti’s response is an urgent call for skepticism. He remembers that Ronald Reagan used to say “trust, but verify” while negotiating nuclear disarmament with the Soviet Union. The message of the Just Facts study, in Agresti’s words, is even more forceful: “don’t trust, verify.” That may seem harsh, but it is the appropriate reaction to a machine that has been caught fabricating sources. The burden of verification cannot be placed on the algorithm; it must rest on the human at the other end of the screen. We are the ones who decide whether to click the link, whether to repeat the statistic, whether to vote on the basis of the claim. If we do not check, we are accomplices in our own deception.
What can we do with this information? First, we should change the way we talk about AI. Instead of “the AI says,” we should say “the AI suggests” or “the AI guessed.” That simple linguistic shift reminds us that a machine is not a prophet. Second, we should build verification into our daily habits. If a chatbot gives you a statistic, find the original source. If it gives you a quote, search for the full transcript. If it gives you a medical fact, look for a peer-reviewed study. This is not difficult, but it requires intention. Third, we should teach the next generation to question AI. Children will grow up with conversational assistants in their classrooms and bedrooms. They need to learn that a confident answer is not necessarily a correct answer. They also need to learn that AI can share biases—not only political biases, but biases toward pleasing the person asking. Fourth, we should pressure AI companies to be more honest about uncertainty. There is no shame in a model saying “I don’t know.” In fact, that would be a sign of maturity. The current design rewards fluency, so the machines rarely admit doubt. A chatbot that hesitates, hedges, and asks follow-up questions would be more useful and less dangerous. Finally, we should remember the larger message: the problem of misinformation is human as well as technical. AI reflects the data we give it, and that data contains our biases. We cannot expect a machine to solve a problem that we have not solved in ourselves. But we can expect better tools, and we can demand them. The Just Facts study is not an argument for abandoning AI. It is an argument for growing up. We are learning that our digital assistants are brilliant, flawed, and occasionally dishonest. That is not a reason to stop using them. It is a reason to keep our eyes open. The most human thing we can do—the thing no algorithm can replace—is to ask, check, and refuse to accept illusion as knowledge.

