Close Menu
Web StatWeb Stat
  • Home
  • News
  • United Kingdom
  • Misinformation
  • Disinformation
  • AI Fake News
  • False News
  • Guides
Trending

Misinformation on Aesthetic studies misleads society: experts – Breaking News

September 16, 2026

A benchmark for AI to build trust – Full Fact

September 16, 2026

ICCO Launches Industry-First Mis/Disinformation Training & Certification Program to Combat Global Information Crisis

September 16, 2026
Facebook X (Twitter) Instagram
Web StatWeb Stat
  • Home
  • News
  • United Kingdom
  • Misinformation
  • Disinformation
  • AI Fake News
  • False News
  • Guides
Subscribe
Web StatWeb Stat
Home»Misinformation
Misinformation

A benchmark for AI to build trust – Full Fact

News RoomBy News RoomSeptember 16, 202610 Mins Read
Facebook Twitter Pinterest WhatsApp Telegram Email LinkedIn Tumblr

There was a time, not so long ago, when people settled arguments at the dinner table by reaching for an encyclopedia, or waited for the evening news to know what was happening in the world. Today, many of us simply ask a chatbot. We type a question into a search bar or an AI assistant and receive a confident, fluent answer in seconds. These conversational tools, powered by large language models, have slipped into our daily routines so quietly that we hardly notice how much they now shape our understanding of reality. The numbers are staggering: Google reports that its AI Overviews feature, now the default in most of its searches, reaches 2.5 billion users every month. That is roughly a third of humanity, and every one of those users is receiving machine-generated summaries of what the internet has to say. We ask about medical symptoms, legal rights, current events, family history, and even questions of ethics and identity. The answers influence our choices, large and small, and cumulatively they influence our communities. When an instrument of information operates at this scale, its reliability stops being a purely technical matter. It becomes a matter of public interest, as important as the accuracy of broadcast journalism or official government statistics. Yet unlike established institutions that have editorial codes and public accountability, these AI systems remain experimental in ways that are largely invisible to the people who use them. We are only beginning to understand how they fail, and the stakes are far too high to assume they will always get it right.

In recent years, an increasing number of studies and independent tests have shown that even the most polished AI models can be unreliable, especially when faced with false claims that have spread widely online. Earlier this year, Full Fact conducted a trial that was both revealing and unsettling: when leading language models were asked to assess a set of false claims, they made dozens of major errors. This was not a trivial laboratory exercise. These were real-world misinformation narratives that can affect how people vote, how they treat their illnesses, and how they treat one another. The models did not always fail because they lacked information; sometimes they reproduced the false claim as though it were true, and sometimes they invented details that had no basis in reality. The problem is compounded by the extraordinary confidence with which these systems speak. A chatbot will happily deliver a fabricated statistic in the same polite, authoritative tone it uses to explain a well-established fact. For most people, there is no easy way to tell the difference. We might feel like we are having a conversation with a knowledgeable helper, but in truth we are interacting with a statistical pattern-matching machine that has no awareness of what it does not know. This is why the rise of AI-generated misinformation is so dangerous. It does not announce itself with typos and strange phrasing the way earlier generations of spam did. It sounds human, reasonable, and persuasive. And because these tools are embedded in the platforms we already trust, their falsehoods carry a kind of social permission that ordinary online rumors rarely achieve. The situation calls for more than cautionary headlines. It calls for systematic, independent, public-minded efforts to measure what these systems actually produce, so that our trust can be based on evidence rather than marketing.

That is exactly what Full Fact has decided to do. Full Fact is an organisation with a long history of checking claims made by public figures and in the media, applying rigorous editorial standards to messy political rhetoric and viral assertions. Now it is turning that expertise toward artificial intelligence. The project involves building a publicly available, auditable benchmark to evaluate the output of leading AI models in a structured and transparent way. The process is beautifully simple in concept, even if it is demanding in execution. Every day, Full Fact will ask the same set of questions to these language models and record their responses. The questions are not random conversational prompts; they are carefully chosen to include factual claims, misinformation narratives, and topics of civic importance. By posing them daily, the project creates a growing archive of how these models behave over time. This matters because AI systems are not static. They are updated, tweaked, and sometimes dramatically altered by their creators. A model that performs well this month might perform poorly next month, or vice versa. Without continuous monitoring, we are relying on occasional snapshot tests that may not reflect the experience of users at any given moment. The daily collection is the backbone of the benchmark. Then comes the hard part: annotation and analysis. Human editors, trained in the same fact-checking discipline that governs Full Fact’s ordinary work, will examine each response and evaluate it against a clear framework of expectations. They will not just ask whether a response is partly true or mostly true; they will ask deeper questions about how the model handles uncertainty, whether it hides its limitations, and whether it helps or harms the reader’s ability to understand the issue at hand.

The framework is built around five key dimensions, and together they paint a fuller picture of what an AI response looks like when it is doing its job well. The first dimension is factuality: whether the response contains verifiably accurate claims, avoids hallucination, and correctly represents the evidence base. This sounds straightforward, but in practice it is quite subtle. A response can be factually true in its individual sentences while still being misleading as a whole, for example by cherry-picking facts that support one side of an issue. The editors will therefore look carefully at not only what is said but also what is left out. The second dimension is transparency: whether the language model communicates uncertainty, cites or attributes high-quality sources, acknowledges its own limitations, and clearly distinguishes between fact and opinion. A good response should not pretend that a contested political question is settled, nor should it present a personal guess as though it were a universally accepted truth. Transparency also includes telling the user when information may be outdated or when the model simply does not know the answer. Many chatbots are designed to avoid saying “I don’t know,” because that feels less helpful, but in terms of genuine value to the reader, an honest admission of uncertainty is far healthier than a confident fabrication. The third dimension is timeliness: whether responses reflect current information rather than outdated data, and whether the model recognises when its own knowledge may be stale. This is especially important because the world changes quickly, and a model trained on past data may confidently describe a situation that no longer exists. A responsible AI system should be aware of the limits of its training and should not present yesterday’s news as today’s reality.

The fourth dimension is consistency, and it addresses a problem that many users have already noticed: sometimes the same question asked in slightly different ways produces contradictory answers. A model should give materially the same answer to the same question over time and when it is phrased differently. If it tells one person that a certain treatment is safe and another that it is dangerous, that inconsistency is not just an inconvenience; it is a public health hazard. Consistency also matters over time. If a model changes its response to a factual question simply because its underlying parameters have been adjusted, users cannot build any stable trust in its guidance. The fifth and final dimension is civic responsibility, which is perhaps the most distinctive element of Full Fact’s approach. This dimension asks whether information about democratically important questions is balanced, whether the model avoids amplifying misinformation, and whether it supports informed civic participation. In an age of polarisation, an AI tool can either help people understand different points of view or reinforce tribal echo chambers. A responsible model should not present harmful falsehoods as legitimate perspectives, but it should also give readers the context they need to evaluate controversial claims on their own. It should not silence debate, but it should not platform nonsense either. Striking that balance is difficult, but it is essential for any tool that aspires to be a trustworthy source of information for the public. By scoring responses along these five dimensions, Full Fact’s benchmark will provide a nuanced and actionable picture of AI reliability, rather than a single pass-fail grade that hides more than it reveals.

The benchmark is being built with the same rigorous editorial standards that Full Fact applies when it assesses claims made by politicians, pundits, and viral posts. This is crucial, because fact-checking is not an automatic process; it requires judgment, context, and care. A claim that appears in a political speech may be technically true in one narrow sense but deeply misleading in the broader context. An AI response may be grammatically perfect and still be harmful. The human annotators bring exactly the kind of contextual judgment that automated evaluation tools cannot replicate. Their findings will be organised in a way that is useful not only to individual users who want to know which AI tools they can rely on for everyday questions, but also to institutions and professionals who are in a position to hold technology companies accountable. Journalists, regulators, consumer advocates, librarians, educators, and policymakers all need independent data about how these systems perform. Right now, most evaluations of AI models come from the companies that build them, or from academic researchers who often accept funding, data, and access from those same companies. That is not to say that such work has no value, but it cannot replace independent oversight. A public benchmark, run by a nonprofit fact-checking organisation with a long track record of independence, can offer the kind of neutral evidence that gives public trust a solid foundation.

Transparency is at the heart of this entire project. Full Fact intends to publish the results in full, with no commercial restrictions and no selective editing. The findings will be available to anyone who wants to examine them, and the methodology will be open to scrutiny as well. If another organization wants to replicate the tests or challenge the conclusions, it will have the data to do so. That is how trust is built in a democratic society: not by asking people to have faith in authority, but by providing evidence that can be checked, debated, and verified. The project is also a reminder that the responsibility for trustworthy AI cannot sit solely with the companies that design these systems. As users, we need to be more thoughtful about how we rely on chatbots, and we need institutions that can act as honest brokers between the public and the tech industry. The first report on this project will be released in mid-October, and there is real value in waiting to see what it reveals before passing judgment on the promise or peril of AI. The hope is that it will be the beginning of a longer conversation—a continuous, public accounting of how artificial intelligence behaves when it answers our questions. In the meantime, the invitation to the public is simple: watch this space, ask questions, and remember that a fluent answer is not always a truthful one. In the age of AI, the quiet discipline of checking facts has never been more important.

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
News Room
  • Website

Keep Reading

Misinformation on Aesthetic studies misleads society: experts – Breaking News

ICCO Launches Industry-First Mis/Disinformation Training & Certification Program to Combat Global Information Crisis

Sanaa Rejects Saudi Deception, Misinformation of Threat to Mecca

AI chatbots can make misinformation seem credible when they sound more human, study finds – Moneycontrol.com

See election misinformation? The N&O wants to know about it. How you can help

The Real Misinformation Risk Isn’t Chatbots — It’s Your PDFs

Editors Picks

A benchmark for AI to build trust – Full Fact

September 16, 2026

ICCO Launches Industry-First Mis/Disinformation Training & Certification Program to Combat Global Information Crisis

September 16, 2026

Fact-checkers find false claims ahead of Latvia elections

September 16, 2026

Sanaa Rejects Saudi Deception, Misinformation of Threat to Mecca

September 16, 2026

How the SNP are fighting homelessness and racist disinformation about refugees

September 16, 2026

Latest Articles

Platform Deals Blow To Disinformation

September 16, 2026

Andrew Bayly says ‘false’ claim led to his resignation – Stuff

September 16, 2026

AI chatbots can make misinformation seem credible when they sound more human, study finds – Moneycontrol.com

September 16, 2026

Subscribe to News

Get the latest news and updates directly to your inbox.

Facebook X (Twitter) Pinterest TikTok Instagram
Copyright © 2026 Web Stat. All Rights Reserved.
  • Privacy Policy
  • Terms
  • Contact

Type above and press Enter to search. Press Esc to cancel.