News Integrity in AI Assistants
Source: European Broadcasting Union / BBC · 2025-10-22
This report is one of the largest cross-market evaluations of its kind. Working with the EBU, 22 public service media organizations in 18 countries – working in 14 languages – assessed how ChatGPT, Copilot, Gemini, and Perplexity answer questions about news and current affairs.
The research built on an earlier BBC study, which exposed inaccuracies and errors in AI assistants’ output. This new study explored whether this had improved and if the issues previously identified were isolated or systemic. Alarmingly, it found that AI routinely misrepresents news content, no matter which language, territory, or AI platform is tested.
The work involved professional journalists from participating PSM evaluating more than 3,000 AI responses against key criteria, including accuracy, sourcing, distinguishing opinion from fact, and providing context.
Almost half of all AI answers had at least one significant issue.
A third of responses showed serious sourcing problems.
A fifth contained major accuracy issues, such as hallucinated and/or outdated information.
Lines of inquiry this paper opens 12
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can AI systems reliably guide voters without introducing political bias?- Does ChatGPT displace search engines or question-and-answer platforms?
- What accuracy do AI chatbots actually provide on election topics?
- Which AI news assistant performs better than the others in this study?
- How often do chatbot news users actually return to original reporting?
- Do other AI assistants perform similarly on voter election questions?
- What makes voting-advice tools like Kieskompas more reliable than chatbots?
- Why do people using AI chatbots for news report higher trust than non-users?
- Do media narratives about AI shape trust in chatbot news answers?
- Does trying AI chatbots for news change people's trust levels over time?
- What specific chatbot features or interactions earn user trust in news contexts?
- Which markets show highest adoption but lowest trust in AI news?
- Does low chatbot trust reflect direct experience or general AI anxiety?