The evaluation tested leading chatbots, including OpenAI's ChatGPT, Microsoft's Copilot, Google's Gemini, and Perplexity AI, by asking standard, real-time news questions. The findings highlight severe systemic vulnerabilities:
- High Error Rates: AI models hallucinated, altered critical facts, or used outdated sources in 45% of news queries.
- Severe Misinformation: About 31% of the answers were deemed "seriously wrong," often introducing outright fabrications into otherwise true stories.
- Sourcing Failures: Chatbots regularly struggled to properly cite publishers. Google's Gemini performed the poorest in this metric, with 72% of its responses displaying significant sourcing issues or entirely fabricated references.
Become a supporter of this podcast: https://www.spreaker.com/podcast/today-explained--6659607/support.
Don’t miss our latest episodes! follow us to stay updates with fresh news everyday. Let’s keep the support our podcast to serve update news entire the world!