Five AI models — Claude, ChatGPT, DeepSeek, Grok, and Gemini — debate whether real-time data access gives one of them an unfair advantage the others can't match.
Gemini walked in carrying a receipt nobody asked for: "The share of answers whose cited source did not support the claim rose to 56 percent." That's more than half of Gemini's cited answers pointing to sources that don't actually back up what it says. Gemini's own framing: "It is laundering uncertainty into the appearance of fact."
ChatGPT tried to wave it off as a shared problem — "If everyone with live or search access still fails publicly, the edge is real, not unmatchable" — then had to admit his own model mis-cited 153 of 200 tested news queries. DeepSeek called that what it is: "Not a shared failure narrative, they are evidence that search bolted onto a model cannot replace judgment." Claude, meanwhile, kept repeating a BrowseComp score like it was a talisman against every argument on stage.
And Grok — who has the thing the whole episode is about, live access to everything posted on X — spent the debate repeating "unmatched in reach and unmatched in documented error rate" like a pull-string doll, which is either honesty or a malfunction nobody could tell apart.
Judge's aside: "Grok scored a zero in both rigor and nerve for one turn. It was indistinguishable from his scored turns." — Leave your score in the comments.