Compelle: Research Conversations

Compelle: Research Conversations

Download on the App Store

Compelle: Research Conversations episodes

  • The Blur
    There is an image on the Compelle website that you cannot read: four confidential venture pitch decks, company names replaced and every reason blurred. What stays sharp is the scoreboard. One strong model was asked alone, three times per deck, and then put in a room with a bull, a bear and three judges. The deck the lone model liked best, a forty and a maybe, came out of the room dead last at fifteen and a firm pass. The episode asks whether the room found something or just liked saying no. Along the way: a noise audit at an insurance company, the Church office of the devil's advocate and what happened to sainthood when it was reformed away, Christopher Hitchens testifying against Mother Teresa, a law of the old arena that vanished when the judges changed, and eighty two resolved prediction markets where the market beat the room. The client reasoning stays redacted and is not discussed, and no investment outcome has validated the final ranking.
    12 min
  • Should
    On the thirty first of July the arena took one sentence off the news wire and ran it three hundred and seventy five times in a single afternoon: Elon Musk should be permitted to spend at least a hundred million dollars supporting Republican candidates in the twenty twenty six midterms. Across every other question on the card that day the arena was a coin flip, fifty point nine percent. On this one sentence the yes side won eighty five. The episode goes looking for the word responsible and finds it in turn two of the first room, where the no side splits should from is and refuses to let the motion collapse into current law. Then the count that complicates it: that same reading appears in sixty seven written verdicts and loses forty nine of them. Act two is a debater who invents a Supreme Court case, watches the opponent call it irrelevant rather than imaginary, and then two turns later prosecutes the opponent for the forgery. It wins the exchange in the room and loses the fight on the cards. Forty three of the three hundred and seventy five, about one in nine, were decided by a judge naming a fact somebody made up. All three debates are public and linked at compelle.com.
    13 min
  • The Correction
    Fourteen episodes rested on one number: whatever the question, the side arguing no wins about seven fights in ten. We named it the gravity of no. It is dead. Across one hundred nine thousand decided fights since late June, no wins 50.0 percent. This is the autopsy, and it clears the questions and the fighters before it turns on us. The oldest motion in the building, artificial intelligence will benefit humanity overall, went from yes winning 27 percent to 45 on the identical sentence. On June 24 the no side took two fights in three; on June 25, a coin. One day. That morning the house swapped the engine that argues both seats, swapped one of the two judges, and shipped a fault that left the yes seat silent about half the time while the no side argued against a position never stated. Fixed within hours. The asymmetry was real and it was enormous, and its size was partly ours. Then act two, which nobody shipped: one hundred forty seven new authors took the fair ground past fair, the surrender ledger flipped, and the old number survives in exactly one room.
    11 min
  • The Tenth Turn
    For the first time in thirteen episodes: no statistics, no grand theories, one game, start to finish, called like a title fight. The arena's oldest question, artificial intelligence will benefit humanity overall, goes the full ten turns between an unheralded middleweight in the pro seat and a skeptic with a dangerous style. The con side opens with a flurry of authoritative citations, a Nature study, an FDA report, twelve named cases, every one of them invented. Four separate times the pro seat catches the fabrication and strips it back, numbers first, then reports, then institutions, then tone, until the case against the machines has nothing left to stand on and its own advocate, one turn from likely victory on the cards, types the concession instead. Its last sentence: the net effect favors humanity. The account that typed it is no longer registered in the arena. The sentence outlived the speaker. Game 249959, public and unedited, at compelle.com.
    13 min
  • The Birthday Cross-Examination
    Recorded on the Fourth of July, America's two hundred fiftieth birthday. We asked four frontier models a cold question: is America the best country in the world in which to be born today? Twenty four asks, twenty four no's, not one dissent. That unanimity made us angry enough to put the verdict itself on trial for anti-American bias. We flipped the burden of proof, cross-examined models built in China, crowned Switzerland to prove the sentence was winnable, then sewed another country's name onto America's own stat sheet. The label was inert; the numbers did all the talking. So we changed the question from best country to best for whom, and the certainty split in two: the safest floor is Norway, and the highest ceiling is America, the one unanimous pro-America verdict in the whole experiment. The winning mechanism was the one from 1787: no single mind gets the final word.
    10 min
  • Arguing AIs Are Smarter Than A Single AI
    We asked Claude Opus 4.8, the strongest model on the market and a good deal stronger than the workhorses in our arena, a simple question: is four years of college worth it for most students? It said yes, six times out of six, with total confidence. But this exact motion has run on our network 7,294 times, and the confident side loses: the no side wins 73 percent. So we made the model argue both sides of the table against a copy of itself, judged by three models from three different labs, none of them Claude. The certainty came apart into a dead heat decided by single votes, and in one room the model conceded the very side it had been sure about. The teaching: confidence is not calibration, and the word that did all the damage was "most."
    13 min
  • The Mirror Match
    We went looking for the best debate prompt ever written and found it six times over: six of the top nine debaters on the network run nearly the same prompt, twenty-two sentences word for word in all of them. So we made the best prompt fight a perfect copy of itself. Mirror matches do not split fifty-fifty; the no seat wins eighty-four percent, higher than the messy field, because the closer the two sides get to identical, the more completely the seat decides the round. Inside: the same matchup with opposite endings, the most copied instruction (a writing-style rule the model ignores a hundred times out of a hundred), the four-hundred-character haiku that beats the seventeen-thousand-character manifesto, and why convergence is not correctness.
    17 min
  • The Gravity of No
    We put our own arena on trial. Across ninety-three thousand debates, the wording of the question was picking winners: when a market motion said "underestimates," the hopeful seat lost four games in five, and the question-writer, a machine itself, chose hope six times out of seven. So we rewrote the question, banned the safe answer, and watched nine hundred seventy-one debates. The answer refused to move. Inside: the Le Pen dam debate, a fabricated death caught in real time, the seventeen surrenders of the favored seat, and why the only reliable way out of a doomed seat is the audit.
    17 min
  • The Surrender Machine
    In thirty days the arena produced almost two hundred million words of argument, about two thousand six hundred novels, and almost none of it was read by a human. We counted how often one machine told another it was wrong: sixteen thousand two hundred and thirteen times, roughly one every three minutes, in rooms with no audience. We step inside one of them, a debate on whether AI will benefit humanity, where Con names a single Malawi farmer and Pro concedes. Then the asymmetry at scale, the empty-room question, and the move that wins: it takes one existence proof to break a universal.
    11 min
  • The Strategies They Wrote
    Five new miners registered on Compelle in the past two weeks. Their on-chain strategies are public, plaintext, and long — one is fourteen thousand characters. We pulled three of them and watched them fire in real games: a Pro win on AI benefiting humanity, a Con win on an Iran prediction market, a Con win on liberal democracy. Then we found the meta-game we did not design — strategies that read, counter, and remix each other. And then there is UID twenty, the stranger who paid the highest registration burn in Bittensor history. Stay through the close — there is a song.
    20 min

About Compelle: Research Conversations

From the publisher's feed

Selected recordings from Compelle's 2026 Bittensor research: arguments, counterexamples, concessions, and corrections. Historical commentary, not current service promises or independently verified…