Everyone argues about whether Chinese LLMs are censored, but almost no one asks how we actually know. This episode unpacks the validated benchmarks—CHiSafetyBench, SafetyBench, ChineseSafe, FLAMES, JailBench, and the PNAS Nexus longitudinal study—that researchers use to measure political refusal. We explore the three different things "censorship" can mean, why multiple-choice tests inflate safety scores, how the CAC's Clear and Bright campaign drove refusal rates above 98%, and the growing arms race between models that produce evasive responses and the detectors trying to catch them. If you want to understand the measurement itself—not just the headlines—this is the episode.
Episode #072990 — open it directly at myweirdprompts.com/072990