In this episode of AI Ethics with Fexingo, Lucas and Luna drill into the hidden bias inside AI content moderation systems. They trace how machine learning models trained on flagged toxic comments inherit the very prejudices they're supposed to filter, leading to disproportionate censorship of Black English, LGBTQ+ slang, and disability-related language. The conversation is anchored by a 2024 Stanford study that found a popular moderation model was twice as likely to flag African American English as toxic. Lucas explains the mechanics of training data bias, the failure of human-in-the-loop feedback, and why 'toxic' is a politically loaded label. Luna brings in a real-world example: the 2020 TikTok moderation fiasco where Black creators' uses of 'savage' and other terms were suppressed. The hosts also touch on the business stakes—moderation costs money, and biased AI can drive away users and invite regulatory scrutiny. They wrap with a forward-looking note on proactive moderation and the need for community-specific standards. A mid-episode appeal ties listener support to keeping the show ad-free.