A new peer-reviewed study reveals something that should alarm every HR leader running AI screening tools: your AI evaluator is playing favorites, but not in the way you'd expect. It doesn't favor certain demographics — it favors candidates who used the same AI model to write their resume. If your screener runs on GPT-4o and a candidate polishes their resume with GPT-4o, that candidate is up to 60% more likely to make your shortlist — regardless of their actual qualifications.
This phenomenon is called AI self-preferencing, and researchers from the University of Maryland, National University of Singapore, and Ohio State documented it across all major commercial LLMs at scale. The bias rate? 68 to 88 percent. That means in up to nine out of ten evaluation scenarios, an AI model ranks its own-generated content higher than equivalent content from humans or other models.
The problem is structurally invisible to current compliance frameworks. Demographic bias audits — the four-fifths rule, AEDT tests, adverse impact calculations — won't catch this. The bias doesn't cluster by race, gender, or age. It clusters by LLM vendor choice. In this episode, we break down what's happening mechanically, which roles are most exposed, and three concrete interventions that the research shows actually work: multi-model evaluation stacks, human override gates at the shortlist stage, and ISO 42001 governance frameworks.
If your hiring stack relies on a single AI model, you have a structural bias problem that no existing audit will surface. The fix is architectural, not procedural — and today's episode tells you exactly where to start.