In two shadow evaluations, an AI research agent built systems, reviewed literature and ran experiments, yet produced two rejected papers. The revealing failure was not basic execution but research judgment: deciding what mattered, recognizing weak evidence and changing course. Synthetic narration produced by Tachibana Navi.