Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models. However, most existing approaches rely on large collections of robot demonstrations, which are costly. This survey covers scalable vision-language-action learning using human-centric data, particularly human videos, as an alternative to expensive robot demonstrations.