Abstract Embodied agents must continuously adapt to the physical world using interaction data collected across varying timescales, controllers, and environmental conditions. However, standard reinforcement learning assumes stationary dynamics and on-policy data collection, which breaks down when agents must learn from heterogeneous data sources with non-stationary dynamics and off-policy behavior.