A 3B parameter unified decoder-only transformer that interleaves vision, text, and action tokens for perception, planning, reasoning, and control in a single model. Trained on the 1.5M-sample EO-1.5M dataset with strong results across manipulation tasks and benchmarks.