A 3B-parameter open-source unified decoder-only transformer for general robot control that jointly handles perception, planning, reasoning, and action via interleaved vision-text-action pretraining on a 1.5M-sample dataset. Achieves strong results across manipulation tasks and benchmarks with fully open-sourced model, code, and dataset.