Reka AI Debuts Rho-1 Omni-Model for Video and Robotics
Reka AI has released Rho-1, a 19-billion-parameter omni-model that unifies text, video, and robot control in a single network to simplify how machines interact with the physical world.

Reka AI has launched a research preview of Rho-1, a new 19-billion-parameter omni-model designed to process and generate text, images, video, and robot control actions within a single neural network. Unlike conventional AI systems that rely on routing tasks to specialized external models or making tool calls, Rho-1 processes all of these diverse modalities as tokens inside a single shared context window. This unified architecture allows the model to generate continuous video in real time and adapt to new instructions on the fly without needing a system restart.
A key innovation of Rho-1 is that the exact same weights used to predict camera images are also responsible for driving physical robot movements. To overcome the persistent industry challenge of scarce robotic training data, Reka AI developed an inverse dynamics model. This auxiliary system extracts viable control signals directly from standard internet videos, allowing the model to learn physical dynamics from passive observation. Training the 19-billion-parameter model required running workloads on 320 Nvidia H100 GPUs for approximately three months.
For AI practitioners and robotics engineers, Rho-1 represents a significant shift toward true world models. By eliminating the latency and complexity of orchestrating multiple specialized models, developers can build more responsive, cohesive systems where visual perception and physical action are deeply integrated. This release follows Reka AI's April 2024 launch of Reka Core, a multimodal language model that positioned itself as a direct competitor to frontier models like GPT-4, Claude 3, and Gemini Ultra. Rho-1 pushes this multimodal expertise further into the physical domain, offering a more streamlined path for embodied AI development.
This is our own summary of reporting by The Decoder



