Rho-1: Collapsing the multimodal stack

Rho-1: Collapsing the multimodal stack

An omni-reasoning model that understands, simulates, and acts.

An omni-reasoning model that understands, simulates, and acts.

Today, we are releasing a research preview of Rho-1, our 19B omni-reasoning model trained from scratch. Within a single neural network, it understands and generates text, images, and video, reasons over them, and takes actions.

Most AI systems today are agentic pipelines: a central model plans and delegates, handing jobs off to modality-specific specialists. Each handoff adds latency, and each specialist sees only a narrow request, not the full context. Rho-1 collapses that stack. Text, vision, and robotic actions are unified as tokens within a single context window.

Transcript of the recorded session. Turn 1. You: Draw a low aerial view of a tall red-and-white striped lighthouse on a rocky headland above the sea, bright midday. Rho-1, thinking: Okay, so the user wants a scenic coastal image featuring a prominent lighthouse. This is all about crisp maritime clarity and striking architectural contrast. The main creative decision for me is composition: balancing the grand lighthouse with the sweeping coastal landscape below. First, the … Shot plan: This image presents a low-angle aerial view of a striking coastal scene, dominated by a tall, red-and-white striped lighthouse standing on a rugged, rocky headland that juts into the sea. The lighthouse, a … It renders image:0. Turn 2. You: Put a box around the lighthouse. Rho-1, thinking: The user wants the lighthouse marked on image:0. I'll mark the lighthouse on image:0. It draws a box around the lighthouse. Turn 3. You: Animate it: the drone slowly flies toward the lighthouse. Rho-1, thinking: The user wants image:0 animated into a video where the drone approaches the lighthouse. I'll generate video:0 starting from the frame of image:0 and animate the drone moving steadily forward, descending slightly as it approaches the tall … Shot plan: An aerial drone slowly and steadily approaches a tall, red and white striped lighthouse standing prominently on a rugged, rocky outcrop, revealing the surrounding ocean and coastal rocks as it descends … It renders video:0. Turn 4. You: Edit the video: make it a heavy snowstorm, with white snow settling on the rocks and grass in natural colors. Keep the lighthouse, the house and the camera move exactly the same. Rho-1, thinking: The user wants video:0 edited into a heavy snowstorm while keeping everything else intact. I'll generate video:1 from the same scene as video:0: a lighthouse on a rocky cliff above the ocean, with a small white building beside it. The … Shot plan: Edit video:0: keep the same scene, lighthouse, white building, and camera movement, but make it a heavy snowstorm with white snowflakes falling steadily, covering the rocky cliff face, grassy vegetation, and … It renders video:1. Turn 5. You: What changed between the two videos? Rho-1, thinking: The user asks what changed between the videos. Answer in text: the first video was clear, sunny, with falling snow; the second is a snowy storm version. Answer: The first video shows a clear, sunny day with the lighthouse, white house, and blue ocean. The second video is a snowy version of the same scene, with snow falling across the sky and accumulating on the rocks, cliff, and grass, and the overall tone is muted and overcast.

Jump to any step
Every response comes from the same model in one conversation. It is not an agentic system. There is no tool call and no second model: text, boxes, images and video are all tokens read from and written into one state. Generation waits are shortened for playback, and denoising is drawn as the distilled model’s 8 passes.

We believe this architecture points toward the future of intelligence: not a patchwork of narrow models wired together, but a single foundation that handles the entire loop. Rho-1 can imagine an environment, track its state, answer questions about it, and set it in motion without ever dropping the thread. With no seams between components, the whole system can be optimized end-to-end, making it a stronger foundation for any application that spans modalities. We see this kind of unification as the core prerequisite for physical AGI.

What Rho-1 produces is not a one-off output but an evolving world state. Steer it with dialogue, and you get fluid conversation. Steer it with real-time controls, and you get a persistent, interactive simulation. Steer it with motor policies, and you get embodied robotic control. This post explores each of these three frontiers in turn.

Rho-1 as a multimodal assistant

Figure 1 shows a single, unedited session with Rho-1. Across five turns, it draws a static scene, locates an object within it, animates the scene into a video, edits the video to change the weather, and explains what changed, reasoning through each step along the way. Work that typically spans several specialized systems is handled within a single model.

Watch the state in the model view grow with each turn. A request to generate a picture routes to the diffusion tower, while a request for a bounding box stays on the understanding side. Either way, every turn reads from and writes to the same KV cache, which acts as a persistent world state:

  • Reasoning: The block above each reply shows the model’s chain of thought. There is no bolt-on prompt enhancer: the same model plans the shot and renders it.

  • Grounding: Rho-1 generated the lighthouse, so the scene is already in its context. The bounding box is emitted as coordinate tokens attending to that state, with no secondary detection model.

  • Animating: The first video frame isn’t a re-encoded image. It’s the original representation held in context, so lighting, geometry, and object identity carry through.

  • Understanding: When asked what changed between clips, Rho-1 isn’t running post-hoc captioning on exported frames. It reads the latent state that produced the change, so its answer is grounded in what it actually generated.

Rho-1 is a fast model. The base Rho-1 model generates video at 0.79× real-time (median), with a watchable stream starting in roughly 6 seconds.

Multi-agent pipeline13.8 seconds, illustrative
route
expand prompt
queue
generate
return
Rho-17.0 seconds, measured
plan
generate
Bars are to scale with each other; the pipeline row is illustrative.

A distilled variant cuts the denoising trajectory from 99 steps to 8 with minimal quality loss, pushing the system into interactive territory. In our internal testing, it is among the fastest models in the world in every modality we measured. It returned a 5.3-second video clip in about a second, faster than any other video model we timed. It made images as quickly as the fastest dedicated image models. And it was the quickest of the models we tested to the first token of a text response.

A world can now be imagined, interrogated, edited, and carried forward within the window of human conversational patience.

Rho-1 Flash edits a video by prompt with a fixed seed. Version 1: a golden retriever wearing a red santa hat sitting in the snow in a pine forest, rendered in 2.01 seconds. Version 2: a black labrador wearing a red santa hat sitting in the snow in a pine forest, rendered in 1.11 seconds. Version 3: a black labrador wearing a red santa hat and a red scarf sitting in the snow in a pine forest, rendered in 1.11 seconds. Version 4: a black labrador wearing a red santa hat and a red scarf sitting in the snow in a pine forest at night with northern lights, rendered in 1.11 seconds. Version 5: a husky puppy wearing a red santa hat and a red scarf sitting in the snow in a pine forest at night with northern lights, rendered in 1.15 seconds.

The shot
1.15 sfrom prompt to video
A distilled Rho-1 enables near-instant video editing through minor prompt modifications.

Rho-1 as a steerable real-time world model

Fast generation alone just yields a faster render farm. But generation that is both real-time and controllable unlocks something fundamentally different: a world model you can simulate and steer.

Rho-1 streams continuously, clip after clip, each picking up exactly where the last left off, so time inside the simulation never stops. New instructions can arrive at any moment. A command enters through the understanding stream, updates the underlying state memory, and moments later the generation stream renders the new trajectory. The world shifts without a single cut.

One continuous real-time session, opened as “green and red ball sitting on a counter top”. A robot arm comes in and grips the red ball when told to, lifts it, then sets it back down and lets go when told to put it down. The panel shows each instruction exactly as it was sent, at the time it was sent.

Because generation is a continuous rollout, a scene can be monitored live and steered at key moments:

Shared · one openingAbank leftBbank right
One opening, two continuations. Left, the rollout told to “bank left” half a second in; right, told to “bank right”. Everything before the split is shared: the first half-second is the same frames on both sides. Both continuations are valid and could be used while planning next steps.

Real-time simulation enables new applications.

Generative environments on demand: An environment that advances under live input is no longer a static asset pipeline, it becomes an interactive simulation. Instead of manually engineering edge-case 3D worlds to benchmark autonomous driving or train embodied robotic policies, developers can spin up reactive environments and perturb them on the fly.

Infinite, steerable livestreams: Running real-time understanding and generation within the same attention context dissolves the boundary between viewer and director. You can watch an uninterrupted stream, adjust the physics or weather mid-frame, inspect object state, and extend the horizon indefinitely, all without a restart.

Rho-1 as a world-language-action model for robotics

Because physical actions and future video frames decode from the same latent state, Rho-1 allows a policy to mentally simulate the next few seconds of an environment and output the trajectory to enact it.

Planning is no longer an external planner querying an auxiliary world model. The network predicting future camera observations is the exact network dictating joint actuation. With actions treated natively as continuous tokens alongside text and pixels, the model does not require an ad-hoc robotics wrapper; the action channels share the core attention space.

One episode with every channel exposed, on a LIBERO simulation task: the observation, the wrist view it predicts against the one it is given, and the seven action channels it emits.

Even under pure text conditioning, Rho-1 exhibits an intuitive grasp of physical contact and rigid-body mechanics. Grippers secure objects that remain solid rather than morphing or clipping into surfaces:

In practice, to scale beyond scarce teleoperated robot logs, Rho-1 can be paired with our Inverse Dynamics Model. While most internet video lacks logged steering angles or joint torques, an IDM observes raw video and infers the underlying control signals. This unlocks web-scale video as action-labeled pretraining data: Reka IDM reconstructs the implicit actions, and Rho-1 ingests those actions as native tokens within its multi-modal state space.

Limitations

The demos above show what Rho-1 can do today, but real-world deployment requires being candid about the failure modes:

  • Long-Horizon Drift: Extended rollouts degrade structurally before they degrade visually. A 30-second stream may keep photorealistic texture and fine detail while drifting into a structurally incompatible room layout.

  • Grounding Across Time: Object detection and coordinate grounding work reliably on static images, but not yet across video.

  • Editing Stability: Targeted visual editing remains nascent. Our conversational demo shows how an edit works, but consistency across diverse prompts is brittle.

  • Resolution: Native video rollouts are currently capped at 672×384.

We believe most of these failures can be attributed to the modest scale of our training compute and data, and we expect them to narrow as we scale both.

One model: a symmetric architecture

Most multimodal systems are asymmetric. Vision-Language Models (VLMs) ingest text, images, and video, but only emit text. Video diffusion models take in multimodal conditioning, but only emit pixels. This input-output imbalance confines each model to half the interactive loop.

Rho-1 restores full symmetry. Every signal entering or leaving the network belongs to one of two native formats:

  • Discrete tokens encode text, symbolic reasoning, and high-level commands.

  • Continuous tokens encode image latents, video frames, robotic actions, and proprioception.

Because outputs share the same formats as inputs, anything the model generates can be fed straight back into its context. And because each modality keeps its natural form, Rho-1 never has to quantize rich continuous dynamics into lossy discrete tokens, or force language into continuous approximations.

video
image
text
actions
proprioception
Rho-1Predicts what happens next
futurevideo
futureimage
futuretext
futureactions
futureproprioception
The same five channels on both sides. Predicting the next frame, the next instruction and the next action are one operation, not three models sharing a bus.

Within each transformer block, compute is split across two expert weight streams: an understanding stream for language and visual parsing, and a generation stream that denoises latents into images and video. Crucially, these streams do not run in isolation:

  • Shared Attention and State: Tokens from both streams attend to one another, operating over the same KV cache. User instructions, chain-of-thought traces, prior video frames, and emitted motor policies all accumulate in one context.

  • Zero-Latency Handoffs: The understanding stream decides the output modality. When a response requires visual synthesis, it emits a discrete handoff token, and the generation stream immediately begins rendering from the full accumulated state.

State the KV cache text:0 user text:1 assistant image:0 assistant text:2 user text:3 assistant video:0 assistant text:4 user Text reasoning, answers, boxes Decoder latents to pixels Reasoning, shot plan, answers<|content_media|> hands off Noise to latentsin 8 passes(99 for the base model) Text head predicts the next token Diffusion head predicts the way back to clean Block 1 Block 1 Block 2 Block 2 Block 3 Block 3 Block 4 Block 4 Block 5 Block 5 ⋮ ⋮ Block N Block N shared attention Understanding stream Generation stream Text tokens yours and its own Visual tokens images, video Noisy latents the image or clip being made ×8
Inputs at the bottom, outputs at the top. Both streams sit in every block and share attention; on the left, the state, colored by the stream that wrote each entry. The generation stream loops from noise to a clean result, in 8 passes for the distilled model and 99 for the base model.

Rho-1 is optimized end-to-end under two concurrent objectives:

  • Next-token prediction for discrete sequences.

  • Flow matching for continuous image, video, and action generation.

Because both streams share attention and a common representation space, improvements do not stay siloed. We hypothesize that when the model sharpens its spatial reasoning during text pretraining, its video consistency improves. When it learns richer physical dynamics from embodied trajectories, its real-time steering becomes more grounded. Progress compounds across the entire network at once.

Early on the compute curve

Every result in this post comes from a checkpoint trained from scratch on just 320 H100 GPUs for three months, a tiny fraction of the compute behind modern frontier language and video models. Rho-1 is a functional proof-of-concept and an architectural direction, not a finished product.

To us, that modest compute budget is the most compelling takeaway. The failure modes highlighted earlier, from structural drift and limited temporal grounding to brittle edits and lower native resolution, are the kinds of problems that have historically improved with scale.

More importantly, scaling a unified architecture yields compounding returns. In a fragmented pipeline, additional compute must be divided among an LLM, an image generator, a video backbone, and an action policy. In Rho-1, every added FLOP goes into one model, so gains in reasoning, visual generation, and physical control can reinforce one another rather than being trained in isolation.

Infrastructure for serving omni models

A unified architecture fundamentally shifts the systems economics of inference. When reasoning, multimodal generation, visual grounding, and motor control collapse into a single set of weights, conventional LLM serving frameworks, designed for autoregressive text generation over a growing KV cache, are no longer enough.

We pursue this research not only to explore frontier model designs, but to build the runtime systems required to serve them at scale. Interleaving continuous diffusion trajectories with discrete autoregressive decoding exposes unprecedented runtime bottlenecks:

  • Heterogeneous Memory Layouts: High-throughput discrete KV caches must coexist in HBM alongside transient continuous latent states without causing fragmentation.

  • Shifting Compute Intensity: Workloads swing between compute-bound dense attention during visual synthesis and memory-bandwidth-bound autoregressive decoding during text generation.

  • Hard Real-Time Latency: Interactive steering leaves little room for pipeline bubbles. The system must ingest new instructions and update trajectories within fractions of a second.

Every finding from Rho-1 feeds directly into our production serving engine on Reka Cloud. The choices behind Rho-1, from token representation to kernel scheduling, shape how we optimize memory bandwidth, schedule execution, and co-locate multimodal workloads in production. Scaling unified models is only half the equation. The other half is building the runtime infrastructure that makes them fast, reliable, and economically viable to deploy.

Sample gallery

Unedited, single-pass outputs across diverse domains, presented at Rho-1’s native runtime resolution. 60 clips, 10 domains, one checkpoint with no per-domain fine-tuning. Each clip represents an active latent state, an entry point into an environment the model can simulate, steer, and carry forward indefinitely.

Real-world driving

05 clips

Rho-1 simulates real driving conditions from the driver’s seat, on highways, city streets at night, and desert and mountain roads.

Video games

17 clips

Rho-1 renders and moves through game worlds in every style, from photoreal open worlds to voxel and low-poly.

Robotics

05 clips

Rho-1 sees from a robot’s point of view: warehouse vehicles, quadrupeds, rovers and industrial arms at work.

Aerial

05 clips

Rho-1 simulates aerial footage for drone navigation and survey, the flight data that autonomous systems train and test on, without flying a mission.

First person

04 clips

Rho-1 simulates egocentric footage and the movement behind it, the first-person view an embodied agent sees as it rides, skis or paddles.

Weather

05 clips

Rho-1 simulates real weather: supercells, blizzards, sandstorms, storm surf and rolling fog.

Nature

05 clips

Rho-1 renders the natural world in motion, from reefs and waterfalls to wildlife and the aurora.

Industrial

05 clips

Rho-1 models industrial sites and processes, including foundries, data centers, mines and production lines.

Urban

04 clips

Rho-1 moves through cities: rain-soaked streets, alleys, rooftops and transit.

Interiors

05 clips

Rho-1 keeps indoor spaces consistent as the camera moves through libraries, kitchens, corridors and cathedrals.

Get in touch

Rho-1 is currently available as a research preview.

If you are building frontier applications—embodied robotics, interactive simulation environments, or closed-loop vision-action systems—or are looking to deploy optimized omni architectures at scale, we’d love to collaborate: contact@reka.ai.

Reka Cloud

Interested in partnering?

What we learn building Rho-1 goes straight into our serving stack on Reka Cloud. If you want to train, customize or deploy multimodal models on infrastructure engineered for performance, efficiency and scale, tell us about your project.

Explore Reka Cloud

Further reading