Researchers have released Astronex-World 1.0, an open controllable video world-model foundation that generates future visual states from text prompts or initial images. The model predicts what comes next in a scene while allowing precise control over camera trajectories, continuous actions, and text events inserted at specific points during generation. Details are available in the paper on arXiv.
Unlike standard text-to-video models that simply render a fixed output, Astronex-World 1.0 treats video generation as a prediction problem over time. Given an initial observation and a set of controls, it forecasts subsequent frames under specified camera movements and embodied actions. This makes it possible to generate coherent, temporally consistent video sequences that respond to both spatial and behavioral inputs.
The model comes in two variants: a bidirectional version for full-context generation and a causal version optimized for persistent, real-time rollout. Both are built on the Wan2.2-TI2V-5B prior and incorporate PRoPE to encode camera intrinsics and extrinsics directly into attention mechanisms. A 64-dimensional action stream provides fine-grained modulation across all Transformer layers, enabling smooth integration of movement and interaction signals.
Training proceeds through a five-stage curriculum. First, the model learns bidirectional camera control. Next, it acquires action control capabilities. Then, the architecture is converted to support block-causal generation using cross-block KV caching. Finally, a compact student model is distilled for efficient inference. This structured approach allows the system to maintain long-range coherence without sacrificing responsiveness.
For creators building interactive experiences, simulation environments, or dynamic content pipelines, Astronex-World 1.0 offers a new level of controllability. Instead of generating isolated clips, users can define how scenes evolve over time through explicit camera paths and action sequences. Text events can be scheduled mid-rollout, opening possibilities for narrative-driven generation or interactive storytelling.
Mina Labs hosts Astronex-World 1.0 and makes it accessible through its inference API. The cost to generate a three-second clip in 4K resolution is 35. This pricing supports iterative experimentation, allowing developers to test multiple camera angles, action profiles, or event timings without prohibitive expense.
Consider using it to prototype virtual environments where camera motion must follow a character’s perspective. Or generate synthetic training data for robotics, where knowing the intended action sequence improves downstream perception accuracy. It also enables rapid exploration of cinematographic ideas—placing a text prompt at frame 15 to trigger an event while the camera orbits a subject.
Astronex-World 1.0 does not replace existing video generation tools. It augments them by adding structure and predictability to the generation process. When temporal consistency and interactive control matter, this model provides a solid foundation. And because it lives on Mina Labs, integrating it into workflows is straightforward.
MINA LABS
Start creating free