Today, we are announcing Echo-2, our best frontier model for generating stunning worlds from text or image inputs.
Echo-2 takes as input an image or a text prompt and produces a highly immersive 3D environment. Rather than a video, our model generates a 3D-consistent scene representation that a user can interact with. Unlike sequential video models, which predict frames one after another, a process prone to high computational demands, geometry drift, and inconsistent outputs across time, Echo-2 generates a spatially persistent scene. The crucial difference lies in the underlying representation.
Our outputs are 3D-grounded by design and can be rendered on any device at real-time rates. This foundational capability of understanding and generating 3D-consistent worlds is the core mechanism that allows Echo-2 to effectively bridge the physical and digital realms. This enables the creation of high-fidelity digital clones of real-world environments, such as homes or factories. This can be achieved from only an input photograph, dramatically simplifying digital twinning for management, editing, and remodeling applications. In addition, it facilitates transfer of knowledge from the virtual to the physical world, allowing robots to be trained in highly realistic simulated environments before acting in reality. This two-way bridge unlocks transformative workflows for industries spanning robotics, architectural visualization, digital twins, and immersive 3D content creation experiences.
Interactive 3D Worlds
Echo-2 is a frontier model that achieves stunning visuals, setting a new state of the art. Given an input image or text prompt, the model predicts a physically-grounded 3D representation of the scene, capturing both geometry and appearance in a consistent spatial layout with high-fidelity appearance. This internal representation is then converted into a renderable format suitable for real-time exploration. For the web demo, we use 3D Gaussian Splatting (3DGS), which provides extremely fast, GPU-friendly rendering and makes interactive viewing possible directly in the browser, even on modest hardware.
3D Worlds for Gaming
Celestial
Echo-2 is a powerful engine for game development, enabling the generation of fully interactive 3D environments directly from high-level descriptions. This drastically reduces the manual effort required for level design and world-building. Dynamic characters can easily be integrated into these generated worlds. Our interactive web demo showcases a walkable character navigating a 3D environment in real time, demonstrating that these are not mere video loops but persistent, navigable spaces. This capability allows developers to rapidly prototype gameplay mechanics and interactive experiences within minutes of 3D world creation.
Robotics Training
We can create digital environments for simulation and training data generation. This world simulation can power large-scale training of robots. In fact, Echo-2 goes even a step further by enabling the creation of scalable digital clones of specific environments, such as a home or factory hall, from minimal input like a single photo.
By eliminating the need for expensive and cumbersome 3D scanning hardware, our model makes it possible to rapidly generate high-fidelity virtual worlds for any space. This capability is critical for embodied AI, as it allows robots to plan, predict, and train in a virtual world before acting in the real one, a process known as Sim2Real knowledge transfer. Through this per-environment training, Echo-2 provides a foundational layer for robots to understand and safely operate within any physical environment.
Scene Understanding
Echo-2 provides a detailed understanding of the environment by predicting high-fidelity scene semantics. Our model decomposes the 3D world into its constituent parts, generating precise semantic segmentation masks that identify specific components such as chairs, tables, floors, and walls. This granular representation enables more localized and sophisticated scene edits, allowing for the manipulation of individual objects while maintaining the global consistency of the spatial layout.
Scene Editing
Echo-2 enables stunning, high-fidelity scene editing by leveraging text prompts along with a precise scene-object decomposition to remove, add, or replace objects within a 3D environment. This capability unlocks transformative workflows for downstream applications such as interior design, building planning, and architectural visualization by allowing professionals to remodel and restyle spaces with geometric accuracy.
Virtual Staging
Explore Echo-2's editing capabilities as a simple room is progressively staged with new objects and furniture that blend seamlessly with the scene's existing style.
Style Transfer
Beyond adding or removing objects, Echo-2 can also holistically restyle an entire scene. This makes it easy to explore alternate looks, moods, and design directions for the same space.
Building Design and Architecture
Echo-2 has broad applications in the architecture and real estate industries by pioneering the generation of entire, high-fidelity 3D environments. For instance, from a 2D floor plan, which acts as a global anchor for the scene's spatial layout, the model generates a fully consistent and navigable 3D scene. This capability dramatically accelerates workflows that were previously manual and time-consuming.
For architects and space planners, Echo-2 provides instant site twins and AI-driven scene generation, converting flat blueprints into 3D models. In real estate, this means replacing costly manual photography with automated, interactive 3D walkthroughs and enabling instant virtual staging for millions of listings, helping buyers to digitally remodel and visualize spaces before purchase.
Comparison to State-of-the-Art Methods
Echo-2 sets a new state of the art, outperforming competing models across a rigorous suite of benchmark evaluations. In our quantitative evaluation on the WorldScore benchmark for world generation [Duan et al., arxiv'25], Echo-2 significantly outperforms World Labs' latest Marble-1.1 model across the three main scene-quality metrics: Content Alignment, which measures how well the generated scene matches the input prompt; Subjective Quality, which captures the scene's aesthetics; and World Score, which summarizes overall scene quality.
What's Next
Echo-2 represents a major leap, but it is only the first critical step toward realizing the ultimate goal: bridging the physical and digital worlds. The next big frontier for us is to bring in a new level of intelligence by enabling the transfer of knowledge from simulation to reality and facilitating Sim2Real training pipelines.
This is fundamentally rooted in the necessity of physical grounding for generated scenes, compelling the model to move beyond mere pixels or abstract representations to natively understand space, distance, and metric size. Echo-2's success lies in establishing the foundational 3D, spatial, and view consistency required for this grounding.
By allowing the generation of editable digital clones of environments, from a text prompt or a few photos without expensive scanning, we enable immediate downstream applications like remodeling, interior design, and digital twin management for homes and factories. However, 3D consistency is only the first step; the ultimate challenge, and our next frontier, is to introduce temporal and physical consistency.
Future versions of the model will incorporate dynamics and physics-based reasoning so that scenes not only look real but also feature physics-based behavior, allowing for interactive simulations and advanced robotics training.
As Richard Feynman said, "What I cannot create, I do not understand"; this capability to generate and simulate the physical world is the missing layer that will ultimately enable embodied AI systems to understand and operate within reality.
