Invention Title:

QUERY-BASED GENERATION OF TRAVERSABLE THREE-DIMENSIONAL SCENES

Publication number:

US20260301328

Publication date:
Section:

Physics

Class:

G06T19/003

Inventors:

Assignee:

Applicant:

Smart overview of the Invention

Techniques for generating traversable three-dimensional scenes from text prompts enable the creation of interactive virtual environments. A text-based query specifying the visual aspects of a desired scene is processed using machine learning models to generate a spherical representation. This representation is then used to construct a traversable three-dimensional scene, allowing navigation and viewing from multiple perspectives. These techniques address limitations of existing methods, such as excessive generation time and lack of spatial consistency.

Background

Advancements in text-to-image generation have facilitated the creation of high-quality two-dimensional images from natural language descriptions. However, generating coherent 3D scenes from text prompts remains challenging. Conventional methods often require significant computational resources due to iterative refinement and complex optimization algorithms, resulting in computational inefficiency and limited scene diversity. These limitations restrict practical applications, as existing techniques struggle with spatial consistency and viewpoint flexibility.

Summary

The described techniques support the creation of immersive and navigable 3D virtual environments from textual descriptions. A processing device receives a text-based query and generates a spherical representation using a generation model. This intermediate representation includes panoramic frames depicting visual aspects from multiple viewpoints. A reconstruction model then constructs a traversable 3D scene based on this representation, outputting an environment that allows comprehensive view coverage and visual consistency.

Detailed Description

Generative AI technologies have diverse applications, but creating coherent 3D scenes is still limited. Conventional methods are computationally expensive and result in incomplete scenes with restricted viewpoints. The described approach leverages a two-stage pipeline involving spherical representation and scene reconstruction to improve visual consistency and navigability. For example, a user can generate a fully navigable 3D scene of an opera house with consistent architectural details and lighting.

Implementation

The process involves generating a spherical representation from a text query using a text-to-video diffusion model. This representation is then processed by a reconstruction model to build a cohesive 3D scene. The approach conserves computational resources by avoiding iterative refinement and complex optimization, resulting in reduced processing time and lower overhead. The generated environments offer improved visual consistency and navigability, overcoming the limitations of conventional techniques.