US20260260429
2026-09-03
Physics
G06T19/00
Methods, apparatus, and systems for text-guided virtual try-on (VTO) of makeup are discussed, focusing on using text inputs to guide makeup applications in a virtual environment. The system extracts text embeddings from user-provided text inputs that describe makeup features. These embeddings are then translated into specific makeup features that a VTO system uses to render an output image. This image shows the user with the makeup applied, determined by the input text. Additionally, methods for training an AI model to facilitate this process are also presented.
The process involves using a pre-trained text encoder to extract embeddings from text inputs describing makeup features. These embeddings are mapped to makeup features using a mapping AI model. The VTO system then uses these features to apply a virtual makeup product to an input image. The system is capable of rendering the output image with the selected makeup applied, offering a more flexible and user-friendly alternative to traditional VTO systems that often require numerical inputs or offer limited product selection.
Training the AI model involves creating a dataset using a large language model (LLM) to generate training texts and corresponding makeup features. The LLM produces training texts with various makeup features, such as color, gloss, and wetness, which are then used to train the AI model. The training process involves minimizing a loss function that combines L1 Loss for color and Cross-Entropy Loss for gloss and wetness, ensuring the model accurately predicts makeup features from text embeddings.
The system comprises a pre-trained text encoder and a mapping AI model. The text encoder extracts embeddings from text inputs, while the mapping model translates these embeddings into makeup features. These features configure a rendering engine within the VTO system to apply the makeup product to an input image. The system supports various makeup products, including lipstick, eyeshadow, nail paint, hair color, and foundation, enhancing user experience by allowing text-based product selection.
Users interact with the VTO system by providing text descriptions of desired makeup features, enabling a more intuitive and expansive selection process compared to traditional methods. This approach eliminates the need for numerical inputs and expands the range of available products for virtual try-on. The system's ability to render images in real-time enhances user satisfaction, providing a seamless and interactive experience. This method leverages advanced AI and language models to offer a sophisticated, user-friendly solution for virtual makeup application.