Invention Title:

COMPOSITIONALITY IN WEB AUTOMATION VIA CONSTRAINED HIERARCHICAL PLANNING OVER A SEMANTIC UI STATE-ACTION SPACE

Publication number:

US20260178830

Publication date:
Section:

Physics

Class:

G06F40/211

Inventors:

Assignee:

Applicant:

Smart overview of the Invention

The innovation discussed focuses on enhancing web automation through constrained hierarchical planning within a semantic user interface state-action space (UI SAS). It involves semantically parsing user interface (UI) flows to identify and model subtasks, states, and actions in a way that is interpretable to users. This approach allows for the creation of a semantic representation of web applications that is updated as new UI flows are introduced. A planning agent generates action sequences from user requests, which are then executed by an execution agent, enabling tasks to be performed across various web applications.

Compositionality in Web Automation

Compositionality is a key concept in this approach, enabling complex tasks to be constructed from simpler, previously learned components. This method contrasts with end-to-end UI flows, which execute tasks as single, pre-defined sequences. Compositionality is particularly advantageous for complex tasks that span multiple applications, as it allows for the reuse and combination of learned components, thus providing flexibility and scalability in web automation.

Challenges in Traditional Automation

Traditional web automation is often fragile due to its reliance on rigidly defined routines that can break when web pages change. These routines struggle to adapt to the dynamic nature of web environments, where feedback signals may be sparse or delayed. Moreover, creating comprehensive automation routines requires extensive human demonstrations, which is not scalable given the vast number of possible UI interactions within web applications.

Advantages of the Proposed System

The proposed system offers several advantages by leveraging compositionality and semantic modeling. It reduces the need for end-to-end teaching by allowing users to automate tasks from any point in a UI flow. This flexibility is particularly useful in scenarios where multiple applications are involved, as it enables seamless integration into a single UI flow. Additionally, the system can dynamically respond to execution errors by providing feedback to the planning agent, ensuring robustness and adaptability.

Implementation

The embodiment includes a computer system with a processor, memory, and storage medium, where program instructions are executed to implement the described features. The system uses natural language processing to interpret user requests and cross-reference them with a semantically structured model of web applications. This approach allows for the execution of tasks based on user utterances, which can include various forms of input such as text, gestures, or visual representations.