US20260260158
2026-09-03
Physics
G06N20/00
The patent application describes innovative methods and apparatuses for adaptive model compression and decompression tailored for inference tasks, specifically designed for dynamic resource-based strategies in networks and systems. It focuses on optimizing performance in edge nodes, which are part of modern networks like 5G and 6G, by addressing the constraints these nodes face in terms of processing power, memory, and energy. Such optimization is crucial as AI/ML models increase in complexity and demand efficient execution in resource-limited environments.
The approach involves identifying resources available at an edge node and the requirements of the communication services it facilitates. Based on these identifications, a suitable compression strategy is determined and applied to the AI/ML model. This results in a compressed model that is efficiently managed and provided to the edge node, ensuring optimal performance without overburdening the limited resources available.
The system processes real-time network and resource data to predict congestion, mobility, and traffic loads. This information helps in deciding the appropriate compression to apply to models that support communication services. The compressed model is then transmitted to the network node, where it is decompressed and used for generating inferences, enhancing the system's responsiveness and efficiency.
The application outlines a comprehensive network infrastructure, including various access types like broadband, wireless, voice, and media. It describes components such as access terminals, base stations, switching devices, and media terminals, all interconnected through a network comprising both wired and wireless links. This infrastructure supports a broad range of devices, from mobile phones to telephony and display devices, ensuring wide applicability of the compression methods.
Addressing challenges such as resource constraints, dynamic network conditions, and energy efficiency, the disclosed methods offer scalable, efficient solutions for model compression. They also explore lossless versus lossy compression, inter-node coordination, and distributed inference generation. The integration of generative AI for further optimization highlights the innovative approach to enhancing current network capabilities.