US20260301749
2026-10-01
Physics
G10L17/24
The patent application describes a system and method for detecting synthetic voices generated by artificial intelligence (AI) during phone calls. This system employs a combination of passive and active techniques to analyze audio inputs from callers. The passive techniques involve extracting acoustic features, spectral patterns, prosodic elements, and lexical characteristics from the audio. Meanwhile, active engagement involves generating interactive challenges through an AI agent to assess the caller's responses.
The need to distinguish between human and AI-generated voices has become increasingly important with advancements in AI technology. Historically, the Turing Test was used to evaluate a machine's ability to mimic human intelligence. Today, AI systems like ChatGPT can convincingly pass such tests, particularly in text-based interactions. However, the integration of speech recognition and generation with language models has enabled AI to engage in real-time voice conversations, posing risks such as fraud and illegal robocalls.
The system's detection method involves receiving an audio input from a caller and analyzing it using passive detection techniques. It then actively engages the caller by presenting interactive challenges designed to probe response timing, content coherence, and complexity. The responses are processed using automatic speech recognition and natural language processing techniques. By combining results from both passive and active analyses, the system generates a composite synthetic voice detection score to determine if the voice is AI-generated.
The system comprises several components: a communications interface for receiving audio input, a processing system with a processor, and a memory storing executable instructions. Key modules include a passive detection module, an active engagement module, an automatic speech recognition module, and a natural language processing module. These modules work together to generate the composite detection score and make determinations about the nature of the caller's voice.
The system can be integrated into various communication networks, including broadband, wireless, and voice access networks. It supports a range of devices from mobile phones to telephony devices, and can be implemented in different network environments such as VoIP, IP, and optical networks. The system is designed to enhance security by identifying and mitigating the risks associated with AI-generated synthetic voices in telecommunication systems.