US20260229224
2026-08-06
Physics
G10L15/063
Techniques are introduced to enhance automatic speech recognition (ASR) systems for code-switched speech, especially in multilingual contexts like Hindi-English. Two primary methodologies are employed. The first combines limited code-switched data with monolingual data, generating code switches during training. The second uses language models to generate code-switched utterances, which are then converted into speech using text-to-speech (TTS) synthesis. Both methods improve ASR performance without relying heavily on real-world code-switched data and are adaptable to various languages and systems.
The field of this innovation is automatic speech recognition, focusing on systems that handle code-switched speech. Code-switching, the practice of alternating between languages, is common in multilingual regions but poses challenges for ASR systems. Current systems excel in monolingual recognition but struggle with code-switched speech, often failing to detect language changes accurately. This issue is pronounced in Indic languages where code-switching is dynamic. Traditional methods like speaker diarization are ineffective due to latency and segmentation issues. The scarcity of labeled training data further complicates the development of effective code-switched ASR systems.
The detailed description outlines methods for training speech recognition models using synthetic data. One approach involves creating synthetic code-switched audio from monolingual data by selecting and combining audio samples from two languages. This generates training examples that simulate real-world language mixing. Another approach uses language models to generate code-switched N-grams and sentences, which are then used to fine-tune ASR models, particularly the decoder component. These methods address the challenge of limited training data and improve ASR systems' ability to recognize code-switched speech.
The process begins with receiving audio samples in two different languages. During training, subsets of these samples are selected and combined to create synthetic code-switched audio. This approach simulates real-world scenarios where speakers mix languages. The synthetic samples are evaluated against training constraints to ensure their quality and effectiveness. Samples meeting these constraints are used to train the ASR model, enabling it to handle code-switching even with limited real-world data.
These methods significantly enhance ASR systems' performance in multilingual environments. By generating synthetic training data, the approaches allow models to learn code-switching patterns without extensive real-world data. This is particularly beneficial in regions where code-switching is prevalent. The techniques offer a systematic way to adapt ASR models, improving their accuracy and reliability in recognizing mixed-language speech. The adaptability of these methods to various language pairs and systems underscores their potential impact on ASR technology.