An awesome spoken LID repository. (Working in progress
-
Updated
Apr 22, 2024 - Python
An awesome spoken LID repository. (Working in progress
End-to-end spoken language identification out of the box.
Spoken Language Identification on Common Voice and AudioSet using Deep Learning
Tidy Tunes is an easy-to-use pipeline for mining high-quality audio data for speech generation models. To do so, it chains multiple open source models while minimizing dependencies.
PHO-LID: A Unified Model to Incorporate Acoustic-Phonetic and Phonotactic Information for Language Identification
A pipeline to isolate and transcribe one language in mixed-language speech
Source code of paper <End-to-End Language Diarization for Bilingual Code-switching Speech>
The official implementation of the method discussed in the paper Improving Spoken Language Identification with Map-Mix(work accepted at ICASSP-2023)
An Web application Language Identification project uses Pytorch and Torchaudio to accurately identify spoken language from audio files.
An object model to the Ethnologue project for Pharo
End-to-end audio ML pipeline for Spoken Language Identification (SLID) across 4 languages using MFCC/Chroma feature extraction, SVM, Random Forest, and LightGBM.
Spoken language identification and clustering across 4 languages (German, Italian, Korean, Spanish) using audio feature extraction (MFCCs, Chroma), SVM, Random Forest, GMM, and t-SNE.
Experiment configuration and metadata
This repository contains the files for the final project of the machine learning course.
End-to-end spoken language identification (German/Italian/Korean/Spanish) from raw audio: MFCC-based feature engineering, leakage-free group splits, KNN/RF/SVM benchmarks and unsupervised clustering analysis.
Spoken language identification DNN implemented in mxnet
A classical machine learning project for identifying the spoken language of audio recordings using engineered acoustic features. The project was developed as a final project for a Machine Learning course and includes both supervised classification and unsupervised clustering of multilingual speech.
Detects the spoken language in an audio clip using Meta's mms-lid-4017, with a microphone demo and a fine-tuning experiment. Built at GRN Hack 2025.
To associate your repository with the spoken-language-identification topic, visit your repo's landing page and select "manage topics."