Spoken Language Identification (LID) is the computational task of automatically classifying the spoken language from raw acoustic utterances. This repository presents an end-to-end Machine Learning system for Spoken Language Identification across four major languages: German, Italian, Korean, and Spanish.
The project encompasses the complete data science lifecycle:
- Dataset Creation & Curation: Collecting audiobook narrations across languages and speaker genders.
- Audio Processing & Cleaning: Standardizing sampling rates (16 kHz Mono), trimming silent pauses via Voice Activity Detection (VAD at 20 dB threshold), and uniform 4-second segmentation yielding 10,862 clean speech samples.
- Spectral & Temporal Feature Engineering: Extracting 44 statistical feature representations per sample, combining 20 Mel-Frequency Cepstral Coefficients (MFCCs), Spectral Centroid, and Zero Crossing Rate (ZCR) aggregated over time by mean and variance.
- Data Leakage Prevention: Assigning unique
group_ididentifiers per source audio recording to guarantee strict train-test separation viaGroupShuffleSplit. - Supervised Classification: Benchmarking K-Nearest Neighbors (KNN), Random Forest (RF), and Support Vector Machines (SVM) on leakage-free splits, reaching up to 99.71% accuracy.
- Cross-Gender Domain Adaptation Experiment: Training exclusively on male speakers and testing on female speakers (34.13% accuracy) to evaluate speaker-dependence vs. language-invariant acoustic cues.
- Unsupervised Clustering & Analysis: Dimensionality reduction via Principal Component Analysis (PCA) alongside K-Means and Hierarchical Agglomerative Clustering (Ward linkage) to uncover intrinsic linguistic groupings without label supervision.
- Institution: University of Tehran, College of Engineering, Department of Electrical and Computer Engineering
- Course: Machine Learning (Final Project)
- Team Members:
- Amir Aref (Student ID:
810102506) - Parsa Saeednia (Student ID:
810102460) - Amirali Dehghani (Student ID:
810102443)
- Amir Aref (Student ID:
Language Identification forms the foundational front-end for multilingual speech processing pipelines, automated transcription systems, call-center routing, and real-time machine translation. The challenge lies in isolating language-specific phonemic and prosodic patterns from speaker-specific vocal traits (e.g., fundamental frequency, pitch, vocal tract geometry) and environmental acoustic conditions.
The dataset evaluates four phonetically distinct target languages:
- German: Germanic branch, heavy consonant clusters, high dynamic range.
- Italian: Romance branch, vowel-ended phonotactics, rhythmic syllabic cadence.
- Spanish: Romance branch, fast syllable-timed rhythm, distinct fricatives.
- Korean: East Asian language, agglutinative grammar, distinct stop consonant voicings.
Raw audio streams were compiled from 720 MP3 audiobook recordings across male and female narrators.
dataset/
βββ German/
β βββ Female/
β βββ Male/
βββ Italian/
β βββ Female/
β βββ Male/
βββ Korean/
β βββ Female/
β βββ Male/
βββ Spanish/
βββ Female/
βββ Male/
Audio processing was performed using librosa and soundfile:
- Sampling Rate Standardization: Resampled all recordings to 16,000 Hz (Nyquist cutoff frequency of 8 kHz, preserving human speech intelligibility while optimizing computation).
- Mono Downmixing: Mixed stereo channels to single-channel mono to eliminate spatial panning variations.
-
Voice Activity Detection (VAD): Trimmed silent pauses and low-energy background signal below a 20 dB threshold (
librosa.effects.trim(y, top_db=20)). -
Fixed-Length Slicing: Sliced continuous speech streams into non-overlapping 4-second segments (
$16,000 \times 4 = 64,000$ samples per audio clip).
| Language | Clean Audio Segments | Dataset Share (%) |
|---|---|---|
| Italian | 2,956 | 27.2% |
| German | 2,677 | 24.6% |
| Korean | 2,620 | 24.1% |
| Spanish | 2,609 | 24.1% |
| Total | 10,862 | 100.0% |
- Gender Distribution: 5,652 Male samples | 5,210 Female samples
- Source Audiobook Files: 720 unique source audiobooks
Because thousands of 4-second segments are derived from 720 source files, a naive random split would place adjacent segments of the same narrator in both training and test sets, leading to artificial accuracy spikes due to speaker memorization ("Clever Hans" effect).
To prevent leakage:
- Each segment receives a
group_idderived from its source file name. - Data splitting uses
GroupShuffleSplit, ensuring that all segments from any source audiobook reside strictly in either the training set or the test set, never both.
Speech waveforms (
-
Mel-Frequency Cepstral Coefficients (MFCCs):
- Extracted 20 coefficients. MFCCs map the linear power spectrum onto the non-linear Mel scale, mimicking the human auditory perception of timbre and phonemic formants.
- Dimensions: 20 Means + 20 Variances = 40 Features.
-
Spectral Centroid:
- Measures the center of mass of the frequency spectrum, representing perceived audio brightness.
- Dimensions: 1 Mean + 1 Variance = 2 Features.
-
Zero Crossing Rate (ZCR):
- Counts how rapidly the time-domain signal changes sign, helping distinguish percussive unvoiced consonants and fricatives from voiced vowels.
- Dimensions: 1 Mean + 1 Variance = 2 Features.
The extracted feature set is stored in features.csv (10,862 rows filename, group_id, and language).
- Split Ratio: 80% Training (8,785 samples) / 20% Testing (2,077 samples).
-
Splitter:
GroupShuffleSplit(n_splits=1, test_size=0.2, random_state=42). Zero group overlap verified. -
Scaler:
StandardScalerfitted on training set and applied to test set ($\mu=0, \sigma=1$ ).
| Classifier | Architecture & Parameters | Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|---|---|
| K-Nearest Neighbors (KNN) |
|
99.71% | 0.99 | 0.99 | 0.99 |
| Support Vector Machine (SVM) | RBF Kernel, |
99.71% | 0.99 | 0.99 | 0.99 |
| Random Forest (RF) | 100 Trees, Gini impurity | 99.13% | 0.99 | 0.99 | 0.99 |
-
KNN (
$k=5$ ): Near-perfect score indicates pure local feature neighborhoods and distinct cluster separation between languages. - SVM (RBF): Achieved 99.71% accuracy, confirming that high-dimensional non-linear transformation creates clear hyperplanes between linguistic spectral features.
- Random Forest: Achieved 99.13% accuracy, demonstrating ensemble decision tree robustness against variance in frame-level spectral energy.
To test whether classifiers learn language-invariant features (phonemes, rhythm, syntax) versus speaker-dependent traits (pitch, vocal tract length), we designed a domain adaptation experiment:
- Training Set: Male voice samples ONLY (5,652 samples).
- Test Set: Female voice samples ONLY (5,210 samples).
- Cross-Gender Test Accuracy: 34.13%
- Random Baseline: 25.00% (4 classes)
While 34.13% accuracy exceeds random guessing (indicating the extraction of some underlying language prosody), the significant drop from 99.71% demonstrates strong Speaker Dependence. Because each language in single-audiobook datasets originates from a limited speaker pool, models partially memorize individual narrator acoustics. True language identification across unseen speakers requires pitch-invariant normalization or multi-speaker corpora.
Unsupervised learning evaluates whether extracted features naturally cluster by language without using target labels during training.
- Variance Retention: 22 Principal Components account for 90% of total feature variance.
- 2D PCA Projection: Used for qualitative scatter plot visualization. Showed clear grouping with partial acoustic overlaps.
-
Cluster Count Selection: Evaluated
$k \in [2, 10]$ using the Elbow Method (Inertia reduction curve). - Silhouette Score: 0.1412 on standardized 44D feature space.
- Cluster Purity Matrix (Normalize by Index):
| Cluster ID | German | Italian | Korean | Spanish | Dominant Language |
|---|---|---|---|---|---|
| Cluster 0 | 59.3% | 10.6% | 0.0% | 30.1% | German |
| Cluster 1 | 7.2% | 62.4% | 11.5% | 18.9% | Italian |
| Cluster 2 | 0.0% | 0.7% | 0.1% | 99.2% | Spanish |
| Cluster 3 | 41.1% | 1.6% | 54.3% | 3.0% | Korean |
- Dendrogram: Plotted on a 1,000-sample subset using Ward linkage to minimize intra-cluster variance. Showed distinct branches at four clusters.
- Silhouette Score: 0.1176 on full dataset.
- Cluster Purity Matrix (Normalize by Index):
| Cluster ID | German | Italian | Korean | Spanish | Dominant Language |
|---|---|---|---|---|---|
| Cluster 0 | 38.6% | 27.3% | 8.3% | 25.8% | Mixed (German/Italian) |
| Cluster 1 | 24.1% | 0.2% | 75.4% | 0.3% | Korean |
| Cluster 2 | 0.0% | 0.0% | 0.0% | 100.0% | Spanish (100% Pure) |
| Cluster 3 | 0.0% | 100.0% | 0.0% | 0.0% | Italian (100% Pure) |
ML-Final-Project/
βββ Data_Cleaning_and_Feature_Extraction.ipynb # Step 1: Preprocessing, VAD, slicing & feature extraction
βββ Classification.ipynb # Step 2: GroupShuffleSplit, KNN/RF/SVM training & cross-gender test
βββ Clustering.ipynb # Step 3: Standardization, PCA, K-Means & Hierarchical clustering
βββ Evaluation.ipynb # Step 4: Final evaluation metrics, confusion matrices & plots
βββ features.csv # Extracted 44-dimensional dataset (10,862 samples x 47 columns)
βββ README.md # Main project documentation
βββ models/ # Trained model binaries & evaluation artifacts
β βββ knn_model.pkl # Serialized KNN model
β βββ rf_model.pkl # Serialized Random Forest model
β βββ svm_model.pkl # Serialized SVM model
β βββ X_test_scaled.pkl # Scaled test dataset
β βββ y_test.pkl # Ground truth test labels
β βββ label_encoder.pkl # Target label encoder
βββ Report/ # Project report documentation & source
βββ ML_Final_Project_Phase2_Report.pdf # Final compiled report PDF
βββ main.tex # Main LaTeX report document
βββ setting.tex # LaTeX package & style configurations
βββ font/ # Font assets
βββ img/ # Visual assets & logos
Clone the repository and install the dependencies:
# Clone repository
git clone https://github.com/YourUsername/ML-Final-Project.git
cd ML-Final-Project
# Install required Python packages
pip install numpy pandas scikit-learn librosa soundfile tqdm matplotlib seaborn joblibRun the Jupyter Notebooks in sequential order:
- Data Cleaning & Feature Extraction:
Run
Data_Cleaning_and_Feature_Extraction.ipynbto process raw audio recordings fromdataset/intoprocessed_data/and generatefeatures.csv. - Classification:
Run
Classification.ipynbto performGroupShuffleSplit, train KNN/RF/SVM models, execute the cross-gender generalization experiment, and export model binaries tomodels/. - Clustering:
Run
Clustering.ipynbto apply feature scaling, PCA dimensionality reduction, K-Means, and Hierarchical Clustering algorithms. - Evaluation:
Run
Evaluation.ipynbto generate confusion matrices, classification reports, silhouette scores, and cluster composition heatmaps.
| Paradigm | Method / Experiment | Key Metric / Result | Acoustic Insight |
|---|---|---|---|
| Supervised | KNN ( |
99.71% Accuracy | Distinct local neighborhoods in feature space |
| Supervised | SVM (RBF Kernel) | 99.71% Accuracy | Optimal non-linear hyperplane separation |
| Supervised | Random Forest (100 trees) | 99.13% Accuracy | Robust ensemble decision boundaries |
| Generalization | Male Train |
34.13% Accuracy | Demonstrates speaker-voice dependence |
| Unsupervised | PCA (22 Components) | 90.0% Variance | High feature redundancy in 44D acoustic space |
| Unsupervised | K-Means ( |
Silhouette: 0.1412 | High purity for Spanish (99.2%) & Italian (62.4%) |
| Unsupervised | Hierarchical ( |
Silhouette: 0.1176 | 100% pure clusters for Italian & Spanish |
- L. R. Rabiner and R. W. Schafer, Digital Processing of Speech Signals. Pearson Education, 1978.
- D. Jurafsky and J. H. Martin, Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. 3rd Edition draft, 2024.
- B. McFee, et al., "librosa: Audio and Music Signal Analysis in Python," Proceedings of the 14th Python in Science Conference, 2015.
- F. Pedregosa, et al., "Scikit-learn: Machine Learning in Python," Journal of Machine Learning Research, vol. 12, pp. 2825-2830, 2011.
- S. Davis and P. Mermelstein, "Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences," IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 28, no. 4, pp. 357-366, 1980.
This project is open-source and released under the MIT License.