Skip to content

About

End-to-end spoken language identification (German/Italian/Korean/Spanish) from raw audio: MFCC-based feature engineering, leakage-free group splits, KNN/RF/SVM benchmarks and unsupervised clustering analysis.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Spoken Language Identification (LID) - Machine Learning Final Project

University of Tehran Department Python License


πŸ“Œ Executive Summary & Overview

Spoken Language Identification (LID) is the computational task of automatically classifying the spoken language from raw acoustic utterances. This repository presents an end-to-end Machine Learning system for Spoken Language Identification across four major languages: German, Italian, Korean, and Spanish.

The project encompasses the complete data science lifecycle:

  1. Dataset Creation & Curation: Collecting audiobook narrations across languages and speaker genders.
  2. Audio Processing & Cleaning: Standardizing sampling rates (16 kHz Mono), trimming silent pauses via Voice Activity Detection (VAD at 20 dB threshold), and uniform 4-second segmentation yielding 10,862 clean speech samples.
  3. Spectral & Temporal Feature Engineering: Extracting 44 statistical feature representations per sample, combining 20 Mel-Frequency Cepstral Coefficients (MFCCs), Spectral Centroid, and Zero Crossing Rate (ZCR) aggregated over time by mean and variance.
  4. Data Leakage Prevention: Assigning unique group_id identifiers per source audio recording to guarantee strict train-test separation via GroupShuffleSplit.
  5. Supervised Classification: Benchmarking K-Nearest Neighbors (KNN), Random Forest (RF), and Support Vector Machines (SVM) on leakage-free splits, reaching up to 99.71% accuracy.
  6. Cross-Gender Domain Adaptation Experiment: Training exclusively on male speakers and testing on female speakers (34.13% accuracy) to evaluate speaker-dependence vs. language-invariant acoustic cues.
  7. Unsupervised Clustering & Analysis: Dimensionality reduction via Principal Component Analysis (PCA) alongside K-Means and Hierarchical Agglomerative Clustering (Ward linkage) to uncover intrinsic linguistic groupings without label supervision.

πŸŽ“ Academic & Team Information

  • Institution: University of Tehran, College of Engineering, Department of Electrical and Computer Engineering
  • Course: Machine Learning (Final Project)
  • Team Members:
    • Amir Aref (Student ID: 810102506)
    • Parsa Saeednia (Student ID: 810102460)
    • Amirali Dehghani (Student ID: 810102443)

🧠 Problem Formulation & Theoretical Background

Why Spoken Language Identification?

Language Identification forms the foundational front-end for multilingual speech processing pipelines, automated transcription systems, call-center routing, and real-time machine translation. The challenge lies in isolating language-specific phonemic and prosodic patterns from speaker-specific vocal traits (e.g., fundamental frequency, pitch, vocal tract geometry) and environmental acoustic conditions.

Language Selection Rationale

The dataset evaluates four phonetically distinct target languages:

  • German: Germanic branch, heavy consonant clusters, high dynamic range.
  • Italian: Romance branch, vowel-ended phonotactics, rhythmic syllabic cadence.
  • Spanish: Romance branch, fast syllable-timed rhythm, distinct fricatives.
  • Korean: East Asian language, agglutinative grammar, distinct stop consonant voicings.

πŸ“ Dataset Architecture & Preprocessing

1. Data Collection & Curation

Raw audio streams were compiled from 720 MP3 audiobook recordings across male and female narrators.

dataset/
β”œβ”€β”€ German/
β”‚   β”œβ”€β”€ Female/
β”‚   └── Male/
β”œβ”€β”€ Italian/
β”‚   β”œβ”€β”€ Female/
β”‚   └── Male/
β”œβ”€β”€ Korean/
β”‚   β”œβ”€β”€ Female/
β”‚   └── Male/
└── Spanish/
    β”œβ”€β”€ Female/
    └── Male/

2. Audio Processing Pipeline

Audio processing was performed using librosa and soundfile:

  1. Sampling Rate Standardization: Resampled all recordings to 16,000 Hz (Nyquist cutoff frequency of 8 kHz, preserving human speech intelligibility while optimizing computation).
  2. Mono Downmixing: Mixed stereo channels to single-channel mono to eliminate spatial panning variations.
  3. Voice Activity Detection (VAD): Trimmed silent pauses and low-energy background signal below a 20 dB threshold (librosa.effects.trim(y, top_db=20)).
  4. Fixed-Length Slicing: Sliced continuous speech streams into non-overlapping 4-second segments ($16,000 \times 4 = 64,000$ samples per audio clip).

3. Final Dataset Statistics

Language Clean Audio Segments Dataset Share (%)
Italian 2,956 27.2%
German 2,677 24.6%
Korean 2,620 24.1%
Spanish 2,609 24.1%
Total 10,862 100.0%
  • Gender Distribution: 5,652 Male samples | 5,210 Female samples
  • Source Audiobook Files: 720 unique source audiobooks

4. Data Leakage Prevention (group_id)

Because thousands of 4-second segments are derived from 720 source files, a naive random split would place adjacent segments of the same narrator in both training and test sets, leading to artificial accuracy spikes due to speaker memorization ("Clever Hans" effect).

To prevent leakage:

  • Each segment receives a group_id derived from its source file name.
  • Data splitting uses GroupShuffleSplit, ensuring that all segments from any source audiobook reside strictly in either the training set or the test set, never both.

πŸ”¬ Feature Extraction Pipeline

Speech waveforms ($64,000$ values) are converted into fixed 44-dimensional feature vectors. Frame-level spectral features are aggregated across time using Mean ($\mu$) and Variance ($\sigma^2$):

$$\text{Vector}_{44} = \left[ \mu(\text{MFCC}_{1..20}), \sigma^2(\text{MFCC}_{1..20}), \mu(\text{Centroid}), \sigma^2(\text{Centroid}), \mu(\text{ZCR}), \sigma^2(\text{ZCR}) \right]$$

Feature Components

  1. Mel-Frequency Cepstral Coefficients (MFCCs):

    • Extracted 20 coefficients. MFCCs map the linear power spectrum onto the non-linear Mel scale, mimicking the human auditory perception of timbre and phonemic formants.
    • Dimensions: 20 Means + 20 Variances = 40 Features.
  2. Spectral Centroid:

    • Measures the center of mass of the frequency spectrum, representing perceived audio brightness.
    • Dimensions: 1 Mean + 1 Variance = 2 Features.
  3. Zero Crossing Rate (ZCR):

    • Counts how rapidly the time-domain signal changes sign, helping distinguish percussive unvoiced consonants and fricatives from voiced vowels.
    • Dimensions: 1 Mean + 1 Variance = 2 Features.

The extracted feature set is stored in features.csv (10,862 rows $\times$ 47 columns including metadata columns filename, group_id, and language).


πŸ€– Supervised Classification Experiments

Data Splitting & Preprocessing

  • Split Ratio: 80% Training (8,785 samples) / 20% Testing (2,077 samples).
  • Splitter: GroupShuffleSplit (n_splits=1, test_size=0.2, random_state=42). Zero group overlap verified.
  • Scaler: StandardScaler fitted on training set and applied to test set ($\mu=0, \sigma=1$).

Model Performance Metrics

Classifier Architecture & Parameters Accuracy Precision Recall F1-Score
K-Nearest Neighbors (KNN) $k=5$, Minkowski metric 99.71% 0.99 0.99 0.99
Support Vector Machine (SVM) RBF Kernel, $C=1.0$, Probability=True 99.71% 0.99 0.99 0.99
Random Forest (RF) 100 Trees, Gini impurity 99.13% 0.99 0.99 0.99

Detailed Classification Analysis

  • KNN ($k=5$): Near-perfect score indicates pure local feature neighborhoods and distinct cluster separation between languages.
  • SVM (RBF): Achieved 99.71% accuracy, confirming that high-dimensional non-linear transformation creates clear hyperplanes between linguistic spectral features.
  • Random Forest: Achieved 99.13% accuracy, demonstrating ensemble decision tree robustness against variance in frame-level spectral energy.

πŸ‘₯ Cross-Gender Generalization Experiment

Motivation & Setup

To test whether classifiers learn language-invariant features (phonemes, rhythm, syntax) versus speaker-dependent traits (pitch, vocal tract length), we designed a domain adaptation experiment:

  • Training Set: Male voice samples ONLY (5,652 samples).
  • Test Set: Female voice samples ONLY (5,210 samples).

Results & Findings

  • Cross-Gender Test Accuracy: 34.13%
  • Random Baseline: 25.00% (4 classes)

Critical Acoustic Takeaway

While 34.13% accuracy exceeds random guessing (indicating the extraction of some underlying language prosody), the significant drop from 99.71% demonstrates strong Speaker Dependence. Because each language in single-audiobook datasets originates from a limited speaker pool, models partially memorize individual narrator acoustics. True language identification across unseen speakers requires pitch-invariant normalization or multi-speaker corpora.


πŸ” Unsupervised Clustering & Analysis

Unsupervised learning evaluates whether extracted features naturally cluster by language without using target labels during training.

1. Principal Component Analysis (PCA)

  • Variance Retention: 22 Principal Components account for 90% of total feature variance.
  • 2D PCA Projection: Used for qualitative scatter plot visualization. Showed clear grouping with partial acoustic overlaps.

2. K-Means Clustering ($k=4$)

  • Cluster Count Selection: Evaluated $k \in [2, 10]$ using the Elbow Method (Inertia reduction curve).
  • Silhouette Score: 0.1412 on standardized 44D feature space.
  • Cluster Purity Matrix (Normalize by Index):
Cluster ID German Italian Korean Spanish Dominant Language
Cluster 0 59.3% 10.6% 0.0% 30.1% German
Cluster 1 7.2% 62.4% 11.5% 18.9% Italian
Cluster 2 0.0% 0.7% 0.1% 99.2% Spanish
Cluster 3 41.1% 1.6% 54.3% 3.0% Korean

3. Hierarchical Agglomerative Clustering ($k=4$, Ward Linkage)

  • Dendrogram: Plotted on a 1,000-sample subset using Ward linkage to minimize intra-cluster variance. Showed distinct branches at four clusters.
  • Silhouette Score: 0.1176 on full dataset.
  • Cluster Purity Matrix (Normalize by Index):
Cluster ID German Italian Korean Spanish Dominant Language
Cluster 0 38.6% 27.3% 8.3% 25.8% Mixed (German/Italian)
Cluster 1 24.1% 0.2% 75.4% 0.3% Korean
Cluster 2 0.0% 0.0% 0.0% 100.0% Spanish (100% Pure)
Cluster 3 0.0% 100.0% 0.0% 0.0% Italian (100% Pure)

πŸ“‚ Repository File Structure

ML-Final-Project/
β”œβ”€β”€ Data_Cleaning_and_Feature_Extraction.ipynb # Step 1: Preprocessing, VAD, slicing & feature extraction
β”œβ”€β”€ Classification.ipynb                       # Step 2: GroupShuffleSplit, KNN/RF/SVM training & cross-gender test
β”œβ”€β”€ Clustering.ipynb                       # Step 3: Standardization, PCA, K-Means & Hierarchical clustering
β”œβ”€β”€ Evaluation.ipynb                       # Step 4: Final evaluation metrics, confusion matrices & plots
β”œβ”€β”€ features.csv                           # Extracted 44-dimensional dataset (10,862 samples x 47 columns)
β”œβ”€β”€ README.md                              # Main project documentation
β”œβ”€β”€ models/                                # Trained model binaries & evaluation artifacts
β”‚   β”œβ”€β”€ knn_model.pkl                      # Serialized KNN model
β”‚   β”œβ”€β”€ rf_model.pkl                       # Serialized Random Forest model
β”‚   β”œβ”€β”€ svm_model.pkl                      # Serialized SVM model
β”‚   β”œβ”€β”€ X_test_scaled.pkl                  # Scaled test dataset
β”‚   β”œβ”€β”€ y_test.pkl                         # Ground truth test labels
β”‚   └── label_encoder.pkl                  # Target label encoder
└── Report/                                # Project report documentation & source
    β”œβ”€β”€ ML_Final_Project_Phase2_Report.pdf # Final compiled report PDF
    β”œβ”€β”€ main.tex                           # Main LaTeX report document
    β”œβ”€β”€ setting.tex                        # LaTeX package & style configurations
    β”œβ”€β”€ font/                              # Font assets
    └── img/                               # Visual assets & logos

πŸ› οΈ Environment Setup & Execution Guide

1. Installation & Requirements

Clone the repository and install the dependencies:

# Clone repository
git clone https://github.com/YourUsername/ML-Final-Project.git
cd ML-Final-Project

# Install required Python packages
pip install numpy pandas scikit-learn librosa soundfile tqdm matplotlib seaborn joblib

2. Running the Pipeline Step-by-Step

Run the Jupyter Notebooks in sequential order:

  1. Data Cleaning & Feature Extraction: Run Data_Cleaning_and_Feature_Extraction.ipynb to process raw audio recordings from dataset/ into processed_data/ and generate features.csv.
  2. Classification: Run Classification.ipynb to perform GroupShuffleSplit, train KNN/RF/SVM models, execute the cross-gender generalization experiment, and export model binaries to models/.
  3. Clustering: Run Clustering.ipynb to apply feature scaling, PCA dimensionality reduction, K-Means, and Hierarchical Clustering algorithms.
  4. Evaluation: Run Evaluation.ipynb to generate confusion matrices, classification reports, silhouette scores, and cluster composition heatmaps.

πŸ“Š Summary of Results

Paradigm Method / Experiment Key Metric / Result Acoustic Insight
Supervised KNN ($k=5$) 99.71% Accuracy Distinct local neighborhoods in feature space
Supervised SVM (RBF Kernel) 99.71% Accuracy Optimal non-linear hyperplane separation
Supervised Random Forest (100 trees) 99.13% Accuracy Robust ensemble decision boundaries
Generalization Male Train $\rightarrow$ Female Test 34.13% Accuracy Demonstrates speaker-voice dependence
Unsupervised PCA (22 Components) 90.0% Variance High feature redundancy in 44D acoustic space
Unsupervised K-Means ($k=4$) Silhouette: 0.1412 High purity for Spanish (99.2%) & Italian (62.4%)
Unsupervised Hierarchical ($k=4$, Ward) Silhouette: 0.1176 100% pure clusters for Italian & Spanish

πŸ“š References

  1. L. R. Rabiner and R. W. Schafer, Digital Processing of Speech Signals. Pearson Education, 1978.
  2. D. Jurafsky and J. H. Martin, Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. 3rd Edition draft, 2024.
  3. B. McFee, et al., "librosa: Audio and Music Signal Analysis in Python," Proceedings of the 14th Python in Science Conference, 2015.
  4. F. Pedregosa, et al., "Scikit-learn: Machine Learning in Python," Journal of Machine Learning Research, vol. 12, pp. 2825-2830, 2011.
  5. S. Davis and P. Mermelstein, "Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences," IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 28, no. 4, pp. 357-366, 1980.

πŸ“„ License

This project is open-source and released under the MIT License.

About

End-to-end spoken language identification (German/Italian/Korean/Spanish) from raw audio: MFCC-based feature engineering, leakage-free group splits, KNN/RF/SVM benchmarks and unsupervised clustering analysis.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages