Skip to content

About

Spoken language identification and clustering across 4 languages (German, Italian, Korean, Spanish) using audio feature extraction (MFCCs, Chroma), SVM, Random Forest, GMM, and t-SNE.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Spoken Language Classification & Clustering

Project Overview

This repository contains the code and report for an end-to-end machine learning project focused on spoken language identification. The goal of the project is to process raw audio recordings in four languages (German, Italian, Korean, and Spanish), extract meaningful acoustic features, and apply both supervised and unsupervised machine learning algorithms to classify the languages and discover hidden acoustic structures.

Methodology & Pipeline

The project is divided into three main phases:

Dataset

The raw audio dataset containing the recordings for German, Italian, Korean, and Spanish used in this project can be accessed here: Spoken Language Audio Dataset

1. Data Cleaning & Feature Extraction

To prepare the raw audio for machine learning models, the following pipeline was applied:

  • Audio Preprocessing: All audio clips were standardized to a 22,050 Hz sampling rate. Silence was trimmed from the start and end of clips (threshold of 20 dB), and amplitude normalization was applied.
  • Segmentation (Data Augmentation): Long audio files were divided into fixed 5-second segments (ignoring segments shorter than 4 seconds), yielding over 8,600 samples.
  • Feature Extraction: For each segment, 68 features were extracted utilizing the librosa library. The features include the Mean and Variance for:
    • MFCCs (20 coefficients)
    • Spectral Centroid
    • Zero Crossing Rate (ZCR)
    • Chroma STFT (12 coefficients)

2. Supervised Learning (Classification)

The extracted features were used to train models to predict the spoken language:

  • Data Splitting Strategy: To prevent critical data leakage (where the model learns the speaker's voice rather than the language), a GroupShuffleSplit was utilized. This ensured that different segments from the same original audio file were not split across both the training and testing sets.
  • Models Evaluated: Support Vector Machine (SVM) with an RBF kernel, Random Forest, and K-Nearest Neighbors (KNN).
  • Feature Standardization: Features were standardized using StandardScaler to ensure distance-based models (like SVM and KNN) performed optimally.

3. Unsupervised Learning (Clustering)

Clustering algorithms were applied to discover natural groupings in the data without relying on the language labels:

  • Models Evaluated: K-Means, DBSCAN, and Gaussian Mixture Models (GMM).
  • Dimensionality Reduction: PCA and t-SNE were used to visualize the high-dimensional feature space in 2D.
  • Evaluation Metrics: The Elbow Method and Silhouette Score were used to find the optimal number of clusters ($k$). Purity Score, Normalized Mutual Information (NMI), and Adjusted Rand Index (ARI) were used to evaluate clustering quality against the ground truth.

Repository Structure

  • Data_Cleaning_and_Feature_Extraction.ipynb Handles the audio preprocessing, silence removal, segmentation, and extraction of the 68 time and frequency domain features using librosa. Outputs the cleaned dataset as a CSV file.

  • Classification.ipynb Contains the supervised learning pipeline. It handles feature scaling, group-based train/test splitting, model training (SVM, Random Forest, KNN), cross-validation, and performance evaluation (Accuracy, Confusion Matrices, Feature Importance).

  • Clustring.ipynb Contains the unsupervised learning pipeline. Implements K-Means, DBSCAN, and GMM algorithms. Includes grid searches for hyperparameter tuning and generates PCA and t-SNE visualizations to map cluster formations.

Key Results

  • Classification: The supervised models achieved exceptional performance (up to 100% accuracy with SVM). A security audit and 5-Fold Group Cross-Validation confirmed that this accuracy was robust and not a result of data leakage, highlighting the high discriminative power of the MFCC and Chroma features.
  • Clustering: GMM outperformed K-Means and DBSCAN, achieving a Purity Score of 73.08% with $k=8$ clusters. The optimal cluster count being higher than the number of languages (4) suggests the model successfully identified acoustic sub-groups within the languages (e.g., distinguishing between male and female speakers or recording environments).

You can see the complete report and analysis in the Documentation PDF.

Dependencies

To run the scripts in this repository, you will need the following Python libraries:

  • numpy
  • pandas
  • librosa
  • scikit-learn
  • matplotlib
  • seaborn
  • tqdm

About

Spoken language identification and clustering across 4 languages (German, Italian, Korean, Spanish) using audio feature extraction (MFCCs, Chroma), SVM, Random Forest, GMM, and t-SNE.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages