The complete audio AI preprocessing pipeline — from raw samples to spectrograms, MFCCs, and beat detection.
Waveform · STFT · Mel Spectrogram · MFCC · Beat Tracking
Whisper, Spotify, Shazam, MusicGen — they all convert audio into the same mathematical representation before doing anything intelligent.
This project implements that representation from scratch.
git clone https://github.com/hey-shiv/Audio-Explorer.git && cd Audio-Explorer
pip install -r requirements.txt
streamlit run app.pyOpens at localhost:8501. Drop a WAV, MP3, or FLAC.
Every major audio AI system preprocesses sound through the same chain of transformations. Audio Explorer implements each stage as an interactive visualization you can zoom, pan, and inspect sample-by-sample.
| Waveform | Decode audio into a 1D float array. 22,050 samples per second of audio. Compute signal health — RMS loudness, peak amplitude, DC offset, clipping detection, dynamic range in dB. |
| Spectrogram | Slide a 2048-sample Hann window with 512-sample hops. FFT each window. Stack magnitude spectra into a time-frequency matrix. 1025 frequency bins, one column per 23ms. Convert to dB. |
| Mel Spectrogram | Human hearing is logarithmic — 200→400 Hz is an octave, 8000→8200 Hz is nothing. Apply 128 triangular Mel filters to warp the spectrogram to match perception. 8x compression, zero perceptual loss. This is what Whisper sees. |
| MFCCs | DCT on the log-mel spectrum. 13 coefficients per frame. Coefficient 0 = energy. Coefficients 1–12 = spectral envelope (vocal tract shape). Delta MFCCs capture rate of change. The standard speech recognition feature since 1980. |
| Beat Tracking | Onset strength envelope → autocorrelation → dynamic programming. Outputs tempo in BPM + individual beat timestamps. Beat interval analysis reveals metronomic vs. human timing. |
Four ideas, two centuries:
| Year | Idea | What it does |
|---|---|---|
| 1807 | Fourier Transform | Any signal = sum of sine waves. Find them. |
| 1937 | Mel Scale | m = 2595 · log₁₀(1 + f/700) — warp frequency to match human ears. |
| 1965 | FFT (Cooley-Tukey) | Compute DFT in O(N log N) instead of O(N²). 200,000x speedup. |
| 1980 | MFCCs | DCT on log-mel spectrum → 13 numbers that describe the spectral shape. |
Every major audio AI system takes a mel spectrogram as input. The model architectures differ. The preprocessing is the same.
| System | Organization | Input |
|---|---|---|
| Whisper | OpenAI | 80-band mel spectrogram |
| AudioMAE | Meta | 128-band mel spectrogram |
| CLAP | LAION | 64-band mel spectrogram |
| BEATs | Microsoft | mel spectrogram |
| MusicGen | Meta | encoded spectrogram |
app.py entry point
src/audio_explorer/
core/loader.py decode audio, extract metadata
visualization/
waveform.py amplitude vs time + signal stats
spectrogram.py STFT → frequency vs time
mel_spectrogram.py mel filterbank warping
mfcc.py cepstral coefficients + deltas
rhythm.py tempo estimation + beat detection
Each module: compute_*() returns arrays, plot_*() returns a Plotly figure. Computation and rendering fully separated.
Stack: Python 3.10 · Librosa · NumPy · SciPy · Plotly · Streamlit
MIT License · Built by Shivashant
