Learn AI Series (#92) - Audio Fundamentals for AI
What will I learn
- You will learn how sound works as a signal: waveforms, frequency, amplitude, and phase;
- sampling rate, bit depth, and the Nyquist theorem that governs digital audio;
- the Fourier Transform: decomposing audio into its constituent frequencies;
- spectrograms: transforming 1D audio into 2D images that CNNs can process;
- the Mel scale: mapping frequencies to match how humans actually perceive pitch;
- MFCCs: the classic feature extraction technique that dominated speech processing for decades;
- loading and processing audio in Python with librosa and torchaudio;
- building a complete audio preprocessing pipeline for any downstream AI task.
Requirements
- A working modern computer running macOS, Windows or Ubuntu;
- An installed Python 3(.11+) distribution;
- The ambition to learn AI and machine learning.
Difficulty
- Beginner
Curriculum (of the Learn AI Series):
- Learn AI Series (#1) - What Machine Learning Actually Is
- Learn AI Series (#2) - Setting Up Your AI Workbench - Python and NumPy
- Learn AI Series (#3) - Your Data Is Just Numbers - How Machines See the World
- Learn AI Series (#4) - Your First Prediction - No Math, Just Intuition
- Learn AI Series (#5) - Patterns in Data - What "Learning" Actually Looks Like
- Learn AI Series (#6) - From Intuition to Math - Why We Need Formulas
- Learn AI Series (#7) - The Training Loop - See It Work Step by Step
- Learn AI Series (#8) - The Math You Actually Need (Part 1) - Linear Algebra
- Learn AI Series (#9) - The Math You Actually Need (Part 2) - Calculus and Probability
- Learn AI Series (#10) - Your First ML Model - Linear Regression From Scratch
- Learn AI Series (#11) - Making Linear Regression Real
- Learn AI Series (#12) - Classification - Logistic Regression From Scratch
- Learn AI Series (#13) - Evaluation - How to Know If Your Model Actually Works
- Learn AI Series (#14) - Data Preparation - The 80% Nobody Talks About
- Learn AI Series (#15) - Feature Engineering and Selection
- Learn AI Series (#16) - Scikit-Learn - The Standard Library of ML
- Learn AI Series (#17) - Decision Trees - How Machines Make Decisions
- Learn AI Series (#18) - Random Forests - Wisdom of Crowds
- Learn AI Series (#19) - Gradient Boosting - The Kaggle Champion
- Learn AI Series (#20) - Support Vector Machines - Drawing the Perfect Boundary
- Learn AI Series (#21) - Mini Project - Predicting Crypto Market Regimes
- Learn AI Series (#22) - K-Means Clustering - Finding Groups
- Learn AI Series (#23) - Advanced Clustering - Beyond K-Means
- Learn AI Series (#24) - Dimensionality Reduction - PCA
- Learn AI Series (#25) - Advanced Dimensionality Reduction - t-SNE and UMAP
- Learn AI Series (#26) - Anomaly Detection - Finding What Doesn't Belong
- Learn AI Series (#27) - Recommendation Systems - "Users Like You Also Liked..."
- Learn AI Series (#28) - Time Series Fundamentals - When Order Matters
- Learn AI Series (#29) - Time Series Forecasting - Predicting What Comes Next
- Learn AI Series (#30) - Natural Language Processing - Text as Data
- Learn AI Series (#31) - Word Embeddings - Meaning in Numbers
- Learn AI Series (#32) - Bayesian Methods - Thinking in Probabilities
- Learn AI Series (#33) - Ensemble Methods Deep Dive - Stacking and Blending
- Learn AI Series (#34) - ML Engineering - From Notebook to Production
- Learn AI Series (#35) - Data Ethics and Bias in ML
- Learn AI Series (#36) - Mini Project - Complete ML Pipeline
- Learn AI Series (#37) - The Perceptron - Where It All Started
- Learn AI Series (#38) - Neural Networks From Scratch - Forward Pass
- Learn AI Series (#39) - Neural Networks From Scratch - Backpropagation
- Learn AI Series (#40) - Training Neural Networks - Practical Challenges
- Learn AI Series (#41) - Optimization Algorithms - SGD, Momentum, Adam
- Learn AI Series (#42) - PyTorch Fundamentals - Tensors and Autograd
- Learn AI Series (#43) - PyTorch Data and Training
- Learn AI Series (#44) - PyTorch nn.Module - Building Real Networks
- Learn AI Series (#45) - Convolutional Neural Networks - Theory
- Learn AI Series (#46) - CNNs in Practice - Classic to Modern Architectures
- Learn AI Series (#47) - CNN Applications - Detection, Segmentation, Style Transfer
- Learn AI Series (#48) - Recurrent Neural Networks - Sequences
- Learn AI Series (#49) - LSTM and GRU - Solving the Memory Problem
- Learn AI Series (#50) - Sequence-to-Sequence Models
- Learn AI Series (#51) - Attention Mechanisms
- Learn AI Series (#52) - The Transformer Architecture (Part 1)
- Learn AI Series (#53) - The Transformer Architecture (Part 2)
- Learn AI Series (#54) - Vision Transformers
- Learn AI Series (#55) - Generative Adversarial Networks
- Learn AI Series (#56) - Mini Project - Building a Transformer From Scratch
- Learn AI Series (#57) - Language Modeling - Predicting the Next Word
- Learn AI Series (#58) - GPT Architecture - Decoder-Only Transformers
- Learn AI Series (#59) - BERT and Encoder Models
- Learn AI Series (#60) - Training Large Language Models
- Learn AI Series (#61) - Instruction Tuning and Alignment
- Learn AI Series (#62) - Prompt Engineering - Getting the Most from LLMs
- Learn AI Series (#63) - Embeddings and Vector Search
- Learn AI Series (#64) - Retrieval-Augmented Generation (RAG) - Basics
- Learn AI Series (#65) - RAG - Advanced Techniques
- Learn AI Series (#66) - Working with LLM APIs
- Learn AI Series (#67) - Building AI Agents (Part 1) - Foundations
- Learn AI Series (#68) - Building AI Agents (Part 2) - Advanced Patterns
- Learn AI Series (#69) - Fine-Tuning Language Models
- Learn AI Series (#70) - Running Local Models
- Learn AI Series (#71) - Text Generation Techniques
- Learn AI Series (#72) - Tokenization Deep Dive
- Learn AI Series (#73) - LLM Evaluation
- Learn AI Series (#74) - The Hugging Face Ecosystem
- Learn AI Series (#75) - Multimodal Models - Text Meets Vision
- Learn AI Series (#76) - Mini Project - Your Own AI Assistant
- Learn AI Series (#77) - Image Processing Fundamentals
- Learn AI Series (#78) - Object Detection (Part 1) - Foundations
- Learn AI Series (#79) - Object Detection (Part 2) - Modern Approaches
- Learn AI Series (#80) - Image Segmentation
- Learn AI Series (#81) - Pose Estimation and Tracking
- Learn AI Series (#82) - Optical Character Recognition
- Learn AI Series (#83) - Video Understanding
- Learn AI Series (#84) - Generative Images - Diffusion Models (Part 1)
- Learn AI Series (#85) - Generative Images - Diffusion Models (Part 2)
- Learn AI Series (#86) - Image-to-Image and Editing
- Learn AI Series (#87) - 3D Vision
- Learn AI Series (#88) - Face Analysis
- Learn AI Series (#89) - Medical and Scientific Imaging
- Learn AI Series (#90) - Self-Supervised Learning for Vision
- Learn AI Series (#91) - Mini Project - Building a Visual AI System
- Learn AI Series (#92) - Audio Fundamentals for AI (this post)
Learn AI Series (#92) - Audio Fundamentals for AI
On to today's episode
Here we go! Episode ninety-two, and we're entering completely new territory. For the past fifteen episodes (#77 through #91) we've been deep in the world of computer vision -- pixels, convolutions, object detection, segmentation, generation, 3D, faces, medical imaging, self-supervised learning, and a full end-to-end visual AI system. That was a massive arc and I'm genuinely proud of how thorough it was.
But vision is only ONE sensory modality. Humans experience the world through sight, sound, touch, smell, and taste. And while AI has made enormous progress on both vision and language (which we covered in episodes #57-76), there's a third modality that's equally important and increasingly central to how we interact with machines: sound.
Speech recognition, music generation, audio classification, speaker identification, voice cloning -- all of these start with understanding how sound becomes data, and how that data becomes features a neural network can learn from. The good news? Many of the techniques will feel familiar. The spectrogram -- the single most important audio representation for AI -- turns audio into a 2D image. Once you have an image, you can use CNNs, transformers, and all the vision techniques we've already built. But getting TO that image requires understanding the physics of sound and the mathematics of signal processing ;-)
This episode lays the foundation. We'll go from raw pressure waves all the way through to a complete preprocessing pipeline that produces fixed-size tensors ready for any downstream model. If you followed along with episode #3 (your data is just numbers), the same principle applies here: sound is just numbers. Lots and lots of numbers sampled very quickly.
Sound as a signal
Sound is pressure waves traveling through air. A microphone converts those pressure changes into an electrical signal, and an analog-to-digital converter (ADC) converts that signal into a sequence of numbers. The result: a 1D array of amplitude values sampled at regular intervals.
import numpy as np
class WaveformGenerator:
"""Generate and analyze basic audio
waveforms from first principles."""
def __init__(self, sample_rate=16000):
self.sr = sample_rate
def sine_wave(self, freq, duration=1.0,
amplitude=0.5):
"""Pure sine wave at given frequency."""
t = np.linspace(
0, duration,
int(self.sr * duration),
endpoint=False)
return amplitude * np.sin(
2 * np.pi * freq * t), t
def complex_tone(self, fundamental,
n_harmonics=8,
duration=1.0):
"""Real instrument sound: fundamental
+ harmonics with decreasing amplitude."""
t = np.linspace(
0, duration,
int(self.sr * duration),
endpoint=False)
signal = np.zeros_like(t)
for i in range(n_harmonics):
harmonic_freq = fundamental * (i + 1)
amplitude = 0.5 / (i + 1)
signal += amplitude * np.sin(
2 * np.pi * harmonic_freq * t)
return signal, t
def analyze(self, signal):
"""Basic signal statistics."""
return {
"samples": len(signal),
"duration": len(signal) / self.sr,
"min": float(signal.min()),
"max": float(signal.max()),
"rms": float(np.sqrt(
np.mean(signal ** 2))),
}
gen = WaveformGenerator(sample_rate=16000)
# Pure 440 Hz tone (concert A)
pure, t = gen.sine_wave(440)
stats = gen.analyze(pure)
print("Pure 440 Hz sine wave:")
print(f" Samples: {stats['samples']}")
print(f" Duration: {stats['duration']:.1f}s")
print(f" Range: [{stats['min']:.3f}, "
f"{stats['max']:.3f}]")
print(f" RMS energy: {stats['rms']:.4f}")
# Complex tone (like a real instrument)
rich, t = gen.complex_tone(440, n_harmonics=8)
stats = gen.analyze(rich)
print(f"\nComplex tone (8 harmonics):")
print(f" Samples: {stats['samples']}")
print(f" RMS energy: {stats['rms']:.4f}")
print(f" Range: [{stats['min']:.3f}, "
f"{stats['max']:.3f}]")
A pure sine wave is a single frequency -- mathematically clean but acoustically boring. Real sounds (speech, instruments, environmental noise) are complex mixtures of dozens or hundreds of simultaneous frequencies, all changing over time. The fundamental frequency determines the perceived pitch, while the mix of harmonics (integer multiples of the fundamental) determines the timbre -- it's why a violin and a piano playing the same note sound completely different even though the fundamental frequency is identical.
Three parameters define a digital audio signal:
Sample rate: how many measurements per second. CD quality is 44,100 Hz. Phone calls use 8,000 Hz. Speech AI models typically use 16,000 Hz -- high enough to capture all speech frequencies, low enough to keep data manageable. When you load an audio file for AI processing, the first step is almost always resampling to a standard rate.
Bit depth: the precision of each sample. 16-bit audio has 65,536 possible amplitude values. 24-bit has 16.7 million. For AI purposes, we convert to 32-bit floating point and normalize the waveform to the range [-1.0, 1.0].
Channels: mono (1 channel) or stereo (2 channels). Most AI models work with mono audio -- we average the channels down if the input is stereo. Spatial audio information is rarely useful for tasks like speech recognition or classification.
The Nyquist theorem
Here's a fundamental constraint that governs ALL digital audio. To accurately capture a frequency, you need at least 2 samples per cycle. A 16,000 Hz sample rate can capture frequencies up to 8,000 Hz (the Nyquist frequency). Human speech fundamentals range from ~85 Hz (deep male voice) to ~255 Hz (high female voice), with important harmonics and consonant sounds reaching up to ~8,000 Hz. That's exactly why 16 kHz is the standard for speech processing -- it captures everything that matters.
If a frequency above the Nyquist limit is present in the signal, it "folds" back and appears as a lower frequency. This is aliasing, and it's not just a theoretical problem -- it's a real distortion that corrupts your data. Anti-aliasing filters in the ADC hardware prevent this during recording. When you resample audio in software (say, downsample from 44.1 kHz to 16 kHz), the resampling algorithm applies a digital anti-aliasing filter automatically:
import numpy as np
class NyquistDemonstrator:
"""Demonstrate the Nyquist theorem and
aliasing effects on audio signals."""
def __init__(self):
self.sr_high = 44100 # CD quality
self.sr_low = 8000 # telephone
def check_capturable(self, freq, sr):
"""Can this frequency be captured
at the given sample rate?"""
nyquist = sr / 2
capturable = freq <= nyquist
alias_freq = None
if not capturable:
alias_freq = abs(sr - freq)
if alias_freq > nyquist:
alias_freq = abs(
alias_freq - sr)
return capturable, nyquist, alias_freq
def demonstrate_aliasing(self, true_freq,
sr, duration=0.01):
"""Show that a high frequency appears
as a lower frequency when undersampled."""
n = int(sr * duration)
t = np.arange(n) / sr
sampled = np.sin(2 * np.pi * true_freq * t)
alias = abs(sr - true_freq)
alias_signal = np.sin(
2 * np.pi * alias * t)
mse = np.mean(
(sampled - alias_signal) ** 2)
return alias, mse
def run(self):
print("Nyquist frequency limits:")
print(f" CD quality (44100 Hz): "
f"up to {44100 / 2:,.0f} Hz")
print(f" Speech AI (16000 Hz): "
f"up to {16000 / 2:,.0f} Hz")
print(f" Telephone (8000 Hz): "
f"up to {8000 / 2:,.0f} Hz")
freqs = [440, 4000, 7000, 10000, 20000]
print(f"\n{'Freq':>8} {'@44.1k':>8} "
f"{'@16k':>8} {'@8k':>8}")
print("-" * 36)
for f in freqs:
results = []
for sr in [44100, 16000, 8000]:
ok, nyq, alias = (
self.check_capturable(f, sr))
if ok:
results.append("OK")
else:
results.append(
f"alias={alias}")
print(f"{f:>8} {results[0]:>8} "
f"{results[1]:>8} "
f"{results[2]:>8}")
alias_f, mse = self.demonstrate_aliasing(
7000, 8000)
print(f"\n7000 Hz sampled at 8000 Hz:")
print(f" Appears as: {alias_f} Hz")
print(f" MSE vs alias: {mse:.6f}")
print(f" (Near zero = confirmed alias)")
demo = NyquistDemonstrator()
demo.run()
This table makes it concrete. A 440 Hz concert A is fine at any sample rate. A 7,000 Hz consonant sound (think the "s" in "sit") is captured at 44.1 kHz and 16 kHz but aliases to 1,000 Hz at 8 kHz -- which is why telephone-quality speech sounds muffled. The sibilant consonants are literally lost. And a 20,000 Hz tone (top of human hearing) requires at least 44.1 kHz to capture.
From time domain to frequency domain
A raw waveform shows amplitude over time. Useful for seeing the overall shape, but terrible for understanding content. What frequencies are actually present in this signal? To answer that, we need the Fourier Transform.
The Discrete Fourier Transform (DFT) decomposes a signal into its constituent frequencies. The result: for each frequency bin, a complex number whose magnitude tells you how strong that frequency is and whose phase tells you its timing offset. In practice we use the Fast Fourier Transform (FFT), which computes the DFT in O(N log N) instead of O(N^2):
import numpy as np
class SpectrumAnalyzer:
"""Analyze the frequency content of
audio signals using the FFT."""
def __init__(self, sample_rate=16000):
self.sr = sample_rate
def compute_spectrum(self, signal):
"""Compute magnitude spectrum using
the real-valued FFT."""
fft_result = np.fft.rfft(signal)
freqs = np.fft.rfftfreq(
len(signal), d=1.0 / self.sr)
magnitudes = np.abs(fft_result)
phases = np.angle(fft_result)
return freqs, magnitudes, phases
def find_peaks(self, freqs, magnitudes,
threshold=0.1):
"""Find dominant frequencies above
a threshold fraction of the max."""
cutoff = magnitudes.max() * threshold
peak_mask = magnitudes > cutoff
peak_freqs = freqs[peak_mask]
peak_mags = magnitudes[peak_mask]
order = np.argsort(peak_mags)[::-1]
return peak_freqs[order], (
peak_mags[order])
def run(self):
duration = 0.1
n = int(self.sr * duration)
t = np.arange(n) / self.sr
signal = (0.5 * np.sin(
2 * np.pi * 440 * t)
+ 0.3 * np.sin(
2 * np.pi * 880 * t)
+ 0.1 * np.sin(
2 * np.pi * 1320 * t))
freqs, mags, phases = (
self.compute_spectrum(signal))
print(f"Signal: {n} samples, "
f"{duration}s at {self.sr} Hz")
print(f"FFT bins: {len(freqs)}")
print(f"Frequency resolution: "
f"{freqs[1] - freqs[0]:.1f} Hz")
peaks_f, peaks_m = self.find_peaks(
freqs, mags)
print(f"\nDominant frequencies:")
for f, m in zip(peaks_f[:5],
peaks_m[:5]):
print(f" {f:>8.1f} Hz "
f"magnitude: {m:.1f}")
analyzer = SpectrumAnalyzer()
analyzer.run()
The frequency resolution of the FFT depends on the window length: resolution = sample_rate / N_samples. A 0.1-second window at 16 kHz gives 10 Hz resolution -- sufficient for most purposes. Longer windows give finer frequency resolution but poorer time resolution. This is the fundamental time-frequency tradeoff in signal processing, and it's not a technical limitation you can engineer around. It's baked into the mathematics (related to the Heisenberg uncertainty principle, if you want to go down that rabbit hole).
Spectrograms: the bridge to vision
The FFT gives you the frequency content of the entire signal at once. But audio is dynamic -- the frequencies change over time. A spectrogram solves this by computing the FFT over short overlapping windows (typically 25ms with 10ms hop), producing a 2D representation: frequency x time. This is the Short-Time Fourier Transform (STFT):
import numpy as np
class SpectrogramComputer:
"""Compute spectrograms from raw
waveforms using the STFT."""
def __init__(self, sr=16000, n_fft=1024,
hop_length=256,
win_length=1024):
self.sr = sr
self.n_fft = n_fft
self.hop_length = hop_length
self.win_length = win_length
def stft(self, signal):
"""Short-Time Fourier Transform.
Returns complex-valued STFT matrix."""
window = np.hanning(self.win_length)
n_frames = 1 + (
(len(signal) - self.win_length)
// self.hop_length)
n_freqs = self.n_fft // 2 + 1
result = np.zeros(
(n_freqs, n_frames),
dtype=np.complex128)
for i in range(n_frames):
start = i * self.hop_length
frame = signal[
start:start + self.win_length]
windowed = frame * window
padded = np.zeros(self.n_fft)
padded[:len(windowed)] = windowed
result[:, i] = np.fft.rfft(padded)
return result
def to_magnitude(self, stft_matrix):
"""Convert complex STFT to magnitude
spectrogram."""
return np.abs(stft_matrix)
def to_db(self, magnitude, ref=None,
min_db=-80.0):
"""Convert magnitude to decibels
(log scale). This matches human
perception of loudness."""
if ref is None:
ref = magnitude.max()
db = 20.0 * np.log10(
magnitude / (ref + 1e-10) + 1e-10)
return np.maximum(db, min_db)
def run(self):
duration = 2.0
n = int(self.sr * duration)
t = np.arange(n) / self.sr
freq_start, freq_end = 200, 4000
phase = 2 * np.pi * (
freq_start * t
+ (freq_end - freq_start)
/ (2 * duration) * t ** 2)
signal = 0.5 * np.sin(phase)
stft_matrix = self.stft(signal)
magnitude = self.to_magnitude(
stft_matrix)
spec_db = self.to_db(magnitude)
print(f"Input: {n} samples "
f"({duration:.1f}s at "
f"{self.sr} Hz)")
print(f"STFT shape: "
f"{stft_matrix.shape}")
print(f" Freq bins: "
f"{stft_matrix.shape[0]} "
f"(0 to {self.sr // 2} Hz)")
print(f" Time frames: "
f"{stft_matrix.shape[1]}")
print(f" Window: {self.win_length} "
f"samples "
f"({self.win_length / self.sr * 1000:.1f}ms)")
print(f" Hop: {self.hop_length} "
f"samples "
f"({self.hop_length / self.sr * 1000:.1f}ms)")
print(f"Spectrogram (dB) range: "
f"[{spec_db.min():.1f}, "
f"{spec_db.max():.1f}]")
print(f"\nThis IS a 2D image -- "
f"feed it to a CNN!")
computer = SpectrogramComputer()
computer.run()
The spectrogram is the fundamental representation for audio AI. It converts a 1D signal into a 2D image where the x-axis is time, the y-axis is frequency, and the pixel brightness is energy (in decibels). A CNN trained on spectrograms can detect speech, identify instruments, classify environmental sounds -- all using the same architectures we studied in the vision arc. The conversion to decibels (logarithmic scale) is essential because it matches human perception of loudness: the difference between 10 dB and 20 dB sounds the same as the difference between 50 dB and 60 dB.
The Mel scale: hearing like humans
Human pitch perception is logarithmic, not linear. The perceptual difference between 100 Hz and 200 Hz (an octave) is HUGE. The difference between 5,000 Hz and 5,100 Hz? Barely noticeable. The Mel scale maps physical frequencies to perceptual pitch, giving more resolution to the low frequencies where human hearing is most sensitive:
import numpy as np
class MelFilterBank:
"""Build a Mel-scale filter bank from
scratch. Converts linear-frequency
spectrograms to perceptual Mel scale."""
def __init__(self, sr=16000, n_fft=1024,
n_mels=80, fmin=0, fmax=8000):
self.sr = sr
self.n_fft = n_fft
self.n_mels = n_mels
self.fmin = fmin
self.fmax = fmax
def hz_to_mel(self, hz):
"""Convert Hz to Mel scale."""
return 2595.0 * np.log10(
1.0 + hz / 700.0)
def mel_to_hz(self, mel):
"""Convert Mel scale back to Hz."""
return 700.0 * (
10.0 ** (mel / 2595.0) - 1.0)
def build_filterbank(self):
"""Create triangular filter bank
on the Mel scale."""
n_freqs = self.n_fft // 2 + 1
mel_min = self.hz_to_mel(self.fmin)
mel_max = self.hz_to_mel(self.fmax)
mel_points = np.linspace(
mel_min, mel_max,
self.n_mels + 2)
hz_points = self.mel_to_hz(mel_points)
bin_indices = np.floor(
(self.n_fft + 1)
* hz_points / self.sr).astype(int)
filterbank = np.zeros(
(self.n_mels, n_freqs))
for i in range(self.n_mels):
left = bin_indices[i]
center = bin_indices[i + 1]
right = bin_indices[i + 2]
for j in range(left, center):
if center > left:
filterbank[i, j] = (
(j - left)
/ (center - left))
for j in range(center, right):
if right > center:
filterbank[i, j] = (
(right - j)
/ (right - center))
return filterbank
def apply(self, magnitude_spec):
"""Apply Mel filterbank to magnitude
spectrogram."""
fb = self.build_filterbank()
return fb @ magnitude_spec
def run(self):
fb = self.build_filterbank()
print(f"Mel filterbank shape: {fb.shape}")
print(f" {self.n_mels} Mel bands, "
f"{fb.shape[1]} freq bins")
mel_min = self.hz_to_mel(self.fmin)
mel_max = self.hz_to_mel(self.fmax)
mel_pts = np.linspace(
mel_min, mel_max,
self.n_mels + 2)
hz_pts = self.mel_to_hz(mel_pts)
print(f"\nFilter center frequencies:")
print(f" First 5: "
+ ", ".join(f"{h:.0f}" for h in
hz_pts[1:6]) + " Hz")
print(f" Last 5: "
+ ", ".join(f"{h:.0f}" for h in
hz_pts[-6:-1]) + " Hz")
low_bw = hz_pts[2] - hz_pts[1]
high_bw = hz_pts[-2] - hz_pts[-3]
print(f"\n Low-freq bandwidth: "
f"{low_bw:.0f} Hz")
print(f" High-freq bandwidth: "
f"{high_bw:.0f} Hz")
print(f" Ratio: {high_bw / low_bw:.1f}x "
f"wider at high frequencies")
print(f"\nCompression:")
print(f" STFT: 513 freq bins")
print(f" Mel: {self.n_mels} bands")
print(f" Reduction: "
f"{513 / self.n_mels:.1f}x")
mel = MelFilterBank()
mel.run()
The non-uniform bandwidth is the key insight. Low-frequency filters are narrow (fine pitch resolution where it matters), high-frequency filters are wide (coarse resolution where human hearing is less sensitive). The Mel spectrogram compresses the frequency axis from 513 linear bins to 80 perceptually-spaced bands. This is the input representation used by virtually every modern audio AI model, including Whisper (speech recognition), AudioLDM (audio generation), and HuBERT (speech representation learning).
MFCCs: the classic features
Mel-Frequency Cepstral Coefficients take the Mel spectrogram one step further. They apply a Discrete Cosine Transform (DCT) to the log Mel spectrogram, producing a compact set of coefficients that capture the "shape" of the spectrum -- which is what distinguishes different speech sounds from each other:
import numpy as np
class MFCCComputer:
"""Compute MFCCs from scratch. The classic
audio feature for speech processing."""
def __init__(self, sr=16000, n_mfcc=13,
n_mels=80, n_fft=1024,
hop_length=256):
self.sr = sr
self.n_mfcc = n_mfcc
self.n_mels = n_mels
self.n_fft = n_fft
self.hop_length = hop_length
def dct_matrix(self, n_filters, n_input):
"""Type-II DCT basis matrix."""
basis = np.zeros((n_filters, n_input))
for k in range(n_filters):
for n in range(n_input):
basis[k, n] = np.cos(
np.pi * k
* (2 * n + 1)
/ (2 * n_input))
basis[0] *= np.sqrt(1.0 / n_input)
basis[1:] *= np.sqrt(2.0 / n_input)
return basis
def compute_deltas(self, features, width=2):
"""Compute delta (velocity) features
using finite differences."""
padded = np.pad(
features,
((0, 0), (width, width)),
mode='edge')
deltas = np.zeros_like(features)
denom = 2 * sum(
i ** 2 for i in range(1, width + 1))
for t in range(features.shape[1]):
for n in range(1, width + 1):
deltas[:, t] += n * (
padded[:, t + width + n]
- padded[:, t + width - n])
deltas[:, t] /= denom
return deltas
def run(self):
rng = np.random.RandomState(42)
n_frames = 100
log_mel = rng.randn(
self.n_mels, n_frames) * 10 - 20
dct = self.dct_matrix(
self.n_mfcc, self.n_mels)
mfccs = dct @ log_mel
deltas = self.compute_deltas(mfccs)
delta_deltas = self.compute_deltas(
deltas)
full_features = np.concatenate(
[mfccs, deltas, delta_deltas],
axis=0)
print(f"Log-Mel input: "
f"{log_mel.shape}")
print(f"MFCCs: {mfccs.shape}")
print(f" (13 coefficients x "
f"{n_frames} frames)")
print(f"Deltas: {deltas.shape}")
print(f"Delta-deltas: "
f"{delta_deltas.shape}")
print(f"Full features: "
f"{full_features.shape}")
print(f" (39 = 13 + 13 + 13)")
print(f"\nFeature timeline:")
print(f" Raw waveform: "
f"{n_frames * self.hop_length} "
f"samples")
print(f" Mel spectrogram: "
f"{self.n_mels} x {n_frames}")
print(f" MFCCs: "
f"{self.n_mfcc} x {n_frames}")
print(f" Full: 39 x {n_frames}")
print(f" Compression: "
f"{n_frames * self.hop_length}"
f" -> {39 * n_frames} "
f"({n_frames * self.hop_length / (39 * n_frames):.1f}x)")
mfcc = MFCCComputer()
mfcc.run()
The delta and delta-delta features capture how the spectrum is changing over time -- they encode the velocity and acceleration of spectral features. Think of it this way: the MFCCs tell you what the spectrum looks like right now, the deltas tell you where it's going, and the delta-deltas tell you whether it's speeding up or slowing down. Together, the 39-dimensional feature vector (13 + 13 + 13) captures both static and dynamic spectral properties at each time frame.
MFCCs dominated speech processing for decades (roughly 1980s through 2010s). Hidden Markov Models with MFCC features were the standard approach for automatic speech recognition before deep learning. Modern models like Whisper work directly on Mel spectrograms, letting the neural network learn its own features instead of relying on hand-crafted MFCCs. But MFCCs remain useful for lightweight applications and as a baseline -- and understanding them gives you insight into what modern networks are learning to extract ;-)
Audio data pipeline: from file to tensor
For any audio AI task, you need a preprocessing pipeline that handles the common gotchas: variable sample rates, stereo vs mono, variable-length recordings, and normalization. Here's a complete pipeline that produces fixed-size tensors ready for batching:
import numpy as np
class AudioPipeline:
"""Complete audio preprocessing pipeline.
Handles loading, resampling, normalization,
and feature extraction."""
def __init__(self, target_sr=16000,
n_mels=80, n_fft=1024,
hop_length=256,
max_duration=10.0):
self.target_sr = target_sr
self.max_samples = int(
target_sr * max_duration)
self.n_mels = n_mels
self.n_fft = n_fft
self.hop_length = hop_length
def normalize(self, signal):
"""Peak normalization to [-1, 1]."""
peak = np.abs(signal).max()
if peak > 0:
signal = signal / peak
return signal
def pad_or_truncate(self, signal):
"""Fixed-length output."""
if len(signal) > self.max_samples:
return signal[:self.max_samples]
elif len(signal) < self.max_samples:
padding = self.max_samples - len(
signal)
return np.pad(signal, (0, padding))
return signal
def mel_spectrogram(self, signal):
"""Compute log-Mel spectrogram
from waveform."""
window = np.hanning(self.n_fft)
n_frames = 1 + (
(len(signal) - self.n_fft)
// self.hop_length)
n_freqs = self.n_fft // 2 + 1
spec = np.zeros(
(n_freqs, n_frames))
for i in range(n_frames):
start = i * self.hop_length
frame = signal[
start:start + self.n_fft]
if len(frame) < self.n_fft:
frame = np.pad(
frame,
(0, self.n_fft - len(frame)))
spec[:, i] = np.abs(
np.fft.rfft(frame * window))
mel_fb = self._build_mel_fb(n_freqs)
mel = mel_fb @ (spec ** 2)
log_mel = 10.0 * np.log10(
mel + 1e-10)
return log_mel
def _build_mel_fb(self, n_freqs):
"""Simplified Mel filterbank."""
fmin, fmax = 0, self.target_sr // 2
mel_min = 2595 * np.log10(
1 + fmin / 700)
mel_max = 2595 * np.log10(
1 + fmax / 700)
mel_pts = np.linspace(
mel_min, mel_max,
self.n_mels + 2)
hz_pts = 700 * (
10 ** (mel_pts / 2595) - 1)
bins = np.floor(
(self.n_fft + 1)
* hz_pts / self.target_sr
).astype(int)
fb = np.zeros(
(self.n_mels, n_freqs))
for i in range(self.n_mels):
for j in range(
bins[i], bins[i + 1]):
if bins[i + 1] > bins[i]:
fb[i, j] = (
(j - bins[i])
/ (bins[i + 1]
- bins[i]))
for j in range(
bins[i + 1], bins[i + 2]):
if bins[i + 2] > bins[i + 1]:
fb[i, j] = (
(bins[i + 2] - j)
/ (bins[i + 2]
- bins[i + 1]))
return fb
def process(self, signal, sr=None):
"""Full pipeline: normalize -> pad/
truncate -> mel spectrogram."""
signal = self.normalize(signal)
signal = self.pad_or_truncate(signal)
mel = self.mel_spectrogram(signal)
return mel
def run_demo(self):
"""Process a synthetic signal to
demonstrate the full pipeline."""
rng = np.random.RandomState(42)
duration = 3.0
n = int(self.target_sr * duration)
t = np.arange(n) / self.target_sr
signal = np.zeros(n)
for freq in [150, 300, 450, 1200,
2400, 3600]:
amp = rng.uniform(0.1, 0.5)
signal += amp * np.sin(
2 * np.pi * freq * t)
signal += rng.randn(n) * 0.02
mel = self.process(signal)
print(f"Input: {n} samples "
f"({duration:.1f}s)")
print(f"After pad/truncate: "
f"{self.max_samples} samples "
f"({self.max_samples / self.target_sr:.1f}s)")
print(f"Mel spectrogram: {mel.shape}")
print(f" {mel.shape[0]} Mel bands "
f"x {mel.shape[1]} time frames")
print(f" dB range: [{mel.min():.1f}, "
f"{mel.max():.1f}]")
n_expected = 1 + (
self.max_samples - self.n_fft
) // self.hop_length
print(f" Expected frames: "
f"{n_expected}")
print(f"\nReady for CNN/transformer "
f"input!")
pipeline = AudioPipeline(max_duration=10.0)
pipeline.run_demo()
This pipeline handles the common gotchas: variable-length audio (padded or truncated to 10 seconds), amplitude normalization (peak normalization to [-1, 1]), and the conversion from waveform to log-Mel spectrogram. The output is a fixed-size 2D array that can be batched and fed to any model.
Augmentation for audio
Just like images benefit from data augmentation (episode #14), audio AI benefits enormously from audio-specific augmentations. The key augmentations that improve robustness without distorting the content:
import numpy as np
class AudioAugmenter:
"""Audio augmentation strategies for
training robust models."""
def __init__(self, sr=16000, seed=42):
self.sr = sr
self.rng = np.random.RandomState(seed)
def add_noise(self, signal, snr_db=20):
"""Add Gaussian noise at a given
signal-to-noise ratio."""
signal_power = np.mean(signal ** 2)
noise_power = signal_power / (
10 ** (snr_db / 10))
noise = self.rng.randn(
len(signal)) * np.sqrt(noise_power)
return signal + noise
def time_shift(self, signal,
max_shift_sec=0.1):
"""Shift the signal in time
(circular shift)."""
max_shift = int(
max_shift_sec * self.sr)
shift = self.rng.randint(
-max_shift, max_shift)
return np.roll(signal, shift)
def speed_change(self, signal,
factor_range=(0.9, 1.1)):
"""Change playback speed (affects
both tempo and pitch)."""
factor = self.rng.uniform(
*factor_range)
indices = np.arange(
0, len(signal), factor)
indices = indices[
indices < len(signal)].astype(int)
return signal[indices]
def spec_augment(self, mel_spec,
n_freq_masks=2,
n_time_masks=2,
freq_mask_width=10,
time_mask_width=20):
"""SpecAugment: mask random frequency
and time bands in the spectrogram.
The single most important augmentation
for speech recognition."""
augmented = mel_spec.copy()
n_mels, n_frames = augmented.shape
for _ in range(n_freq_masks):
f = self.rng.randint(
0, freq_mask_width)
f0 = self.rng.randint(
0, max(n_mels - f, 1))
augmented[f0:f0 + f, :] = 0
for _ in range(n_time_masks):
t = self.rng.randint(
0, time_mask_width)
t0 = self.rng.randint(
0, max(n_frames - t, 1))
augmented[:, t0:t0 + t] = 0
return augmented
def run(self):
n = self.sr * 2
t = np.arange(n) / self.sr
signal = 0.5 * np.sin(
2 * np.pi * 440 * t)
noisy = self.add_noise(signal, snr_db=20)
shifted = self.time_shift(signal)
fast = self.speed_change(
signal, (1.1, 1.1))
print(f"Original: {len(signal)} "
f"samples, RMS={np.sqrt(np.mean(signal ** 2)):.4f}")
print(f"+ noise: {len(noisy)} "
f"samples, RMS={np.sqrt(np.mean(noisy ** 2)):.4f}")
print(f"Time-shifted:{len(shifted)} "
f"samples")
print(f"Speed 1.1x: {len(fast)} "
f"samples ({len(fast) / self.sr:.2f}s)")
mel = self.rng.randn(80, 200) * 10
augmented = self.spec_augment(mel)
masked_freq = np.sum(
augmented.sum(axis=1) == 0)
masked_time = np.sum(
augmented.sum(axis=0) == 0)
print(f"\nSpecAugment on (80, 200):")
print(f" Freq bands zeroed: "
f"{masked_freq}")
print(f" Time frames zeroed: "
f"{masked_time}")
aug = AudioAugmenter()
aug.run()
SpecAugment (Park et al., 2019) deserves special attention. It's the audio equivalent of random erasing in vision -- mask out random frequency bands and time segments in the spectrogram during training. It was a key ingredient in Google's state-of-the-art speech recognition systems and is now used in virtually every serious ASR pipeline. The beauty is that it operates on the spectrogram (after feature extraction), so it's trivially easy to implement as a training-time transform.
The representation landscape
Let me put all the representations we've covered in context, because understanding when to use which representation is half the battle in audio AI:
import numpy as np
class RepresentationComparator:
"""Compare audio representations
in terms of size, information, and
typical use cases."""
def __init__(self, sr=16000,
duration=10.0):
self.sr = sr
self.duration = duration
self.n_samples = int(sr * duration)
def compute_sizes(self):
"""Calculate data sizes for each
representation of a 10-second clip."""
n = self.n_samples
n_fft = 1024
hop = 256
n_frames = 1 + (n - n_fft) // hop
n_freqs = n_fft // 2 + 1
reps = {
"Raw waveform": n,
"STFT magnitude": n_freqs * n_frames,
"Mel spectrogram (80)": 80 * n_frames,
"Mel spectrogram (128)": 128 * n_frames,
"MFCC (13)": 13 * n_frames,
"MFCC + deltas (39)": 39 * n_frames,
}
print(f"Audio: {self.duration:.0f}s "
f"at {self.sr} Hz")
print(f"Time frames (hop={hop}): "
f"{n_frames}")
print(f"\n{'Representation':<24} "
f"{'Shape':>16} {'Values':>10} "
f"{'Ratio':>8}")
print("-" * 62)
baseline = n
for name, size in reps.items():
if "waveform" in name:
shape = f"({n},)"
elif "STFT" in name:
shape = (f"({n_freqs}, "
f"{n_frames})")
elif "Mel" in name:
bands = 80 if "80" in name else 128
shape = (f"({bands}, "
f"{n_frames})")
elif "39" in name:
shape = f"(39, {n_frames})"
else:
shape = f"(13, {n_frames})"
ratio = baseline / size
print(f"{name:<24} {shape:>16} "
f"{size:>10,} "
f"{ratio:>7.1f}x")
print(f"\nTypical use cases:")
print(f" Waveform: WaveNet, "
f"SampleRNN (raw audio gen)")
print(f" STFT: phase-sensitive "
f"tasks (source separation)")
print(f" Mel 80: Whisper, most "
f"modern speech/audio models")
print(f" Mel 128: higher detail "
f"for music analysis")
print(f" MFCC 13: lightweight "
f"classifiers, keyword spotting")
print(f" MFCC 39: traditional ASR "
f"(pre-deep-learning)")
comp = RepresentationComparator()
comp.compute_sizes()
The compression ratio tells a clear story. Going from raw waveform (160,000 values) to Mel spectrogram (80 x 625 = 50,000 values) is a 3.2x compression, but the Mel spectrogram retains the perceptually important information. MFCCs compress further to 13 x 625 = 8,125 values (19.7x compression), but at the cost of discarding fine spectral detail that modern deep networks can exploit. The trend in the field is clear: use the least-compressed representation your compute budget allows, and let the network learn what to extract.
Samengevat
- Sound is pressure waves converted to a 1D array of amplitude samples; sample rate (Hz) determines the maximum capturable frequency via the Nyquist theorem (max frequency = sample_rate / 2), and aliasing corrupts any frequency above this limit;
- the Fourier Transform decomposes a signal into its frequency components; the Short-Time Fourier Transform (STFT) does this over sliding windows (typically 25ms with 10ms hop) to produce a time-frequency representation called a spectrogram;
- spectrograms are 2D images (frequency x time) of audio energy in decibels -- the bridge between audio processing and the vision models we've already built; they are THE standard input for modern audio AI;
- the Mel scale maps physical frequencies to human perceptual pitch (logarithmic), giving more resolution to low frequencies where hearing is most sensitive; Mel spectrograms compress 513 linear frequency bins to 80-128 perceptually-spaced bands;
- MFCCs apply DCT to log-Mel spectrograms for compact spectral shape descriptors (39 dimensions with deltas); they dominated pre-deep-learning speech processing but modern models learn directly from Mel spectrograms;
- SpecAugment (masking random frequency bands and time segments) is the most important audio augmentation technique, directly analogous to random erasing in vision;
- the standard pipeline: load -> resample to 16 kHz -> mono -> normalize -> pad/truncate -> Mel spectrogram -> log scale -> feed to model.
We've laid the foundation for the entire audio AI arc in this episode. Sound is now data, and we know how to transform that data into representations neural networks can learn from. The audio domain has its own unique challenges (temporal structure, the importance of phase for some tasks, the massive compression from waveform to features), but the core machine learning techniques from the previous arcs -- CNNs, transformers, attention mechanisms -- all apply directly once you have the right representation.
Exercises
Exercise 1: Build a frequency content analyzer. Create a class FrequencyAnalyzer that: (a) generates three synthetic signals of 2 seconds each at 16,000 Hz sample rate: (1) a pure 440 Hz sine wave, (2) a chord with 440 Hz + 554 Hz + 659 Hz (A major chord), (3) white noise from numpy's random generator with seed=42, (b) implements compute_spectrum(signal) that computes the FFT magnitude spectrum, (c) implements spectral_centroid(freqs, magnitudes) that computes the weighted average frequency: sum(f * m) / sum(m) where f is frequency and m is magnitude, (d) implements spectral_bandwidth(freqs, magnitudes, centroid) that computes the weighted standard deviation of frequencies around the centroid, (e) implements spectral_rolloff(freqs, magnitudes, percentile=85) that finds the frequency below which the given percentile of total spectral energy is contained, (f) prints a comparison table showing centroid, bandwidth, and rolloff for all three signals. Verify that white noise has the highest centroid (energy spread across all frequencies), the chord has intermediate values, and the pure tone has the narrowest bandwidth (all energy at one frequency).
Exercise 2: Build a Mel filterbank visualizer and verifier. Create a class MelFilterbankAnalyzer that: (a) implements the Hz-to-Mel and Mel-to-Hz conversion formulas (Mel = 2595 * log10(1 + Hz/700)), (b) builds a triangular Mel filterbank for sr=16000, n_fft=1024, n_mels=40, fmin=0, fmax=8000, (c) computes for each filter: center frequency (Hz), bandwidth (Hz), and peak height, (d) verifies two properties: (1) that every frequency bin from fmin to fmax is covered by at least one filter with nonzero weight (no gaps), (2) that the sum of all filters at any frequency bin is approximately 1.0 (energy conservation), (e) computes the average bandwidth for the lowest 10 filters vs the highest 10 filters, (f) prints a summary table showing filter index, center frequency, bandwidth, and peak height for filters 0, 9, 19, 29, 39. Verify that low-frequency filters have narrow bandwidth (fine pitch resolution) and high-frequency filters have wide bandwidth (coarse resolution), and that the bandwidth ratio between highest and lowest filters is approximately 10:1 or more.
Exercise 3: Build a time-frequency resolution tradeoff analyzer. Create a class ResolutionAnalyzer that: (a) generates a 1-second synthetic signal at 16,000 Hz that contains: a 500 Hz tone in the first 0.5 seconds ONLY, and a 600 Hz tone in the last 0.5 seconds ONLY (with zero-crossings at the transition point to avoid clicks), (b) computes the STFT (magnitude spectrogram in dB) with four different window sizes: 256, 512, 1024, 2048 samples (all with hop_length = window_size // 4), (c) for each window size, measures: (1) frequency resolution (sr / n_fft Hz), (2) time resolution (hop_length / sr seconds), (3) whether the two tones (500 Hz and 600 Hz) are resolvable as separate peaks in the frequency axis (check if the frequency resolution is smaller than the 100 Hz gap between them), (4) whether the time boundary (0.5 seconds) is clearly localized (check how many time frames span the transition region), (d) prints a comparison table showing window size, freq resolution, time resolution, frequency-resolvable (yes/no), and number of frames in the transition zone. Verify that small windows give good time resolution but poor frequency resolution (the two tones blur together), while large windows give good frequency resolution but poor time resolution (the transition is smeared across many frames).