Learn AI Series (#92) - Audio Fundamentals for AI

Words
3716
Reading
17 min
Listen
Play
4M

Learn AI Series (#92) - Audio Fundamentals for AI

variant-c-02-gold.png

What will I learn

  • You will learn how sound works as a signal: waveforms, frequency, amplitude, and phase;
  • sampling rate, bit depth, and the Nyquist theorem that governs digital audio;
  • the Fourier Transform: decomposing audio into its constituent frequencies;
  • spectrograms: transforming 1D audio into 2D images that CNNs can process;
  • the Mel scale: mapping frequencies to match how humans actually perceive pitch;
  • MFCCs: the classic feature extraction technique that dominated speech processing for decades;
  • loading and processing audio in Python with librosa and torchaudio;
  • building a complete audio preprocessing pipeline for any downstream AI task.

Requirements

  • A working modern computer running macOS, Windows or Ubuntu;
  • An installed Python 3(.11+) distribution;
  • The ambition to learn AI and machine learning.

Difficulty

  • Beginner

Curriculum (of the Learn AI Series):

Learn AI Series (#92) - Audio Fundamentals for AI

On to today's episode

Here we go! Episode ninety-two, and we're entering completely new territory. For the past fifteen episodes (#77 through #91) we've been deep in the world of computer vision -- pixels, convolutions, object detection, segmentation, generation, 3D, faces, medical imaging, self-supervised learning, and a full end-to-end visual AI system. That was a massive arc and I'm genuinely proud of how thorough it was.

But vision is only ONE sensory modality. Humans experience the world through sight, sound, touch, smell, and taste. And while AI has made enormous progress on both vision and language (which we covered in episodes #57-76), there's a third modality that's equally important and increasingly central to how we interact with machines: sound.

Speech recognition, music generation, audio classification, speaker identification, voice cloning -- all of these start with understanding how sound becomes data, and how that data becomes features a neural network can learn from. The good news? Many of the techniques will feel familiar. The spectrogram -- the single most important audio representation for AI -- turns audio into a 2D image. Once you have an image, you can use CNNs, transformers, and all the vision techniques we've already built. But getting TO that image requires understanding the physics of sound and the mathematics of signal processing ;-)

This episode lays the foundation. We'll go from raw pressure waves all the way through to a complete preprocessing pipeline that produces fixed-size tensors ready for any downstream model. If you followed along with episode #3 (your data is just numbers), the same principle applies here: sound is just numbers. Lots and lots of numbers sampled very quickly.

Sound as a signal

Sound is pressure waves traveling through air. A microphone converts those pressure changes into an electrical signal, and an analog-to-digital converter (ADC) converts that signal into a sequence of numbers. The result: a 1D array of amplitude values sampled at regular intervals.

import numpy as np


class WaveformGenerator:
    """Generate and analyze basic audio
    waveforms from first principles."""

    def __init__(self, sample_rate=16000):
        self.sr = sample_rate

    def sine_wave(self, freq, duration=1.0,
                  amplitude=0.5):
        """Pure sine wave at given frequency."""
        t = np.linspace(
            0, duration,
            int(self.sr * duration),
            endpoint=False)
        return amplitude * np.sin(
            2 * np.pi * freq * t), t

    def complex_tone(self, fundamental,
                     n_harmonics=8,
                     duration=1.0):
        """Real instrument sound: fundamental
        + harmonics with decreasing amplitude."""
        t = np.linspace(
            0, duration,
            int(self.sr * duration),
            endpoint=False)
        signal = np.zeros_like(t)
        for i in range(n_harmonics):
            harmonic_freq = fundamental * (i + 1)
            amplitude = 0.5 / (i + 1)
            signal += amplitude * np.sin(
                2 * np.pi * harmonic_freq * t)
        return signal, t

    def analyze(self, signal):
        """Basic signal statistics."""
        return {
            "samples": len(signal),
            "duration": len(signal) / self.sr,
            "min": float(signal.min()),
            "max": float(signal.max()),
            "rms": float(np.sqrt(
                np.mean(signal ** 2))),
        }


gen = WaveformGenerator(sample_rate=16000)

# Pure 440 Hz tone (concert A)
pure, t = gen.sine_wave(440)
stats = gen.analyze(pure)
print("Pure 440 Hz sine wave:")
print(f"  Samples: {stats['samples']}")
print(f"  Duration: {stats['duration']:.1f}s")
print(f"  Range: [{stats['min']:.3f}, "
      f"{stats['max']:.3f}]")
print(f"  RMS energy: {stats['rms']:.4f}")

# Complex tone (like a real instrument)
rich, t = gen.complex_tone(440, n_harmonics=8)
stats = gen.analyze(rich)
print(f"\nComplex tone (8 harmonics):")
print(f"  Samples: {stats['samples']}")
print(f"  RMS energy: {stats['rms']:.4f}")
print(f"  Range: [{stats['min']:.3f}, "
      f"{stats['max']:.3f}]")

A pure sine wave is a single frequency -- mathematically clean but acoustically boring. Real sounds (speech, instruments, environmental noise) are complex mixtures of dozens or hundreds of simultaneous frequencies, all changing over time. The fundamental frequency determines the perceived pitch, while the mix of harmonics (integer multiples of the fundamental) determines the timbre -- it's why a violin and a piano playing the same note sound completely different even though the fundamental frequency is identical.

Three parameters define a digital audio signal:

Sample rate: how many measurements per second. CD quality is 44,100 Hz. Phone calls use 8,000 Hz. Speech AI models typically use 16,000 Hz -- high enough to capture all speech frequencies, low enough to keep data manageable. When you load an audio file for AI processing, the first step is almost always resampling to a standard rate.

Bit depth: the precision of each sample. 16-bit audio has 65,536 possible amplitude values. 24-bit has 16.7 million. For AI purposes, we convert to 32-bit floating point and normalize the waveform to the range [-1.0, 1.0].

Channels: mono (1 channel) or stereo (2 channels). Most AI models work with mono audio -- we average the channels down if the input is stereo. Spatial audio information is rarely useful for tasks like speech recognition or classification.

The Nyquist theorem

Here's a fundamental constraint that governs ALL digital audio. To accurately capture a frequency, you need at least 2 samples per cycle. A 16,000 Hz sample rate can capture frequencies up to 8,000 Hz (the Nyquist frequency). Human speech fundamentals range from ~85 Hz (deep male voice) to ~255 Hz (high female voice), with important harmonics and consonant sounds reaching up to ~8,000 Hz. That's exactly why 16 kHz is the standard for speech processing -- it captures everything that matters.

If a frequency above the Nyquist limit is present in the signal, it "folds" back and appears as a lower frequency. This is aliasing, and it's not just a theoretical problem -- it's a real distortion that corrupts your data. Anti-aliasing filters in the ADC hardware prevent this during recording. When you resample audio in software (say, downsample from 44.1 kHz to 16 kHz), the resampling algorithm applies a digital anti-aliasing filter automatically:

import numpy as np


class NyquistDemonstrator:
    """Demonstrate the Nyquist theorem and
    aliasing effects on audio signals."""

    def __init__(self):
        self.sr_high = 44100  # CD quality
        self.sr_low = 8000    # telephone

    def check_capturable(self, freq, sr):
        """Can this frequency be captured
        at the given sample rate?"""
        nyquist = sr / 2
        capturable = freq <= nyquist
        alias_freq = None
        if not capturable:
            alias_freq = abs(sr - freq)
            if alias_freq > nyquist:
                alias_freq = abs(
                    alias_freq - sr)
        return capturable, nyquist, alias_freq

    def demonstrate_aliasing(self, true_freq,
                             sr, duration=0.01):
        """Show that a high frequency appears
        as a lower frequency when undersampled."""
        n = int(sr * duration)
        t = np.arange(n) / sr
        sampled = np.sin(2 * np.pi * true_freq * t)
        alias = abs(sr - true_freq)
        alias_signal = np.sin(
            2 * np.pi * alias * t)
        mse = np.mean(
            (sampled - alias_signal) ** 2)
        return alias, mse

    def run(self):
        print("Nyquist frequency limits:")
        print(f"  CD quality (44100 Hz): "
              f"up to {44100 / 2:,.0f} Hz")
        print(f"  Speech AI (16000 Hz):  "
              f"up to {16000 / 2:,.0f} Hz")
        print(f"  Telephone (8000 Hz):   "
              f"up to {8000 / 2:,.0f} Hz")

        freqs = [440, 4000, 7000, 10000, 20000]
        print(f"\n{'Freq':>8} {'@44.1k':>8} "
              f"{'@16k':>8} {'@8k':>8}")
        print("-" * 36)
        for f in freqs:
            results = []
            for sr in [44100, 16000, 8000]:
                ok, nyq, alias = (
                    self.check_capturable(f, sr))
                if ok:
                    results.append("OK")
                else:
                    results.append(
                        f"alias={alias}")
            print(f"{f:>8} {results[0]:>8} "
                  f"{results[1]:>8} "
                  f"{results[2]:>8}")

        alias_f, mse = self.demonstrate_aliasing(
            7000, 8000)
        print(f"\n7000 Hz sampled at 8000 Hz:")
        print(f"  Appears as: {alias_f} Hz")
        print(f"  MSE vs alias: {mse:.6f}")
        print(f"  (Near zero = confirmed alias)")


demo = NyquistDemonstrator()
demo.run()

This table makes it concrete. A 440 Hz concert A is fine at any sample rate. A 7,000 Hz consonant sound (think the "s" in "sit") is captured at 44.1 kHz and 16 kHz but aliases to 1,000 Hz at 8 kHz -- which is why telephone-quality speech sounds muffled. The sibilant consonants are literally lost. And a 20,000 Hz tone (top of human hearing) requires at least 44.1 kHz to capture.

From time domain to frequency domain

A raw waveform shows amplitude over time. Useful for seeing the overall shape, but terrible for understanding content. What frequencies are actually present in this signal? To answer that, we need the Fourier Transform.

The Discrete Fourier Transform (DFT) decomposes a signal into its constituent frequencies. The result: for each frequency bin, a complex number whose magnitude tells you how strong that frequency is and whose phase tells you its timing offset. In practice we use the Fast Fourier Transform (FFT), which computes the DFT in O(N log N) instead of O(N^2):

import numpy as np


class SpectrumAnalyzer:
    """Analyze the frequency content of
    audio signals using the FFT."""

    def __init__(self, sample_rate=16000):
        self.sr = sample_rate

    def compute_spectrum(self, signal):
        """Compute magnitude spectrum using
        the real-valued FFT."""
        fft_result = np.fft.rfft(signal)
        freqs = np.fft.rfftfreq(
            len(signal), d=1.0 / self.sr)
        magnitudes = np.abs(fft_result)
        phases = np.angle(fft_result)
        return freqs, magnitudes, phases

    def find_peaks(self, freqs, magnitudes,
                   threshold=0.1):
        """Find dominant frequencies above
        a threshold fraction of the max."""
        cutoff = magnitudes.max() * threshold
        peak_mask = magnitudes > cutoff
        peak_freqs = freqs[peak_mask]
        peak_mags = magnitudes[peak_mask]
        order = np.argsort(peak_mags)[::-1]
        return peak_freqs[order], (
            peak_mags[order])

    def run(self):
        duration = 0.1
        n = int(self.sr * duration)
        t = np.arange(n) / self.sr
        signal = (0.5 * np.sin(
            2 * np.pi * 440 * t)
            + 0.3 * np.sin(
                2 * np.pi * 880 * t)
            + 0.1 * np.sin(
                2 * np.pi * 1320 * t))

        freqs, mags, phases = (
            self.compute_spectrum(signal))

        print(f"Signal: {n} samples, "
              f"{duration}s at {self.sr} Hz")
        print(f"FFT bins: {len(freqs)}")
        print(f"Frequency resolution: "
              f"{freqs[1] - freqs[0]:.1f} Hz")

        peaks_f, peaks_m = self.find_peaks(
            freqs, mags)
        print(f"\nDominant frequencies:")
        for f, m in zip(peaks_f[:5],
                        peaks_m[:5]):
            print(f"  {f:>8.1f} Hz  "
                  f"magnitude: {m:.1f}")


analyzer = SpectrumAnalyzer()
analyzer.run()

The frequency resolution of the FFT depends on the window length: resolution = sample_rate / N_samples. A 0.1-second window at 16 kHz gives 10 Hz resolution -- sufficient for most purposes. Longer windows give finer frequency resolution but poorer time resolution. This is the fundamental time-frequency tradeoff in signal processing, and it's not a technical limitation you can engineer around. It's baked into the mathematics (related to the Heisenberg uncertainty principle, if you want to go down that rabbit hole).

Spectrograms: the bridge to vision

The FFT gives you the frequency content of the entire signal at once. But audio is dynamic -- the frequencies change over time. A spectrogram solves this by computing the FFT over short overlapping windows (typically 25ms with 10ms hop), producing a 2D representation: frequency x time. This is the Short-Time Fourier Transform (STFT):

import numpy as np


class SpectrogramComputer:
    """Compute spectrograms from raw
    waveforms using the STFT."""

    def __init__(self, sr=16000, n_fft=1024,
                 hop_length=256,
                 win_length=1024):
        self.sr = sr
        self.n_fft = n_fft
        self.hop_length = hop_length
        self.win_length = win_length

    def stft(self, signal):
        """Short-Time Fourier Transform.
        Returns complex-valued STFT matrix."""
        window = np.hanning(self.win_length)
        n_frames = 1 + (
            (len(signal) - self.win_length)
            // self.hop_length)
        n_freqs = self.n_fft // 2 + 1
        result = np.zeros(
            (n_freqs, n_frames),
            dtype=np.complex128)
        for i in range(n_frames):
            start = i * self.hop_length
            frame = signal[
                start:start + self.win_length]
            windowed = frame * window
            padded = np.zeros(self.n_fft)
            padded[:len(windowed)] = windowed
            result[:, i] = np.fft.rfft(padded)
        return result

    def to_magnitude(self, stft_matrix):
        """Convert complex STFT to magnitude
        spectrogram."""
        return np.abs(stft_matrix)

    def to_db(self, magnitude, ref=None,
              min_db=-80.0):
        """Convert magnitude to decibels
        (log scale). This matches human
        perception of loudness."""
        if ref is None:
            ref = magnitude.max()
        db = 20.0 * np.log10(
            magnitude / (ref + 1e-10) + 1e-10)
        return np.maximum(db, min_db)

    def run(self):
        duration = 2.0
        n = int(self.sr * duration)
        t = np.arange(n) / self.sr
        freq_start, freq_end = 200, 4000
        phase = 2 * np.pi * (
            freq_start * t
            + (freq_end - freq_start)
            / (2 * duration) * t ** 2)
        signal = 0.5 * np.sin(phase)

        stft_matrix = self.stft(signal)
        magnitude = self.to_magnitude(
            stft_matrix)
        spec_db = self.to_db(magnitude)

        print(f"Input: {n} samples "
              f"({duration:.1f}s at "
              f"{self.sr} Hz)")
        print(f"STFT shape: "
              f"{stft_matrix.shape}")
        print(f"  Freq bins: "
              f"{stft_matrix.shape[0]} "
              f"(0 to {self.sr // 2} Hz)")
        print(f"  Time frames: "
              f"{stft_matrix.shape[1]}")
        print(f"  Window: {self.win_length} "
              f"samples "
              f"({self.win_length / self.sr * 1000:.1f}ms)")
        print(f"  Hop: {self.hop_length} "
              f"samples "
              f"({self.hop_length / self.sr * 1000:.1f}ms)")
        print(f"Spectrogram (dB) range: "
              f"[{spec_db.min():.1f}, "
              f"{spec_db.max():.1f}]")
        print(f"\nThis IS a 2D image -- "
              f"feed it to a CNN!")


computer = SpectrogramComputer()
computer.run()

The spectrogram is the fundamental representation for audio AI. It converts a 1D signal into a 2D image where the x-axis is time, the y-axis is frequency, and the pixel brightness is energy (in decibels). A CNN trained on spectrograms can detect speech, identify instruments, classify environmental sounds -- all using the same architectures we studied in the vision arc. The conversion to decibels (logarithmic scale) is essential because it matches human perception of loudness: the difference between 10 dB and 20 dB sounds the same as the difference between 50 dB and 60 dB.

The Mel scale: hearing like humans

Human pitch perception is logarithmic, not linear. The perceptual difference between 100 Hz and 200 Hz (an octave) is HUGE. The difference between 5,000 Hz and 5,100 Hz? Barely noticeable. The Mel scale maps physical frequencies to perceptual pitch, giving more resolution to the low frequencies where human hearing is most sensitive:

import numpy as np


class MelFilterBank:
    """Build a Mel-scale filter bank from
    scratch. Converts linear-frequency
    spectrograms to perceptual Mel scale."""

    def __init__(self, sr=16000, n_fft=1024,
                 n_mels=80, fmin=0, fmax=8000):
        self.sr = sr
        self.n_fft = n_fft
        self.n_mels = n_mels
        self.fmin = fmin
        self.fmax = fmax

    def hz_to_mel(self, hz):
        """Convert Hz to Mel scale."""
        return 2595.0 * np.log10(
            1.0 + hz / 700.0)

    def mel_to_hz(self, mel):
        """Convert Mel scale back to Hz."""
        return 700.0 * (
            10.0 ** (mel / 2595.0) - 1.0)

    def build_filterbank(self):
        """Create triangular filter bank
        on the Mel scale."""
        n_freqs = self.n_fft // 2 + 1

        mel_min = self.hz_to_mel(self.fmin)
        mel_max = self.hz_to_mel(self.fmax)
        mel_points = np.linspace(
            mel_min, mel_max,
            self.n_mels + 2)
        hz_points = self.mel_to_hz(mel_points)

        bin_indices = np.floor(
            (self.n_fft + 1)
            * hz_points / self.sr).astype(int)

        filterbank = np.zeros(
            (self.n_mels, n_freqs))
        for i in range(self.n_mels):
            left = bin_indices[i]
            center = bin_indices[i + 1]
            right = bin_indices[i + 2]
            for j in range(left, center):
                if center > left:
                    filterbank[i, j] = (
                        (j - left)
                        / (center - left))
            for j in range(center, right):
                if right > center:
                    filterbank[i, j] = (
                        (right - j)
                        / (right - center))
        return filterbank

    def apply(self, magnitude_spec):
        """Apply Mel filterbank to magnitude
        spectrogram."""
        fb = self.build_filterbank()
        return fb @ magnitude_spec

    def run(self):
        fb = self.build_filterbank()
        print(f"Mel filterbank shape: {fb.shape}")
        print(f"  {self.n_mels} Mel bands, "
              f"{fb.shape[1]} freq bins")

        mel_min = self.hz_to_mel(self.fmin)
        mel_max = self.hz_to_mel(self.fmax)
        mel_pts = np.linspace(
            mel_min, mel_max,
            self.n_mels + 2)
        hz_pts = self.mel_to_hz(mel_pts)

        print(f"\nFilter center frequencies:")
        print(f"  First 5: "
              + ", ".join(f"{h:.0f}" for h in
                          hz_pts[1:6]) + " Hz")
        print(f"  Last 5:  "
              + ", ".join(f"{h:.0f}" for h in
                          hz_pts[-6:-1]) + " Hz")

        low_bw = hz_pts[2] - hz_pts[1]
        high_bw = hz_pts[-2] - hz_pts[-3]
        print(f"\n  Low-freq bandwidth:  "
              f"{low_bw:.0f} Hz")
        print(f"  High-freq bandwidth: "
              f"{high_bw:.0f} Hz")
        print(f"  Ratio: {high_bw / low_bw:.1f}x "
              f"wider at high frequencies")

        print(f"\nCompression:")
        print(f"  STFT: 513 freq bins")
        print(f"  Mel:   {self.n_mels} bands")
        print(f"  Reduction: "
              f"{513 / self.n_mels:.1f}x")


mel = MelFilterBank()
mel.run()

The non-uniform bandwidth is the key insight. Low-frequency filters are narrow (fine pitch resolution where it matters), high-frequency filters are wide (coarse resolution where human hearing is less sensitive). The Mel spectrogram compresses the frequency axis from 513 linear bins to 80 perceptually-spaced bands. This is the input representation used by virtually every modern audio AI model, including Whisper (speech recognition), AudioLDM (audio generation), and HuBERT (speech representation learning).

MFCCs: the classic features

Mel-Frequency Cepstral Coefficients take the Mel spectrogram one step further. They apply a Discrete Cosine Transform (DCT) to the log Mel spectrogram, producing a compact set of coefficients that capture the "shape" of the spectrum -- which is what distinguishes different speech sounds from each other:

import numpy as np


class MFCCComputer:
    """Compute MFCCs from scratch. The classic
    audio feature for speech processing."""

    def __init__(self, sr=16000, n_mfcc=13,
                 n_mels=80, n_fft=1024,
                 hop_length=256):
        self.sr = sr
        self.n_mfcc = n_mfcc
        self.n_mels = n_mels
        self.n_fft = n_fft
        self.hop_length = hop_length

    def dct_matrix(self, n_filters, n_input):
        """Type-II DCT basis matrix."""
        basis = np.zeros((n_filters, n_input))
        for k in range(n_filters):
            for n in range(n_input):
                basis[k, n] = np.cos(
                    np.pi * k
                    * (2 * n + 1)
                    / (2 * n_input))
        basis[0] *= np.sqrt(1.0 / n_input)
        basis[1:] *= np.sqrt(2.0 / n_input)
        return basis

    def compute_deltas(self, features, width=2):
        """Compute delta (velocity) features
        using finite differences."""
        padded = np.pad(
            features,
            ((0, 0), (width, width)),
            mode='edge')
        deltas = np.zeros_like(features)
        denom = 2 * sum(
            i ** 2 for i in range(1, width + 1))
        for t in range(features.shape[1]):
            for n in range(1, width + 1):
                deltas[:, t] += n * (
                    padded[:, t + width + n]
                    - padded[:, t + width - n])
            deltas[:, t] /= denom
        return deltas

    def run(self):
        rng = np.random.RandomState(42)
        n_frames = 100
        log_mel = rng.randn(
            self.n_mels, n_frames) * 10 - 20

        dct = self.dct_matrix(
            self.n_mfcc, self.n_mels)
        mfccs = dct @ log_mel

        deltas = self.compute_deltas(mfccs)
        delta_deltas = self.compute_deltas(
            deltas)
        full_features = np.concatenate(
            [mfccs, deltas, delta_deltas],
            axis=0)

        print(f"Log-Mel input: "
              f"{log_mel.shape}")
        print(f"MFCCs: {mfccs.shape}")
        print(f"  (13 coefficients x "
              f"{n_frames} frames)")
        print(f"Deltas: {deltas.shape}")
        print(f"Delta-deltas: "
              f"{delta_deltas.shape}")
        print(f"Full features: "
              f"{full_features.shape}")
        print(f"  (39 = 13 + 13 + 13)")

        print(f"\nFeature timeline:")
        print(f"  Raw waveform: "
              f"{n_frames * self.hop_length} "
              f"samples")
        print(f"  Mel spectrogram: "
              f"{self.n_mels} x {n_frames}")
        print(f"  MFCCs: "
              f"{self.n_mfcc} x {n_frames}")
        print(f"  Full: 39 x {n_frames}")
        print(f"  Compression: "
              f"{n_frames * self.hop_length}"
              f" -> {39 * n_frames} "
              f"({n_frames * self.hop_length / (39 * n_frames):.1f}x)")


mfcc = MFCCComputer()
mfcc.run()

The delta and delta-delta features capture how the spectrum is changing over time -- they encode the velocity and acceleration of spectral features. Think of it this way: the MFCCs tell you what the spectrum looks like right now, the deltas tell you where it's going, and the delta-deltas tell you whether it's speeding up or slowing down. Together, the 39-dimensional feature vector (13 + 13 + 13) captures both static and dynamic spectral properties at each time frame.

MFCCs dominated speech processing for decades (roughly 1980s through 2010s). Hidden Markov Models with MFCC features were the standard approach for automatic speech recognition before deep learning. Modern models like Whisper work directly on Mel spectrograms, letting the neural network learn its own features instead of relying on hand-crafted MFCCs. But MFCCs remain useful for lightweight applications and as a baseline -- and understanding them gives you insight into what modern networks are learning to extract ;-)

Audio data pipeline: from file to tensor

For any audio AI task, you need a preprocessing pipeline that handles the common gotchas: variable sample rates, stereo vs mono, variable-length recordings, and normalization. Here's a complete pipeline that produces fixed-size tensors ready for batching:

import numpy as np


class AudioPipeline:
    """Complete audio preprocessing pipeline.
    Handles loading, resampling, normalization,
    and feature extraction."""

    def __init__(self, target_sr=16000,
                 n_mels=80, n_fft=1024,
                 hop_length=256,
                 max_duration=10.0):
        self.target_sr = target_sr
        self.max_samples = int(
            target_sr * max_duration)
        self.n_mels = n_mels
        self.n_fft = n_fft
        self.hop_length = hop_length

    def normalize(self, signal):
        """Peak normalization to [-1, 1]."""
        peak = np.abs(signal).max()
        if peak > 0:
            signal = signal / peak
        return signal

    def pad_or_truncate(self, signal):
        """Fixed-length output."""
        if len(signal) > self.max_samples:
            return signal[:self.max_samples]
        elif len(signal) < self.max_samples:
            padding = self.max_samples - len(
                signal)
            return np.pad(signal, (0, padding))
        return signal

    def mel_spectrogram(self, signal):
        """Compute log-Mel spectrogram
        from waveform."""
        window = np.hanning(self.n_fft)
        n_frames = 1 + (
            (len(signal) - self.n_fft)
            // self.hop_length)
        n_freqs = self.n_fft // 2 + 1
        spec = np.zeros(
            (n_freqs, n_frames))
        for i in range(n_frames):
            start = i * self.hop_length
            frame = signal[
                start:start + self.n_fft]
            if len(frame) < self.n_fft:
                frame = np.pad(
                    frame,
                    (0, self.n_fft - len(frame)))
            spec[:, i] = np.abs(
                np.fft.rfft(frame * window))

        mel_fb = self._build_mel_fb(n_freqs)
        mel = mel_fb @ (spec ** 2)
        log_mel = 10.0 * np.log10(
            mel + 1e-10)
        return log_mel

    def _build_mel_fb(self, n_freqs):
        """Simplified Mel filterbank."""
        fmin, fmax = 0, self.target_sr // 2
        mel_min = 2595 * np.log10(
            1 + fmin / 700)
        mel_max = 2595 * np.log10(
            1 + fmax / 700)
        mel_pts = np.linspace(
            mel_min, mel_max,
            self.n_mels + 2)
        hz_pts = 700 * (
            10 ** (mel_pts / 2595) - 1)
        bins = np.floor(
            (self.n_fft + 1)
            * hz_pts / self.target_sr
        ).astype(int)

        fb = np.zeros(
            (self.n_mels, n_freqs))
        for i in range(self.n_mels):
            for j in range(
                    bins[i], bins[i + 1]):
                if bins[i + 1] > bins[i]:
                    fb[i, j] = (
                        (j - bins[i])
                        / (bins[i + 1]
                           - bins[i]))
            for j in range(
                    bins[i + 1], bins[i + 2]):
                if bins[i + 2] > bins[i + 1]:
                    fb[i, j] = (
                        (bins[i + 2] - j)
                        / (bins[i + 2]
                           - bins[i + 1]))
        return fb

    def process(self, signal, sr=None):
        """Full pipeline: normalize -> pad/
        truncate -> mel spectrogram."""
        signal = self.normalize(signal)
        signal = self.pad_or_truncate(signal)
        mel = self.mel_spectrogram(signal)
        return mel

    def run_demo(self):
        """Process a synthetic signal to
        demonstrate the full pipeline."""
        rng = np.random.RandomState(42)
        duration = 3.0
        n = int(self.target_sr * duration)
        t = np.arange(n) / self.target_sr

        signal = np.zeros(n)
        for freq in [150, 300, 450, 1200,
                     2400, 3600]:
            amp = rng.uniform(0.1, 0.5)
            signal += amp * np.sin(
                2 * np.pi * freq * t)
        signal += rng.randn(n) * 0.02

        mel = self.process(signal)
        print(f"Input: {n} samples "
              f"({duration:.1f}s)")
        print(f"After pad/truncate: "
              f"{self.max_samples} samples "
              f"({self.max_samples / self.target_sr:.1f}s)")
        print(f"Mel spectrogram: {mel.shape}")
        print(f"  {mel.shape[0]} Mel bands "
              f"x {mel.shape[1]} time frames")
        print(f"  dB range: [{mel.min():.1f}, "
              f"{mel.max():.1f}]")
        n_expected = 1 + (
            self.max_samples - self.n_fft
        ) // self.hop_length
        print(f"  Expected frames: "
              f"{n_expected}")
        print(f"\nReady for CNN/transformer "
              f"input!")


pipeline = AudioPipeline(max_duration=10.0)
pipeline.run_demo()

This pipeline handles the common gotchas: variable-length audio (padded or truncated to 10 seconds), amplitude normalization (peak normalization to [-1, 1]), and the conversion from waveform to log-Mel spectrogram. The output is a fixed-size 2D array that can be batched and fed to any model.

Augmentation for audio

Just like images benefit from data augmentation (episode #14), audio AI benefits enormously from audio-specific augmentations. The key augmentations that improve robustness without distorting the content:

import numpy as np


class AudioAugmenter:
    """Audio augmentation strategies for
    training robust models."""

    def __init__(self, sr=16000, seed=42):
        self.sr = sr
        self.rng = np.random.RandomState(seed)

    def add_noise(self, signal, snr_db=20):
        """Add Gaussian noise at a given
        signal-to-noise ratio."""
        signal_power = np.mean(signal ** 2)
        noise_power = signal_power / (
            10 ** (snr_db / 10))
        noise = self.rng.randn(
            len(signal)) * np.sqrt(noise_power)
        return signal + noise

    def time_shift(self, signal,
                   max_shift_sec=0.1):
        """Shift the signal in time
        (circular shift)."""
        max_shift = int(
            max_shift_sec * self.sr)
        shift = self.rng.randint(
            -max_shift, max_shift)
        return np.roll(signal, shift)

    def speed_change(self, signal,
                     factor_range=(0.9, 1.1)):
        """Change playback speed (affects
        both tempo and pitch)."""
        factor = self.rng.uniform(
            *factor_range)
        indices = np.arange(
            0, len(signal), factor)
        indices = indices[
            indices < len(signal)].astype(int)
        return signal[indices]

    def spec_augment(self, mel_spec,
                     n_freq_masks=2,
                     n_time_masks=2,
                     freq_mask_width=10,
                     time_mask_width=20):
        """SpecAugment: mask random frequency
        and time bands in the spectrogram.
        The single most important augmentation
        for speech recognition."""
        augmented = mel_spec.copy()
        n_mels, n_frames = augmented.shape

        for _ in range(n_freq_masks):
            f = self.rng.randint(
                0, freq_mask_width)
            f0 = self.rng.randint(
                0, max(n_mels - f, 1))
            augmented[f0:f0 + f, :] = 0

        for _ in range(n_time_masks):
            t = self.rng.randint(
                0, time_mask_width)
            t0 = self.rng.randint(
                0, max(n_frames - t, 1))
            augmented[:, t0:t0 + t] = 0

        return augmented

    def run(self):
        n = self.sr * 2
        t = np.arange(n) / self.sr
        signal = 0.5 * np.sin(
            2 * np.pi * 440 * t)

        noisy = self.add_noise(signal, snr_db=20)
        shifted = self.time_shift(signal)
        fast = self.speed_change(
            signal, (1.1, 1.1))

        print(f"Original:    {len(signal)} "
              f"samples, RMS={np.sqrt(np.mean(signal ** 2)):.4f}")
        print(f"+ noise:     {len(noisy)} "
              f"samples, RMS={np.sqrt(np.mean(noisy ** 2)):.4f}")
        print(f"Time-shifted:{len(shifted)} "
              f"samples")
        print(f"Speed 1.1x:  {len(fast)} "
              f"samples ({len(fast) / self.sr:.2f}s)")

        mel = self.rng.randn(80, 200) * 10
        augmented = self.spec_augment(mel)
        masked_freq = np.sum(
            augmented.sum(axis=1) == 0)
        masked_time = np.sum(
            augmented.sum(axis=0) == 0)
        print(f"\nSpecAugment on (80, 200):")
        print(f"  Freq bands zeroed: "
              f"{masked_freq}")
        print(f"  Time frames zeroed: "
              f"{masked_time}")


aug = AudioAugmenter()
aug.run()

SpecAugment (Park et al., 2019) deserves special attention. It's the audio equivalent of random erasing in vision -- mask out random frequency bands and time segments in the spectrogram during training. It was a key ingredient in Google's state-of-the-art speech recognition systems and is now used in virtually every serious ASR pipeline. The beauty is that it operates on the spectrogram (after feature extraction), so it's trivially easy to implement as a training-time transform.

The representation landscape

Let me put all the representations we've covered in context, because understanding when to use which representation is half the battle in audio AI:

import numpy as np


class RepresentationComparator:
    """Compare audio representations
    in terms of size, information, and
    typical use cases."""

    def __init__(self, sr=16000,
                 duration=10.0):
        self.sr = sr
        self.duration = duration
        self.n_samples = int(sr * duration)

    def compute_sizes(self):
        """Calculate data sizes for each
        representation of a 10-second clip."""
        n = self.n_samples
        n_fft = 1024
        hop = 256
        n_frames = 1 + (n - n_fft) // hop
        n_freqs = n_fft // 2 + 1

        reps = {
            "Raw waveform": n,
            "STFT magnitude": n_freqs * n_frames,
            "Mel spectrogram (80)": 80 * n_frames,
            "Mel spectrogram (128)": 128 * n_frames,
            "MFCC (13)": 13 * n_frames,
            "MFCC + deltas (39)": 39 * n_frames,
        }

        print(f"Audio: {self.duration:.0f}s "
              f"at {self.sr} Hz")
        print(f"Time frames (hop={hop}): "
              f"{n_frames}")
        print(f"\n{'Representation':<24} "
              f"{'Shape':>16} {'Values':>10} "
              f"{'Ratio':>8}")
        print("-" * 62)
        baseline = n
        for name, size in reps.items():
            if "waveform" in name:
                shape = f"({n},)"
            elif "STFT" in name:
                shape = (f"({n_freqs}, "
                         f"{n_frames})")
            elif "Mel" in name:
                bands = 80 if "80" in name else 128
                shape = (f"({bands}, "
                         f"{n_frames})")
            elif "39" in name:
                shape = f"(39, {n_frames})"
            else:
                shape = f"(13, {n_frames})"
            ratio = baseline / size
            print(f"{name:<24} {shape:>16} "
                  f"{size:>10,} "
                  f"{ratio:>7.1f}x")

        print(f"\nTypical use cases:")
        print(f"  Waveform: WaveNet, "
              f"SampleRNN (raw audio gen)")
        print(f"  STFT: phase-sensitive "
              f"tasks (source separation)")
        print(f"  Mel 80: Whisper, most "
              f"modern speech/audio models")
        print(f"  Mel 128: higher detail "
              f"for music analysis")
        print(f"  MFCC 13: lightweight "
              f"classifiers, keyword spotting")
        print(f"  MFCC 39: traditional ASR "
              f"(pre-deep-learning)")


comp = RepresentationComparator()
comp.compute_sizes()

The compression ratio tells a clear story. Going from raw waveform (160,000 values) to Mel spectrogram (80 x 625 = 50,000 values) is a 3.2x compression, but the Mel spectrogram retains the perceptually important information. MFCCs compress further to 13 x 625 = 8,125 values (19.7x compression), but at the cost of discarding fine spectral detail that modern deep networks can exploit. The trend in the field is clear: use the least-compressed representation your compute budget allows, and let the network learn what to extract.

Samengevat

  • Sound is pressure waves converted to a 1D array of amplitude samples; sample rate (Hz) determines the maximum capturable frequency via the Nyquist theorem (max frequency = sample_rate / 2), and aliasing corrupts any frequency above this limit;
  • the Fourier Transform decomposes a signal into its frequency components; the Short-Time Fourier Transform (STFT) does this over sliding windows (typically 25ms with 10ms hop) to produce a time-frequency representation called a spectrogram;
  • spectrograms are 2D images (frequency x time) of audio energy in decibels -- the bridge between audio processing and the vision models we've already built; they are THE standard input for modern audio AI;
  • the Mel scale maps physical frequencies to human perceptual pitch (logarithmic), giving more resolution to low frequencies where hearing is most sensitive; Mel spectrograms compress 513 linear frequency bins to 80-128 perceptually-spaced bands;
  • MFCCs apply DCT to log-Mel spectrograms for compact spectral shape descriptors (39 dimensions with deltas); they dominated pre-deep-learning speech processing but modern models learn directly from Mel spectrograms;
  • SpecAugment (masking random frequency bands and time segments) is the most important audio augmentation technique, directly analogous to random erasing in vision;
  • the standard pipeline: load -> resample to 16 kHz -> mono -> normalize -> pad/truncate -> Mel spectrogram -> log scale -> feed to model.

We've laid the foundation for the entire audio AI arc in this episode. Sound is now data, and we know how to transform that data into representations neural networks can learn from. The audio domain has its own unique challenges (temporal structure, the importance of phase for some tasks, the massive compression from waveform to features), but the core machine learning techniques from the previous arcs -- CNNs, transformers, attention mechanisms -- all apply directly once you have the right representation.

Exercises

Exercise 1: Build a frequency content analyzer. Create a class FrequencyAnalyzer that: (a) generates three synthetic signals of 2 seconds each at 16,000 Hz sample rate: (1) a pure 440 Hz sine wave, (2) a chord with 440 Hz + 554 Hz + 659 Hz (A major chord), (3) white noise from numpy's random generator with seed=42, (b) implements compute_spectrum(signal) that computes the FFT magnitude spectrum, (c) implements spectral_centroid(freqs, magnitudes) that computes the weighted average frequency: sum(f * m) / sum(m) where f is frequency and m is magnitude, (d) implements spectral_bandwidth(freqs, magnitudes, centroid) that computes the weighted standard deviation of frequencies around the centroid, (e) implements spectral_rolloff(freqs, magnitudes, percentile=85) that finds the frequency below which the given percentile of total spectral energy is contained, (f) prints a comparison table showing centroid, bandwidth, and rolloff for all three signals. Verify that white noise has the highest centroid (energy spread across all frequencies), the chord has intermediate values, and the pure tone has the narrowest bandwidth (all energy at one frequency).

Exercise 2: Build a Mel filterbank visualizer and verifier. Create a class MelFilterbankAnalyzer that: (a) implements the Hz-to-Mel and Mel-to-Hz conversion formulas (Mel = 2595 * log10(1 + Hz/700)), (b) builds a triangular Mel filterbank for sr=16000, n_fft=1024, n_mels=40, fmin=0, fmax=8000, (c) computes for each filter: center frequency (Hz), bandwidth (Hz), and peak height, (d) verifies two properties: (1) that every frequency bin from fmin to fmax is covered by at least one filter with nonzero weight (no gaps), (2) that the sum of all filters at any frequency bin is approximately 1.0 (energy conservation), (e) computes the average bandwidth for the lowest 10 filters vs the highest 10 filters, (f) prints a summary table showing filter index, center frequency, bandwidth, and peak height for filters 0, 9, 19, 29, 39. Verify that low-frequency filters have narrow bandwidth (fine pitch resolution) and high-frequency filters have wide bandwidth (coarse resolution), and that the bandwidth ratio between highest and lowest filters is approximately 10:1 or more.

Exercise 3: Build a time-frequency resolution tradeoff analyzer. Create a class ResolutionAnalyzer that: (a) generates a 1-second synthetic signal at 16,000 Hz that contains: a 500 Hz tone in the first 0.5 seconds ONLY, and a 600 Hz tone in the last 0.5 seconds ONLY (with zero-crossings at the transition point to avoid clicks), (b) computes the STFT (magnitude spectrogram in dB) with four different window sizes: 256, 512, 1024, 2048 samples (all with hop_length = window_size // 4), (c) for each window size, measures: (1) frequency resolution (sr / n_fft Hz), (2) time resolution (hop_length / sr seconds), (3) whether the two tones (500 Hz and 600 Hz) are resolvable as separate peaks in the frequency axis (check if the frequency resolution is smaller than the 100 Hz gap between them), (4) whether the time boundary (0.5 seconds) is clearly localized (check how many time frames span the transition region), (d) prints a comparison table showing window size, freq resolution, time resolution, frequency-resolvable (yes/no), and number of frames in the transition zone. Verify that small windows give good time resolution but poor frequency resolution (the two tones blur together), while large windows give good frequency resolution but poor time resolution (the transition is smeared across many frames).

Bedankt en tot de volgende keer!

scipio@scipio