Learn Zig Series (#169) - Audio Fundamentals: PCM and Buffers
Learn Zig Series (#169) - Audio Fundamentals: PCM and Buffers
What will I learn?
- What sound actually is to a computer -- PCM samples in a buffer -- and how sample rate, bit depth and channel count turn a continuous wave into a finite array of numbers;
- How to model a clean, allocator-owned audio buffer in Zig, and the interleaved-vs-planar layout decision that every audio codebase has to make;
- How to synthesize a real signal (a sine wave) into that buffer, and convert between the two formats that matter: 16-bit integer PCM and 32-bit float;
- How to write and read a WAV file by hand -- the RIFF header byte by byte -- so your generated sound lands on disk and plays in any media player;
- How to test audio code sanely (round-trips, RMS levels, invariants) when "listen to it" is not an option in a test suite;
- A few honest performance notes on the audio hot path, and how C, Rust and Go tackle the same PCM work.
Requirements
- A working modern computer running macOS, Windows or Ubuntu;
- An installed Zig 0.14+ distribution (download from ziglang.org) -- the code here is written against Zig 0.16;
- The slices and buffers from episode 5, allocators from episode 7, error unions from episode 4, file I/O from episode 10, and the testing habits from episode 12;
- It helps to remember the
Framebufferidea from episode 156 -- audio is the same trick, one dimension down; - The ambition to learn Zig programming.
Difficulty
- Advanced
Curriculum (of the Learn Zig Series):
- Zig Programming Tutorial - ep001 - Intro
- Learn Zig Series (#2) - Hello Zig, Variables and Types
- Learn Zig Series (#3) - Functions and Control Flow
- Learn Zig Series (#4) - Error Handling (Zig's Best Feature)
- Learn Zig Series (#5) - Arrays, Slices, and Strings
- Learn Zig Series (#6) - Structs, Enums, and Tagged Unions
- Learn Zig Series (#7) - Memory Management and Allocators
- Learn Zig Series (#8) - Pointers and Memory Layout
- Learn Zig Series (#9) - Comptime (Zig's Superpower)
- Learn Zig Series (#10) - Project Structure, Modules, and File I/O
- Learn Zig Series (#11) - Mini Project: Building a Step Sequencer
- Learn Zig Series (#12) - Testing and Test-Driven Development
- Learn Zig Series (#13) - Interfaces via Type Erasure
- Learn Zig Series (#14) - Generics with Comptime Parameters
- Learn Zig Series (#15) - The Build System (build.zig)
- Learn Zig Series (#16) - Sentinel-Terminated Types and C Strings
- Learn Zig Series (#17) - Packed Structs and Bit Manipulation
- Learn Zig Series (#18b) - Addendum: Async Returns in Zig 0.16
- Learn Zig Series (#19) - SIMD with @Vector
- Learn Zig Series (#20) - Working with JSON
- Learn Zig Series (#21) - Networking and TCP Sockets
- Learn Zig Series (#22) - Hash Maps and Data Structures
- Learn Zig Series (#23) - Iterators and Lazy Evaluation
- Learn Zig Series (#24) - Logging, Formatting, and Debug Output
- Learn Zig Series (#25) - Mini Project: HTTP Status Checker
- Learn Zig Series (#26) - Writing a Custom Allocator
- Learn Zig Series (#27) - C Interop: Calling C from Zig
- Learn Zig Series (#28) - C Interop: Exposing Zig to C
- Learn Zig Series (#29) - Inline Assembly and Low-Level Control
- Learn Zig Series (#30) - Thread Safety and Atomics
- Learn Zig Series (#31) - Memory-Mapped I/O and Files
- Learn Zig Series (#32) - Compile-Time Reflection with @typeInfo
- Learn Zig Series (#33) - Building a State Machine with Tagged Unions
- Learn Zig Series (#34) - Performance Profiling and Optimization
- Learn Zig Series (#35) - Cross-Compilation and Target Triples
- Learn Zig Series (#36) - Mini Project: CLI Task Runner
- Learn Zig Series (#37) - Markdown to HTML: Tokenizer and Lexer
- Learn Zig Series (#38) - Markdown to HTML: Parser and AST
- Learn Zig Series (#39) - Markdown to HTML: Renderer and CLI
- Learn Zig Series (#40) - Key-Value Store: In-Memory Store
- Learn Zig Series (#41) - Key-Value Store: Write-Ahead Log
- Learn Zig Series (#42) - Key-Value Store: TCP Server
- Learn Zig Series (#43) - Key-Value Store: Client Library and Benchmarks
- Learn Zig Series (#44) - Image Tool: Reading and Writing PPM/BMP
- Learn Zig Series (#45) - Image Tool: Pixel Operations
- Learn Zig Series (#46) - Image Tool: CLI Pipeline
- Learn Zig Series (#47) - Build a Shell: Parsing Commands
- Learn Zig Series (#48) - Build a Shell: Process Spawning
- Learn Zig Series (#49) - Build a Shell: Built-in Commands
- Learn Zig Series (#50) - Build a Shell: Job Control and Signals
- Learn Zig Series (#51) - HTTP Server: Accept Loop and Parsing
- Learn Zig Series (#52) - HTTP Server: Router and Responses
- Learn Zig Series (#53) - HTTP Server: Static Files and MIME
- Learn Zig Series (#54) - HTTP Server: Middleware and Logging
- Learn Zig Series (#55) - ECS Game Engine: Architecture
- Learn Zig Series (#56) - ECS Game Engine: Component Storage
- Learn Zig Series (#57) - ECS Game Engine: Systems and Queries
- Learn Zig Series (#58) - ECS Game Engine: Terminal Rendering
- Learn Zig Series (#59) - Assembler: Instruction Encoding
- Learn Zig Series (#60) - Assembler: Two-Pass Assembly
- Learn Zig Series (#61) - Assembler: Disassembler and Binary Inspector
- Learn Zig Series (#62) - File Systems: Reading Directories and Metadata
- Learn Zig Series (#63) - File Watching: Detecting Changes
- Learn Zig Series (#64) - Process Management: Fork, Exec, Wait
- Learn Zig Series (#65) - Pipes and Inter-Process Communication
- Learn Zig Series (#66) - Shared Memory and Semaphores
- Learn Zig Series (#67) - Signal Handling Deep Dive
- Learn Zig Series (#68) - Unix Domain Sockets
- Learn Zig Series (#69) - Daemonization: Background Services
- Learn Zig Series (#70) - Timers and Scheduling
- Learn Zig Series (#71) - Resource Limits and Capabilities
- Learn Zig Series (#72) - System Call Wrappers
- Learn Zig Series (#73) - seccomp and Sandboxing
- Learn Zig Series (#74) - ptrace: Process Tracing
- Learn Zig Series (#75) - Reading Kernel State from /proc and /sys
- Learn Zig Series (#76) - Mini Project: Process Monitor
- Learn Zig Series (#77) - Mini Project: File Sync Tool - Part 1
- Learn Zig Series (#78) - Mini Project: File Sync Tool - Part 2: Delta Transfer
- Learn Zig Series (#79) - Mini Project: File Sync Tool - Part 3: Network Protocol
- Learn Zig Series (#80) - Mini Project: File Sync Tool - Part 4: Polish
- Learn Zig Series (#81) - UDP Sockets and Datagrams
- Learn Zig Series (#82) - DNS Resolver from Scratch
- Learn Zig Series (#83) - DNS Server Implementation
- Learn Zig Series (#84) - HTTP/1.1 Deep Dive
- Learn Zig Series (#85) - HTTP/2 Frames and Streams
- Learn Zig Series (#86) - TLS via C Interop
- Learn Zig Series (#87) - WebSocket Protocol
- Learn Zig Series (#88) - WebSocket Server
- Learn Zig Series (#89) - MQTT Messaging Protocol
- Learn Zig Series (#90) - Protocol Buffers Serialization
- Learn Zig Series (#91) - MessagePack Format
- Learn Zig Series (#92) - gRPC Service in Zig
- Learn Zig Series (#93) - SOCKS5 Proxy
- Learn Zig Series (#94) - NAT Traversal and Hole Punching
- Learn Zig Series (#95) - Mini Project: Chat Server - Protocol Design
- Learn Zig Series (#96) - Mini Project: Chat Server - Server Core
- Learn Zig Series (#97) - Mini Project: Chat Server - Client TUI
- Learn Zig Series (#98) - Mini Project: Chat Server - Rooms and History
- Learn Zig Series (#99) - Mini Project: DNS-over-HTTPS Proxy
- Learn Zig Series (#100) - Mini Project: Port Scanner
- Learn Zig Series (#101) - Mini Project: HTTP Load Tester - Part 1
- Learn Zig Series (#102) - Mini Project: HTTP Load Tester - Part 2
- Learn Zig Series (#103) - Mini Project: Reverse Proxy - Routing
- Learn Zig Series (#104) - Mini Project: Reverse Proxy - Load Balancing
- Learn Zig Series (#105) - Mini Project: Reverse Proxy - Health Checks
- Learn Zig Series (#106) - Linked Lists: Singly and Doubly
- Learn Zig Series (#107) - Skip Lists
- Learn Zig Series (#108) - B-Trees
- Learn Zig Series (#109) - Red-Black Trees
- Learn Zig Series (#110) - Tries: Prefix Trees
- Learn Zig Series (#111) - Bloom Filters
- Learn Zig Series (#112) - Cuckoo Filters
- Learn Zig Series (#113) - Ring Buffers: Lock-Free
- Learn Zig Series (#114) - Memory Pools
- Learn Zig Series (#115) - Slab Allocators
- Learn Zig Series (#116) - Sorting Algorithms in Zig
- Learn Zig Series (#117) - Binary Search Variations
- Learn Zig Series (#118) - Graph Representation
- Learn Zig Series (#119) - BFS and DFS
- Learn Zig Series (#120) - Dijkstra and A*
- Learn Zig Series (#121) - Topological Sort
- Learn Zig Series (#122) - Union-Find
- Learn Zig Series (#123) - LRU Cache
- Learn Zig Series (#124) - Consistent Hashing
- Learn Zig Series (#125) - Mini Project: Search Engine - Inverted Index
- Learn Zig Series (#126) - Mini Project: Search Engine - TF-IDF
- Learn Zig Series (#127) - Mini Project: Search Engine - Query Parser
- Learn Zig Series (#128) - Mini Project: Database Engine - Page Storage
- Learn Zig Series (#129) - Mini Project: Database Engine - B-Tree Index
- Learn Zig Series (#130) - Mini Project: Database Engine - SQL Parser
- Learn Zig Series (#131) - Lexing a Simple Language
- Learn Zig Series (#132) - Recursive Descent Parsing
- Learn Zig Series (#133) - AST Design and Traversal
- Learn Zig Series (#134) - Type Checking
- Learn Zig Series (#135) - Bytecode Design
- Learn Zig Series (#136) - Stack-Based Virtual Machine
- Learn Zig Series (#137) - Closures and Upvalues
- Learn Zig Series (#138) - Garbage Collection: Mark and Sweep
- Learn Zig Series (#139) - Garbage Collection: Generational
- Learn Zig Series (#140) - JIT Compilation Basics
- Learn Zig Series (#141) - Regex: Thompson NFA
- Learn Zig Series (#142) - Regex: NFA to DFA
- Learn Zig Series (#143) - Regex: Matching Engine
- Learn Zig Series (#144) - Code Generation: AST to Machine Code
- Learn Zig Series (#145) - Register Allocation
- Learn Zig Series (#146) - Mini Project: Calculator - Lexer/Parser
- Learn Zig Series (#147) - Mini Project: Calculator - Interpreter
- Learn Zig Series (#148) - Mini Project: Calculator - Bytecode Compiler
- Learn Zig Series (#149) - Mini Project: Calculator - VM with Debugger
- Learn Zig Series (#150) - Mini Project: Lisp - Reader
- Learn Zig Series (#151) - Mini Project: Lisp - Evaluator
- Learn Zig Series (#152) - Mini Project: Lisp - Special Forms and Macros
- Learn Zig Series (#153) - Mini Project: Lisp - Standard Library
- Learn Zig Series (#154) - Mini Project: Regex Engine - NFA
- Learn Zig Series (#155) - Mini Project: Regex Engine - Matching
- Learn Zig Series (#156) - Framebuffer Basics
- Learn Zig Series (#157) - Line Drawing: Bresenham
- Learn Zig Series (#158) - Circle and Ellipse Rasterization
- Learn Zig Series (#159) - Polygon Filling: Scanline
- Learn Zig Series (#160) - 2D Transform Matrices
- Learn Zig Series (#161) - Double Buffering and Vsync
- Learn Zig Series (#162) - Sprite Rendering and Tile Maps
- Learn Zig Series (#163) - Bitmap Font Rendering
- Learn Zig Series (#164) - TrueType Parsing
- Learn Zig Series (#165) - Color Spaces: RGB, HSV, sRGB
- Learn Zig Series (#166) - Alpha Blending and Compositing
- Learn Zig Series (#167) - PNG Decoder in Zig
- Learn Zig Series (#168) - JPEG Decoder Basics
- Learn Zig Series (#169) - Audio Fundamentals: PCM and Buffers (this post)
Learn Zig Series (#169) - Audio Fundamentals: PCM and Buffers
I closed the JPEG episode with a small tease: there is a whole other sense we have not touched yet, one that is "also just numbers in a buffer if you look at it right." That sense is hearing, and today we start looking at it right. For the last twenty-five-odd episodes we have made the machine show things -- pixels, shapes, glyphs, decoded photos. Now we make it speak. And the beautiful part, the part that makes this a natural continuation and not a hard left turn, is that sound in a computer is exactly the same idea as a framebuffer (episode 156), just one dimension down: a big flat array of numbers, sampled at regular intervals, that a device turns into something your body perceives. A framebuffer is samples across space; an audio buffer is samples across time. Master the buffer and you have the foundation for everything that follows -- synthesis, mixing, effects, the lot. Here we go!
But first, the three loose ends from the JPEG episode.
Solutions to Episode 168 Exercises
Exercise 1 -- finish the DHT parser and Huffman decode. JPEG's DHT segment gives, per bit-length 1..16, how many codes exist, then the symbols in code order. That is the canonical Huffman layout, so we can rebuild the codes with a running counter in stead of storing a tree. We also need a bit reader that unstuffs the 0xFF 0x00 escape and stops cleanly when it hits a real marker:
const std = @import("std");
pub const HuffTable = struct {
counts: [17]u8 = [_]u8{0} ** 17, // counts[len] = number of codes of that length
symbols: [256]u8 = undefined,
min_code: [17]i32 = [_]i32{0} ** 17,
max_code: [17]i32 = [_]i32{-1} ** 17, // -1 = no codes of this length
val_ptr: [17]usize = [_]usize{0} ** 17,
};
fn parseDht(payload: []const u8, tables: *[4]?HuffTable) !void {
var i: usize = 0;
while (i < payload.len) {
const tc_th = payload[i];
const id = tc_th & 0x0F;
if (id >= 4) return error.InvalidData;
i += 1;
var t = HuffTable{};
var total: usize = 0;
for (1..17) |len| {
t.counts[len] = payload[i];
total += payload[i];
i += 1;
}
if (i + total > payload.len) return error.Truncated;
for (0..total) |k| t.symbols[k] = payload[i + k];
i += total;
// Assign canonical codes: first code of each length, and where its symbols start.
var code: i32 = 0;
var k: usize = 0;
for (1..17) |len| {
if (t.counts[len] != 0) {
t.val_ptr[len] = k;
t.min_code[len] = code;
code += t.counts[len];
t.max_code[len] = code - 1;
k += t.counts[len];
}
code <<= 1;
}
tables[id] = t;
}
}
pub const BitReader = struct {
data: []const u8,
pos: usize = 0,
buf: u8 = 0,
bits_left: u4 = 0,
fn nextByte(self: *BitReader) ?u8 {
if (self.pos >= self.data.len) return null;
const b = self.data[self.pos];
self.pos += 1;
if (b == 0xFF) {
const stuff = if (self.pos < self.data.len) self.data[self.pos] else 0xD9;
if (stuff == 0x00) {
self.pos += 1; // 0xFF 0x00 -> literal 0xFF
} else {
self.pos -= 1; // a real marker: rewind, the scan is over
return null;
}
}
return b;
}
fn getBit(self: *BitReader) ?u1 {
if (self.bits_left == 0) {
self.buf = self.nextByte() orelse return null;
self.bits_left = 8;
}
self.bits_left -= 1;
return @intCast((self.buf >> @intCast(self.bits_left)) & 1);
}
};
fn decodeSymbol(br: *BitReader, t: *const HuffTable) !u8 {
var code: i32 = 0;
for (1..17) |len| {
const bit = br.getBit() orelse return error.Truncated;
code = (code << 1) | @as(i32, bit);
if (t.max_code[len] >= 0 and code <= t.max_code[len]) {
const idx = t.val_ptr[len] + @as(usize, @intCast(code - t.min_code[len]));
return t.symbols[idx];
}
}
return error.InvalidData;
}
The key insight is that a canonical Huffman code needs no tree at all: because codes of the same length are consecutive integers, a min_code/max_code pair per length lets us decode by accumulating bits and checking, at each length, "does this fit the range for this length yet?" The byte-stuffing check in nextByte is the JPEG-specific wrinkle -- a literal 0xFF in the entropy stream is written as 0xFF 0x00, and any other 0xFF xx is a marker that ends the scan.
Exercise 2 -- add 4:2:0 chroma upsampling. Baseline JPEGs usually store Cb/Cr at half width and half height, so before color conversion we parse the frame header's per-component sampling factors and blow the chroma planes back up:
const std = @import("std");
pub const Component = struct { id: u8, h: u4, v: u4, quant_id: u8 };
fn parseSof0(payload: []const u8, comps: *[3]Component) !struct { w: u16, h: u16, n: u8 } {
if (payload.len < 6) return error.Truncated;
if (payload[0] != 8) return error.UnsupportedFormat; // 8-bit precision only
const h = std.mem.readInt(u16, payload[1..3], .big);
const w = std.mem.readInt(u16, payload[3..5], .big);
const n = payload[5];
if (n > 3 or 6 + @as(usize, n) * 3 > payload.len) return error.InvalidData;
var i: usize = 6;
for (0..n) |c| {
const hv = payload[i + 1];
comps[c] = .{ .id = payload[i], .h = @intCast(hv >> 4), .v = @intCast(hv & 0x0F), .quant_id = payload[i + 2] };
i += 3;
}
return .{ .w = w, .h = h, .n = n };
}
// Nearest-neighbour 2x upsample of a half-resolution plane.
fn upsample2x(src: []const u8, sw: usize, sh: usize, dst: []u8) void {
const dw = sw * 2;
for (0..sh * 2) |y| {
for (0..dw) |x| dst[y * dw + x] = src[(y / 2) * sw + (x / 2)];
}
}
test "upsample2x doubles each sample in both directions" {
const src = [_]u8{ 10, 20, 30, 40 }; // 2x2
var dst: [16]u8 = undefined; // 4x4
upsample2x(&src, 2, 2, &dst);
try std.testing.expectEqual(@as(u8, 10), dst[0]);
try std.testing.expectEqual(@as(u8, 20), dst[3]);
try std.testing.expectEqual(@as(u8, 40), dst[15]);
}
Parsing the h/v factors is where you learn a component is subsampled: luma typically has h=2, v=2 while chroma has h=1, v=1, meaning chroma covers a 2x2 luma area with one sample. Nearest-neighbour doubling is the honest starting point (block-y but correct); bilinear is the quality upgrade, and comparing the two by PSNR was the stretch goal.
Exercise 3 -- build a comptime IDCT cosine table. Those cos(...) calls in the naive IDCT are constant for every block, so we compute an 8x8 table once, at compile time, and the hot path becomes multiplies and adds:
const std = @import("std");
const idct_cos = blk: {
@setEvalBranchQuota(20000);
var table: [8][8]f32 = undefined;
for (0..8) |x| {
for (0..8) |u| {
const cu: f32 = if (u == 0) std.math.sqrt1_2 else 1.0;
const angle = @as(f32, @floatFromInt(2 * x + 1)) * @as(f32, @floatFromInt(u)) * std.math.pi / 16.0;
table[x][u] = cu * @cos(angle);
}
}
break :blk table;
};
// One 1D inverse DCT row using the precomputed cosines.
fn idct1d(in: [8]f32, out: *[8]f32) void {
for (0..8) |x| {
var sum: f32 = 0;
for (0..8) |u| sum += idct_cos[x][u] * in[u];
out[x] = sum / 2.0;
}
}
test "DC-only input yields a flat row" {
var out: [8]f32 = undefined;
idct1d([_]f32{ 8, 0, 0, 0, 0, 0, 0, 0 }, &out);
for (out) |v| try std.testing.expect(@abs(v - out[0]) < 1e-4);
}
The blk: { ... break :blk table; } is Zig's labeled-block-as-expression form, and because it runs at comptime the table is baked straight into the binary -- no runtime cos at all. The @setEvalBranchQuota just tells the comptime interpreter it is allowed to loop this many times. Notice I dropped from the four-nested-loop 2D IDCT to two 1D passes (rows then columns): that separability is the real speedup, the table is the cherry on top. Right, the image arc is properly behind us. Now, sound.
What sound actually is to a computer
Sound in the physical world is a continuous pressure wave wobbling through air. A microphone turns that wobble into a continuously varying voltage. But a computer cannot store "continuous" anything -- it stores numbers, discretely. So we do two discretizations. First, sampling: we measure the voltage at fixed instants, thousands of times a second. The sample rate is how many measurements per second: 44100 Hz (CD quality) and 48000 Hz (video/pro audio) are the two you will meet constantly. Second, quantization: each measurement is rounded to a fixed-precision number. The bit depth is how many bits that number gets -- 16-bit signed integers are the classic (values from -32768 to 32767), and 32-bit float (values conventionally in the range -1.0 to 1.0) is what modern audio engines compute in.
That is the whole of PCM -- Pulse Code Modulation -- which is a grand name for a dead-simple idea: uncompressed audio is just an array of samples. No prediction, no frequency transform, none of the JPEG cleverness. A one-second mono clip at 44100 Hz is literally 44100 numbers. Stereo doubles it, because you have two channels (left and right). The Nyquist theorem tells us why 44100 is not arbitrary: to represent frequencies up to F, you must sample at least 2 * F, and human hearing tops out around 20000 Hz, so 44100 leaves headroom for the anti-aliasing filter. Let me pin the vocabulary in a tiny module of constants and helpers:
const std = @import("std");
pub const Sample = i16; // one 16-bit PCM sample, -32768 .. 32767
pub const AudioSpec = struct {
sample_rate: u32 = 44_100,
channels: u8 = 2,
bits_per_sample: u16 = 16,
/// Bytes one frame occupies (all channels for a single instant in time).
pub fn frameBytes(self: AudioSpec) usize {
return @as(usize, self.channels) * (self.bits_per_sample / 8);
}
/// How many sample-frames represent the given duration.
pub fn framesForSeconds(self: AudioSpec, seconds: f32) usize {
return @intFromFloat(@round(@as(f32, @floatFromInt(self.sample_rate)) * seconds));
}
};
test "one second of CD-quality stereo is the size we expect" {
const spec = AudioSpec{};
try std.testing.expectEqual(@as(usize, 44_100), spec.framesForSeconds(1.0));
try std.testing.expectEqual(@as(usize, 4), spec.frameBytes()); // 2 ch * 2 bytes
}
I want to draw a sharp line between a sample and a frame early, because mixing them up is the classic audio bug. A sample is one number for one channel. A frame is all the channels for one instant -- so a stereo frame is two samples (left, right). Durations and playback positions are counted in frames, never in samples, because a frame is what advances the clock. Getting this wrong gives you audio that plays at half speed or double speed, or with the channels swapped -- ask me how I know. ;-)
An owned audio buffer
Now the workhorse: a buffer type that owns its samples through an allocator (episode 7), just like our Framebuffer owned its pixels. I will store samples interleaved -- L, R, L, R, ... -- because that is what sound card APIs and WAV files expect. (The alternative, planar, stores all the lefts then all the rights; DSP-heavy engines prefer it because it vectorizes cleaner. More on that trade in the performance section.) The buffer knows its spec, so it can answer "how many frames?" and hand you one frame at a time:
const std = @import("std");
pub const AudioBuffer = struct {
samples: []f32, // interleaved, one f32 per sample, range roughly -1.0 .. 1.0
channels: u8,
sample_rate: u32,
allocator: std.mem.Allocator,
pub fn init(allocator: std.mem.Allocator, channels: u8, sample_rate: u32, frames: usize) !AudioBuffer {
const samples = try allocator.alloc(f32, frames * channels);
@memset(samples, 0.0); // silence is all zeros
return .{
.samples = samples,
.channels = channels,
.sample_rate = sample_rate,
.allocator = allocator,
};
}
pub fn deinit(self: *AudioBuffer) void {
self.allocator.free(self.samples);
self.* = undefined;
}
pub fn frameCount(self: AudioBuffer) usize {
return self.samples.len / self.channels;
}
/// The slice of samples belonging to frame `i` (length == channels).
pub fn frame(self: AudioBuffer, i: usize) []f32 {
const start = i * self.channels;
return self.samples[start .. start + self.channels];
}
pub fn durationSeconds(self: AudioBuffer) f32 {
return @as(f32, @floatFromInt(self.frameCount())) / @as(f32, @floatFromInt(self.sample_rate));
}
};
test "an audio buffer reports its geometry correctly" {
var buf = try AudioBuffer.init(std.testing.allocator, 2, 48_000, 24_000);
defer buf.deinit();
try std.testing.expectEqual(@as(usize, 24_000), buf.frameCount());
try std.testing.expectApproxEqAbs(@as(f32, 0.5), buf.durationSeconds(), 1e-6);
try std.testing.expectEqual(@as(usize, 2), buf.frame(10).len);
}
Notice @memset(samples, 0.0) -- silence is literally all zeros, both in float and in integer PCM, which is a small mercy: a freshly-allocated-and-zeroed buffer is already valid, playable silence. The frame(i) accessor returning a slice into the interleaved data (not a copy) is the interleaved layout paying rent: writing buf.frame(i)[0] = left; buf.frame(i)[1] = right; reads naturally and compiles to a couple of stores. And deinit sets self.* = undefined so a use-after-free trips the safety checks in debug builds in stead of quietly reading freed memory (episode 7's discipline).
Making a sound: the sine wave
An empty buffer is not very exciting, so let us fill one. The sine wave is the "hello world" of audio -- a pure tone at one frequency, the building block every other timbre is made from. The math is sample = amplitude * sin(2 * pi * frequency * t), where t is the time in seconds of that sample. Since sample n happens at time n / sample_rate, the whole generator is one loop:
const std = @import("std");
/// Fill a mono f32 buffer with a sine tone. Amplitude 0.0 .. 1.0.
pub fn fillSine(samples: []f32, freq: f32, sample_rate: u32, amplitude: f32) void {
const two_pi = 2.0 * std.math.pi;
const sr: f32 = @floatFromInt(sample_rate);
for (samples, 0..) |*s, n| {
const t = @as(f32, @floatFromInt(n)) / sr;
s.* = amplitude * @sin(two_pi * freq * t);
}
}
test "a 1 Hz sine over one second returns to zero and stays in range" {
var samples: [8]f32 = undefined;
fillSine(&samples, 1.0, 8, 1.0); // 8 samples == exactly one period
try std.testing.expectApproxEqAbs(@as(f32, 0.0), samples[0], 1e-5);
for (samples) |s| try std.testing.expect(s >= -1.0 and s <= 1.0);
}
Computing t fresh from n each iteration is clean and correct, and it never drifts. There is a faster idiom -- keep a running phase and add a per-sample increment 2 * pi * freq / sample_rate each step -- which we will lean on hard when we build oscillators later, because it lets the frequency change smoothly over time. For a fixed tone, though, the t = n / sample_rate form is the one to teach: it is obviously right, and "obviously right" beats "clever" until a profiler says otherwise (episode 34, forever). One thing to watch: keep amplitude below 1.0 with some headroom. A full-scale sine that then gets summed with anything else will clip -- exceed the representable range -- and clipping sounds like a nasty crackle, not a warning dialog.
Converting between float and integer PCM
We compute in float (easy math, no overflow while mixing) but store and transmit in 16-bit integer (half the size, what WAV and sound cards want). So we need two honest conversions. Float-to-int scales the -1.0..1.0 range onto -32768..32767 and must clamp, because a value that wandered past 1.0 during mixing would otherwise wrap around from a loud positive to a loud negative -- an audible pop. Int-to-float is the clean inverse:
const std = @import("std");
pub fn floatToI16(x: f32) i16 {
const scaled = std.math.clamp(x, -1.0, 1.0) * 32767.0;
return @intFromFloat(@round(scaled));
}
pub fn i16ToFloat(s: i16) f32 {
return @as(f32, @floatFromInt(s)) / 32768.0;
}
/// Convert a whole interleaved f32 buffer to i16, in a caller-owned slice.
pub fn convertF32ToI16(src: []const f32, dst: []i16) void {
std.debug.assert(src.len == dst.len);
for (src, dst) |x, *out| out.*= floatToI16(x);
}
test "float/int round-trip stays close and clipping is contained" {
const values = [_]f32{ 0.0, 0.5, -0.5, 0.999, 1.5, -2.0 };
for (values) |v| {
const back = i16ToFloat(floatToI16(v));
const expected = std.math.clamp(v, -1.0, 1.0);
try std.testing.expect(@abs(back - expected) < 0.001);
}
try std.testing.expectEqual(@as(i16, 32767), floatToI16(2.0)); // clamped high
try std.testing.expectEqual(@as(i16, -32767), floatToI16(-2.0)); // clamped low
}
The asymmetry in the scale factors (* 32767.0 going out, / 32768.0 coming in) is a genuine subtlety, not a typo: signed 16-bit runs from -32768 to +32767, so there is one more negative code than positive. Different codecs pick different conventions here; I scale up by 32767 so the loudest positive maps exactly to the max code and nothing overflows, and down by 32768 so a full-scale negative maps cleanly. Whichever pair you pick, pick one and be consistent, because mixing conventions across a pipeline gives you a tiny DC offset that you will chase for an afternoon. The clamp is the real guardrail: without it, @intFromFloat on an out-of-range value is illegal behavior, and with it, overdriven audio degrades into flat-topped distortion in stead of exploding.
Writing a WAV file by hand
Time for the payoff -- getting sound onto disk in a format any media player understands. WAV is a RIFF container, and for uncompressed PCM it is refreshingly simple: a 44-byte header followed by the raw interleaved samples. That is it. No compression, no tables, no zig-zag. The header is three chunks -- the RIFF wrapper, a fmt chunk describing the format, and a data chunk length -- all little-endian. Let me build that header explicitly with std.mem.writeInt, then dump the samples:
const std = @import("std");
fn writeWavHeader(w: anytype, spec: AudioSpec, data_bytes: u32) !void {
const byte_rate = spec.sample_rate * @as(u32, spec.channels) * (spec.bits_per_sample / 8);
const block_align: u16 = @intCast(spec.channels * (spec.bits_per_sample / 8));
var buf: [4]u8 = undefined;
try w.writeAll("RIFF");
std.mem.writeInt(u32, &buf, 36 + data_bytes, .little); // total size minus 8
try w.writeAll(&buf);
try w.writeAll("WAVE");
try w.writeAll("fmt ");
std.mem.writeInt(u32, &buf, 16, .little); // PCM fmt chunk is 16 bytes
try w.writeAll(&buf);
var h: [2]u8 = undefined;
std.mem.writeInt(u16, &h, 1, .little); // format 1 == PCM
try w.writeAll(&h);
std.mem.writeInt(u16, &h, spec.channels, .little);
try w.writeAll(&h);
std.mem.writeInt(u32, &buf, spec.sample_rate, .little);
try w.writeAll(&buf);
std.mem.writeInt(u32, &buf, byte_rate, .little);
try w.writeAll(&buf);
std.mem.writeInt(u16, &h, block_align, .little);
try w.writeAll(&h);
std.mem.writeInt(u16, &h, spec.bits_per_sample, .little);
try w.writeAll(&h);
try w.writeAll("data");
std.mem.writeInt(u32, &buf, data_bytes, .little);
try w.writeAll(&buf);
}
/// Write interleaved i16 samples to `path` as a PCM WAV.
pub fn writeWav(path: []const u8, spec: AudioSpec, samples: []const i16) !void {
const file = try std.fs.cwd().createFile(path, .{});
defer file.close();
var wbuf: [4096]u8 = undefined;
var fw = file.writer(&wbuf);
const w = &fw.interface;
const data_bytes: u32 = @intCast(samples.len * 2);
try writeWavHeader(w, spec, data_bytes);
for (samples) |s| {
var two: [2]u8 = undefined;
std.mem.writeInt(i16, &two, s, .little);
try w.writeAll(&two);
}
try w.flush();
}
Every multi-byte field is written little-endian because that is what RIFF mandates (an accident of its x86 origins), which is why I reach for std.mem.writeInt(..., .little) on every number in stead of a packed struct -- being explicit about endianness here means the file is correct on a big-endian machine too, for free. The two magic constants worth memorizing: 36 + data_bytes is "everything after the first 8 bytes" (the file size minus the RIFF tag and this very length field), and the fmt chunk is 16 bytes for plain PCM. Get the data length wrong and players either truncate your audio or read garbage past the end -- so I compute it once, from the sample count, and never by hand. Writing the samples one i16 at a time is clear but not fast; for a big buffer you would @memcpy the whole slice reinterpreted as bytes, which we will do the moment it matters.
Reading a WAV back
A writer with no reader is only half a tool, and reading is where we practice the same defensive parsing the image episodes drilled into us: an untrusted file is a hostile file, so validate the magic, and never trust a length field to fit in the bytes you actually have:
const std = @import("std");
pub const WavError = error{ BadMagic, Truncated, Unsupported };
pub const LoadedWav = struct {
spec: AudioSpec,
samples: []i16,
allocator: std.mem.Allocator,
pub fn deinit(self: *LoadedWav) void {
self.allocator.free(self.samples);
self.* = undefined;
}
};
pub fn readWav(allocator: std.mem.Allocator, bytes: []const u8) !LoadedWav {
if (bytes.len < 44) return WavError.Truncated;
if (!std.mem.eql(u8, bytes[0..4], "RIFF") or !std.mem.eql(u8, bytes[8..12], "WAVE"))
return WavError.BadMagic;
const audio_format = std.mem.readInt(u16, bytes[20..22], .little);
if (audio_format != 1) return WavError.Unsupported; // PCM only
const channels = bytes[22];
const sample_rate = std.mem.readInt(u32, bytes[24..28], .little);
const bits = std.mem.readInt(u16, bytes[34..36], .little);
if (bits != 16) return WavError.Unsupported;
const data_len = std.mem.readInt(u32, bytes[40..44], .little);
if (44 + @as(usize, data_len) > bytes.len) return WavError.Truncated;
const raw = bytes[44 .. 44 + data_len];
const count = raw.len / 2;
const samples = try allocator.alloc(i16, count);
errdefer allocator.free(samples);
for (0..count) |i| samples[i] = std.mem.readInt(i16, raw[i * 2 ..][0..2], .little);
return .{
.spec = .{ .sample_rate = sample_rate, .channels = channels, .bits_per_sample = bits },
.samples = samples,
.allocator = allocator,
};
}
I have hard-coded the classic 44-byte layout on purpose, and I want to be honest that this is the teaching reader, not the bullet-proof one: real WAV files can carry extra chunks (LIST, fact, a padded fmt) between the header and the data, so a production reader walks the chunk list looking for fmt and data in stead of assuming fixed offsets -- exactly the marker-walk pattern from the JPEG episode. But the shape is identical and the defensive habits are the same: check the magic, reject non-PCM and non-16-bit up front, and clamp the data length to the bytes we truly hold before slicing. The errdefer frees the samples if a later step fails, so an error never leaks the allocation (episode 7 again).
Testing audio you cannot listen to
"Run the tests and listen" is not a thing a CI pipeline can do, so audio testing leans on the same instinct the JPEG episode used for a lossy codec: test the invariant, not the ear. Three tools carry you a long way. First, round-trips: write a buffer to WAV bytes and read them back; the samples must match exactly, because WAV PCM is lossless. Second, level metrics like RMS (root-mean-square, the perceived loudness): a sine of amplitude A has a known RMS of A / sqrt(2), so you can assert your generator produced the loudness you asked for. Third, structural invariants: silence is all zeros, a mono buffer has one sample per frame, durations line up. Here is RMS and a round-trip check:
const std = @import("std");
pub fn rms(samples: []const f32) f32 {
if (samples.len == 0) return 0;
var sum_sq: f64 = 0;
for (samples) |s| sum_sq += @as(f64, s) * @as(f64, s);
return @floatCast(@sqrt(sum_sq / @as(f64, @floatFromInt(samples.len))));
}
test "a full-scale sine has RMS near amplitude/sqrt(2)" {
var samples: [44_100]f32 = undefined;
fillSine(&samples, 440.0, 44_100, 0.8);
const expected = 0.8 / std.math.sqrt2;
try std.testing.expectApproxEqAbs(expected, rms(&samples), 0.01);
}
test "i16 sample round-trips through the buffer conversion" {
const original = [_]i16{ 0, 100, -100, 32767, -32768, 12345 };
var back: [6]i16 = undefined;
for (original, 0..) |s, i| back[i] = floatToI16(i16ToFloat(s));
// Every value survives except the asymmetric extreme, which clamps by one code.
for (original, back) |a, b| try std.testing.expect(@abs(@as(i32, a) - @as(i32, b)) <= 1);
}
RMS is the honest single number for "how loud is this", the audio cousin of PSNR: it collapses a whole buffer into one figure you can assert a floor or a target on. Computing the sum of squares in f64 even though the samples are f32 matters more than it looks -- a full second is 44100 terms, and accumulating tiny squares in f32 loses precision fast (episode 34's numerical-care footnote). The round-trip test even documents the one-code asymmetry we discussed: -32768 cannot come back exactly through a * 32767 scale, and asserting "within one code" states that honestly in stead of pretending the conversion is perfect.
Performance: the audio hot path
Audio has a performance constraint most code does not: it is soft-real-time. The sound card asks for the next block of samples on a strict clock (every few milliseconds), and if you are late, the listener hears a glitch -- a click or a dropout. There is no "just render the frame a bit late" grace like graphics sometimes has. So the rules of the hot path are strict, and mostly they are about what you do not do. The callback that fills the sound card's buffer must never allocate, never take a lock that a slow thread holds, never touch the filesystem, and never call anything that might block. You prepare buffers ahead of time on other threads and hand them over through a lock-free queue -- which is exactly why we built a lock-free ring buffer back in episode 113. A few concrete levers, in episode-34 measure-first order:
- Interleaved vs planar is a real choice. Interleaved matches the hardware and WAV, so it wins for I/O. Planar (all lefts, then all rights) lets a per-channel gain or filter run as a straight
@Vectorsweep (episode 19) with no stride, so DSP-heavy engines convert to planar internally and back at the edges. - Batch the sample math. Multiplying a whole buffer by a gain is a textbook SIMD case; do it over slices with
@Vector, not one float at a time in a hot callback. - Never allocate in the callback. Size your buffers once, up front, from the spec. A GC pause or a malloc in the audio thread is a guaranteed glitch -- this is where Zig's explicit, allocate-outside-the-hot-path model is a genuine advantage over a garbage-collected runtime.
- Prefer
@memcpyfor format-flat conversions. When two layouts differ only by element type reinterpretation, a bulk copy crushes a per-sample loop.
The meta-point is the same as always: for writing a WAV once, none of this matters and the per-sample loop is fine; for a live callback at 48 kHz, the whole game is do the work early, hand it over cheaply, and touch nothing slow on the audio thread.
The same job in C, Rust and Go
The concepts -- samples, frames, sample rate, interleaving, a PCM buffer -- are universal; what differs is the posture toward memory and who owns the buffer. In C, a PCM buffer is a raw pointer and a length you carry around by hand, and you talk to the hardware through a library like PortAudio or miniaudio (or ALSA directly on Linux), passing a callback that must obey the exact same no-allocation hot-path rules:
// C with miniaudio-style callback: raw interleaved f32, you track frame count yourself.
void data_callback(ma_device* dev, void* out, const void* in, ma_uint32 frame_count) {
float* samples = (float*)out;
for (ma_uint32 f = 0; f < frame_count; f++) {
samples[f * 2 + 0] = 0.0f; // left
samples[f * 2 + 1] = 0.0f; // right
}
}
Rust models the buffer as a Vec<f32> (or a slice) and reaches for cpal for device I/O; the samples are the same interleaved floats, but ownership and bounds are checked, and the hound crate reads and writes WAV so nobody hand-rolls the RIFF header:
// Rust: hound writes a WAV; the sample loop is the same math we wrote in Zig.
let spec = hound::WavSpec { channels: 2, sample_rate: 44_100,
bits_per_sample: 16, sample_format: hound::SampleFormat::Int };
let mut writer = hound::WavWriter::create("out.wav", spec)?;
for n in 0..44_100 {
let t = n as f32 / 44_100.0;
let s = (0.8 * (2.0 * std::f32::consts::PI * 440.0 * t).sin() * 32767.0) as i16;
writer.write_sample(s)?; // left
writer.write_sample(s)?; // right
}
Go does not ship audio in its standard library, so you pull a package like oto for playback or go-audio/wav for files; the buffer is a []int16 or []float32 slice and the code reads much like our Zig, garbage-collected rather than allocator-owned -- which is convenient for offline processing and a liability in a real-time callback, where a GC pause is the very glitch we work to avoid. Where does Zig sit? For real device I/O you would bind a C library through the zero-cost interop of episodes 27 and 28 (PortAudio and miniaudio are both single-header-friendly). But the buffer, the synthesis, and the file format -- the parts that teach -- we just wrote from scratch, and Zig's explicitness (an owned buffer with an allocator, endianness spelled out on every field, conversions that clamp) makes the audio pipeline unusually legible.
Exercises
Stereo panning. Extend
fillSineinto afillSineStereothat writes an interleaved stereoAudioBufferand takes apanparameter from -1.0 (hard left) to +1.0 (hard right). Use equal-power panning (left = cos(theta),right = sin(theta)wherethetamaps the pan to 0..pi/2) so the perceived loudness stays constant as the sound moves across the stereo field. Test thatpan = 0gives equal channels and that the summed power is roughly constant across pan positions.A fade envelope. Write
applyFade(samples, sample_rate, fade_ms)that ramps the amplitude linearly from 0 to full over the firstfade_msmilliseconds and back down to 0 over the lastfade_ms. This kills the click you get when a buffer starts or ends on a non-zero sample (an instant jump is a high-frequency pop). Test that the very first and very last samples are near zero and the middle is untouched.Robust WAV reader. Upgrade
readWavto walk the chunk list in stead of assuming fixed offsets: afterWAVE, loop over(4-byte id, 4-byte little-endian length, payload)triples, capturing thefmtanddatachunks wherever they appear and skipping any others (remember chunks are padded to an even length). Test it against a file that has an extraLISTchunk wedged beforedata.
What we learned
- Sound to a computer is PCM: an array of samples taken at a fixed sample rate and rounded to a fixed bit depth -- the exact same "sampled numbers in a buffer" idea as a framebuffer, one dimension down into time;
- The sample vs frame distinction is the one to never blur -- a frame is all channels for one instant, and durations and playback positions are always counted in frames;
- An allocator-owned
AudioBuffer(interleaved samples, plus its spec) is the workhorse, and interleaved-vs-planar is a real layout trade-off between I/O friendliness and SIMD friendliness; - Synthesis starts with the sine wave (
amplitude * sin(2*pi*f*t)), and the two conversions that matter -- f32 to i16 and back -- must clamp to avoid the wrap-around pop, with an honest one-code asymmetry between the scale factors; - A WAV file is a 44-byte RIFF header plus raw little-endian samples, simple enough to write and read by hand, and the reading side wants the same defensive parsing (check magic, clamp lengths) the image episodes taught;
- Testing audio means round-trips (lossless PCM must match exactly), level metrics like RMS, and structural invariants -- test the invariant, not the ear -- while the soft-real-time hot path forbids allocating or blocking, which is where Zig's explicit memory model shines.
That lands us with real, playable sound on disk built from nothing but numbers and a header. We have a buffer, we can fill it, and we can save it -- but so far the only way to hear it is to open the file in a player. The obvious next step is to skip the file and push those samples straight at the speaker in real time, which means talking to the operating system's audio device -- and that is a job for the C interop we sharpened back in episodes 27 and 28. After that, the buffer stops being a static sine and starts becoming an instrument. Plenty still to build.
De groeten, en tot de volgende keer! ;-)