Sound Race — every song generates a 3-lane hazard course, decoded, analyzed, and driven entirely in the browser
How we ship a rhythm-driven WebGL racer with no backend. A Web Audio decode + a Web Worker that runs FFT off the main thread, a beatmap synthesizer that turns spectral flux into a curved elevated road, an audio-clock game loop that never drifts against the music, and a Jamendo-only music source that keeps CORS out of the failure list. All from the production code of /gamecenter/sound-race.
TL;DR — Sound Race is a browser rhythm-runner where the track you drive on — its curves, hills, hazard placements, cluster payoffs — is generated at load time from an FFT of the audio you picked. Four moving parts make it feel right and stay under 60 fps: (1) a Web Worker owns the FFT so the main thread stays free for the pipeline UI; (2) the beatmap synthesizer maps spectral flux → color palette, low/mid balance → curvature, RMS → elevation, so drops feel like hill climbs and breakdowns feel like valleys; (3) the game loop reads
AudioContext.currentTime— neverperformance.now()— so ±10 ms drift never opens up between what you hear and what you see; (4) Jamendo is the only music source because its CDN sends CORS headers that let the browserdecodeAudioDatathe bytes without a proxy. This post walks through the actual code.Live demo: /gamecenter/sound-race · engine tree:
src/features/gamecenter/games/sound-race/
Two of our earlier posts covered games with a fixed world — the Snooker table is 22 balls and six pockets, the 3D Race marble is a randomized tower of blocks. Sound Race is the opposite: the world is the audio. Load an MP3, run 20 seconds of DSP, and now you have a track. Load a different MP3, run the same DSP, and now you have a completely different track — same code, different music, different game.
That’s a fun idea. It also has a very specific set of failure modes. This post walks the four decisions that separate the version that shipped from the four earlier prototypes that didn’t.
The tree that ended up mattering:
sound-race/
├── audio/
│ ├── decoder.ts ← File | URL → AudioBuffer
│ ├── analyzer.worker.ts ← FFT off the main thread
│ ├── beatmap.ts ← features → TrackData (curves, elevation, events)
│ ├── player.ts ← AudioContext + AnalyserNode
│ └── jamendo.ts ← CORS-safe music source
└── game/
├── engine.ts ← audio-clock game loop
└── scene.ts ← Three.js world reads TrackData sample-by-sample
1. FFT belongs in a Worker, and the buffer belongs in a transfer
The naive shape of “beatmap generation” is: read PCM, hop through it 512 samples at a time, run FFT, compute band energies, done. On a 3-minute track at 44.1 kHz that’s ~15,000 FFTs. On the main thread that’s ~800 ms of frozen UI. On a Chromebook it’s five seconds. Nobody stays through five seconds of a stalled loading spinner.
The Worker version isn’t dramatically faster — the same math has to happen — but the main thread stays alive for the loading progress bar to render, the menu backdrop to keep animating, and the audio itself to keep decoding on its own thread. The user perceives ~3 seconds of engaged loading, not ~3 seconds of unresponsive tab.
The instantiation is the one line that has to be right for Next.js:
// beatmap.ts
const worker = new Worker(
new URL("./analyzer.worker.ts", import.meta.url),
{ type: "module" },
);
Not ?worker (that’s Vite). Not new Worker("/analyzer-worker.js") (bundler doesn’t know to emit it). The new URL(..., import.meta.url) form is the one webpack (and Turbopack) recognize as “emit this module as a worker chunk and hand me back the URL”. Every other shape either ships nothing or ships the wrong thing.
The other detail worth calling out is the transfer:
const req: AnalyzeRequest = { type: "analyze", pcm, sampleRate };
worker.postMessage(req, [pcm.buffer]); // ← transfer, don't copy
pcm is an ~8 MB Float32Array for a 3-minute stereo track. Without the second argument, postMessage structured-clones the whole thing into the worker — 8 MB of memory allocation and copying, on the main thread, right where the user is waiting. With [pcm.buffer] in the transfer list the underlying ArrayBuffer is moved to the worker; the main thread’s view of it becomes zero-length. That’s a memcpy you don’t pay for and a GC pressure spike you don’t cause. On phones this is the difference between “smooth” and “hitch.”
The worker is small and single-purpose:
// analyzer.worker.ts (shape)
self.onmessage = (e: MessageEvent<AnalyzeRequest>) => {
const { pcm, sampleRate } = e.data;
const { fft, hop } = createFFT(1024, 256, sampleRate);
const frames = Math.floor((pcm.length - 1024) / hop);
const low = new Float32Array(frames);
const mid = new Float32Array(frames);
const high = new Float32Array(frames);
const flux = new Float32Array(frames);
const rms = new Float32Array(frames);
// ... loop, writing per-frame band energies, streaming
// progress messages every ~5% so the UI can show a bar.
self.postMessage({ type: "result", features: { low, mid, high, flux, rms, ... } });
};
Progress messages are cheap because they’re small ({ type: "progress", ratio: 0.42 }) — no PCM in them. The result message is one big postMessage at the end.
2. The beatmap synthesizer is where “song → track” actually happens
If you plotted low[i] (bass energy) against time for a synthwave track, you’d see something that looks nothing like a driveable road. It’s spiky, it’s noisy, it swings from 0.02 to 0.9 in one 5.8 ms hop. If you fed it directly to a Three.js curve you’d get a track that turns 90° every 20 ms.
The trick is that a driveable road is what you want the audio to feel like on the eyes, not what the audio literally is. The beatmap synthesizer is a series of small deliberate lies.
Curvature — smoothed low/mid balance drives left-right sway:
function buildCurvature(features: AudioFeatures): Float32Array {
const lo = smooth(features.low, 24); // 24-frame moving avg (~140 ms)
const md = smooth(features.mid, 24);
const out = new Float32Array(features.frameCount);
let phase = 0;
let dir = 1;
for (let i = 0; i < features.frameCount; i++) {
const bias = md[i]! - lo[i]!; // mids "pull" the curve
phase += 0.015 + lo[i]! * 0.03; // bass speeds up the sway
if (phase > Math.PI) { phase -= Math.PI; dir = -dir; }
out[i] = Math.max(-1, Math.min(1, dir * Math.sin(phase) + bias * 0.4));
}
return out;
}
The base motion is a sine — the road always sways left, right, left, right. Bass speeds up the sway period (heavier drops → tighter turns), the mid/low balance biases the sine’s midpoint (bass-heavy sections drift one way, mid-heavy sections drift the other). The result: the general shape of the road tracks the general shape of the mix, but you never get a hairpin from a single frame’s energy spike.
Elevation — bass + RMS drive hill climbs:
function buildElevation(features: AudioFeatures): Float32Array {
const lo = smooth(features.low, 32);
const rms = smooth(features.rms, 32);
let phase = 0;
const out = new Float32Array(features.frameCount);
for (let i = 0; i < features.frameCount; i++) {
phase += 0.008 + lo[i]! * 0.015;
const base = Math.sin(phase);
const bias = (rms[i]! - 0.4) * 0.45; // ← the whole trick
out[i] = Math.max(-1, Math.min(1, base + bias));
}
return out;
}
The bias line is the entire game feel. rms > 0.4 (loud sections — drops, choruses) pushes the sine up: the road climbs. rms < 0.4 (breakdowns, intros) pushes it down: the road sinks. So a drop literally lifts you, and a breakdown literally drops you into a valley. Nobody explicitly asks for that pattern — but every player recognizes it, because it’s how music already feels physically.
Hazard placement — flux peaks pick colored cubes and gray spikes:
// beat-aligned event stream, one cube-or-spike per beat
for (let i = 1; i < beats.length; i++) {
const t = beats[i]!;
const f = frameIndexFor(t);
const rms = features.rms[f]!;
if (rms < 0.03) continue; // skip near-silent beats
const color = pickBlockColor(low[f]!, mid[f]!, high[f]!, rms);
const lane = biasedLane(curve[f]! - 0.5); // curve-biased placement
events.push({ type: rms > hazardThreshold ? "hazard" : "pickup",
t, lane, color });
}
pickBlockColor picks red for bass-dominant frames, cyan for treble-dominant, pink for mid, yellow for balanced-and-loud (drops). So the palette of the level tracks the timbre of the song. Bass-heavy tracks give you a sea of red. Sparkly synth pads give you cyan. A drop gives you a burst of yellow. It’s not quite synesthesia, but it’s close enough that players read the color of the incoming cube as what the music is about to do.
hazardThreshold scales inversely with the difficulty slider — peaceful mode never spawns hazards, nightmare mode makes half the beats hostile. The same DSP output, different playback experience.
3. Never trust performance.now() — the audio clock is the game clock
If you write a naive game loop, you use requestAnimationFrame and either its timestamp or performance.now():
// WRONG for audio-driven games
requestAnimationFrame(function tick(t) {
scene.update(t / 1000);
renderer.render();
requestAnimationFrame(tick);
});
Fine for physics. Wrong for rhythm. performance.now() is a monotonically increasing wall clock, but the audio pipeline is not on that clock — it’s on the audio hardware’s clock, which drifts. On some devices the audio card runs at 44.1006 kHz instead of 44.1000. On others the tab throttles to 15 fps in the background, performance.now() skips ahead but audio keeps going, and now the visuals are 200 ms behind the beat. On phones with battery-saver mode kicking in mid-track, the drift is worse.
The fix is dead simple, and it’s the entire premise of AudioPlayer:
// player.ts
get currentTime(): number {
if (!this.playing) return this.pausedAt;
return this.ctx.currentTime - this.startedAt;
}
ctx.currentTime is the AudioContext’s clock, which is the same clock the actual audio playback is running against. There is no drift by construction — it’s not a wall clock trying to guess where the audio is, it’s the audio’s own timeline.
The game engine’s tick uses this and nothing else:
// engine.ts (shape)
const raf = () => {
const t = this.player.currentTime; // audio-locked seconds
const dt = t - this.lastT;
this.lastT = t;
this.hooks.update(dt, t);
this.hooks.render(t, dt);
this.rafId = requestAnimationFrame(raf);
};
Every position in the world — where the next hazard is right now, how far the road has scrolled, whether the beat pulse is at 0.9 or 0.1 — is derived from t = player.currentTime and the pre-computed TrackData arrays. The renderer looks up sampleSeries(this.centerline, hopSec, t) and gets the road’s X offset at exactly the right moment. Pause the audio → currentTime freezes → the world freezes. Resume → they resume together, no re-sync code needed.
There’s a small penalty. AudioContext.currentTime is a double whose precision degrades slightly for long tracks (the granularity is ~5.4 μs at song time = 30 minutes — still 200× finer than a 60 fps frame). And on Firefox, iOS Safari 15, and Chrome under strict privacy settings, the value is quantized to 1 ms as an anti-fingerprinting measure. You still can’t do timing-attack analysis, but 1 ms drift is the ceiling — imperceptible for gameplay, invisible for sync.
The AnalyserNode on the same graph is what feeds the “live pulse” visuals — the sun bloom that flares on a bass hit, the road grid that brightens on treble. Same node, no extra latency:
// player.ts
this.analyser = this.ctx.createAnalyser();
this.analyser.fftSize = 1024;
this.analyser.connect(this.ctx.destination); // audible + analysis in one graph
The scene calls player.getLiveBands() once per frame, gets { low, mid, high, rms } computed from this frame’s audible samples, and folds them into shader uniforms. Because the analyser is on the same graph as the destination, its readings are what the user is hearing, not what a re-decoded parallel stream would produce. This matters when the audio has effects (compression, DRM, the browser’s own Loudness Enhancement) — the analyser sees them, so the visuals stay in step with the actual output.
4. Music sourcing is a CORS problem, not a music problem
The first version of this game shipped with six featured MP3s in public/. That worked, and it was 26 MB of audio in the Docker image. It also meant “Sound Race” was really “Sound Race with these six songs” — and rhythm games live and die on whether you can bring your own music.
The obvious fix is to let the user paste a URL. It doesn’t work. Almost no music CDN sends Access-Control-Allow-Origin: * on their MP3s, and AudioContext.decodeAudioData needs the raw bytes to construct an AudioBuffer. <audio src> streams won’t do — you can play them, but you can’t read the samples for FFT, so no beatmap. No beatmap, no game.
We tried three sources before landing on Jamendo:
- Pixabay music — search page returns 403 to any non-browser fetch; even from a real browser their CDN doesn’t send CORS. Dead.
- Freesound — CORS works, but the catalog is sound effects, not music. Dead for our use.
- iTunes Search API — no key, CORS-friendly, but the previewUrl is a 30-second clip. Not enough for a race.
Jamendo has a public API keyed by a browser-safe client_id (not a secret), a documented tracks endpoint, and — crucially — MP3 URLs served from prod-1.storage.jamendo.com with Access-Control-Allow-Origin: *. Two lines to search:
// jamendo.ts
const params = new URLSearchParams({
client_id: CLIENT_ID,
format: "json",
limit: String(opts.limit ?? 12),
search: opts.query,
audioformat: "mp32",
order: "popularity_total",
});
const res = await fetch(`https://api.jamendo.com/v3.0/tracks/?${params}`, { signal: opts.signal });
And zero lines to fetch the MP3 — decodeAudioSource(track.audioUrl) in the existing pipeline just works. No proxy, no server-side rewrite, no API key on our backend. The client_id is NEXT_PUBLIC_JAMENDO_CLIENT_ID because it’s meant to be public. The whole “music sourcing” feature is one API call plus a <button onClick>.
The one production gotcha isn’t visible in this code — it’s Docker. NEXT_PUBLIC_* values are inlined into the client bundle at build time, not read at runtime. So the docker-compose environment: block reaches the running Node process, but the browser bundle was built without seeing the value. Browser reads "". Search UI hides itself. We spent a genuinely embarrassing amount of debugging time on this before adding ARG NEXT_PUBLIC_JAMENDO_CLIENT_ID to the Dockerfile builder stage and passing it via --build-arg. If you’re reading this with a similar setup: check that the value is in your image at build time, not just at runtime. The value in the running container is not the value your users are reading.
What we didn’t build
On-the-fly BPM change / tempo changes. Real songs modulate. We use web-audio-beat-detector to pick one dominant BPM and lay a rigid grid over the whole track. Songs with a genuine tempo change (metal breakdowns, prog rock time signatures) get a beatmap that goes out of sync in the second half. The right fix is a rolling BPM detector that re-locks every ~30 seconds, and it’s a real DSP project. For 95% of Jamendo tracks — synthwave, electronic, hip-hop, most pop — a fixed BPM is fine.
Server-side transcoding. If a user uploads a .wav, we let the browser decode it. If it’s an obscure .opus variant Chrome doesn’t support, we let it fail. We had a prototype that hit an FFmpeg WASM transcoder as a fallback, and it worked, but it added 25 MB to the client bundle and changed the loading time from “seconds” to “twenty seconds” for anyone who didn’t need it. We killed the fallback and made the error message clearer instead. Most of the time when you cut a feature, the right thing to add is a better error message about the missing feature.
Multiplayer sync. People asked for “race a friend on the same song.” Even a two-player version needs deterministic beatmap generation across two browsers (fine — same PCM in, same beats out) and clock-synchronized playback (very much not fine — network jitter, one player pauses to answer a Slack message, WebRTC’s clock drifts against the audio hardware clock, etc.). The design work for this is bigger than the design work for the whole singleplayer game. We wrote it down and moved on.
Shape of the whole system
If you’re building something like this, this is the topology that ends up working:
user picks a Jamendo track
│
▼
decodeAudioSource(url) ─── streaming fetch + AudioContext.decodeAudioData
│
▼
generateBeatmap(buffer) ┌─── AnalyzerWorker (FFT off main thread)
├── toMono(buffer) → Float32Array PCM ───┤
├── detectBpm(buffer) [main thread] └─── features { low, mid, high, flux, rms, ... }
├── buildCurvature(features) ─────► Float32Array (per-frame X sway)
├── buildElevation(features) ─────► Float32Array (per-frame Y hills)
├── buildIntensity(features) ─────► Float32Array (per-frame speed multiplier)
└── buildEvents(features, beats, difficulty) ───► GameEvent[]
│ (t, lane, type, color)
▼
TrackData
│
▼
new AudioPlayer(buffer) + new GameEngine(player, scene)
│
▼
┌───────────────┴──────────────┐
▼ ▼
player.start() (audio graph) engine.start() (rAF loop, reads player.currentTime)
│ │
▼ ▼
AnalyserNode → live bands scene.update(dt, t) samples TrackData at time t
│ │
└──────────► visuals ◄─────────┘
The whole game is that graph. Everything upstream of AudioPlayer is offline: the DSP runs once, at load time, and produces immutable Float32Array buffers. Everything downstream is online: the game loop reads player.currentTime, samples the pre-baked arrays, and draws. There’s no shared mutable state between them, which is why pause / resume / restart just works — the offline data doesn’t care, and the online loop reads from a clock that pauses cleanly.
The scene itself is standard Three.js — a PerspectiveCamera, a road mesh whose geometry is generated from the curvature+elevation arrays, an Entities3D group that spawns pickup gems and hazards ~600 ms of look-ahead time before their event t. Nothing exotic. All the interesting stuff already happened during the ~2-second generateBeatmap phase.
Coda
The parts of this game that felt hard when we started weren’t. The parts that felt easy — sourcing music, keeping visuals in sync — turned out to be the ones we spent the most time on.
The lesson, if there is one, is that browser audio games are almost never bottlenecked on the audio side. Web Audio is fast, decodeAudioData is fast, AudioContext.currentTime is exactly what you want. The bottlenecks are always at the seams: worker instantiation for the bundler you happen to be using, CORS on someone else’s CDN, a build-arg that didn’t make it into a Docker layer. The DSP is the fun part; the shipping is the hard part.
Live demo: /gamecenter/sound-race — pick a track, hit play, and watch the drop in the second chorus lift you over a hill.