Sound has a position and a clock#
A bird behind the player and a waterfall in front should not sound identical. AudioWorld uses a listener transform and source transforms to calculate distance and direction. It produces commands for a mixer that runs at the audio device’s sample rate, independently of the display frame rate. A SoundAsset stores decoded sample metadata or a streaming decoder. A Voice owns playback position, gain, looping, and source state. Scene components reference voice handles; they do not own the mixer’s buffers directly. Destroying an entity requests a voice stop at an audio-safe boundary.
Move a source around a listener This illustration uses canvas. The explanation below describes the same process.
Listener and source relationships#
The active camera usually supplies listener position and orientation, though cutscenes may choose another listener. A mono spatial source is positioned relative to that listener. Stereo ambience may be mixed as a bed rather than placed like a single point source. Channel layout and headphones affect the spatialization technique. Distance attenuation converts separation into gain using a defined curve and minimum distance. A maximum distance can reduce or virtualize voices that contribute little. Doppler effects need consistent velocity and units; teleporting should reset velocity estimates to avoid a sudden unrealistic pitch shift.
| Relationship | Data | Result |
|---|---|---|
| Source to listener | Relative direction | Spatial panning or HRTF inputs |
| Distance | Meters and attenuation curve | Gain |
| Relative velocity | Smoothed source/listener velocity | Optional Doppler shift |
| Environment | Zones and obstruction queries | Filter, reverb, or occlusion gain |
Timing and updates#
The game thread submits a compact audio command buffer after simulation. The mixer consumes it without waiting on rendering or asset decoding. Parameters can be ramped over sample intervals to avoid clicks. A frame stall should not stop the audio callback from supplying samples. Short effects can be fully decoded. Long music or dialogue can stream through a ring buffer with decoder work on another thread. If streaming cannot keep up, diagnostics record underruns; the callback must still follow the audio backend’s rules rather than block on file access.
Implementation: a voice command
struct VoiceUpdate {
VoiceHandle voice;
float position[3];
float velocity[3];
float gain;
bool stop;
};VoiceHandle is a generation-checked index owned by AudioWorld. The command queue carries values, not borrowed pointers to mutable Scene components. A buffer referenced by the mixer remains alive until an audio-thread acknowledgment retires it. This is a separate completion mechanism from GPU fences.
Budgeting voices#
A voice budget ranks sounds by audibility and gameplay importance. Virtualized voices may advance their playback clock without expensive mixing, depending on policy. When a voice becomes audible again, it resumes at the appropriate position. A looping engine noise and a short UI notification can require different policies. Reverb and obstruction are approximations with costs. Performing a physics ray test for every voice every frame may be unnecessary; spread updates and reuse stable results. Smooth filter changes so a moving source does not produce abrupt timbral jumps.
Diagnose silence and glitches#
Check the output device, voice state, gain, channel format, and asset readiness before debugging spatial math. A source can be inaudible because the listener is wrong or the attenuation curve assumes different units. Clicking suggests discontinuous sample values or parameter changes. Repeated bursts of silence suggest decoder starvation or a callback doing blocking work. Follow asset ownership and frame lifecycle to see where sound updates enter the simulation. Audio playback is a native engine responsibility; the website illustrates relationships without claiming to run the mixer.