Binaural Recording Basics

Imagine you are standing in a busy city park with your eyes closed. You can pinpoint exactly where a bird chirps or where a distant car engine hums. This ability to map sound in space relies on how your ears capture vibrations from the world. We use this natural biological process to create immersive audio experiences that mimic real life. By understanding how our anatomy shapes incoming sound waves, we can simulate a three-dimensional space inside a pair of simple headphones.
The Mechanics of Sound Localization
Your brain determines the location of a sound source by analyzing tiny differences between your two ears. When a sound arrives from the right side, it reaches your right ear slightly before it hits the left ear. This brief delay is known as an interaural time difference, which provides your brain with a primary cue for direction. Additionally, your head acts as a physical barrier that blocks high-frequency sounds from reaching the far ear. This creates an interaural level difference where the sound is louder in the closer ear. These two cues work together to build a mental map of your surroundings.
Key term: Head-related transfer function — a mathematical model describing how an individual’s physical anatomy filters and modifies sound waves before they reach the inner ear.
Beyond these basic timing and volume cues, your outer ears play a vital role in vertical sound localization. The complex folds of your ears, called the pinnae, reflect sound waves in unique ways depending on the angle of arrival. These reflections create tiny notches in the frequency spectrum that your brain interprets as height information. If you were to wear a mold that changed the shape of your ears, your brain would struggle to tell if a sound came from above or below. This shows that your anatomy is an essential part of your hearing system.
Applying HRTF Principles to Mixing
To recreate this effect for a listener, we use a binaural recording technique that captures sound using two microphones placed inside a dummy head. This dummy head is shaped like a human skull, complete with realistic ear canals and pinnae. By recording through these structures, the audio captures the same frequency changes and time delays that your own ears would naturally experience. When a listener plays this recording back through headphones, their brain receives the same spatial cues as if they were present in the room. It is like an acoustic photograph that preserves the exact spatial geometry of the original performance.
We can also simulate these effects digitally using a process called spatial audio processing. Instead of using dummy heads, sound engineers apply a digital filter to standard audio tracks to mimic the natural filtering of the human head. This process requires a complex mathematical formula that accounts for the shape of the head and ears. The following table highlights how different components of your anatomy contribute to spatial perception:
| Anatomical Feature | Primary Function | Spatial Cue Provided |
|---|---|---|
| Pinnae | Reflects high frequencies | Vertical localization |
| Head Mass | Blocks high frequencies | Level differences |
| Ear Distance | Creates arrival delays | Time differences |
This technology allows creators to place sound objects anywhere in a virtual sphere around the listener. You can make a whisper feel as if it is happening right behind the listener’s neck. This level of precision is impossible with traditional stereo mixing, which only pans sound left or right. By mastering these principles, you can transform a flat recording into a vibrant, three-dimensional soundscape that feels completely real.
Human spatial hearing relies on the physical shape of the head and ears to filter sound, which we can replicate through digital processing or specialized recording techniques.
The next Station introduces Ambisonics Explained, which determines how spatial audio fields are encoded for various playback systems.