WAV or MP3? 44.1 or 48 kHz? What's a LUFS target? Here's a plain-English guide to the delivery specs that decide whether your voice over drops straight in — or bounces back for a re-export.
In this article
A great performance can still cause headaches if it arrives in the wrong format. Editors lose time re-exporting, broadcasters reject audio that misses their loudness spec, and telephone systems mangle files recorded at the wrong sample rate. The good news: the specs that matter are few, and once you know them you can brief any studio with confidence.
WAV vs MP3: lossless vs lossy
The single most important choice is compressed or uncompressed.
- WAV (or AIFF) is uncompressed and lossless. It holds the full quality of the recording, which is exactly what you want for editing, mixing and broadcast. It is the professional delivery standard.
- MP3 is compressed and lossy — it throws away data to shrink the file. That is perfect for emailing a preview or a reference, but every time a lossy file is re-exported it degrades a little further. Never master a project from an MP3 if you can help it.
Our rule of thumb: request WAV for the master, and MP3s alongside only if you need lightweight files for review or web.
Sample rate and bit depth
Two numbers describe the resolution of a recording. Sample rate (measured in kHz) is how many times per second the audio is captured; bit depth is how much detail each sample holds.
| Use case | Recommended spec |
|---|---|
| Video, broadcast, most pro work | 48 kHz / 24-bit WAV |
| Audio-only, music, podcasts | 44.1 kHz / 16 or 24-bit |
| IVR / telephone systems | 8 kHz (often µ-law or A-law) |
When in doubt, 48 kHz / 24-bit is the safe default — it comfortably covers video and broadcast, and can be converted down for anything that needs a lower spec. Telephony is the notable exception, and getting it wrong there means crackly, distorted prompts; see our IVR & telephone voice over page for how we handle it.
Loudness: what LUFS means for you
A recording can be technically perfect and still be rejected for being too loud or too quiet relative to everything around it. Modern delivery is measured in LUFS — a standard for perceived loudness — and each platform sets a target:
- Broadcast (EBU R128 / ATSC A/85) — around −23 LUFS in Europe, −24 LKFS in the US.
- Streaming & online video — commonly around −14 LUFS.
- Podcasts — typically −16 to −14 LUFS.
You do not need to become an audio engineer — you just need to tell your studio where the audio will run. We master to the correct target so your voice over sits at a consistent, compliant level next to music, other ads or programme audio.
One mix or separate stems?
A mixed file is finished and ready to use. Stems keep the voice, music and effects on separate tracks so your editor can re-balance later. For single-language ads a mix is usually enough; for multi-language work we typically deliver stems so the mix can be rebuilt cleanly in each language — a small step that saves a lot of rework. This is closely tied to voice over translation and dubbing projects.
File naming and organisation
On big projects, naming is not an afterthought — it is the difference between a smooth import and hours of manual sorting. Hundreds of eLearning clips, game lines or IVR prompts need names that map exactly to your script or asset list. Send us your convention, or we will propose one, and every file arrives labelled and ordered so it drops straight into your system.
Brief it once, get it right
To deliver perfectly the first time, a studio needs four things: the format (WAV/MP3), the sample rate and bit depth, the loudness target, and whether you want a mix or stems. If you are not sure, just tell us where the audio will run and we will recommend the spec. Explore our voice over services, or send us your brief and we will handle the technical detail from there.

