MP3, WAV, or OGG for two-second clips
By TrendyMemez Editorial · · 7 min read
At this length the usual format advice inverts. What actually matters is the encoder delay nobody warns you about.
The standard advice about audio formats is written for songs and podcasts — things that run for minutes. At two seconds, most of that advice stops applying, and one specific problem that is irrelevant for a song becomes the dominant issue.
The short version
- MP3 — plays everywhere without exception. Has a gap problem, described below. Correct default for distribution.
- WAV — uncompressed, no gap problem, sample-accurate. Correct choice while you are editing. Large, but at two seconds “large” means a few hundred kilobytes.
- OGG / Opus — better quality per byte than MP3 and no gap problem. Not universally supported, particularly by older hardware and some editing software.
- M4A / AAC — good quality, wide support on Apple platforms, has its own version of the gap problem.
The gap nobody warns you about
MP3 encoding works on fixed-size blocks of samples. If your audio does not divide evenly into those blocks, the encoder pads it — and it pads the beginning as well as the end. Every MP3 therefore starts with a short stretch of silence that was not in your original file — the same encoder-delay behaviour documented in Hydrogenaudio’s gapless playback reference.
The amount is small, typically a few milliseconds. For a four-minute song this is completely irrelevant. For a sound effect you are trying to land precisely on a video cut, it is the difference between tight and slightly late — and it is consistent, so it does not average out.
Decoders can compensate if the file carries the right metadata, but whether they do depends entirely on the software. Some editors honour it, some ignore it, and you generally cannot tell which without testing.
Bitrate matters less than the source
A common instinct is to re-encode everything at the highest available bitrate. This does not recover anything. Encoding is lossy and one-way — a 128 kbps file re-encoded at 320 kbps is a 128 kbps file in a larger container, with an extra generation of loss on top.
Two practical consequences:
- Keep the highest-quality version you have as your working copy and export from that.
- Avoid round-tripping. Every import-edit-export cycle through a lossy format costs something, and on short percussive clips the damage shows up as smearing on the transient — precisely the part that makes the sound work.
At very short lengths the quality difference between a well-encoded 192 kbps MP3 and a WAV is genuinely hard to hear on the kind of playback these clips get. The transient integrity matters more than the bitrate number.
Level, not format, is what usually sounds wrong
When people say a downloaded clip “sounds bad”, the cause is far more often level than format. Clips sourced from different places arrive at wildly different loudness, and a clip that was already limited to within an inch of its life will sound harsh no matter what container it is in.
If a sound is clipping in your edit, turning it down in the editor does not fix it — the distortion is already baked into the samples. You need a different source or a limiter before the damage, not after.
What to actually use
- Downloading to use in an edit: MP3 is fine. Trim the leading silence on import.
- Building a soundboard: check what your software wants. Discord’s built-in soundboard has a tight size limit, which pushes you toward MP3 or OGG over WAV — the soundboard guide covers the specifics.
- Producing something you will re-edit repeatedly: convert to WAV once, work in WAV, export to a lossy format only at the end.
- Serving audio on a website: MP3 for compatibility. Opus if you control the playback environment and care about bandwidth, because at these lengths it is substantially smaller for the same perceived quality.
Everything in the sound effects library here is served as MP3 for the compatibility reason — it is the only format with genuinely no exceptions.