Those piercing S sounds that jab at your ears on a voice recording have a name, sibilance, and a proper fix that does not flatten the rest of the voice. The harshness lives in a narrow band of high frequencies, and the tool built for it, a de-esser, clamps down on exactly that band only when the harshness flares up. It leaves the brightness of the voice intact the rest of the time, which is precisely why it beats reaching for a blunt EQ cut.
Quick Answer
Use a de-esser set to the 5kHz to 8kHz band, where sibilant energy concentrates. It only attenuates when a harsh S peak hits, so it tames the hiss without dulling the voice. Aim for 3dB to 6dB of reduction, and never cut more than 6dB or the S vanishes and leaves a hole.
Where Sibilance Lives and Why It Grates
Sibilance is essentially a volume spike in the high frequencies, typically between 5kHz and 8kHz, produced by S, SH and similar sounds. Human hearing is naturally sensitive up there, so even a modest spike reads as harsh and tiring. Recorded through a microphone, those peaks often get exaggerated, and a listener subjected to constant sharp S sounds quietly fatigues without knowing why. The goal is not to remove the S entirely. It is to take the edge off the worst peaks so speech sounds natural.
Why a De-Esser Beats a Static EQ Cut
You could carve out 6kHz with an equaliser and the harshness would drop, but so would every bit of brightness in the voice, even on words with no S at all. That leaves the recording dull and lifeless. A de-esser is dynamic. It watches the targeted band and only reduces it when a sibilant peak crosses your threshold, then releases the moment the S passes. So the voice keeps its natural air and clarity through the rest of the sentence, and only the offending spikes get pulled down. That selectivity is the whole point.
Setting It Up Step by Step
Start by finding the frequency. For most male voices the harsh energy sits in the 5kHz to 8kHz region. For female voices or very bright recordings it can run higher, into the 7kHz to 12kHz range. Many de-essers have a listen or solo mode that lets you hear only the band being affected, which makes finding the exact spot quick.
Next, set the threshold. Begin somewhere moderate, around -15dB, and lower it gradually until the de-esser engages on the harsh S sounds but ignores normal speech. You want it reacting to the spikes, not riding the whole vocal. Keep the timing fast, roughly a 1ms to 2ms attack and a 10ms to 20ms release, so it catches the quick sharp peaks and lets go cleanly.
Finally, set the amount. A reduction of 3dB to 6dB on the offending peaks is usually plenty. Resist the urge to crush harder. Cut beyond about 6dB and the S disappears completely, which sounds just as wrong as the harshness did, like a lisp punched into the audio. Subtle is the target.
A note on the source
A de-esser cleans up what you already have, but quieter sibilance starts at capture. Angling the mic slightly off to the side of your mouth and not crowding it reduces the harsh peaks before they are ever recorded, leaving the de-esser far less to do. Monitoring on a decent pair of closed-back cans from the headphones and headsets range also lets you actually hear the sibilance you are trying to fix, which cheap earbuds will hide from you. The headset best sellers at Evetech are a practical reference if you want to see which models other creators rely on.
Frequently Asked Questions
What frequency should I set a de-esser to?
Start in the 5kHz to 8kHz band for most voices, where sibilant energy concentrates. Bright or female voices may need a higher target, around 7kHz to 12kHz. Use the de-esser's listen mode to pinpoint the exact harsh frequency for your recording.
Why not just cut the highs with an EQ?
A static EQ cut dulls the entire voice, including words with no harsh S, leaving it lifeless. A de-esser only reduces the targeted band when a sibilant peak appears, preserving the voice's natural brightness everywhere else. That dynamic behaviour is its advantage.
How much reduction should I apply?
Around 3dB to 6dB on the harsh peaks is usually right. Cutting more than 6dB tends to remove the S entirely, which sounds as unnatural as the original harshness. Aim for taming, not erasing.
Can I prevent sibilance while recording instead?
Yes, partly. Angling the microphone slightly off-axis from your mouth and keeping a little distance reduces the harsh peaks at the source, so the de-esser has less work to do afterwards. Good capture and good processing work together.