Skip to main content

Podcast Audio Editing: The 5 Stages of a Professional Pass

An audio editor in headphones working on a podcast episode, with two speech waveforms on a widescreen monitor and one hand on a fader control surface

Podcast audio editing covers cleanup, levels, content cuts, structure, and a final listen. Here is what a professional pass includes and how to judge one.

Written by The HiveCast Team Published Last updated

When you pay someone for podcast audio editing, most of what you are buying is a person’s attention. Noise reduction and leveling are the parts hosts picture, and software now does most of that work. The part that costs money is the listening: someone plays the episode from the first second to the last and decides what stays.

That difference shows up in what comes back to you. A shop working at volume runs your file through noise reduction, pulls the levels toward a target, trims the longest silences, and sends it back. Nobody there has heard your episode, so the problems that actually lose listeners survive the edit:

  • A story you already told two episodes ago.
  • A guest’s name pronounced two different ways.
  • An answer that starts before the question finishes.
  • A level that drifts twenty minutes in when someone leans back in their chair.

We have produced more than 5,000 episodes, and the full listen-through is the last step we would give up. Below is what a professional podcast audio editing pass contains, in the order it happens, so you can price the work and judge the person doing it.

What podcast audio editing covers, in order

Each stage depends on the one before it being finished, so doing them out of sequence means redoing work.

Numbered flowchart of the five stages of a professional podcast audio editing pass: 1 cleanup, mostly software, removing broadband noise, room tone, plosives and desk bumps with background sound kept at least 20 dB below the speech; 2 levels, mostly software, balancing speakers and episodes to a podcast target of around -16 LUFS; 3 content edits, where a person listens to the full episode and cuts tangents, restarts and repeated stories; 4 structure, the repeatable assembly of cold open, intro, music bed, transitions, ad slot and outro; and 5 final QA, where a person listens to the rendered file end to end

Cleanup

Cleanup is the mechanical pass, and it comes first because every later decision depends on hearing the recording clearly. Four kinds of problem get handled here:

  • Broadband noise, meaning air-conditioning hum, computer fans, and traffic behind a window.
  • Room tone, the ambient sound of the room you recorded in, which has to be matched wherever a cut joins two pieces of audio.
  • Plosives, the thumps on hard P and B sounds that arrive when a mouth is too close to the microphone.
  • Incidental noise: mouth clicks, coughs, chair creaks, and the desk bumps a host stops noticing by the third episode.

Two things are worth knowing about this stage:

  • Cleanup cannot fix the source recording. If the mic sat across the room, the reverb is baked into the file and no plugin takes it back out, which is why the setup advice for remote recordings is worth more than any plugin you could buy afterwards.
  • Noise reduction has a cost. Pushed too hard, voices go thin and watery, which listeners hear as cheap even when they cannot name what is wrong, so a good editor removes less noise than the software can. The accessibility standards put a number on how much is enough: the W3C’s guidance on background audio asks for background sound at least 20 decibels below the speech, which is roughly four times quieter.

Levels

Levels is the stage most people mean when they say an episode sounds professional, and it solves three separate problems:

  1. Balance between speakers. You and your guest almost never arrive at the same volume, and a listener should not have to reach for the volume control when the conversation switches sides.
  2. Consistency inside the episode. People lean in when they get excited and lean back when they are thinking, and the level moves with them.
  3. Consistency across episodes. A new episode should land at the same loudness as one from last year.

Loudness targets

Loudness is measured in LUFS, the unit broadcasters standardized on in the EBU R 128 recommendation.

Where it playsLoudness target
Broadcast television and radio-23 LUFS
Podcasting, stereo episodesAround -16 LUFS

The podcast number is a convention rather than a rule, and platforms apply their own normalization on top of it, but it exists so your show sits at about the same volume as everything else in a listener’s queue. Hitting the same target every week matters more than the target itself.

Dynamics

Leveling also covers dynamics. Compression and limiting reduce the gap between the loudest and quietest moments, which is what lets someone hear you clearly in a car without the peaks turning harsh.

Content edits

This is the full listen-through, and it separates editing from processing. A person plays the episode start to finish and makes calls no automated tool can make. What comes out at this stage:

  • Tangents that go somewhere and never come back. Ten minutes of pleasant conversation that does not serve the episode still costs you listeners.
  • Restarts and do-overs, where you stopped, said “let me say that again”, and delivered the better version. Only the better version ships.
  • Repeated stories, since hosts reuse their best anecdotes and the listener who has heard one twice notices before the host does.
  • Answers that began before the question finished, where the overlap needs tightening so the exchange reads as a conversation.
  • Dead ends: a question that landed badly, a name nobody could remember, half a minute of two people each waiting for the other to speak.

Cuts also have to be inaudible. A hard seam, an abrupt change in room tone, or a word clipped at the front tells the listener they are hearing an edited file. Crossfades, matched room tone, and cutting at a natural breath keep the change from registering.

Structure

Structure is assembly. The edited conversation gets dropped into the shape your show uses every week:

  • The cold open, if you run one
  • The intro, and the music bed fading under your first words
  • Transitions between segments
  • The house ad, in its pre-roll, mid-roll, or post-roll slot
  • The outro

Most of this is repeatable, which is the point. A listener who knows your show recognizes where they are inside it, and a consistent episode structure is what makes an hour of talk feel organized rather than long. The part worth revisiting is the opening, since that is where people decide whether to stay, and our intro guidance covers what belongs in it.

This stage also produces the assets that ship with the episode:

  • The title and description your editor writes feed straight into search, which is why podcast SEO work belongs in the same pass rather than bolted on weeks later.
  • The finished master is what any podcast clips get cut from, so the levelling done here saves the same work being redone per clip.

Final QA

The last step is listening to the rendered file end to end, after every edit and every level change is in place. It catches the problems that only exist in the finished version:

  • An intro that starts on top of your first word.
  • A music bed that never faded.
  • An ad read sitting louder than the conversation.
  • A cut that clicks on export.
  • An episode that renders short because a section got dropped in assembly.

Most audible mistakes in a published episode come from skipping this step. It is also the easiest one to skip, since by then every edit has been made and the file already sounds finished.

What filler-word removal is actually worth

Most hosts assume editing is mainly about removing “um”, and it is the thing they ask about first, but filler removal is close to the least valuable thing an editor does.

Two reasons it matters less than you would expect:

  • Listeners do not hear individual filler words in a conversation they are following. What they hear is hesitation when a speaker is struggling, which is a pacing problem rather than a word problem.
  • Removing every instance makes people sound wrong. A track with every “um”, “you know”, and “sort of” taken out sounds clipped and robotic.

What we do instead is selective. Filler comes out where it stacks up, where three of them sit in front of one sentence, or where a long “uhhhh” stalls an otherwise good answer. The rest stays, because it is how people talk. If you want the heavier pass, with filler words, long pauses, and stumbles removed at the producer’s discretion and no timestamps required from you, that is an enhanced editing add-on on top of either plan.

Why remote recordings are harder

Most of the shows we produce are recorded remotely, and a remote episode arrives with problems a single-room recording does not have: two rooms with different acoustics, two microphones with different characters, and two internet connections that can each fail. One side might be on a good dynamic mic in a carpeted office while the other is on a laptop in a kitchen.

That changes the work in three specific ways:

  • Room tone has to be handled per speaker rather than per file, because the two sides never match.
  • Leveling gets harder when one side is clipping and the other sits too quiet.
  • Dropouts and digital artifacts have to be found and either repaired or cut around, and a repair is only possible when there is clean audio nearby to work from.

Almost all of it is cheaper to prevent than to fix. Local recording on each side, headphones on both ends, and separate tracks per speaker give an editor something to work with, and they cost nothing but a minute of setup.

What professional editing costs, and why per-minute pricing misleads

Podcast audio editing is sold one of two ways, and the pricing model tells you something real about the work you will get:

  • By the minute of finished audio. You pay per delivered minute, so the incentive is to move faster.
  • As part of a monthly production subscription. You pay for a slot in a team’s schedule, so the incentive is to keep your show sounding right month after month.

Where per-minute pricing cuts corners

Per-minute pricing rewards speed. The margin comes from handling more minutes per hour of labour, which makes the full listen-through the first thing to go, because it is the one step that cannot run faster than real time. An hour-long episode takes at least an hour to hear, while everything else in the workflow can be batched or automated.

What a flat monthly plan covers

Our own pricing is a flat monthly subscription with editing included:

PlanCadencePriceVideo add-on
Bi-weeklyAn episode every other week$500/month+$100/month
WeeklyAn episode every week$875/month+$200/month

Both plans cover the same work:

  • The full listen-through and the final QA listen
  • SEO show notes
  • Two promotional graphics per episode
  • Publishing to the podcast directories
  • A five business day turnaround

Video and YouTube production sits outside audio production as a paid add-on, at the rates in the last column above.

A subscription also suits editing because the work improves as the editor learns your show. Someone who has cut your episodes for months knows which stories you repeat and how you sound just before you restart a sentence.

How to tell if your editor is doing this

You can check most of this yourself, on your own last episode.

  • The full listen. Play your most recent episode start to finish at normal volume, in a car or on headphones. If you hit a moment that makes you wince, your editor either did not hear it or left it in.
  • The seams. Find three or four cuts you know about and listen closely. A clean edit sits on a breath and keeps the background even on both sides of the join; a rough one clips a word or shifts the room tone.
  • Volume across episodes. Queue your newest episode after one from last year and do not touch the volume control. A noticeable jump means nobody is mastering to a consistent target.
  • The guest side. On a remote interview, check whether both voices sit at the same level and seem to occupy the same space. Untreated guest audio next to clean host audio is the most common sign of a fast edit.
  • What your editor asks you. An editor doing the content pass comes back with notes: a tangent they cut, a repeated story they flagged, a name they need spelled. An editor who never sends notes probably did not do a content pass.

None of this is visible to a listener when it is done properly, which is part of why it gets under-bought. If you would rather hand the whole pass to a team that does it every week, our podcast production service covers podcast audio editing, show notes, graphics, and publishing, and the production library here has more on the recording side. You can book a call and we will listen to your latest episode before we talk, so the conversation is about your show.

Frequently Asked Questions

What does podcast audio editing include?

A professional pass has five stages: cleanup (noise, room tone, plosives, clicks, coughs), leveling across speakers and across episodes, content edits made during a full listen-through, structural assembly with intro, outro, music, and ad slots, and a final QA listen to the rendered file. Cleanup and leveling are largely technical and partly automated. The content edits and the final listen are where a person makes judgment calls, and they are the stages that cheaper arrangements leave out.

How long does it take to edit a podcast episode?

It always takes longer than the episode itself. A full listen-through runs the length of the recording on its own, and the final QA listen runs it again, with cleanup, leveling, assembly, and show notes on top of that. The commitment we make to clients is turnaround rather than hours: five business days from the day we receive your audio.

What is the difference between editing and mastering?

Editing decides what the episode contains: what gets cut, what order it runs in, and where the intro, transitions, and ad slots sit. Mastering is the final technical treatment of the assembled file, mainly loudness and dynamics, so the episode lands at a consistent level and holds up on phone speakers, car stereos, and headphones. You need both, because mastering a badly edited episode only gives you a consistent version of the wrong cut.

How loud should a podcast be?

Most stereo podcast episodes are mastered to around -16 LUFS, the loudness convention the industry settled on so that shows sit at a similar volume in a listener's queue. Platforms handle loudness their own way, so treat it as a target rather than a rule. Consistency matters more than the number: every episode of your show should land at the same level as the last one.

Should I remove every “um”?

No. Listeners do not notice individual filler words in a conversation they are following, and stripping all of them out makes speech sound clipped and robotic. The useful version is selective: cut filler where it stacks up in front of a sentence or stalls a good answer, and leave the rest. If you want the heavier pass without sending timestamps, that is an enhanced editing add-on rather than the standard edit.

Is editing included in HiveCast's plans?

Yes. Professional editing, including the full listen-through and final QA, is part of both plans: $500 a month for an episode every other week, and $875 a month for a weekly episode. Both plans also include SEO show notes, two promotional graphics per episode, and scheduling and publishing to the podcast directories. Video and YouTube production is a paid add-on and is not part of audio production.

podcast editing audio editing podcast production
Share

Want this handled for you?

You record. HiveCast handles the editing, show notes, graphics, and scheduling. 5,000+ episodes produced, no long-term contracts.

More on Podcast Production

Related Articles