SoundForgePro Official 15-day trial

Workflows

Sound Forge Pro for Voice-Over Recording

In this guideSections
    A voice-over booth microphone and the resulting take open in Sound Forge

    Sound Forge Pro 2026 is a strong voice-over editor when the job centers on one narration file: record, punch in a correction, clean the take, measure it, and export. It can also record multiple input channels, but it is not a multitrack DAW with separate timeline tracks and a mixer. That distinction matters more than any generic “professional” label. For a broader orientation, use the Sound Forge Pro 2026 guide; for other task-specific setups, open the workflows hub.

    Quick answer: create a mono 24-bit file, choose the interface and input under Options → Preferences → Audio, open View → Record Options, choose Manual, test the loudest line with safe headroom, Arm, and press Ctrl+R. Edit on a copy, process only what the recording needs, then verify the actual delivery specification before export.

    A voice-over take in Sound Forge with a punched-in retake and the delivery file measured
    Voice-over in Sound Forge: arm the input, punch in a precise retake, then measure the exported delivery file.

    Where Sound Forge Fits in a Voice-Over Workflow

    For one narrator, Sound Forge keeps the work close to the file: the waveform, Record Options, regions, plug-in chain, Statistics, and export are available without building a DAW session. That is useful for commercials, e-learning, podcast inserts, auditions, and audiobook chapters.

    Don’t confuse multichannel recording with multitrack production. The current multichannel recording help confirms that Sound Forge can capture several hardware inputs into a multichannel data window. It also states that Sound Forge is not a multitrack editor. Use a DAW when each speaker needs an independent timeline track, routing, automation, or later mix revisions.

    For long-form narration, work chapter by chapter and keep an untouched source plus a processed copy. That avoids turning one large destructive edit history into a single point of failure and matches ACX's requirement for one chapter or section per delivery file.

    Audio Device Setup: ASIO Driver and Input Configuration

    Before recording anything, set up the audio device correctly in Sound Forge Pro. The driver choice affects latency during monitoring (the gap between speaking and hearing yourself in your headphones), and that matters for VO because performers need to monitor naturally without delay.

    Go to Options → Preferences → Audio. Under Audio Device Type, choose the manufacturer's ASIO driver when the interface provides one; the 2026 help describes ASIO as the low-latency option. Then open the Record tab and map the hardware input to the file channel. Don’t publish a latency number unless it was measured on the named driver, buffer, sample rate, and device.

    Set the buffer low enough for comfortable monitoring and raise it only if you hear gaps or glitches. Direct hardware monitoring, when the interface offers it, avoids a software round trip. The site methodology doesn't document a Focusrite test bench, so the earlier ownership and 4 ms/28 ms comparison has been removed.

    If no native ASIO driver exists, test the Windows driver first. A generic wrapper can help some devices, but it adds another configuration layer and doesn't guarantee lower latency or lower noise. Confirm the input in the Record tab, click Apply, Arm the device, and check for a moving meter before recording the take.

    On the Audio preference page, open the Record tab and assign the interface input to the data-window channel. For one microphone, create a mono file unless the client explicitly requires stereo. A stereo file isn't automatically two identical channels; the result depends on routing, so verify the meter and a short test recording instead of assuming.

    Recording Levels: Where to Set Your Gain

    Sound Forge's current recording help says to capture a strong signal without clipping; its digital-level help suggests 3–6 dB of headroom only when the loudest section is unknown. For speech with unpredictable emphasis, peaks around -12 to -6 dBFS are a conservative editorial starting range, not a platform rule.

    That range gives processing headroom: the EQ, compression, and de-esser you apply later need room to work without clipping the output. A recording peaking at -3 dBFS has almost none left; adding 2 dB of presence boost in the EQ clips the output before you can do anything else. A recording peaking at -8 dBFS has room for all of that plus normalization at the end.

    Set the gain on your audio interface hardware rather than in software. Sound Forge's input meter shows what's arriving, but the control that matters is the interface's physical gain knob before the signal reaches the software. Turn it up until loud passages peak around -6 dBFS in Sound Forge's recording meters, then back off a notch. Check for clipping (flat-topped peaks in the meter or red indicators) and reduce gain until it disappears.

    Record the loudest line and a quiet line before the real take. Listen for room noise, reflections, plosives, cable faults, and clipping. A clean test file is more useful than chasing a fixed meter number; once samples clip at 0 dBFS, lowering the file later does not restore them.

    The Record Dialog: Key Settings for Voice-Over

    Open View → Record Options. The current Sound Forge Pro 2026 Record Options help documents the voice-over path directly: Manual is the general-purpose and punch-in method, while Normal, Create regions, and Create new windows control what happens after a stop or restart.

    Choose Manual for normal narration and punch-ins. Automatic: Threshold is for sound-triggered capture; Automatic: Time and MIDI Timecode serve scheduled or synchronized jobs. Manual gives the narrator deliberate control and is the method the current help names for voice-over.

    For a punch-in, use enough pre-roll to hear the previous phrase and match pace and tone. Configure pre-roll, post-roll, and the prerecord buffer from the Record Options Settings control. The right duration depends on the sentence and performer; 2–5 seconds is a practical starting range rather than a requirement.

    A prerecord buffer can retain audio from just before recording begins, but it's no substitute for confirming that the device is armed. Set it in Record Options Settings, test it on a disposable file, and verify exactly what your build retains before trusting it during a paid session.

    Normal leaves one continuous recording sequence. Create regions adds a new region whenever recording is restarted or resumed; it doesn't create separate non-destructive takes on independent tracks. Create new windows makes a separate window after each restart but disables punch-in recording.

    Keep the Script Visible Without Relying on an Old Menu Path

    Older tutorials call the floating control OTR or Remote Recording. The current 2026 help index still mentions Remote Recording, but the primary documented workflow is View → Record Options plus the main Arm, Record, Pause, and Stop controls. Don’t promise a View → OTR command without verifying the installed build.

    Use a second monitor, Windows Snap, or a narrow script window beside Sound Forge. Keep the input meter visible and test keyboard focus before the take; a shortcut sent to the PDF or browser instead of Sound Forge can lose a recording cue.

    The goal is simple: the script remains readable while record status and peaks remain visible. Use the layout your hardware actually supports and save it as a Sound Forge workspace only after a short test take confirms focus, monitoring, and input routing.

    Marking Mistakes While Recording: The M Key

    During recording, press M to drop a marker at the current recording position without stopping the take. Use it to flag a stumble, mispronunciation, cough, or section that needs a retake, then confirm the marker after recording.

    The workflow: record a full pass without stopping. Every time you stumble, mis-read, or cough, press M. Don't stop and restart; keep going. At the end of the session you have a complete recording with markers at every problem point. Navigate between markers using the Regions List (View → Regions List) and zoom in to each marked section for the retake.

    This keeps the performance moving and gives the editor a visible correction map. It doesn't remove the retake work: each marker still needs a precise edit, a matched replacement, and a listen across both boundaries.

    Before relying on M during a long session, record a short test, press M, stop, and confirm that the marker is stored where expected in the current build. If it is not, note the time display or pause and create a region. The earlier 4,200-word/14-marker/47-minute claim had no session log and has been removed.

    Punch-In Recording for Retakes

    Punch-in records over a specific section of an existing recording without affecting the rest of the file. In SF Pro, place the cursor at the start of the section to replace, or select the region to overwrite.

    Set Method to Manual. Select the exact range when the correction must stop at a defined point; the 2026 help states that recording stops automatically at the end of a selection. Without a selection, recording overwrites from the cursor until you stop. Keep the original file so every punch-in is recoverable after the session closes.

    Match microphone distance, angle, gain, room, and performance before the punch. Listen across both edit boundaries before processing. Don’t use noise reduction to hide a mismatched retake by default; re-recording under the same conditions is usually cleaner than forcing two different rooms to match.

    EQ for Voice-Over: The Specific Moves

    Start with the current Modern Equalizer as a VST effect; keep Paragraphic EQ only for a legacy Process → EQ workflow. Fix the recording and microphone position before EQ. Then make the smallest change that improves intelligibility at a matched output level.

    Use a high-pass filter only when there is unwanted low-frequency energy. Move the cutoff upward while previewing and stop before the voice loses body. There's no universal 80, 100, or 120 Hz setting, and microphone type alone doesn't determine the correct cutoff.

    If the low end is already clean, leave it alone. A filter applied to every file by habit can thin voices, remove useful chest resonance, and make chapter-to-chapter consistency harder to maintain.

    For boxiness or resonance, listen first, then use Spectrum Analysis and a temporary narrow sweep to confirm the frequency. Cut gently and widen the band until the correction sounds natural. A graph peak isn't automatically a defect, and 200–400 Hz isn't a preset for every room or voice.

    A broad presence boost can improve intelligibility, but it can also exaggerate sibilance, mouth noise, and fatigue. Compare on headphones, small speakers, and the client's reference. Skip the boost when articulation is already clear.

    An air shelf is optional. Use it only after de-essing and noise checks because it can lift hiss and mouth noise with the voice. The earlier unnamed-client RE20 recipe was not auditable and has been removed; the EQ guide explains current controls without treating another session's curve as a preset.

    De-essing: Controlling Sibilance

    Sibilance is level-dependent and changes with voice, microphone angle, distance, and room. The current core help does not document a dedicated native de-esser. Load a compatible VST2/VST3 de-esser through the Plug-In Chainer, or use automation for a few isolated syllables.

    A static EQ cut reduces that band for the entire file, including non-sibilant words, so use it only for mild, consistent harshness. For changing sibilance, a de-esser or gain automation is more selective. Level-match the comparison so a quieter result isn't mistaken for a better one.

    Mouth De-click and de-essing solve different problems: one targets mouth transients, the other reduces sibilant energy. Old MAGIX bundles and current Boris FX packages aren't interchangeable, so confirm the plug-in license you actually own, then test it through the Plug-In Chainer before building a repeatable preset.

    Set the detector on the actual narrator, audition the difference signal if the plug-in offers it, and listen for softened consonants. Save a preset only as a starting point; microphone position and performance can change enough between sessions to require a new threshold.

    Editing Breaths and Room Tone

    Breath sounds between sentences divide opinion in VO production. Some clients want all breaths removed; others want natural breathing kept because breath removal makes recordings sound clinical. Know your client's preference before starting the editing pass.

    Reduce only distracting breaths. Select conservatively so the next consonant and the phrase timing remain intact. Lowering gain often sounds more natural than deletion. When a replacement is necessary, use room tone from the same session rather than digital zero and cross-check both edit boundaries on headphones.

    Capture room tone at the start of every session without moving the microphone or changing gain. Keep a clean section for edits. For ACX delivery, the current requirement recommends 1–5 seconds of room tone at the beginning and end of each file; internal pauses still need natural timing rather than a fixed duration.

    Don’t batch-delete breaths across an audiobook. The risk is clipped phrasing, digital silence, and inconsistent pacing. Keyboard shortcuts can speed up reviewed edits, but every cut should remain audible in context.

    ACX and Audiobook Delivery Standards

    For Audible/ACX delivery, master and inspect every finished file against the current ACX audio submission requirements. The official page was updated on April 15, 2026 and is the source of truth if a number changes after this guide is published.

    ACX currently requires each file to measure between -23 and -18 dB RMS, keep peak values no higher than -3 dB, and keep the noise floor no higher than -60 dB RMS. Submit a 44.1 kHz MP3 at 192 kbps or higher CBR. Use the same mono or stereo format for every file, put one chapter or section in each file, keep each file at 120 minutes or less, and leave 1–5 seconds of room tone at the beginning and end without exceeding five seconds.

    Measure representative room tone, not digital silence. If the noise floor fails, fix the room, microphone position, gain staging, or source recording before aggressive processing. Noise reduction has no guaranteed dB improvement and can create artifacts; use the noise-reduction guide for a controlled pass, then remeasure the exported chapter.

    Measure each finished chapter across the whole file. Adjust level, re-check RMS, then verify that peaks remain below -3 dB. ACX's current wording says peak level, not “true peak,” so do not silently substitute one metric for the other. Listen after every limiting or gain change; passing numbers does not guarantee clean narration.

    Sound Forge Pro 2026 documents ACX Check and ACX Export in View → Instant Actions. Run ACX Check on the finished chapter, export with ACX Export, then inspect the MP3 itself and use ACX's own audio analysis tool before submission. A wizard result does not replace a full listen-through.

    Frequently Asked Questions

    Can Sound Forge Pro be used for professional voice-over recording?

    Yes. Sound Forge Pro 2026 can record mono or multichannel audio, edit at sample level, run VST2/VST3 processing, and export delivery files. It is not a multitrack DAW: use a DAW when speakers need separate timeline tracks, routing, or later mix revisions.

    What recording levels should I use for voice-over in Sound Forge Pro?

    Use the interface or device control to set a strong signal that never reaches 0 dBFS. For unpredictable narration, peaks around -12 to -6 dBFS are a conservative editorial starting range, not a Sound Forge requirement. Record the loudest line before the take and leave more margin if the delivery is dynamic.

    Where are the voice-over recording controls in Sound Forge Pro 2026?

    Open View → Record Options. Choose Manual for general voice-over or punch-in work, choose Normal or Create regions as the recording mode, select monitoring, then use Arm and Record on the main toolbar. The current 2026 help does not document View → OTR as the primary path.

    How do I do punch-in recording in Sound Forge Pro for voice-over?

    Select the exact range to replace, set Method to Manual in View → Record Options, configure pre-roll in Settings, then Arm and record. Recording stops at the end of an active selection; without a selection it continues from the cursor and overwrites existing data until you stop.

    What are the ACX audio standards for audiobooks in Sound Forge Pro?

    ACX currently requires each file to measure between -23 and -18 dB RMS, stay below -3 dB peak, and have a noise floor below -60 dB RMS. Submit 44.1 kHz MP3 at 192 kbps or higher CBR, one chapter or section per file, no longer than 120 minutes, with 1–5 seconds of room tone at the beginning and end. Sound Forge 2026 includes ACX Check and ACX Export in Instant Actions.

    How do I remove breaths from voice-over in Sound Forge Pro?

    Don’t remove every breath automatically. Reduce or replace only distracting breaths, keep consonants and phrase timing intact, and use matching room tone instead of digital zero when a cut exposes silence. Follow the client's style guide for commercials, e-learning, or audiobooks.

    Does Sound Forge Pro have a de-esser for voice-over?

    The current core help does not document a dedicated native de-esser. Load a compatible VST2/VST3 de-esser through the Plug-In Chainer, or automate a narrow EQ only for mild cases. Mouth De-click removes mouth transients; it is not a de-esser and should not be presented as one.

    For the next pass, use the recording setup guide if the meter never moves, the playback troubleshooting guide if the take records but cannot be heard, and the hiss and hum workflow only when the source actually needs repair.

    Last fact-checked: August 2, 2026.