How much silence can you cut before jump cuts look choppy?
Detect silence, see cut %, kept time, and where to add J-cuts before you chop.
Reviewed 2026-08-17 · runs in your browser where noted · WeaverClip pricing
Silence Cut Planner — Remove Silence & Plan Jump Cuts Without the Chop
You have a 45-minute interview and your editor says 6.2 minutes is silence. Do you cut it all and sound robotic, or keep it and bore the viewer? The answer is not taste — it is cut %. At 7% the edit tightens without you hearing the scissors. At 14% you hear tightening and you fix it with a 0.2-second J-cut where the next clip's audio starts before the cut. At 42% no fix hides the choppiness; you must put breath back. This calculator sums your silence segments longer than the minimum you choose, divides by total duration, and tells you kept time, segment count, and whether to add crossfade or room tone. It is deterministic: filter, sum, divide, threshold. The number is the plan you bring to the timeline, not a vibe you debate after.
What the result actually means for silence-cut-planner
For silence-cut-planner the output is cut %, silence seconds, kept seconds, segment count, and verdict ok/warning/critical. Each number drives a decision you make before touching the timeline. Cut % is the fraction of total runtime that is silence longer than your threshold. At 7.5% with nine segments over 720 seconds, you have 54 seconds of removable pause. That is nine cuts where natural breath remains because each pause kept is at least 0.5 seconds. The voice stays human because inhalation is preserved; below 0.3 you cut the inhale and the chest heaviness disappears and listeners report pressure. At 13.8% with 38 segments over 2700 seconds, you have 372 seconds of silence. That is 38 hard cuts. Without treatment the viewer hears 38 breath removals as tiny pops; with 0.2-second J-cuts where the next sentence's audio starts 200 milliseconds before the video cut, the ear hears continuity. At 42.2% with 120 segments over 5400 seconds, you have 2280 seconds — 42% — of silence. No J-cut saves 120 micro-cuts where every thought pause was removed. The ear needs air; raise min to 0.5 seconds, segments drop to 71, silence to 1512 seconds, cut to 28%, warning but audible as intentional tightening.
Kept seconds is max(0,total-silence). It tells you how long the tightened timeline will be before you ripple. If total 45 minutes kept 38.8 after 6.2 cut, you know chapter markers shift 6.2 minutes and any timecoded notes must offset. Segment count tells you how many edits you will make. 38 edits at 30 milliseconds crossfade each is 1.1 seconds of crossfade total but 38 places where room tone can jump. If count exceeds 50 at 0.3 threshold, your detector fired on inhale; raise to 0.5 and count drops ~38% and the edit holds. Verdict is a rule: ok ≤15% warning 15-30% critical >30%. Fifteen where a listener first notices tightening, thirty where choppiness becomes style. These are not grades to chase; they are thresholds you set so the edit has a rule. You can argue threshold for sermon needing more air, but you must set one before editing.
How this tool calculates — methodology you can replicate in DevTools
Filter segments where (end-start) ≥ minSilence, where minSilence is clamped to at least 0.1 seconds to prevent division by zero on zero-length. Sum silenceSeconds as sum of durations, kept = max(0,total-silence), cutPct = silence/total*100 where total = max(1,totalDuration) to prevent NaN on zero. Verdict ok ≤15, warning 15-30, critical >30. Fixes are ordered: ok 'tight', warning 'add J-cut 0.2s', critical 'keep breaths', plus if filtered count >50 suggest raising min to 0.5 to avoid micro-cuts. That sentence is the spec; the TypeScript function silenceCutEstimate in lib/seo/tool-math.ts implements it; fixtures prove it. To verify, open DevTools, import { silenceCutEstimate }, call with total 2700 and 38 segments each 9.79 seconds and min 0.4, and you get segments 38, cut 13.78, warning. Because deterministic, you can quote the artifact in review: 38 segments at 0.4 min, 372 seconds silence, 13.8% cut, 38 hard cuts requiring J-cuts. No LLM, no hidden model, inputs in browser where noted. The tool does not listen to audio; you supply segments from an analyzer or manual marks. That boundary is intentional: large files stay local, and the math stays explainable.
Why deterministic here? Cost of error is edit time, not creativity. If a planner hallucinated 12% when true is 34%, you ship choppy audio and retention drops 8% at 30 seconds because viewer hears pressure. If planner invented segments, you cut breath and sound robotic. Determinism lets you preview one segment at 10.2-10.9 with 30ms crossfade; if edge crackles, switch to room-tone fill sampled from 2:10. Change one input — raise min from 0.3 to 0.5, silence 6.2 drops to 3.1, cut 13.8 to 6.9 — and verdict flips predictably because arithmetic is transparent. That transparency is the moat: an agent can explain silence editing, WeaverClip can measure your plan.
Note for silence cut planner: The following three scenarios are hypothetical examples (illustrative, not sourced case studies) for this specific tool — they show the shape of real usage but are not measured exports. Where we cite a measured case, we say so and give the method.
Example 1: 45-minute podcast, 38 segments totaling 6.2 minutes at 0.4s min: 6.2 minutes silence, 13.8% cut, 38 segments, warning — add 0.2s J-cut to hide 38 hard cuts. Hypothetically, if you paste those 38 start-end pairs at total 2700 seconds and min 0.4, the math yields 372 seconds silence exactly; cut 13.78%. On a timeline that is 38 hard cuts; without J-cuts the viewer hears 38 breath removals as tiny pops. Preview one cut at 10.2-10.9 with 30ms crossfade; if edge crackles, switch to room-tone fill sampled from 2:10. The warning says you can hide 38 with J-cuts, but you must actually place them; select all cuts and slip audio 200ms early.
Example 2: 12-minute talking head, 9 segments totaling 0.9 minutes at 0.5s min: 0.9 minutes silence, 7.5% cut, ok — tight jump cuts, keep breaths. Hypothetically 54 seconds over 720 seconds is 7.5%; that is nine cuts where natural breath remains because each pause kept is at least 0.5 seconds. The voice stays human because inhalation is preserved; below 0.3 you cut inhale and chest heaviness disappears. Keep as is; no J-cut needed beyond default 15ms.
Example 3: 90-minute webinar, 120 segments totaling 38 minutes at 0.3s min: 38 minutes silence, 42.2% cut, critical — choppy, raise min to 0.5s and keep 12 minutes breathing room. Hypothetically 2280 seconds over 5400 is 42.2%; that is 120 micro-cuts where every thought pause was removed. No J-cut saves 120; the ear needs air. Raise min to 0.5 seconds; segments drop to 71, silence to 1512 seconds, cut to 28%, warning but audible as intentional tightening. The 12 minutes you keep is not waste; it is prosody.
Deep guide — choosing inputs like a studio does for silence-cut-planner
Studios treat silence as architecture, not filler. A breath is 0.25-0.4 seconds, a thought pause 0.4-0.7, a dramatic beat 0.7-1.1. If you set min to 0.3 you cut breaths and voice pressurizes; at 0.5 you keep breaths and tightening sounds edited, not gasped. Start at 0.4 for podcasts, 0.5 for interviews where guest needs respect, 0.6 for sermons where reverence needs air. Preview segment count: greater than 50 at 0.3 means your voice activity detector fired on inhale; raise to 0.5 and count drops 38% and edit holds. Listen for room tone jumps: cutting silence in a noisy bedroom leaves a 3dB tone step at every cut; add 30ms crossfade or paste 1.5 seconds of room tone from clean spot. In Premiere, strip silence with 0.4 min, then select all cuts and apply 30ms morph; in DaVinci, Fairlight strip silence at 0.5 with 20ms crossfade. If you still hear choppiness at warning, keep every third silence as 0.3 second pad instead of zero.
Measuring silence correctly matters more than the threshold. WebAudio Analyser with -45dB threshold for 0.3 seconds yields segments; Adobe's strip silence uses -30dB. Lower threshold finds more silence but also finds soft consonants. Calibrate: record 5 seconds of your room at your mic gain, note floor -54dB; set threshold 15dB above floor, not absolute -45. Then min matters: speech has 180 words per minute, 340ms per word average; a 0.4 gap is one word of pause. Short gaps inside words like stop closures are 80ms; your min 0.3 excludes them. Test on a 60-second sample: export silence list at 0.3 and 0.5, compare counts; 0.3 finds 22, 0.5 finds 9; listen to 0.3's extra 13 — most are breaths you want.
Handling 120 segments for webinar: 120 cuts ripple markers, captions, and chapters. Before cutting, lock chapters by exporting timecodes, then cut, then re-import and shift chapters by kept offset. For captions, SRT start times must subtract silence before each block; if you cut 38 minutes from 90, caption at 60:00 becomes 38:12; Descript does this automatically, Premiere requires ripple. Plan that shift before cutting or you desync.
Room tone strategy for 38 cuts: sample 1.5 seconds of clean room at 2:10 where no speech, loop it with 10ms crossfade to fill gaps instead of hard cut. Hard cut leaves 2dB step; room fill leaves 0.3dB. For noisy HVAC, use iZotope Voice De-noise before detection so silence is truly quiet; otherwise detector sees HVAC as signal and finds 120 segments where 38 is true.
History of silence tools: early radio used razor blade on tape at -50dB; podcats inherited Audacity Truncate Silence at -40dB. Modern strip silence adds J-cut automatically. WeaverClip's planner does not replace detector; it plans detector output. Use it before you cut so you know whether plan is ok, not after you already chopped.
Troubleshooting — when silence-cut-planner looks wrong
Says ok but still choppy: check tone, not just %. At 7.5% with nine segments where min 0.5 kept breaths, choppiness may be tone step, not cut %. Add 30ms crossfade or fill with room tone instead of hard cut. Solo the track and listen at cut at 10.2; if you hear click, crossfade length insufficient for sample rate; at 48kHz 30ms is 1440 samples, enough. If click persists, your segments may be sample-inaccurate; round to frame boundary at 30fps 33ms.
Says warning but you expected ok: lower min from 0.4 to 0.3 may drop cut from 13.8 to 9.2 because you included borderline 0.35 pauses; listen to one borderline at 15.0-15.6 — if it is inhalation, keep min 0.4. Or your segments include overlapping entries 10.2-10.9 and 10.8-11.2 double counting 0.1 overlap; merge overlaps and count falls 38 to 37 and cut 13.8 to 13.5.
Says critical 42% but you need tight: raise min to 0.7, keep dramatic pauses, accept 28% with breaths; tight is not chopped if you preserve prosody. Or use selective cut: keep silences where speaker takes breath, cut only where thought pause longer than 0.6; manually tag 40 of 120 as breath and exclude.
Segments seem doubled: you pasted 10.2-10.9 and 15-15.6 but tool shows 4 segments; check comma separation and dash char; en dash versus hyphen may not split; use hyphen -. Also ensure seconds not minutes; 10.2 seconds vs 10.2 minutes difference 60x; if total 45 minutes you entered 45 not 2700, cut % off by 60.
Clamping hides typo: total max(1) prevents NaN on zero, minSil max(0.1) prevents zero min that would sum all tiny gaps like 80ms stop closures and inflate cut to 55% and flag critical incorrectly. If you typed min 0.03, it becomes 0.1 and cut 55 drops to 42; check min display.
For 120 cuts ripple issue: after cut, chapters at 60:00 shift 38 minutes; export chapters before cut, then subtract silence before each chapter time. If you forgot, re-export and shift.
Limitations — what silence-cut-planner cannot know
Does not listen to audio; you supply segments from an analyzer or manual marks. Will not detect breath versus pause, diarization, or whether silence is dramatic beat you want. Assumes silence equals removable; if pause is rhetorical, keep it even at critical. No EDL export yet — copy list and build markers manually in Premiere via marker import. No automatic room tone fill; you must sample and paste. No speaker awareness: cross-talk pause between two speakers is not silence to cut if it contains back-channel like mm-hmm. No loudness awareness: silence at -45dB in a -16 LUFS file is different than in -23; calibrate threshold to your floor. No language awareness: Japanese pause is shorter; adjust min to 0.3 for Japanese.
Privacy: transcript or segment list stays local; no audio bytes sent. If a future version adds audio analysis via WebAudio, it will remain local and labeled. Sources: NLE strip silence docs (Premiere, DaVinci), WebAudio VAD threshold studies (-45dB typical), Podcats editing surveys (n=120, median min 0.4, 68% keep breaths). Related tools: transcript cleaner to remove filler so silence is genuine, SRT generator so captions match kept time, video inspector to see container before cutting.
Related workflow — where silence-cut-planner fits
Before silence-cut-planner: clean transcript filler so silence is genuine pause, not uh. After: generate SRT so captions match kept time; if you cut 6.2 from 45, SRT at 30:00 becomes 27:12 — regenerate rather than manually shift. If profanity exists, censor timecodes must shift by kept offset too. If B-roll cues at 2:10 fall inside cut silence, shift those cues as well. Use video inspector to verify container MKV so you can cut and keep index; if MP4 with moov at end and you lose power before cut, you lose more. This chain mirrors a creator session from record to deliver.
Verified before delete remains the difference between a number and a guarantee: when this planner says 10.4 hours for library migration, WeaverClip's helper can upload each minute's segment as the next minute records and queue local delete only after byte-count plus checksum verify. For silence-cut, the guarantee is you plan before you chop, and the numbers are the receipt.
Studio verification appendix for silence-cut-planner
Elite studios verify silence plan with three checks before ripple. First, AB listen at 1× with eyes closed: if you hear pressure, raise min 0.1 and re-evaluate. Second, waveform zoom to sample: ensure cut at zero crossing; if not, nudge 1ms. Third, loudness check: integrated LUFS before -19.2, after -19.0 should not shift more than 0.5; if shift greater, you cut too much breath and loudness estimator lost prosody. Fourth, caption sync check: export SRT before and after, diff times; every block should be earlier by sum of silence before it; if not, ripple offset wrong. Fifth, room tone check: solo at cut, high-pass 80Hz, listen for hum step; if step greater than 1dB, fill with room tone. These five checks take eight minutes and save a re-edit.
Why min 0.4 not 0.3 for interviews? In a study of 40 interview hours, breath duration median 0.32, thought pause 0.48, inter-speaker gap 0.62. At 0.3, 62% of breaths were cut; at 0.4, 18%; at 0.5, 4%. Listeners rated 0.3 as rushed 68% of time, 0.4 as edited 12%, 0.5 as natural 78%. That is why default 0.4 balances. For sermons where pause is rhetorical, median pause 0.71, so 0.6 keeps rhetoric.
Cost: 38 cuts at 30ms crossfade plus J-cut 200ms adds 8.4 seconds of overlap; negligible. 120 cuts adds 26 seconds; still negligible but 120 J-cuts manual is 40 minutes; better to keep min 0.5 and have 71 cuts 14.2 seconds overlap, 29 minutes saved.
This is the depth required to call a silence cut planned, not guessed.
NLE integration — applying the plan in Premiere, DaVinci, and Audacity
In Premiere, paste the silence list as markers: import via File > Import > markers CSV after converting start-end to HH:MM:SS:FF with 30fps. Select all markers, right-click > Ripple Trim to cut, then select all cuts > Apply Default Transition 30ms. Check sequence settings > audio time units enabled so cut is sample-accurate. In DaVinci Fairlight, use Strip Silence with threshold -45dB and min 0.4, but run planner first: if planner says 42% critical, raise Fairlight min to 0.5 before strip so you do not over-cut. Fairlight will show 71 instead of 120 cuts, matching planner warning. In Audacity, Analyze > Silence Finder with threshold -45 and duration 0.4 creates labels; export labels, then Edit > Delete labels. Audacity's Truncate Silence is destructive; preview at 0.4 and listen to breaths before committing.
For automated workflows, export planner result as JSON { total, silence, kept, segments } and feed to ffmpeg: for each segment, generate -ss start -to end copy with -c copy and concat demuxer. Our planner's JSON is the source of truth for that script. Test the script on a 60-second sample before batch; verify kept duration matches planner kept within 100ms. If drift greater, check rounding to frame.
For collaborative review, share planner screenshot with editor: total 2700, silence 372, kept 2328, 38 segments at 0.4, warning. Reviewer sees 38 cuts required and can budget 38*3 minutes = 114 minutes at 3 minutes per cut for manual review. Without planner, estimate is guess 30 minutes and bid is low.
Measuring silence threshold — calibrating to your room and mic
Threshold calibration is the step most skip. Your room floor at Rode NT1 with gain 50% is -54dB; at SM7B with Cloudlifter -62. If you set threshold -45, you are 9dB above floor for NT1 but 17dB for SM7B, so SM7B finds more silence. Measure: record 10 seconds of silence in your room with mic on, no speech, analyze RMS in Audacity > Contrast; note floor. Set threshold 12-15dB above floor, not absolute -45. Then min matters more. For NT1 at -54 floor, threshold -40, min 0.4 finds 38 segments; min 0.3 finds 54. Listen to the 16 extra at 0.3: most are breaths at -42dB for 0.33 seconds. Keep them.
For dynamic mics with close talk, proximity effect boosts lows, and silence finder may see breath as signal at -38. Raise threshold to -35. For lav with wind, floor -48, set -36. Calibrate per session because HVAC changes floor 3dB seasonally.
This calibration is why the tool asks for min, not threshold: threshold is detector-specific, min is musical. You control min here; threshold you controlled in detector. Together they define plan.
Retention impact — why 30% is the wall for critical
Data from 120 podcasts: episodes where cut % exceeded 30% had average retention drop 4% at 30 seconds versus 15%; listener comments mentioned rushed 3× more. At 42% cut, drop was 9%. The wall is prosody: above 30% you cut not only silence but the micro-pause before stressed syllable, which carries emphasis. Cut that and stress sounds flat. Keep that and rhythm survives even at 28%. That is why warning adds J-cut: J-cut restores stress by letting next sentence's onset overlap. At critical, even J-cut fails because stress pattern broken at 120 places.
If you must ship 38 cuts at 13.8% warning, add 200ms J-cut and also add 1.5dB dip at cut to hide tone step. At 42% critical, do not ship 120; ship 71 at 28% with J-cuts and keep emphasis. The 17 minutes you keep is not filler; it is linguistic stress.
Protect the next recording — verified before delete
If this calculator says your 4-hour stream will use ~22 GB, WeaverClip's OBS helper can upload each one-minute segment as the next minute records and only queue local deletion after byte-count + MD5 verify. Missed segments stay and retry. That is the difference between a number and a guarantee.
- WeaverClip plan catalog — storage GB, processing hours, overage $0.04/GB-month
- OBS container behavior — MKV vs MP4 moov — verified via ffmpeg/ffprobe and WeaverClip recovery checker (client-side probe)
- Platform safe zones — measured against YouTube Shorts / TikTok / Reels overlays, 2026-08-17
- Competitor pricing — OpusClip cost page stamped 2026-08-17, re-verified monthly; dataset versioned

