Paste a transcript. Get the best clips, ranked.
Hook strength, standalone context, payoff and quotability scored — 5 to 20 candidate moments with titles and timings.
Reviewed 2026-08-17 · runs in your browser where noted · WeaverClip pricing
Find Best Clips From Transcript — mine a full transcript for candidate moments
A finished hour of recording contains maybe five moments strangers would watch to the end. Finding them by scrubbing video is slow and biased toward whatever you remember; finding them in text is fast, because a transcript compresses an hour of footage into something you can scan in four minutes. This tool does the scanning for you. Paste a transcript, pick a mode, and it returns up to five ranked candidates — each with a title, a score, and the sentence that earned the rank. It is a mining tool: it surveys many moments and surfaces the strongest, rather than judging one moment you already chose.
The pipeline from paste to ranked list
Deterministic is the point. The same paste produces the same ranking every time, which lets you compare edits of a transcript, share results with an editor, and trust that a candidate did not appear because of a random seed. Here is exactly what happens when you press a mode button.
Step 1 — split into sentences
The text splits on sentence boundaries: a period, question mark, or exclamation point followed by whitespace. Fragments shorter than 21 characters drop out, because a sentence that small cannot carry a clip on its own. This step is why punctuation quality matters so much for transcript mining — a wall of text without sentence marks is one giant sentence to the splitter, and the ranking degrades to whatever the first eight fragments happen to be.
Step 2 — score the first eight sentences
The tool examines the first eight qualifying sentences. Each starts at a base of 68 points and earns additions from two places:
- Structure markers (+8). A colon, an em-dash, or the words why, how, never, or stop. These markers correlate with a sentence that sets up a claim and then delivers it — the shape of a standalone thought.
- Length spread (0–16). The sentence's character count modulo 17. This deliberately rewards variety rather than raw length: two adjacent sentences of similar length get different contributions, which keeps the list from crowding around one rhythm.
The score caps at 96. No candidate is ever presented as perfect, because a text score cannot see delivery, pacing, or audio quality.
Step 3 — apply the mode bias
Two modes add a topical bonus of 12 points when a sentence matches their subject matter:
- funny boosts sentences containing funny, laugh, or joke.
- controversial boosts sentences containing hot take, unpopular, or wrong.
The other modes change the lens without changing the arithmetic: educational labels each pick "Useful, tactical explanation," story labels each pick "Complete story with payoff," and best and emotional label picks by the general structure test. Switching modes is therefore two different operations wearing one control: for funny and controversial you re-rank; for the rest you re-label the same underlying candidates.
Step 4 — sort, cap, and keep five
Candidates sort by score, descending, and the list truncates to five. If two candidates tie, the one that appeared earlier in the transcript stays ahead. Five is a deliberate ceiling: a mining pass that returns twenty "best" moments is not a decision, it is a second inbox.
What the ranking rewards, translated to writing
The arithmetic above is a proxy for four qualities that make a sentence clip-worthy. Knowing them lets you improve the transcript before you paste it.
Hook. Sentences that open with why or how, or that plant a prohibition (never, stop), read as openings. They create a question the listener needs answered. If your episode's best moment starts with "So anyway," the miner sees a weak opening even if the payoff is gold.
Structure. Colons and em-dashes are setup→delivery punctuation. "Here is the mistake: we shipped before testing" has the shape of a complete mini-argument in one breath. The +8 marker bonus exists because that shape survives being cut out of context more often than run-on narration does.
Self-containment. The 21-character floor and the sentence-boundary split both push toward thoughts that resolve within themselves. A sentence that refers back to "what I said earlier" or forward to "the thing coming up" scores like any other sentence but fails as a clip; that gap is yours to catch at review time, and it is the main reason the tool is a shortlist generator rather than a final editor.
Quotability. Short, complete, marker-bearing sentences are the ones people repeat back. The miner cannot measure whether a stranger would quote the line, but the same features that make a line repeatable — brevity, a verb, a stance — are the ones it can see.
Preparing a transcript that mines well
Garbage pacing in, garbage ranking out. Ten minutes of preparation changes results more than any mode choice.
- Restore sentence punctuation. ASR output often arrives as comma soup or one giant line. Add periods where thoughts end. The splitter only knows what you tell it.
- Keep speaker labels, but put them outside sentences. A leading "Sarah:" survives the split fine, while "[crosstalk]" interjections inside a sentence create fragments that pollute the candidate pool. Move stage directions to their own lines; they fall out as sub-21-character fragments.
- One thought per sentence where possible. Long compound sentences that contain two ideas split badly and score as neither. If a guest's best answer is one 90-word run-on, break it at the natural hinge before pasting.
- Trim the dead open. Most episodes spend their first ninety seconds on housekeeping. The tool scores the first eight qualifying sentences, so a long cold-open of "can everyone hear me" consumes candidate slots that the actual content should hold. Cut the warm-up from the paste, or paste starting at the first real topic.
- Keep numbers and names in. Specifics ("we lost 4,000 subscribers in March") give candidates the texture that makes them watch; they also change the length-modulo contribution, which is a legitimate way to nudge a true-but-flat sentence upward by making it more precise.
Reading the results honestly
Scores live in a 68–96 band by construction. That band has meaning:
- 90s — strong structure plus mode fit. Treat as first-cut candidates, then verify by ear.
- 80s — workable. Usually needs one of: a sharper first line, a tighter ending, or a context sentence prepended in the edit.
- High 60s to 70s — the sentence is complete but unremarkable. Useful as B-roll narration or chapter anchors rather than standalone clips.
Two honest caveats about the output columns. First, the time labels beside each candidate (00:00, 00:01, …) are ordering markers the tool assigns so you can see pick order at a glance — they are not parsed clock positions from your recording. To locate a pick in footage, match its text against your timestamped transcript. Second, the "why" line describes the scoring reason (structure, story shape, tactical value), not a content judgment; a candidate can score high and still be a moment you personally would not post. The ranking tells you where to look, never what to publish.
Worked example: one episode, five candidates
Note for find-best-clips-from-transcript: This walkthrough is a hypothetical example, illustrative rather than a sourced case — an invented episode used to show the mechanics.
Suppose you paste this (trimmed) transcript of a fictional pricing episode:
"Welcome back to the show. Today I want to explain why we killed our unlimited plan. Two years ago we promised unlimited everything and priced for average use. The average user does not exist: one customer recorded 340 hours in a month. Here is the math that changed our minds — serving that one account cost more than twelve subscriptions bought. We killed the plan on a Tuesday, and churn dropped by the end of the quarter. The lesson: price for the extreme user, because the extreme user is who you actually serve. That is the whole story."
The splitter yields eight qualifying sentences. The colon sentences ("Here is the math…" / "The lesson:…") earn the marker bonus; "why we killed our unlimited plan" earns the why-marker; the 340-hours sentence carries a number and an unusual length. After sorting, the top of the list is the math sentence and the lesson sentence, with the origin sentence ("Two years ago…") in third — which is exactly the editorial call a human would make, because those are the three lines that survive being quoted without context. The mode choice then tunes what you do with the list: story keeps all three as one arc, educational foregrounds the math, controversial would boost any "unpopular opinion" framing if the episode contained one.
From there, the workflow is: open the recording at each pick's location, listen to ten seconds either side, and mark the actual in/out points with a little padding. The text told you where; the footage tells you how.
From candidate to cut: closing the gap between text and footage
A ranked list is a map, not the territory. The distance between a good-looking sentence and a good clip is closed in four checks, in order:
- Find the moment in the recording. Match the candidate's text against your timestamped transcript, then jump to that point in the footage. Keep a margin: the sentence you ranked usually sits inside a longer spoken beat that starts earlier and resolves later than the text suggests.
- Listen to the delivery. Text flattens emphasis. A sentence that reads flat can be the best moment of the episode because of how it was said — and a sentence that reads brilliantly can have been mumbled into a cough. The score never heard any of that; you are the tie-breaker.
- Check the edges. Does the thought start inside the previous sentence? Does the speaker add a clarifying "and that matters because…" right after the ranked sentence ends? Clip boundaries chosen at sentence edges in text usually want two to five seconds of extra padding on each side in footage.
- Test for crosstalk and overlap. Transcripts often drop or merge overlapping speech. If two voices were talking at once, the text looks clean while the audio is a mess — and no edit fixes audio you never recorded cleanly. Keep the candidate only if the underlying audio survives being heard alone.
Candidates that fail checks three or four are not wasted: they make excellent chapter titles, newsletter quotes, or B-roll narration, and the miner will find them again next time because the ranking is deterministic.
Choosing a mode for your content type
- best — the default lens. Structure-only scoring, no topical bias. Start here on any transcript you have not seen before.
- educational — for tutorials, lectures, and explainer episodes. Same underlying scores, labeled to keep you oriented toward tactical sentences when you skim the list.
- story — for narrative shows and interview arcs. Labels picks as complete stories so you assemble one arc instead of five disconnected zingers.
- funny — the only mode that re-ranks for humor. Use it on comedy-leaning shows; the 12-point bonus genuinely reshuffles the top five when jokes are present.
- controversial — re-ranks toward hot takes and dissent. Genuinely useful for debate formats; use it knowingly, because optimizing for controversy is an editorial stance, not a neutral filter.
- emotional — same candidates, oriented toward the affecting moments. Pair with story mode when an episode has one emotional peak you want to find fast.
A practical pattern: run best first for the honest structural ranking, then run the topical mode that matches your show. Candidates that survive both passes are your shortlist.
When the tool returns one candidate instead of five
Paste fewer than three qualifying sentences and the miner changes shape: it returns a single candidate titled "One complete thought," scored 82, showing the first 220 characters of what you pasted. That is the tool telling you the input is a moment, not a mine. Two situations end up here:
- The paste is genuinely short — a single answer, a voicemail, a cold open. The fallback is the right answer: this is one clip candidate, and its job is to stand alone.
- The paste is long but unpunctuated — the splitter saw one or two giant sentences. Fix the punctuation and re-paste; the real candidate list appears.
An empty paste returns the same single card with empty text rather than an error, which is the signal to check that the paste actually landed.
Where mining fits: shortlist first, judge second
Mining and judging are different operations, and this site keeps them on separate tools on purpose. This page finds many candidates from a long transcript. The Viral Clip Score tool takes one candidate you already have and scores whether it stands alone — hook, clarity, payoff, quotability — with a suggested hook rewrite. The productive loop runs in that order:
- Mine the full transcript here → five candidates.
- Score each finalist with the standalone checker → kill the ones that need context you cannot give them.
- Cut the survivors with padding, verify edges, and publish.
Running the loop backwards — scoring moments you picked by memory — reproduces the exact bias that mining exists to remove: you only ever test what you already remembered. The transcript contains moments you forgot saying. That is the whole value proposition of the miner, and why a deterministic, text-level pass beats re-watching the footage for discovery.
For format-specific lenses — podcast story arcs, gaming events, sermon context, interview answers, webinar lessons — the sibling tools on this site apply the same mining engine with mode sets tuned to each format's definition of a highlight.
.
The scoring arithmetic on one sentence, end to end
Because the ranking is deterministic, you can audit any candidate by hand. Take the fictional sentence "Here is the math that changed our minds — serving that one account cost more than twelve subscriptions bought." and walk it through:
- Base score: 68.
- Structure markers: the sentence contains an em-dash, so +8 → 76.
- Length spread: the sentence is 111 characters; 111 modulo 17 is 9 → 85.
- Mode bonus: none in best mode; in controversial mode this sentence has no hot-take vocabulary, so still 85.
- Cap check: 85 is under 96, so it stands.
Change one word and the audit changes with it. Replace the em-dash with a comma and the marker bonus vanishes (−8). Shorten the sentence under 21 characters and it drops out of the pool entirely. The point of showing the math is calibration: once you know a colon is worth eight points and a why-hook does not move this stage at all, you stop being surprised by the list and start using it.
What text-level mining cannot see
The miner reads words. Everything else about a moment is invisible to it, and naming the blind spots keeps the output in perspective:
- Delivery. Sarcasm, emphasis, a pause before the punchline, a voice crack — none of it reaches the text. A flat sentence can be the moment the room went quiet, and a brilliant sentence can have been read while someone chewed.
- Laughter and reaction. The funniest moment of an episode is often the three seconds after the ranked sentence, where the hosts lose it. The miner points at the setup; the payoff may live in audio it never heard.
- Visual events. A demo failing on screen, a prop appearing, a guest's face — mining text cannot see any of it. For event-driven content, that is why a gaming-specific lens matters more than any transcript score.
- Context the speaker assumes. "As I said before the break" is a structurally fine sentence and a broken clip. The 21-character floor filters fragments, not references.
None of this argues against mining; it defines the division of labor. Text finds the where; ears and eyes confirm the what. Tools that pretend text can do the whole job end up shipping clips with no audio review — which is exactly the thin-clip failure mode this workflow exists to prevent.
A weekly mining workflow that scales
Note for find-best-clips-from-transcript: The schedule below is an illustrative, hypothetical workflow for a fictional two-show operation — a template to adapt, not a measured case study.
Assume a fictional creator producing one 60-minute podcast and one 45-minute interview per week, each transcribed within an hour of recording. A realistic mining pass costs about 25 minutes per episode:
- Minute 0–5: clean the transcript. Restore periods, pull bracketed noise onto its own lines, cut the housekeeping open.
- Minute 5–8: run best mode on the first half of the transcript; note the five candidates.
- Minute 8–11: run best on the second half; note those five. Merge the two lists.
- Minute 11–15: run the show's topical mode — story for the podcast, emotional or expert-leaning review for the interview — and mark any candidate that appears on both passes. Double-listed candidates are the priority shortlist.
- Minute 15–25: score each shortlisted candidate with the standalone clip checker, kill the ones that fail self-containment, and mark in/out points in the recording with padding.
That is roughly 50 minutes of discovery per week for two shows, replacing what would otherwise be two-plus hours of scrubbing. The deterministic ranking means an editor can rerun the same passes and get the same list, which turns clip selection from a debate about taste into a review of evidence. When a candidate list feels wrong twice in a row, the transcript prep is almost always the cause: bad punctuation splits thoughts at the wrong places, and the miner faithfully ranks what the split gave it.
Why a shortlist beats a firehose
Some clipping tools advertise dozens of suggestions per upload. For a miner that works on text, more output is less useful, because every extra candidate beyond the top few costs more review time than the ranking saved. Five candidates is the point where a human can verify each one by ear inside ten minutes; twenty candidates means the verification gets skipped, and unverified candidates become clips that cut mid-thought. The five-cap is therefore a quality control, not a limitation: it forces the workflow to stay mine → verify → cut, in that order, and it keeps the cost of discovery proportional to the number of clips you can actually finish.
FAQ
Does the tool upload my transcript anywhere? No. The scoring runs locally in your browser tab; pasted text is not stored or transmitted. Disconnect the network and the ranking works identically.
Why does my paste get only five candidates when the catalog copy mentions up to twenty? The implemented ranking returns the top five from the first eight qualifying sentences in one pass. The practical way to mine deeper is to paste the transcript in sections — each pass returns its own top five, and you merge the lists yourself.
Are the time labels real timestamps? They are ordering markers showing pick sequence, not clock positions. Locate picks by matching their text against your own timestamped transcript.
Can it mine a transcript with speaker labels? Yes, as long as labels do not break sentence punctuation. Leading labels like "Host:" are fine; bracketed noise inside sentences should move to its own line so it drops out.
The top candidate is a sentence I would never post. Is the tool broken? No — the score measures structural strength in text, not taste or delivery. The miner's contract is to surface candidates worth checking; your ear makes the final call.
What is the highest possible score? 96. The cap is deliberate: a text-only score can never certify a clip as perfect, because delivery, audio, and pacing are invisible at this stage.
Do modes change scores or just labels? funny and controversial change scores via their topical bonuses. educational, story, emotional, and best change the explanation labels while the underlying ranking stays identical
Protect the next recording — verified before delete
If this calculator says your 4-hour stream will use ~22 GB, WeaverClip's OBS helper can upload each one-minute segment as the next minute records and only queue local deletion after byte-count + MD5 verify. Missed segments stay and retry. That is the difference between a number and a guarantee.
- WeaverClip plan catalog — storage GB, processing hours, overage $0.04/GB-month
- OBS container behavior — MKV vs MP4 moov — verified via ffmpeg/ffprobe and WeaverClip recovery checker (client-side probe)
- Platform safe zones — measured against YouTube Shorts / TikTok / Reels overlays, 2026-08-17
- Competitor pricing — OpusClip cost page stamped 2026-08-17, re-verified monthly; dataset versioned

