Every minute transcribed counts once. Captions reuse it forever.
WeaverClip transcribes complete 1-minute segments with word-timing. A minute that finishes counts once; failed minutes, re-watching, and searching don't. Dual 16:9+9:16 pair shares one transcript. Free 2h total, Creator 4h/mo, Pro 8h/mo, Studio 70h/mo.
Most tools charge you to watch your own video again. Weaver charges you once to understand it. A 2-hour sermon at 6 Mbps is 2.0 transcript hours — whether you make 1 clip or 18, whether you watch it 3 times, whether you download the full SRT a week later. That single transcript powers 20 ranked clip suggestions per hour (max 200), word-timed captions with 5 presets, and 300ms search without downloading video. The nuance streamers feel at 2am: you don't pay to store the 24h, you pay to understand the hours you actually clip.
Word-timed transcript: every word has startMs/endMs, not jus
Counts once: re-watching, scrubbing, searching, downloading
Dual canvas shares one transcript: 2h of show is 2h, not 4h,
Schedule transcribe: pick the last 6h of a 24h, not all 24h
SRT ⇄ VTT ⇄ TXT, overlap fixer, 2-line linter
All in browser — your file never leaves the page.
See the 3 moments we'd pull before you sign up
MCP weaverclip_search_transcript — free taste.
How much silence are you transcribing?
Is your old VOD already gone?
Twitch deletes fast. WeaverClip keeps it.
Does your lane keep up live?
Record bitrate vs upstream — will you finish uploading after you stop?
Two numbers match or you don't touch the file
MCP weaverclip_get_usage — verify before delete.
One canvas or two?
1.4 TB freed — zero files lost
Will your webcam get cropped?
Do people actually read this?
*Hours are at 6 Mbps (1080p30 webinar, slides, sermons) ≈2.7 GB/hour. Actual hours vary with bitrate — 5 Mbps ~2.25 GB/h, 12 Mbps ~5.4 GB/h, 25 Mbps ~11.25 GB/h. Vault is always bytes; hour label is a guide. Hard limit = 2× vault (e.g., Creator 75→150 GB), overage $0.04/GB-month. Transcript never bills overage.
Sharp nuances you didn't know you needed
How it works — 1-minute segments, word by word
Your helper or import splits nothing — OBS already writes 1-minute MKVs. Each segment is uploaded, verified byte count + MD5, then the worker slices the audio and sends it to Modal. Modal returns words with startMs/endMs, not just text. We stitch those words into transcript_chunks (seq, start_ms, end_ms, text, words JSON, search_vector). A 2-hour VOD is ~120 chunks, each searchable. Failed chunks don't count; they stay pending and retry. Already-transcribed minutes never count again — we check transcript_chunks before billing.
The nuance no one tells you: transcript hours are scarcer than storage
Your plan gives you far more gigabytes of storage than hours of transcription, so the transcript allowance is the one that runs out first. Choosing which hours to transcribe is the decision that matters. A 12h gaming week at 12 Mbps is ~64.8 GB of vault, but you don't need 12h of transcript. Use weaverclip_schedule_transcribe with startMs=4*3600000, endMs=12*3600000 to transcribe only the last 8h where chat was hype. For a 24h subathon, schedule the 6h window with the sponsor read, not all 24h. The metrics strip on /dashboard shows '1.2h left' before you pick, so you never burn 4h on loading screens.
Dead-air reclaim — the 38% you didn't know you paid for
38% of a 6-hour stream is loading, BRB, or silent. That's 2.3h of transcript you don't need to spend if you pick the window. Weaver shows dead-air percentage on the Safety Rail and lets you skip it. Opus would bill the whole 6h of credits; we let you bill 3.7h. For a podcaster with 12 eps × 90 min = 18h, that's 6.8h of dead air saved — almost a full Pro month.
Search without downloading video — the agent superpower
weaverclip_search_transcript does not download video. It queries the transcript_chunks search_vector (tsvector) server-side and returns [{ startMs, endMs, text }] in 300ms for 'pricing' across 6h. The agent then calls weaverclip_bulk_create_edits with 12 hits → 12 queued renders, without ever pulling 6h of MP4 through the chat. That's why agents can handle 24h without piling ffmpeg on your Mac — the transcript is the index, not the video. Try it live: paste any 2-hour transcript into Hook Finder on /features/transcription and see the 10 moments ranked before you sign up.
Word-timed vs blob — why it matters for captions
A blob transcript ('Here is a long paragraph') can't do word-timed captions. Word-timed means every word has timing: 'burnout' startMs=1836000, endMs=1837200, punch:true. The caption renderer builds 2–3 word cards (max 16 chars for 9:16, 34 for 16:9), leadIn 100ms, hold 150ms, min 400ms, baseline 0.74. Change from Clean to Space without re-transcribing — same words, new font, new accent, new tracking. The full-stream SRT/VTT download is just those words stitched: 1 → 00:00:01,200 --> 00:00:03,400. Same transcript powers clip suggestions, captions, and search.
Questions that decide the purchase — answered before you ask
Do failed minutes count? What if my internet drops at 11pm?
No. Only minutes that finish transcription count. If OBS keeps recording to disk while upload waits, those segments stay 'pending' and retry with backoff. Nothing is deleted until byte count + MD5 match R2. If Modal fails on a 1-minute chunk, that minute stays pending, you retry, and we don't bill until it succeeds. Already-transcribed minutes never count again — we check transcript_chunks before billing, so re-watching or re-searching tomorrow is free.
Can I download the full transcript as SRT/VTT and make clips later?
Yes. On any session with transcript, the FullStreamCaptions component stitches transcript_chunks into SRT (00:00:01,200 --> 00:00:03,400) and VTT (WEBVTT + 00:00:01.200 -->). Download the full .srt tonight, import to Descript or your editor, and come back next week to make 18 clips — no re-transcribe. That's the 'plus the clips later' you asked for: upload 24h once, download .srt now, clip later.
How is this different from Descript's transcript?
Descript edits video by editing text — brilliant for fine cuts. Weaver finds which text is worth editing. Upload 12h to Weaver, search 'pricing' in 300ms, get 12 hits with timecodes, bulk-create 12 edits, then open Descript with those timecodes. Use Weaver as the vault that ranks 20/h (max 200), Descript as the scalpel. Both count minutes, but Weaver counts each minute once and shares it across dual canvas.
What about a 24h stream-a-thon — do I need 24h of transcript?
No. Even Studio's 70h would handle one 24h fully, but you'd waste hours on dead air. Instead, schedule transcribe for the window that matters: last 6h where donations peak, via weaverclip_schedule_transcribe startMs=18*3600000, endMs=24*3600000. You spend 6h, save 18h. If you split 24h into two 12h sessions, each gets 20/h ranked moments (capped 200 total, so ~100 each if split) — better density than one 24h at 8.3/h.
Try it with the video you already have.
Free holds a full 2–3h webinar at 6 Mbps for 30 days. Upload the whole thing, download the SRT, make clips later — no re-transcribe.

