Pick the source
Drop a file, paste captions, or reference a link for attribution. Nothing starts until you confirm you have the rights to it.
Your browser does the heavy work on the video file. Cloudflare Workers AI transcribes and writes. A deterministic pass checks the result. Here is each step, in order.
Drop a file, paste captions, or reference a link for attribution. Nothing starts until you confirm you have the rights to it.
The file is decoded on your device. It is never uploaded or stored — only what the next steps need leaves your machine.
Audio becomes 30-second parts. Key frames are picked where the screen changes, and their text is read by OCR — slides, UI labels, numbers.
Each audio part is transcribed once with timestamps and language detection. A part is never transcribed — or paid for — twice.
Speech and screens are merged into ~45-second evidence windows, then written into an article in one structured call. Every paragraph must cite its windows.
A deterministic pass checks citations, verbatim quotes and numbers against the source, and points at the exact paragraph to review.
From the reviewed article: LinkedIn post, X thread, newsletter section, YouTube description and quotes — same citations, one call.
Upload a video or audio file (MP4, MOV, WebM, MP3, M4A, WAV), or paste captions or a timestamped transcript. Links from YouTube, Vimeo, Instagram or TikTok are kept for attribution and timestamp links; the media itself comes from your file. Duplicate files are detected so you never pay twice for the same video.
Your browser decodes the file locally (WebCodecs for MP4/MOV, Web Audio for other formats). The audio is resampled into 16 kHz, 30-second parts. Key frames are picked where the screen changes — a frame counts as a duplicate only if structure, colour and local-region signals all agree — with gap-filling so slow screen recordings still get illustrated.
Each key frame is read by OCR in the browser (Tesseract, served from our own domain). Lines below 70% confidence are dropped, and text repeated on most frames — toolbars, watermarks — is removed.
Each 30-second part is transcribed by Whisper large-v3-turbo on Cloudflare Workers AI, with timestamps and language detection. Parts are idempotent: re-sending a part (after a refresh or an interruption) returns the stored result without a new model call.
Speech and frames are merged into windows of about 45 seconds. Uncertain speech is flagged, never “fixed”.
The article is written in one structured call to Gemma on Workers AI, under rules the model must follow: no fact, number or name that is not in the speech or on screen; every paragraph cites its windows; inferred context and editorial transitions are labelled. Identical requests reuse the existing article instead of generating it again.
A deterministic quality pass — no AI — verifies citations exist, quotes appear word-for-word, numbers appear in the source, readability, redundancy and SEO basics. Problems point at the exact paragraph.
The editor keeps the evidence beside the text. AI rewrites of a section are re-grounded in the source and shown as a suggestion you accept or discard. Export to Markdown, HTML, WordPress blocks, plain text, JSON, or a .zip with images.
Drop in a demo, a webinar or a tutorial. Review the evidence, not a wall of AI text.
Free · 20 minutes of video · no card