How it works

One evidence timeline. Everything else is built on it.

Your browser does the heavy work on the video file. Cloudflare Workers AI transcribes and writes. A deterministic pass checks the result. Here is each step, in order.

  1. Your videoMP4 · MOV · WebM · audio

    Drop a file, paste captions, or reference a link for attribution. Nothing starts until you confirm you have the rights to it.

  2. Local prepin your browser

    The file is decoded on your device. It is never uploaded or stored — only what the next steps need leaves your machine.

  3. Audio · frames · OCRin your browser

    Audio becomes 30-second parts. Key frames are picked where the screen changes, and their text is read by OCR — slides, UI labels, numbers.

  4. TranscriptionWorkers AI · Whisper

    Each audio part is transcribed once with timestamps and language detection. A part is never transcribed — or paid for — twice.

  5. AI writingWorkers AI · Gemma

    Speech and screens are merged into ~45-second evidence windows, then written into an article in one structured call. Every paragraph must cite its windows.

  6. Articlechecked without AI

    A deterministic pass checks citations, verbatim quotes and numbers against the source, and points at the exact paragraph to review.

  7. Derived contentLinkedIn · X · newsletter

    From the reviewed article: LinkedIn post, X thread, newsletter section, YouTube description and quotes — same citations, one call.

Add a source

Upload a video or audio file (MP4, MOV, WebM, MP3, M4A, WAV), or paste captions or a timestamped transcript. Links from YouTube, Vimeo, Instagram or TikTok are kept for attribution and timestamp links; the media itself comes from your file. Duplicate files are detected so you never pay twice for the same video.

Prepare on your device

Your browser decodes the file locally (WebCodecs for MP4/MOV, Web Audio for other formats). The audio is resampled into 16 kHz, 30-second parts. Key frames are picked where the screen changes — a frame counts as a duplicate only if structure, colour and local-region signals all agree — with gap-filling so slow screen recordings still get illustrated.

On-screen text

Each key frame is read by OCR in the browser (Tesseract, served from our own domain). Lines below 70% confidence are dropped, and text repeated on most frames — toolbars, watermarks — is removed.

Transcribe

Each 30-second part is transcribed by Whisper large-v3-turbo on Cloudflare Workers AI, with timestamps and language detection. Parts are idempotent: re-sending a part (after a refresh or an interruption) returns the stored result without a new model call.

Build the evidence timeline

Speech and frames are merged into windows of about 45 seconds. Uncertain speech is flagged, never “fixed”.

Write

The article is written in one structured call to Gemma on Workers AI, under rules the model must follow: no fact, number or name that is not in the speech or on screen; every paragraph cites its windows; inferred context and editorial transitions are labelled. Identical requests reuse the existing article instead of generating it again.

Check

A deterministic quality pass — no AI — verifies citations exist, quotes appear word-for-word, numbers appear in the source, readability, redundancy and SEO basics. Problems point at the exact paragraph.

Edit and export

The editor keeps the evidence beside the text. AI rewrites of a section are re-grounded in the source and shown as a suggestion you accept or discard. Export to Markdown, HTML, WordPress blocks, plain text, JSON, or a .zip with images.

Where your data goes

Your browser

Your file stays with you

  • Video decoded locally (WebCodecs / Web Audio)
  • Key frames chosen on-device
  • On-screen text read on-device (Tesseract, self-hosted)
  • Only 30-s audio parts + up to 60 screenshots are sent
Cloudflare Workers AI

One model call per job step

  • Whisper large-v3-turbo transcribes each part once
  • Gemma writes the article in one structured call
  • Results cached by input: retries never pay twice
  • No third-party AI provider, no keys to manage
Your workspace

Evidence you can check

  • Transcript, screenshots and on-screen text kept for editing
  • Audio parts discarded after transcription
  • Export everything as JSON, delete anytime
  • Content is never used to train models
30 saudio parts, each transcribed exactly once
60key frames at most — chosen, not sampled
1AI call to write an article (formats: one more)
0AI calls in the quality check
Start with one video

Your next article is already recorded.

Drop in a demo, a webinar or a tutorial. Review the evidence, not a wall of AI text.

Free · 20 minutes of video · no card