Video → PDF → AI
ChatGPT and Claude do not watch video. ClipContext turns any recording into a PDF with the frames that matter and a synced transcript — ready to attach to the conversation.
No sign-up, no install, no size limit. The video is processed inside your browser and is never uploaded to any server.
Language models read text and images — not video. Either they reject the file, or its size makes uploading impractical. And when you take screenshots by hand, you end up sending ten near-identical screens and none of the moment that mattered, with nothing linking what is shown to what was being said.
How it works
Drag the file onto the page. It is read straight from your disk — 300 MB or 3 GB makes no difference, because nothing is uploaded anywhere.
Transcribe it right there, with Whisper running on your own machine, or drag in a .vtt / .srt caption file you already have.
Review the frames, discard the bad ones and download the document. Along with it comes the text that tells the model how to read that PDF — that is what changes the quality of the answer.
Instead of grabbing a frame every ten seconds, it scans the video and keeps the moments where the screen actually changes. In a presentation, that becomes one frame per slide instead of forty repeats.
In the PDF, each frame appears next to what is said at that point, with the timestamp marked. The model can connect what was on screen with what was being said.
There is no server. The video, the audio and the transcript are processed in your browser — including confidential meeting recordings.
The ready-made text tells the model to cite timestamps, separate what it saw from what it heard, and admit when the information is not in the document instead of inventing it.
Privacy
Most online video tools require an upload: your file goes to someone's machine, is processed there and stays stored for a period you do not control. ClipContext does none of that — the page is a static file, and all processing happens in your browser.
There is no account, no database, no tracking. The only external connection is downloading the transcription model, and only if you use automatic transcription. Read the full policy.
No. It opens in the browser and works. An up-to-date Chrome or Edge gives the best experience, because automatic transcription uses the graphics card through WebGPU.
There is no limit imposed by us, precisely because there is no upload. The practical limit is your computer's memory. Long videos take longer to scan, and you can stop the scan whenever you want.
MP4 (H.264), WebM and MOV in most cases. Formats the browser cannot decode — HEVC and some MKV files — will not open; those need converting to MP4 first.
It uses Whisper running locally. The fast model is reasonable and the accurate one is considerably better, at the cost of a larger download. If you already have captions from another service, drag the file in and skip this step.
The tool as it stands today is free and the code is open under the MIT licence. Since there is no server cost, there is no reason to charge for it.
Any model that accepts attachments and reads images — Claude, ChatGPT and Gemini all work. Attach the PDF and paste the prompt the tool generates alongside it.