
/watch (Claude Video)
github.com/bradautomates/claude-video- Category
- Developer Tools
- Rank
- No. 752Tools index
Previous survey · No. 794 ·
- Pricing
- Open Source
- Type
- TOOL
- Use case
- Data Processing
- Interfaces
- Agent Skill / Plugin
- Builder
- bradautomates
- GitHub
- 17.9k stars
- Latest release
- v0.3.2
- Date
About
A Claude Code skill/plugin that lets Claude 'watch' any video by downloading it, extracting frames with scene-aware sampling, pulling captions or a Whisper-generated transcript, and reading the frames as images to answer questions grounded in what's actually shown and said. Includes tunable detail modes and frame deduplication to manage token cost on long videos.
What it does
This is a Claude Code skill that gives Claude something it lacks by default: the ability to actually look at and listen to a video before answering a question about it. Point it at a link or a local file with what you want to know, and it pulls subtitles when they already exist, downloads only the media it needs, picks out a handful of representative moments instead of every frame, and transcribes the audio itself when there are no subtitles. Claude reads those moments as images alongside the transcript, so the answer reflects what actually happened on screen, not a guess from the title.
Why it's ranked here
The pitch is unusually well engineered for a single-purpose skill. The frame budget scales with video length instead of blindly sampling at a fixed rate, a visual near-duplicate check keeps a ninety-second static slide from burning the same tokens as ninety seconds of dense motion, and the same package installs cleanly across Claude Code and other agent hosts, not just one. Licensed MIT, with a documented fix that closed a subprocess argument-injection path in how it shells out to the downloader. That is the kind of detail that separates a maintained utility from a one-off script.
What's good
The engineering shows in details that matter for a tool that shells out to external binaries. The near-duplicate frame filter compares each candidate to the last frame actually kept, not the previous one, so it still catches a slow fade even though a simple frame-to-frame threshold would miss it. Frame counts are capped by minute of runtime rather than one flat number, so a thirty-second clip gets dense coverage and a fifty-minute one does not silently blow past a reasonable budget. Users can also pin exact timestamps a transcript calls out, so a moment the speaker flags by name is never sampled away by chance.
Tradeoffs
It leans on two external binaries and, for videos without captions, a paid speech-to-text API key, so a fresh install is not truly zero-effort: the automatic installer only covers macOS, and Linux or Windows users have to run printed commands themselves. The most thorough detail mode is explicitly uncapped, so a long, high-motion video can spend far more in image tokens than expected unless the user knows to ask for a narrower window first. Coverage of a long video is inherently sparse in every mode except that priciest one, which trades the sparsity for cost.
How to use it well
This fits best as an on-demand tool for a specific clip: debugging a screen recording someone sent you, checking what a competitor's demo actually shows, or pulling the real structure out of a talk instead of trusting its title. Start cheap, with captions only, and step up to frame sampling only when the question needs to see the screen. For anything over ten minutes, naming the section you care about gets far better answers than asking for the whole video, since full coverage past that length gets sparse by design. It is not a fit for exhaustive frame-by-frame motion analysis unless you accept the uncapped mode's cost.
Technical notes+
skills/watch/scripts/watch.py is the entry point: it loads settings via config.get_config(), resolves a frame cap from frame_cap(), then calls into download.py, frames.py, and transcribe.py before handing frame paths and a formatted transcript to Claude to read. frames.py hardcodes the operating limits: MAX_FPS=2.0, a SCENE_THRESHOLD of 0.20 for ffmpeg's scene-change filter, and a dedup pass (DEDUP_THUMB=16, DEDUP_THRESHOLD=2.0) that downscales each frame to a 16x16 grayscale thumbnail and drops it when its mean absolute pixel difference from the last kept frame is at or below the threshold; auto_fps() and auto_fps_focus() compute the target frame count from clip duration, and merge_frames() folds any --timestamps cue frames into the sampled set without letting them be evicted by the cap. download.py's is_url() rejects sources with an empty netloc or a leading hyphen, and both fetch_captions() and download_url() insert a literal -- before the URL in the yt-dlp argv, a documented fix in CHANGELOG.md against option injection. config.py reads ~/.config/watch/.env for WATCH_DETAIL and strips inline comments only when the # is preceded by whitespace, so an API key containing # is not truncated. transcribe.py's parse_vtt() regex-parses WebVTT cues and _dedupe() collapses YouTube's rolling-duplicate auto-caption lines. hooks/hooks.json registers a SessionStart hook that runs a setup-check script on launch. CHANGELOG.md also documents a pytest suite covering the script modules, though the test files themselves were not part of this read.
Observed
- License
- MIT licensed
- Packaging
- Packaged as a Claude Code plugin/skill, and separately installable into other Agent-Skills-compatible hosts such as Codex, Cursor, Copilot, and Gemini CLI via a dedicated installer CLI
- Interfaces
- Ships as a Python command-line script invoked from within the skill, not exposed as an MCP server or hosted API
- Dependencies
- Requires the external binaries yt-dlp and ffmpeg on the host machine; the bundled installer auto-installs both via Homebrew on macOS and prints manual install commands on Linux and Windows
- Speech-to-text
- Speech-to-text fallback supports two interchangeable backends, Groq and OpenAI, selected by which API key is configured
- Frame processing
- Frame near-duplicate detection uses only the Python standard library after an initial ffmpeg downscale step, with no image-processing library dependency
- Lifecycle hook
- Registers a SessionStart hook that runs a setup-check script on session launch
Read from README.md, .claude-plugin/plugin.json, skills/watch/SKILL.md, skills/watch/scripts/watch.py, skills/watch/scripts/download.py, skills/watch/scripts/config.py, skills/watch/scripts/frames.py, skills/watch/scripts/transcribe.py, hooks/hooks.json, CHANGELOG.md.
What it can do
Download a video from a URL
Video URL → Video file
Extract frames from video using scene-aware sampling
Video file → Set of image frames
Deduplicate extracted frames to reduce token cost
Set of image frames → Reduced set of unique frames
Fetch existing captions or generate a transcript via Whisper
Video file or audio track → Text transcript/captions
Answer questions about a video grounded in its visual frames and transcript
Video frames and transcript, user question → Answer text
Tags
Comments (0)
No comments yet
Editorially curated, with community endorsements as a secondary signal. Corrections welcome.