From Watch Skill
Watches any video (URL, stream, or local path) by downloading, extracting scene-aware deduped frames, OCR, transcription, and indexing. Answers follow-up questions from a persistent index without re-processing.
How this skill is triggered — by the user, by Claude, or both
Slash command
/watch-skill:watch <video-url-or-path> [question]<video-url-or-path> [question]This skill is limited to the following tools:
The summary Claude sees in its skill listing — used to decide when to auto-load this skill
You don't have a video input; this skill gives you one. It is a thin wrapper
You don't have a video input; this skill gives you one. It is a thin wrapper
around the watch-skill CLI — all logic lives in the engine, so this skill
works identically on every harness (Claude Code, Codex, Cursor, ...).
This is a drop-in upgrade of the classic claude-video /watch skill:
same invocation shape, plus a persistent index (ask answers follow-ups
without re-processing), OCR on frames, scene-aware sampling with perceptual
dedup, local Whisper (offline by default, no API key needed), and THE LOOP
(capture -> critique -> fix -> re-capture) for iterating on your own output.
watch-skill doctor --json
fix. doctor
auto-bootstraps ffmpeg and yt-dlp into a managed bin dir on Windows/macOS/
Linux; re-run once after it reports fixes. Only involve the user when a
check still fails after remediation.watch-skill itself is not on PATH: pip install watch-skill (or
uv tool install watch-skill), then re-run the doctor.No API key is required for acquisition, transcription, OCR, indexing, or
search: transcription falls back to local faster-whisper. Visual synthesis
and verification can use the user's existing Anthropic, OpenAI, Gemini, or
OpenRouter key, or an optional local Ollama model. The agent and provider are
independent; see the configuring-vision skill. Cloud STT is opt-in
(--cloud-stt) and only ever uploads extracted mono audio — the video file
never leaves the machine.
Parse the user input into source + optional question, then:
watch-skill watch "<source>" [--start T --end T] [--max-frames N] [--transcript-only]
--duration 60 bounds live streams), and local files all work.--start / --end (SS, MM:SS, HH:MM:SS) switch to dense focused
sampling of that window — use for "what happens at 2:30?" questions and for
any video over ~10 minutes when the user cares about one section.--timestamps T1,T2,... pins frames at transcript-flagged moments
("look here", "as you can see") that visual selection may miss.--transcript-only skips frames entirely (fastest; no video download when
captions exist).--max-frames N tightens the token budget (default: duration-tiered,
hard cap 100, max 2 fps).The report prints an Indexed: video_id ... line, frames with t=MM:SS
timestamps, OCR text, and the transcript.
Read every frame path the report lists, in a single message (parallel Read calls), so you see them together in chronological order.
Answer from frames + OCR + transcript, citing timestamps. No question → summarize structure, key moments, notable visuals, spoken content.
The video is already indexed. For any follow-up question in this or a LATER session:
watch-skill ask <video_id> "<question>" # self-healing answer + evidence
watch-skill search "<phrase>" # across every video ever watched
ask (v0.6) answers text-first with timestamped evidence, a confidence
score, and a ~N tokens saved line. It escalates on its own when unsure
(dense re-sampling, zoom-crop re-OCR) and states plainly when the video
does not clearly show the answer — trust that refusal; do NOT invent an
answer past it. Frame paths are listed only when the engine wants you to
look yourself (or pass --frames); Read them then. Never re-run watch
for a follow-up on an already-indexed video.
If the user corrects one of your video answers, report it so the system learns (locally):
watch-skill lessons add <video_id> "<question>" "<your wrong answer>" "<the correction>"
When the user asks you to fix UI/visual output and verify the fix:
watch-skill loop start "<url-or-screen:-or-file>" "<pass criteria>" [--script '<json steps>']
# ... you apply the suggested fixes ...
watch-skill loop iterate <loop_id>
The critique returns structured issues with timestamps and suggested fixes.
YOU change the code; the loop only observes. On pass it renders a
before/after MP4+GIF proof. watch-skill capture "<target>" records without
critiquing.
--cloud-stt..env; they are never logged or echoed.~/.watch-skill/cache (LRU, size-capped).npx claudepluginhub oxbshw/watch-skill --plugin watch-skillDownloads videos, extracts frames and transcripts, and lets Claude answer questions about video content.
Downloads and analyzes videos from URLs (YouTube, TikTok, streams) or local files: extracts scene-aware frames, OCRs text, transcribes audio with offline Whisper. Answers questions about video content without guessing.
Transcribes and analyzes any video (YouTube, Loom, Vimeo, Zoom, local files) with three depth modes: transcript, visual (with frame extraction and Claude vision), and multimodal (Gemini native video).