agentsclimarketplace

Overcast pinpoint

Skill kdr/overcast/skills/overcast-pinpoint

Video OSINT agent: senses + OSINT reach for any agent.

Install
npx -y skills add kdr/overcast --skill overcast-pinpoint

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 4 stars4 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Pinpoint WHEN a specific thing happens in a video — funnel from cheap coarse candidates (shots / CLIP / grid) to a frame-verified time window.

SKILL.md

3.2 KB, as published. Nobody here has run it

overcast-pinpoint

Use this skill to answer "exactly when does X happen?" in a clip and back it with evidence the model actually looked at. It mirrors the temporal-search pattern from the VLM video literature (T* / VideoAgent): score cheaply over the whole clip, then spend expensive VLM calls only on a few candidate frames and zoom in. Use the broad overcast skill and overcast/reference/verbs.md for exact flags.

Two rules that make the answer trustworthy:

  • Report a window, not a frame. Frame-exact localization is unreliable; emit [t1-t2] plus one verified keyframe.
  • Every timestamp must trace to a frame you see-verified. Never emit a time the model merely guessed — models answer correctly while grounding on the wrong moment, so confirm by looking at the frame.

Workflow

  1. Make the clip local and get a record id (see frame:// needs media on disk — capture a remote clip first). watch also gives per-shot timestamped content to search:
overcast doctor --json
overcast case init --json
overcast watch ./clip.mp4 --json         # -> video.analysis record id (REC)
  1. Get COARSE candidates cheaply (pick what's available):
overcast ask "moments where <X> happens, with timestamps" --json      # over watch shots/notes
overcast grid ./clip.mp4 --count 16 --json                            # one contact sheet ...
overcast see <montage-path> --prompt "which numbered cells show <X>? give cell numbers" --json
overcast similar search "<X>" --index <basic-clip-id> --json          # if a local CLIP index exists
overcast ask "moments <X> happens" --index <media-descriptions-id> --probe --json  # remote index

For grid, translate the chosen cell number to a time via the grid record's payload.cells[n].at (don't trust a model-guessed time). CLIP/shots only SHORTLIST — CLIP is weak on actions/order — so verify next.

  1. VERIFY + zoom on each candidate time T (expensive, precise):
overcast see frame://REC@T --prompt "Is <X> happening here? answer yes/no and what you see" --json
# refine: sample T-d and T+d, halve d each round until adjacent frames flip yes<->no
overcast see frame://REC@<T-2> --prompt "Is <X> happening?" --json
overcast see frame://REC@<T+2> --prompt "Is <X> happening?" --json
  1. Record the verified window and eyeball it:
overcast note "<X> occurs" --ref REC --at <t1-t2> --confidence medium --json
overcast view REC --at <t1-t2> --json
overcast brief --export ./pinpoint.md --json

Output

The tightest [t1-t2] window, the single see keyframe that confirms it (its record.id), and how it was found (shots / grid / CLIP / probe). Cite record.id + media.at for every claim; state the window, not a false-precise single frame.

Caveats

Needs the video local (a URL-only watch record can't extract frames). CLIP and shot text shortlist but don't decide — the see frame check does. If X recurs, pinpoint each occurrence separately; don't collapse them into one window.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.