agentsclimarketplace

Grok imagine video

Skill satasuk03/media-gen-skills/grok-imagine-video

Generate, edit, and extend videos with xAI's grok-imagine-video model. Use when the user wants to create a video from text, animate a still image, build a reference-driven video (virtual try-on, character consistency), edit an existing video, or extend one. Asynchronous — the script handles polling.From its SKILL.md

Install
npx -y skills add satasuk03/media-gen-skills --skill grok-imagine-video

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

6.5 KB, ~1.6k tokens by cl100k_base, as published. Nobody here has run it

grok-imagine-video — video generation, editing, and extension via xAI

Calls xAI's video API (grok-imagine-video). Five modes, all in scripts/video.py:

ModeWhen to useCLI
text-to-videoPure prompt → video--prompt "..."
image-to-videoAnimate a still image (image becomes the first frame)--prompt "..." --image PATH_OR_URL
reference-to-videoMake a video featuring people / clothing / objects from reference images (no first-frame lock)--prompt "..." --reference-image URL [URL ...]
edit-videoModify an existing video while preserving the rest--prompt "..." --edit-video URL
extend-videoContinue a video from its last frame--prompt "..." --extend-video URL

For full parameters, status flow, and gotchas, see references/api-reference.md. For Python recipes (xai-sdk, manual polling, async batches), see references/examples.md.

Prerequisites

  1. XAI_API_KEY exported in the environment.
  2. Python 3.9+ with stdlib only — no SDK required (the script uses urllib). If the user prefers the official SDK, pip install xai-sdk and use the snippets in references/examples.md.

Quickstart

# Text-to-video
python scripts/video.py \
  --prompt "Glowing crystal-powered rocket launching from red Mars dunes, ancient alien ruins lighting up" \
  --duration 10 --aspect-ratio 16:9 --resolution 720p \
  --output rocket.mp4

# Image-to-video (local file → animated)
python scripts/video.py \
  --prompt "Camera pushes in slowly. The leaves rustle in a gentle breeze." \
  --image ./still.jpg \
  --output animated.mp4

# Reference-to-video (virtual try-on / character consistency)
python scripts/video.py \
  --prompt "The model from <IMAGE_1> walks the runway wearing the shirt from <IMAGE_2>, slow motion, dramatic lighting" \
  --reference-image https://.../model.jpg https://.../shirt.jpg \
  --duration 10 --aspect-ratio 16:9 --resolution 720p \
  --output runway.mp4

# Edit an existing video (preserves duration / aspect / resolution)
python scripts/video.py \
  --prompt "Give the woman a silver necklace" \
  --edit-video https://.../source.mp4 \
  --output edited.mp4

# Extend an existing video (duration = length of NEW portion only)
python scripts/video.py \
  --prompt "The cat notices a butterfly and leaps off the windowsill" \
  --extend-video https://.../cat.mp4 --duration 6 \
  --output extended.mp4

The script polls until status=done, downloads the resulting MP4 to --output, and prints both the local path and the source URL on stderr.

Decision rules

When the user wants a video, route by which inputs they have:

  1. No inputs, just a description → text-to-video.
  2. One image, "make this move" → image-to-video. The image becomes the first frame; aspect ratio defaults to the image's unless overridden.
  3. One or more images, "use these people/clothes/objects" → reference-to-video. Refer to images in the prompt as <IMAGE_1>, <IMAGE_2>, etc.
  4. An existing video, "change X about it" → edit-video. Duration / aspect / resolution are inherited and cannot be overridden (capped at 720p).
  5. An existing video, "show what happens next" → extend-video. --duration is the length of the extension only; output total = original + extension.

Mutually exclusive — don't combine --image with --reference-image, and don't pass --edit-video with --extend-video. The script enforces this.

Parameters cheat sheet

FlagValuesDefaultNotes
--promptstrrequiredUse <IMAGE_1>, <IMAGE_2> markers when referencing images
--durationint 1–15 (seconds)model defaultFor extend, this is the added length only. Edit mode ignores.
--aspect-ratio1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:316:9Edit mode ignores; image-to-video defaults to image's ratio
--resolution480p, 720p480pEdit mode ignores (caps at 720p)
--imagepath or https:// URL or data: URIImage-to-video. Local files auto-encoded as base64 data URIs.
--reference-imageone or more URLs / pathsReference-to-video. Local files auto-encoded.
--edit-videoURLEdit-video mode. Must be an HTTPS URL (xAI-hosted is fine).
--extend-videoURLExtend-video mode. Source must be 2–15 s.
--poll-intervalseconds5How often to check status
--poll-timeoutseconds900 (15 min)Bail out if not done by then
--outputpathrequiredWhere to save the downloaded MP4
--keep-urlflagoffSkip the download, only print the xAI-hosted URL

Things to know

  • Asynchronous. Generation typically takes a couple minutes; long / 720p / edit jobs take longer. The script handles polling — don't introduce your own loop on top.
  • URLs are ephemeral. xAI returns a temporary hosted MP4 URL. The script downloads it immediately. Pass --keep-url only if you'll fetch within minutes.
  • Edit mode is constrained. Editing inherits the source video's duration, aspect, and resolution; passing those flags is silently ignored by the API. Editing input is also capped at 8.7 seconds.
  • Extend mode bookkeeping. --duration is the length of the new portion. A 10s source + --duration 5 → 15s output.
  • Mode collisions error out 400. Don't combine --image and --reference-image, and don't combine --edit-video and --extend-video.
  • Moderation. Output may be filtered. The response carries a respect_moderation flag; if the model declines the prompt, rewrite rather than retry verbatim.
  • Status enum. pending → keep polling, done → download, expired → request lifecycle ended (re-submit), failed → permanent failure (rewrite prompt or check inputs).

When to escalate

  • Need still images, not video → use the sibling grok-image or gpt-image-2 skills.
  • Need precise editing of specific regions / masks → not supported here; use gpt-image-2 for image masking and consider re-animating.
  • Need >15 s of fully novel content → chain extends (each adds up to 10s) or generate multiple clips and stitch them externally.

What ships with it: 3 files

20.4 KB alongside SKILL.md, 1 of them executable

references/

scripts/

Keep looking

Skills are one crate of 325,949. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.