agentsclimarketplace

Media gen orchestrator

Skill Sheshiyer/skill-clusters/skills/media-gen-orchestrator

Hub-and-spoke agent-skill clusters, one per stack (Astro·GSAP·Remotion, Tauri, …). Installable via skills.sh.

Install
npx -y skills add Sheshiyer/skill-clusters --skill media-gen-orchestrator

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Route a media-generation task to the right tool among 11 specialists — image generation (GPT Image 2, Nano Banana 2, OpenAI Images), image enhancement, image-to-3D, AI video direction, FFmpeg processing, frame extraction, and GIF search/authoring. USE WHEN a user wants to generate, edit, upscale, convert, or assemble an image, 3D model, video, or GIF but hasn't named the specific tool or backend.

SKILL.md

5.0 KB, ~1.2k tokens by cl100k_base, as published. Nobody here has run it

Media-Gen Orchestrator

The single entry skill for generative + transformational media work. It locates the task on the modality × stage map — what asset (image / 3D / video / GIF) and what stage (generate → transform → assemble) — and delegates to one of 11 specialist spokes. The cross-cutting decision every request shares — which backend to route to, given its access/billing model and the modality — lives in media-gen-core; read it before picking an image or video generator.

Routing map (intent → spoke)

Generate an image

  • Via your ChatGPT Plus/Pro plan, no per-image billing → gpt-image-2
  • Via Google Gemini 3.1 Flash Image (multi-image input, Search grounding) → nano-banana-2
  • Batch generation against the OpenAI Images API (+ gallery) → openai-image-gen
  • Headers, icons, diagrams, mermaid, infographics (a full visual-content system) → art

Transform / clean up an image

  • Upscale, sharpen, de-noise screenshots & assets → image-enhancer
  • Restyle / edit an existing image (image-to-image) → gpt-image-2, nano-banana-2 (both edit, not just generate)

Make a 3D asset

  • Image(s) → textured mesh (GLB/OBJ/PLY/STL) via Hunyuan3D-2.1 → hunyuan3d

Generate video

  • Shot-by-shot AI video (Kling / Sora / Grok), character-consistent series → ai-video-director

Process / assemble video

  • Convert, resize, compress, extract audio, GIF↔MP4 → ffmpeg
  • Pull frames or short clips out of a video → video-frames

Work with GIFs

  • Search GIF providers, download, extract stills/sheets → gifgrep
  • Author a size-constrained animated GIF for Slack → slack-gif-creator

Picked-up spokes

Additional image-generation spokes (picked up from antigravity-awesome-skills, MIT). Route to these alongside the generators above:

  • Standalone Gemini image CLI — generate an image from a prompt with only a GEMINI_API_KEY and stdlib Python, no MCP server (good for UI placeholders, docs illustrations, icons) → imagen
  • SEO-optimized image assets — OG/social cards, blog heroes, schema ImageObject visuals, product shots, infographics, each with per-use-case aspect ratio + resolution and a post-gen SEO checklist (alt text, WebP, schema, og:image tags) → seo-image-gen

Routing note: prefer seo-image-gen when the image is destined for a page/social card and must be search/social optimized; prefer imagen for a quick subscription-free local generate; prefer the plan-backed generators (gpt-image-2, nano-banana-2, openai-image-gen) when their specific backend/capability (multi-image input, character consistency, batch gallery) is required.

Standard Operating Flow

  1. Locate the task: which modality (image / 3D / video / GIF) and which stage (generate → transform → assemble).
  2. If it generates an image or video, pull the routing model from media-gen-core first — the right backend depends on access (subscription vs API key vs local GPU/HF), cost, and the specific capability needed (multi-image input, character consistency, format).
  3. Delegate to the spoke(s). Multi-step asks fan out in modality order — generators produce assets, then processors transform/assemble them. E.g. "make an animated logo intro"art/gpt-image-2 (frames) → ai-video-director or ffmpeg (motion) → ffmpeg (encode).
  4. Return: chosen spoke(s), the backend/model selected and why (access + capability), the output format, and the next action.

Guardrails

See media-gen-core. In short: confirm the backend's prerequisite before generating — a subscription-backed CLI (gpt-image-2), an API key + billing (openai-image-gen, paid image APIs), or a local/hosted GPU (hunyuan3d via HF Gradio) are not interchangeable, and silently picking the wrong one wastes a paid call or fails. Generation is non-deterministic and may incur cost or rate limits — state the cost/quota implication before a batch. Respect provider content policies and the license/attribution of any downloaded GIF or reference image. Keep ffmpeg/ video-frames re-encodes lossless-aware (don't silently transcode a master).

Loading spokes on demand

To keep CLI startup context lean, this cluster's spokes are not separately registered as skills — only this orchestrator and its *-core are enumerated. When you route to a spoke named above, load it on demand by reading its file:

~/.agents/skill-clusters/skills/<spoke-name>/SKILL.md (or skills/<spoke-name>/SKILL.md inside the skill-clusters repo).

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.