Visual prompt
Skill NovateStudioGit/novate-studio-skills/creative-production/visual-prompt
57 agent skills for Claude Code — creative production, paid growth, copywriting, ecommerce, email marketing & knowledge ops. By Novate Studio.
npx -y skills add NovateStudioGit/novate-studio-skills --skill visual-promptAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Expand a short visual brief into a Shlabu-density, paste-ready prompt for any modern image or video generation tool — Kling, Veo, Sora, Seedance, Higgsfield, Flux, GPT Image, Nano Banana, or Midjourney. Trigger this skill in TWO modes. (1) Explicit — the user types `/visual-prompt` or natural-language like "expand this prompt", "make this prompt detailed", "give me the long version", "Shlabu this", "write the detailed prompt for [scene]". (2) Auto-detect — the user's last message LOOKS like a raw visual brief: a short declarative scene description (1–4 sentences), photographic/cinematic vocabulary like "shot of", "POV", "angle", "frame", "lit by", "hero shot", a subject + setting + intent shape, and no other clear request type (not a question, not a feedback note, not a code request, not a meta-conversation). When triggered in auto-detect mode, do NOT compose the prompt yet — first respond with a one-line gate asking Will to confirm by typing `/visual-prompt`, then wait. Only compose after explicit confirmation. Output is a single comma-spliced prose paragraph at Shlabu density (250–350 words) for prose tools, or a comma-separated keyword + parameter line for Midjourney.
SKILL.md
12.9 KB, ~2.9k tokens by cl100k_base, as published. Nobody here has run it
Visual Prompt
Takes a short brief from the user and returns a paste-ready visual generation prompt at the level of detail that actually pulls 4K-quality realism out of modern image/video tools — applying the 9 Shlabu craft patterns from skills/references/shlabu-realism.md.
The output is a prompt, not media. Print it to chat in a fenced block. Save only if Will asks.
Step 0 — Triggering and the confirmation gate
This skill fires in two modes. The first thing to do every time is figure out which.
Mode A — Explicit invocation. Will typed /visual-prompt (with or without a brief on the same line), OR a natural-language phrase that obviously asks for prompt expansion ("expand this for Kling", "give me the long version of [X]", "Shlabu this", "make this prompt detailed").
→ Skip ahead to Step 1. Compose immediately.
Mode B — Auto-detect. Will's last message wasn't an explicit invocation, but it looks like a raw visual brief — short, declarative, photographic/cinematic vocabulary, subject + scene + intent shape, no other clear request type. Examples that should trigger Mode B:
- "Gorgeous blonde girl aged 27 laying on her bed, white bed sheets, white tee. POV shot of her."
- "Hero shot of the watch face slowly rotating on a black velvet pad, single key light."
- "Marcus walking out of a record shop in Tokyo at night, neon, light rain on the pavement."
- "Close-up of espresso pulling into a cup, warm overhead light, steam."
→ Do not compose yet. Reply with one line:
Looks like a visual brief — type
/visual-promptand I'll build it out. (or sayimage/videoto lock the medium.)
Then stop. Wait. The next message from Will is the confirmation. When he sends /visual-prompt (alone, or with corrections), carry forward the previous brief as the input and proceed to Step 1.
When NOT to trigger Mode B — leave these alone, don't gate them:
- Questions ("how should I shoot this?")
- Feedback on an existing prompt or image ("the lighting feels off")
- Conversation about a brief ("what tool should we use?")
- Code, file, or system requests
- Single-word replies ("yes", "go", "image", etc — these are usually responses, not new briefs)
If you're not sure whether something is a brief or just brainstorming, lean toward gating in Mode B — the cost of a one-line confirmation question is much smaller than the cost of generating a 300-word prompt for something that wasn't a brief.
Step 1 — Receive the brief
The brief is whatever Will gives you after /visual-prompt or in natural language. Free-form, 1–3 sentences typically. Examples:
- "blonde woman in a Paris café at golden hour, looking at her phone, raw iPhone selfie energy"
- "hero shot of the watch face slowly rotating on a black velvet pad, single key light"
- "young surfer walking out of the ocean at dusk, board under arm, wet hair, salt on the skin"
If the brief is missing a critical anchor (no subject, no scene, or no intent at all), ask one tight clarifying question. Otherwise just go.
Step 2 — Detect optional inputs
Scan the brief for these. Don't ask if not mentioned — defaults are fine.
- Target tool — keywords like "for Kling", "Veo prompt", "Midjourney version". Default:
genericlong prose. - Medium — "image", "video", "motion", "selfie". If unclear, infer from the verbs (a "shot of X rotating" is video; a "portrait of X" is image). When still ambiguous, ask once.
- Model anchor — a name like "using ExampleBrand-male-01" or "with jade". If present, read
brands/[brand-name]/models/[name]/characteristics.mdand pull hair / eye / skin / freckles / build into prose anchors. Pick the brand by matching the model name acrossbrands/*/models/*/. - Dialog (video only) — quoted lines like "she says 'hey there'". Pull the line and any delivery notes the user provided.
Step 3 — Read the realism reference
Read skills/references/shlabu-realism.md before composing. The 9 patterns described there are the spec for this skill — apply them mechanically. Don't paraphrase from memory; the file is the source of truth.
Step 4 — Compose the prompt
Pick one of two grammars based on target tool:
Grammar A — Long prose (default for everything except Midjourney)
Used for: kling, veo, sora, seedance, higgsfield, flux, gpt-image, nano-banana, generic.
Format:
- Single paragraph, 250–350 words for image tools
- Comma-spliced with em-dashes (
—), not bullets, not headers, not labels - No "Subject:" "Camera:" "Lighting:" labels — write it as continuous prose
- Reads like screenplay direction, not a fill-in-the-blank form
Tool-specific length caps — these are hard limits, not soft suggestions:
| Tool | Hard cap | Implied word ceiling |
|---|---|---|
| Kling 3.0 (FAL) | 2500 characters | ~380 words (verified — anything over this gets a 422 from FAL) |
| Veo 3.1 | ~2000 characters (recommended) | ~320 words |
| Sora | ~2000 characters | ~320 words |
| Seedance, Higgsfield, Flux, GPT Image, Nano Banana | no documented hard cap | 250–350 words |
When the target tool is kling, count characters before producing the final prompt and trim aggressively if over 2500. The 422 error from FAL costs Will an iteration cycle, not money — but it's still friction worth avoiding. Cuts to make first when over-cap: redundant scene-setting (the start frame already shows it), repeated anti-synthetic phrases (4× is enough, not 6×), trailing colour-grade caveats. Preserve dialog, delivery direction, camera-as-choreography, and high-frequency detail callouts — those are load-bearing.
Mechanical checklist (every prompt must hit all that apply):
- Anti-synthetic vocabulary repeated 4–6× — handheld, organic micro-shake, loose, unhurried, unforced, not performing, just messing around on her phone, completely unscripted, casual grip. Sprinkle throughout, not all in one place.
- High-frequency detail callouts — visible pores, scattered freckles, baby hairs along the hairline, dewy skin, individual strands of hair catching the light, soft flush of natural color, fine skin texture. At least 3 per prompt.
- Camera as verb-beat choreography (video only) — "camera held loosely at arm's length → drifts closer → tightens to cheek-macro until skin texture dominates → holds for a beat → slowly pulls back." Beats with verbs, not shot names like "closeup → medium."
- Hair-as-light-catcher — "soft diffused light catching the individual strands of her hair, each strand picking up its own highlight." One mention.
- Background described as soft / blurred — "city blurred softly in the background bokeh", "white linen soft in the background", "pedestrians blurred behind the glass." Tells the tool not to spend detail budget there.
- One subject, one location, one continuous moment. No montages, no scene cuts, no second character. If Will's brief implies a montage, collapse it to the single most evocative beat.
- Quoted dialog with delivery direction inline (video only, if dialog present) —
"line" — tone, gesture, beat. Multiple short lines beat one long monologue. - Re-pin character anchors — even when the model anchor is set, restate hair colour / eye colour / freckles in the prose. Belt-and-suspenders against drift.
- Single comma-spliced paragraph with em-dashes. No bullets, no headers, no labels. One continuous run.
Skip patterns 1, 2 (skin/hair detail), 3, 4 when the brief is for a clean studio packshot, product rotation, or commercial campaign hero — those want clean directed cinematography, not handheld organic chaos. Use patterns 5, 6, 9 always.
Grammar B — Midjourney
Used only when target tool is midjourney.
Format:
- Comma-separated keyword phrases on a single line
- Append parameters:
--ar [aspect]--v 6.1(or--v 7if newer)--style raw(for photoreal) - Use
--sref [url]if a style reference image is mentioned - Do NOT produce 300-word prose
- Front-load the most distinctive nouns; descriptors trail
- Example shape:
blonde woman in Paris café, golden hour through window, raw iPhone selfie, freckles, dewy skin, soft window light catching individual hair strands, blurred street bokeh, candid unposed, photographic authenticity --ar 9:16 --v 6.1 --style raw
Step 5 — Output
Print the prompt to chat as plain visible prose — not in a code fence, not blockquoted, not collapsed. Will needs to see and read the prose directly in the message, not click-to-expand a code block.
Format:
**Prompt:**
[the full prose paragraph here, written out in full, no truncation]
[ generic prose · 287 words · image ]
Rules:
- Always write the entire prompt in full — no "..." truncation, no "see above," no referencing a previous message. Every run produces a complete written-out prompt, every time.
- Lead with
**Prompt:**on its own line so it's visually anchored. - One blank line, then the prose paragraph.
- One blank line, then the meta line in square brackets:
[ {tool grammar} · {word count} words · {medium} ]. - For Midjourney (Grammar B), the keyword + parameter line replaces the prose paragraph but the structure is the same.
After the meta line, stop. Do not ask whether to save. Do not offer follow-up options. The skill ends with the prompt printed in chat — Will copies it from there. If he wants to save it later, he'll ask explicitly with something like "save that" or "save it under [name]", which routes to Step 6.
Step 6 — Optional save
Only on user request. Save to:
brands/[brand-name]/visual-prompts/[slug]/
prompt.txt the prose / keyword line
brief.json the original input + detected options
Slug derivation: short kebab-case from the most distinctive nouns in the brief. e.g. blonde-paris-cafe-selfie, watch-rotation-velvet. Ask Will to confirm or rename if ambiguous.
brief.json shape:
{
"brief": "blonde woman in a Paris café at golden hour, looking at her phone, raw iPhone selfie energy",
"target_tool": "generic",
"medium": "video",
"model_anchor": null,
"dialog": null,
"word_count": 287,
"created": "2026-05-09T22:14:00Z"
}
If no brand is active in the project (no brands/*/brand-identity/visual-guidelines.md), save to visual-prompts/[slug]/ at project root instead, and tell Will.
Notes
- Density is the whole point. A 90-word prompt that "covers" the brief is failure. The skill exists because hand-writing 300-word Shlabu-density prose is hard and slow — if the output is short, you've defeated the purpose.
- No headers, no bullets, no JSON inside the prompt. Tools like Kling and Veo parse continuous prose dramatically better than structured input. Resist the urge to organise.
- Don't fabricate facts about the subject. If Will doesn't specify hair colour, don't invent one — use generic phrasing ("hair catching the window light") instead. Per his standing rule, never fabricate when output is paste-ready.
- Re-running with tweaks — if Will says "do it again but with [change]", take the previous brief, apply the change, regenerate. Don't ask the full set of inputs again.
- Skip the realism patterns when Will's brief is clearly for a studio packshot, product rotation, ghost mannequin, or other commercial-campaign-clean shot. Anti-synthetic vocab and pore callouts pull those renders in the wrong direction. Use patterns 5, 6, 9 always; the rest are conditional on selfie/UGC/handheld POV intent.
- Source for the 9 patterns:
skills/references/shlabu-realism.md(distilled from Shlabu's "Kling 4K Prompts" tutorial). Always read that file before composing — don't compose from memory.