Geo multimodal tagger
GEO skills for OpenClaw, including site audit, llms-txt optimization, schema generation and more.
npx -y skills add geoly-ai/geo-skills --skill geo-multimodal-taggerAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
- 13 stars13 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Generate AI-optimized Alt Text, file names, captions, and Schema markup for images, videos, and audio assets. Improves AI discoverability on Google Lens, ChatGPT Vision, and Perplexity. Use whenever the user mentions optimizing images for AI, writing Alt Text, generating video Schema, tagging assets for AI discoverability, or making images visible in ChatGPT Vision and Google Lens.
SKILL.md
3.7 KB, 867 tokens by cl100k_base, as published. Nobody here has run it
Multimodal Asset Tagger
Methodology by GEOly AI (geoly.ai) — every image and video is a citation opportunity AI can either read or miss.
Generate optimized metadata for images, videos, and audio files for AI platforms.
Quick Start
python scripts/optimize_asset.py --type image --description "dashboard showing metrics" --output optimized.md
Why Multimodal Matters
AI platforms increasingly read visual content:
| Platform | Visual Capability | Citation Type |
|---|---|---|
| Google Lens | Image search | Direct image citation |
| ChatGPT Vision | Image understanding | Contextual reference |
| Perplexity | Video transcripts | Transcript citations |
| Gemini | Native image processing | Multimodal answers |
Image Optimization
Alt Text Formula
[Descriptive subject] + [Brand if relevant] + [Context/use case]
Examples:
❌ alt="image1.jpg"
❌ alt="product photo"
✅ alt="GEOly AI dashboard showing AIGVR score trend over 30 days"
✅ alt="Brand visibility comparison chart across ChatGPT and Perplexity — GEOly AI"
Filename Formula
[primary-keyword]-[secondary-keyword]-[brand]-[descriptor].jpg
Examples:
❌ IMG_3847.jpg
✅ geo-brand-visibility-dashboard-geoly-ai.png
✅ aigvr-score-chart-ai-search-monitoring.jpg
ImageObject Schema
{
"@context": "https://schema.org",
"@type": "ImageObject",
"name": "AIGVR Score Dashboard",
"description": "Dashboard showing brand visibility scores across AI platforms",
"contentUrl": "https://example.com/images/dashboard.jpg",
"author": {
"@type": "Organization",
"name": "GEOly AI"
},
"keywords": "AIGVR, brand visibility, AI search, dashboard"
}
Video Optimization
Checklist
- Title contains primary keyword
- Description: first 150 chars = keyword + brand
- Transcript/captions attached (SRT/VTT)
- Chapters/timestamps for long videos
- Thumbnail: keyword-rich filename
- VideoObject Schema added
VideoObject Schema
{
"@context": "https://schema.org",
"@type": "VideoObject",
"name": "How to Optimize for AI Search",
"description": "Complete guide to GEO strategies...",
"thumbnailUrl": "https://example.com/thumbs/geo-guide.jpg",
"uploadDate": "2024-01-15",
"duration": "PT12M30S",
"contentUrl": "https://example.com/videos/geo-guide.mp4"
}
Audio/Podcast Optimization
- Descriptive episode titles (not "Episode 47")
- 150+ word descriptions, keyword-rich
- Full transcript as page content
- Guest names and topics as entities
Asset Optimization Tool
python scripts/optimize_asset.py \
--type [image|video|audio] \
--description "Asset description" \
--brand "BrandName" \
--keywords "keyword1,keyword2"
Output:
- Optimized Alt Text
- Recommended filename
- Schema markup
- Discoverability score (Before/After)
Scoring
| Factor | Weight | Best Practice |
|---|---|---|
| Descriptiveness | 30% | Specific, detailed |
| Keyword presence | 25% | Natural inclusion |
| Brand mention | 20% | When relevant |
| Context | 15% | Use case clear |
| Length | 10% | 100-150 chars for Alt |
Discoverability Score: 0-10
- 8-10: Excellent
- 6-7: Good
- 4-5: Fair
- <4: Poor
Gives 0 of the 12 instructions most video audio skills give in 867 tokens
Counted across 621 of the 795 authors here whose files we hold, read 2026-08-06
- read individual rule files for detailed explanationsin 21 of 621, across 9 files
- Use WAV PCM 16kHz mono audio formatin 13 of 621, across 4 files
- render final videoin 13 of 621, across 6 files
- use this skill when dealing with Remotion codein 11 of 621, across 4 files
- save generated audio to a WAV filein 11 of 621, across 4 files
- handle conversion errors gracefullyin 10 of 621, across 6 files
- add captions to videos alwaysin 10 of 621, across 4 files
- generate music from text descriptions using MusicGenin 9 of 621, across 2 files
- do not skip pipeline layersin 9 of 621, across 3 files
- do not make one tool do everythingin 9 of 621, across 3 files
- never ask the user to paste their full API keyin 9 of 621, across 3 files
- use azure document intelligence for complex pdfsin 9 of 621, across 4 files
Said here and by no other author read
- combine subject brand and context for alt text
- format filenames as primary keyword secondary keyword brand descriptor
- include a primary keyword in video titles
- put keywords and brand in video description first 150 characters
- attach transcript or caption files to videos
- add timestamps or chapters to long videos
Grouped from the skills themselves: near-identical wordings counted once, and counted by distinct author, so one author publishing three of these counts once. Length counted with cl100k_base; the agent that loads this file may tokenize it differently.