Elevenlabs v3 text optimization
Skill msdanyg/elevenlabs-v3-text-optimization/elevenlabs-v3-text-optimization
Free Claude Skill that rewrites any text into ElevenLabs V3-optimized scripts — audio tags, pacing, phonetic respelling & acronym/number normalization for better AI voice synthesis.
npx -y skills add msdanyg/elevenlabs-v3-text-optimization --skill elevenlabs-v3-text-optimizationAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 3 stars3 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Transforms raw text into ElevenLabs V3 Studio-optimized scripts using audio tags, punctuation engineering, creative spelling and contextual priming. Automatically analyzes document type, emotional context and delivery intent to apply V3-specific formatting techniques that improve voice synthesis quality.
SKILL.md
23.6 KB, as published. Nobody here has run it
ElevenLabs V3 Text Optimization
What This Skill Does
Transforms raw text into ElevenLabs V3 Studio-optimized scripts using audio tags, punctuation engineering, creative spelling and contextual priming. This skill automatically analyzes document type, emotional context and delivery intent to apply V3-specific formatting techniques that improve voice synthesis quality.
Key capabilities:
- Adds V3 audio tags for emotional direction and pacing control
- Applies creative spelling to fix pronunciation issues
- Normalizes numbers, dates, acronyms and technical terms
- Engineers punctuation for optimal rhythm and breathing
- Removes incompatible SSML tags (V3 doesn't support them)
- Adapts strategy based on content type (dialogue, narration, technical, dramatic)
When to Use This Skill
Trigger phrases:
- "Optimize for ElevenLabs V3"
- "Make this V3-ready"
- "Fix this for V3 Studio"
- "Convert to V3 format"
- "Add V3 audio tags"
- "Improve this for voice synthesis"
- "Prepare this for TTS"
Use this skill when:
- User provides text intended for ElevenLabs V3 voice generation
- User mentions ElevenLabs, V3, TTS or voice synthesis
- User asks to improve pacing, pronunciation or emotional delivery
- User shares scripts needing audio enhancement
- User wants to fix pronunciation of names, brands or technical terms
Skip this skill when:
- User asks about V2 models or SSML formatting specifically
- User wants general writing advice unrelated to voice synthesis
- User requests document creation without voice/audio context
- User is working with non-ElevenLabs TTS systems
V3-Specific Technical Knowledge
Critical V3 Constraints
V3 DOES NOT support:
- SSML
<break time="1.0s" />tags - SSML
<phoneme>tags - Any XML-style markup
V3 DOES support:
- Audio tags in square brackets:
[excited],[pause],[whispers] - Punctuation-based pacing control
- Creative spelling for pronunciation
- Natural language pause tags:
[short pause],[long pause]
Audio Tags Reference
Emotional States:
[excited], [angry], [sad], [nervous], [cheerful], [calm], [frustrated], [tired], [sorrowful], [tense], [curious]
Volume/Intensity:
[whispers], [shouting], [quietly], [loudly]
Performance Modifiers:
[rushed], [slowly], [dramatic tone], [narrator]
Reactions/Non-verbals:
[laughs], [giggle], [sighs], [clears throat], [gasps], [gulps], [exhales]
Pacing:
[pause], [short pause], [long pause]
Character/Voice:
[British accent], [American accent], [deep voice], [feminine tone]
Tag Placement Rules:
- Place immediately before the phrase it should affect
- Use one strong tag instead of multiple weak ones
- Avoid contradictory tags back-to-back
- Refresh emotional tags to sustain effect over long passages
- No "closing tags" required
Punctuation Engineering
Comma (,): Micro-pause (~200ms), continuing intonation. Use to separate clauses.
Period (.): Full stop (~500ms), pitch resolution. Essential for ending thoughts.
Ellipsis (...): Hesitation pause (~500-800ms). Triggers "trailing off" effect, vocal fry or breathiness.
Em-dash (—): Abrupt break or interruption. Simulates sudden thought change.
Double dash ( -- ): Significant separation longer than period without trailing-off prosody. Resets pacing.
Question mark (?): Forces upspeak (rising intonation). Makes statements sound uncertain.
Exclamation mark (!): Increases loudness and projection.
No punctuation: Creates "stream of consciousness" or panicked effect where words run together.
Creative Spelling Strategies
Consonant Hardening:
- Replace ambiguous "c" with "k" or "s"
- Replace ambiguous "g" with hard sound
- Examples: "Celtic" → "Seltic" or "Keltic"
Vowel Disambiguation:
- Long A: use "ay" → "Data" becomes "Day-tuh"
- Long I: use "eye" or "y" → "Finance" becomes "Fye-nance"
- Long E: use "ee" → "Kubernetes" becomes "Koober-net-eez"
Syllabic Segmentation:
- Use hyphens to force slower articulation
- Example: "Revoicer" → "Ree-voy-ser"
- Example: "ActivTrak" → "Ak-tiv-trak"
Stress Control:
- Capitalize stressed syllables or words
- Verb: "pro-JECT" vs Noun: "PRO-ject"
- Emphasis: "I did NOT say that"
Common Problem Words:
| Word/Type | Issue | V3 Solution |
|---|---|---|
| Data | Dah-ta vs Day-ta | Day-tuh or Dah-tuh |
| Linux | Lye-nux vs Lin-ucks | Lin-ucks |
| SQL | Sequel vs S-Q-L | Sequel or Ess-Cue-Ell |
| Siobhan | Unknown name | Shiv-awn |
| Nguyen | Unknown name | Win |
| Live (broadcast) | Heteronym | Lyve |
| Live (exist) | Heteronym | Liv |
| Read (past) | Heteronym | Red |
| SaaS | Acronym | Sass or S A A S |
| API | Acronym | A P I (with spaces) |
Text Normalization Rules
Always expand in the text itself:
Numbers:
- "1998" → "nineteen ninety-eight"
- "2026" → "twenty twenty-six"
- "$50" → "fifty dollars"
- "45.67" → "forty-five point six seven"
Phone Numbers:
- "555-1234" → "five five five, one two three four"
Dates:
- "1/15/26" → "January fifteenth, twenty twenty-six"
- "Q3 2025" → "Q three, twenty twenty-five"
URLs:
- "elevenlabs.io" → "eleven labs dot io"
- "activtrak.com" → "activ trak dot com"
Acronyms:
- Letter-by-letter: Add spaces "SOC 2" → "S O C 2"
- Or expand once: "SOC 2 (ess oh see two)"
- Word-like: "NASA" → "Nasa" or "NAY-sa"
Measurements:
- "5km" → "five kilometers"
- "100GB" → "one hundred gigabytes"
Contextual Priming Technique
For short lines lacking emotional context, add "ghost text" that sets the scene, generate audio, then trim the ghost text in post-production.
Example:
Raw: "Open the door."
Primed for generation:
[scared] He backed away into the corner, his voice trembling with absolute terror as he whispered: [whispers] Open the door.
Post-production: User trims everything except "Open the door" which now carries the embedded terror from the priming context.
When to use:
- Short lines (Yes, No, What, single sentences)
- Lines where emotional intent isn't clear from surrounding text
- Situations where tone is critical but context is minimal
Document Type Optimization Strategies
Explainer/Educational Content
Goal: Clear, measured delivery with strategic pauses for comprehension.
Techniques:
- One idea per paragraph
- Use
[pause]or[short pause]between key concepts - Add emphasis through capitalization sparingly
- Maintain
[calm]or[narrator]tone - Break complex sentences with commas
Example:
[calm] Today we're covering three key areas: value, workflow and rollout.
First, value. [short pause] What changes for customers is speed and clarity.
Second, workflow. [short pause] We reduce manual steps and standardize the process.
Third, rollout. [pause] We start with a pilot, then expand.
Dramatic/Emotional Narrative
Goal: Maximum emotional impact with layered performance direction.
Techniques:
- Use contextual priming for short impactful lines
- Layer audio tags: emotional state + pacing + reaction
- Employ ellipses for weighted pauses
- Use em-dashes for interrupted thoughts
- Keep emotional tags close to affected phrases
- Refresh emotional tags to sustain effect
Example:
[quietly] I thought I understood the risk... [long pause] [tense] I didn't.
[nervous] The door opened slowly. [gasps] [whispers] Someone was inside.
The room was empty... or so it seemed. [short pause] [scared] That's when I heard the breathing.
Dialogue/Multi-speaker Scripts
Goal: Distinct character voices and natural conversational flow.
Techniques:
- Use "Speaker:" format for character distinction
- Place audio tags immediately before affected dialogue
- Keep each speaker turn in separate paragraph
- Use character-appropriate tags consistently
- Simulate interruptions with hyphens and
[interrupting]
Example:
Host: [excited] Welcome back! Today's episode is special.
Guest: [cheerful] Thanks for having me. [short pause] I'm thrilled to be here.
Host: [curious] Let's dive right in. What inspired this project?
Guest: [thoughtful] Well, it started with a simple observation...
Technical/Business Content
Goal: Professional, authoritative delivery with perfect pronunciation.
Techniques:
- Normalize ALL acronyms, numbers, measurements
- Use creative spelling for brand names and technical terms
- Keep pacing measured with strategic pauses
- Minimize audio tags (maintain professional tone)
- Spell out metrics and data points
- One concept per paragraph for clarity
Example:
Ak-tiv-trak provides workforce analytics for over nine thousand five hundred brands globally. [short pause] Our S O C 2 compliance ensures enterprise-grade security.
The platform measures productivity signals, schedule adherence and workforce management metrics. [pause] This data drives operational efficiency across distributed teams.
For twenty twenty-six, we're launching our enterprise packaging with two core solutions: workforce management and productivity optimization.
Conversational/Podcast Style
Goal: Natural, engaging flow that mimics authentic human speech.
Techniques:
- Vary sentence length (short punchy + longer flowing)
- Use natural reactions:
[laughs],[sighs],[hmm] - Include verbal fillers when appropriate
- Let punctuation handle most pacing
- Minimal audio tags for authenticity
Example:
So here's the thing about workforce analytics... [short pause] it's not just about tracking time.
[thoughtful] It's really about understanding how work actually happens. [pause] The patterns, the rhythms, the friction points.
And when you see those patterns? [excited] That's when you can make real improvements.
Workflow Instructions
Standard Optimization Process
-
Analyze the text:
- Identify document type (explainer, dialogue, technical, dramatic, conversational)
- Determine emotional context (neutral, energetic, somber, urgent, calm)
- Assess delivery intent (inform, persuade, entertain, instruct)
- Flag problem areas (names, acronyms, numbers, technical terms)
-
Apply V3 techniques:
- Add appropriate audio tags based on document type
- Fix pronunciation issues with creative spelling
- Normalize all numbers, dates, acronyms, URLs
- Engineer punctuation for optimal pacing
- Break into logical paragraphs (one idea each)
- Apply contextual priming if needed for short lines
-
Quality check:
- Verify no SSML tags present
- Confirm audio tags placed immediately before affected phrases
- Ensure paragraph length allows easy regeneration
- Check that character count >250 (V3 needs context to avoid hallucinations)
- Validate all normalizations are complete
-
Deliver output:
- Provide optimized text in copy-paste ready format
- Explain key optimization decisions
- Note any areas requiring testing or alternatives
- Include confidence level for pronunciation fixes
Override Handling
Explicit user instructions always override default optimization:
- "Keep it neutral" → Minimize audio tags, focus on pacing
- "Make it dramatic" → Add layered emotional tags and reactions
- "Professional tone only" → Avoid whispers, shouts, non-verbals
- "Fix pronunciation only" → Focus on creative spelling, skip pacing
- "Add more pauses" → Increase pause density throughout
- "No audio tags" → Use only punctuation and creative spelling
- "Explain the changes" → Include detailed annotation of all modifications
Output Format Templates
Standard Output:
# V3-Optimized Script
[Optimized text with all V3 techniques applied]
---
## Key Optimizations Made
- [Pronunciation fixes applied]
- [Emotional/pacing strategy used]
- [Text normalizations completed]
## Confidence Level: XX%
[Note any uncertainties or areas requiring testing]
Detailed Output (when requested):
# V3-Optimized Script
[Optimized text]
---
## Optimization Report
**Document Type:** [Classification]
**Delivery Intent:** [Purpose]
**Emotional Context:** [Tone]
**Changes Applied:**
1. **Audio Tags Used:**
- [List specific tags and placement rationale]
2. **Pronunciation Fixes:**
- Original → Fixed (rationale)
3. **Pacing Strategy:**
- [Explanation of pause placement logic]
4. **Text Normalizations:**
- [Numbers, acronyms, dates expanded]
**Testing Recommendations:**
- [Specific areas to listen for]
- [Alternative approaches if current doesn't work]
## Confidence Level: XX%
Common Problems & V3 Solutions
| Problem | V3 Solution |
|---|---|
| Audio sounds too fast | Add commas, periods, [short pause] tags between clauses |
| No emotional weight | Add audio tags: [excited], [sad], [calm] immediately before affected phrases |
| Name mispronounced | Use creative spelling: "Siobhan" → "Shiv-awn" |
| Wrong acronym reading | Add spaces: "API" → "A P I" or expand once then use normally |
| Unnatural pauses | Use punctuation first (commas, periods), audio tags second |
| Flat delivery on short line | Apply contextual priming technique with ghost text |
| Number misread | Always normalize: "2026" → "twenty twenty-six" |
| Sounds robotic | Vary sentence length, add natural reactions [sighs], [hmm] |
| Emotional tag fades | Refresh the tag every 2-3 sentences to sustain effect |
| Word slurred together | Use hyphens for syllabic segmentation: "ActivTrak" → "Ak-tiv-trak" |
| Wrong emphasis | Capitalize the stressed syllable or word: "I did NOT say that" |
Quality Optimization Checklist
Before finalizing any V3 script:
- ✓ One idea per paragraph (enables easy regeneration)
- ✓ Pauses primarily from punctuation (minimal audio tags)
- ✓ NO SSML tags (
<break>,<phoneme>don't work in V3) - ✓ All numbers, dates, URLs, acronyms normalized to words
- ✓ Tricky names resolved with creative spelling
- ✓ Audio tags placed immediately before affected phrases
- ✓ Emotional tags refreshed every 2-3 sentences for sustained effect
- ✓ No contradictory tags (avoid
[whispers][shouts]) - ✓ No excessive capitalization (sounds like yelling)
- ✓ Character count >250 (V3 needs context to avoid hallucinations)
- ✓ Each speaker turn in separate paragraph for dialogue
- ✓ Technical terms spelled phonetically when needed
User-Specific Preferences
Critical formatting rules:
- Always use "users" or "accounts" (never "customers" for ActivTrak end users)
- Always use "flexible hours schedules" and "schedule adherence" (not "flex work")
- Never use Oxford comma
- Never use em-dashes unless absolutely necessary
- Century Gothic font in artifacts when applicable
- White backgrounds in artifacts by default
- Copy-paste friendly formats always
Quality standards:
- Ground responses in accurate, factual information
- Never include low-confidence information
- Use
[TBD],[Missing info]or ask for clarification instead of guessing - Provide confidence level for pronunciation fixes and strategic recommendations
ActivTrak terminology:
- Product areas: productivity signals, schedule adherence, workforce management, contractor billing reconciliation
- Navigation: "Productivity Optimization" (not "Performance Optimization")
- Key competitors: Teramind, Hubstaff, Insightful, Time Doctor
Examples
Example 1: Technical Explainer (Raw → Optimized)
Raw:
ActivTrak's new AI Staffing Advisor uses ML algorithms to analyze workforce data and predict staffing needs. The system processes 10,000+ data points per user and provides recommendations in <5 seconds. It's SOC2 compliant and integrates with major HRIS platforms.
V3-Optimized:
[calm] Ak-tiv-trak's new A I Staffing Advisor uses machine learning algorithms to analyze workforce data and predict staffing needs. [short pause] The system processes over ten thousand data points per user and provides recommendations in under five seconds.
[pause] It's S O C 2 compliant and integrates with major H R I S platforms.
Changes made:
- Fixed "ActivTrak" → "Ak-tiv-trak" (syllabic segmentation)
- Normalized "AI" → "A I" (letter-by-letter)
- Expanded "ML" → "machine learning" (first use)
- Normalized "10,000+" → "over ten thousand"
- Normalized "<5" → "under five"
- Normalized "SOC2" → "S O C 2" (letter-by-letter)
- Normalized "HRIS" → "H R I S" (letter-by-letter)
- Added
[calm]for professional tone - Added strategic pauses for comprehension
- Split into two paragraphs for clarity
Example 2: Dramatic Dialogue (Raw → Optimized)
Raw:
"I told you not to open that door," Sarah said.
"I...I didn't think—"
"You never do!" she shouted.
V3-Optimized:
Sarah: [angry] I told you NOT to open that door.
Mark: [nervous] I... I didn't think— [short pause]
Sarah: [shouting] You never do!
Changes made:
- Added speaker labels for clarity
- Added
[angry]to set Sarah's emotional state - Capitalized "NOT" for emphasis
- Added
[nervous]for Mark's delivery - Kept ellipsis and em-dash for natural interruption
- Added
[short pause]to emphasize the cut-off - Added
[shouting]to intensify Sarah's final line - Separated into distinct paragraphs per speaker
Example 3: Educational Content (Raw → Optimized)
Raw:
There are 3 key benefits of schedule adherence tracking: 1) improved operational efficiency 2) better workforce planning 3) enhanced employee accountability. Let's examine each one.
V3-Optimized:
[narrator] There are three key benefits of schedule adherence tracking. [short pause]
First, improved operational efficiency. [pause] When employees follow their assigned schedules, workflows run smoothly and deadlines are met consistently.
Second, better workforce planning. [pause] Historical adherence data helps managers forecast staffing needs more accurately.
Third, enhanced employee accountability. [pause] Clear expectations and tracking create transparency for both employees and managers.
Changes made:
- Normalized "3" → "three"
- Removed numbered list format (converted to prose)
- Added
[narrator]for professional educational tone - Added strategic pauses for comprehension
- Expanded each point into full explanatory sentences
- Split into logical paragraphs (one concept each)
- Used "First, Second, Third" pattern for clarity
Example 4: Conversational Podcast Style (Raw → Optimized)
Raw:
So basically what we're seeing in 2026 is this massive shift in how companies think about productivity. It's not just about time tracking anymore. Companies want insights - they want to understand WHY things are happening, not just WHAT is happening.
V3-Optimized:
[conversational] So basically what we're seeing in twenty twenty-six is this massive shift in how companies think about productivity. [short pause] It's not just about time tracking anymore.
[thoughtful] Companies want insights. [pause] They want to understand WHY things are happening, not just WHAT is happening.
Changes made:
- Normalized "2026" → "twenty twenty-six"
- Added
[conversational]to set natural tone - Added
[thoughtful]for contemplative shift - Added strategic pauses for emphasis
- Split into two paragraphs for better pacing
- Kept capitalized "WHY" and "WHAT" for natural emphasis
- Removed em-dash, split into two sentences instead
Example 5: Brand Name Pronunciation (Raw → Optimized)
Raw:
Teramind, Insightful and Time Doctor are ActivTrak's primary competitors in the workforce analytics space.
V3-Optimized:
Tair-uh-mind, In-syte-full and Time Doctor are Ak-tiv-trak's primary competitors in the workforce analytics space.
Changes made:
- Fixed "Teramind" → "Tair-uh-mind" (phonetic respelling)
- Fixed "Insightful" → "In-syte-full" (syllabic segmentation + phonetic)
- Fixed "ActivTrak" → "Ak-tiv-trak" (syllabic segmentation)
- Kept "Time Doctor" as-is (clear pronunciation)
- No audio tags needed (neutral informational tone)
Confidence Levels Guide
When providing optimizations, always include a confidence level:
90-100%: Standard normalizations, common words, established patterns
- Numbers, dates, basic acronyms
- Common punctuation engineering
- Standard audio tag usage
70-89%: Brand names with unclear pronunciation, technical jargon, unusual names
- Custom brand names without official pronunciation guide
- Technical terms with multiple possible pronunciations
- Names from unfamiliar languages
Below 70%: Omit uncertain information, use [TBD] or ask for clarification
- Obscure proper nouns without reference
- Highly technical domain-specific jargon
- When user intent is ambiguous
Example confidence statements:
- "Confidence: 95% - All normalizations use standard conventions"
- "Confidence: 75% - 'Teramind' pronunciation based on phonetic analysis, recommend testing"
- "Confidence: 60% - Unable to verify 'Nguyen' pronunciation preference, using common English approximation. Please confirm or provide guidance."
Advanced Techniques
The "Sandwich" Method for Sustained Emotion
Emotional tags decay over time. To sustain emotion across multiple sentences:
Instead of:
[angry] I told you to stop! Why won't you listen? This is the last time!
Use:
[angry] I told you to stop! [angry] Why won't you listen? [angry] This is the last time!
The "Neutralizing Tag" Reset
To end an emotional performance, introduce a new contrasting tag:
[excited] We did it! We actually did it! [pause] [calm] Now let's review what comes next.
Stutter Engineering with Hyphens
Create realistic stutters or hesitations:
I-I-I don't think that's right. [nervous] We should probably double-check.
The "Silent Actor" Technique
Use [pause] with context to create dramatic silence:
[tense] She opened the envelope. [long pause] [whispers] It was empty.
Accent Layering for Characters
Combine accent tags with emotional tags for distinct characters:
Guard 1: [British accent] [authoritative] State your business.
Guard 2: [American accent] [casual] Hey, take it easy on 'em.
Integration Notes
This skill works alongside other skills:
- activtrak-brand-guidelines: Applies brand terminology corrections
- linkedin-post-generator: Can optimize voice-over scripts for LinkedIn videos
- activtrak-release-notes: Can add V3 optimization to release note voice-overs
- doc-coauthoring: Can optimize finalized documentation for audio versions
Limitations & Disclaimers
This skill cannot:
- Guarantee perfect pronunciation without testing (V3 is voice-dependent)
- Replace actual audio testing and iteration
- Control global settings (Stability, Clarity, Voice Selection)
- Edit or trim audio files (post-production only)
- Support V2 SSML requirements (different skill needed)
Important notes:
- V3 is nondeterministic (same text may produce slight variations)
- Audio tag effectiveness varies by voice model selected
- Always test critical pronunciations in actual generation
- Some voices may add "uh" or "ah" during pauses (voice-dependent behavior)
- Character count below 250 may cause hallucinations in V3
Version Notes
Skill version: 1.0
Based on: ElevenLabs V3 documentation (January 2025)
Compatible with: Eleven V3 Alpha, Eleven V3
Not compatible with: Eleven Multilingual v2, Turbo v2.5, Flash v2 (use SSML skill instead)
Last updated: January 2025
Maintained for: Daniel's ActivTrak product marketing workflows