agentsclimarketplace

Elevenlabs performance tuning

Skill jeremylongshore/claude-code-plugins-plus-skills/skills/.curated/elevenlabs-performance-tuning

425 plugins, 2,810 skills, 200 agents for Claude Code. Open-source marketplace at tonsofskills.com with the ccpi CLI package manager.

Install
npx -y skills add jeremylongshore/claude-code-plugins-plus-skills --skill elevenlabs-performance-tuning

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

What its author says it does

Copied from the file, not written here

Optimize ElevenLabs TTS latency with model selection, streaming, caching, and audio format tuning. Use when experiencing slow TTS responses, implementing real-time voice features, or optimizing audio generation throughput. Trigger with "elevenlabs performance", "optimize elevenlabs", "elevenlabs latency", "elevenlabs slow", "fast TTS", "reduce elevenlabs latency", or "TTS streaming".

The file declares its own license as MIT. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

7.9 KB, as published. Nobody here has run it

ElevenLabs Performance Tuning

Overview

Optimize ElevenLabs TTS latency and throughput through model selection, streaming strategies, audio format tuning, and caching. Latency ranges from ~75ms (Flash) to ~500ms (v3) depending on configuration.

The two highest-leverage, lowest-effort levers — model choice (Step 1) and output format (Step 2) — are documented inline below. The four deeper integrations (HTTP streaming, WebSocket streaming, caching, parallel generation) are summarized here with copy-ready code in the full implementation walkthrough.

Prerequisites

  • ElevenLabs SDK installed (@elevenlabs/elevenlabs-js)
  • An ElevenLabs API key exported as ELEVENLABS_API_KEY (used by the SDK and passed as xi_api_key on the WebSocket handshake)
  • Understanding of your latency requirements
  • Audio playback infrastructure (browser, mobile, server-side)

Instructions

Step 1: Model Selection for Latency

The single biggest performance lever is model choice:

ModelAvg LatencyQualityLanguagesUse Case
eleven_flash_v2_5~75msGood32Real-time chat, IVR, gaming
eleven_turbo_v2_5~150msGood32Balanced speed/quality
eleven_multilingual_v2~300msHigh29Narration, content creation
eleven_v3~500msHighest70+Maximum expressiveness
// Select model based on use case
function selectModel(useCase: "realtime" | "balanced" | "quality" | "max_quality"): string {
  const models = {
    realtime:    "eleven_flash_v2_5",
    balanced:    "eleven_turbo_v2_5",
    quality:     "eleven_multilingual_v2",
    max_quality: "eleven_v3",
  };
  return models[useCase];
}

Step 2: Output Format Optimization

Smaller formats = faster transfer:

FormatSize/SecondQualityBest For
mp3_44100_128~16 KB/sHighDownloads, archival
mp3_22050_32~4 KB/sMediumStreaming, mobile
pcm_16000~32 KB/sRawServer-side processing
pcm_44100~88 KB/sRawHigh-quality processing
ulaw_8000~8 KB/sPhoneTelephony/IVR
// Use smaller format for streaming, higher quality for downloads
const streamingConfig = {
  output_format: "mp3_22050_32",  // 4 KB/s — fast streaming
  model_id: "eleven_flash_v2_5",   // ~75ms first byte
};

const downloadConfig = {
  output_format: "mp3_44100_128", // 16 KB/s — high quality
  model_id: "eleven_multilingual_v2",
};

Step 3: HTTP Streaming for Time-to-First-Byte

Call client.textToSpeech.stream() instead of .convert() and write each chunk to the response as it arrives, so playback starts before generation finishes — roughly halving time-to-first-byte. Set style: 0.0 in voice_settings to shave another 10–20%. Full server handler: implementation.md § Step 3.

Step 4: WebSocket Streaming for Lowest Latency

For interactive apps where text arrives incrementally (e.g., an LLM token stream), open a stream-input WebSocket, sendText() chunks as they arrive, and tune chunk_length_schedule — fewer characters per chunk means lower latency but less prosody context. Full bidirectional client: implementation.md § Step 4.

Step 5: Audio Caching

Cache generated audio for repeated content (greetings, prompts, errors) in an LRU cache keyed by a SHA-256 of voiceId:modelId:text, so a changed voice or model never serves stale audio. This eliminates ~99% of latency for repeated phrases. Full cachedTTS helper: implementation.md § Step 5.

Step 6: Parallel Generation

Generate multiple segments concurrently with a p-queue whose concurrency matches your plan's request limit (going higher returns 429s, not more throughput). Full chapter-generator: implementation.md § Step 6.

Output

Applying these levers produces:

  • A model + output-format choice matched to the use case (Steps 1–2).
  • A streaming code path (HTTP or WebSocket) that logs measured time-to-first-byte, e.g. Time to first byte: 78ms / WebSocket TTFB: 91ms.
  • An LRU audio cache emitting [Cache HIT] / [Cache MISS] telemetry for repeated content.
  • A concurrency-bounded batch path that logs per-segment generation time.

Expected latency after tuning: ~75–150ms first byte on Flash/Turbo with streaming, versus ~300–500ms for a blocking convert() call on a higher-quality model.

Performance Optimization Checklist

OptimizationLatency ImpactImplementation
Flash model-60% vs v2, -85% vs v3Change model_id
Streaming endpoint-50% time-to-first-byteUse .stream() instead of .convert()
WebSocket streamingBest for LLM integrationSee Step 4
Smaller output format-30% transfer timemp3_22050_32 vs mp3_44100_128
Audio caching-99% for repeated contentLRU cache with SHA-256 keys
style: 0-10-20% latencyRemove style exaggeration
Concurrency queueMaximize throughputp-queue matching plan limit

Error Handling

IssueCauseSolution
High TTFBWrong modelSwitch to eleven_flash_v2_5
Choppy streamingNetwork bufferingUse pcm_16000 for direct playback
Cache miss stormTTL expired for popular contentUse stale-while-revalidate pattern
WebSocket dropsNetwork instabilityReconnect with buffered text
Memory pressureAudio cache too largeSet maxSize limit on LRU cache
HTTP 429Concurrency above plan limitLower p-queue concurrency

Examples

Real-time IVR (lowest latency). Pick eleven_flash_v2_5 + ulaw_8000 via selectModel("realtime"), then stream over HTTP:

await streamToResponse(greeting, voiceId, res); // logs "Time to first byte: 78ms"

LLM voice agent (incremental text). Open a WebSocket and forward tokens as they stream from the model, ending with finish():

const stream = await createTTSStream({ voiceId, chunkLengthSchedule: [50, 100, 150] });
stream.sendText("Hello, "); stream.sendText("how are you?");
const audio = await stream.finish();

Audiobook batch (throughput). Cache repeated phrases and generate chapters concurrently:

const buffers = await generateChapters(chapters, voiceId); // 5-wide, cache-backed

Full, runnable versions of every snippet above are in the implementation walkthrough.

Resources

Next Steps

For cost optimization once latency is tuned, see the elevenlabs-cost-tuning skill, which covers character-usage budgeting, model-tier cost tradeoffs, and cache-hit-rate targets.

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.