agentsclimarketplace

Openai whisper

Skill coco-research/coco/skills/openai-whisper

Meet Coco. A superintelligent agent framework powered by an advisory board of 389 world-class minds. Scale your AI assistant into a complete engineering department with 142 skills, 277 commands, and persistent state. Universal compatibility. Local privacy. Free and open source.

Install
npx -y skills add coco-research/coco --skill openai-whisper

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.

What its author says it does

Copied from the file, not written here

Speech-to-text transcription via OpenAI Whisper. Supports two modes — Local CLI (no API key, runs on-device) and Cloud API (fast, scalable, requires OPENAI_API_KEY). Use when the user needs to transcribe audio files, translate speech, or convert audio to text.

SKILL.md

3.8 KB, as published. Nobody here has run it

OpenAI Whisper — Speech-to-Text

Transcribe audio files using OpenAI's Whisper model. Two modes available depending on your needs:

ModeLatencyCostPrivacySetup
Local CLISlower (on-device GPU/CPU)FreeAudio never leaves machineInstall whisper binary
Cloud APIFastPer-minute pricingAudio sent to OpenAIOPENAI_API_KEY required

Mode 1: Local CLI

Run Whisper locally with no API key required. Models download to ~/.cache/whisper on first run.

Quick Start

whisper /path/audio.mp3 --model medium --output_format txt --output_dir .

Common Commands

# Transcribe to text file
whisper /path/audio.mp3 --model medium --output_format txt --output_dir .

# Transcribe with translation to English
whisper /path/audio.m4a --task translate --output_format srt

# Transcribe with specific language
whisper /path/audio.wav --model large --language en --output_format json

Model Selection

ModelSpeedAccuracyVRAM
tinyFastestLowest~1 GB
baseFastLow~1 GB
smallMediumGood~2 GB
mediumSlowBetter~5 GB
largeSlowestBest~10 GB
turboFastGood (default)~6 GB

Output Formats

  • txt — Plain text transcript
  • srt — SubRip subtitle format with timestamps
  • vtt — WebVTT subtitle format
  • json — Detailed JSON with word-level timestamps
  • tsv — Tab-separated values

Notes

  • --model defaults to turbo on most installs
  • Use smaller models for speed, larger for accuracy
  • GPU acceleration used automatically when available

Mode 2: Cloud API

Transcribe via OpenAI's /v1/audio/transcriptions endpoint. Faster for large batches, no local GPU needed.

Quick Start

{baseDir}/scripts/transcribe.sh /path/to/audio.m4a

Defaults:

  • Model: whisper-1
  • Output: <input>.txt

Common Commands

# Basic transcription
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a

# Specify model and output
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model whisper-1 --out /tmp/transcript.txt

# With language hint
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --language en

# With speaker name hints (improves accuracy)
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --prompt "Speaker names: Peter, Daniel"

# JSON output with timestamps
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --json --out /tmp/transcript.json

Raw curl Example

curl https://api.openai.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -H "Content-Type: multipart/form-data" \
  -F file="@/path/to/audio.m4a" \
  -F model="whisper-1" \
  -F response_format="text"

API Key Setup

Set OPENAI_API_KEY environment variable, or configure in ~/.clawdbot/clawdbot.json:

{
  skills: {
    "openai-whisper-api": {
      apiKey: "OPENAI_KEY_HERE"
    }
  }
}

Choosing Between Modes

ConsiderationLocal CLICloud API
Privacy-sensitive audioBestAudio sent to OpenAI
Large batch processingSlow without GPUFast and parallel
Offline usageWorks offlineRequires internet
CostFree (hardware cost)Per-minute pricing
Setup complexityInstall binary + modelsAPI key only
Audio format supportMost formatsMost formats

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.