agentsclimarketplace

Nvidia nemotron voice agent deploy

Skill autohandai/community-skills/nvidia-nemotron-voice-agent-deploy

Deploy Nemotron Voice Agent on Workstation (x86), Jetson Thor, or Cloud NIMs. Real-time speech-to-speech using NVIDIA ASR, TTS, LLM with WebRTC/WebSocket transport.From its SKILL.md

Install
npx -y skills add autohandai/community-skills --skill nvidia-nemotron-voice-agent-deploy

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 9 stars9 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its file declares

Copied from the file, not written here

The file declares its own license as Apache-2.0 AND CC-BY-4.0. That is the author’s claim about this one file, and it is not the same thing as the license GitHub reports for the repository, which is listed with the other numbers below.

SKILL.md

2.3 KB, 615 tokens by cl100k_base, as published. Nobody here has run it

Nemotron Voice Agent Deployment

Real-time conversational AI voice agent using NVIDIA NIMs (ASR, TTS, LLM) with WebRTC (default) or WebSocket transport.

Deployment Flow

Always verify hardware first, even if user mentions a specific platform.

STEP 1: Hardware Detection

nvidia-smi --query-gpu=name,memory.total --format=csv,noheader 2>/dev/null
ResultAction
Command fails / No outputCloud NIMs
GPU detectedSTEP 2: Platform Detection

Cloud NIMs (No GPU)

cd nemotron-voice-agent
git submodule update --init
cp config/env.example .env

Export your NVIDIA API key:

export NVIDIA_API_KEY=your-api-key  # Get from https://build.nvidia.com

Then edit .env:

NVIDIA_LLM_MODEL=nvidia/nemotron-3-nano-30b-a3b  # Cloud model name

If user requests WebSocket transport, also add to .env:

TRANSPORT=WEBSOCKET
docker compose up --build --no-deps -d python-app ui-app
# WebRTC: http://localhost:9000
# WebSocket: http://localhost:7860/static/index.html

Note: Deployment may take 30-60 minutes on first run.

If user requests Multilingual mode, also add to .env:

ENABLE_MULTILINGUAL=true
ASR_CLOUD_FUNCTION_ID=71203149-d3b7-4460-8231-1be2543a1fca
ASR_MODEL_NAME=parakeet-rnnt-1.1b-unified-ml-cs-universal-multi-asr-streaming

Remote Access: ssh -L 9000:localhost:9000 user@host or http://<HOST_IP>:9000


STEP 2: Platform Detection (if GPU detected)

uname -m  # x86_64 → Workstation, aarch64 → Jetson
cat /etc/nv_tegra_release 2>/dev/null && echo "Jetson"
PlatformReferenceRequirements
Workstation (x86_64)workstation-deployment.md2x GPU (24GB+ VRAM), NIM containers
Jetson Thor (aarch64)jetson-deployment.mdJetPack 7.0, Nemotron Speech ASR and TTS, vLLM

Note: Multilingual mode available on Workstation with WebRTC transport only.

What ships with it: 3 files

17.0 KB alongside SKILL.md

Keep looking

Skills are one crate of 326,422. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.