Deploy a streaming speech recognition (ASR) and voice synthesis (TTS) service on your edge device — Jetson Orin, RK3576, RK3588, or a Pi 5.
What you'll get:
- Real-time streaming speech recognition (WebSocket)
- Low-latency voice synthesis (HTTP streaming + batch)
- Multiple language modes: Chinese+English, English-only, or 52-language Qwen3
- HTTP + WebSocket API on port 8621
Requirements: SSH access to device · Internet to pull Docker image and download models