Clear Audio
Clear in noise, and you can interrupt while it speaks
Local Compute
Fully local: 0.9 s to reply, measured; works offline
Tuned, Ready to Run
Tuned and benchmarked per board, open source

Scenario details

Interaction and deployment options for desktop voice robots, industrial voice control and smart home assistants.

Understand and reply on the device

Recognizes multiple languages on the device and replies by voice; fits desktop devices, meeting terminals and guide kiosks for two-way voice conversation.


Core Advantages

  • Multilingual Recognition
  • Barge-in
  • Tiered Voice Quality
Scene Feature
Multilingual Recognition
30 conversation languages, one chosen at deployment; if the device cannot serve the chosen language, deployment stops with an error.
Scene Feature
Barge-in
The microphone keeps listening during the reply; when the user speaks, playback stops and the next turn starts.
Scene Feature
Voice Persona
Machine, simulated and natural voice tiers; the natural-voice model on J4012 can clone a voice from reference audio.

Voice device control & field data entry

Voice replaces UIs & scanners in warehouses & workshops. Workers use natural language for logging, inspection, patrol forms & alerts. Local ASR outputs structured text for WMS, MES & IoT.


Core Advantages

  • Lower Op Barrier
  • Weak-Network Ready
  • Structured Output
Scene Feature
Warehouse Inbound/Outbound
Voice verify SKU → direct WMS write
Scene Feature
Equipment Inspection
Voice equip check auto-fills forms
Scene Feature
Patrol Reporting & Alerts
Voice patrol forms & hazard alerts

Wake, then handle it locally — even offline

XIAO ESP32S3 is the low-power wake frontend; after wake-up the main unit handles recognition, conversation and the spoken reply, and controls devices through local protocols such as Matter, HomeAssistant and Mi Home. With the on-device LLM, commands are processed locally and work offline.


Core Advantages

  • Low standby power
  • Custom wake word
  • Local Control
Scene Feature
Low-Power Wake
Low-power wake word on ESP32S3
Scene Feature
Custom Wake Word
Replace the wake word with a short Chinese or English brand phrase; sensitivity has strict, balanced and sensitive settings.
Scene Feature
Local IoT Orchestration
Matter/HA/Mi Home local control

Deployment & Selection

The configurator works through use case, voice tier, concurrency, on-device LLM, microphone setup and latency target to reach specific models.

Not sure which to choose? See the full selection guide →

Use Case

Three Models: Frontend, Hybrid, All-in-One

Voice compute placement determines capability ceiling & per-unit BOM. Three common models:


Core Advantages

  • Frontend only
  • Hybrid
  • All-in-One (Frontend + Large AI Box)
ImageHardwareVoice Concurrency (voice-only)Voice CapabilitiesLocal LLM
reRouter CM4 SeriesreRouter CM4 Series1 channelMachine VoiceCloud optional
reComputer RK3576 SeriesreComputer RK3576 Series1 channelSimulated Voice4B–7B (w/ RK1828 card)
reComputer RK3588 SeriesreComputer RK3588 Series1 channelSimulated Voice4B–7B (w/ RK1828 card)
reComputer J30 SeriesreComputer J30 Series1 channelNatural Voice~2B
reComputer J40 SeriesreComputer J40 Series2–3 channelsNatural Voice4B–7B
reComputer J50 SeriesreComputer J50 Series3+ channelsNatural Voice7B–14B
NVIDIA Jetson AGX Thor Developer KitNVIDIA Jetson AGX Thor Developer Kit6+ concurrent (voice-only)Realistic voice14B–32B

Choose AI Compute Box by Capability

Compute boxes are tiered by voice capabilities. The table below is the device list for this scenario (by family): voice-only concurrency, voice capability, and deployable local-LLM size per family. Pick exact models, configurations, and accessories in the configurator below (see next tab for mic & speaker selection). Note: running a local LLM on the same device as the voice pipeline reduces performance and increases latency; for high concurrency plus large models, choose a higher compute tier.


Core Advantages

  • Wake & Commands
  • Standalone
  • Pro-tier natural voice
  • High concurrency + advanced LLM
ProductTypeChipPickup
Range
Coverage
Angle
Built-in
Amp
Core Algorithms
reSpeaker LiteLinear
2-Mic
XMOS XU3163m180°5WAEC · DoA
reSpeaker XVF3800Circular
4-Mic
XMOS XVF38005m360°5WAEC · DoA · Multi-beamforming
reSpeaker ClipWearable
2-Mic
nRF5340+nRF70023m360°—Noise reduction (SpeexDSP)
reSpeaker Flex Circular-4Circular
4-Mic
XMOS XVF38005m360°10WAEC · DoA · Multi-beamforming
reSpeaker Flex Linear-4Linear
4-Mic
XMOS XVF38005m180°10WAEC · DoA · Multi-beamforming

Three Core Advantages of the reSpeaker Series


Core Advantages

  • Hardware pickup
  • On-board acoustics
  • Open SDK

FAQ

Didn't find your answer?
Contact Us
We Are Glad to Be Your Hardware Partner !
Next