New open-source voice model listens nonstop and decides every 0.4 seconds whether to speak or stay silent
Signal
72
Hype
35
In three linesAn open-source voice model listens continuously and decides every 0.4 seconds whether to speak or stay silent. Unlike GPT-4o or Qwen3.5-Omni, Audio Interaction handles transcription, translation, and chat in a single stream, detecting ambient noise (coughing). Code, weights, and training data available on GitHub under Apache 2.0 license.Read source
Your take?
Summary generated by Claude — human-verified