Hugging Face speech-to-speech: Voice Agents That Run on Your Own Hardware
A four-stage voice pipeline — VAD, speech-to-text, a language model, text-to-speech — behind an OpenAI Realtime-compatible API. Every stage is swappable, and the whole thing can run without a single network request.