Real-Time Voice Assistant (voice-rt)

Status: Phase 2 complete & verified 2026-09-04 — live Tags: Service Deployment Infrastructure Multimodal Agentic Created: 2026-09-05 Updated: 2026-09-05 Related: local-llm-stack, cloudflare-pages-deploy, hermes-cron-operations

Definition

A live real-time, full-duplex voice assistant (speech-in → speech-out) running on this host, exposed to phones as a PWA at https://voice.meiyoucheveux.com. Built 2026-09-04: Phase 1 = offline STT→LLM→TTS round-trip (e2e.py, sample replies out_en.wav / out_zh.wav); Phase 2 = live streaming server with barge-in. Supports English + 中文 and multiple simultaneous phones over one websocket.

How It Works

Client-to-server path:

Phone PWA (mic, 16 kHz int16 PCM)
  -> wss://voice.meiyoucheveux.com/ws (Cloudflare edge)
  -> cloudflared.service  (ingress voice.meiyoucheveux.com -> http://localhost:8090)
  -> voice-rt/server.py (aiohttp, port 8090)
       state machine: IDLE -> LISTENING -> PROCESSING -> SPEAKING

Per-utterance pipeline:

  1. VAD — Silero VAD: onset = 2 consecutive 32 ms chunks ≥ 0.5; end = 25 silent chunks; max 625 chunks.
  2. STT — faster-whisper medium, CPU int8 (voice-rt/whisper/medium/).
  3. LLM — llama-server-voice4b on :8020 (GPU 1, Qwen3-4B Q4_K_M, 4 slots) — a dedicated serving instance, see local-llm-stack.
  4. TTS — Piper 1.7.0 (CPU), en_US-lessac-medium / zh_CN-huayan-medium, streamed chunk-by-chunk (first audio ~20 ms after the LLM reply is ready).

Barge-in: speech detected during SPEAKING cancels the in-flight TTS+LLM and flushes the audio queue (~50 ms interrupt → IDLE).

Service / boot: runs as systemd voice-rt.service (enabled at boot, Restart=on-failure), pinning /home/romain/.hermes/hermes-agent/venv/bin/python (has aiohttp, faster_whisper, piper, onnxruntime). Boot order: docker.service → llama-server-voice4b container (unless-stopped, :8020) → cloudflared.service → voice-rt.service (:8090). Whisper + Piper models load on first use; the PWA serves immediately.

Key Parameters

ItemValue
PWAhttps://voice.meiyoucheveux.com (tap 📱 Install to add to home screen)
Servervoice-rt/server.py, aiohttp, port 8090
TunnelCloudflare, cloudflared.service, tunnel id 2350ecdf-b6f6-4b72-87d9-48f46b8108a5, config /home/romain/.cloudflared/config.yml (DNS voice is a plain CF-edge A record, not a CNAME)
STTfaster-whisper medium, CPU int8
LLMQwen3-4B Q4_K_M on :8020 (GPU 1, 4 slots) — llama-server-voice4b
TTSPiper 1.7.0, en_US-lessac + zh_CN-huayan (medium)
LanguagesEnglish + 中文
Measured (LAN, 2026-09-04)EN round-trip ~6.2 s wall; ZH ~8.1 s (STT of the full clip ~3.7 s dominates); first TTS ~20–60 ms after LLM ready; barge-in ~50 ms; 2 concurrent sessions — no cross-stall (total ≈ max, not sum)

When To Use

  • Live hands-free voice assistant on a phone (mobile Safari/Chrome PWA), EN/ZH, multiple devices at once.
  • Reference design for a full-duplex speech pipeline on this host’s local stack: VAD + faster-whisper + a small GPU LLM + Piper, fronted by a Cloudflare websocket tunnel.

Risks & Pitfalls

  • Silent bot on phone (transcript shows, no voice): mobile Safari/Chrome keep the playback AudioContext suspended until a user tap. Unlock it inside the 🎤 tap (playCtx().resume() in start()), not later when audio arrives. Also: volume up, not on the silent switch. Diagnose via the bottom status line — “speaking…” = audio blocked (client); stuck on “thinking…” = server/tunnel.
  • Invariants — never touch: the 27 B llama-server on :8000 and vllm-06b on :8010 must stay untouched by voice-rt. Do not create a .venv-voice; all deps live in the hermes venv. ffmpeg = voice-rt/bin/ffmpeg (symlink to the imageio-ffmpeg static binary).
  • edge-tts test clips have internal pauses → VAD fragments them and may fire a barge_in mid-test (correct behaviour, not a bug).
  • write_file truncates content >~1 KB — build big files via execute_code or small part-files.
  • Model downloads only via https://hf-mirror.com (huggingface.co is blocked; the python huggingface_hub API 401s on CAS — use wget).
  • PyPI file downloads are slow — run pip installs as background jobs.
  • Server stdout is buffered when backgrounded — trust test-client output, not the server log.

Tests

cd /home/romain/workspace/voice-rt
python3 test_ws.py in_en.mp3 en              # EN round-trip (LAN)
python3 test_ws.py in_zh.mp3 zh              # 中文 round-trip
python3 test_ws.py in_en.mp3 en interrupt    # barge-in (~50 ms)
python3 conc_test.py 2                        # 2 concurrent sessions
# public (through Cloudflare):
sed 's|ws://127.0.0.1:8090/ws|wss://voice.meiyoucheveux.com/ws|' test_ws.py > /tmp/pub.py
python3 /tmp/pub.py in_en.mp3 en
  • local-llm-stack — the llama.cpp / vLLM serving stack, including the dedicated :8020 voice4b LLM that this assistant calls
  • cloudflare-pages-deploy — Cloudflare tooling on this host (Python deployer); the tunnel here is the separate cloudflared.service
  • hermes-cron-operations — general host infra / ops context

Sources

  • /home/romain/workspace/voice-rt/RUNBOOK.md (runbook; Phase 2 verified 2026-09-04)
  • /home/romain/workspace/voice-rt/ — server.py, www/index.html (PWA), test_ws.py, conc_test.py, e2e.py, bench.py, voice-rt.service
  • Quick health check: curl http://127.0.0.1:8090/, :8020/health, :8000/health, https://voice.meiyoucheveux.com/ (all 200) + systemctl is-active cloudflared.service voice-rt.service