Voice mode latency/broken in chat #1

Closed
opened 2026-09-24 14:39:35 +00:00 by ntnsndr · 1 comment

What do you need help with?

In the chat interface, voice mode isn't really working. Perhaps we need to set up a dedicated TTS model or the like to facilitate this.

What happened?

  • Dictation works, but it is very slow
  • The interactive voice mode doesn't seem to work at all
## What do you need help with? In the chat interface, voice mode isn't really working. Perhaps we need to set up a dedicated TTS model or the like to facilitate this. ## What happened? - Dictation works, but it is very slow - The interactive voice mode doesn't seem to work at all
Owner

Resolved. Root cause was two layers:

  1. No STT/TTS engine configured. Open WebUI's audio settings were empty (fallback = slow local whisper base for dictation; no TTS at all, so interactive voice mode couldn't work).

  2. Model name mismatch (my error). The first wiring attempt set the STT model to whisper-large-v3-turbo without the tinfoil/ prefix. LiteLLM requires exact names, so every transcription request was rejected ("Invalid model name"), surfacing as 500 External on both dictation and voice mode.

What's now in place:

  • STT: tinfoil/whisper-large-v3-turbo via the co-op gateway — hosted, TEE-protected (private tier), ~2s per transcription vs. slow local whisper.
  • TTS: tinfoil/voxtral-tts (voice neutral_female) — GreenPT has no TTS offering, and Tinfoil's is listed at $0.
  • The audio models are registered on the gateway and visible in the API model list for developers calling /v1/audio/*, but filtered out of the chat model picker (they're not chat models) via a portal-side filter.
  • Verified: transcription returns HTTP 200 end-to-end; TTS returns real WAV audio through the EHBP-encrypted sidecar.

Commits: gateway 8f42be9, portal 54fa4d3. Voice settings: STT model prefixed correctly after the mismatch fix.

Resolved. Root cause was two layers: 1. **No STT/TTS engine configured.** Open WebUI's audio settings were empty (fallback = slow local `whisper base` for dictation; no TTS at all, so interactive voice mode couldn't work). 2. **Model name mismatch (my error).** The first wiring attempt set the STT model to `whisper-large-v3-turbo` without the `tinfoil/` prefix. LiteLLM requires exact names, so every transcription request was rejected ("Invalid model name"), surfacing as `500 External` on both dictation and voice mode. **What's now in place:** - STT: `tinfoil/whisper-large-v3-turbo` via the co-op gateway — hosted, TEE-protected (private tier), ~2s per transcription vs. slow local whisper. - TTS: `tinfoil/voxtral-tts` (voice `neutral_female`) — GreenPT has no TTS offering, and Tinfoil's is listed at $0. - The audio models are registered on the gateway and visible in the **API** model list for developers calling `/v1/audio/*`, but filtered out of the **chat** model picker (they're not chat models) via a portal-side filter. - Verified: transcription returns HTTP 200 end-to-end; TTS returns real WAV audio through the EHBP-encrypted sidecar. Commits: gateway `8f42be9`, portal `54fa4d3`. Voice settings: STT model prefixed correctly after the mismatch fix.
Sign in to join this conversation.
No labels
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: co-op/support#1