Document audio models (STT/TTS) with API examples

This commit is contained in:
inference-bot committed 2026-09-27 11:25:50 -06:00
1 parent 40925b82b5
commit 34ce3b6753
1 file changed
+32
+32
View File
@@ -52,6 +52,38 @@ provider — your own cost is your monthly allowance, not per-token billing).
| `publicai/apertus-v1.5-8b` | PublicAI | policy | undisclosed | $0.10 / $0.20 | Fully open sovereign model, small/fast. |
| `publicai/apertus-v1.5-70b` | PublicAI | policy | undisclosed | $0.82 / $2.92 | Fully open sovereign model, large/capable. |
### Audio models
These are **not chat models** — they power the chat's voice mode (dictation
and spoken replies) and are available through the API on `/v1/audio/*` routes
for members building speech into their own tools. They're hidden from the chat
model picker to avoid confusion.
| Model | Provider | Use | Pricing |
|---|---|---|---|
| `tinfoil/whisper-large-v3-turbo` | Tinfoil | Speech-to-text | $0.05 / 1M input tokens |
| `tinfoil/voxtral-tts` | Tinfoil | Text-to-speech | listed $0 (verify on invoice) |
Both run inside Tinfoil's enclaves, so voice data gets the same architectural
privacy as the private chat models. TTS voices available: `neutral_female`,
`neutral_male`, `casual_female/male`, `cheerful_female`, plus French, German,
Spanish, Italian, Portuguese, Dutch, and Hindi variants.
```bash
# Speech-to-text
curl https://gateway.inference.coop/v1/audio/transcriptions \
-H "Authorization: Bearer sk-your-key" \
-F file=@your-audio.mp3 \
-F model=tinfoil/whisper-large-v3-turbo
# Text-to-speech (returns WAV audio)
curl https://gateway.inference.coop/v1/audio/speech \
-H "Authorization: Bearer sk-your-key" \
-H "Content-Type: application/json" \
-d '{"model": "tinfoil/voxtral-tts", "input": "Hello from the co-op", "voice": "neutral_female"}' \
--output speech.wav
```
### About the Apertus models
Apertus is the **Swiss AI Initiative's** open foundation model, built by a