Commit Graph
16 Commits
Author SHA1 Message Date
Inferencebot b6e8542146 config: remove duplicate broken audio-model entries (audio_transcription//audio_speech/ prefixes); keep the correct openai/ routing 2026-10-02 15:39:07 -06:00
inference-bot 8f42be988e Add Tinfoil audio models (whisper STT, voxtral-tts) for voice mode + API use 2026-09-27 11:10:51 -06:00
inference-bot 5ea2256776 Add greenpt/glm-5.3 (full flagship, $1.10/$4.40 per 1M) 2026-09-26 16:52:21 -06:00
inference-bot 057507d67c Add PublicAI Apertus models (publicai/apertus-v1.5-8b, -70b) 2026-09-25 15:56:08 -06:00
inference-bot 2369c3e848 Add greenpt/glm-5.3-flash (/usr/bin/bash.11//usr/bin/bash.44 per 1M) 2026-09-24 10:43:14 -06:00
inference-bot f86e005062 Rename Tinfoil models to tinfoil/* provider prefix 2026-09-24 08:16:16 -06:00
inference-bot 67192de232 Document pricing sync with co-op/docs/models.md 2026-09-23 21:08:33 -06:00
inference-bot 09ee36014c Add GreenPT models (greenpt/green-r, greenpt/green-l) via provider/model naming convention 2026-09-23 20:10:18 -06:00
inference-bot d1f57c45d5 Retry connection errors but never re-bill timeouts
num_retries=2 (retries transient connection errors), retry_policy with
TimeoutErrorRetries=0 (a timed-out generation is already billed by Tinfoil, so
re-hitting it re-bills). Combined with timeout=180, this hedges connection
errors without the silent re-billing that caused the Tinfoil spend gap.
2026-09-22 18:18:07 -06:00
inference-bot 576885ed21 Fix Tinfoil spend gap: raise timeout 30->180s, num_retries 2->0
DeepSeek V4.1 Flash is a reasoning model; 30s timeout caused LiteLLM to
abandon in-flight generations (falling back to GPT-OSS) while Tinfoil still
completed and billed them. num_retries re-hit Tinfoil up to 3x. Result: Tinfoil
billed ~2x the tokens LiteLLM logged. Generous timeout + zero retries stops
the silent re-billing.
2026-09-22 18:05:27 -06:00
inference-bot 700154c3d4 Add custom per-token pricing (model_info) so spend and budget enforcement work 2026-09-14 14:28:13 -06:00
inference-bot da50b0438f Replace deprecated deepseek-v4-flash with deepseek-v4-1-flash (default + fallback) 2026-09-14 14:06:23 -06:00
inference-bot 878ab8fa8d Add GLM-5.3 Flash model 2026-09-09 14:18:14 -06:00
inference-bot e05fec0ca8 Switch to Tinfoil (TEE): add tinfoil-proxy sidecar for EHBP encryption, point LiteLLM at local proxy 2026-09-09 13:21:01 -06:00
inference-bot 2ce5c66b9b Switch to Ollama Cloud (OpenAI-compatible /v1 endpoint) for plumbing smoke test 2026-09-04 14:15:49 -06:00
inference-bot b9b4620e52 Initial commit: LiteLLM Cloudron app package for Inference Cooperative 2026-09-03 21:55:54 -06:00