Commit Graph
12 Commits
Author SHA1 Message Date
Inferencebot 23f73cf6ae README: current architecture (3 providers, portal key injection, hardening), remove MVP model list (point to co-op/docs as source of truth), all provider env vars, member-key verification kept 2026-10-02 01:08:34 -06:00
inference-co-op-bot e8c7a55933 Add upgrade runbook + bake memory limit and LITELLM_MIGRATION_DIR into packaging
The 2026-10-02 outage had three causes the runbook now prevents:
- floating tag drift (README now mandates exact-tag pinning)
- 256MB default cgroup cap OOM-killing the prisma migration engine
  (manifest now sets memoryLimit: 2147483648)
- no _prisma_migrations ledger -> P3005 crash-loop on boot
  (start.sh now exports LITELLM_MIGRATION_DIR=/app/data/migrations)

README 'Updating LiteLLM' section rewritten from the old (wrong)
floating-tag procedure into a runbook with rules + verify checklist.
2026-10-02 00:52:50 -06:00
inference-bot fd34e33508 Sync packaging to live v1.84.0 deployment (CVE-2026-35029/59822 fix line)
- Dockerfile: pin ghcr.io/berriai/litellm:v1.84.0 (was floating
  main-v1.74.0-stable, which silently moved to a newer digest).
  v1.84.0 is the fix line for CVE-2026-35029 (auth bypass on
  /config/update, fixed 1.83.0) and CVE-2026-59822 (MCP session
  auth bypass, fixed 1.84.0, CISA KEV).
- Drop the v1.74.0-era sed CORS patch: upstream v1.84.0 moved the
  file and now reads LITELLM_CORS_ORIGINS natively.
- start.sh: export LITELLM_CORS_ORIGINS (the native v1.84.0 var)
  instead of LITELLM_CORS_ALLOWED_ORIGINS (the old sed-patch var
  that v1.84.0 ignores) — keeps CORS locked to chat.inference.coop.
- CloudronManifest: upstreamVersion 1.74.0 -> 1.84.0.

Live gateway already runs v1.84.0 (image digest
sha256:dc532d896ba8..., built 2026-10-02 04:50 UTC); this commit
makes the repo mirror the deployed packaging per the no-drift rule.
LITELLM_CORS_ORIGINS also set as a Cloudron env var on the app so
the running container honours it without a rebuild.
2026-10-01 23:32:43 -06:00
inference-bot 8f42be988e Add Tinfoil audio models (whisper STT, voxtral-tts) for voice mode + API use 2026-09-27 11:10:51 -06:00
inference-bot 5ea2256776 Add greenpt/glm-5.3 (full flagship, $1.10/$4.40 per 1M) 2026-09-26 16:52:21 -06:00
inference-bot 057507d67c Add PublicAI Apertus models (publicai/apertus-v1.5-8b, -70b) 2026-09-25 15:56:08 -06:00
inference-bot 2369c3e848 Add greenpt/glm-5.3-flash (/usr/bin/bash.11//usr/bin/bash.44 per 1M) 2026-09-24 10:43:14 -06:00
inference-bot f86e005062 Rename Tinfoil models to tinfoil/* provider prefix 2026-09-24 08:16:16 -06:00
inference-bot 67192de232 Document pricing sync with co-op/docs/models.md 2026-09-23 21:08:33 -06:00
inference-bot 09ee36014c Add GreenPT models (greenpt/green-r, greenpt/green-l) via provider/model naming convention 2026-09-23 20:10:18 -06:00
inference-bot d1f57c45d5 Retry connection errors but never re-bill timeouts
num_retries=2 (retries transient connection errors), retry_policy with
TimeoutErrorRetries=0 (a timed-out generation is already billed by Tinfoil, so
re-hitting it re-bills). Combined with timeout=180, this hedges connection
errors without the silent re-billing that caused the Tinfoil spend gap.
2026-09-22 18:18:07 -06:00
inference-bot 576885ed21 Fix Tinfoil spend gap: raise timeout 30->180s, num_retries 2->0
DeepSeek V4.1 Flash is a reasoning model; 30s timeout caused LiteLLM to
abandon in-flight generations (falling back to GPT-OSS) while Tinfoil still
completed and billed them. num_retries re-hit Tinfoil up to 3x. Result: Tinfoil
billed ~2x the tokens LiteLLM logged. Generous timeout + zero retries stops
the silent re-billing.
2026-09-22 18:05:27 -06:00