The 2026-10-02 outage had three causes the runbook now prevents: - floating tag drift (README now mandates exact-tag pinning) - 256MB default cgroup cap OOM-killing the prisma migration engine (manifest now sets memoryLimit: 2147483648) - no _prisma_migrations ledger -> P3005 crash-loop on boot (start.sh now exports LITELLM_MIGRATION_DIR=/app/data/migrations) README 'Updating LiteLLM' section rewritten from the old (wrong) floating-tag procedure into a runbook with rules + verify checklist.
7.5 KiB
LiteLLM Gateway — Cloudron App
A Cloudron app package for LiteLLM, the open-source AI gateway, configured for the Inference Cooperative.
LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.
What this provides
- LiteLLM proxy with an OpenAI-compatible API on port 4000
- Per-member virtual API keys with budget caps and rate limits
- Usage / spend tracking per key, per model
- Model routing to Tinfoil (TEE-protected inference)
- Cloudron PostgreSQL for keys, teams, and spend logs
- Cloudron Redis for rate limiting and caching
- Automatic SSL, backups, and sandboxing (Cloudron-managed)
- Master key + salt key auto-generated on first start, persisted in
/app/data
Architecture
Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE)
LiteLLM is the control plane (keys, metering, routing). It does not run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs).
Files
CloudronManifest.json— Cloudron app manifest (addons, ports, metadata)Dockerfile— wraps LiteLLM's official image for Cloudron's environmentstart.sh— startup script that wires up Cloudron's Postgres/Redis and writes the default configconfig.yaml— default LiteLLM config (Tinfoil-only, two models)logo.png— app icon
Models (MVP)
Two models, both TEE-protected via Tinfoil:
| Model | Role | Price (in/out per M tokens) |
|---|---|---|
deepseek-v4-flash |
Default — best agent performance, 1M context | $0.30 / $0.70 |
gpt-oss-120b |
Fallback — cheapest, built for agentic workflows | $0.15 / $0.60 |
If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS.
Building and installing
This package uses Cloudron's on-server build (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server.
Prerequisites
- The Cloudron CLI installed on a machine with access to the Cloudron server
- A Cloudron instance with the
inference.coopdomain
Install
# 1. Install the Cloudron CLI (once)
npm install -g cloudron-cli
# 2. Log in to the Cloudron
cloudron login my.inference.coop
# 3. From this directory, install (builds on the server)
cloudron install --location gateway
--location gateway installs the app at gateway.inference.coop. Omit --location to be prompted.
Configure
After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables:
TINFOIL_API_KEY— your Tinfoil API key (from https://tinfoil.sh/)
The following are auto-generated and should not be set manually:
LITELLM_MASTER_KEY— generated on first start, stored in/app/data/.master_keyLITELLM_SALT_KEY— generated on first start, stored in/app/data/.salt_keyDATABASE_URL— from Cloudron's PostgreSQL addonREDIS_HOST,REDIS_PORT,REDIS_PASSWORD— from Cloudron's Redis addon
Access the admin UI
- Open
https://gateway.inference.coop/ui - Log in with username
adminand the master key as password - Find the master key:
cloudron exec --app <app-id> cat /app/data/.master_key
Updating LiteLLM
This is the routine maintenance path. LiteLLM releases frequently. Upgrading is a code change (pin a new version) plus a DB migration that runs on first boot — expect a 15–20 min maintenance window and never panic-restart mid-migration (each restart resets the clock).
Rules (learned from the 2026-10-02 outage)
- Pin the exact version tag, never a floating tag. Registries silently move
floating tags (
main-v1.74.0-stable,latest) to newer digests, so a routine rebuild changes the deployed version with no pin change and triggers a full schema migration against an unprepared DB. AlwaysFROM ghcr.io/berriai/litellm:vX.Y.Z. - Give the app a real memory limit. Cloudron's default when
memoryLimit: 0is a 256 MB cgroup cap; the proxy fits in 256 MB but Prisma's schema engine does NOT, and gets SIGKILL'd (OOM, exit 137) mid-migration. This package setsmemoryLimit: 2147483648(2 GB) inCloudronManifest.json— don't lower it. LITELLM_MIGRATION_DIR=/app/data/migrationsis exported instart.sh. The database was created byprisma db push, so it has no_prisma_migrationsledger; without a writable migration dir the boot migration hits P3005 ("schema is not empty") and crash-loops. Keep it pointed at the persistent volume.- Verify with a member-scope key, not only the master key. Master-key requests bypass the member auth path, so they can pass even on a half-migrated schema while every member key 401s. See the checklist below.
Steps
-
Pick a target version from the LiteLLM changelog. Stay at or above v1.84.0 (the fix line for CVE-2026-35029 and CVE-2026-59822, the latter on CISA's KEV list).
-
Bump the pin in
Dockerfile:FROM ghcr.io/berriai/litellm:v1.84.0 # ^^^^^^ exact tag, never main-*-stable -
Bump the manifest in
CloudronManifest.json:"upstreamVersion": "1.84.0" -
Redeploy:
cloudron update --app <app-id>First boot after the version change runs
prisma migrate deploy(or baselines +db pushif the ledger is missing) before binding port 4000. Expect 15–20 min of ECONNREFUSED healthcheck errors — this is normal.
Verify after upgrading
GET https://gateway.inference.coop/health/liveliness→ 200- Generate a temp key with the master key → if
/key/generate500s, the schema is incomplete; fix the migration before anything else - Use the temp key (not master) for
GET /v1/models→ full catalog - Use the temp key for a small
POST /v1/chat/completions→ content + usage - OPTIONS preflight from a foreign origin → 400; from
chat.inference.coop→ allowed - Cloudron app health = healthy, then delete the temp key
Before updating
- Check the LiteLLM changelog for breaking changes
- The
start.shandconfig.yamlmay need adjustment if LiteLLM changes its env-var or config schema - The CORS env var is
LITELLM_CORS_ORIGINS(upstream-native since 1.84.0); the oldLITELLM_CORS_ALLOWED_ORIGINS+ sed patch are dead — don't re-add them - Test on a staging instance if the version jump is large
Environment variables
| Variable | Required | Purpose |
|---|---|---|
TINFOIL_API_KEY |
Yes | Tinfoil API key for TEE inference |
Connecting Open WebUI
In Open WebUI settings:
- API Base URL:
https://gateway.inference.coop/v1 - API Key: a virtual key created in LiteLLM's Admin UI
OIDC / SSO (future)
The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the "oidc": {} addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso.
Related repos
co-op/org-dev— the full inference.coop design doc and deploy scriptsco-op/website— the public landing pageco-op/design-assets— logos, fonts, brand materials