# LiteLLM Gateway — Cloudron App A Cloudron app package for [LiteLLM](https://litellm.ai/), the open-source AI gateway, configured for the **Inference Cooperative**. LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint. ## What this provides - **LiteLLM proxy** with an OpenAI-compatible API on port 4000 - **Per-member virtual API keys** with budget caps and rate limits - **Usage / spend tracking** per key, per model - **Model routing** to Tinfoil (TEE-protected inference) - **Cloudron PostgreSQL** for keys, teams, and spend logs - **Cloudron Redis** for rate limiting and caching - **Automatic SSL, backups, and sandboxing** (Cloudron-managed) - **Master key + salt key** auto-generated on first start, persisted in `/app/data` ## Architecture ``` Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE) ``` LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs). ## Files - `CloudronManifest.json` — Cloudron app manifest (addons, ports, metadata) - `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment - `start.sh` — startup script that wires up Cloudron's Postgres/Redis and writes the default config - `config.yaml` — default LiteLLM config (Tinfoil-only, two models) - `logo.png` — app icon ## Models (MVP) Two models, both TEE-protected via Tinfoil: | Model | Role | Price (in/out per M tokens) | |-------|------|------------------------------| | `deepseek-v4-flash` | Default — best agent performance, 1M context | $0.30 / $0.70 | | `gpt-oss-120b` | Fallback — cheapest, built for agentic workflows | $0.15 / $0.60 | If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS. --- ## Building and installing This package uses Cloudron's **on-server build** (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server. ### Prerequisites - The [Cloudron CLI](https://docs.cloudron.io/packaging/cli/) installed on a machine with access to the Cloudron server - A Cloudron instance with the `inference.coop` domain ### Install ```bash # 1. Install the Cloudron CLI (once) npm install -g cloudron-cli # 2. Log in to the Cloudron cloudron login my.inference.coop # 3. From this directory, install (builds on the server) cloudron install --location gateway ``` `--location gateway` installs the app at `gateway.inference.coop`. Omit `--location` to be prompted. ### Configure After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables: - `TINFOIL_API_KEY` — your Tinfoil API key (from https://tinfoil.sh/) The following are auto-generated and should **not** be set manually: - `LITELLM_MASTER_KEY` — generated on first start, stored in `/app/data/.master_key` - `LITELLM_SALT_KEY` — generated on first start, stored in `/app/data/.salt_key` - `DATABASE_URL` — from Cloudron's PostgreSQL addon - `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon ### Access the admin UI 1. Open `https://gateway.inference.coop/ui` 2. Log in with username `admin` and the master key as password 3. Find the master key: `cloudron exec --app cat /app/data/.master_key` --- ## Updating LiteLLM This is the routine maintenance path. LiteLLM releases frequently. Upgrading is a code change (pin a new version) plus a DB migration that runs on first boot — expect a 15–20 min maintenance window and **never panic-restart mid-migration** (each restart resets the clock). ### Rules (learned from the 2026-10-02 outage) - **Pin the exact version tag, never a floating tag.** Registries silently move floating tags (`main-v1.74.0-stable`, `latest`) to newer digests, so a routine rebuild changes the deployed version with no pin change and triggers a full schema migration against an unprepared DB. Always `FROM ghcr.io/berriai/litellm:vX.Y.Z`. - **Give the app a real memory limit.** Cloudron's default when `memoryLimit: 0` is a 256 MB cgroup cap; the proxy fits in 256 MB but Prisma's schema engine does NOT, and gets SIGKILL'd (OOM, exit 137) mid-migration. This package sets `memoryLimit: 2147483648` (2 GB) in `CloudronManifest.json` — don't lower it. - **`LITELLM_MIGRATION_DIR=/app/data/migrations` is exported in `start.sh`.** The database was created by `prisma db push`, so it has no `_prisma_migrations` ledger; without a writable migration dir the boot migration hits P3005 ("schema is not empty") and crash-loops. Keep it pointed at the persistent volume. - **Verify with a member-scope key, not only the master key.** Master-key requests bypass the member auth path, so they can pass even on a half-migrated schema while every member key 401s. See the checklist below. ### Steps 1. **Pick a target version** from the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases). Stay at or above v1.84.0 (the fix line for CVE-2026-35029 and CVE-2026-59822, the latter on CISA's KEV list). 2. **Bump the pin** in `Dockerfile`: ```dockerfile FROM ghcr.io/berriai/litellm:v1.84.0 # ^^^^^^ exact tag, never main-*-stable ``` 3. **Bump the manifest** in `CloudronManifest.json`: ```json "upstreamVersion": "1.84.0" ``` 4. **Redeploy:** ```bash cloudron update --app ``` First boot after the version change runs `prisma migrate deploy` (or baselines + `db push` if the ledger is missing) before binding port 4000. Expect 15–20 min of ECONNREFUSED healthcheck errors — this is normal. ### Verify after upgrading 1. `GET https://gateway.inference.coop/health/liveliness` → 200 2. Generate a temp key with the master key → if `/key/generate` 500s, the schema is incomplete; fix the migration before anything else 3. Use the **temp key** (not master) for `GET /v1/models` → full catalog 4. Use the **temp key** for a small `POST /v1/chat/completions` → content + usage 5. OPTIONS preflight from a foreign origin → 400; from `chat.inference.coop` → allowed 6. Cloudron app health = healthy, then delete the temp key ### Before updating - Check the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases) for breaking changes - The `start.sh` and `config.yaml` may need adjustment if LiteLLM changes its env-var or config schema - The CORS env var is `LITELLM_CORS_ORIGINS` (upstream-native since 1.84.0); the old `LITELLM_CORS_ALLOWED_ORIGINS` + sed patch are dead — don't re-add them - Test on a staging instance if the version jump is large --- ## Environment variables | Variable | Required | Purpose | |----------|----------|---------| | `TINFOIL_API_KEY` | Yes | Tinfoil API key for TEE inference | ## Connecting Open WebUI In Open WebUI settings: - **API Base URL:** `https://gateway.inference.coop/v1` - **API Key:** a virtual key created in LiteLLM's Admin UI ## OIDC / SSO (future) The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the `"oidc": {}` addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso. ## Related repos - `co-op/org-dev` — the full inference.coop design doc and deploy scripts - `co-op/website` — the public landing page - `co-op/design-assets` — logos, fonts, brand materials