Files
litellm/README.md
T
inference-co-op-bot e8c7a55933 Add upgrade runbook + bake memory limit and LITELLM_MIGRATION_DIR into packaging
The 2026-10-02 outage had three causes the runbook now prevents:
- floating tag drift (README now mandates exact-tag pinning)
- 256MB default cgroup cap OOM-killing the prisma migration engine
  (manifest now sets memoryLimit: 2147483648)
- no _prisma_migrations ledger -> P3005 crash-loop on boot
  (start.sh now exports LITELLM_MIGRATION_DIR=/app/data/migrations)

README 'Updating LiteLLM' section rewritten from the old (wrong)
floating-tag procedure into a runbook with rules + verify checklist.
2026-10-02 00:52:50 -06:00

186 lines
7.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# LiteLLM Gateway — Cloudron App
A Cloudron app package for [LiteLLM](https://litellm.ai/), the open-source AI gateway, configured for the **Inference Cooperative**.
LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.
## What this provides
- **LiteLLM proxy** with an OpenAI-compatible API on port 4000
- **Per-member virtual API keys** with budget caps and rate limits
- **Usage / spend tracking** per key, per model
- **Model routing** to Tinfoil (TEE-protected inference)
- **Cloudron PostgreSQL** for keys, teams, and spend logs
- **Cloudron Redis** for rate limiting and caching
- **Automatic SSL, backups, and sandboxing** (Cloudron-managed)
- **Master key + salt key** auto-generated on first start, persisted in `/app/data`
## Architecture
```
Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE)
```
LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs).
## Files
- `CloudronManifest.json` — Cloudron app manifest (addons, ports, metadata)
- `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment
- `start.sh` — startup script that wires up Cloudron's Postgres/Redis and writes the default config
- `config.yaml` — default LiteLLM config (Tinfoil-only, two models)
- `logo.png` — app icon
## Models (MVP)
Two models, both TEE-protected via Tinfoil:
| Model | Role | Price (in/out per M tokens) |
|-------|------|------------------------------|
| `deepseek-v4-flash` | Default — best agent performance, 1M context | $0.30 / $0.70 |
| `gpt-oss-120b` | Fallback — cheapest, built for agentic workflows | $0.15 / $0.60 |
If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS.
---
## Building and installing
This package uses Cloudron's **on-server build** (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server.
### Prerequisites
- The [Cloudron CLI](https://docs.cloudron.io/packaging/cli/) installed on a machine with access to the Cloudron server
- A Cloudron instance with the `inference.coop` domain
### Install
```bash
# 1. Install the Cloudron CLI (once)
npm install -g cloudron-cli
# 2. Log in to the Cloudron
cloudron login my.inference.coop
# 3. From this directory, install (builds on the server)
cloudron install --location gateway
```
`--location gateway` installs the app at `gateway.inference.coop`. Omit `--location` to be prompted.
### Configure
After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables:
- `TINFOIL_API_KEY` — your Tinfoil API key (from https://tinfoil.sh/)
The following are auto-generated and should **not** be set manually:
- `LITELLM_MASTER_KEY` — generated on first start, stored in `/app/data/.master_key`
- `LITELLM_SALT_KEY` — generated on first start, stored in `/app/data/.salt_key`
- `DATABASE_URL` — from Cloudron's PostgreSQL addon
- `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon
### Access the admin UI
1. Open `https://gateway.inference.coop/ui`
2. Log in with username `admin` and the master key as password
3. Find the master key: `cloudron exec --app <app-id> cat /app/data/.master_key`
---
## Updating LiteLLM
This is the routine maintenance path. LiteLLM releases frequently. Upgrading is
a code change (pin a new version) plus a DB migration that runs on first boot —
expect a 15–20 min maintenance window and **never panic-restart mid-migration**
(each restart resets the clock).
### Rules (learned from the 2026-10-02 outage)
- **Pin the exact version tag, never a floating tag.** Registries silently move
floating tags (`main-v1.74.0-stable`, `latest`) to newer digests, so a routine
rebuild changes the deployed version with no pin change and triggers a full
schema migration against an unprepared DB. Always
`FROM ghcr.io/berriai/litellm:vX.Y.Z`.
- **Give the app a real memory limit.** Cloudron's default when `memoryLimit: 0`
is a 256 MB cgroup cap; the proxy fits in 256 MB but Prisma's schema engine
does NOT, and gets SIGKILL'd (OOM, exit 137) mid-migration. This package sets
`memoryLimit: 2147483648` (2 GB) in `CloudronManifest.json` — don't lower it.
- **`LITELLM_MIGRATION_DIR=/app/data/migrations` is exported in `start.sh`.** The
database was created by `prisma db push`, so it has no `_prisma_migrations`
ledger; without a writable migration dir the boot migration hits P3005
("schema is not empty") and crash-loops. Keep it pointed at the persistent
volume.
- **Verify with a member-scope key, not only the master key.** Master-key
requests bypass the member auth path, so they can pass even on a half-migrated
schema while every member key 401s. See the checklist below.
### Steps
1. **Pick a target version** from the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases). Stay at or above v1.84.0 (the fix line for CVE-2026-35029 and CVE-2026-59822, the latter on CISA's KEV list).
2. **Bump the pin** in `Dockerfile`:
```dockerfile
FROM ghcr.io/berriai/litellm:v1.84.0
# ^^^^^^ exact tag, never main-*-stable
```
3. **Bump the manifest** in `CloudronManifest.json`:
```json
"upstreamVersion": "1.84.0"
```
4. **Redeploy:**
```bash
cloudron update --app <app-id>
```
First boot after the version change runs `prisma migrate deploy` (or
baselines + `db push` if the ledger is missing) before binding port 4000.
Expect 15–20 min of ECONNREFUSED healthcheck errors — this is normal.
### Verify after upgrading
1. `GET https://gateway.inference.coop/health/liveliness` → 200
2. Generate a temp key with the master key → if `/key/generate` 500s, the schema is incomplete; fix the migration before anything else
3. Use the **temp key** (not master) for `GET /v1/models` → full catalog
4. Use the **temp key** for a small `POST /v1/chat/completions` → content + usage
5. OPTIONS preflight from a foreign origin → 400; from `chat.inference.coop` → allowed
6. Cloudron app health = healthy, then delete the temp key
### Before updating
- Check the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases) for breaking changes
- The `start.sh` and `config.yaml` may need adjustment if LiteLLM changes its env-var or config schema
- The CORS env var is `LITELLM_CORS_ORIGINS` (upstream-native since 1.84.0); the old `LITELLM_CORS_ALLOWED_ORIGINS` + sed patch are dead — don't re-add them
- Test on a staging instance if the version jump is large
---
## Environment variables
| Variable | Required | Purpose |
|----------|----------|---------|
| `TINFOIL_API_KEY` | Yes | Tinfoil API key for TEE inference |
## Connecting Open WebUI
In Open WebUI settings:
- **API Base URL:** `https://gateway.inference.coop/v1`
- **API Key:** a virtual key created in LiteLLM's Admin UI
## OIDC / SSO (future)
The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the `"oidc": {}` addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso.
## Related repos
- `co-op/org-dev` — the full inference.coop design doc and deploy scripts
- `co-op/website` — the public landing page
- `co-op/design-assets` — logos, fonts, brand materials