191 lines
9.2 KiB
Markdown
191 lines
9.2 KiB
Markdown
# LiteLLM Gateway — Cloudron App
|
||
|
||
A Cloudron app package for [LiteLLM](https://litellm.ai/), the open-source AI gateway, configured for the **Inference Cooperative**.
|
||
|
||
LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.
|
||
|
||
**Current base image:** LiteLLM **v1.84.0**, pinned to the exact tag — the fix line for CVE-2026-35029 and CVE-2026-59822 (CISA KEV-listed). See [Updating LiteLLM](#updating-litellm).
|
||
|
||
## What this provides
|
||
|
||
- **LiteLLM proxy** with an OpenAI-compatible API on port 4000
|
||
- **Per-member virtual API keys** with budget caps and rate limits
|
||
- **Usage / spend tracking** per key, per model
|
||
- **Model routing** to multiple providers, chosen by member values (private / green / public)
|
||
- **Cloudron PostgreSQL** for keys, teams, and spend logs
|
||
- **Cloudron Redis** for rate limiting and caching
|
||
- **Automatic SSL, backups, and sandboxing** (Cloudron-managed)
|
||
- **Master key + salt key** auto-generated on first start, persisted in `/app/data`
|
||
- Hardened for the public internet: admin UI and API docs disabled at the proxy; CORS locked to the chat origin
|
||
|
||
## Architecture
|
||
|
||
```
|
||
Members → Open WebUI (chat) / own tools (API) → member portal (key injection)
|
||
→ LiteLLM gateway (this app)
|
||
├→ Tinfoil (TEE enclaves — architectural privacy)
|
||
├→ GreenPT (100% renewable energy)
|
||
└→ PublicAI (publicly developed, sovereign models)
|
||
```
|
||
|
||
LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to providers, each chosen by members for a different value. A [tinfoil-proxy](https://github.com/tinfoilsh/tinfoil-proxy) sidecar provides EHBP end-to-end encryption to Tinfoil's enclaves.
|
||
|
||
## Files
|
||
|
||
- `CloudronManifest.json` — Cloudron app manifest (addons, ports, memory limit, metadata)
|
||
- `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment (pinned to an exact tag)
|
||
- `start.sh` — startup script: wires Cloudron's Postgres/Redis, downloads the tinfoil-proxy sidecar, exports `LITELLM_MIGRATION_DIR`
|
||
- `config.yaml` — LiteLLM config: model catalog, provider routing, retry policy, custom per-token pricing
|
||
- `logo.png` — app icon
|
||
|
||
## Models
|
||
|
||
The catalog is defined in `config.yaml` and documented for members in [co-op/docs models](https://git.inference.coop/co-op/docs/src/branch/main/models.md) — that page is the single source of truth; this repo's `config.yaml` mirrors it. Models follow the `provider/model-name` convention so members can choose not just which model but which provider values it embodies (privacy, renewable energy, public development).
|
||
|
||
The catalog changes by member decision — expect it to evolve.
|
||
|
||
---
|
||
|
||
## Building and installing
|
||
|
||
This package uses Cloudron's **on-server build** (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server.
|
||
|
||
### Prerequisites
|
||
|
||
- The [Cloudron CLI](https://docs.cloudron.io/packaging/cli/) installed on a machine with access to the Cloudron server
|
||
- A Cloudron instance with the `inference.coop` domain
|
||
|
||
### Install
|
||
|
||
```bash
|
||
# 1. Install the Cloudron CLI (once)
|
||
npm install -g cloudron-cli
|
||
|
||
# 2. Log in to the Cloudron
|
||
cloudron login my.inference.coop
|
||
|
||
# 3. From this directory, install (builds on the server)
|
||
cloudron install --location gateway
|
||
```
|
||
|
||
`--location gateway` installs the app at `gateway.inference.coop`. Omit `--location` to be prompted.
|
||
|
||
### Configure
|
||
|
||
After install, set the provider API keys in Cloudron → app → Settings → Environment Variables:
|
||
|
||
- `TINFOIL_API_KEY` — Tinfoil (https://tinfoil.sh/)
|
||
- `GREENPT_API_KEY` — GreenPT (https://greenpt.ai/)
|
||
- `PUBLICAI_API_KEY` — PublicAI (https://publicai.co/)
|
||
|
||
The following are auto-generated and should **not** be set manually:
|
||
|
||
- `LITELLM_MASTER_KEY` — generated on first start, stored in `/app/data/.master_key`
|
||
- `LITELLM_SALT_KEY` — generated on first start, stored in `/app/data/.salt_key`
|
||
- `DATABASE_URL` — from Cloudron's PostgreSQL addon
|
||
- `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon
|
||
|
||
### Admin access
|
||
|
||
The LiteLLM admin UI and API docs are **disabled** on the public gateway (hardened by default; see `start.sh`). Administrative operations run through the [member portal](https://git.inference.coop/code/member-portal), which holds the master key server-side. If direct LiteLLM admin access is ever needed, re-enable the UI in `start.sh`, redeploy, and use the master key from `/app/data/.master_key` — then disable it again.
|
||
|
||
---
|
||
|
||
## Updating LiteLLM
|
||
|
||
This is the routine maintenance path. LiteLLM releases frequently. Upgrading is
|
||
a code change (pin a new version) plus a DB migration that runs on first boot —
|
||
expect a 15–20 min maintenance window and **never panic-restart mid-migration**
|
||
(each restart resets the clock).
|
||
|
||
### Rules (learned from the 2026-10-02 outage)
|
||
|
||
- **Pin the exact version tag, never a floating tag.** Registries silently move
|
||
floating tags (`main-v1.74.0-stable`, `latest`) to newer digests, so a routine
|
||
rebuild changes the deployed version with no pin change and triggers a full
|
||
schema migration against an unprepared DB. Always
|
||
`FROM ghcr.io/berriai/litellm:vX.Y.Z`.
|
||
- **Give the app a real memory limit.** Cloudron's default when `memoryLimit: 0`
|
||
is a 256 MB cgroup cap; the proxy fits in 256 MB but Prisma's schema engine
|
||
does NOT, and gets SIGKILL'd (OOM, exit 137) mid-migration. This package sets
|
||
`memoryLimit: 2147483648` (2 GB) in `CloudronManifest.json` — don't lower it.
|
||
- **`LITELLM_MIGRATION_DIR=/app/data/migrations` is exported in `start.sh`.** The
|
||
database was created by `prisma db push`, so it has no `_prisma_migrations`
|
||
ledger; without a writable migration dir the boot migration hits P3005
|
||
("schema is not empty") and crash-loops. Keep it pointed at the persistent
|
||
volume.
|
||
- **Verify with a member-scope key, not only the master key.** Master-key
|
||
requests bypass the member auth path, so they can pass even on a half-migrated
|
||
schema while every member key 401s. See the checklist below.
|
||
|
||
### Steps
|
||
|
||
1. **Pick a target version** from the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases). Stay at or above v1.84.0 (the fix line for CVE-2026-35029 and CVE-2026-59822, the latter on CISA's KEV list).
|
||
|
||
2. **Bump the pin** in `Dockerfile`:
|
||
|
||
```dockerfile
|
||
FROM ghcr.io/berriai/litellm:v1.84.0
|
||
# ^^^^^^ exact tag, never main-*-stable
|
||
```
|
||
|
||
3. **Bump the manifest** in `CloudronManifest.json`:
|
||
|
||
```json
|
||
"upstreamVersion": "1.84.0"
|
||
```
|
||
|
||
4. **Redeploy:**
|
||
|
||
```bash
|
||
cloudron update --app <app-id>
|
||
```
|
||
|
||
First boot after the version change runs `prisma migrate deploy` (or
|
||
baselines + `db push` if the ledger is missing) before binding port 4000.
|
||
Expect 15–20 min of ECONNREFUSED healthcheck errors — this is normal.
|
||
|
||
### Verify after upgrading
|
||
|
||
1. `GET https://gateway.inference.coop/health/liveliness` → 200
|
||
2. Generate a temp key with the master key → if `/key/generate` 500s, the schema is incomplete; fix the migration before anything else
|
||
3. Use the **temp key** (not master) for `GET /v1/models` → full catalog
|
||
4. Use the **temp key** for a small `POST /v1/chat/completions` → content + usage
|
||
5. OPTIONS preflight from a foreign origin → 400; from `chat.inference.coop` → allowed
|
||
6. Cloudron app health = healthy, then delete the temp key
|
||
|
||
### Before updating
|
||
|
||
- Check the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases) for breaking changes
|
||
- The `start.sh` and `config.yaml` may need adjustment if LiteLLM changes its env-var or config schema
|
||
- The CORS env var is `LITELLM_CORS_ORIGINS` (upstream-native since 1.84.0); the old `LITELLM_CORS_ALLOWED_ORIGINS` + sed patch are dead — don't re-add them
|
||
- Test on a staging instance if the version jump is large
|
||
|
||
---
|
||
|
||
## Environment variables
|
||
|
||
| Variable | Required | Purpose |
|
||
|----------|----------|---------|
|
||
| `TINFOIL_API_KEY` | Yes | Tinfoil provider (TEE inference) |
|
||
| `GREENPT_API_KEY` | Yes | GreenPT provider (renewable energy) |
|
||
| `PUBLICAI_API_KEY` | Yes | PublicAI provider (public models) |
|
||
| `LITELLM_CORS_ORIGINS` | Set by start.sh | CORS allowlist (chat origin only) |
|
||
| `LITELLM_MIGRATION_DIR` | Set by start.sh | Persistent Prisma migrations ledger |
|
||
|
||
## Connecting chat clients
|
||
|
||
Open WebUI does **not** talk to this gateway directly — it routes through the [member portal](https://git.inference.coop/code/member-portal), which authenticates the member and injects their personal LiteLLM key per request. Members building their own tools create API keys on the [member dashboard](https://dashboard.inference.coop) and use:
|
||
|
||
- **API Base URL:** `https://gateway.inference.coop/v1` (or the portal's proxy endpoint for chat-integrated tools)
|
||
- **API Key:** their own virtual key from the dashboard
|
||
|
||
## Related repos
|
||
|
||
- `code/member-portal` — membership middleware: provisioning, billing webhooks, key injection
|
||
- `code/member-dashboard` — member-facing usage and API-key management
|
||
- `code/admin-panel` — admin overview
|
||
- `co-op/website` — the public landing page
|
||
- `co-op/docs` — documentation, including the member-facing model list
|
||
- `co-op/design-assets` — logos, fonts, brand materials
|