Files

191 lines
9.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# LiteLLM Gateway — Cloudron App
A Cloudron app package for [LiteLLM](https://litellm.ai/), the open-source AI gateway, configured for the **Inference Cooperative**.
LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.
**Current base image:** LiteLLM **v1.103.2**, pinned to the exact tag (fix line for CVE-2026-35029 / CVE-2026-59822, CISA KEV-listed; also carries the streaming-usage fix for prompt_tokens_details). See [Updating LiteLLM](#updating-litellm).
## What this provides
- **LiteLLM proxy** with an OpenAI-compatible API on port 4000
- **Per-member virtual API keys** with budget caps and rate limits
- **Usage / spend tracking** per key, per model
- **Model routing** to multiple providers, chosen by member values (private / green / public)
- **Cloudron PostgreSQL** for keys, teams, and spend logs
- **Cloudron Redis** for rate limiting and caching
- **Automatic SSL, backups, and sandboxing** (Cloudron-managed)
- **Master key + salt key** auto-generated on first start, persisted in `/app/data`
- Hardened for the public internet: admin UI and API docs disabled at the proxy; CORS locked to the chat origin
## Architecture
```
Members → Open WebUI (chat) / own tools (API) → member portal (key injection)
→ LiteLLM gateway (this app)
├→ Tinfoil (TEE enclaves — architectural privacy)
├→ GreenPT (100% renewable energy)
└→ PublicAI (publicly developed, sovereign models)
```
LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to providers, each chosen by members for a different value. A [tinfoil-proxy](https://github.com/tinfoilsh/tinfoil-proxy) sidecar provides EHBP end-to-end encryption to Tinfoil's enclaves.
## Files
- `CloudronManifest.json` — Cloudron app manifest (addons, ports, memory limit, metadata)
- `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment (pinned to an exact tag)
- `start.sh` — startup script: wires Cloudron's Postgres/Redis, downloads the tinfoil-proxy sidecar, exports `LITELLM_MIGRATION_DIR`
- `config.yaml` — LiteLLM config: model catalog, provider routing, retry policy, custom per-token pricing
- `logo.png` — app icon
## Models
The catalog is defined in `config.yaml` and documented for members in [co-op/docs models](https://git.inference.coop/co-op/docs/src/branch/main/models.md) — that page is the single source of truth; this repo's `config.yaml` mirrors it. Models follow the `provider/model-name` convention so members can choose not just which model but which provider values it embodies (privacy, renewable energy, public development).
The catalog changes by member decision — expect it to evolve.
---
## Building and installing
This package uses Cloudron's **on-server build** (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server.
### Prerequisites
- The [Cloudron CLI](https://docs.cloudron.io/packaging/cli/) installed on a machine with access to the Cloudron server
- A Cloudron instance with the `inference.coop` domain
### Install
```bash
# 1. Install the Cloudron CLI (once)
npm install -g cloudron-cli
# 2. Log in to the Cloudron
cloudron login my.inference.coop
# 3. From this directory, install (builds on the server)
cloudron install --location gateway
```
`--location gateway` installs the app at `gateway.inference.coop`. Omit `--location` to be prompted.
### Configure
After install, set the provider API keys in Cloudron → app → Settings → Environment Variables:
- `TINFOIL_API_KEY` — Tinfoil (https://tinfoil.sh/)
- `GREENPT_API_KEY` — GreenPT (https://greenpt.ai/)
- `PUBLICAI_API_KEY` — PublicAI (https://publicai.co/)
The following are auto-generated and should **not** be set manually:
- `LITELLM_MASTER_KEY` — generated on first start, stored in `/app/data/.master_key`
- `LITELLM_SALT_KEY` — generated on first start, stored in `/app/data/.salt_key`
- `DATABASE_URL` — from Cloudron's PostgreSQL addon
- `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon
### Admin access
The LiteLLM admin UI and API docs are **disabled** on the public gateway (hardened by default; see `start.sh`). Administrative operations run through the [member portal](https://git.inference.coop/code/member-portal), which holds the master key server-side. If direct LiteLLM admin access is ever needed, re-enable the UI in `start.sh`, redeploy, and use the master key from `/app/data/.master_key` — then disable it again.
---
## Updating LiteLLM
This is the routine maintenance path. LiteLLM releases frequently. Upgrading is
a code change (pin a new version) plus a DB migration that runs on first boot —
expect a 15–20 min maintenance window and **never panic-restart mid-migration**
(each restart resets the clock).
### Rules (learned from the 2026-10-02 outage)
- **Pin the exact version tag, never a floating tag.** Registries silently move
floating tags (`main-v1.74.0-stable`, `latest`) to newer digests, so a routine
rebuild changes the deployed version with no pin change and triggers a full
schema migration against an unprepared DB. Always
`FROM ghcr.io/berriai/litellm:vX.Y.Z`.
- **Give the app a real memory limit.** Cloudron's default when `memoryLimit: 0`
is a 256 MB cgroup cap; the proxy fits in 256 MB but Prisma's schema engine
does NOT, and gets SIGKILL'd (OOM, exit 137) mid-migration. This package sets
`memoryLimit: 2147483648` (2 GB) in `CloudronManifest.json` — don't lower it.
- **`LITELLM_MIGRATION_DIR=/app/data/migrations` is exported in `start.sh`.** The
database was created by `prisma db push`, so it has no `_prisma_migrations`
ledger; without a writable migration dir the boot migration hits P3005
("schema is not empty") and crash-loops. Keep it pointed at the persistent
volume.
- **Verify with a member-scope key, not only the master key.** Master-key
requests bypass the member auth path, so they can pass even on a half-migrated
schema while every member key 401s. See the checklist below.
### Steps
1. **Pick a target version** from the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases). Stay at or above v1.84.0 (the fix line for CVE-2026-35029 and CVE-2026-59822, the latter on CISA's KEV list).
2. **Bump the pin** in `Dockerfile`:
```dockerfile
FROM ghcr.io/berriai/litellm:v1.84.0
# ^^^^^^ exact tag, never main-*-stable
```
3. **Bump the manifest** in `CloudronManifest.json`:
```json
"upstreamVersion": "1.84.0"
```
4. **Redeploy:**
```bash
cloudron update --app <app-id>
```
First boot after the version change runs `prisma migrate deploy` (or
baselines + `db push` if the ledger is missing) before binding port 4000.
Expect 15–20 min of ECONNREFUSED healthcheck errors — this is normal.
### Verify after upgrading
1. `GET https://gateway.inference.coop/health/liveliness` → 200
2. Generate a temp key with the master key → if `/key/generate` 500s, the schema is incomplete; fix the migration before anything else
3. Use the **temp key** (not master) for `GET /v1/models` → full catalog
4. Use the **temp key** for a small `POST /v1/chat/completions` → content + usage
5. OPTIONS preflight from a foreign origin → 400; from `chat.inference.coop` → allowed
6. Cloudron app health = healthy, then delete the temp key
### Before updating
- Check the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases) for breaking changes
- The `start.sh` and `config.yaml` may need adjustment if LiteLLM changes its env-var or config schema
- The CORS env var is `LITELLM_CORS_ORIGINS` (upstream-native since 1.84.0); the old `LITELLM_CORS_ALLOWED_ORIGINS` + sed patch are dead — don't re-add them
- Test on a staging instance if the version jump is large
---
## Environment variables
| Variable | Required | Purpose |
|----------|----------|---------|
| `TINFOIL_API_KEY` | Yes | Tinfoil provider (TEE inference) |
| `GREENPT_API_KEY` | Yes | GreenPT provider (renewable energy) |
| `PUBLICAI_API_KEY` | Yes | PublicAI provider (public models) |
| `LITELLM_CORS_ORIGINS` | Set by start.sh | CORS allowlist (chat origin only) |
| `LITELLM_MIGRATION_DIR` | Set by start.sh | Persistent Prisma migrations ledger |
## Connecting chat clients
Open WebUI does **not** talk to this gateway directly — it routes through the [member portal](https://git.inference.coop/code/member-portal), which authenticates the member and injects their personal LiteLLM key per request. Members building their own tools create API keys on the [member dashboard](https://dashboard.inference.coop) and use:
- **API Base URL:** `https://gateway.inference.coop/v1` (or the portal's proxy endpoint for chat-integrated tools)
- **API Key:** their own virtual key from the dashboard
## Related repos
- `code/member-portal` — membership middleware: provisioning, billing webhooks, key injection
- `code/member-dashboard` — member-facing usage and API-key management
- `code/admin-panel` — admin overview
- `co-op/website` — the public landing page
- `co-op/docs` — documentation, including the member-facing model list
- `co-op/design-assets` — logos, fonts, brand materials