Add upgrade runbook + bake memory limit and LITELLM_MIGRATION_DIR into packaging

The 2026-10-02 outage had three causes the runbook now prevents:
- floating tag drift (README now mandates exact-tag pinning)
- 256MB default cgroup cap OOM-killing the prisma migration engine
  (manifest now sets memoryLimit: 2147483648)
- no _prisma_migrations ledger -> P3005 crash-loop on boot
  (start.sh now exports LITELLM_MIGRATION_DIR=/app/data/migrations)

README 'Updating LiteLLM' section rewritten from the old (wrong)
floating-tag procedure into a runbook with rules + verify checklist.
This commit is contained in:
inference-co-op-bot committed 2026-10-02 00:52:50 -06:00
1 parent fd34e33508
commit e8c7a55933
3 files changed
+60 -18

No files matched your search

+1
View File
@@ -8,6 +8,7 @@
"upstreamVersion": "1.84.0",
"healthCheckPath": "/health/readiness",
"httpPort": 4000,
"memoryLimit": 2147483648,
"manifestVersion": 2,
"website": "https://litellm.ai",
"contactEmail": "botbot@hermes",
+52 -18
View File
@@ -91,38 +91,72 @@ The following are auto-generated and should **not** be set manually:
## Updating LiteLLM
This is the routine maintenance path. LiteLLM releases frequently; updating is a two-step change.
This is the routine maintenance path. LiteLLM releases frequently. Upgrading is
a code change (pin a new version) plus a DB migration that runs on first boot —
expect a 15–20 min maintenance window and **never panic-restart mid-migration**
(each restart resets the clock).
### 1. Bump the base image tag
### Rules (learned from the 2026-10-02 outage)
In `Dockerfile`, change the pinned version:
- **Pin the exact version tag, never a floating tag.** Registries silently move
floating tags (`main-v1.74.0-stable`, `latest`) to newer digests, so a routine
rebuild changes the deployed version with no pin change and triggers a full
schema migration against an unprepared DB. Always
`FROM ghcr.io/berriai/litellm:vX.Y.Z`.
- **Give the app a real memory limit.** Cloudron's default when `memoryLimit: 0`
is a 256 MB cgroup cap; the proxy fits in 256 MB but Prisma's schema engine
does NOT, and gets SIGKILL'd (OOM, exit 137) mid-migration. This package sets
`memoryLimit: 2147483648` (2 GB) in `CloudronManifest.json` — don't lower it.
- **`LITELLM_MIGRATION_DIR=/app/data/migrations` is exported in `start.sh`.** The
database was created by `prisma db push`, so it has no `_prisma_migrations`
ledger; without a writable migration dir the boot migration hits P3005
("schema is not empty") and crash-loops. Keep it pointed at the persistent
volume.
- **Verify with a member-scope key, not only the master key.** Master-key
requests bypass the member auth path, so they can pass even on a half-migrated
schema while every member key 401s. See the checklist below.
```dockerfile
FROM ghcr.io/berriai/litellm:main-v1.74.0-stable
# ^^^^^^ bump this
```
### Steps
### 2. Bump the manifest version
1. **Pick a target version** from the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases). Stay at or above v1.84.0 (the fix line for CVE-2026-35029 and CVE-2026-59822, the latter on CISA's KEV list).
In `CloudronManifest.json`, update:
2. **Bump the pin** in `Dockerfile`:
```json
"version": "1.0.0",
"upstreamVersion": "1.74.0"
```
```dockerfile
FROM ghcr.io/berriai/litellm:v1.84.0
# ^^^^^^ exact tag, never main-*-stable
```
### 3. Redeploy
3. **Bump the manifest** in `CloudronManifest.json`:
```bash
cloudron update --app <app-id>
```
```json
"upstreamVersion": "1.84.0"
```
Cloudron rebuilds from source and applies the update. Data (keys, spend logs) persists in PostgreSQL.
4. **Redeploy:**
```bash
cloudron update --app <app-id>
```
First boot after the version change runs `prisma migrate deploy` (or
baselines + `db push` if the ledger is missing) before binding port 4000.
Expect 15–20 min of ECONNREFUSED healthcheck errors — this is normal.
### Verify after upgrading
1. `GET https://gateway.inference.coop/health/liveliness` → 200
2. Generate a temp key with the master key → if `/key/generate` 500s, the schema is incomplete; fix the migration before anything else
3. Use the **temp key** (not master) for `GET /v1/models` → full catalog
4. Use the **temp key** for a small `POST /v1/chat/completions` → content + usage
5. OPTIONS preflight from a foreign origin → 400; from `chat.inference.coop` → allowed
6. Cloudron app health = healthy, then delete the temp key
### Before updating
- Check the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases) for breaking changes
- The `start.sh` and `config.yaml` may need adjustment if LiteLLM changes its env-var or config schema
- The CORS env var is `LITELLM_CORS_ORIGINS` (upstream-native since 1.84.0); the old `LITELLM_CORS_ALLOWED_ORIGINS` + sed patch are dead — don't re-add them
- Test on a staging instance if the version jump is large
---
+7
View File
@@ -61,6 +61,13 @@ export STORE_MODEL_IN_DB="True"
# (for single-instance deployment this is fine)
export DISABLE_SCHEMA_UPDATE="false"
# The DB was created by `prisma db push` (no _prisma_migrations ledger), so on
# first boot after a version change LiteLLM hits P3005 ("schema is not empty")
# and must baseline the existing schema. That baseline write needs a WRITABLE
# directory — the package dir is read-only under Cloudron's rootfs, so point it
# at the persistent volume. Without this the boot migration crash-loops.
export LITELLM_MIGRATION_DIR="/app/data/migrations"
# UI configuration — use Cloudron's domain
export UI_USERNAME="admin"
export UI_PASSWORD="${LITELLM_MASTER_KEY}"