From e8c7a559332933dc96b7d044886ad4312545f836 Mon Sep 17 00:00:00 2001 From: inference-co-op-bot Date: Fri, 2 Oct 2026 00:52:50 -0600 Subject: [PATCH] Add upgrade runbook + bake memory limit and LITELLM_MIGRATION_DIR into packaging The 2026-10-02 outage had three causes the runbook now prevents: - floating tag drift (README now mandates exact-tag pinning) - 256MB default cgroup cap OOM-killing the prisma migration engine (manifest now sets memoryLimit: 2147483648) - no _prisma_migrations ledger -> P3005 crash-loop on boot (start.sh now exports LITELLM_MIGRATION_DIR=/app/data/migrations) README 'Updating LiteLLM' section rewritten from the old (wrong) floating-tag procedure into a runbook with rules + verify checklist. --- CloudronManifest.json | 1 + README.md | 70 ++++++++++++++++++++++++++++++++----------- start.sh | 7 +++++ 3 files changed, 60 insertions(+), 18 deletions(-) diff --git a/CloudronManifest.json b/CloudronManifest.json index 256ab41..b4f5c8f 100644 --- a/CloudronManifest.json +++ b/CloudronManifest.json @@ -8,6 +8,7 @@ "upstreamVersion": "1.84.0", "healthCheckPath": "/health/readiness", "httpPort": 4000, + "memoryLimit": 2147483648, "manifestVersion": 2, "website": "https://litellm.ai", "contactEmail": "botbot@hermes", diff --git a/README.md b/README.md index e96d558..2ccf6d8 100644 --- a/README.md +++ b/README.md @@ -91,38 +91,72 @@ The following are auto-generated and should **not** be set manually: ## Updating LiteLLM -This is the routine maintenance path. LiteLLM releases frequently; updating is a two-step change. +This is the routine maintenance path. LiteLLM releases frequently. Upgrading is +a code change (pin a new version) plus a DB migration that runs on first boot — +expect a 15–20 min maintenance window and **never panic-restart mid-migration** +(each restart resets the clock). -### 1. Bump the base image tag +### Rules (learned from the 2026-10-02 outage) -In `Dockerfile`, change the pinned version: +- **Pin the exact version tag, never a floating tag.** Registries silently move + floating tags (`main-v1.74.0-stable`, `latest`) to newer digests, so a routine + rebuild changes the deployed version with no pin change and triggers a full + schema migration against an unprepared DB. Always + `FROM ghcr.io/berriai/litellm:vX.Y.Z`. +- **Give the app a real memory limit.** Cloudron's default when `memoryLimit: 0` + is a 256 MB cgroup cap; the proxy fits in 256 MB but Prisma's schema engine + does NOT, and gets SIGKILL'd (OOM, exit 137) mid-migration. This package sets + `memoryLimit: 2147483648` (2 GB) in `CloudronManifest.json` — don't lower it. +- **`LITELLM_MIGRATION_DIR=/app/data/migrations` is exported in `start.sh`.** The + database was created by `prisma db push`, so it has no `_prisma_migrations` + ledger; without a writable migration dir the boot migration hits P3005 + ("schema is not empty") and crash-loops. Keep it pointed at the persistent + volume. +- **Verify with a member-scope key, not only the master key.** Master-key + requests bypass the member auth path, so they can pass even on a half-migrated + schema while every member key 401s. See the checklist below. -```dockerfile -FROM ghcr.io/berriai/litellm:main-v1.74.0-stable -# ^^^^^^ bump this -``` +### Steps -### 2. Bump the manifest version +1. **Pick a target version** from the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases). Stay at or above v1.84.0 (the fix line for CVE-2026-35029 and CVE-2026-59822, the latter on CISA's KEV list). -In `CloudronManifest.json`, update: +2. **Bump the pin** in `Dockerfile`: -```json -"version": "1.0.0", -"upstreamVersion": "1.74.0" -``` + ```dockerfile + FROM ghcr.io/berriai/litellm:v1.84.0 + # ^^^^^^ exact tag, never main-*-stable + ``` -### 3. Redeploy +3. **Bump the manifest** in `CloudronManifest.json`: -```bash -cloudron update --app -``` + ```json + "upstreamVersion": "1.84.0" + ``` -Cloudron rebuilds from source and applies the update. Data (keys, spend logs) persists in PostgreSQL. +4. **Redeploy:** + + ```bash + cloudron update --app + ``` + + First boot after the version change runs `prisma migrate deploy` (or + baselines + `db push` if the ledger is missing) before binding port 4000. + Expect 15–20 min of ECONNREFUSED healthcheck errors — this is normal. + +### Verify after upgrading + +1. `GET https://gateway.inference.coop/health/liveliness` → 200 +2. Generate a temp key with the master key → if `/key/generate` 500s, the schema is incomplete; fix the migration before anything else +3. Use the **temp key** (not master) for `GET /v1/models` → full catalog +4. Use the **temp key** for a small `POST /v1/chat/completions` → content + usage +5. OPTIONS preflight from a foreign origin → 400; from `chat.inference.coop` → allowed +6. Cloudron app health = healthy, then delete the temp key ### Before updating - Check the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases) for breaking changes - The `start.sh` and `config.yaml` may need adjustment if LiteLLM changes its env-var or config schema +- The CORS env var is `LITELLM_CORS_ORIGINS` (upstream-native since 1.84.0); the old `LITELLM_CORS_ALLOWED_ORIGINS` + sed patch are dead — don't re-add them - Test on a staging instance if the version jump is large --- diff --git a/start.sh b/start.sh index 4168295..e792afd 100644 --- a/start.sh +++ b/start.sh @@ -61,6 +61,13 @@ export STORE_MODEL_IN_DB="True" # (for single-instance deployment this is fine) export DISABLE_SCHEMA_UPDATE="false" +# The DB was created by `prisma db push` (no _prisma_migrations ledger), so on +# first boot after a version change LiteLLM hits P3005 ("schema is not empty") +# and must baseline the existing schema. That baseline write needs a WRITABLE +# directory — the package dir is read-only under Cloudron's rootfs, so point it +# at the persistent volume. Without this the boot migration crash-loops. +export LITELLM_MIGRATION_DIR="/app/data/migrations" + # UI configuration — use Cloudron's domain export UI_USERNAME="admin" export UI_PASSWORD="${LITELLM_MASTER_KEY}"