Upgrade LiteLLM v1.84.0 → v1.103.2 with ~1–2 min downtime (pre-migrate + swap) #2

Closed
opened 2026-10-02 18:15:01 +00:00 by inference-bot · 2 comments
Owner

LiteLLM upgrade plan: v1.84.0 → v1.103.2

Goal: upgrade the gateway with ~1–2 minutes of downtime instead of the multi-hour outage we got on 2026-10-02. Status: plan only — nothing executed yet. Two decisions needed from the operator before execution (see "Open decisions").

Why upgrade now

  • We sit 50 stable releases / 63 schema migrations behind (v1.84.0 → v1.103.2, latest stable, 2026-10-01).
  • Not security-driven: v1.84.0 is already the fix line for CVE-2026-35029, CVE-2026-59822 (KEV), CVE-2026-49468, and CVE-2026-42208 (9.3 pre-auth SQLi). This is currency, not urgency.

Verified current state

  • Deployed: v1.84.0 exact-tag pinned, 2 GB memory limit, LITELLM_MIGRATION_DIR=/app/data/migrations (all three 2026-10-02 outage fixes baked into packaging, commit e8c7a55).
  • Prisma ledger complete through 20260501195714_managed_resource_team_owner = exactly v1.84.0's last migration. No drift.
  • DB sizes (why this is easy for us): SpendLogs 5,291 rows · DailyTeamSpend 232 · TeamTable 37 · VerificationToken 57. Migration runtime will be seconds-to-a-minute, not hours.
  • Model catalog: 11 models served from config.yaml (DB ModelTable empty; STORE_MODEL_IN_DB note in start.sh is vestigial). Deployed config matches repo config.
  • New migrations audit (all 63 read): every one is additive (ADD COLUMN IF NOT EXISTS / CREATE TABLE IF NOT EXISTS / CREATE INDEX IF NOT EXISTS). No DROP, no row-rewriting DML. Upstream CI banned row-rewriting DML from migrations as of v1.100.0 (PR 37899).

Why last time took hours (all fixed)

  1. Floating tag main-v1.74.0-stable drifted → now exact tags only.
  2. 256 MB cgroup → Prisma OOM (exit 137) mid-migration → 2 GB limit.
  3. Unwritable migration dir → P3005 crash-loop → LITELLM_MIGRATION_DIR on persistent volume.

Plan: decouple migration from the version swap

Since all 63 migrations are additive and backward-compatible, the old v1.84.0 keeps serving while the new schema lands:

  1. Pre-migrate while live (the downtime win): run the v1.103.2 image's prisma migrate deploy against the live DB from a throwaway container while v1.84.0 continues serving. Old app ignores the new columns/tables.
  2. Swap: bump Dockerfile pin to v1.103.2, upstreamVersion to 1.103.2, cloudron update. Boot finds the ledger already advanced → no migration on boot → binds port 4000 in ~1–2 min (image pull + start only).
  3. Verify with a member-scope key (master-key requests bypass member auth and can pass on a half-broken schema):
    • /v1/models lists all 11
    • one completion on each provider (tinfoil / greenpt / publicai)
    • portal /broker/usage, dashboard, chat (chat.inference.coop) all render spend/balance
  4. Rollback: additive-only means revert = re-pin Dockerfile to v1.84.0 + redeploy; new columns simply go unused. No destructive rollback.

Watch item

v1.99–v1.101 churned the migration resolver (v1→v2, revert, deadlock fixes in 1.101.3/1.102). By 1.103.2 it's stable — this is why we target 1.103.2 (patch release) and not an older midpoint.

Open decisions

  1. Target v1.103.2 (recommended) vs a conservative midpoint (~1.95.0)?
  2. Sign-off on pre-migrating the live DB while v1.84.0 serves (the core downtime win; verified safe for these specific migrations).

Staged commands (for review — DO NOT run until decisions confirmed)

# 0. Preflight (read-only)
#    - app healthy, ledger at 20260501195714, config.yaml matches repo, image v1.84.0

# 1. Pre-migrate live DB with the new image's Prisma (v1.84.0 keeps serving)
docker run --rm --network cloudron \
  -e DATABASE_URL=<from gateway app env> \
  -v /app/data/migrations:/migrations \
  ghcr.io/berriai/litellm:v1.103.2 prisma migrate deploy
# NOTE: exact invocation (prisma CLI path inside that image, migration dir mount)
# to be confirmed against the v1.103.2 image layout before running.

# 2. Verify ledger advanced (read-only psql)
#    SELECT max(migration_name) FROM _prisma_migrations;  -- expect 20260918000000_add_managed_file_content_table

# 3. Swap the pin
#    Dockerfile:  FROM ghcr.io/berriai/litellm:v1.103.2
#    CloudronManifest.json: "upstreamVersion": "1.103.2"
#    commit + push + cloudron update --app adababd4-2427-4bdd-9bf7-6bdad0302610

# 4. Verify (member-scope key)
#    curl /v1/models; one completion per provider; portal /broker/usage;
#    dashboard + chat render balance/spend

# 5. On any anomaly: re-pin v1.84.0, redeploy, re-verify (additive-only rollback)

Post-upgrade follow-ups

  • Re-sync repo README "Current base image" line + this issue closes.
  • Remove the cosmetic duplicate entries for the two audio models in deployed config.yaml (voxtral-tts, whisper listed twice) — harmless, cleanup-only.
  • Re-check budget_reset_at direct-Postgres write path still honored (portal pins reset dates on OC payments via direct DB write).
# LiteLLM upgrade plan: v1.84.0 → v1.103.2 **Goal:** upgrade the gateway with ~1–2 minutes of downtime instead of the multi-hour outage we got on 2026-10-02. **Status: plan only — nothing executed yet.** Two decisions needed from the operator before execution (see "Open decisions"). ## Why upgrade now - We sit 50 stable releases / 63 schema migrations behind (v1.84.0 → v1.103.2, latest stable, 2026-10-01). - Not security-driven: v1.84.0 is already the fix line for CVE-2026-35029, CVE-2026-59822 (KEV), CVE-2026-49468, and CVE-2026-42208 (9.3 pre-auth SQLi). This is currency, not urgency. ## Verified current state - Deployed: v1.84.0 exact-tag pinned, 2 GB memory limit, `LITELLM_MIGRATION_DIR=/app/data/migrations` (all three 2026-10-02 outage fixes baked into packaging, commit e8c7a55). - Prisma ledger complete through `20260501195714_managed_resource_team_owner` = exactly v1.84.0's last migration. No drift. - DB sizes (why this is easy for us): SpendLogs 5,291 rows · DailyTeamSpend 232 · TeamTable 37 · VerificationToken 57. Migration runtime will be seconds-to-a-minute, not hours. - Model catalog: 11 models served from config.yaml (DB ModelTable empty; `STORE_MODEL_IN_DB` note in start.sh is vestigial). Deployed config matches repo config. - New migrations audit (all 63 read): every one is additive (`ADD COLUMN IF NOT EXISTS` / `CREATE TABLE IF NOT EXISTS` / `CREATE INDEX IF NOT EXISTS`). No DROP, no row-rewriting DML. Upstream CI banned row-rewriting DML from migrations as of v1.100.0 (PR 37899). ## Why last time took hours (all fixed) 1. Floating tag `main-v1.74.0-stable` drifted → now exact tags only. 2. 256 MB cgroup → Prisma OOM (exit 137) mid-migration → 2 GB limit. 3. Unwritable migration dir → P3005 crash-loop → `LITELLM_MIGRATION_DIR` on persistent volume. ## Plan: decouple migration from the version swap Since all 63 migrations are additive and backward-compatible, the old v1.84.0 keeps serving while the new schema lands: 1. **Pre-migrate while live** (the downtime win): run the v1.103.2 image's `prisma migrate deploy` against the live DB from a throwaway container while v1.84.0 continues serving. Old app ignores the new columns/tables. 2. **Swap**: bump Dockerfile pin to `v1.103.2`, `upstreamVersion` to `1.103.2`, `cloudron update`. Boot finds the ledger already advanced → no migration on boot → binds port 4000 in ~1–2 min (image pull + start only). 3. **Verify with a member-scope key** (master-key requests bypass member auth and can pass on a half-broken schema): - `/v1/models` lists all 11 - one completion on each provider (tinfoil / greenpt / publicai) - portal `/broker/usage`, dashboard, chat (chat.inference.coop) all render spend/balance 4. **Rollback**: additive-only means revert = re-pin Dockerfile to `v1.84.0` + redeploy; new columns simply go unused. No destructive rollback. ### Watch item v1.99–v1.101 churned the migration resolver (v1→v2, revert, deadlock fixes in 1.101.3/1.102). By 1.103.2 it's stable — this is why we target 1.103.2 (patch release) and not an older midpoint. ## Open decisions 1. Target v1.103.2 (recommended) vs a conservative midpoint (~1.95.0)? 2. Sign-off on pre-migrating the live DB while v1.84.0 serves (the core downtime win; verified safe for these specific migrations). ## Staged commands (for review — DO NOT run until decisions confirmed) ```bash # 0. Preflight (read-only) # - app healthy, ledger at 20260501195714, config.yaml matches repo, image v1.84.0 # 1. Pre-migrate live DB with the new image's Prisma (v1.84.0 keeps serving) docker run --rm --network cloudron \ -e DATABASE_URL=<from gateway app env> \ -v /app/data/migrations:/migrations \ ghcr.io/berriai/litellm:v1.103.2 prisma migrate deploy # NOTE: exact invocation (prisma CLI path inside that image, migration dir mount) # to be confirmed against the v1.103.2 image layout before running. # 2. Verify ledger advanced (read-only psql) # SELECT max(migration_name) FROM _prisma_migrations; -- expect 20260918000000_add_managed_file_content_table # 3. Swap the pin # Dockerfile: FROM ghcr.io/berriai/litellm:v1.103.2 # CloudronManifest.json: "upstreamVersion": "1.103.2" # commit + push + cloudron update --app adababd4-2427-4bdd-9bf7-6bdad0302610 # 4. Verify (member-scope key) # curl /v1/models; one completion per provider; portal /broker/usage; # dashboard + chat render balance/spend # 5. On any anomaly: re-pin v1.84.0, redeploy, re-verify (additive-only rollback) ``` ## Post-upgrade follow-ups - Re-sync repo README "Current base image" line + this issue closes. - Remove the cosmetic duplicate entries for the two audio models in deployed config.yaml (voxtral-tts, whisper listed twice) — harmless, cleanup-only. - Re-check `budget_reset_at` direct-Postgres write path still honored (portal pins reset dates on OC payments via direct DB write).
Author
Owner

Follow-up from member report (2026-10-02, prompt-cache reporting): streaming responses through the gateway drop usage.prompt_tokens_details.cached_tokens from the final usage chunk (non-streaming passes it through; GreenPT direct also passes it through). Verified empirically:

  • Billing is correct: streaming calls ARE billed at the cache-read rate internally (test: 16k-token cached prompt billed $0.000357 vs $0.001764 full price). SpendLogs reflect discounts — over the last 7 days, 1729/1786 glm-5.3-flash calls >15k input tokens got cache discounts ($34.65 saved vs full price).
  • Client-visible usage is wrong on streaming — known upstream bug: BerriAI/litellm#39088 (open, affects the streaming translation layer; provider sends the field).

Action for this upgrade: after moving to v1.103.2, re-test streaming usage passthrough:

# warm cache, then stream with stream_options.include_usage and check the final chunk for prompt_tokens_details.cached_tokens

If still dropped in 1.103.2, consider a fix upstream or an interim note in member docs (costs ARE discounted; only the response object under-reports savings).

**Follow-up from member report (2026-10-02, prompt-cache reporting):** streaming responses through the gateway drop `usage.prompt_tokens_details.cached_tokens` from the final usage chunk (non-streaming passes it through; GreenPT direct also passes it through). Verified empirically: - **Billing is correct**: streaming calls ARE billed at the cache-read rate internally (test: 16k-token cached prompt billed $0.000357 vs $0.001764 full price). SpendLogs reflect discounts — over the last 7 days, 1729/1786 glm-5.3-flash calls >15k input tokens got cache discounts ($34.65 saved vs full price). - **Client-visible usage is wrong on streaming** — known upstream bug: BerriAI/litellm#39088 (open, affects the streaming translation layer; provider sends the field). **Action for this upgrade**: after moving to v1.103.2, re-test streaming usage passthrough: ```bash # warm cache, then stream with stream_options.include_usage and check the final chunk for prompt_tokens_details.cached_tokens ``` If still dropped in 1.103.2, consider a fix upstream or an interim note in member docs (costs ARE discounted; only the response object under-reports savings).
Author
Owner

Executed 2026-10-02 ~21:25 UTC — v1.103.2 live. Downtime: under 2 minutes.

What actually happened:

  1. Pre-migration (live, zero downtime): since no docker access from the agent host, the new image's Prisma couldn't be run directly — instead the 63 v1.103.2 migration.sql files were fetched from the upstream tag, safety-audited (all additive; only guarded DROP CONSTRAINT/INDEX replacements + 2 UPDATEs on a table created in the same set), and applied via psycopg2 from the portal container in exact prisma-migrate-deploy semantics: per-migration transaction, sha256 checksums, ledger INSERTs. One special case: 20260831120001 (CREATE INDEX CONCURRENTLY) applied in autocommit outside a txn block, as Prisma itself does. v1.84.0 served traffic throughout.
  2. Backup: full app backup taken first (snapshot/app_adababd4...).
  3. Swap: Dockerfile/manifest pinned v1.103.2 (commit a2864ce), cloudron update. Boot found the ledger already at 20260918000000 → no boot migration → healthcheck green in <2 min.
  4. Verification (all pass): litellm dist 1.103.2 · /v1/models 11 · member-scope key: /v1/models + completion ('ok') · one completion per provider (tinfoil ok / greenpt ok / publicai ok) · spend logs writing with cache discounts · chat + dashboard healthy · streaming usage now reports prompt_tokens_details.cached_tokens: 16000 (the #39088-style reporting bug is fixed in 1.103.2 — the new cost field also appears).

Rollback available if anything degrades: app backup from step 2 + re-pin v1.84.0 (additive-only schema, old image ignores new columns).

Closing. Post-upgrade follow-ups from the issue remain: remove duplicated audio model entries in deployed config.yaml (cosmetic), watch GreenPT tail latency.

**Executed 2026-10-02 ~21:25 UTC — v1.103.2 live. Downtime: under 2 minutes.** What actually happened: 1. **Pre-migration (live, zero downtime):** since no docker access from the agent host, the new image's Prisma couldn't be run directly — instead the 63 v1.103.2 migration.sql files were fetched from the upstream tag, safety-audited (all additive; only guarded DROP CONSTRAINT/INDEX replacements + 2 UPDATEs on a table created in the same set), and applied via psycopg2 from the portal container in exact prisma-migrate-deploy semantics: per-migration transaction, sha256 checksums, ledger INSERTs. One special case: 20260831120001 (CREATE INDEX CONCURRENTLY) applied in autocommit outside a txn block, as Prisma itself does. v1.84.0 served traffic throughout. 2. **Backup:** full app backup taken first (snapshot/app_adababd4...). 3. **Swap:** Dockerfile/manifest pinned v1.103.2 (commit a2864ce), cloudron update. Boot found the ledger already at 20260918000000 → **no boot migration** → healthcheck green in <2 min. 4. **Verification (all pass):** litellm dist 1.103.2 · /v1/models 11 · member-scope key: /v1/models + completion ('ok') · one completion per provider (tinfoil ok / greenpt ok / publicai ok) · spend logs writing with cache discounts · chat + dashboard healthy · **streaming usage now reports prompt_tokens_details.cached_tokens: 16000** (the #39088-style reporting bug is fixed in 1.103.2 — the new cost field also appears). Rollback available if anything degrades: app backup from step 2 + re-pin v1.84.0 (additive-only schema, old image ignores new columns). Closing. Post-upgrade follow-ups from the issue remain: remove duplicated audio model entries in deployed config.yaml (cosmetic), watch GreenPT tail latency.
Sign in to join this conversation.
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: code/litellm#2