Goal: upgrade the gateway with ~1–2 minutes of downtime instead of the multi-hour outage we got on 2026-10-02. Status: plan only — nothing executed yet. Two decisions needed from the operator before execution (see "Open decisions").
Why upgrade now
We sit 50 stable releases / 63 schema migrations behind (v1.84.0 → v1.103.2, latest stable, 2026-10-01).
Not security-driven: v1.84.0 is already the fix line for CVE-2026-35029, CVE-2026-59822 (KEV), CVE-2026-49468, and CVE-2026-42208 (9.3 pre-auth SQLi). This is currency, not urgency.
Verified current state
Deployed: v1.84.0 exact-tag pinned, 2 GB memory limit, LITELLM_MIGRATION_DIR=/app/data/migrations (all three 2026-10-02 outage fixes baked into packaging, commit e8c7a55).
Prisma ledger complete through 20260501195714_managed_resource_team_owner = exactly v1.84.0's last migration. No drift.
DB sizes (why this is easy for us): SpendLogs 5,291 rows · DailyTeamSpend 232 · TeamTable 37 · VerificationToken 57. Migration runtime will be seconds-to-a-minute, not hours.
Model catalog: 11 models served from config.yaml (DB ModelTable empty; STORE_MODEL_IN_DB note in start.sh is vestigial). Deployed config matches repo config.
New migrations audit (all 63 read): every one is additive (ADD COLUMN IF NOT EXISTS / CREATE TABLE IF NOT EXISTS / CREATE INDEX IF NOT EXISTS). No DROP, no row-rewriting DML. Upstream CI banned row-rewriting DML from migrations as of v1.100.0 (PR 37899).
Why last time took hours (all fixed)
Floating tag main-v1.74.0-stable drifted → now exact tags only.
Unwritable migration dir → P3005 crash-loop → LITELLM_MIGRATION_DIR on persistent volume.
Plan: decouple migration from the version swap
Since all 63 migrations are additive and backward-compatible, the old v1.84.0 keeps serving while the new schema lands:
Pre-migrate while live (the downtime win): run the v1.103.2 image's prisma migrate deploy against the live DB from a throwaway container while v1.84.0 continues serving. Old app ignores the new columns/tables.
Swap: bump Dockerfile pin to v1.103.2, upstreamVersion to 1.103.2, cloudron update. Boot finds the ledger already advanced → no migration on boot → binds port 4000 in ~1–2 min (image pull + start only).
Verify with a member-scope key (master-key requests bypass member auth and can pass on a half-broken schema):
/v1/models lists all 11
one completion on each provider (tinfoil / greenpt / publicai)
portal /broker/usage, dashboard, chat (chat.inference.coop) all render spend/balance
Rollback: additive-only means revert = re-pin Dockerfile to v1.84.0 + redeploy; new columns simply go unused. No destructive rollback.
Watch item
v1.99–v1.101 churned the migration resolver (v1→v2, revert, deadlock fixes in 1.101.3/1.102). By 1.103.2 it's stable — this is why we target 1.103.2 (patch release) and not an older midpoint.
Open decisions
Target v1.103.2 (recommended) vs a conservative midpoint (~1.95.0)?
Sign-off on pre-migrating the live DB while v1.84.0 serves (the core downtime win; verified safe for these specific migrations).
Staged commands (for review — DO NOT run until decisions confirmed)
# 0. Preflight (read-only)# - app healthy, ledger at 20260501195714, config.yaml matches repo, image v1.84.0# 1. Pre-migrate live DB with the new image's Prisma (v1.84.0 keeps serving)
docker run --rm --network cloudron \
-e DATABASE_URL=<from gateway app env> \
-v /app/data/migrations:/migrations \
ghcr.io/berriai/litellm:v1.103.2 prisma migrate deploy
# NOTE: exact invocation (prisma CLI path inside that image, migration dir mount)# to be confirmed against the v1.103.2 image layout before running.# 2. Verify ledger advanced (read-only psql)# SELECT max(migration_name) FROM _prisma_migrations; -- expect 20260918000000_add_managed_file_content_table# 3. Swap the pin# Dockerfile: FROM ghcr.io/berriai/litellm:v1.103.2# CloudronManifest.json: "upstreamVersion": "1.103.2"# commit + push + cloudron update --app adababd4-2427-4bdd-9bf7-6bdad0302610# 4. Verify (member-scope key)# curl /v1/models; one completion per provider; portal /broker/usage;# dashboard + chat render balance/spend# 5. On any anomaly: re-pin v1.84.0, redeploy, re-verify (additive-only rollback)
Post-upgrade follow-ups
Re-sync repo README "Current base image" line + this issue closes.
Remove the cosmetic duplicate entries for the two audio models in deployed config.yaml (voxtral-tts, whisper listed twice) — harmless, cleanup-only.
Re-check budget_reset_at direct-Postgres write path still honored (portal pins reset dates on OC payments via direct DB write).
# LiteLLM upgrade plan: v1.84.0 → v1.103.2
**Goal:** upgrade the gateway with ~1–2 minutes of downtime instead of the multi-hour outage we got on 2026-10-02. **Status: plan only — nothing executed yet.** Two decisions needed from the operator before execution (see "Open decisions").
## Why upgrade now
- We sit 50 stable releases / 63 schema migrations behind (v1.84.0 → v1.103.2, latest stable, 2026-10-01).
- Not security-driven: v1.84.0 is already the fix line for CVE-2026-35029, CVE-2026-59822 (KEV), CVE-2026-49468, and CVE-2026-42208 (9.3 pre-auth SQLi). This is currency, not urgency.
## Verified current state
- Deployed: v1.84.0 exact-tag pinned, 2 GB memory limit, `LITELLM_MIGRATION_DIR=/app/data/migrations` (all three 2026-10-02 outage fixes baked into packaging, commit e8c7a55).
- Prisma ledger complete through `20260501195714_managed_resource_team_owner` = exactly v1.84.0's last migration. No drift.
- DB sizes (why this is easy for us): SpendLogs 5,291 rows · DailyTeamSpend 232 · TeamTable 37 · VerificationToken 57. Migration runtime will be seconds-to-a-minute, not hours.
- Model catalog: 11 models served from config.yaml (DB ModelTable empty; `STORE_MODEL_IN_DB` note in start.sh is vestigial). Deployed config matches repo config.
- New migrations audit (all 63 read): every one is additive (`ADD COLUMN IF NOT EXISTS` / `CREATE TABLE IF NOT EXISTS` / `CREATE INDEX IF NOT EXISTS`). No DROP, no row-rewriting DML. Upstream CI banned row-rewriting DML from migrations as of v1.100.0 (PR 37899).
## Why last time took hours (all fixed)
1. Floating tag `main-v1.74.0-stable` drifted → now exact tags only.
2. 256 MB cgroup → Prisma OOM (exit 137) mid-migration → 2 GB limit.
3. Unwritable migration dir → P3005 crash-loop → `LITELLM_MIGRATION_DIR` on persistent volume.
## Plan: decouple migration from the version swap
Since all 63 migrations are additive and backward-compatible, the old v1.84.0 keeps serving while the new schema lands:
1. **Pre-migrate while live** (the downtime win): run the v1.103.2 image's `prisma migrate deploy` against the live DB from a throwaway container while v1.84.0 continues serving. Old app ignores the new columns/tables.
2. **Swap**: bump Dockerfile pin to `v1.103.2`, `upstreamVersion` to `1.103.2`, `cloudron update`. Boot finds the ledger already advanced → no migration on boot → binds port 4000 in ~1–2 min (image pull + start only).
3. **Verify with a member-scope key** (master-key requests bypass member auth and can pass on a half-broken schema):
- `/v1/models` lists all 11
- one completion on each provider (tinfoil / greenpt / publicai)
- portal `/broker/usage`, dashboard, chat (chat.inference.coop) all render spend/balance
4. **Rollback**: additive-only means revert = re-pin Dockerfile to `v1.84.0` + redeploy; new columns simply go unused. No destructive rollback.
### Watch item
v1.99–v1.101 churned the migration resolver (v1→v2, revert, deadlock fixes in 1.101.3/1.102). By 1.103.2 it's stable — this is why we target 1.103.2 (patch release) and not an older midpoint.
## Open decisions
1. Target v1.103.2 (recommended) vs a conservative midpoint (~1.95.0)?
2. Sign-off on pre-migrating the live DB while v1.84.0 serves (the core downtime win; verified safe for these specific migrations).
## Staged commands (for review — DO NOT run until decisions confirmed)
```bash
# 0. Preflight (read-only)
# - app healthy, ledger at 20260501195714, config.yaml matches repo, image v1.84.0
# 1. Pre-migrate live DB with the new image's Prisma (v1.84.0 keeps serving)
docker run --rm --network cloudron \
-e DATABASE_URL=<from gateway app env> \
-v /app/data/migrations:/migrations \
ghcr.io/berriai/litellm:v1.103.2 prisma migrate deploy
# NOTE: exact invocation (prisma CLI path inside that image, migration dir mount)
# to be confirmed against the v1.103.2 image layout before running.
# 2. Verify ledger advanced (read-only psql)
# SELECT max(migration_name) FROM _prisma_migrations; -- expect 20260918000000_add_managed_file_content_table
# 3. Swap the pin
# Dockerfile: FROM ghcr.io/berriai/litellm:v1.103.2
# CloudronManifest.json: "upstreamVersion": "1.103.2"
# commit + push + cloudron update --app adababd4-2427-4bdd-9bf7-6bdad0302610
# 4. Verify (member-scope key)
# curl /v1/models; one completion per provider; portal /broker/usage;
# dashboard + chat render balance/spend
# 5. On any anomaly: re-pin v1.84.0, redeploy, re-verify (additive-only rollback)
```
## Post-upgrade follow-ups
- Re-sync repo README "Current base image" line + this issue closes.
- Remove the cosmetic duplicate entries for the two audio models in deployed config.yaml (voxtral-tts, whisper listed twice) — harmless, cleanup-only.
- Re-check `budget_reset_at` direct-Postgres write path still honored (portal pins reset dates on OC payments via direct DB write).
Follow-up from member report (2026-10-02, prompt-cache reporting): streaming responses through the gateway drop usage.prompt_tokens_details.cached_tokens from the final usage chunk (non-streaming passes it through; GreenPT direct also passes it through). Verified empirically:
Billing is correct: streaming calls ARE billed at the cache-read rate internally (test: 16k-token cached prompt billed $0.000357 vs $0.001764 full price). SpendLogs reflect discounts — over the last 7 days, 1729/1786 glm-5.3-flash calls >15k input tokens got cache discounts ($34.65 saved vs full price).
Client-visible usage is wrong on streaming — known upstream bug: BerriAI/litellm#39088 (open, affects the streaming translation layer; provider sends the field).
Action for this upgrade: after moving to v1.103.2, re-test streaming usage passthrough:
# warm cache, then stream with stream_options.include_usage and check the final chunk for prompt_tokens_details.cached_tokens
If still dropped in 1.103.2, consider a fix upstream or an interim note in member docs (costs ARE discounted; only the response object under-reports savings).
**Follow-up from member report (2026-10-02, prompt-cache reporting):** streaming responses through the gateway drop `usage.prompt_tokens_details.cached_tokens` from the final usage chunk (non-streaming passes it through; GreenPT direct also passes it through). Verified empirically:
- **Billing is correct**: streaming calls ARE billed at the cache-read rate internally (test: 16k-token cached prompt billed $0.000357 vs $0.001764 full price). SpendLogs reflect discounts — over the last 7 days, 1729/1786 glm-5.3-flash calls >15k input tokens got cache discounts ($34.65 saved vs full price).
- **Client-visible usage is wrong on streaming** — known upstream bug: BerriAI/litellm#39088 (open, affects the streaming translation layer; provider sends the field).
**Action for this upgrade**: after moving to v1.103.2, re-test streaming usage passthrough:
```bash
# warm cache, then stream with stream_options.include_usage and check the final chunk for prompt_tokens_details.cached_tokens
```
If still dropped in 1.103.2, consider a fix upstream or an interim note in member docs (costs ARE discounted; only the response object under-reports savings).
Executed 2026-10-02 ~21:25 UTC — v1.103.2 live. Downtime: under 2 minutes.
What actually happened:
Pre-migration (live, zero downtime): since no docker access from the agent host, the new image's Prisma couldn't be run directly — instead the 63 v1.103.2 migration.sql files were fetched from the upstream tag, safety-audited (all additive; only guarded DROP CONSTRAINT/INDEX replacements + 2 UPDATEs on a table created in the same set), and applied via psycopg2 from the portal container in exact prisma-migrate-deploy semantics: per-migration transaction, sha256 checksums, ledger INSERTs. One special case: 20260831120001 (CREATE INDEX CONCURRENTLY) applied in autocommit outside a txn block, as Prisma itself does. v1.84.0 served traffic throughout.
Backup: full app backup taken first (snapshot/app_adababd4...).
Swap: Dockerfile/manifest pinned v1.103.2 (commit a2864ce), cloudron update. Boot found the ledger already at 20260918000000 → no boot migration → healthcheck green in <2 min.
Verification (all pass): litellm dist 1.103.2 · /v1/models 11 · member-scope key: /v1/models + completion ('ok') · one completion per provider (tinfoil ok / greenpt ok / publicai ok) · spend logs writing with cache discounts · chat + dashboard healthy · streaming usage now reports prompt_tokens_details.cached_tokens: 16000 (the #39088-style reporting bug is fixed in 1.103.2 — the new cost field also appears).
Rollback available if anything degrades: app backup from step 2 + re-pin v1.84.0 (additive-only schema, old image ignores new columns).
Closing. Post-upgrade follow-ups from the issue remain: remove duplicated audio model entries in deployed config.yaml (cosmetic), watch GreenPT tail latency.
**Executed 2026-10-02 ~21:25 UTC — v1.103.2 live. Downtime: under 2 minutes.**
What actually happened:
1. **Pre-migration (live, zero downtime):** since no docker access from the agent host, the new image's Prisma couldn't be run directly — instead the 63 v1.103.2 migration.sql files were fetched from the upstream tag, safety-audited (all additive; only guarded DROP CONSTRAINT/INDEX replacements + 2 UPDATEs on a table created in the same set), and applied via psycopg2 from the portal container in exact prisma-migrate-deploy semantics: per-migration transaction, sha256 checksums, ledger INSERTs. One special case: 20260831120001 (CREATE INDEX CONCURRENTLY) applied in autocommit outside a txn block, as Prisma itself does. v1.84.0 served traffic throughout.
2. **Backup:** full app backup taken first (snapshot/app_adababd4...).
3. **Swap:** Dockerfile/manifest pinned v1.103.2 (commit a2864ce), cloudron update. Boot found the ledger already at 20260918000000 → **no boot migration** → healthcheck green in <2 min.
4. **Verification (all pass):** litellm dist 1.103.2 · /v1/models 11 · member-scope key: /v1/models + completion ('ok') · one completion per provider (tinfoil ok / greenpt ok / publicai ok) · spend logs writing with cache discounts · chat + dashboard healthy · **streaming usage now reports prompt_tokens_details.cached_tokens: 16000** (the #39088-style reporting bug is fixed in 1.103.2 — the new cost field also appears).
Rollback available if anything degrades: app backup from step 2 + re-pin v1.84.0 (additive-only schema, old image ignores new columns).
Closing. Post-upgrade follow-ups from the issue remain: remove duplicated audio model entries in deployed config.yaml (cosmetic), watch GreenPT tail latency.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
LiteLLM upgrade plan: v1.84.0 → v1.103.2
Goal: upgrade the gateway with ~1–2 minutes of downtime instead of the multi-hour outage we got on 2026-10-02. Status: plan only — nothing executed yet. Two decisions needed from the operator before execution (see "Open decisions").
Why upgrade now
Verified current state
LITELLM_MIGRATION_DIR=/app/data/migrations(all three 2026-10-02 outage fixes baked into packaging, commite8c7a55).20260501195714_managed_resource_team_owner= exactly v1.84.0's last migration. No drift.STORE_MODEL_IN_DBnote in start.sh is vestigial). Deployed config matches repo config.ADD COLUMN IF NOT EXISTS/CREATE TABLE IF NOT EXISTS/CREATE INDEX IF NOT EXISTS). No DROP, no row-rewriting DML. Upstream CI banned row-rewriting DML from migrations as of v1.100.0 (PR 37899).Why last time took hours (all fixed)
main-v1.74.0-stabledrifted → now exact tags only.LITELLM_MIGRATION_DIRon persistent volume.Plan: decouple migration from the version swap
Since all 63 migrations are additive and backward-compatible, the old v1.84.0 keeps serving while the new schema lands:
prisma migrate deployagainst the live DB from a throwaway container while v1.84.0 continues serving. Old app ignores the new columns/tables.v1.103.2,upstreamVersionto1.103.2,cloudron update. Boot finds the ledger already advanced → no migration on boot → binds port 4000 in ~1–2 min (image pull + start only)./v1/modelslists all 11/broker/usage, dashboard, chat (chat.inference.coop) all render spend/balancev1.84.0+ redeploy; new columns simply go unused. No destructive rollback.Watch item
v1.99–v1.101 churned the migration resolver (v1→v2, revert, deadlock fixes in 1.101.3/1.102). By 1.103.2 it's stable — this is why we target 1.103.2 (patch release) and not an older midpoint.
Open decisions
Staged commands (for review — DO NOT run until decisions confirmed)
Post-upgrade follow-ups
budget_reset_atdirect-Postgres write path still honored (portal pins reset dates on OC payments via direct DB write).Follow-up from member report (2026-10-02, prompt-cache reporting): streaming responses through the gateway drop
usage.prompt_tokens_details.cached_tokensfrom the final usage chunk (non-streaming passes it through; GreenPT direct also passes it through). Verified empirically:Action for this upgrade: after moving to v1.103.2, re-test streaming usage passthrough:
If still dropped in 1.103.2, consider a fix upstream or an interim note in member docs (costs ARE discounted; only the response object under-reports savings).
Executed 2026-10-02 ~21:25 UTC — v1.103.2 live. Downtime: under 2 minutes.
What actually happened:
a2864ce), cloudron update. Boot found the ledger already at 20260918000000 → no boot migration → healthcheck green in <2 min.Rollback available if anything degrades: app backup from step 2 + re-pin v1.84.0 (additive-only schema, old image ignores new columns).
Closing. Post-upgrade follow-ups from the issue remain: remove duplicated audio model entries in deployed config.yaml (cosmetic), watch GreenPT tail latency.