On 2026-10-02 the LiteLLM gateway (gateway.inference.coop) went down for members. Chat, the member dashboard, and the admin panel all depend on the gateway, so all three surfaces were unavailable.
Root cause
Two compounding problems:
Floating image tag drift. The Dockerfile pinned ghcr.io/berriai/litellm:main-v1.74.0-stable. Upstream silently re-pointed that floating tag (and v1.74.0-stable) to v1.84.0 bits, so a cloudron update rebuild pulled v1.84.0 code against a database still at the v1.74.0 schema. The running code expected tables (LiteLLM_ProjectTable, LiteLLM_ConfigOverrides, LiteLLM_SSOConfig) that did not exist, so every authenticated API call returned 401 / TableNotFound errors.
256 MB memory cap silently killed the migration engine. The app had memoryLimit: 0 (Cloudron's default = 256 MB). The running proxy fits in 256 MB, but Prisma's schema engine got SIGKILL'd (exit 137 / OOM) every time it tried to run the migration — which is why every boot-time and manual migration attempt failed for hours.
No migration ledger. The database was originally created via prisma db push, so there was no _prisma_migrations history. prisma migrate deploy refused with P3005 ("schema is not empty"), and LiteLLM's boot-time baseline recovery tried to write to the read-only rootfs and crashed, restarting the container in a loop.
Resolution
Raised the app memory limit to 2 GB (configure/memory_limit).
Set LITELLM_MIGRATION_DIR=/app/data/migrations (writable, persistent) so LiteLLM's P3005 baseline recovery had a place to write.
On the next boot, LiteLLM's own resolver: baselined the existing schema, generated the diff to v1.84.0's schema.prisma, and applied it (prisma db execute — "Migration diff applied successfully"), then began recording the 122 historical migrations into the ledger (~10s each).
Follow-ups / hardening
Pin exact image tags, never floating.main-v1.74.0-stable moved out from under us. Pin v1.84.0 (or a digest).
Set an explicit memory limit in the manifest or per-app rather than relying on Cloudron's 256 MB default, so Prisma migrations can run.
Decide the migration strategy going forward (migrate vs db push) and set LITELLM_MIGRATION_DIR once, so future upgrades don't re-hit P3005.
Re-verify CORS after the upgrade (env var renamed LITELLM_CORS_ALLOWED_ORIGINS -> LITELLM_CORS_ORIGINS).
Timeline (UTC)
04:50 - floating-tag rebuild pulled v1.84.0; gateway down
~05:34 - brief self-recovery, then down again (migration never completed)
~06:12-06:34 - manual migration attempts (blocked by 256 MB cap + P3005)
~06:28 - memory raised to 2 GB
~06:34 - LITELLM_MIGRATION_DIR set
~06:36 - schema diff applied; ledger backfill began
~06:5x - ledger backfill completing; gateway binds port 4000
## Summary
On 2026-10-02 the LiteLLM gateway (gateway.inference.coop) went down for members. Chat, the member dashboard, and the admin panel all depend on the gateway, so all three surfaces were unavailable.
## Root cause
Two compounding problems:
1. **Floating image tag drift.** The Dockerfile pinned `ghcr.io/berriai/litellm:main-v1.74.0-stable`. Upstream silently re-pointed that floating tag (and `v1.74.0-stable`) to v1.84.0 bits, so a `cloudron update` rebuild pulled v1.84.0 code against a database still at the v1.74.0 schema. The running code expected tables (`LiteLLM_ProjectTable`, `LiteLLM_ConfigOverrides`, `LiteLLM_SSOConfig`) that did not exist, so every authenticated API call returned 401 / TableNotFound errors.
2. **256 MB memory cap silently killed the migration engine.** The app had `memoryLimit: 0` (Cloudron's default = 256 MB). The running proxy fits in 256 MB, but Prisma's schema engine got SIGKILL'd (exit 137 / OOM) every time it tried to run the migration — which is why every boot-time and manual migration attempt failed for hours.
3. **No migration ledger.** The database was originally created via `prisma db push`, so there was no `_prisma_migrations` history. `prisma migrate deploy` refused with P3005 ("schema is not empty"), and LiteLLM's boot-time baseline recovery tried to write to the read-only rootfs and crashed, restarting the container in a loop.
## Resolution
1. Raised the app memory limit to 2 GB (`configure/memory_limit`).
2. Set `LITELLM_MIGRATION_DIR=/app/data/migrations` (writable, persistent) so LiteLLM's P3005 baseline recovery had a place to write.
3. On the next boot, LiteLLM's own resolver: baselined the existing schema, generated the diff to v1.84.0's schema.prisma, and applied it (`prisma db execute` — "Migration diff applied successfully"), then began recording the 122 historical migrations into the ledger (~10s each).
## Follow-ups / hardening
- **Pin exact image tags, never floating.** `main-v1.74.0-stable` moved out from under us. Pin `v1.84.0` (or a digest).
- **Set an explicit memory limit** in the manifest or per-app rather than relying on Cloudron's 256 MB default, so Prisma migrations can run.
- **Decide the migration strategy going forward** (migrate vs db push) and set `LITELLM_MIGRATION_DIR` once, so future upgrades don't re-hit P3005.
- **Re-verify CORS** after the upgrade (env var renamed `LITELLM_CORS_ALLOWED_ORIGINS` -> `LITELLM_CORS_ORIGINS`).
## Timeline (UTC)
- 04:50 - floating-tag rebuild pulled v1.84.0; gateway down
- ~05:34 - brief self-recovery, then down again (migration never completed)
- ~06:12-06:34 - manual migration attempts (blocked by 256 MB cap + P3005)
- ~06:28 - memory raised to 2 GB
- ~06:34 - `LITELLM_MIGRATION_DIR` set
- ~06:36 - schema diff applied; ledger backfill began
- ~06:5x - ledger backfill completing; gateway binds port 4000
Closing — resolved and verified, with hardening baked in so it can't recur the same way:
Root cause fixed: Dockerfile now pins the exact tag ghcr.io/berriai/litellm:v1.84.0 (commit fd34e33) — no more floating-tag drift
Upgrade path hardened: 2 GB memory limit + LITELLM_MIGRATION_DIR on the persistent volume baked into the packaging (commit e8c7a55), and the README upgrade runbook documents the full procedure including member-key verification
Verified healthy: gateway serving the full catalog, completions working through member-scope keys, all dependent surfaces (chat, dashboard, admin panel) back
Bonus: the version we landed on is the fix line for CVE-2026-35029 and CVE-2026-59822 (CISA KEV), so this incident also closed a security gap.
Closing — resolved and verified, with hardening baked in so it can't recur the same way:
- **Root cause fixed:** Dockerfile now pins the exact tag `ghcr.io/berriai/litellm:v1.84.0` (commit fd34e33) — no more floating-tag drift
- **Upgrade path hardened:** 2 GB memory limit + `LITELLM_MIGRATION_DIR` on the persistent volume baked into the packaging (commit e8c7a55), and the [README upgrade runbook](https://git.inference.coop/code/litellm) documents the full procedure including member-key verification
- **Verified healthy:** gateway serving the full catalog, completions working through member-scope keys, all dependent surfaces (chat, dashboard, admin panel) back
Bonus: the version we landed on is the fix line for CVE-2026-35029 and CVE-2026-59822 (CISA KEV), so this incident also closed a security gap.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Summary
On 2026-10-02 the LiteLLM gateway (gateway.inference.coop) went down for members. Chat, the member dashboard, and the admin panel all depend on the gateway, so all three surfaces were unavailable.
Root cause
Two compounding problems:
Floating image tag drift. The Dockerfile pinned
ghcr.io/berriai/litellm:main-v1.74.0-stable. Upstream silently re-pointed that floating tag (andv1.74.0-stable) to v1.84.0 bits, so acloudron updaterebuild pulled v1.84.0 code against a database still at the v1.74.0 schema. The running code expected tables (LiteLLM_ProjectTable,LiteLLM_ConfigOverrides,LiteLLM_SSOConfig) that did not exist, so every authenticated API call returned 401 / TableNotFound errors.256 MB memory cap silently killed the migration engine. The app had
memoryLimit: 0(Cloudron's default = 256 MB). The running proxy fits in 256 MB, but Prisma's schema engine got SIGKILL'd (exit 137 / OOM) every time it tried to run the migration — which is why every boot-time and manual migration attempt failed for hours.No migration ledger. The database was originally created via
prisma db push, so there was no_prisma_migrationshistory.prisma migrate deployrefused with P3005 ("schema is not empty"), and LiteLLM's boot-time baseline recovery tried to write to the read-only rootfs and crashed, restarting the container in a loop.Resolution
configure/memory_limit).LITELLM_MIGRATION_DIR=/app/data/migrations(writable, persistent) so LiteLLM's P3005 baseline recovery had a place to write.prisma db execute— "Migration diff applied successfully"), then began recording the 122 historical migrations into the ledger (~10s each).Follow-ups / hardening
main-v1.74.0-stablemoved out from under us. Pinv1.84.0(or a digest).LITELLM_MIGRATION_DIRonce, so future upgrades don't re-hit P3005.LITELLM_CORS_ALLOWED_ORIGINS->LITELLM_CORS_ORIGINS).Timeline (UTC)
LITELLM_MIGRATION_DIRsetResolved at last!
Closing — resolved and verified, with hardening baked in so it can't recur the same way:
ghcr.io/berriai/litellm:v1.84.0(commit fd34e33) — no more floating-tag driftLITELLM_MIGRATION_DIRon the persistent volume baked into the packaging (commit e8c7a55), and the README upgrade runbook documents the full procedure including member-key verificationBonus: the version we landed on is the fix line for CVE-2026-35029 and CVE-2026-59822 (CISA KEV), so this incident also closed a security gap.