Files
litellm/README.md
T
inference-co-op-bot e8c7a55933 Add upgrade runbook + bake memory limit and LITELLM_MIGRATION_DIR into packaging
The 2026-10-02 outage had three causes the runbook now prevents:
- floating tag drift (README now mandates exact-tag pinning)
- 256MB default cgroup cap OOM-killing the prisma migration engine
  (manifest now sets memoryLimit: 2147483648)
- no _prisma_migrations ledger -> P3005 crash-loop on boot
  (start.sh now exports LITELLM_MIGRATION_DIR=/app/data/migrations)

README 'Updating LiteLLM' section rewritten from the old (wrong)
floating-tag procedure into a runbook with rules + verify checklist.
2026-10-02 00:52:50 -06:00

7.5 KiB
Raw Blame History

LiteLLM Gateway — Cloudron App

A Cloudron app package for LiteLLM, the open-source AI gateway, configured for the Inference Cooperative.

LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.

What this provides

  • LiteLLM proxy with an OpenAI-compatible API on port 4000
  • Per-member virtual API keys with budget caps and rate limits
  • Usage / spend tracking per key, per model
  • Model routing to Tinfoil (TEE-protected inference)
  • Cloudron PostgreSQL for keys, teams, and spend logs
  • Cloudron Redis for rate limiting and caching
  • Automatic SSL, backups, and sandboxing (Cloudron-managed)
  • Master key + salt key auto-generated on first start, persisted in /app/data

Architecture

Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE)

LiteLLM is the control plane (keys, metering, routing). It does not run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs).

Files

  • CloudronManifest.json — Cloudron app manifest (addons, ports, metadata)
  • Dockerfile — wraps LiteLLM's official image for Cloudron's environment
  • start.sh — startup script that wires up Cloudron's Postgres/Redis and writes the default config
  • config.yaml — default LiteLLM config (Tinfoil-only, two models)
  • logo.png — app icon

Models (MVP)

Two models, both TEE-protected via Tinfoil:

Model Role Price (in/out per M tokens)
deepseek-v4-flash Default — best agent performance, 1M context $0.30 / $0.70
gpt-oss-120b Fallback — cheapest, built for agentic workflows $0.15 / $0.60

If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS.


Building and installing

This package uses Cloudron's on-server build (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server.

Prerequisites

  • The Cloudron CLI installed on a machine with access to the Cloudron server
  • A Cloudron instance with the inference.coop domain

Install

# 1. Install the Cloudron CLI (once)
npm install -g cloudron-cli

# 2. Log in to the Cloudron
cloudron login my.inference.coop

# 3. From this directory, install (builds on the server)
cloudron install --location gateway

--location gateway installs the app at gateway.inference.coop. Omit --location to be prompted.

Configure

After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables:

The following are auto-generated and should not be set manually:

  • LITELLM_MASTER_KEY — generated on first start, stored in /app/data/.master_key
  • LITELLM_SALT_KEY — generated on first start, stored in /app/data/.salt_key
  • DATABASE_URL — from Cloudron's PostgreSQL addon
  • REDIS_HOST, REDIS_PORT, REDIS_PASSWORD — from Cloudron's Redis addon

Access the admin UI

  1. Open https://gateway.inference.coop/ui
  2. Log in with username admin and the master key as password
  3. Find the master key: cloudron exec --app <app-id> cat /app/data/.master_key

Updating LiteLLM

This is the routine maintenance path. LiteLLM releases frequently. Upgrading is a code change (pin a new version) plus a DB migration that runs on first boot — expect a 15–20 min maintenance window and never panic-restart mid-migration (each restart resets the clock).

Rules (learned from the 2026-10-02 outage)

  • Pin the exact version tag, never a floating tag. Registries silently move floating tags (main-v1.74.0-stable, latest) to newer digests, so a routine rebuild changes the deployed version with no pin change and triggers a full schema migration against an unprepared DB. Always FROM ghcr.io/berriai/litellm:vX.Y.Z.
  • Give the app a real memory limit. Cloudron's default when memoryLimit: 0 is a 256 MB cgroup cap; the proxy fits in 256 MB but Prisma's schema engine does NOT, and gets SIGKILL'd (OOM, exit 137) mid-migration. This package sets memoryLimit: 2147483648 (2 GB) in CloudronManifest.json — don't lower it.
  • LITELLM_MIGRATION_DIR=/app/data/migrations is exported in start.sh. The database was created by prisma db push, so it has no _prisma_migrations ledger; without a writable migration dir the boot migration hits P3005 ("schema is not empty") and crash-loops. Keep it pointed at the persistent volume.
  • Verify with a member-scope key, not only the master key. Master-key requests bypass the member auth path, so they can pass even on a half-migrated schema while every member key 401s. See the checklist below.

Steps

  1. Pick a target version from the LiteLLM changelog. Stay at or above v1.84.0 (the fix line for CVE-2026-35029 and CVE-2026-59822, the latter on CISA's KEV list).

  2. Bump the pin in Dockerfile:

    FROM ghcr.io/berriai/litellm:v1.84.0
    #                              ^^^^^^ exact tag, never main-*-stable
    
  3. Bump the manifest in CloudronManifest.json:

    "upstreamVersion": "1.84.0"
    
  4. Redeploy:

    cloudron update --app <app-id>
    

    First boot after the version change runs prisma migrate deploy (or baselines + db push if the ledger is missing) before binding port 4000. Expect 15–20 min of ECONNREFUSED healthcheck errors — this is normal.

Verify after upgrading

  1. GET https://gateway.inference.coop/health/liveliness → 200
  2. Generate a temp key with the master key → if /key/generate 500s, the schema is incomplete; fix the migration before anything else
  3. Use the temp key (not master) for GET /v1/models → full catalog
  4. Use the temp key for a small POST /v1/chat/completions → content + usage
  5. OPTIONS preflight from a foreign origin → 400; from chat.inference.coop → allowed
  6. Cloudron app health = healthy, then delete the temp key

Before updating

  • Check the LiteLLM changelog for breaking changes
  • The start.sh and config.yaml may need adjustment if LiteLLM changes its env-var or config schema
  • The CORS env var is LITELLM_CORS_ORIGINS (upstream-native since 1.84.0); the old LITELLM_CORS_ALLOWED_ORIGINS + sed patch are dead — don't re-add them
  • Test on a staging instance if the version jump is large

Environment variables

Variable Required Purpose
TINFOIL_API_KEY Yes Tinfoil API key for TEE inference

Connecting Open WebUI

In Open WebUI settings:

  • API Base URL: https://gateway.inference.coop/v1
  • API Key: a virtual key created in LiteLLM's Admin UI

OIDC / SSO (future)

The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the "oidc": {} addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso.

  • co-op/org-dev — the full inference.coop design doc and deploy scripts
  • co-op/website — the public landing page
  • co-op/design-assets — logos, fonts, brand materials