LiteLLM Gateway — Cloudron App

A Cloudron app package for LiteLLM, the open-source AI gateway, configured for the Inference Cooperative.

LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.

Current base image: LiteLLM v1.84.0, pinned to the exact tag — the fix line for CVE-2026-35029 and CVE-2026-59822 (CISA KEV-listed). See Updating LiteLLM.

What this provides

  • LiteLLM proxy with an OpenAI-compatible API on port 4000
  • Per-member virtual API keys with budget caps and rate limits
  • Usage / spend tracking per key, per model
  • Model routing to multiple providers, chosen by member values (private / green / public)
  • Cloudron PostgreSQL for keys, teams, and spend logs
  • Cloudron Redis for rate limiting and caching
  • Automatic SSL, backups, and sandboxing (Cloudron-managed)
  • Master key + salt key auto-generated on first start, persisted in /app/data
  • Hardened for the public internet: admin UI and API docs disabled at the proxy; CORS locked to the chat origin

Architecture

Members → Open WebUI (chat) / own tools (API) → member portal (key injection)
        → LiteLLM gateway (this app)
            ├→ Tinfoil (TEE enclaves — architectural privacy)
            ├→ GreenPT  (100% renewable energy)
            └→ PublicAI (publicly developed, sovereign models)

LiteLLM is the control plane (keys, metering, routing). It does not run models itself — it forwards to providers, each chosen by members for a different value. A tinfoil-proxy sidecar provides EHBP end-to-end encryption to Tinfoil's enclaves.

Files

  • CloudronManifest.json — Cloudron app manifest (addons, ports, memory limit, metadata)
  • Dockerfile — wraps LiteLLM's official image for Cloudron's environment (pinned to an exact tag)
  • start.sh — startup script: wires Cloudron's Postgres/Redis, downloads the tinfoil-proxy sidecar, exports LITELLM_MIGRATION_DIR
  • config.yaml — LiteLLM config: model catalog, provider routing, retry policy, custom per-token pricing
  • logo.png — app icon

Models

The catalog is defined in config.yaml and documented for members in co-op/docs models — that page is the single source of truth; this repo's config.yaml mirrors it. Models follow the provider/model-name convention so members can choose not just which model but which provider values it embodies (privacy, renewable energy, public development).

The catalog changes by member decision — expect it to evolve.


Building and installing

This package uses Cloudron's on-server build (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server.

Prerequisites

  • The Cloudron CLI installed on a machine with access to the Cloudron server
  • A Cloudron instance with the inference.coop domain

Install

# 1. Install the Cloudron CLI (once)
npm install -g cloudron-cli

# 2. Log in to the Cloudron
cloudron login my.inference.coop

# 3. From this directory, install (builds on the server)
cloudron install --location gateway

--location gateway installs the app at gateway.inference.coop. Omit --location to be prompted.

Configure

After install, set the provider API keys in Cloudron → app → Settings → Environment Variables:

The following are auto-generated and should not be set manually:

  • LITELLM_MASTER_KEY — generated on first start, stored in /app/data/.master_key
  • LITELLM_SALT_KEY — generated on first start, stored in /app/data/.salt_key
  • DATABASE_URL — from Cloudron's PostgreSQL addon
  • REDIS_HOST, REDIS_PORT, REDIS_PASSWORD — from Cloudron's Redis addon

Admin access

The LiteLLM admin UI and API docs are disabled on the public gateway (hardened by default; see start.sh). Administrative operations run through the member portal, which holds the master key server-side. If direct LiteLLM admin access is ever needed, re-enable the UI in start.sh, redeploy, and use the master key from /app/data/.master_key — then disable it again.


Updating LiteLLM

This is the routine maintenance path. LiteLLM releases frequently. Upgrading is a code change (pin a new version) plus a DB migration that runs on first boot — expect a 15–20 min maintenance window and never panic-restart mid-migration (each restart resets the clock).

Rules (learned from the 2026-10-02 outage)

  • Pin the exact version tag, never a floating tag. Registries silently move floating tags (main-v1.74.0-stable, latest) to newer digests, so a routine rebuild changes the deployed version with no pin change and triggers a full schema migration against an unprepared DB. Always FROM ghcr.io/berriai/litellm:vX.Y.Z.
  • Give the app a real memory limit. Cloudron's default when memoryLimit: 0 is a 256 MB cgroup cap; the proxy fits in 256 MB but Prisma's schema engine does NOT, and gets SIGKILL'd (OOM, exit 137) mid-migration. This package sets memoryLimit: 2147483648 (2 GB) in CloudronManifest.json — don't lower it.
  • LITELLM_MIGRATION_DIR=/app/data/migrations is exported in start.sh. The database was created by prisma db push, so it has no _prisma_migrations ledger; without a writable migration dir the boot migration hits P3005 ("schema is not empty") and crash-loops. Keep it pointed at the persistent volume.
  • Verify with a member-scope key, not only the master key. Master-key requests bypass the member auth path, so they can pass even on a half-migrated schema while every member key 401s. See the checklist below.

Steps

  1. Pick a target version from the LiteLLM changelog. Stay at or above v1.84.0 (the fix line for CVE-2026-35029 and CVE-2026-59822, the latter on CISA's KEV list).

  2. Bump the pin in Dockerfile:

    FROM ghcr.io/berriai/litellm:v1.84.0
    #                              ^^^^^^ exact tag, never main-*-stable
    
  3. Bump the manifest in CloudronManifest.json:

    "upstreamVersion": "1.84.0"
    
  4. Redeploy:

    cloudron update --app <app-id>
    

    First boot after the version change runs prisma migrate deploy (or baselines + db push if the ledger is missing) before binding port 4000. Expect 15–20 min of ECONNREFUSED healthcheck errors — this is normal.

Verify after upgrading

  1. GET https://gateway.inference.coop/health/liveliness → 200
  2. Generate a temp key with the master key → if /key/generate 500s, the schema is incomplete; fix the migration before anything else
  3. Use the temp key (not master) for GET /v1/models → full catalog
  4. Use the temp key for a small POST /v1/chat/completions → content + usage
  5. OPTIONS preflight from a foreign origin → 400; from chat.inference.coop → allowed
  6. Cloudron app health = healthy, then delete the temp key

Before updating

  • Check the LiteLLM changelog for breaking changes
  • The start.sh and config.yaml may need adjustment if LiteLLM changes its env-var or config schema
  • The CORS env var is LITELLM_CORS_ORIGINS (upstream-native since 1.84.0); the old LITELLM_CORS_ALLOWED_ORIGINS + sed patch are dead — don't re-add them
  • Test on a staging instance if the version jump is large

Environment variables

Variable Required Purpose
TINFOIL_API_KEY Yes Tinfoil provider (TEE inference)
GREENPT_API_KEY Yes GreenPT provider (renewable energy)
PUBLICAI_API_KEY Yes PublicAI provider (public models)
LITELLM_CORS_ORIGINS Set by start.sh CORS allowlist (chat origin only)
LITELLM_MIGRATION_DIR Set by start.sh Persistent Prisma migrations ledger

Connecting chat clients

Open WebUI does not talk to this gateway directly — it routes through the member portal, which authenticates the member and injects their personal LiteLLM key per request. Members building their own tools create API keys on the member dashboard and use:

  • API Base URL: https://gateway.inference.coop/v1 (or the portal's proxy endpoint for chat-integrated tools)
  • API Key: their own virtual key from the dashboard
  • code/member-portal — membership middleware: provisioning, billing webhooks, key injection
  • code/member-dashboard — member-facing usage and API-key management
  • code/admin-panel — admin overview
  • co-op/website — the public landing page
  • co-op/docs — documentation, including the member-facing model list
  • co-op/design-assets — logos, fonts, brand materials
S
Description
LiteLLM AI gateway packaged as a Cloudron app for the Inference Cooperative
Readme
119 KiB
0 Stars 2 Watchers 0 Forks
Languages
Shell 82.4%
Dockerfile 17.6%