LiteLLM Gateway — Cloudron App

A Cloudron app package for LiteLLM, the open-source AI gateway, configured for the Inference Cooperative.

LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.

What this provides

  • LiteLLM proxy with an OpenAI-compatible API on port 4000
  • Per-member virtual API keys with budget caps and rate limits
  • Usage / spend tracking per key, per model
  • Model routing to Tinfoil (TEE-protected inference)
  • Cloudron PostgreSQL for keys, teams, and spend logs
  • Cloudron Redis for rate limiting and caching
  • Automatic SSL, backups, and sandboxing (Cloudron-managed)
  • Master key + salt key auto-generated on first start, persisted in /app/data

Architecture

Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE)

LiteLLM is the control plane (keys, metering, routing). It does not run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs).

Files

  • CloudronManifest.json — Cloudron app manifest (addons, ports, metadata)
  • Dockerfile — wraps LiteLLM's official image for Cloudron's environment
  • start.sh — startup script that wires up Cloudron's Postgres/Redis and writes the default config
  • config.yaml — default LiteLLM config (Tinfoil-only, two models)
  • logo.png — app icon

Models (MVP)

Two models, both TEE-protected via Tinfoil:

Model Role Price (in/out per M tokens)
deepseek-v4-flash Default — best agent performance, 1M context $0.30 / $0.70
gpt-oss-120b Fallback — cheapest, built for agentic workflows $0.15 / $0.60

If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS.


Building and installing

This package uses Cloudron's on-server build (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server.

Prerequisites

  • The Cloudron CLI installed on a machine with access to the Cloudron server
  • A Cloudron instance with the inference.coop domain

Install

# 1. Install the Cloudron CLI (once)
npm install -g cloudron-cli

# 2. Log in to the Cloudron
cloudron login my.inference.coop

# 3. From this directory, install (builds on the server)
cloudron install --location gateway

--location gateway installs the app at gateway.inference.coop. Omit --location to be prompted.

Configure

After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables:

The following are auto-generated and should not be set manually:

  • LITELLM_MASTER_KEY — generated on first start, stored in /app/data/.master_key
  • LITELLM_SALT_KEY — generated on first start, stored in /app/data/.salt_key
  • DATABASE_URL — from Cloudron's PostgreSQL addon
  • REDIS_HOST, REDIS_PORT, REDIS_PASSWORD — from Cloudron's Redis addon

Access the admin UI

  1. Open https://gateway.inference.coop/ui
  2. Log in with username admin and the master key as password
  3. Find the master key: cloudron exec --app <app-id> cat /app/data/.master_key

Updating LiteLLM

This is the routine maintenance path. LiteLLM releases frequently; updating is a two-step change.

1. Bump the base image tag

In Dockerfile, change the pinned version:

FROM ghcr.io/berriai/litellm:main-v1.74.0-stable
#                              ^^^^^^ bump this

2. Bump the manifest version

In CloudronManifest.json, update:

"version": "1.0.0",
"upstreamVersion": "1.74.0"

3. Redeploy

cloudron update --app <app-id>

Cloudron rebuilds from source and applies the update. Data (keys, spend logs) persists in PostgreSQL.

Before updating

  • Check the LiteLLM changelog for breaking changes
  • The start.sh and config.yaml may need adjustment if LiteLLM changes its env-var or config schema
  • Test on a staging instance if the version jump is large

Environment variables

Variable Required Purpose
TINFOIL_API_KEY Yes Tinfoil API key for TEE inference

Connecting Open WebUI

In Open WebUI settings:

  • API Base URL: https://gateway.inference.coop/v1
  • API Key: a virtual key created in LiteLLM's Admin UI

OIDC / SSO (future)

The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the "oidc": {} addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso.

  • co-op/org-dev — the full inference.coop design doc and deploy scripts
  • co-op/website — the public landing page
  • co-op/design-assets — logos, fonts, brand materials
S
Description
LiteLLM AI gateway packaged as a Cloudron app for the Inference Cooperative
Readme
119 KiB
0 Stars 2 Watchers 0 Forks
Languages
Shell 82.4%
Dockerfile 17.6%