inference-bot fd34e33508 Sync packaging to live v1.84.0 deployment (CVE-2026-35029/59822 fix line)
- Dockerfile: pin ghcr.io/berriai/litellm:v1.84.0 (was floating
  main-v1.74.0-stable, which silently moved to a newer digest).
  v1.84.0 is the fix line for CVE-2026-35029 (auth bypass on
  /config/update, fixed 1.83.0) and CVE-2026-59822 (MCP session
  auth bypass, fixed 1.84.0, CISA KEV).
- Drop the v1.74.0-era sed CORS patch: upstream v1.84.0 moved the
  file and now reads LITELLM_CORS_ORIGINS natively.
- start.sh: export LITELLM_CORS_ORIGINS (the native v1.84.0 var)
  instead of LITELLM_CORS_ALLOWED_ORIGINS (the old sed-patch var
  that v1.84.0 ignores) — keeps CORS locked to chat.inference.coop.
- CloudronManifest: upstreamVersion 1.74.0 -> 1.84.0.

Live gateway already runs v1.84.0 (image digest
sha256:dc532d896ba8..., built 2026-10-02 04:50 UTC); this commit
makes the repo mirror the deployed packaging per the no-drift rule.
LITELLM_CORS_ORIGINS also set as a Cloudron env var on the app so
the running container honours it without a rebuild.
2026-10-01 23:32:43 -06:00

LiteLLM Gateway — Cloudron App

A Cloudron app package for LiteLLM, the open-source AI gateway, configured for the Inference Cooperative.

LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.

What this provides

  • LiteLLM proxy with an OpenAI-compatible API on port 4000
  • Per-member virtual API keys with budget caps and rate limits
  • Usage / spend tracking per key, per model
  • Model routing to Tinfoil (TEE-protected inference)
  • Cloudron PostgreSQL for keys, teams, and spend logs
  • Cloudron Redis for rate limiting and caching
  • Automatic SSL, backups, and sandboxing (Cloudron-managed)
  • Master key + salt key auto-generated on first start, persisted in /app/data

Architecture

Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE)

LiteLLM is the control plane (keys, metering, routing). It does not run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs).

Files

  • CloudronManifest.json — Cloudron app manifest (addons, ports, metadata)
  • Dockerfile — wraps LiteLLM's official image for Cloudron's environment
  • start.sh — startup script that wires up Cloudron's Postgres/Redis and writes the default config
  • config.yaml — default LiteLLM config (Tinfoil-only, two models)
  • logo.png — app icon

Models (MVP)

Two models, both TEE-protected via Tinfoil:

Model Role Price (in/out per M tokens)
deepseek-v4-flash Default — best agent performance, 1M context $0.30 / $0.70
gpt-oss-120b Fallback — cheapest, built for agentic workflows $0.15 / $0.60

If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS.


Building and installing

This package uses Cloudron's on-server build (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server.

Prerequisites

  • The Cloudron CLI installed on a machine with access to the Cloudron server
  • A Cloudron instance with the inference.coop domain

Install

# 1. Install the Cloudron CLI (once)
npm install -g cloudron-cli

# 2. Log in to the Cloudron
cloudron login my.inference.coop

# 3. From this directory, install (builds on the server)
cloudron install --location gateway

--location gateway installs the app at gateway.inference.coop. Omit --location to be prompted.

Configure

After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables:

The following are auto-generated and should not be set manually:

  • LITELLM_MASTER_KEY — generated on first start, stored in /app/data/.master_key
  • LITELLM_SALT_KEY — generated on first start, stored in /app/data/.salt_key
  • DATABASE_URL — from Cloudron's PostgreSQL addon
  • REDIS_HOST, REDIS_PORT, REDIS_PASSWORD — from Cloudron's Redis addon

Access the admin UI

  1. Open https://gateway.inference.coop/ui
  2. Log in with username admin and the master key as password
  3. Find the master key: cloudron exec --app <app-id> cat /app/data/.master_key

Updating LiteLLM

This is the routine maintenance path. LiteLLM releases frequently; updating is a two-step change.

1. Bump the base image tag

In Dockerfile, change the pinned version:

FROM ghcr.io/berriai/litellm:main-v1.74.0-stable
#                              ^^^^^^ bump this

2. Bump the manifest version

In CloudronManifest.json, update:

"version": "1.0.0",
"upstreamVersion": "1.74.0"

3. Redeploy

cloudron update --app <app-id>

Cloudron rebuilds from source and applies the update. Data (keys, spend logs) persists in PostgreSQL.

Before updating

  • Check the LiteLLM changelog for breaking changes
  • The start.sh and config.yaml may need adjustment if LiteLLM changes its env-var or config schema
  • Test on a staging instance if the version jump is large

Environment variables

Variable Required Purpose
TINFOIL_API_KEY Yes Tinfoil API key for TEE inference

Connecting Open WebUI

In Open WebUI settings:

  • API Base URL: https://gateway.inference.coop/v1
  • API Key: a virtual key created in LiteLLM's Admin UI

OIDC / SSO (future)

The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the "oidc": {} addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso.

  • co-op/org-dev — the full inference.coop design doc and deploy scripts
  • co-op/website — the public landing page
  • co-op/design-assets — logos, fonts, brand materials
S
Description
LiteLLM AI gateway packaged as a Cloudron app for the Inference Cooperative
Readme
119 KiB
0 Stars 2 Watchers 0 Forks
Languages
Shell 82.4%
Dockerfile 17.6%