Initial commit: LiteLLM Cloudron app package for Inference Cooperative
This commit is contained in:
commit
b9b4620e52
7 files changed
+375
No files matched your search
@@ -0,0 +1,5 @@
|
|||||||
|
node_modules
|
||||||
|
.git
|
||||||
|
*.md
|
||||||
|
Dockerfile
|
||||||
|
.dockerignore
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
{
|
||||||
|
"id": "ai.coop.litellm",
|
||||||
|
"title": "LiteLLM Gateway",
|
||||||
|
"author": "inference.coop",
|
||||||
|
"description": "LiteLLM AI Gateway — OpenAI-compatible proxy with per-user API keys, usage tracking, rate limiting, and model routing. Packaged for Cloudron with SSO/OIDC and PostgreSQL.",
|
||||||
|
"tagline": "AI gateway for cooperative inference",
|
||||||
|
"version": "1.0.0",
|
||||||
|
"upstreamVersion": "1.74.0",
|
||||||
|
"healthCheckPath": "/health/readiness",
|
||||||
|
"httpPort": 4000,
|
||||||
|
"manifestVersion": 2,
|
||||||
|
"website": "https://litellm.ai",
|
||||||
|
"contactEmail": "botbot@hermes",
|
||||||
|
"icon": "file://logo.png",
|
||||||
|
"addons": {
|
||||||
|
"postgresql": {},
|
||||||
|
"redis": {},
|
||||||
|
"localstorage": {}
|
||||||
|
},
|
||||||
|
"tags": ["ai", "gateway", "proxy", "api"],
|
||||||
|
"mediaLinks": [],
|
||||||
|
"changelog": "Initial package for inference.coop"
|
||||||
|
}
|
||||||
+25
@@ -0,0 +1,25 @@
|
|||||||
|
FROM ghcr.io/berriai/litellm:main-v1.74.0-stable
|
||||||
|
|
||||||
|
# Cloudron runs apps as the 'cloudron' user (uid 1000) by default.
|
||||||
|
# LiteLLM's official image runs as root; we switch to cloudron for security.
|
||||||
|
# However, LiteLLM needs to write to /app/data for its database/config.
|
||||||
|
|
||||||
|
USER root
|
||||||
|
|
||||||
|
# Install supervisor to manage processes (LiteLLM + optional cron)
|
||||||
|
RUN pip install --no-cache-dir supervisor
|
||||||
|
|
||||||
|
# Create the data directory and set ownership
|
||||||
|
RUN mkdir -p /app/data && chown -R cloudron:cloudron /app/data
|
||||||
|
|
||||||
|
# Create a startup script that configures LiteLLM with Cloudron addons
|
||||||
|
COPY start.sh /app/code/start.sh
|
||||||
|
RUN chmod +x /app/code/start.sh && chown cloudron:cloudron /app/code/start.sh
|
||||||
|
|
||||||
|
# Default config — will be overridden by Cloudron env vars at runtime
|
||||||
|
COPY config.yaml /app/data/config.yaml
|
||||||
|
RUN chown cloudron:cloudron /app/data/config.yaml
|
||||||
|
|
||||||
|
USER cloudron
|
||||||
|
|
||||||
|
CMD ["/app/code/start.sh"]
|
||||||
@@ -0,0 +1,151 @@
|
|||||||
|
# LiteLLM Gateway — Cloudron App
|
||||||
|
|
||||||
|
A Cloudron app package for [LiteLLM](https://litellm.ai/), the open-source AI gateway, configured for the **Inference Cooperative**.
|
||||||
|
|
||||||
|
LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.
|
||||||
|
|
||||||
|
## What this provides
|
||||||
|
|
||||||
|
- **LiteLLM proxy** with an OpenAI-compatible API on port 4000
|
||||||
|
- **Per-member virtual API keys** with budget caps and rate limits
|
||||||
|
- **Usage / spend tracking** per key, per model
|
||||||
|
- **Model routing** to Tinfoil (TEE-protected inference)
|
||||||
|
- **Cloudron PostgreSQL** for keys, teams, and spend logs
|
||||||
|
- **Cloudron Redis** for rate limiting and caching
|
||||||
|
- **Automatic SSL, backups, and sandboxing** (Cloudron-managed)
|
||||||
|
- **Master key + salt key** auto-generated on first start, persisted in `/app/data`
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
```
|
||||||
|
Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE)
|
||||||
|
```
|
||||||
|
|
||||||
|
LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs).
|
||||||
|
|
||||||
|
## Files
|
||||||
|
|
||||||
|
- `CloudronManifest.json` — Cloudron app manifest (addons, ports, metadata)
|
||||||
|
- `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment
|
||||||
|
- `start.sh` — startup script that wires up Cloudron's Postgres/Redis and writes the default config
|
||||||
|
- `config.yaml` — default LiteLLM config (Tinfoil-only, two models)
|
||||||
|
- `logo.png` — app icon
|
||||||
|
|
||||||
|
## Models (MVP)
|
||||||
|
|
||||||
|
Two models, both TEE-protected via Tinfoil:
|
||||||
|
|
||||||
|
| Model | Role | Price (in/out per M tokens) |
|
||||||
|
|-------|------|------------------------------|
|
||||||
|
| `deepseek-v4-flash` | Default — best agent performance, 1M context | $0.30 / $0.70 |
|
||||||
|
| `gpt-oss-120b` | Fallback — cheapest, built for agentic workflows | $0.15 / $0.60 |
|
||||||
|
|
||||||
|
If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Building and installing
|
||||||
|
|
||||||
|
This package uses Cloudron's **on-server build** (no local Docker, no registry). The Cloudron CLI uploads the source and Cloudron builds the image on the server.
|
||||||
|
|
||||||
|
### Prerequisites
|
||||||
|
|
||||||
|
- The [Cloudron CLI](https://docs.cloudron.io/packaging/cli/) installed on a machine with access to the Cloudron server
|
||||||
|
- A Cloudron instance with the `inference.coop` domain
|
||||||
|
|
||||||
|
### Install
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 1. Install the Cloudron CLI (once)
|
||||||
|
npm install -g cloudron-cli
|
||||||
|
|
||||||
|
# 2. Log in to the Cloudron
|
||||||
|
cloudron login my.inference.coop
|
||||||
|
|
||||||
|
# 3. From this directory, install (builds on the server)
|
||||||
|
cloudron install --location gateway
|
||||||
|
```
|
||||||
|
|
||||||
|
`--location gateway` installs the app at `gateway.inference.coop`. Omit `--location` to be prompted.
|
||||||
|
|
||||||
|
### Configure
|
||||||
|
|
||||||
|
After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables:
|
||||||
|
|
||||||
|
- `TINFOIL_API_KEY` — your Tinfoil API key (from https://tinfoil.sh/)
|
||||||
|
|
||||||
|
The following are auto-generated and should **not** be set manually:
|
||||||
|
|
||||||
|
- `LITELLM_MASTER_KEY` — generated on first start, stored in `/app/data/.master_key`
|
||||||
|
- `LITELLM_SALT_KEY` — generated on first start, stored in `/app/data/.salt_key`
|
||||||
|
- `DATABASE_URL` — from Cloudron's PostgreSQL addon
|
||||||
|
- `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon
|
||||||
|
|
||||||
|
### Access the admin UI
|
||||||
|
|
||||||
|
1. Open `https://gateway.inference.coop/ui`
|
||||||
|
2. Log in with username `admin` and the master key as password
|
||||||
|
3. Find the master key: `cloudron exec --app <app-id> cat /app/data/.master_key`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Updating LiteLLM
|
||||||
|
|
||||||
|
This is the routine maintenance path. LiteLLM releases frequently; updating is a two-step change.
|
||||||
|
|
||||||
|
### 1. Bump the base image tag
|
||||||
|
|
||||||
|
In `Dockerfile`, change the pinned version:
|
||||||
|
|
||||||
|
```dockerfile
|
||||||
|
FROM ghcr.io/berriai/litellm:main-v1.74.0-stable
|
||||||
|
# ^^^^^^ bump this
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2. Bump the manifest version
|
||||||
|
|
||||||
|
In `CloudronManifest.json`, update:
|
||||||
|
|
||||||
|
```json
|
||||||
|
"version": "1.0.0",
|
||||||
|
"upstreamVersion": "1.74.0"
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. Redeploy
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cloudron update --app <app-id>
|
||||||
|
```
|
||||||
|
|
||||||
|
Cloudron rebuilds from source and applies the update. Data (keys, spend logs) persists in PostgreSQL.
|
||||||
|
|
||||||
|
### Before updating
|
||||||
|
|
||||||
|
- Check the [LiteLLM changelog](https://github.com/BerriAI/litellm/releases) for breaking changes
|
||||||
|
- The `start.sh` and `config.yaml` may need adjustment if LiteLLM changes its env-var or config schema
|
||||||
|
- Test on a staging instance if the version jump is large
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Environment variables
|
||||||
|
|
||||||
|
| Variable | Required | Purpose |
|
||||||
|
|----------|----------|---------|
|
||||||
|
| `TINFOIL_API_KEY` | Yes | Tinfoil API key for TEE inference |
|
||||||
|
|
||||||
|
## Connecting Open WebUI
|
||||||
|
|
||||||
|
In Open WebUI settings:
|
||||||
|
|
||||||
|
- **API Base URL:** `https://gateway.inference.coop/v1`
|
||||||
|
- **API Key:** a virtual key created in LiteLLM's Admin UI
|
||||||
|
|
||||||
|
## OIDC / SSO (future)
|
||||||
|
|
||||||
|
The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the `"oidc": {}` addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso.
|
||||||
|
|
||||||
|
## Related repos
|
||||||
|
|
||||||
|
- `co-op/org-dev` — the full inference.coop design doc and deploy scripts
|
||||||
|
- `co-op/website` — the public landing page
|
||||||
|
- `co-op/design-assets` — logos, fonts, brand materials
|
||||||
+42
@@ -0,0 +1,42 @@
|
|||||||
|
# LiteLLM config for inference.coop
|
||||||
|
# MVP config: Tinfoil only (TEE-protected, privacy-first)
|
||||||
|
# Two models: DeepSeek V4 Flash (default, best agents) + GPT-OSS 120B (cheapest)
|
||||||
|
# Provider key set via environment variable or Admin UI
|
||||||
|
|
||||||
|
model_list:
|
||||||
|
# DeepSeek V4 Flash — default model
|
||||||
|
# Best agent performance (Terminal Bench 2.1: 82.7, DeepSWE: 54.4)
|
||||||
|
# MoE, 1M context, tool calling, $0.30/$0.70 per M tokens
|
||||||
|
- model_name: deepseek-v4-flash
|
||||||
|
litellm_params:
|
||||||
|
model: openai/deepseek-v4-flash
|
||||||
|
api_base: https://api.tinfoil.sh/v1
|
||||||
|
api_key: os.environ/TINFOIL_API_KEY
|
||||||
|
|
||||||
|
# GPT-OSS 120B — lightweight fallback
|
||||||
|
# Built for agentic workflows, web search + code execution
|
||||||
|
# Apache 2.0, $0.15/$0.60 per M tokens
|
||||||
|
- model_name: gpt-oss-120b
|
||||||
|
litellm_params:
|
||||||
|
model: openai/gpt-oss-120b
|
||||||
|
api_base: https://api.tinfoil.sh/v1
|
||||||
|
api_key: os.environ/TINFOIL_API_KEY
|
||||||
|
|
||||||
|
# Fallback: if DeepSeek is down, use GPT-OSS
|
||||||
|
router_settings:
|
||||||
|
num_retries: 2
|
||||||
|
timeout: 30
|
||||||
|
fallbacks:
|
||||||
|
- deepseek-v4-flash:
|
||||||
|
- gpt-oss-120b
|
||||||
|
|
||||||
|
general_settings:
|
||||||
|
master_key: os.environ/LITELLM_MASTER_KEY
|
||||||
|
database_url: os.environ/DATABASE_URL
|
||||||
|
store_model_in_db: true
|
||||||
|
|
||||||
|
litellm_settings:
|
||||||
|
salt_key: os.environ/LITELLM_SALT_KEY
|
||||||
|
drop_params: true
|
||||||
|
num_threads: 4
|
||||||
|
request_timeout: 30
|
||||||
@@ -0,0 +1,129 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
set -eu
|
||||||
|
|
||||||
|
# ============================================================================
|
||||||
|
# LiteLLM Cloudron Startup Script
|
||||||
|
# Configures LiteLLM with Cloudron's PostgreSQL and Redis addons
|
||||||
|
# ============================================================================
|
||||||
|
|
||||||
|
echo "=== LiteLLM Cloudron Startup ==="
|
||||||
|
|
||||||
|
# --- Cloudron PostgreSQL addon provides:
|
||||||
|
# CLOUDRON_POSTGRESQL_URL - postgresql://user:pass@host:port/dbname
|
||||||
|
# CLOUDRON_POSTGRESQL_USERNAME
|
||||||
|
# CLOUDRON_POSTGRESQL_PASSWORD
|
||||||
|
# CLOUDRON_POSTGRESQL_HOST
|
||||||
|
# CLOUDRON_POSTGRESQL_PORT
|
||||||
|
# CLOUDRON_POSTGRESQL_DATABASE
|
||||||
|
|
||||||
|
# --- Cloudron Redis addon provides:
|
||||||
|
# CLOUDRON_REDIS_URL - redis://user:pass@host:port
|
||||||
|
# CLOUDRON_REDIS_USERNAME
|
||||||
|
# CLOUDRON_REDIS_PASSWORD
|
||||||
|
# CLOUDRON_REDIS_HOST
|
||||||
|
# CLOUDRON_REDIS_PORT
|
||||||
|
|
||||||
|
# --- Cloudron provides:
|
||||||
|
# CLOUDRON_APP_DOMAIN - the app's domain (e.g. gateway.inference.coop)
|
||||||
|
# CLOUDRON_API_ORIGIN - origin for API calls
|
||||||
|
# CLOUDRON_WEB_ORIGIN - origin for web UI
|
||||||
|
|
||||||
|
# --- LiteLLM required env vars ---
|
||||||
|
|
||||||
|
# Generate a master key if not already in /app/data
|
||||||
|
MASTER_KEY_FILE="/app/data/.master_key"
|
||||||
|
if [ ! -f "$MASTER_KEY_FILE" ]; then
|
||||||
|
echo "Generating master key..."
|
||||||
|
python3 -c "import secrets; print('sk-' + secrets.token_urlsafe(32))" > "$MASTER_KEY_FILE"
|
||||||
|
fi
|
||||||
|
export LITELLM_MASTER_KEY=$(cat "$MASTER_KEY_FILE")
|
||||||
|
|
||||||
|
# Generate a salt key if not already in /app/data
|
||||||
|
SALT_KEY_FILE="/app/data/.salt_key"
|
||||||
|
if [ ! -f "$SALT_KEY_FILE" ]; then
|
||||||
|
echo "Generating salt key..."
|
||||||
|
python3 -c "import secrets; print(secrets.token_urlsafe(32))" > "$SALT_KEY_FILE"
|
||||||
|
fi
|
||||||
|
export LITELLM_SALT_KEY=$(cat "$SALT_KEY_FILE")
|
||||||
|
|
||||||
|
# Use Cloudron's PostgreSQL
|
||||||
|
export DATABASE_URL="${CLOUDRON_POSTGRESQL_URL}"
|
||||||
|
|
||||||
|
# Use Cloudron's Redis (for rate limiting, caching)
|
||||||
|
export REDIS_HOST="${CLOUDRON_REDIS_HOST}"
|
||||||
|
export REDIS_PORT="${CLOUDRON_REDIS_PORT}"
|
||||||
|
export REDIS_PASSWORD="${CLOUDRON_REDIS_PASSWORD}"
|
||||||
|
|
||||||
|
# Store models in DB (manage from Admin UI)
|
||||||
|
export STORE_MODEL_IN_DB="True"
|
||||||
|
|
||||||
|
# Don't run schema migrations on every start — let LiteLLM handle it
|
||||||
|
# (for single-instance deployment this is fine)
|
||||||
|
export DISABLE_SCHEMA_UPDATE="false"
|
||||||
|
|
||||||
|
# UI configuration — use Cloudron's domain
|
||||||
|
export UI_USERNAME="admin"
|
||||||
|
export UI_PASSWORD="${LITELLM_MASTER_KEY}"
|
||||||
|
|
||||||
|
# Proxy settings for Cloudron's reverse proxy
|
||||||
|
export LITELLM_PROXY_BASE_URL="https://${CLOUDRON_APP_DOMAIN}"
|
||||||
|
|
||||||
|
# Allow requests from Cloudron's Open WebUI, LobeChat, etc.
|
||||||
|
export LITELLM_CORS_ALLOWED_ORIGINS="${CLOUDRON_WEB_ORIGIN:-*}"
|
||||||
|
|
||||||
|
echo "PostgreSQL: ${CLOUDRON_POSTGRESQL_HOST}:${CLOUDRON_POSTGRESQL_PORT}"
|
||||||
|
echo "Redis: ${CLOUDRON_REDIS_HOST}:${CLOUDRON_REDIS_PORT}"
|
||||||
|
echo "Domain: ${CLOUDRON_APP_DOMAIN}"
|
||||||
|
|
||||||
|
# --- Write config.yaml with Cloudron provider keys if set ---
|
||||||
|
|
||||||
|
CONFIG_FILE="/app/data/config.yaml"
|
||||||
|
|
||||||
|
if [ ! -f "$CONFIG_FILE" ] || [ ! -s "$CONFIG_FILE" ]; then
|
||||||
|
echo "Writing default config.yaml..."
|
||||||
|
cat > "$CONFIG_FILE" << 'YAML'
|
||||||
|
# LiteLLM config for inference.coop
|
||||||
|
# MVP config: Tinfoil only (TEE-protected, privacy-first)
|
||||||
|
# Two models: DeepSeek V4 Flash (default, best agents) + GPT-OSS 120B (cheapest)
|
||||||
|
# Provider key set via environment variable or Admin UI
|
||||||
|
|
||||||
|
model_list:
|
||||||
|
# DeepSeek V4 Flash — default model
|
||||||
|
- model_name: deepseek-v4-flash
|
||||||
|
litellm_params:
|
||||||
|
model: openai/deepseek-v4-flash
|
||||||
|
api_base: https://api.tinfoil.sh/v1
|
||||||
|
api_key: os.environ/TINFOIL_API_KEY
|
||||||
|
|
||||||
|
# GPT-OSS 120B — lightweight fallback
|
||||||
|
- model_name: gpt-oss-120b
|
||||||
|
litellm_params:
|
||||||
|
model: openai/gpt-oss-120b
|
||||||
|
api_base: https://api.tinfoil.sh/v1
|
||||||
|
api_key: os.environ/TINFOIL_API_KEY
|
||||||
|
|
||||||
|
# Fallback: if DeepSeek is down, use GPT-OSS
|
||||||
|
router_settings:
|
||||||
|
num_retries: 2
|
||||||
|
timeout: 30
|
||||||
|
fallbacks:
|
||||||
|
- deepseek-v4-flash:
|
||||||
|
- gpt-oss-120b
|
||||||
|
|
||||||
|
general_settings:
|
||||||
|
master_key: os.environ/LITELLM_MASTER_KEY
|
||||||
|
database_url: os.environ/DATABASE_URL
|
||||||
|
store_model_in_db: true
|
||||||
|
|
||||||
|
litellm_settings:
|
||||||
|
salt_key: os.environ/LITELLM_SALT_KEY
|
||||||
|
drop_params: true
|
||||||
|
num_threads: 4
|
||||||
|
request_timeout: 30
|
||||||
|
YAML
|
||||||
|
echo "Config written to $CONFIG_FILE"
|
||||||
|
fi
|
||||||
|
|
||||||
|
# --- Start LiteLLM ---
|
||||||
|
echo "Starting LiteLLM on port 4000..."
|
||||||
|
exec litellm --config "$CONFIG_FILE" --port 4000 --host 0.0.0.0
|
||||||
Reference in new issue
Block a user