Compare commits

...
4 changed files with 45 additions and 52 deletions

No files matched your search

+8 -3
View File
@@ -2,10 +2,10 @@
"id": "ai.coop.litellm", "id": "ai.coop.litellm",
"title": "LiteLLM Gateway", "title": "LiteLLM Gateway",
"author": "inference.coop", "author": "inference.coop",
"description": "LiteLLM AI Gateway — OpenAI-compatible proxy with per-user API keys, usage tracking, rate limiting, and model routing. Packaged for Cloudron with SSO/OIDC and PostgreSQL.", "description": "LiteLLM AI Gateway \u2014 OpenAI-compatible proxy with per-user API keys, usage tracking, rate limiting, and model routing. Packaged for Cloudron with SSO/OIDC and PostgreSQL.",
"tagline": "AI gateway for cooperative inference", "tagline": "AI gateway for cooperative inference",
"version": "1.0.0", "version": "1.0.0",
"upstreamVersion": "1.84.0", "upstreamVersion": "1.103.2",
"healthCheckPath": "/health/readiness", "healthCheckPath": "/health/readiness",
"httpPort": 4000, "httpPort": 4000,
"memoryLimit": 2147483648, "memoryLimit": 2147483648,
@@ -18,7 +18,12 @@
"redis": {}, "redis": {},
"localstorage": {} "localstorage": {}
}, },
"tags": ["ai", "gateway", "proxy", "api"], "tags": [
"ai",
"gateway",
"proxy",
"api"
],
"mediaLinks": [], "mediaLinks": [],
"changelog": "Initial package for inference.coop" "changelog": "Initial package for inference.coop"
} }
+1 -1
View File
@@ -1,4 +1,4 @@
FROM ghcr.io/berriai/litellm:v1.84.0 FROM ghcr.io/berriai/litellm:v1.103.2
# v1.84.0 is the fix line for CVE-2026-35029 (auth bypass on # v1.84.0 is the fix line for CVE-2026-35029 (auth bypass on
# /config/update, fixed 1.83.0) and CVE-2026-59822 (MCP session auth # /config/update, fixed 1.83.0) and CVE-2026-59822 (MCP session auth
+36 -31
View File
@@ -4,43 +4,45 @@ A Cloudron app package for [LiteLLM](https://litellm.ai/), the open-source AI ga
LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint. LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.
**Current base image:** LiteLLM **v1.103.2**, pinned to the exact tag (fix line for CVE-2026-35029 / CVE-2026-59822, CISA KEV-listed; also carries the streaming-usage fix for prompt_tokens_details). See [Updating LiteLLM](#updating-litellm).
## What this provides ## What this provides
- **LiteLLM proxy** with an OpenAI-compatible API on port 4000 - **LiteLLM proxy** with an OpenAI-compatible API on port 4000
- **Per-member virtual API keys** with budget caps and rate limits - **Per-member virtual API keys** with budget caps and rate limits
- **Usage / spend tracking** per key, per model - **Usage / spend tracking** per key, per model
- **Model routing** to Tinfoil (TEE-protected inference) - **Model routing** to multiple providers, chosen by member values (private / green / public)
- **Cloudron PostgreSQL** for keys, teams, and spend logs - **Cloudron PostgreSQL** for keys, teams, and spend logs
- **Cloudron Redis** for rate limiting and caching - **Cloudron Redis** for rate limiting and caching
- **Automatic SSL, backups, and sandboxing** (Cloudron-managed) - **Automatic SSL, backups, and sandboxing** (Cloudron-managed)
- **Master key + salt key** auto-generated on first start, persisted in `/app/data` - **Master key + salt key** auto-generated on first start, persisted in `/app/data`
- Hardened for the public internet: admin UI and API docs disabled at the proxy; CORS locked to the chat origin
## Architecture ## Architecture
``` ```
Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE) Members → Open WebUI (chat) / own tools (API) → member portal (key injection)
→ LiteLLM gateway (this app)
├→ Tinfoil (TEE enclaves — architectural privacy)
├→ GreenPT (100% renewable energy)
└→ PublicAI (publicly developed, sovereign models)
``` ```
LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs). LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to providers, each chosen by members for a different value. A [tinfoil-proxy](https://github.com/tinfoilsh/tinfoil-proxy) sidecar provides EHBP end-to-end encryption to Tinfoil's enclaves.
## Files ## Files
- `CloudronManifest.json` — Cloudron app manifest (addons, ports, metadata) - `CloudronManifest.json` — Cloudron app manifest (addons, ports, memory limit, metadata)
- `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment - `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment (pinned to an exact tag)
- `start.sh` — startup script that wires up Cloudron's Postgres/Redis and writes the default config - `start.sh` — startup script: wires Cloudron's Postgres/Redis, downloads the tinfoil-proxy sidecar, exports `LITELLM_MIGRATION_DIR`
- `config.yaml` — default LiteLLM config (Tinfoil-only, two models) - `config.yaml` — LiteLLM config: model catalog, provider routing, retry policy, custom per-token pricing
- `logo.png` — app icon - `logo.png` — app icon
## Models (MVP) ## Models
Two models, both TEE-protected via Tinfoil: The catalog is defined in `config.yaml` and documented for members in [co-op/docs models](https://git.inference.coop/co-op/docs/src/branch/main/models.md) — that page is the single source of truth; this repo's `config.yaml` mirrors it. Models follow the `provider/model-name` convention so members can choose not just which model but which provider values it embodies (privacy, renewable energy, public development).
| Model | Role | Price (in/out per M tokens) | The catalog changes by member decision — expect it to evolve.
|-------|------|------------------------------|
| `deepseek-v4-flash` | Default — best agent performance, 1M context | $0.30 / $0.70 |
| `gpt-oss-120b` | Fallback — cheapest, built for agentic workflows | $0.15 / $0.60 |
If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS.
--- ---
@@ -70,9 +72,11 @@ cloudron install --location gateway
### Configure ### Configure
After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables: After install, set the provider API keys in Cloudron → app → Settings → Environment Variables:
- `TINFOIL_API_KEY` — your Tinfoil API key (from https://tinfoil.sh/) - `TINFOIL_API_KEY` — Tinfoil (https://tinfoil.sh/)
- `GREENPT_API_KEY` — GreenPT (https://greenpt.ai/)
- `PUBLICAI_API_KEY` — PublicAI (https://publicai.co/)
The following are auto-generated and should **not** be set manually: The following are auto-generated and should **not** be set manually:
@@ -81,11 +85,9 @@ The following are auto-generated and should **not** be set manually:
- `DATABASE_URL` — from Cloudron's PostgreSQL addon - `DATABASE_URL` — from Cloudron's PostgreSQL addon
- `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon - `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon
### Access the admin UI ### Admin access
1. Open `https://gateway.inference.coop/ui` The LiteLLM admin UI and API docs are **disabled** on the public gateway (hardened by default; see `start.sh`). Administrative operations run through the [member portal](https://git.inference.coop/code/member-portal), which holds the master key server-side. If direct LiteLLM admin access is ever needed, re-enable the UI in `start.sh`, redeploy, and use the master key from `/app/data/.master_key` — then disable it again.
2. Log in with username `admin` and the master key as password
3. Find the master key: `cloudron exec --app <app-id> cat /app/data/.master_key`
--- ---
@@ -165,21 +167,24 @@ expect a 15–20 min maintenance window and **never panic-restart mid-migration*
| Variable | Required | Purpose | | Variable | Required | Purpose |
|----------|----------|---------| |----------|----------|---------|
| `TINFOIL_API_KEY` | Yes | Tinfoil API key for TEE inference | | `TINFOIL_API_KEY` | Yes | Tinfoil provider (TEE inference) |
| `GREENPT_API_KEY` | Yes | GreenPT provider (renewable energy) |
| `PUBLICAI_API_KEY` | Yes | PublicAI provider (public models) |
| `LITELLM_CORS_ORIGINS` | Set by start.sh | CORS allowlist (chat origin only) |
| `LITELLM_MIGRATION_DIR` | Set by start.sh | Persistent Prisma migrations ledger |
## Connecting Open WebUI ## Connecting chat clients
In Open WebUI settings: Open WebUI does **not** talk to this gateway directly — it routes through the [member portal](https://git.inference.coop/code/member-portal), which authenticates the member and injects their personal LiteLLM key per request. Members building their own tools create API keys on the [member dashboard](https://dashboard.inference.coop) and use:
- **API Base URL:** `https://gateway.inference.coop/v1` - **API Base URL:** `https://gateway.inference.coop/v1` (or the portal's proxy endpoint for chat-integrated tools)
- **API Key:** a virtual key created in LiteLLM's Admin UI - **API Key:** their own virtual key from the dashboard
## OIDC / SSO (future)
The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the `"oidc": {}` addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso.
## Related repos ## Related repos
- `co-op/org-dev` — the full inference.coop design doc and deploy scripts - `code/member-portal` — membership middleware: provisioning, billing webhooks, key injection
- `code/member-dashboard` — member-facing usage and API-key management
- `code/admin-panel` — admin overview
- `co-op/website` — the public landing page - `co-op/website` — the public landing page
- `co-op/docs` — documentation, including the member-facing model list
- `co-op/design-assets` — logos, fonts, brand materials - `co-op/design-assets` — logos, fonts, brand materials
-17
View File
@@ -103,24 +103,7 @@ model_list:
# (Open WebUI) by their `mode` field, so they never appear as pickable chat # (Open WebUI) by their `mode` field, so they never appear as pickable chat
# options. Pricing per Tinfoil catalog: whisper $0.05/1M in; voxtral-tts # options. Pricing per Tinfoil catalog: whisper $0.05/1M in; voxtral-tts
# listed $0/$0 (verify on invoice — may be beta-free). # listed $0/$0 (verify on invoice — may be beta-free).
- model_name: tinfoil/whisper-large-v3-turbo
litellm_params:
model: audio_transcription/whisper-large-v3-turbo
api_base: http://127.0.0.1:3301/v1
api_key: os.environ/TINFOIL_API_KEY
model_info:
mode: audio_transcription
input_cost_per_token: 0.00000005
- model_name: tinfoil/voxtral-tts
litellm_params:
model: audio_speech/voxtral-tts
api_base: http://127.0.0.1:3301/v1
api_key: os.environ/TINFOIL_API_KEY
model_info:
mode: audio_speech
input_cost_per_token: 0.0
output_cost_per_token: 0.0
# --- PublicAI (publicly developed / sovereign models) --- # --- PublicAI (publicly developed / sovereign models) ---
# OpenAI-compatible gateway for public open models. Apertus is the Swiss AI # OpenAI-compatible gateway for public open models. Apertus is the Swiss AI