Compare commits
4
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
b6e8542146 | ||
|
|
b9af62a614 | ||
|
|
a2864ce34a | ||
|
|
23f73cf6ae |
No files matched your search
@@ -2,10 +2,10 @@
|
|||||||
"id": "ai.coop.litellm",
|
"id": "ai.coop.litellm",
|
||||||
"title": "LiteLLM Gateway",
|
"title": "LiteLLM Gateway",
|
||||||
"author": "inference.coop",
|
"author": "inference.coop",
|
||||||
"description": "LiteLLM AI Gateway — OpenAI-compatible proxy with per-user API keys, usage tracking, rate limiting, and model routing. Packaged for Cloudron with SSO/OIDC and PostgreSQL.",
|
"description": "LiteLLM AI Gateway \u2014 OpenAI-compatible proxy with per-user API keys, usage tracking, rate limiting, and model routing. Packaged for Cloudron with SSO/OIDC and PostgreSQL.",
|
||||||
"tagline": "AI gateway for cooperative inference",
|
"tagline": "AI gateway for cooperative inference",
|
||||||
"version": "1.0.0",
|
"version": "1.0.0",
|
||||||
"upstreamVersion": "1.84.0",
|
"upstreamVersion": "1.103.2",
|
||||||
"healthCheckPath": "/health/readiness",
|
"healthCheckPath": "/health/readiness",
|
||||||
"httpPort": 4000,
|
"httpPort": 4000,
|
||||||
"memoryLimit": 2147483648,
|
"memoryLimit": 2147483648,
|
||||||
@@ -18,7 +18,12 @@
|
|||||||
"redis": {},
|
"redis": {},
|
||||||
"localstorage": {}
|
"localstorage": {}
|
||||||
},
|
},
|
||||||
"tags": ["ai", "gateway", "proxy", "api"],
|
"tags": [
|
||||||
|
"ai",
|
||||||
|
"gateway",
|
||||||
|
"proxy",
|
||||||
|
"api"
|
||||||
|
],
|
||||||
"mediaLinks": [],
|
"mediaLinks": [],
|
||||||
"changelog": "Initial package for inference.coop"
|
"changelog": "Initial package for inference.coop"
|
||||||
}
|
}
|
||||||
+1
-1
@@ -1,4 +1,4 @@
|
|||||||
FROM ghcr.io/berriai/litellm:v1.84.0
|
FROM ghcr.io/berriai/litellm:v1.103.2
|
||||||
|
|
||||||
# v1.84.0 is the fix line for CVE-2026-35029 (auth bypass on
|
# v1.84.0 is the fix line for CVE-2026-35029 (auth bypass on
|
||||||
# /config/update, fixed 1.83.0) and CVE-2026-59822 (MCP session auth
|
# /config/update, fixed 1.83.0) and CVE-2026-59822 (MCP session auth
|
||||||
|
|||||||
@@ -4,43 +4,45 @@ A Cloudron app package for [LiteLLM](https://litellm.ai/), the open-source AI ga
|
|||||||
|
|
||||||
LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.
|
LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint.
|
||||||
|
|
||||||
|
**Current base image:** LiteLLM **v1.103.2**, pinned to the exact tag (fix line for CVE-2026-35029 / CVE-2026-59822, CISA KEV-listed; also carries the streaming-usage fix for prompt_tokens_details). See [Updating LiteLLM](#updating-litellm).
|
||||||
|
|
||||||
## What this provides
|
## What this provides
|
||||||
|
|
||||||
- **LiteLLM proxy** with an OpenAI-compatible API on port 4000
|
- **LiteLLM proxy** with an OpenAI-compatible API on port 4000
|
||||||
- **Per-member virtual API keys** with budget caps and rate limits
|
- **Per-member virtual API keys** with budget caps and rate limits
|
||||||
- **Usage / spend tracking** per key, per model
|
- **Usage / spend tracking** per key, per model
|
||||||
- **Model routing** to Tinfoil (TEE-protected inference)
|
- **Model routing** to multiple providers, chosen by member values (private / green / public)
|
||||||
- **Cloudron PostgreSQL** for keys, teams, and spend logs
|
- **Cloudron PostgreSQL** for keys, teams, and spend logs
|
||||||
- **Cloudron Redis** for rate limiting and caching
|
- **Cloudron Redis** for rate limiting and caching
|
||||||
- **Automatic SSL, backups, and sandboxing** (Cloudron-managed)
|
- **Automatic SSL, backups, and sandboxing** (Cloudron-managed)
|
||||||
- **Master key + salt key** auto-generated on first start, persisted in `/app/data`
|
- **Master key + salt key** auto-generated on first start, persisted in `/app/data`
|
||||||
|
- Hardened for the public internet: admin UI and API docs disabled at the proxy; CORS locked to the chat origin
|
||||||
|
|
||||||
## Architecture
|
## Architecture
|
||||||
|
|
||||||
```
|
```
|
||||||
Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE)
|
Members → Open WebUI (chat) / own tools (API) → member portal (key injection)
|
||||||
|
→ LiteLLM gateway (this app)
|
||||||
|
├→ Tinfoil (TEE enclaves — architectural privacy)
|
||||||
|
├→ GreenPT (100% renewable energy)
|
||||||
|
└→ PublicAI (publicly developed, sovereign models)
|
||||||
```
|
```
|
||||||
|
|
||||||
LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs).
|
LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to providers, each chosen by members for a different value. A [tinfoil-proxy](https://github.com/tinfoilsh/tinfoil-proxy) sidecar provides EHBP end-to-end encryption to Tinfoil's enclaves.
|
||||||
|
|
||||||
## Files
|
## Files
|
||||||
|
|
||||||
- `CloudronManifest.json` — Cloudron app manifest (addons, ports, metadata)
|
- `CloudronManifest.json` — Cloudron app manifest (addons, ports, memory limit, metadata)
|
||||||
- `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment
|
- `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment (pinned to an exact tag)
|
||||||
- `start.sh` — startup script that wires up Cloudron's Postgres/Redis and writes the default config
|
- `start.sh` — startup script: wires Cloudron's Postgres/Redis, downloads the tinfoil-proxy sidecar, exports `LITELLM_MIGRATION_DIR`
|
||||||
- `config.yaml` — default LiteLLM config (Tinfoil-only, two models)
|
- `config.yaml` — LiteLLM config: model catalog, provider routing, retry policy, custom per-token pricing
|
||||||
- `logo.png` — app icon
|
- `logo.png` — app icon
|
||||||
|
|
||||||
## Models (MVP)
|
## Models
|
||||||
|
|
||||||
Two models, both TEE-protected via Tinfoil:
|
The catalog is defined in `config.yaml` and documented for members in [co-op/docs models](https://git.inference.coop/co-op/docs/src/branch/main/models.md) — that page is the single source of truth; this repo's `config.yaml` mirrors it. Models follow the `provider/model-name` convention so members can choose not just which model but which provider values it embodies (privacy, renewable energy, public development).
|
||||||
|
|
||||||
| Model | Role | Price (in/out per M tokens) |
|
The catalog changes by member decision — expect it to evolve.
|
||||||
|-------|------|------------------------------|
|
|
||||||
| `deepseek-v4-flash` | Default — best agent performance, 1M context | $0.30 / $0.70 |
|
|
||||||
| `gpt-oss-120b` | Fallback — cheapest, built for agentic workflows | $0.15 / $0.60 |
|
|
||||||
|
|
||||||
If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS.
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -70,9 +72,11 @@ cloudron install --location gateway
|
|||||||
|
|
||||||
### Configure
|
### Configure
|
||||||
|
|
||||||
After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables:
|
After install, set the provider API keys in Cloudron → app → Settings → Environment Variables:
|
||||||
|
|
||||||
- `TINFOIL_API_KEY` — your Tinfoil API key (from https://tinfoil.sh/)
|
- `TINFOIL_API_KEY` — Tinfoil (https://tinfoil.sh/)
|
||||||
|
- `GREENPT_API_KEY` — GreenPT (https://greenpt.ai/)
|
||||||
|
- `PUBLICAI_API_KEY` — PublicAI (https://publicai.co/)
|
||||||
|
|
||||||
The following are auto-generated and should **not** be set manually:
|
The following are auto-generated and should **not** be set manually:
|
||||||
|
|
||||||
@@ -81,11 +85,9 @@ The following are auto-generated and should **not** be set manually:
|
|||||||
- `DATABASE_URL` — from Cloudron's PostgreSQL addon
|
- `DATABASE_URL` — from Cloudron's PostgreSQL addon
|
||||||
- `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon
|
- `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon
|
||||||
|
|
||||||
### Access the admin UI
|
### Admin access
|
||||||
|
|
||||||
1. Open `https://gateway.inference.coop/ui`
|
The LiteLLM admin UI and API docs are **disabled** on the public gateway (hardened by default; see `start.sh`). Administrative operations run through the [member portal](https://git.inference.coop/code/member-portal), which holds the master key server-side. If direct LiteLLM admin access is ever needed, re-enable the UI in `start.sh`, redeploy, and use the master key from `/app/data/.master_key` — then disable it again.
|
||||||
2. Log in with username `admin` and the master key as password
|
|
||||||
3. Find the master key: `cloudron exec --app <app-id> cat /app/data/.master_key`
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -165,21 +167,24 @@ expect a 15–20 min maintenance window and **never panic-restart mid-migration*
|
|||||||
|
|
||||||
| Variable | Required | Purpose |
|
| Variable | Required | Purpose |
|
||||||
|----------|----------|---------|
|
|----------|----------|---------|
|
||||||
| `TINFOIL_API_KEY` | Yes | Tinfoil API key for TEE inference |
|
| `TINFOIL_API_KEY` | Yes | Tinfoil provider (TEE inference) |
|
||||||
|
| `GREENPT_API_KEY` | Yes | GreenPT provider (renewable energy) |
|
||||||
|
| `PUBLICAI_API_KEY` | Yes | PublicAI provider (public models) |
|
||||||
|
| `LITELLM_CORS_ORIGINS` | Set by start.sh | CORS allowlist (chat origin only) |
|
||||||
|
| `LITELLM_MIGRATION_DIR` | Set by start.sh | Persistent Prisma migrations ledger |
|
||||||
|
|
||||||
## Connecting Open WebUI
|
## Connecting chat clients
|
||||||
|
|
||||||
In Open WebUI settings:
|
Open WebUI does **not** talk to this gateway directly — it routes through the [member portal](https://git.inference.coop/code/member-portal), which authenticates the member and injects their personal LiteLLM key per request. Members building their own tools create API keys on the [member dashboard](https://dashboard.inference.coop) and use:
|
||||||
|
|
||||||
- **API Base URL:** `https://gateway.inference.coop/v1`
|
- **API Base URL:** `https://gateway.inference.coop/v1` (or the portal's proxy endpoint for chat-integrated tools)
|
||||||
- **API Key:** a virtual key created in LiteLLM's Admin UI
|
- **API Key:** their own virtual key from the dashboard
|
||||||
|
|
||||||
## OIDC / SSO (future)
|
|
||||||
|
|
||||||
The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the `"oidc": {}` addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso.
|
|
||||||
|
|
||||||
## Related repos
|
## Related repos
|
||||||
|
|
||||||
- `co-op/org-dev` — the full inference.coop design doc and deploy scripts
|
- `code/member-portal` — membership middleware: provisioning, billing webhooks, key injection
|
||||||
|
- `code/member-dashboard` — member-facing usage and API-key management
|
||||||
|
- `code/admin-panel` — admin overview
|
||||||
- `co-op/website` — the public landing page
|
- `co-op/website` — the public landing page
|
||||||
|
- `co-op/docs` — documentation, including the member-facing model list
|
||||||
- `co-op/design-assets` — logos, fonts, brand materials
|
- `co-op/design-assets` — logos, fonts, brand materials
|
||||||
-17
@@ -103,24 +103,7 @@ model_list:
|
|||||||
# (Open WebUI) by their `mode` field, so they never appear as pickable chat
|
# (Open WebUI) by their `mode` field, so they never appear as pickable chat
|
||||||
# options. Pricing per Tinfoil catalog: whisper $0.05/1M in; voxtral-tts
|
# options. Pricing per Tinfoil catalog: whisper $0.05/1M in; voxtral-tts
|
||||||
# listed $0/$0 (verify on invoice — may be beta-free).
|
# listed $0/$0 (verify on invoice — may be beta-free).
|
||||||
- model_name: tinfoil/whisper-large-v3-turbo
|
|
||||||
litellm_params:
|
|
||||||
model: audio_transcription/whisper-large-v3-turbo
|
|
||||||
api_base: http://127.0.0.1:3301/v1
|
|
||||||
api_key: os.environ/TINFOIL_API_KEY
|
|
||||||
model_info:
|
|
||||||
mode: audio_transcription
|
|
||||||
input_cost_per_token: 0.00000005
|
|
||||||
|
|
||||||
- model_name: tinfoil/voxtral-tts
|
|
||||||
litellm_params:
|
|
||||||
model: audio_speech/voxtral-tts
|
|
||||||
api_base: http://127.0.0.1:3301/v1
|
|
||||||
api_key: os.environ/TINFOIL_API_KEY
|
|
||||||
model_info:
|
|
||||||
mode: audio_speech
|
|
||||||
input_cost_per_token: 0.0
|
|
||||||
output_cost_per_token: 0.0
|
|
||||||
|
|
||||||
# --- PublicAI (publicly developed / sovereign models) ---
|
# --- PublicAI (publicly developed / sovereign models) ---
|
||||||
# OpenAI-compatible gateway for public open models. Apertus is the Swiss AI
|
# OpenAI-compatible gateway for public open models. Apertus is the Swiss AI
|
||||||
|
|||||||
Reference in new issue
Block a user