From 23f73cf6ae4b8bd8aef97dd1a77a74a256d49ef3 Mon Sep 17 00:00:00 2001 From: Inferencebot Date: Fri, 2 Oct 2026 01:08:34 -0600 Subject: [PATCH] README: current architecture (3 providers, portal key injection, hardening), remove MVP model list (point to co-op/docs as source of truth), all provider env vars, member-key verification kept --- README.md | 67 ++++++++++++++++++++++++++++++------------------------- 1 file changed, 36 insertions(+), 31 deletions(-) diff --git a/README.md b/README.md index 2ccf6d8..fb35a1c 100644 --- a/README.md +++ b/README.md @@ -4,43 +4,45 @@ A Cloudron app package for [LiteLLM](https://litellm.ai/), the open-source AI ga LiteLLM is the single gateway between members and the inference backends. It handles per-member API keys, usage metering, rate limiting (via budget caps), and model routing — all behind one OpenAI-compatible endpoint. +**Current base image:** LiteLLM **v1.84.0**, pinned to the exact tag — the fix line for CVE-2026-35029 and CVE-2026-59822 (CISA KEV-listed). See [Updating LiteLLM](#updating-litellm). + ## What this provides - **LiteLLM proxy** with an OpenAI-compatible API on port 4000 - **Per-member virtual API keys** with budget caps and rate limits - **Usage / spend tracking** per key, per model -- **Model routing** to Tinfoil (TEE-protected inference) +- **Model routing** to multiple providers, chosen by member values (private / green / public) - **Cloudron PostgreSQL** for keys, teams, and spend logs - **Cloudron Redis** for rate limiting and caching - **Automatic SSL, backups, and sandboxing** (Cloudron-managed) - **Master key + salt key** auto-generated on first start, persisted in `/app/data` +- Hardened for the public internet: admin UI and API docs disabled at the proxy; CORS locked to the chat origin ## Architecture ``` -Members → Open WebUI / LobeChat → LiteLLM (this app) → Tinfoil (TEE) +Members → Open WebUI (chat) / own tools (API) → member portal (key injection) + → LiteLLM gateway (this app) + ├→ Tinfoil (TEE enclaves — architectural privacy) + ├→ GreenPT (100% renewable energy) + └→ PublicAI (publicly developed, sovereign models) ``` -LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to Tinfoil, which runs open models inside trusted execution environments (TEEs). +LiteLLM is the control plane (keys, metering, routing). It does **not** run models itself — it forwards to providers, each chosen by members for a different value. A [tinfoil-proxy](https://github.com/tinfoilsh/tinfoil-proxy) sidecar provides EHBP end-to-end encryption to Tinfoil's enclaves. ## Files -- `CloudronManifest.json` — Cloudron app manifest (addons, ports, metadata) -- `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment -- `start.sh` — startup script that wires up Cloudron's Postgres/Redis and writes the default config -- `config.yaml` — default LiteLLM config (Tinfoil-only, two models) +- `CloudronManifest.json` — Cloudron app manifest (addons, ports, memory limit, metadata) +- `Dockerfile` — wraps LiteLLM's official image for Cloudron's environment (pinned to an exact tag) +- `start.sh` — startup script: wires Cloudron's Postgres/Redis, downloads the tinfoil-proxy sidecar, exports `LITELLM_MIGRATION_DIR` +- `config.yaml` — LiteLLM config: model catalog, provider routing, retry policy, custom per-token pricing - `logo.png` — app icon -## Models (MVP) +## Models -Two models, both TEE-protected via Tinfoil: +The catalog is defined in `config.yaml` and documented for members in [co-op/docs models](https://git.inference.coop/co-op/docs/src/branch/main/models.md) — that page is the single source of truth; this repo's `config.yaml` mirrors it. Models follow the `provider/model-name` convention so members can choose not just which model but which provider values it embodies (privacy, renewable energy, public development). -| Model | Role | Price (in/out per M tokens) | -|-------|------|------------------------------| -| `deepseek-v4-flash` | Default — best agent performance, 1M context | $0.30 / $0.70 | -| `gpt-oss-120b` | Fallback — cheapest, built for agentic workflows | $0.15 / $0.60 | - -If DeepSeek is down, LiteLLM automatically falls back to GPT-OSS. +The catalog changes by member decision — expect it to evolve. --- @@ -70,9 +72,11 @@ cloudron install --location gateway ### Configure -After install, set the Tinfoil API key in Cloudron → app → Settings → Environment Variables: +After install, set the provider API keys in Cloudron → app → Settings → Environment Variables: -- `TINFOIL_API_KEY` — your Tinfoil API key (from https://tinfoil.sh/) +- `TINFOIL_API_KEY` — Tinfoil (https://tinfoil.sh/) +- `GREENPT_API_KEY` — GreenPT (https://greenpt.ai/) +- `PUBLICAI_API_KEY` — PublicAI (https://publicai.co/) The following are auto-generated and should **not** be set manually: @@ -81,11 +85,9 @@ The following are auto-generated and should **not** be set manually: - `DATABASE_URL` — from Cloudron's PostgreSQL addon - `REDIS_HOST`, `REDIS_PORT`, `REDIS_PASSWORD` — from Cloudron's Redis addon -### Access the admin UI +### Admin access -1. Open `https://gateway.inference.coop/ui` -2. Log in with username `admin` and the master key as password -3. Find the master key: `cloudron exec --app cat /app/data/.master_key` +The LiteLLM admin UI and API docs are **disabled** on the public gateway (hardened by default; see `start.sh`). Administrative operations run through the [member portal](https://git.inference.coop/code/member-portal), which holds the master key server-side. If direct LiteLLM admin access is ever needed, re-enable the UI in `start.sh`, redeploy, and use the master key from `/app/data/.master_key` — then disable it again. --- @@ -165,21 +167,24 @@ expect a 15–20 min maintenance window and **never panic-restart mid-migration* | Variable | Required | Purpose | |----------|----------|---------| -| `TINFOIL_API_KEY` | Yes | Tinfoil API key for TEE inference | +| `TINFOIL_API_KEY` | Yes | Tinfoil provider (TEE inference) | +| `GREENPT_API_KEY` | Yes | GreenPT provider (renewable energy) | +| `PUBLICAI_API_KEY` | Yes | PublicAI provider (public models) | +| `LITELLM_CORS_ORIGINS` | Set by start.sh | CORS allowlist (chat origin only) | +| `LITELLM_MIGRATION_DIR` | Set by start.sh | Persistent Prisma migrations ledger | -## Connecting Open WebUI +## Connecting chat clients -In Open WebUI settings: +Open WebUI does **not** talk to this gateway directly — it routes through the [member portal](https://git.inference.coop/code/member-portal), which authenticates the member and injects their personal LiteLLM key per request. Members building their own tools create API keys on the [member dashboard](https://dashboard.inference.coop) and use: -- **API Base URL:** `https://gateway.inference.coop/v1` -- **API Key:** a virtual key created in LiteLLM's Admin UI - -## OIDC / SSO (future) - -The current package uses LiteLLM's built-in admin/master-key auth, which is sufficient for the MVP. Full OIDC SSO (admin UI behind Cloudron login) is a TODO — add the `"oidc": {}` addon to the manifest and configure LiteLLM's SSO hooks per https://docs.litellm.ai/docs/proxy/custom_sso. +- **API Base URL:** `https://gateway.inference.coop/v1` (or the portal's proxy endpoint for chat-integrated tools) +- **API Key:** their own virtual key from the dashboard ## Related repos -- `co-op/org-dev` — the full inference.coop design doc and deploy scripts +- `code/member-portal` — membership middleware: provisioning, billing webhooks, key injection +- `code/member-dashboard` — member-facing usage and API-key management +- `code/admin-panel` — admin overview - `co-op/website` — the public landing page +- `co-op/docs` — documentation, including the member-facing model list - `co-op/design-assets` — logos, fonts, brand materials