197 lines
14 KiB
Markdown
197 lines
14 KiB
Markdown
# Implementation Plan: API Access, Member Dashboard & Admin Dashboard
|
||
|
||
**Status:** Planning (not yet implemented)
|
||
**Last updated:** September 13, 2026
|
||
|
||
## Goal
|
||
|
||
Give members direct API access to the cooperative's models (alongside chat), a dashboard to view their usage and manage their own API keys, and an admin dashboard for co-op-wide monitoring.
|
||
|
||
## Guiding principles
|
||
|
||
1. **One unified budget.** Chat and API draw from the same $15/month credit pool. No separate accounting.
|
||
2. **The member surface holds no master key.** A member-facing app must never touch the LiteLLM master key, the Cloudron token, or the Open Collective token. Key creation happens through a tightly-scoped broker.
|
||
3. **Leverage LiteLLM's native primitives.** Teams for budgets/spend, spend-log endpoints for usage. Don't reinvent what the gateway already does.
|
||
4. **No commercial dependencies for now.** Stay on the open-source LiteLLM build.
|
||
|
||
## Verified facts (from research)
|
||
|
||
- **LiteLLM Teams** (open-source) provide: per-team budget, per-team spend tracking, and key attribution to both user and team. This is the abstraction layer for "unified account + multiple keys."
|
||
- **Clarification on "teams":** we use LiteLLM's open-source **team budget** feature (a budget container, referenced as `team_id`). We do **not** use the Premium **`team_admin` role / team-member permissions** — those are the paid features that would let members self-manage keys. We act as the "team admin" ourselves via the broker/master key instead.
|
||
- **Member-facing vocabulary:** members see "your account" / "your budget" / "your API keys" — never "team" (that's an internal implementation term).
|
||
- **Spend endpoints** are live on our gateway: `/global/spend`, `/spend/logs`, `/user/info`, `/team/list` (all return 200).
|
||
- **LiteLLM is v1.74.0, with Postgres** — so budgets and teams are fully supported.
|
||
- **Key self-serve requires Premium.** The `team_admin` role and "team member can create keys" permission are both **✨ Premium Features**. On the OSS build, only the master key can mint keys. Hence the broker.
|
||
- **Team members can still *view* key info on OSS** — read-only usage/key visibility is viable without Premium.
|
||
|
||
## Target architecture
|
||
|
||
```
|
||
Member (browser) ──SSO──▶ Member Dashboard (new Cloudron app, "members" group)
|
||
│ read-only: my spend, my keys, my usage
|
||
│ writes: "create/revoke key" → Broker
|
||
▼
|
||
Broker (portal or admin app; holds master key)
|
||
│ mints/revokes a key scoped to the member's team
|
||
▼
|
||
LiteLLM Gateway (Teams + budgets + spend logs)
|
||
```
|
||
|
||
- **Chat key:** portal-held, per member, under their team, hidden from the member.
|
||
- **API keys:** member-managed, under their team, visible in the member dashboard.
|
||
- Both draw from the same team's $15/30d budget.
|
||
|
||
## Phases
|
||
|
||
### Phase 1 — Teams migration (foundation)
|
||
|
||
- Create a **Team per member** (`team_alias: member:<email>`) with `$15/30d` budget and the three models.
|
||
- Two keys per team:
|
||
- **Chat key** — created/handled by the portal injector (existing `member:<email>` key becomes this).
|
||
- **API key(s)** — created via the broker on member request.
|
||
- Update `provision_member()` to create the team + chat key (instead of a bare key).
|
||
- Migrate existing members (Nathan, Liz, Dan, Ed, Lee).
|
||
- Verify: chat still works end-to-end; spend now attributed per-team.
|
||
|
||
### Phase 2 — Broker (key management)
|
||
|
||
- A token-protected endpoint that, given a member's email + a requested key name:
|
||
- Verifies the caller (SSO identity or portal JWT).
|
||
- Mints a key scoped to that member's team (via master key).
|
||
- Revokes a key on request.
|
||
- Lives in the **portal** (which already holds the master key) or the **admin app**.
|
||
- This is the only place the master key is used for key creation.
|
||
|
||
### Phase 3 — Member dashboard (new Cloudron app)
|
||
|
||
- Installed in the Cloudron **"members"** group → SSO-gated (same pattern as OpenWebUI/Loomio). No custom auth to build; trust the forwarded user identity.
|
||
- **Read-only views** (via team-scoped endpoints, no master key):
|
||
- My current spend vs. $15.
|
||
- My API keys (list, with the chat key hidden).
|
||
- Recent usage (from `/spend/logs` scoped to my key/team).
|
||
- **Write actions** (call the broker):
|
||
- Create API key.
|
||
- Revoke API key.
|
||
- Holds **no** master key, Cloudron token, or OC token.
|
||
|
||
### Phase 4 — Admin dashboard
|
||
|
||
- Co-op-wide view (admin-gated, holds master key):
|
||
- All members + status (active/inactive/founder).
|
||
- Aggregate spend + per-model breakdown.
|
||
- Budget adjustment, key/team management.
|
||
- Either a section of the portal or a distinct app. Decision deferred until Phase 3 is built (the admin view is lower urgency — LiteLLM's built-in UI already covers some of it).
|
||
|
||
### Phase 5 — Documentation
|
||
|
||
- Document the API endpoint (`https://gateway.inference.coop/v1`), OpenAI-compatible format, and "shares your $15 budget with chat."
|
||
- Document the member dashboard (how to get/rotate keys).
|
||
- Add an open question to `open-questions.md`: how to split/communicate the chat-vs-API budget if members want separate limits (future governance decision).
|
||
|
||
## Security model (the load-bearing part)
|
||
|
||
| Component | Holds master key? | Notes |
|
||
|-----------|-------------------|-------|
|
||
| Portal (control plane) | Yes | Existing: provisioning, key injection, webhook, admin |
|
||
| Broker | Yes | Scoped: only mints/revokes team keys |
|
||
| Member dashboard | **No** | SSO-gated, team-scoped, read-only + broker calls |
|
||
| Admin dashboard | Yes | Admin-only |
|
||
|
||
The invariant: **a member's session can only ever (a) read their own team's data, or (b) ask the broker to act on their own team.** No member-path code can enumerate or touch another member's keys.
|
||
|
||
## Open decisions (to resolve before/at build time)
|
||
|
||
1. **Broker location** — portal (simplest, already has master key) vs. admin app. Leaning portal.
|
||
2. **Member dashboard framework** — a small custom app (FastAPI, reusing the portal's patterns) vs. an app-store tool. Leaning custom: the app is a thin read/write wrapper, not a BI dashboard, and existing tools (Metabase/Superset/Baserow) don't fit "create my key."
|
||
3. **Key naming/metadata** — what a member sees when they create a key (name, purpose tag) for their own cost tracking.
|
||
4. **Premium later?** If the co-op grows an enterprise offering, LiteLLM Premium unlocks true key self-serve (drop the broker). Deferred.
|
||
|
||
## Risks
|
||
|
||
- **Broker is a new attack surface** — must be narrowly scoped and rate-limited.
|
||
- **Migration touches existing members** — needs careful idempotency (reuse existing keys by alias).
|
||
- **Spend-log scoping** — must verify `/spend/logs` can be filtered per-key/per-team so the member dashboard can't see others' usage.
|
||
|
||
## Stepwise development process
|
||
|
||
Build order with dependencies and verification gates. Each step must be verified in isolation before the next begins — no batching that obscures which step broke.
|
||
|
||
### Step 1 — Confirm spend-log scoping (the riskiest unknown)
|
||
|
||
**COMPLETED (2026-09-13). Outcome: the member dashboard cannot read LiteLLM directly — all reads AND writes must go through the broker/portal.**
|
||
|
||
Findings:
|
||
|
||
- **Member keys cannot read spend endpoints.** A member's own key (both `sk-` virtual keys and legacy raw-hash tokens) gets `401 "Only proxy admin"` on `/spend/logs`, `/spend/keys`, and `/user/info`. Only the master key can query spend.
|
||
- **Consequence:** the member dashboard must route *reads* through the portal (which holds the master key) too, not just writes. This actually *simplifies* the security model — the member dashboard holds **no LiteLLM credential at all**; it just calls the portal, which already authenticates members and already holds the master key.
|
||
- **`/spend/keys` is the clean per-member read primitive** — it returns one row per key with `key_alias` (`member:<email>`), `spend`, and `team_id`. Filtering by alias gives a member's spend directly.
|
||
- **`/spend/logs` filtering is inconsistent:** `api_key=<hash>` filters correctly, but `team_id=<x>` is silently ignored (returned all 92 rows for a bogus team id). Do not rely on `team_id` filtering on `/spend/logs`; prefer `/spend/keys` + `key_alias`.
|
||
- **Data inconsistency found (must fix in Step 2):** three members (Dan, Ed, Nathan) have **raw-hash key tokens** (not valid LiteLLM virtual keys), while four (Lee, Joseph, Liz, Roz) have `sk-` keys. The raw-hash tokens are a legacy artifact; the Teams migration will recreate all keys consistently as `sk-` virtual keys.
|
||
- **`/user/info?user_id=`** works with the master key but shows `spend: 0` and empty `keys`/`teams` because current keys have `user_id: null` (key-only, not user-bound). After the Teams migration binds keys to teams, per-team spend becomes queryable.
|
||
|
||
Revised architecture (all LiteLLM access via the portal):
|
||
|
||
```
|
||
Member (browser) ──SSO──▶ Member Dashboard (new Cloudron app, "members" group)
|
||
│ holds NO LiteLLM/Cloudron/OC credentials
|
||
│ calls the portal for everything
|
||
▼
|
||
Portal (holds master key; authenticates member)
|
||
│ read: /spend/keys?key_alias=member:<email>
|
||
│ write: mint/revoke key via master key
|
||
▼
|
||
LiteLLM Gateway (Teams + budgets + spend logs)
|
||
```
|
||
|
||
### Step 2 — Teams migration (Phase 1)
|
||
|
||
**COMPLETED (2026-09-13).** All 7 members now have a LiteLLM team + a fresh `sk-` key under it.
|
||
|
||
- Added `litellm_get_or_create_team()` (idempotent team lookup by alias — `/team/new` is *not* idempotent, so we check `/team/list` first) and rewrote `litellm_create_key()` to always generate a fresh `sk-` key (never "reuse").
|
||
- **Critical bug fixed:** the old `litellm_find_key_by_alias` returned SHA256 hashes from `/key/list` and stored them as the member's key. LiteLLM only reveals the `sk-` token at generation time, so three members (Dan, Ed, Nathan) had unusable hash "keys" → their chat was returning 401. All rekeyed; verified Dan can now complete inference.
|
||
- Added a one-off `/admin/migrate-teams/{token}` endpoint + `rekey_member()` (preserves `cloudron_user_id`/`slug`). Ran it: all 7 members rekeyed.
|
||
- Result: `/team/list` shows 7 teams (one per member, `$15/30d`), and `/spend/keys` now reports `team_id` on every key — spend is attributable per-member/team.
|
||
|
||
### Step 3 — Broker (Phase 2)
|
||
|
||
**COMPLETED (2026-09-13).** The portal now serves as the broker for member-managed API keys.
|
||
|
||
- Added `member_api_keys` table (email + name → sk_token + hash) and three endpoints under `/broker/keys` (GET list, POST create, DELETE revoke).
|
||
- Authentication: the member dashboard passes `X-Broker-Secret` (shared secret, `BROKER_SECRET` env) + `X-Member-Email`. The portal verifies the secret and that the email is an active member. Fail-closed: missing/wrong secret → 401, non-member → 403.
|
||
- Each API key is a LiteLLM virtual key under the member's team (`alias: api:<email>:<name>`), so it draws from the same $15 team budget as chat. The `sk-` token is returned once and stored; the hash is kept for revoke.
|
||
- Verified: create/list/revoke all work; the API key completes real inference; security boundaries hold (wrong secret 401, non-member 403, no-secret 401).
|
||
|
||
### Step 4 — Member dashboard (Phase 3)
|
||
|
||
**COMPLETED (2026-09-14).** New Cloudron app `code/member-dashboard`, deployed at `dashboard.inference.coop`.
|
||
|
||
- **App**: FastAPI, `proxyAuth` SSO (identity via `X-Forwarded-User` header), holds NO privileged credentials.
|
||
- **Access**: restricted to the `members` group via `POST /apps/:id/configure/access_restriction` (same model as Loomio).
|
||
- **Calls the portal broker** (`GET/POST/DELETE /broker/*`) with `BROKER_SECRET`, scoped to the logged-in member's email.
|
||
- **UI**: brand-matched Monitor surface — balance vs. spend bar, API key list (create/revoke), self-contained (no external requests).
|
||
- **Verified**: `/api/usage` and create→list→revoke all work end-to-end against Dan's real team (balance $15, spend $0).
|
||
|
||
**Note:** proxyAuth header name (`X-Forwarded-User`) confirmed via Cloudron forum + source; verified empirically by exec'ing into the app. The app also accepts `X-Remote-User`/`X-Auth-Request-Email` fallbacks.
|
||
|
||
### Step 5 — Admin dashboard (Phase 4)
|
||
|
||
**COMPLETED (2026-09-14).** New Cloudron app `code/admin-panel`, deployed at `panel.inference.coop`.
|
||
|
||
- **App**: FastAPI, `proxyAuth` SSO + an `ADMIN_USERNAMES` allowlist (so only `ntnsndr` / `inference-admin` reach it, not ordinary members).
|
||
- **Data**: calls the portal's `GET /admin/overview/{token}` (member list + balance + spend + reset date + co-op aggregates) and `POST /admin/set-balance/{token}` (top-up / adjust a member's balance).
|
||
- **UI**: co-op totals (members / active / total spend) + a member table with per-row balance, spend, remaining, reset date, and an inline "add balance" action.
|
||
- **Gate:** admin identity reaches the overview; a non-admin member gets 403.
|
||
|
||
### Step 6 — Documentation (Phase 5)
|
||
|
||
- README now has an **API access** section (endpoint, auth, example curl, same-balance note) and lists the Admin Panel in the components table.
|
||
|
||
### Rollback
|
||
|
||
- Every phase is additive; a failed phase can be reverted by re-pointing the injector at the prior key scheme (Step 2) or removing the new app (Steps 4–5). No data loss: keys are re-creatable, teams are idempotent by alias.
|
||
|
||
### Definition of done
|
||
|
||
- A member can: chat, retrieve an API key, make a direct `curl` call to `gateway.inference.coop/v1`, see their own usage, and create/revoke keys — all within one $15 budget, with no member path touching the master key.
|
||
- An admin can see co-op-wide membership, spend, and per-model breakdown.
|