Files
docs/infrastructure.md
T

159 lines
8.4 KiB
Markdown

# Infrastructure
How the Inference Cooperative is built, for members who want to understand it —
and contributors who want to work on it.
## The big picture
Members reach the co-op through two front doors: the **chat** (Open WebUI) and
the **API** (the LiteLLM gateway). Both route through a single gateway to cloud
LLM providers running inside **TEE-protected enclaves** (Tinfoil), so inference
is private end-to-end. A thin layer of custom middleware — the member portal —
ties Open Collective membership to Cloudron accounts and LiteLLM keys.
```
Members ──> Open WebUI (chat) ─┐
├──> LiteLLM gateway ──> Tinfoil (TEE) ──> models
Members ──> your own tools ────┘
(API)
```
The whole stack runs on [Cloudron](https://cloudron.io), which handles hosting,
single sign-on, and per-app isolation.
## Components
| Component | Purpose | URL | Code |
|-----------|---------|-----|------|
| **Open WebUI** | Member-facing chat (web search, file uploads) | chat.inference.coop | — |
| **LiteLLM** | AI gateway — keys, metering, model routing | gateway.inference.coop | [code/litellm](https://git.inference.coop/code/litellm) |
| **Member Portal** | Membership middleware — onboarding, key provisioning, Loomio sync | portal.inference.coop | [code/member-portal](https://git.inference.coop/code/member-portal) |
| **Member Dashboard** | Member-facing usage view + API key management | dashboard.inference.coop | [code/member-dashboard](https://git.inference.coop/code/member-dashboard) |
| **Admin Panel** | Admin-only member overview, spend, balance management | panel.inference.coop | [code/admin-panel](https://git.inference.coop/code/admin-panel) |
| **SearXNG** | Self-hosted web search (feeds the chat's search) | search.inference.coop | — |
| **Loomio** | Member governance | forum.inference.coop | — |
| **Gitea** | Git hosting (code + docs) | git.inference.coop | — |
| **Surfer** | Static hosting (landing page, thanks page) | inference.coop | [co-op/website](https://git.inference.coop/co-op/website) |
| **Listmonk** | Member newsletter (auto-synced from membership) | newsletter.inference.coop | — |
| **Vaultwarden** | Password manager (admin credentials) | vault.inference.coop | — |
| **SnappyMail** | Webmail | mail.inference.coop | — |
Open Collective (billing and fiscal sponsorship) is the one external service in
the flow — see [opencollective.com/inference-cooperative](https://opencollective.com/inference-cooperative).
## How membership is wired
Membership flows through the member portal, which is the single source of truth
for "who is a member":
1. A member contributes on Open Collective.
2. Open Collective fires a webhook to the portal, which provisions a Cloudron
account, a LiteLLM key, and a welcome email.
3. A nightly sweep reconciles against the live Open Collective list, so a missed
webhook self-heals within 24 hours.
The portal database is authoritative for member emails and balances; LiteLLM's
team budget is the source of truth for a member's allowance, shared between chat
and API.
## The code
The cooperative's software is open source, in the [`code`](https://git.inference.coop/code)
organization on Gitea:
- [**code/member-portal**](https://git.inference.coop/code/member-portal) — membership middleware (Python).
- [**code/member-dashboard**](https://git.inference.coop/code/member-dashboard) — member-facing usage + API keys (Python).
- [**code/admin-panel**](https://git.inference.coop/code/admin-panel) — admin overview (Python).
- [**code/litellm**](https://git.inference.coop/code/litellm) — the gateway packaged as a Cloudron app.
Plus the `co-op` organization:
- [**co-op/website**](https://git.inference.coop/co-op/website) — landing page and static site.
- [**co-op/docs**](https://git.inference.coop/co-op/docs) — this documentation.
- [**co-op/design-assets**](https://git.inference.coop/co-op/design-assets) — logos, fonts, brand materials.
## Contributing (the Git is yours, too)
Everything the co-op runs is open source and lives at
[git.inference.coop](https://git.inference.coop) — the website, this
documentation, the member portal, the dashboard, the admin panel, and the
gateway. As a member, you can browse all of it and propose changes.
**Who can sign in.** Git access is tied to your membership: you log into the Git
with the same single sign-on you use for the chat. Members are in automatically
(through the `members` group); when a membership lapses, Git access lapses with it.
**How to contribute.** Anyone can read the repos without an account. To make a
change, fork the repo, edit in your fork, and open a **pull request**:
1. Log into [git.inference.coop](https://git.inference.coop) with your co-op account.
2. **Fork** the repo you want to change (e.g. `co-op/docs`).
3. Make your edits in your fork and commit them.
4. Open a **pull request** back to the original repo, describing your change.
**All changes are reviewed.** Nothing merges directly into `main` — every change
comes in as a pull request that an admin reviews and approves first. You don't
need to be able to create repos or organizations to contribute; you just fork,
edit, and ask for review.
Questions about the Git itself? Write to info@inference.coop.
## Models and providers
The co-op serves models from multiple providers, named `provider/model-name` so
members can choose both *which* model and *where* it runs. Each provider carries
a value you can route by:
| Provider | Value | How it works |
|---|---|---|
| **Tinfoil** | *private* | Models run inside hardware enclaves (TEEs); request/response bodies are encrypted end-to-end (EHBP), so even Tinfoil's infrastructure can't read them. Verifiable via remote attestation. |
| **GreenPT** | *green* | Models served on 100% renewable energy in the EU. Standard (non-enclave) hosting — cleaner energy, but privacy is a policy commitment, not architectural. |
| **PublicAI** | *public* | (coming soon — publicly developed models.) |
Five models are currently exposed:
- **DeepSeek V4.1 Flash** (`deepseek-v4-1-flash`) — default, agentic tasks. *(Tinfoil)*
- **GPT-OSS 120B** (`gpt-oss-120b`) — lightweight fallback. *(Tinfoil)*
- **GLM-5.3 Flash** (`glm-5-3-flash`) — fast, efficient. *(Tinfoil)*
- **Green-R** (`greenpt/green-r`) — reasoning, renewable energy. *(GreenPT)*
- **Green-L** (`greenpt/green-l`) — lightweight, renewable energy. *(GreenPT)*
The three Tinfoil models currently use their bare names (no `tinfoil/` prefix);
they'll be renamed to `tinfoil/…` in an upcoming, announced change.
A local Tinfoil proxy sidecar handles the enclave encryption on our side of the
gateway.
The honest caveat: only **Tinfoil** offers architectural privacy. The other
providers are chosen for their value (renewable energy, public models) but do
not run in enclaves — prompts and responses pass through them in the ordinary
way. Your **chat history** is also stored on our server so you can revisit it,
and that stored history is not encrypted in a way that prevents us from
technically reading it — we commit not to. The full distinction — what's
architecturally private versus what's a policy commitment — is in the
[Privacy Policy](privacy-policy.md).
## Security posture (plain language)
- **Token-gated control plane.** The portal is publicly reachable, but every
sensitive surface requires a secret token (admin, broker, webhook). All fail
closed — unauthenticated requests get `401`.
- **Separated secrets.** Chat, broker, admin, and model-layer secrets are
independent and independently rotatable, so a leak has bounded scope.
- **Member isolation.** Each API key is scoped to its owner's LiteLLM team and
budget; a member can only ever act on their own keys.
- **Rate limiting** on control-plane endpoints; the inference path is capped by
each member's budget instead.
- **No secrets in version control.** Credentials live in Cloudron env vars, never
in the repos.
- **Accepted trade-offs** (documented, not hidden): member keys are stored in
plaintext in the portal database (functionally necessary), and the gateway
must stay publicly reachable for API access (its admin UI and docs are
disabled; only the `/v1/*` API is exposed).
## For operators
Maintaining the co-op's infrastructure — rotating Cloudron tokens, editing the
portal's environment, and troubleshooting the provisioning pipeline — is
documented in the [operator runbook](operator-token-runbook.md).