6.9 KiB
Infrastructure
How the Inference Cooperative is built, for members who want to understand it — and contributors who want to work on it.
The big picture
Members reach the co-op through two front doors: the chat (Open WebUI) and the API (the LiteLLM gateway). Both route through a single gateway to cloud LLM providers running inside TEE-protected enclaves (Tinfoil), so inference is private end-to-end. A thin layer of custom middleware — the member portal — ties Open Collective membership to Cloudron accounts and LiteLLM keys.
Members ──> Open WebUI (chat) ─┐
├──> LiteLLM gateway ──> Tinfoil (TEE) ──> models
Members ──> your own tools ────┘
(API)
The whole stack runs on Cloudron, which handles hosting, single sign-on, and per-app isolation.
Components
| Component | Purpose | URL | Code |
|---|---|---|---|
| Open WebUI | Member-facing chat (web search, file uploads) | chat.inference.coop | — |
| LiteLLM | AI gateway — keys, metering, model routing | gateway.inference.coop | code/litellm |
| Member Portal | Membership middleware — onboarding, key provisioning, Loomio sync | portal.inference.coop | code/member-portal |
| Member Dashboard | Member-facing usage view + API key management | dashboard.inference.coop | code/member-dashboard |
| Admin Panel | Admin-only member overview, spend, balance management | panel.inference.coop | code/admin-panel |
| SearXNG | Self-hosted web search (feeds the chat's search) | search.inference.coop | — |
| Loomio | Member governance | forum.inference.coop | — |
| Gitea | Git hosting (code + docs) | git.inference.coop | — |
| Surfer | Static hosting (landing page, thanks page) | inference.coop | co-op/website |
| Listmonk | Member newsletter (auto-synced from membership) | newsletter.inference.coop | — |
| Vaultwarden | Password manager (admin credentials) | vault.inference.coop | — |
| SnappyMail | Webmail | mail.inference.coop | — |
Open Collective (billing and fiscal sponsorship) is the one external service in the flow — see opencollective.com/inference-cooperative.
How membership is wired
Membership flows through the member portal, which is the single source of truth for "who is a member":
- A member contributes on Open Collective.
- Open Collective fires a webhook to the portal, which provisions a Cloudron account, a LiteLLM key, and a welcome email.
- A nightly sweep reconciles against the live Open Collective list, so a missed webhook self-heals within 24 hours.
The portal database is authoritative for member emails and balances; LiteLLM's team budget is the source of truth for a member's allowance, shared between chat and API.
The code
The cooperative's software is open source, in the code
organization on Gitea:
- code/member-portal — membership middleware (Python).
- code/member-dashboard — member-facing usage + API keys (Python).
- code/admin-panel — admin overview (Python).
- code/litellm — the gateway packaged as a Cloudron app.
Plus the co-op organization:
- co-op/website — landing page and static site.
- co-op/docs — this documentation.
- co-op/design-assets — logos, fonts, brand materials.
Contributing (the Git is yours, too)
Everything the co-op runs is open source and lives at git.inference.coop — the website, this documentation, the member portal, the dashboard, the admin panel, and the gateway. As a member, you can browse all of it and propose changes.
Who can sign in. Git access is tied to your membership: you log into the Git
with the same single sign-on you use for the chat. Members are in automatically
(through the members group); when a membership lapses, Git access lapses with it.
How to contribute. Anyone can read the repos without an account. To make a change, fork the repo, edit in your fork, and open a pull request:
- Log into git.inference.coop with your co-op account.
- Fork the repo you want to change (e.g.
co-op/docs). - Make your edits in your fork and commit them.
- Open a pull request back to the original repo, describing your change.
All changes are reviewed. Nothing merges directly into main — every change
comes in as a pull request that an admin reviews and approves first. You don't
need to be able to create repos or organizations to contribute; you just fork,
edit, and ask for review.
Questions about the Git itself? Write to info@inference.coop.
Models and providers
The co-op serves models from multiple providers (Tinfoil for privacy, GreenPT for renewable energy, PublicAI coming soon). See the model list for the current catalog — that's the single source of truth for which models are available and how they're named.
One infra detail worth knowing: a local Tinfoil proxy sidecar handles the enclave encryption on our side of the gateway — only Tinfoil traffic goes through it; the other providers are called directly.
Security posture (plain language)
- Token-gated control plane. The portal is publicly reachable, but every
sensitive surface requires a secret token (admin, broker, webhook). All fail
closed — unauthenticated requests get
401. - Separated secrets. Chat, broker, admin, and model-layer secrets are independent and independently rotatable, so a leak has bounded scope.
- Member isolation. Each API key is scoped to its owner's LiteLLM team and budget; a member can only ever act on their own keys.
- Rate limiting on control-plane endpoints; the inference path is capped by each member's budget instead.
- No secrets in version control. Credentials live in Cloudron env vars, never in the repos.
- Accepted trade-offs (documented, not hidden): member keys are stored in
plaintext in the portal database (functionally necessary), and the gateway
must stay publicly reachable for API access (its admin UI and docs are
disabled; only the
/v1/*API is exposed).
For operators
Maintaining the co-op's infrastructure — rotating Cloudron tokens, editing the portal's environment, and troubleshooting the provisioning pipeline — is documented in the operator runbook.