Files
docs/infrastructure.md
T

8.4 KiB

Infrastructure

How the Inference Cooperative is built, for members who want to understand it — and contributors who want to work on it.

The big picture

Members reach the co-op through two front doors: the chat (Open WebUI) and the API (the LiteLLM gateway). Both route through a single gateway to cloud LLM providers running inside TEE-protected enclaves (Tinfoil), so inference is private end-to-end. A thin layer of custom middleware — the member portal — ties Open Collective membership to Cloudron accounts and LiteLLM keys.

Members ──> Open WebUI (chat) ─┐
                               ├──> LiteLLM gateway ──> Tinfoil (TEE) ──> models
Members ──> your own tools ────┘
                          (API)

The whole stack runs on Cloudron, which handles hosting, single sign-on, and per-app isolation.

Components

Component Purpose URL Code
Open WebUI Member-facing chat (web search, file uploads) chat.inference.coop —
LiteLLM AI gateway — keys, metering, model routing gateway.inference.coop code/litellm
Member Portal Membership middleware — onboarding, key provisioning, Loomio sync portal.inference.coop code/member-portal
Member Dashboard Member-facing usage view + API key management dashboard.inference.coop code/member-dashboard
Admin Panel Admin-only member overview, spend, balance management panel.inference.coop code/admin-panel
SearXNG Self-hosted web search (feeds the chat's search) search.inference.coop —
Loomio Member governance forum.inference.coop —
Gitea Git hosting (code + docs) git.inference.coop —
Surfer Static hosting (landing page, thanks page) inference.coop co-op/website
Listmonk Member newsletter (auto-synced from membership) newsletter.inference.coop —
Vaultwarden Password manager (admin credentials) vault.inference.coop —
SnappyMail Webmail mail.inference.coop —

Open Collective (billing and fiscal sponsorship) is the one external service in the flow — see opencollective.com/inference-cooperative.

How membership is wired

Membership flows through the member portal, which is the single source of truth for "who is a member":

  1. A member contributes on Open Collective.
  2. Open Collective fires a webhook to the portal, which provisions a Cloudron account, a LiteLLM key, and a welcome email.
  3. A nightly sweep reconciles against the live Open Collective list, so a missed webhook self-heals within 24 hours.

The portal database is authoritative for member emails and balances; LiteLLM's team budget is the source of truth for a member's allowance, shared between chat and API.

The code

The cooperative's software is open source, in the code organization on Gitea:

Plus the co-op organization:

Contributing (the Git is yours, too)

Everything the co-op runs is open source and lives at git.inference.coop — the website, this documentation, the member portal, the dashboard, the admin panel, and the gateway. As a member, you can browse all of it and propose changes.

Who can sign in. Git access is tied to your membership: you log into the Git with the same single sign-on you use for the chat. Members are in automatically (through the members group); when a membership lapses, Git access lapses with it.

How to contribute. Anyone can read the repos without an account. To make a change, fork the repo, edit in your fork, and open a pull request:

  1. Log into git.inference.coop with your co-op account.
  2. Fork the repo you want to change (e.g. co-op/docs).
  3. Make your edits in your fork and commit them.
  4. Open a pull request back to the original repo, describing your change.

All changes are reviewed. Nothing merges directly into main — every change comes in as a pull request that an admin reviews and approves first. You don't need to be able to create repos or organizations to contribute; you just fork, edit, and ask for review.

Questions about the Git itself? Write to info@inference.coop.

Models and providers

The co-op serves models from multiple providers, named provider/model-name so members can choose both which model and where it runs. Each provider carries a value you can route by:

Provider Value How it works
Tinfoil private Models run inside hardware enclaves (TEEs); request/response bodies are encrypted end-to-end (EHBP), so even Tinfoil's infrastructure can't read them. Verifiable via remote attestation.
GreenPT green Models served on 100% renewable energy in the EU. Standard (non-enclave) hosting — cleaner energy, but privacy is a policy commitment, not architectural.
PublicAI public (coming soon — publicly developed models.)

Five models are currently exposed:

  • DeepSeek V4.1 Flash (deepseek-v4-1-flash) — default, agentic tasks. (Tinfoil)
  • GPT-OSS 120B (gpt-oss-120b) — lightweight fallback. (Tinfoil)
  • GLM-5.3 Flash (glm-5-3-flash) — fast, efficient. (Tinfoil)
  • Green-R (greenpt/green-r) — reasoning, renewable energy. (GreenPT)
  • Green-L (greenpt/green-l) — lightweight, renewable energy. (GreenPT)

The three Tinfoil models currently use their bare names (no tinfoil/ prefix); they'll be renamed to tinfoil/… in an upcoming, announced change.

A local Tinfoil proxy sidecar handles the enclave encryption on our side of the gateway.

The honest caveat: only Tinfoil offers architectural privacy. The other providers are chosen for their value (renewable energy, public models) but do not run in enclaves — prompts and responses pass through them in the ordinary way. Your chat history is also stored on our server so you can revisit it, and that stored history is not encrypted in a way that prevents us from technically reading it — we commit not to. The full distinction — what's architecturally private versus what's a policy commitment — is in the Privacy Policy.

Security posture (plain language)

  • Token-gated control plane. The portal is publicly reachable, but every sensitive surface requires a secret token (admin, broker, webhook). All fail closed — unauthenticated requests get 401.
  • Separated secrets. Chat, broker, admin, and model-layer secrets are independent and independently rotatable, so a leak has bounded scope.
  • Member isolation. Each API key is scoped to its owner's LiteLLM team and budget; a member can only ever act on their own keys.
  • Rate limiting on control-plane endpoints; the inference path is capped by each member's budget instead.
  • No secrets in version control. Credentials live in Cloudron env vars, never in the repos.
  • Accepted trade-offs (documented, not hidden): member keys are stored in plaintext in the portal database (functionally necessary), and the gateway must stay publicly reachable for API access (its admin UI and docs are disabled; only the /v1/* API is exposed).

For operators

Maintaining the co-op's infrastructure — rotating Cloudron tokens, editing the portal's environment, and troubleshooting the provisioning pipeline — is documented in the operator runbook.