# Infrastructure How the Inference Cooperative is built, for members who want to understand it — and contributors who want to work on it. ## The big picture Members reach the co-op through two front doors: the **chat** (Open WebUI) and the **API** (the LiteLLM gateway). Both route through a single gateway to cloud LLM providers running inside **TEE-protected enclaves** (Tinfoil), so inference is private end-to-end. A thin layer of custom middleware — the member portal — ties Open Collective membership to Cloudron accounts and LiteLLM keys. ``` Members ──> Open WebUI (chat) ─┐ ├──> LiteLLM gateway ──> Tinfoil (TEE) ──> models Members ──> your own tools ────┘ (API) ``` The whole stack runs on [Cloudron](https://cloudron.io), which handles hosting, single sign-on, and per-app isolation. ## Components | Component | Purpose | URL | Code | |-----------|---------|-----|------| | **Open WebUI** | Member-facing chat (web search, file uploads) | chat.inference.coop | — | | **LiteLLM** | AI gateway — keys, metering, model routing | gateway.inference.coop | [code/litellm](https://git.inference.coop/code/litellm) | | **Member Portal** | Membership middleware — onboarding, key provisioning, Loomio sync | portal.inference.coop | [code/member-portal](https://git.inference.coop/code/member-portal) | | **Member Dashboard** | Member-facing usage view + API key management | dashboard.inference.coop | [code/member-dashboard](https://git.inference.coop/code/member-dashboard) | | **Admin Panel** | Admin-only member overview, spend, balance management | panel.inference.coop | [code/admin-panel](https://git.inference.coop/code/admin-panel) | | **SearXNG** | Self-hosted web search (feeds the chat's search) | search.inference.coop | — | | **Loomio** | Member governance | forum.inference.coop | — | | **Gitea** | Git hosting (code + docs) | git.inference.coop | — | | **Surfer** | Static hosting (landing page, thanks page) | inference.coop | [co-op/website](https://git.inference.coop/co-op/website) | | **Listmonk** | Member newsletter (auto-synced from membership) | newsletter.inference.coop | — | | **Vaultwarden** | Password manager (admin credentials) | vault.inference.coop | — | | **SnappyMail** | Webmail | mail.inference.coop | — | Open Collective (billing and fiscal sponsorship) is the one external service in the flow — see [opencollective.com/inference-cooperative](https://opencollective.com/inference-cooperative). ## How membership is wired Membership flows through the member portal, which is the single source of truth for "who is a member": 1. A member contributes on Open Collective. 2. Open Collective fires a webhook to the portal, which provisions a Cloudron account, a LiteLLM key, and a welcome email. 3. A nightly sweep reconciles against the live Open Collective list, so a missed webhook self-heals within 24 hours. The portal database is authoritative for member emails and balances; LiteLLM's team budget is the source of truth for a member's allowance, shared between chat and API. ### The bot account on Open Collective You may notice **"Inference Co-op Bot"** listed as an admin of the collective on Open Collective. This is infrastructure, not a member or a person: - **What it is:** a dedicated Open Collective account owned by the co-op. It holds a **Personal Token** that the member portal uses to read membership data — the roster, tiers (needed to route credit-pack purchases correctly), and contributor emails (including guests, which only admins can see). - **What it can do:** read member and tier data via Open Collective's API. It cannot move money, issue refunds, or change the collective — Open Collective restricts those to dashboard actions by human admins. Its API access was verified read-mostly during the October 2026 security review. - **Why it has admin:** reading guest contributor emails requires admin on Open Collective; without it, membership reconciliation and credit-pack routing can't function. - **Who supervises it:** all of its code is open source in [code/member-portal](https://git.inference.coop/code/member-portal) — every API call it makes is documented there. Nathan Schneider (collective admin) oversees its use and can revoke or rotate its token at any time from the Open Collective dashboard. Nothing on Open Collective is charged, transferred, or refunded by automation: payments always come from a human's own contribution action, and any automated system only *reads* the resulting records. ## The code The cooperative's software is open source, in the [`code`](https://git.inference.coop/code) organization on Gitea: - [**code/member-portal**](https://git.inference.coop/code/member-portal) — membership middleware (Python). - [**code/member-dashboard**](https://git.inference.coop/code/member-dashboard) — member-facing usage + API keys (Python). - [**code/admin-panel**](https://git.inference.coop/code/admin-panel) — admin overview (Python). - [**code/litellm**](https://git.inference.coop/code/litellm) — the gateway packaged as a Cloudron app. Plus the `co-op` organization: - [**co-op/website**](https://git.inference.coop/co-op/website) — landing page and static site. - [**co-op/docs**](https://git.inference.coop/co-op/docs) — this documentation. - [**co-op/design-assets**](https://git.inference.coop/co-op/design-assets) — logos, fonts, brand materials. ## Contributing (the Git is yours, too) Everything the co-op runs is open source and lives at [git.inference.coop](https://git.inference.coop) — the website, this documentation, the member portal, the dashboard, the admin panel, and the gateway. As a member, you can browse all of it and propose changes. **Who can sign in.** Git access is tied to your membership: you log into the Git with the same single sign-on you use for the chat. Members are in automatically (through the `members` group); when a membership lapses, Git access lapses with it. **How to contribute.** Anyone can read the repos without an account. To make a change, fork the repo, edit in your fork, and open a **pull request**: 1. Log into [git.inference.coop](https://git.inference.coop) with your co-op account. 2. **Fork** the repo you want to change (e.g. `co-op/docs`). 3. Make your edits in your fork and commit them. 4. Open a **pull request** back to the original repo, describing your change. **All changes are reviewed.** Nothing merges directly into `main` — every change comes in as a pull request that an admin reviews and approves first. You don't need to be able to create repos or organizations to contribute; you just fork, edit, and ask for review. Questions about the Git itself? Write to info@inference.coop. ## Models and providers The co-op serves models from multiple providers (Tinfoil for privacy, GreenPT for renewable energy, PublicAI coming soon). See the [model list](models.md) for the current catalog — that's the single source of truth for which models are available and how they're named. One infra detail worth knowing: a local Tinfoil proxy sidecar handles the enclave encryption on our side of the gateway — only Tinfoil traffic goes through it; the other providers are called directly. ## Security posture (plain language) - **Token-gated control plane.** The portal is publicly reachable, but every sensitive surface requires a secret token (admin, broker, webhook). All fail closed — unauthenticated requests get `401`. - **Separated secrets.** Chat, broker, admin, and model-layer secrets are independent and independently rotatable, so a leak has bounded scope. - **Member isolation.** Each API key is scoped to its owner's LiteLLM team and budget; a member can only ever act on their own keys. - **Rate limiting** on control-plane endpoints; the inference path is capped by each member's budget instead. - **No secrets in version control.** Credentials live in Cloudron env vars, never in the repos. - **Accepted trade-offs** (documented, not hidden): member keys are stored in plaintext in the portal database (functionally necessary), and the gateway must stay publicly reachable for API access (its admin UI and docs are disabled; only the `/v1/*` API is exposed). ## For operators Maintaining the co-op's infrastructure — rotating Cloudron tokens, editing the portal's environment, and troubleshooting the provisioning pipeline — is documented in the [operator runbook](operator-token-runbook.md). ## Gateway version and security update policy The LiteLLM gateway is pinned to an **exact upstream tag** in [`code/litellm`](https://git.inference.coop/code/litellm) (currently `v1.84.0`). Two reasons: 1. **CVE fix line.** A member security review (Oct 2026) found the prior pin (`main-v1.74.0-stable`) sat below the fix line for two published vulnerabilities: CVE-2026-35029 (authorization bypass on `/config/update`, fixed in 1.83.0, with documented in-the-wild probing) and CVE-2026-59822 (MCP session auth bypass, fixed in 1.84.0, listed in the CISA Known Exploited Vulnerabilities catalog). `v1.84.0` is at or above both fix lines. 2. **Floating-tag drift.** `main-v1.74.0-stable` is a *floating* tag — the registry silently moved it to a newer build. A rebuild intended as a no-op therefore pulled different bits and triggered a full schema migration, which is how the 2026-10-02 outage happened (see below). Rule: **never pin a floating tag** (`main-*` / `latest`); always the exact version tag. When upgrading, expect the first boot after the upgrade to take 10–15 minutes (prisma migration with internal retries) — plan the restart window accordingly. CORS: the gateway answers browser preflights only for `https://chat.inference.coop` (via LiteLLM's native `LITELLM_CORS_ORIGINS`). The old `LITELLM_CORS_ALLOWED_ORIGINS` name (our v1.74.0-era sed-patch variable) is dead — v1.84.0 reads the native name only. ## Incident log **2026-10-02 — gateway outage after security-upgrade rebuild.** A Cloudron update triggered by the CVE report rebuilt the app image; the floating `main-v1.74.0-stable` tag had moved to newer bits (v1.84.0-era), so the rebuild changed the running version and kicked off a full prisma DB migration that takes 10–15 minutes per boot with internal timeouts/retries. The gateway was down ~50 minutes (04:50–05:34 UTC); no data loss. Repairs: version pinned to the exact tag in the repo (PR [code/litellm#1](https://git.inference.coop/code/litellm/pulls/1) — pending merge), `LITELLM_CORS_ORIGINS` set as a Cloudron env var (the upgrade had silently reverted CORS to `*`), packaging synced to mirror the live deployment. Also during the incident: the domain's authoritative nameservers (yoursrs.com) were intermittently unreachable, briefly failing resolution of some hostnames (git, portal) — transient registrar-side issue, self-recovered, no records were wrong.