Inference Cooperative

Private AI, governed together.

The Inference Cooperative is a member-governed project providing private AI inference. We are fiscally sponsored by Metagov, a nonprofit, and funded through our Open Collective.

The model

The Inference Cooperative is a member-governed AI inference utility. Members contribute on a sliding scale (currently $10–20/month) and, in return, get access to private AI inference through a shared, cooperatively governed infrastructure.

Because we are legally under a nonprofit—Metagov is our fiscal sponsor—we are member-governed rather than member-owned: members direct the project through governance, while the nonprofit holds the legal and financial structure.

Membership

  • Sliding scale dues: $10–20/month
  • Access: private AI inference through the cooperative's gateway
  • Governance: members participate in decisions about models, priorities, and direction

Important: your email address on Open Collective and on Cloudron must match. Membership is verified by matching the email you use to contribute on Open Collective against the email you use to log in. If they differ, you won't be recognized as a member. (Guest contributors — those who contribute without an Open Collective account — are recognized by the email they entered at checkout.)

Access control: the chat app (Open WebUI) and the governance app (Loomio) are restricted to Cloudron's members group. Provisioned members are added to this group; lapsed members are moved to the inactive group (which has no app access). Both apps must be listed in the members group's app list — if either is missing, members can't reach it.

Founder membership (free, invite-only)

There is a free "Co-founder" membership for early adopters — people who have been involved in earlier work on the project. It is invite-only: we reach out to trusted contributors directly rather than accepting requests from the general public.

Founder members are onboarded manually (Open Collective cannot process a $0 recurring subscription — recurring contributions require an automatic payment method, and a $0 tier has none):

  1. We invite the founder and collect their name and email.
  2. Their email is added to the portal's MANUAL_MEMBERS allowlist (a comma-separated env var). This exempts them from the daily sweep's deactivation, since they have no Open Collective membership to lapse.
  3. They are provisioned via the portal's /admin/provision endpoint (Cloudron user + LiteLLM key + welcome email + Loomio sync).

The stack

The Inference Cooperative runs a self-hosted stack on Cloudron, with all inference routed through a single gateway to cloud-based LLM providers.

Components

Component Purpose URL
Open WebUI Member-facing chat interface (with web search) chat.inference.coop
LiteLLM AI gateway — keys, metering, model routing gateway.inference.coop
Member Portal Membership middleware — onboarding, key provisioning, Loomio sync portal.inference.coop
Member Dashboard Member-facing usage view + API key management dashboard.inference.coop
Admin Panel Admin-only member overview, spend, and balance management panel.inference.coop
SearXNG Self-hosted web search (feeds the chat's search tool) search.inference.coop
Gitea Git hosting (code + docs) git.inference.coop
Loomio Member governance forum.inference.coop
Open Collective Membership billing + fiscal sponsorship opencollective.com/inference-cooperative
Listmonk Member newsletter (list auto-synced from membership) newsletter.inference.coop

The cooperative's software is open source and lives in the code organization on our Gitea instance.

Newsletter

Member newsletters are sent via Listmonk (self-hosted, at newsletter.inference.coop).

  • Single source of truth: the portal database. Listmonk is not managed by hand.
  • Auto-sync: the portal subscribes a member to the "Members" list on provisioning, and unsubscribes them on deactivation (both in the webhook path and the nightly sweep). It's best-effort — a newsletter outage can't break onboarding.
  • Opt-in model: single opt-in. Members consent by joining, so they're auto-confirmed rather than double-opted-in (no second confirmation email at signup).
  • Sender: info@inference.coop via Cloudron's mail addon, inheriting the domain's SPF/DKIM/DMARC alignment.
  • Machine credential: the portal authenticates to listmonk using a dedicated API user (portal-sync) with a token stored in the portal's Cloudron env vars (LISTMONK_API_TOKEN), not a password. Human admins log into listmonk via Cloudron SSO; password login is disabled.

Models

The gateway currently exposes three models (all served through Tinfoil's TEE-protected enclaves):

  • DeepSeek V4.1 Flash — default model, best for agentic tasks (1M context, tool calling)
  • GPT-OSS 120B — lightweight fallback
  • GLM-5.3 Flash — fast, efficient MoE model

Features

  • Web search — enabled by default, backed by a self-hosted SearXNG instance (no external search API or scraper key required).
  • File uploads (RAG) — works out of the box via Open WebUI's bundled local embedding model (sentence-transformers/all-MiniLM-L6-v2), no external embedding API.

API access

Members can use the co-op's models programmatically through an OpenAI-compatible API:

  • Endpoint (base URL): https://gateway.inference.coop/v1
  • Auth: a bearer API key (create one in the Member Dashboard)
  • Format: OpenAI chat completions — POST https://gateway.inference.coop/v1/chat/completions
  • Models: deepseek-v4-1-flash (default), gpt-oss-120b, glm-5-3-flash

API usage draws from the same monthly balance as chat — there is no separate quota. The gateway is publicly reachable (member keys authenticate it), but its admin UI and API documentation are disabled; admin access is via the internal dashboard or the gateway's admin endpoints.

Plain curl

curl https://gateway.inference.coop/v1/chat/completions \
  -H "Authorization: Bearer sk-your-key" \
  -H "Content-Type: application/json" \
  -d '{"model": "deepseek-v4-1-flash", "messages": [{"role": "user", "content": "Hello"}]}'

OpenAI SDK (Python, Node, etc.)

Point the SDK at our base URL and use the member key:

from openai import OpenAI

client = OpenAI(
    base_url="https://gateway.inference.coop/v1",
    api_key="sk-your-key",
)
resp = client.chat.completions.create(
    model="deepseek-v4-1-flash",
    messages=[{"role": "user", "content": "Hello"}],
)

Any OpenAI-compatible client works the same way: set the base URL to https://gateway.inference.coop/v1 and the API key to your member key.

Coding agents and other OpenAI-compatible clients

Any client that speaks the OpenAI API works the same way. Where a tool asks for a provider/endpoint configuration:

  • Base URL / endpoint: https://gateway.inference.coop/v1
  • API type / protocol: "OpenAI compatible" (or openai-completions where a protocol must be named) — not "Ollama", "Anthropic", or "Responses"
  • API key: your member key (the sk-… string)
  • Model: one of deepseek-v4-1-flash, gpt-oss-120b, glm-5-3-flash

A typical JSON provider entry looks like:

{
  "providers": {
    "inference-coop": {
      "baseUrl": "https://gateway.inference.coop/v1",
      "api": "openai-completions",
      "apiKey": "sk-your-key",
      "models": [
        { "id": "deepseek-v4-1-flash" },
        { "id": "gpt-oss-120b" },
        { "id": "glm-5-3-flash" }
      ]
    }
  }
}

If a tool's own docs show a local-server example (e.g. baseUrl: http://localhost:11434/v1 with api: "openai-completions"), substitute our base URL and your member key for apiKey — the protocol value stays the same. Some tools append /chat/completions to the base URL themselves; if your base URL already ends in /v1, that's correct and you should not add /chat/completions.

Privacy

Privacy is a core value. Inference now runs through Tinfoil, which provides architectural privacy: models run inside hardware enclaves (TEEs), and request/response bodies are encrypted end-to-end with the Encrypted HTTP Body Protocol (EHBP), so even Tinfoil's own infrastructure cannot read them. This is verifiable via remote attestation — not just a policy promise.

Because Tinfoil requires EHBP-encrypted request bodies (it rejects plaintext with 426 EHBP_REQUIRED), the LiteLLM gateway routes through a local Tinfoil proxy sidecar, which verifies the enclave attestation and handles the encryption. LiteLLM talks plaintext OpenAI to the proxy; the proxy encrypts and forwards to the enclave.

Security posture

A plain-language summary of how the co-op's infrastructure is secured, and the trade-offs we've deliberately accepted.

  • Public-but-token-gated control plane. The member portal (portal.inference.coop) is publicly reachable but every sensitive surface is gated: the admin endpoints require an X-Admin-Token header, the broker requires X-Broker-Secret, and the Open Collective webhook requires a secret token in its URL (Open Collective's webhooks aren't HMAC-signed, so a secret URL is the standard mechanism). All endpoints fail closed — unauthenticated requests get 401. The portal stays public because the webhook is an external caller that can't go through Cloudron SSO.
  • Secret separation. Chat, broker, admin, and model-layer secrets are independent and independently rotatable, so a leak in one has bounded scope.
  • Member isolation. A member's API key is scoped to their own LiteLLM team and budget; the broker can only ever operate on the authenticated member's own keys.
  • Rate limiting. Control-plane endpoints (admin, broker, webhook, the model list) are rate-limited per IP. The member inference path is not rate-limited — it's capped by each member's budget instead, so legitimate long generations aren't throttled.
  • No secrets in version control. All credentials live in Cloudron environment variables, never in the repos.
  • Accepted trade-off — keys at rest. Member chat keys and API keys are stored in plaintext in the portal's database. This is functionally necessary (the injector needs the chat key per request; the broker must display a freshly-created API key once), and Cloudron app volumes aren't encrypted at rest. Access to that volume is limited to the portal app itself.
  • Accepted trade-off — gateway surface. The LiteLLM gateway must stay publicly reachable for member API access. Its admin UI and docs (/docs, /redoc, Swagger) are disabled; only the /v1/* API is exposed, which requires a bearer key.

Governance

Members govern the project through Loomio and the Open Collective. See the Charter for the cooperative's values, membership terms, and governance structure, and Open Questions for the decisions currently open for member discussion.

During the pilot phase, Nathan Schneider serves as Managing Director and has sole final discretion on decisions.

  • Terms of Service — the agreement between the cooperative and its members about the service.
  • Privacy Policy — how we handle your data, including what we can and can't see.

By creating your account, you agree to these terms.

Planning


This documentation lives in a public git repository. Contributions welcome.

S
Description
Public documentation for the Inference Cooperative
Readme
321 KiB
0 Stars 2 Watchers 0 Forks
Languages
Markdown 100%