Files
docs/mvp-planning.md
T

55 KiB
Raw Blame History

type, function, region, maturity, ownership
type function region maturity ownership
research-project inference global planned cooperative

inference.coop — Minimum Viable Cooperative Inference Utility

Status: Design/planning phase Category: Cooperative AI infrastructure — proof of concept Goal: A working, membership-based inference cooperative providing API and chatbot access to open models, hosted on renewable energy, with unified account management and cooperative governance.

Design goals

  1. Leverage existing open-source software for the full stack — no custom infrastructure to build
  2. Renewable energy hosting — inference on providers powered by renewables
  3. Unified account management across services (chat, API, billing, governance)
  4. Membership-based payment through a service that supports recurring payments and cooperative governance
  5. Small scale first — viable for a handful of user-testers, then scale

Component 1: Inference engine

The backend that loads and serves open models via an OpenAI-compatible API.

Options

Engine Throughput Key features Best for
vLLM Highest (PagedAttention, continuous batching) OpenAI-compatible API, multi-LoRA, production-grade multi-tenant Production serving, many concurrent users
Ollama Lower (single-user optimized) One-command setup, model library, extreme ease of use MVP/prototyping, small deployments
SGLang ~30% faster than vLLM on some workloads RadixAttention, structured outputs Latency-sensitive workloads
TGI (HuggingFace) Competitive HF ecosystem integration HF-centric workflows

Recommendation

Two paths to consider — and Tinfoil changes the calculus:

Path A: Self-hosted Ollama → vLLM (original plan)

  • Start with Ollama (Cloudron app), migrate to vLLM at scale
  • Full control over model selection, no per-token costs for self-hosted models
  • But: no TEE privacy at MVP, and vLLM migration required at scale

Path B: Tinfoil Inference API from day one (new option)

  • OpenAI-compatible API — drop-in for Open WebUI and LobeChat
  • TEE-protected inference from the start — verifiable privacy, no pinky promises
  • No GPU hosting, no model management, no Cloudron GPU app needed
  • Pay-as-you-go per token (see pricing below)
  • Open-source, SOC 2 compliant, audited by Trail of Bits
  • Used by Duck.ai, UC Berkeley, Stanford, Red Hat

Tinfoil model lineup and pricing (as of Aug 2026):

Model Context Parameters Input $/M tok Output $/M tok
DeepSeek V4 Flash 1M 304B (MoE) $0.30 $0.70
GPT-OSS 120B 131k 117B (5.1B active) $0.15 $0.60
Gemma 4 31B 256k 31B $0.40 $1.00
Llama 3.3 70B 128k 70B $1.75 $2.75
GLM-5.2 384k 754B (40B active) $1.50 $5.25
Kimi K3 256k 2.8T (MoE) $4.00 $20.00
Nomic Embed Text 8k 137M $0.05 $0
Whisper Large V3 Turbo 30s audio 809M $0.01/request

Source: Tinfoil models page

Cost comparison for 10 members, ~20 queries/day:

  • Path A (self-hosted): €36-66/month fixed (VPS + Cloudron + OpenRouter for frontier)
  • Path B (Tinfoil): ~$15-30/month variable (at GPT-OSS 120B rates, ~500K tokens/day for 10 users) + €15/month Cloudron for chat frontend/governance/SSO. No GPU server needed.

The hybrid approach (recommended):

Use both — Tinfoil as the default inference backend, with self-hosted Ollama for lightweight models when cost or latency matters:

  1. Tinfoil API for all models — gives TEE privacy from day one, no GPU management
  2. Ollama (Cloudron) as an optional backend for small models (7B range) when members want zero per-token cost and are OK without TEE
  3. OpenRouter as fallback for models Tinfoil doesn't offer yet (e.g. Claude, GPT)

Open WebUI and LobeChat both support multiple OpenAI-compatible endpoints — members can choose which backend serves their request. The cooperative's middleware routes and meters usage across all three.

This means the MVP launches with verifiable privacy (TEE) from day one without any GPU hosting or TEE configuration — Tinfoil handles the hard part. The cooperative's infrastructure is just Cloudron (for chat, governance, SSO) + a thin middleware for billing/access control + Tinfoil's API.

When to add self-hosted inference: When per-token costs exceed the cost of renting GPU hardware (probably at 50+ active members), add self-hosted vLLM on H100 GPUs with Confidential Computing enabled for the most-used models. Use Tinfoil's open-source framework (GitHub) as reference for TEE deployment.

Model selection

Start with 2-3 models covering capability tiers:

  • Lightweight: Qwen 3.5 7B or Llama 4 Scout 13B — fast, cheap, handles 80% of queries
  • Mid-range: DeepSeek V4 Lite or Qwen 3.5 32B — research assistance, complex reasoning
  • Frontier (brokered): Access via OpenRouter API for models too large to self-host

Recommendation: use OpenRouter for frontier access

OpenRouter provides a single API to 300+ models from 60+ providers with pay-as-you-go pricing. For the MVP, self-host lightweight/mid-range models and broker frontier access through OpenRouter. This avoids the cost of renting H100 GPUs for frontier models while still offering full model coverage.


Component 2: Chat frontend

The user-facing web interface for chatting with models.

Options

Frontend Stars Key features Cloudron app?
Open WebUI 141K+ ChatGPT-style, multi-user, RAG, model switching, document upload, pipelines, Python extensibility ✅ Yes
LibreChat — ChatGPT-like with model switching, multi-user ✅ Yes
Lobster — Clean chat interface ❌

Recommendation: Open WebUI

Open WebUI is the leading self-hosted chat interface (141K+ GitHub stars). It's available as a one-click Cloudron app, supports multi-user accounts with role-based access, model switching, document upload for RAG, and connects to both Ollama and any OpenAI-compatible API (including vLLM and OpenRouter). It also has a Python pipeline system for extensibility.


Component 3: Hosting (renewable energy)

Options

Provider Renewable claim GPU types Location Notes
Scaleway 100% renewable (wind + hydro, GO certified) at DC5 (Paris) L4, L40S Paris, Amsterdam, Warsaw EU data residency, GDPR. No H100. Best for lightweight models.
Hetzner Certified data centers in Germany/Finland (mixed energy) GPU dedicated servers Germany, Finland Best value CPU/GPU servers, but not 100% renewable
IREN 100% renewable-powered data centers H100, H200, B200 Global (AU, US, CA) For GPU-heavy workloads
Lambda Labs Not renewable-specific H100, A100 US Best price/reliability for GPU rental
GreenPT (backend) 100% renewable (uses Scaleway infra) Not disclosed Netherlands Could white-label, but €17.50/seat/month is expensive

Recommendation

Three tiers of green hosting, depending on whether the MVP uses GPU:

Tier 1: CPU-only VPS (no GPU) — for Ollama small models or Cloudron-only (no self-hosted inference)

If using Tinfoil/Ollama Cloud for inference and Cloudron only for chat/governance/SSO, no GPU is needed. A basic VPS suffices:

Provider Renewable claim GPU? Price Location Notes
Hetzner Cloud 100% hydropower (Germany since 2008), hydro+wind (Finland since 2018). Founded HT Clean Energy GmbH for solar parks. EMAS/ISO 14001 certified. Cloud: no GPU. Dedicated: yes CPX31 (4 vCPU, 8GB RAM): ~€13/mo Germany, Finland Best value. 100% renewable.
Scaleway 100% renewable at DC5 (Paris): wind + hydro, GO-certified. PUE 1.16. Free + adiabatic cooling. Cloud: L4, L40S. No A100/H100 DEV1-L (4 vCPU, 8GB): ~€15/mo Paris, Amsterdam, Warsaw Strongest renewable claims in EU.

Tier 2: GPU server for self-hosted Ollama (lightweight models, 7B-13B)

For self-hosting small models on Ollama, an L4 or L40S GPU is sufficient:

Provider Renewable claim GPU Price Location Notes
Scaleway 100% renewable (DC5 Paris) L4 ~$0.79/hr on-demand, ~$289/mo reserved Paris Greenest GPU option in Europe
Scaleway 100% renewable (DC5 Paris) L40S (48GB) ~$1.68/hr on-demand, ~$614/mo reserved Paris Best for 13B models
Hetzner 100% hydropower (Germany/Finland) L4 €199/mo dedicated Germany, Finland Fixed monthly, no A100/H100
Hetzner 100% hydropower L40S €349/mo dedicated Germany, Finland Fixed monthly, good value
Hetzner 100% hydropower RTX 6000 Ada (96GB) €889/mo dedicated Germany, Finland Largest VRAM, can run 70B quantized

Tier 3: H100 for frontier models + future TEE migration

For H100/H200 with Confidential Computing (TEE), needed when serving 70B+ models or implementing Tier 3 privacy:

Provider Renewable claim GPU Price Location TEE capable?
Scaleway 100% renewable (DC5) H100 SXM ~$2.80/hr on-demand Paris Unknown
Lambda Labs Not renewable-specific H100 $2.21/hr reserved US H100 CC capable
IREN 100% renewable-powered DCs H100, H200, B200 Custom pricing Global (AU, US, CA) H100 CC capable

Key findings

  1. Hetzner is the greenest practical choice for the MVP. 100% hydropower since 2008 (Germany) and 2018 (Finland). Now building own solar parks via HT Clean Energy GmbH. EMAS/ISO 14001 certified. GPU servers available from €199/month (L4) or €349/month (L40S). No H100/H200 — tops out at RTX 6000 Ada (96GB VRAM, can run 70B quantized).

  2. Scaleway DC5 is the greenest GPU cloud. 100% renewable (wind + hydro, GO-certified), PUE 1.16, ultra-efficient cooling. Offers L4, L40S, and H100. On-demand pricing means you pay only when serving requests.

  3. Neither Hetzner nor Scaleway explicitly supports NVIDIA Confidential Computing (TEE). For the TEE migration (Phase 3 privacy), IREN or Lambda Labs with H100 CC mode would be needed. This is a future concern, not an MVP requirement.

  4. Ollama on CPU is viable for the MVP. For 5-10 user-testers using 7B models (Qwen 3.5 7B, Gemma 4 9B), Ollama runs on CPU — slowly (5-15 tokens/sec) but usable for a proof of concept. A Hetzner CPX31 (4 vCPU, 8GB RAM, ~€13/mo) can serve lightweight models without a GPU. When faster response is needed, add a GPU server or switch to Tinfoil/Ollama Cloud for inference.

Option A: Minimal (CPU only, no GPU)

  • Hetzner CPX31 (4 vCPU, 8GB RAM): €13/mo — runs Cloudron + Ollama (7B models on CPU, slow but functional)
  • Inference: Ollama self-hosted (lightweight models) + OpenRouter (frontier models)
  • Total infra: ~€28/mo (VPS + Cloudron)
  • Renewable: ✅ 100% hydropower

Option B: With GPU (recommended for better UX)

  • Hetzner CCX23 (dedicated vCPU, 8GB RAM) for Cloudron: ~€25/mo
  • Hetzner GEX44 (L4 GPU dedicated server): €199/mo
  • Inference: Ollama self-hosted on GPU (fast, 7B-13B models) + OpenRouter (frontier)
  • Total infra: ~€239/mo (VPS + Cloudron + GPU server)
  • Renewable: ✅ 100% hydropower
  • Or: Scaleway L4 on-demand (~$0.79/hr = ~$230/mo if running 24/7, but can shut down off-hours)

Option C: Cloud inference, no GPU hosting

  • Hetzner CPX31 for Cloudron only: €13/mo
  • Inference: Tinfoil API (TEE privacy, per-token) or Ollama Cloud ($25/seat/mo)
  • Total infra: ~€28/mo + per-token/seat costs
  • Renewable: ✅ VPS on hydropower; inference provider's energy mix unknown

Sustainability tracking

GreenPT's real-time per-conversation CO₂ display is a feature the cooperative should replicate. This requires:

  • Measuring GPU/CPU energy per inference request
  • Multiplying by the grid carbon intensity of the data center's location
  • Displaying in the chat interface (Open WebUI supports custom UI elements via pipelines)

For Hetzner Germany/Finland (100% renewable), the carbon intensity is near-zero. For Scaleway DC5 Paris (100% renewable GO-certified), also near-zero. This is a genuine competitive differentiator: the cooperative can credibly claim near-zero carbon AI inference.


Component 4: Unified account management (Cloudron)

What Cloudron provides

Cloudron is a self-hosting platform that runs on your VPS and provides:

  • One-click app installation from an app store (100+ apps)

  • Single Sign-On (SSO) — users log in once, access all apps

  • Centralized user management — LDAP-based directory, admin/user roles

  • Automatic SSL, DNS, backups, updates

  • Access control — restrict which users can access which apps

  • Cloudron

  • Cloudron docs: user management

  • Cloudron pricing: ~€15/month for the Cloudron license + your VPS cost

Key Cloudron apps for inference.coop

App Purpose In Cloudron store?
Open WebUI Chat frontend ✅
Ollama Model serving ✅
LibreChat Alternative chat frontend ✅
Gitea Git hosting (for skills/agents) ✅
Nextcloud File sharing, collaboration ✅
Loomio Democratic decision-making / governance ✅
Discourse Forum / community discussion ✅
Matrix/Element Chat / communication ✅

This means the entire cooperative platform — inference, chat, governance, communication, file sharing — can be deployed through Cloudron with unified SSO. Members log in once and access all services.

Alternatives to Cloudron

Platform Open source? SSO? Notes
Coolify ✅ Yes (self-hostable) Partial Open-source Cloudron alternative. 280+ one-click services. No built-in SSO/LDAP.
YunoHost ✅ Yes ✅ Yes Free, Debian-based. Simpler than Cloudron. Fewer apps. Already used on Nathan's server.
Cosmo Cloud ✅ Yes ✅ Yes Newer, Cosmos Server based.

Recommendation

Cloudron is the right choice for inference.coop. The security model is materially stronger than YunoHost for running untrusted or semi-trusted software like model inference engines.

Cloudron security model (per Cloudron security docs):

  • Docker app isolation: Each app runs in its own Docker container. One app cannot access another app's database or files.
  • Read-only rootfs: Apps cannot tamper with their own application code.
  • AppArmor profiles: System calls restricted, /proc and /sys filesystems blocked.
  • Non-root execution: All apps run as non-root user by default.
  • Per-app subdomains: Prevents XSS in one app from compromising others (unlike sub-path deployments).
  • Dropped capabilities: CAP_SYS_ADMIN and other dangerous capabilities are dropped.
  • Automatic SSL/TLS: All apps HTTPS-only, A+ SSL Labs rating, HSTS, wildcard certs to avoid certificate transparency leaks.
  • Rate limiting: Built-in protection against brute force on passwords, SSH, email, database.
  • Signed updates: Platform updates GPG-signed, keys maintained offline.
  • No remote access: No mechanism for cloudron.io to access your server.
  • Automatic OS security updates: Ubuntu unattended-upgrades enabled.

YunoHost security model — YunoHost does not provide Docker-based app isolation by default. Apps are installed directly on the system (not in containers), meaning a vulnerability in one app can potentially affect others. There is no AppArmor profiling, no read-only rootfs, no per-app sandboxing. This is acceptable for trusted, well-maintained apps on a personal server, but is a real concern for:

  • Running model inference engines that process untrusted user input
  • Running multiple services for multiple members (not just yourself)
  • A cooperative that has a duty of care to its members' data

Cost comparison:

  • Cloudron: €15/month license + VPS cost
  • YunoHost: free + VPS cost
  • The €15/month is the cost of containerized isolation, automatic security updates, and professional app packaging. For a cooperative handling member data and running inference workloads, this is a necessary expense, not a luxury.

Alternative considered: Coolify (open-source, self-hostable). Coolify deploys apps via Docker (better isolation than YunoHost) but lacks built-in SSO/LDAP user management and the polished app store. It's a middle ground but doesn't solve the unified account management problem as well as Cloudron.


Component 5: Payment and membership (Open Collective)

What Open Collective provides

Open Collective is a platform for transparent funding of open projects:

  • Recurring payments — monthly/yearly tiers

  • Fiscal sponsorship — no need to incorporate immediately; a fiscal host handles legal/banking

  • Transparent budget — all income and expenses visible

  • Tier system — different membership levels with different benefits

  • Expense submission — members can submit expenses for reimbursement

  • Open Collective

  • Open Collective: fiscal hosts FAQ

How it fits inference.coop

  1. Create an Open Collective for inference.coop
  2. Set up tiers with tiered budgets (see below)
  3. Connect to a fiscal host (e.g., Open Collective Foundation, or a cooperative-specific host)
  4. Members pay through Open Collective → funds held transparently
  5. Expenses (server costs, Tinfoil API costs) paid from the collective

Membership tiers and budgets

The cooperative uses tiered budgets — each membership level maps to a LiteLLM budget cap, so "excessive use" is rate-limited natively by LiteLLM (no custom billing logic). A free tier lets people try the service before committing.

Tier Price LiteLLM budget cap Purpose
Free / Trial €0 Small (e.g. €2/month of tokens) Quick try — enough to test the service, not enough for sustained use
Member €10/month Generous (e.g. €15/month of tokens) Standard membership — day-to-day use
Supporter €20+/month Higher (e.g. €30/month of tokens) Heavy users + those who want to support the co-op

The budget cap is just a number set on each member's LiteLLM virtual key. When a member hits their cap, LiteLLM returns a clean "budget exceeded" message — no surprise bills, no proration, no reconciliation logic. The exact cap values are a governance decision (set in Loomio), not a technical constraint.

Key design decision: defer usage-based billing. The MVP uses flat membership with tiered budgets. There is no per-token billing, no overage charges, no proration. If members later want metered billing, that's the point where the middleware grows in complexity — but it's not built speculatively.

Gaps and additional needs

Open Collective handles money in (membership payments) but not:

  • Access provisioning — automatically granting/denying API access based on payment status
  • Usage metering — tracking token usage per member (LiteLLM does this)
  • Rate limiting — enforcing per-tier limits (LiteLLM does this via budget caps)

The only custom code needed is a lightweight middleware (webhook handler) between Open Collective and LiteLLM. LiteLLM handles metering and rate-limiting natively; the middleware just creates/enables/disables keys with the right budget.

Alternatives

Platform Recurring payments Fiscal sponsorship Membership management Cooperative-friendly?
Open Collective ✅ ✅ Partial (tiers, not access control) ✅ Designed for collective governance
Patreon ✅ ❌ Partial ❌ Extractive platform
Stripe + custom ✅ ❌ Custom build Neutral
Memberstack ✅ ❌ ✅ Access control, member portals ❌ SaaS, not cooperative
Outseta ✅ ❌ ✅ Full member management ❌ SaaS

Recommendation

Open Collective for payments + fiscal sponsorship, LiteLLM for metering/rate-limiting, custom middleware for access provisioning. The middleware is a small webhook handler that:

  1. Receives Open Collective webhooks when payments are made/missed
  2. Creates/enables/disables the member's LiteLLM virtual key with the tier's budget cap
  3. On payment lapse, disables the key (LiteLLM then rejects requests)

This is maybe 200-500 lines of Python — the only custom code needed for the MVP. Rate-limiting excessive use is LiteLLM's native budget feature, not custom logic.


Component 6: Governance (missing from original design goals)

A cooperative needs governance infrastructure. Cloudron has Loomio in its app store — an open-source platform for democratic decision-making used by cooperatives and community organizations worldwide.

Recommendation

Install Loomio via Cloudron for member governance:

  • Proposals and voting
  • Discussion threads
  • Working groups
  • Transparent decision history

This gives members a space to decide on model selection, pricing, privacy policies, and platform direction — all within the same SSO ecosystem as the chat and API.


Component 7: API access gateway (LiteLLM)

The cooperative needs a layer between members and the inference backends (Tinfoil, Ollama, OpenRouter) that handles:

  • Per-user API keys — each member gets their own key
  • Usage metering — track token usage per user for billing
  • Rate limiting — enforce per-tier limits (free, member, supporter)
  • Model routing — route requests to the right backend (Tinfoil for TEE, Ollama for free, OpenRouter for frontier)
  • Cost tracking — track spend per user, per model, per key

Solution: LiteLLM Proxy

LiteLLM is an open-source AI gateway that does all of the above. It sits between Open WebUI / LobeChat and the inference backends:

Members → Open WebUI / LobeChat → LiteLLM Proxy → Tinfoil (TEE)
                                                   → Ollama (self-hosted, free)
                                                   → OpenRouter (frontier fallback)

LiteLLM provides:

  • Virtual API keys — per-user, per-team keys with individual budgets and rate limits
  • Spend tracking — per key, per user, per team, per model, across 140+ providers
  • Model access control — restrict which users can access which models
  • Budget enforcement — when a user hits their cap, requests stop automatically
  • Admin dashboard — usage analytics, spend tracking, key management
  • OpenAI-compatible — one endpoint for all backends; Open WebUI and LobeChat see a single API
  • Fallback routing — if Tinfoil is down, automatically fall back to OpenRouter
  • Self-hostable — runs as a Docker container, can be deployed via Cloudron or alongside it

Sources: LiteLLM, LiteLLM + Ollama gateway guide, LiteLLM proxy setup

How it fits the cooperative model

  1. Tinfoil API key is stored only in LiteLLM's config — members never see it. The cooperative holds one Tinfoil account, LiteLLM distributes access via virtual keys.

  2. Per-member virtual keys — each member gets an sk-coop-xxxxx key. LiteLLM tracks their usage and enforces their tier's limits. This is how the cooperative meters who uses what.

  3. Open WebUI integration — Open WebUI supports per-user API keys and has built-in token tracking (via the token tracking pipes feature). It can display per-conversation token counts and costs to users.

  4. Billing connection — LiteLLM's spend tracking API can be queried by the billing middleware (Component 5) to reconcile usage with Open Collective payments. If a member stops paying, the middleware disables their LiteLLM virtual key.

  5. Model governance — LiteLLM's model access control lets the cooperative decide which models each tier can access. Members vote on model selection in Loomio; admins update LiteLLM config to reflect decisions.

Gap: self-serve portal and monetization

LiteLLM is the best AI-native gateway (model routing, token metering, fallback, virtual keys), but it is weak on two capabilities that are load-bearing for a cooperative where members are also API customers:

  1. Self-serve portal — LiteLLM's admin UI is for admins creating keys. There is no clean way for a member to log in, see their own usage, and generate/rotate/revoke their own key.
  2. Monetization — LiteLLM tracks spend but does not bill. No payment, no invoice, no subscription lifecycle.

Why not add a separate API gateway (Zuplo, Tyk, Kong)?

Option Self-hostable? AI-native routing Self-serve portal Monetization
LiteLLM alone ✅ ✅ ❌ (admin UI only) ❌ (tracks, doesn't bill)
LiteLLM + Tyk/Kong ✅ ✅ (LiteLLM) ✅ ✅
LiteLLM + Zuplo ❌ (Zuplo is SaaS) ✅ ✅ ✅
LiteLLM + custom portal ✅ ✅ ✅ (build it) ⚠️ (via Open Collective)
  • Zuplo is out — hosted SaaS, conflicts with the self-hostable, no-external-dependencies design. A cooperative should not route member API keys through a third-party SaaS.
  • Tyk/Kong is a real option but adds a second gateway — self-hostable with proper developer portals and monetization, but general-purpose (not AI-native) and oriented toward commercial API products, not cooperative membership. Two systems to operate.
  • Custom portal is smaller than it sounds — LiteLLM already does the hard part (metering, virtual keys, budgets). The portal is a thin web app: login (Cloudron SSO), "generate/rotate/revoke my key" (calls LiteLLM's key API), "my usage" (reads LiteLLM spend data), "my membership" (Open Collective status). ~300-500 lines, and it's the same middleware already planned for the Open Collective webhook — one build, not two.

Monetization = Open Collective + LiteLLM spend data. Open Collective handles the actual money (recurring membership, fiscal sponsorship). LiteLLM's spend tracking gives per-member usage. The middleware reconciles the two: if a member's usage exceeds their tier, or their payment lapses, the middleware adjusts their LiteLLM budget or disables their key.

Decision: Keep LiteLLM as the single AI gateway. Build the self-serve portal + billing reconciliation as part of the middleware (Component 5). Revisit Tyk/Kong only if the cooperative grows to need a full developer portal (docs, tiered self-serve plans, usage dashboards) — at which point the operational cost of a second gateway becomes justified.

Architecture with LiteLLM

┌──────────────────────────────────────────────────────────┐
│                     Member's browser                       │
│  Chat: Open WebUI or LobeChat → api.inference.coop        │
│  API: Direct → api.inference.coop/v1/chat/completions     │
└────────────────────────┬─────────────────────────────────┘
                         │ sk-coop-xxxxx (per-member key)
┌────────────────────────▼─────────────────────────────────┐
│              LiteLLM Proxy (Docker)                        │
│  • Virtual keys per member                                 │
│  • Rate limits per tier                                    │
│  • Usage tracking per user/model                           │
│  • Spend tracking for billing                              │
│  • Model routing + fallback                                │
└──────┬──────────────────┬──────────────────┬────────────┘
       │                  │                  │
┌──────▼─────┐    ┌──────▼──────┐    ┌──────▼──────┐
│  Tinfoil   │    │   Ollama    │    │ OpenRouter  │
│  (TEE)     │    │ (self-host) │    │ (fallback)  │
│  Default   │    │ Free tier   │    │ Frontier    │
└────────────┘    └─────────────┘    └─────────────┘

Is LiteLLM in Cloudron?

LiteLLM is not in the Cloudron app store, but it can be packaged as a custom Cloudron app. This is the planned approach — it gives LiteLLM the full Cloudron treatment: SSO, sandboxing, SSL, backups, and access control, without a separate Docker container or reverse proxy.

How custom Cloudron apps work:

  1. Create a CloudronManifest.json (app ID, port, OIDC/LDAP addons, health check)
  2. Write a Dockerfile wrapping LiteLLM's image with Cloudron-specific configuration (read-only rootfs, /app/data for persistent DB, OIDC integration)
  3. cloudron build && cloudron install — deploys to Cloudron as a one-click app
  4. LiteLLM appears in the Cloudron dashboard with SSO, SSL, sandboxing, and backups

Manifest (planned):

{
  "id": "ai.coop.litellm",
  "title": "LiteLLM Gateway",
  "httpPort": 4000,
  "addons": {
    "oidc": {},
    "ldap": {},
    "localstorage": {}
  }
}

The oidc addon gives LiteLLM access to Cloudron's SSO. The ldap addon lets it query the user directory. localstorage provides /app/data for the LiteLLM database.

Result: Members log in once via Cloudron SSO → access Open WebUI (chat), Loomio (governance), and LiteLLM admin dashboard (if admin) — all same credentials. API access uses per-member virtual keys managed by LiteLLM.

Build task: Package LiteLLM as a Cloudron app (~1 day). See Cloudron packaging tutorial. Could also publish to the Cloudron community app store for others.

Sources: Cloudron packaging docs, Cloudron manifest docs, Cloudron user directory (OIDC)

Alternatives to LiteLLM

Tool Per-user keys Usage tracking Rate limiting Self-hosted?
LiteLLM ✅ ✅ Per-user, per-model ✅ ✅ Docker
Open WebUI built-in ✅ ✅ Via token tracking pipes Partial (in development, #23323) ✅ (already in Cloudron)
Morph LLM Proxy ✅ ✅ ✅ ✅
Custom middleware Build it Build it Build it ✅

Open WebUI's built-in token tracking is improving and may eventually handle per-user limits natively. For the MVP, LiteLLM is the more complete solution. As Open WebUI's features mature, the cooperative could simplify by dropping LiteLLM and using Open WebUI's native tracking.

Component 8: Privacy and confidential inference

A cooperative handling member data has a stronger duty of care than a corporate provider — members own the cooperative, so the cooperative should not be able to surveil them. This is a structural advantage over corporate AI if we get it right.

Privacy tiers (progressive)

The privacy model should be designed in layers, each stronger than the last:

Tier 1: Self-hosted data sovereignty (MVP baseline)

  • All data stays on cooperative infrastructure — no third-party access
  • Open WebUI stores conversation history in its own database (SQLite or PostgreSQL)
  • Encryption at rest: Use SQLCipher for SQLite or disk-level encryption for PostgreSQL
  • Configurable retention: Open WebUI supports automatic chat deletion after X days (feature request #21027, in development)
  • No training on member data: Models are pre-trained; member conversations are never used for fine-tuning
  • Sources: Open WebUI security docs, Open WebUI hardening guide

Tier 2: Member-controlled retention + browser-local chat history

Two complementary approaches:

Option A: Open WebUI with server-side retention controls

  • Members choose their own data retention: session-only, 7 days, 30 days, permanent
  • Session-only mode: conversations deleted when the session ends — no persistent storage at all
  • API access: no conversation logging for API users, only usage metering (token counts)
  • This requires a small configuration layer on top of Open WebUI's existing retention settings

Option B: LobeChat in browser-local mode (Duck.ai-style)

LobeChat is the better architecture for the Duck.ai approach. It has two deployment modes:

  • Standalone/Lite mode: All conversations stored in browser's IndexedDB. No server-side database. Chat history never leaves the user's device. This is exactly the Duck.ai model.
  • Database mode: PostgreSQL server-side storage with multi-user support for when persistence/sync is needed.

LobeChat features relevant to inference.coop:

  • Browser-local storage: IndexedDB — conversations stay on the user's device, not on cooperative servers
  • CRDT-based sync: Experimental multi-device sync using Conflict-Free Replicated Data Types — no central server needed
  • Multi-provider support: Connect to Ollama, vLLM, OpenAI, Anthropic, OpenRouter from one interface
  • OpenAI-compatible: Points at any OpenAI-compatible API endpoint — works with our inference backend
  • Plugin system: Extensible with web search, data tools, MCP (Model Context Protocol)
  • PWA: Works as a Progressive Web App on mobile and desktop
  • MIT licensed, self-hostable, Docker deployment
  • 55K+ GitHub stars

This means members can use LobeChat in standalone mode pointed at inference.coop's API endpoint — their chat history lives only in their browser, the cooperative never stores conversation content, and the cooperative's servers only see transient inference requests.

Duck.ai's architecture (for reference):

  • Chat history stored locally in browser by default — no server-side storage
  • Optional "Sync & Backup" uses end-to-end encrypted server storage (DuckDuckGo never holds decryption keys)
  • 18-month automatic deletion for synced chats not accessed
  • Up to 100 recent conversations stored locally
  • Notably, Duck.ai now uses Tinfoil.sh to host some models — the same TEE-based confidential inference platform relevant to our Tier 3 design
  • Sources: DuckDuckGo help: recent chats, DuckDuckGo help: AI chat privacy, Factually.co: Duck.ai storage analysis

Recommendation: Offer both interfaces:

  • Open WebUI for members who want server-side chat history (with retention controls) and multi-user features
  • LobeChat in standalone mode for members who want Duck.ai-style browser-local privacy

Both point at the same inference API. Members choose their privacy posture by choosing their frontend.

Tier 3: Encrypted inference via TEEs (Confer.to model)

This is where it gets interesting — and where Confer.to's architecture becomes directly relevant.

What Confer.to does

Confer.to (founded by Moxie Marlinspike, Signal founder) provides end-to-end encrypted AI chat:

  1. Client-side encryption: User's prompt is encrypted with WebAuthn passkeys before it leaves the browser
  2. Trusted Execution Environment (TEE): The encrypted prompt enters a hardware-enforced enclave (TEE) on the server. The server's OS cannot read the TEE's memory or execution state. The TEE decrypts the prompt, runs inference, encrypts the output, and sends it back.
  3. Remote attestation: Users can cryptographically verify that the TEE is running the claimed code — the cooperative cannot read conversations even if it wants to.
  4. Statelessness: LLMs are stateless (input in, output out) — making them ideal for TEE deployment since no state persists after the request.

The TEE landscape in 2026

Confidential computing for AI inference matured significantly in 2026:

Technology Hardware Status Overhead
NVIDIA Confidential Computing H100, H200, B200 (GPU TEE) Production-ready. CC-On mode encrypts GPU memory, protected PCIe between CPU and GPU. <7% throughput
Intel TDX Intel CPUs (Trust Domain Extensions) Mature. Creates encrypted VMs. Low
AMD SEV-SNP AMD EPYC CPUs Mature. Encrypted virtualization. Low
AWS Nitro Enclaves AWS cloud Available. Isolated compute environments. Medium

Key developments:

  • NVIDIA H100/H200 Confidential Computing went production in 2025-2026, enabling GPU-accelerated TEE inference
  • TEE overhead dropped below 7% — negligible for most workloads
  • Open-source frameworks emerging: OpenPCC (open-source confidential LLM serving on commodity TEEs, using vLLM, arXiv June 2026)
  • Tinfoil (tinfoil.sh) — open-source, OpenAI-compatible confidential inference API running in hardware enclaves. Collaborating with Red Hat. Production-ready.
  • Sources: GeniusTechLab (Jul 2026), OpenPCC (arXiv Jun 2026), Tinfoil

How inference.coop could build on this

Phase 1 (MVP): Tier 1 + Tier 2. Self-hosted, encrypted at rest, member-controlled retention. The cooperative could technically read conversations, but has a policy commitment not to. This is already stronger than OpenAI/Anthropic, whose terms allow them to read your data.

Phase 2 (growth): Tier 3 with TEEs. When deploying on GPU hardware (for the vLLM migration), choose H100 GPUs with Confidential Computing enabled. Use a framework like OpenPCC or Tinfoil to serve models inside TEEs. At this point, the cooperative cannot read conversations even if it wants to — the privacy guarantee becomes architectural, not just policy-based.

Phase 3 (long-term): Client-side encryption + TEE (full Confer.to model). User's browser encrypts the prompt before transmission. Only the TEE can decrypt it. Even the network operator can't see the content. This requires custom frontend work on top of Open WebUI, but the architecture is proven by Confer.to.

Tinfoil as a potential shortcut

Tinfoil is particularly interesting for the cooperative:

  • Open-source, OpenAI-compatible API
  • Runs models inside hardware enclaves
  • Verifiably private — users can attest the enclave
  • Collaborating with Red Hat on open-source confidential AI infrastructure
  • Could potentially be used as a backend instead of (or alongside) vLLM, providing TEE-protected inference without building the confidential computing layer from scratch

If Tinfoil's API is OpenAI-compatible, it could be a drop-in replacement for the vLLM backend, providing both inference and privacy in one package. Worth investigating for the Phase 2 migration.

Privacy as cooperative advantage

This is the key strategic point: privacy is where the cooperative can structurally outperform corporate AI providers.

Corporate AI (OpenAI, Anthropic) inference.coop Tier 1 inference.coop Tier 3 (TEE)
Can the provider read your data? Yes (per terms of service) Technically yes, but policy prohibits No — architecturally impossible
Can the provider train on your data? Yes (unless you pay for enterprise) No — never No — never
Can you verify the privacy claim? No (trust the company) No (trust the cooperative) Yes — remote attestation
Data leaves your control? Yes (sent to corporate servers) No (stays on co-op infra) No (encrypted before leaving your device)

The cooperative's structural advantage is that it can deploy TEEs and offer verifiable privacy without a business model that depends on reading user data. OpenAI and Anthropic could deploy TEEs, but their business models (training on user data, targeting, enterprise upsells) create a conflict of interest. The cooperative has no such conflict.


Architecture summary (MVP)

┌─────────────────────────────────────────────────┐
│                 Member's browser                  │
│  Chat: https://chat.inference.coop (Open WebUI)  │
│  Chat: https://chat.inference.coop (LobeChat)   │
│  API:  https://api.inference.coop/v1/...         │
│  Gov:  https://gov.inference.coop (Loomio)       │
│  Pay:  https://opencollective.com/inference-coop│
└──────────────────────┬──────────────────────────┘
                       │ HTTPS + Cloudron SSO (OIDC)
┌──────────────────────▼──────────────────────────┐
│              Cloudron (VPS, green energy)         │
│  ┌─────────┐  ┌──────────┐  ┌─────────────────┐ │
│  │ SSO/OIDC│  │ Open WebUI│  │  Loomio         │ │
│  └─────────┘  └────┬─────┘  └─────────────────┘ │
│               ┌────▼─────┐                       │
│               │ LiteLLM  │ ← custom Cloudron app  │
│               │ Gateway  │   (SSO, sandboxed)     │
│               └────┬─────┘                        │
│     ┌──────────┬───┴──────┬───────────┐           │
│     │          │          │           │           │
│  ┌──▼──┐  ┌──▼───┐  ┌──▼──┐  ┌──▼──────┐        │
│  │Tinfoil│ │Ollama│  │Open  │  │Billing  │        │
│  │(TEE)  │ │(free)│  │Router│  │middleware│       │
│  └───────┘└──────┘  └─────┘  └─────────┘        │
│  (cloud)  (Cloudron)  (cloud)  (custom webhook) │
└──────────────────────────────────────────────────┘

Green hosting: Hetzner (100% hydropower) or Scaleway DC5 (100% renewable) Privacy: Tinfoil TEE for inference; LobeChat browser-local for chat history SSO: Cloudron OIDC — one login for chat, governance, API admin Billing: Open Collective → webhook middleware → LiteLLM key enable/disable


Cost comparison: self-hosting vs cloud inference (10 users, ~12M tokens/month)

Assumptions: 10 user-testers, ~20 queries/day each, ~2000 tokens per conversation = ~400K tokens/day = ~12M tokens/month.

Option Cost/mo Green? TEE? Self-hosted? Inference quality
Tinfoil API (GPT-OSS 120B) ~$18 Unknown ✅ No Full open models, fast
Tinfoil API (DeepSeek V4 Flash) ~$20 Unknown ✅ No 304B MoE, 1M context
Hetzner CPU + Ollama + OpenRouter €38 ✅ Hydro No ✅ 7B only on CPU (slow, 5-15 tok/s)
Ollama Cloud Team (5 seats min) ~$140 Unknown No No Full model library
Scaleway L4 on-demand (8hrs/day) ~$214 ✅ Wind+hydro No ✅ 7B-13B on GPU (fast)
Hetzner L4 GPU + Ollama €222 ✅ Hydro No ✅ 7B-13B on GPU (fast)
Ollama Cloud Pro (10× individual) ~$215 Unknown No No Full model library

Tinfoil wins on price by 2-10x — the cooperative pays only for actual tokens used, not for idle GPU time. At 12M tokens/month, Tinfoil costs $3.96 (GPT-OSS 120B) to $5.52 (DeepSeek V4 Flash) in inference charges, plus €15/month for Cloudron = ~$18-20/month total.

Self-hosting on green energy costs 2-10x more because you're paying for a GPU server 24/7 regardless of whether anyone is querying it. The cheapest green self-hosted option (Hetzner CPU, €38/month) is slow (7B models on CPU, 5-15 tokens/sec) and still costs more than Tinfoil.

Sources: Tinfoil pricing, Hetzner GPU servers, Scaleway GPU pricing, Ollama Cloud pricing


Green + TEE: the gap

There is no option today that is both clearly green AND TEE-protected at MVP scale:

Approach Green? TEE? Cost/mo MVP viable?
Hetzner (hydro) + Ollama ✅ ❌ €38-222 Yes
Scaleway DC5 (wind+hydro) + Ollama ✅ ❌ ~$214 Yes
Tinfoil API (TEE) Unknown ✅ ~$18-20 Yes
IREN (100% renewable) + H100 TEE ✅ ✅ ~$1,800 No (overkill)
Google Cloud A3 Confidential VM + H100 ~67% carbon-free ✅ ~$1,900 No (overkill)

The closest to "both" at MVP scale is Tinfoil — verifiable TEE privacy at $18-20/month, but energy mix unknown. The cooperative can:

  1. Be transparent about the trade-off — "our inference is TEE-private; we cannot yet verify it is green"
  2. Offset the gap — purchase renewable energy certificates (RECs) to cover estimated inference emissions
  3. Migrate when scale justifies it — at 50+ members, self-host on IREN (renewable + TEE) or deploy Tinfoil's open-source framework on H100 hardware at a green data center

The alternative — self-hosting on Hetzner (green, €38/month, no TEE) — costs only €18/month more but requires GPU management and gives weaker privacy (policy-based, not architectural). For a cooperative whose value proposition is member privacy, the TEE advantage is worth more than the green gap.


Hybrid approach (Tinfoil + Cloudron + optional Ollama):

Component Monthly cost Notes
VPS (Cloudron host) €15-20 2-4 vCPU, 4-8GB RAM — no GPU needed
Cloudron license €15 SSO, app management, security sandboxing
Tinfoil Inference API $5-20 Pay-as-you-go, GPT-OSS 120B at $0.15/$0.60 per M tok
OpenRouter (fallback) $2-5 For models Tinfoil doesn't offer
Domain ~€1 inference.coop or similar
Open Collective 0-5% Platform fee depends on fiscal host
Total ~€40-60/month No GPU server needed

At 10 members paying €10/month = €100/month → surplus of €40-60/month for reinvestment.

Note: If Tinfoil per-token costs grow with usage, the cooperative can add self-hosted Ollama (free) for lightweight models or rent GPU hardware for vLLM at the break-even point (~50 members).


What's missing / needs further design

  1. Legal structure — Open Collective fiscal sponsorship for MVP; formal cooperative incorporation later
  2. Data policy — what's logged, retained, shared. FERPA/GDPR considerations. Member-controlled retention (Tier 2).
  3. Model governance — how members decide which models to add/remove
  4. Sustainability metrics — energy per query, CO₂ tracking (à la GreenPT)
  5. API key management — whether to use Open WebUI's built-in keys or a dedicated gateway (LiteLLM)
  6. Migration path — Ollama → vLLM when scale requires it
  7. Agent support — whether to offer agentic capabilities (like Hermes Agent) beyond chat
  8. Confidential inference (Phase 2-3) — TEE deployment via OpenPCC or Tinfoil when GPU hardware is available. Client-side encryption (Confer.to model) as long-term goal. This is the cooperative's structural privacy advantage.

Connections

  • co/core — most architecturally similar project (ATProto, mutual credit)
  • GreenPT — renewable energy AI hosting precedent
  • Public AI Switzerland — consumer-facing AI cooperative model
  • Ollama — inference engine + potential E2C candidate
  • OpenRouter — model brokering for frontier access
  • Campus AI Cooperative — larger-scale version of the same concept
  • AI Potluck — coalition-built public AI stack; see learnings below

Learnings from AI Potluck

AI Potluck (aipotluck.org) is the closest large-scale analogue to inference.coop's goal — "AI made and owned by the public" — but built as a coalition rather than a cooperative. Its design choices are directly instructive:

  1. Provenance transparency as a product feature. AI Potluck's commitment "every response shows its provenance: the model, the organization, the compute, the country" is a concrete, implementable norm. inference.coop could adopt a lighter version: surface which model served each response (already possible via LiteLLM) and which backend (Tinfoil vs. Ollama) — turning the co-op's transparency value into a visible feature rather than a policy statement.

  2. Anti-engagement design as a differentiator. AI Potluck explicitly rejects the extractive engagement model ("AI that wants you to turn it off," "doesn't use your emotional state to extend the conversation," "doesn't sell you sycophancy"). This is a values stance inference.coop shares and could articulate more explicitly in its charter and product copy — a cooperative has no engagement-maximization incentive, which is a structural advantage worth naming.

  3. The "no single point of failure" resilience argument. AI Potluck's core claim — distributed ownership means no single actor's withdrawal is fatal — is a resilience argument, not just a fairness one. inference.coop's cooperative model achieves a version of this (no single member owns the whole), but the framing is worth borrowing: resilience-through-distribution is a selling point for members, not just a governance nicety.

  4. Data ethics as a trust signal. "We don't host models trained on stolen data... you can delete everything in two steps." inference.coop's privacy posture (TEE, no training on member data) is stronger than AI Potluck's, but the communication of it — simple, concrete, verifiable — is a model to follow.

  5. The ownership-model contrast. AI Potluck is a coalition (no single owner, no formal entity); inference.coop is a cooperative (member-owned, formal bylaws). The key difference is accountability: a cooperative has a defined membership, voting, and a legal entity that can be held responsible; a coalition does not. This is inference.coop's structural advantage over the potluck model — and worth articulating when explaining why "cooperative" rather than "coalition."

  6. Government-as-contributor framing. AI Potluck reframes sovereign AI as contribution ("a government can own the parts that matter most to them without building everything else from scratch"). For inference.coop, this suggests a path for institutional members (universities, municipalities) to contribute compute or data without needing to own the whole — a membership tier worth considering.