Add MVP planning doc (moved from wiki inference-coop/README.md)
This commit is contained in:
1 parent
28719fb5ac
commit
54b4603ffa
1 file changed
+801
+801
@@ -0,0 +1,801 @@
|
||||
---
|
||||
type: research-project
|
||||
function: inference
|
||||
region: global
|
||||
maturity: planned
|
||||
ownership: cooperative
|
||||
---
|
||||
|
||||
# inference.coop — Minimum Viable Cooperative Inference Utility
|
||||
|
||||
**Status:** Design/planning phase
|
||||
**Category:** Cooperative AI infrastructure — proof of concept
|
||||
**Goal:** A working, membership-based inference cooperative providing API and chatbot access to open models, hosted on renewable energy, with unified account management and cooperative governance.
|
||||
|
||||
## Design goals
|
||||
|
||||
1. **Leverage existing open-source software** for the full stack — no custom infrastructure to build
|
||||
2. **Renewable energy hosting** — inference on providers powered by renewables
|
||||
3. **Unified account management** across services (chat, API, billing, governance)
|
||||
4. **Membership-based payment** through a service that supports recurring payments and cooperative governance
|
||||
5. **Small scale first** — viable for a handful of user-testers, then scale
|
||||
|
||||
---
|
||||
|
||||
## Component 1: Inference engine
|
||||
|
||||
The backend that loads and serves open models via an OpenAI-compatible API.
|
||||
|
||||
### Options
|
||||
|
||||
| Engine | Throughput | Key features | Best for |
|
||||
|--------|-----------|-------------|----------|
|
||||
| **vLLM** | Highest (PagedAttention, continuous batching) | OpenAI-compatible API, multi-LoRA, production-grade multi-tenant | Production serving, many concurrent users |
|
||||
| **Ollama** | Lower (single-user optimized) | One-command setup, model library, extreme ease of use | MVP/prototyping, small deployments |
|
||||
| **SGLang** | ~30% faster than vLLM on some workloads | RadixAttention, structured outputs | Latency-sensitive workloads |
|
||||
| **TGI (HuggingFace)** | Competitive | HF ecosystem integration | HF-centric workflows |
|
||||
|
||||
### Recommendation
|
||||
|
||||
**Two paths to consider — and Tinfoil changes the calculus:**
|
||||
|
||||
**Path A: Self-hosted Ollama → vLLM (original plan)**
|
||||
- Start with Ollama (Cloudron app), migrate to vLLM at scale
|
||||
- Full control over model selection, no per-token costs for self-hosted models
|
||||
- But: no TEE privacy at MVP, and vLLM migration required at scale
|
||||
|
||||
**Path B: Tinfoil Inference API from day one (new option)**
|
||||
- OpenAI-compatible API — drop-in for Open WebUI and LobeChat
|
||||
- TEE-protected inference from the start — verifiable privacy, no pinky promises
|
||||
- No GPU hosting, no model management, no Cloudron GPU app needed
|
||||
- Pay-as-you-go per token (see pricing below)
|
||||
- Open-source, SOC 2 compliant, audited by Trail of Bits
|
||||
- Used by Duck.ai, UC Berkeley, Stanford, Red Hat
|
||||
|
||||
**Tinfoil model lineup and pricing (as of Aug 2026):**
|
||||
|
||||
| Model | Context | Parameters | Input $/M tok | Output $/M tok |
|
||||
|-------|---------|-----------|--------------|---------------|
|
||||
| DeepSeek V4 Flash | 1M | 304B (MoE) | $0.30 | $0.70 |
|
||||
| GPT-OSS 120B | 131k | 117B (5.1B active) | $0.15 | $0.60 |
|
||||
| Gemma 4 31B | 256k | 31B | $0.40 | $1.00 |
|
||||
| Llama 3.3 70B | 128k | 70B | $1.75 | $2.75 |
|
||||
| GLM-5.2 | 384k | 754B (40B active) | $1.50 | $5.25 |
|
||||
| Kimi K3 | 256k | 2.8T (MoE) | $4.00 | $20.00 |
|
||||
| Nomic Embed Text | 8k | 137M | $0.05 | $0 |
|
||||
| Whisper Large V3 Turbo | 30s audio | 809M | $0.01/request | |
|
||||
|
||||
Source: [Tinfoil models page](https://tinfoil.sh/models)
|
||||
|
||||
**Cost comparison for 10 members, ~20 queries/day:**
|
||||
- Path A (self-hosted): €36-66/month fixed (VPS + Cloudron + OpenRouter for frontier)
|
||||
- Path B (Tinfoil): ~$15-30/month variable (at GPT-OSS 120B rates, ~500K tokens/day for 10 users) + €15/month Cloudron for chat frontend/governance/SSO. No GPU server needed.
|
||||
|
||||
**The hybrid approach (recommended):**
|
||||
|
||||
Use **both** — Tinfoil as the default inference backend, with self-hosted Ollama for lightweight models when cost or latency matters:
|
||||
|
||||
1. **Tinfoil API** for all models — gives TEE privacy from day one, no GPU management
|
||||
2. **Ollama (Cloudron)** as an optional backend for small models (7B range) when members want zero per-token cost and are OK without TEE
|
||||
3. **OpenRouter** as fallback for models Tinfoil doesn't offer yet (e.g. Claude, GPT)
|
||||
|
||||
Open WebUI and LobeChat both support multiple OpenAI-compatible endpoints — members can choose which backend serves their request. The cooperative's middleware routes and meters usage across all three.
|
||||
|
||||
This means the MVP launches with **verifiable privacy (TEE) from day one** without any GPU hosting or TEE configuration — Tinfoil handles the hard part. The cooperative's infrastructure is just Cloudron (for chat, governance, SSO) + a thin middleware for billing/access control + Tinfoil's API.
|
||||
|
||||
**When to add self-hosted inference:** When per-token costs exceed the cost of renting GPU hardware (probably at 50+ active members), add self-hosted vLLM on H100 GPUs with Confidential Computing enabled for the most-used models. Use Tinfoil's open-source framework ([GitHub](https://github.com/tinfoilsh)) as reference for TEE deployment.
|
||||
|
||||
- Tinfoil: [https://tinfoil.sh/](https://tinfoil.sh/)
|
||||
- Tinfoil models: [https://tinfoil.sh/models](https://tinfoil.sh/models)
|
||||
- Tinfoil GitHub: [https://github.com/tinfoilsh](https://github.com/tinfoilsh)
|
||||
- Tinfoil docs: [https://docs.tinfoil.sh/](https://docs.tinfoil.sh/)
|
||||
- Tinfoil JS SDK: [https://github.com/tinfoilsh/tinfoil-js](https://github.com/tinfoilsh/tinfoil-js) — OpenAI-compatible, encrypts requests with HPKE
|
||||
- Duck.ai uses Tinfoil: [DuckDuckGo help](https://duckduckgo.com/duckduckgo-help-pages/duckai)
|
||||
|
||||
### Model selection
|
||||
|
||||
Start with 2-3 models covering capability tiers:
|
||||
- **Lightweight:** Qwen 3.5 7B or Llama 4 Scout 13B — fast, cheap, handles 80% of queries
|
||||
- **Mid-range:** DeepSeek V4 Lite or Qwen 3.5 32B — research assistance, complex reasoning
|
||||
- **Frontier (brokered):** Access via OpenRouter API for models too large to self-host
|
||||
|
||||
### Recommendation: use OpenRouter for frontier access
|
||||
|
||||
OpenRouter provides a single API to 300+ models from 60+ providers with pay-as-you-go pricing. For the MVP, self-host lightweight/mid-range models and broker frontier access through OpenRouter. This avoids the cost of renting H100 GPUs for frontier models while still offering full model coverage.
|
||||
- [OpenRouter](https://openrouter.ai/)
|
||||
|
||||
---
|
||||
|
||||
## Component 2: Chat frontend
|
||||
|
||||
The user-facing web interface for chatting with models.
|
||||
|
||||
### Options
|
||||
|
||||
| Frontend | Stars | Key features | Cloudron app? |
|
||||
|----------|-------|-------------|--------------|
|
||||
| **Open WebUI** | 141K+ | ChatGPT-style, multi-user, RAG, model switching, document upload, pipelines, Python extensibility | ✅ Yes |
|
||||
| **LibreChat** | — | ChatGPT-like with model switching, multi-user | ✅ Yes |
|
||||
| **Lobster** | — | Clean chat interface | ❌ |
|
||||
|
||||
### Recommendation: Open WebUI
|
||||
|
||||
Open WebUI is the leading self-hosted chat interface (141K+ GitHub stars). It's available as a **one-click Cloudron app**, supports multi-user accounts with role-based access, model switching, document upload for RAG, and connects to both Ollama and any OpenAI-compatible API (including vLLM and OpenRouter). It also has a Python pipeline system for extensibility.
|
||||
|
||||
- [Open WebUI](https://openwebui.com/)
|
||||
- [Open WebUI on GitHub](https://github.com/open-webui/open-webui)
|
||||
- [Open WebUI on Cloudron App Store](https://www.cloudron.io/store/index.html) (listed as "Open WebUI — ChatGPT-Style WebUI for LLMs")
|
||||
- Open WebUI + vLLM integration guide: [Markaicode (2026)](https://markaicode.com/integrate/open-webui-with-vllm/)
|
||||
- Self-hosted AI stack guide (Ollama + Open WebUI): [DEV Community (May 2026)](https://dev.to/jovan_chan_9500711396d4e6a/self-hosted-ai-privacy-stack-2026-49n8)
|
||||
|
||||
---
|
||||
|
||||
## Component 3: Hosting (renewable energy)
|
||||
|
||||
### Options
|
||||
|
||||
| Provider | Renewable claim | GPU types | Location | Notes |
|
||||
|----------|----------------|-----------|----------|-------|
|
||||
| **Scaleway** | 100% renewable (wind + hydro, GO certified) at DC5 (Paris) | L4, L40S | Paris, Amsterdam, Warsaw | EU data residency, GDPR. No H100. Best for lightweight models. |
|
||||
| **Hetzner** | Certified data centers in Germany/Finland (mixed energy) | GPU dedicated servers | Germany, Finland | Best value CPU/GPU servers, but not 100% renewable |
|
||||
| **IREN** | 100% renewable-powered data centers | H100, H200, B200 | Global (AU, US, CA) | For GPU-heavy workloads |
|
||||
| **Lambda Labs** | Not renewable-specific | H100, A100 | US | Best price/reliability for GPU rental |
|
||||
| **GreenPT** (backend) | 100% renewable (uses Scaleway infra) | Not disclosed | Netherlands | Could white-label, but €17.50/seat/month is expensive |
|
||||
|
||||
### Recommendation
|
||||
|
||||
**Three tiers of green hosting, depending on whether the MVP uses GPU:**
|
||||
|
||||
**Tier 1: CPU-only VPS (no GPU) — for Ollama small models or Cloudron-only (no self-hosted inference)**
|
||||
|
||||
If using Tinfoil/Ollama Cloud for inference and Cloudron only for chat/governance/SSO, no GPU is needed. A basic VPS suffices:
|
||||
|
||||
| Provider | Renewable claim | GPU? | Price | Location | Notes |
|
||||
|----------|----------------|------|-------|----------|-------|
|
||||
| **Hetzner Cloud** | 100% hydropower (Germany since 2008), hydro+wind (Finland since 2018). Founded HT Clean Energy GmbH for solar parks. EMAS/ISO 14001 certified. | Cloud: no GPU. Dedicated: yes | CPX31 (4 vCPU, 8GB RAM): ~€13/mo | Germany, Finland | Best value. 100% renewable. |
|
||||
| **Scaleway** | 100% renewable at DC5 (Paris): wind + hydro, GO-certified. PUE 1.16. Free + adiabatic cooling. | Cloud: L4, L40S. No A100/H100 | DEV1-L (4 vCPU, 8GB): ~€15/mo | Paris, Amsterdam, Warsaw | Strongest renewable claims in EU. |
|
||||
|
||||
- Sources: [Hetzner sustainability](https://www.hetzner.com/unternehmen/nachhaltigkeit/), [Hetzner green energy](https://cloudnews.tech/hetzner-invests-in-renewable-energy-with-the-creation-of-ht-clean-energy-gmbh/), [Scaleway L40S (renewable)](https://www.scaleway.com/en/l40s-gpu-instance/)
|
||||
|
||||
**Tier 2: GPU server for self-hosted Ollama (lightweight models, 7B-13B)**
|
||||
|
||||
For self-hosting small models on Ollama, an L4 or L40S GPU is sufficient:
|
||||
|
||||
| Provider | Renewable claim | GPU | Price | Location | Notes |
|
||||
|----------|----------------|-----|-------|----------|-------|
|
||||
| **Scaleway** | 100% renewable (DC5 Paris) | L4 | ~$0.79/hr on-demand, ~$289/mo reserved | Paris | Greenest GPU option in Europe |
|
||||
| **Scaleway** | 100% renewable (DC5 Paris) | L40S (48GB) | ~$1.68/hr on-demand, ~$614/mo reserved | Paris | Best for 13B models |
|
||||
| **Hetzner** | 100% hydropower (Germany/Finland) | L4 | €199/mo dedicated | Germany, Finland | Fixed monthly, no A100/H100 |
|
||||
| **Hetzner** | 100% hydropower | L40S | €349/mo dedicated | Germany, Finland | Fixed monthly, good value |
|
||||
| **Hetzner** | 100% hydropower | RTX 6000 Ada (96GB) | €889/mo dedicated | Germany, Finland | Largest VRAM, can run 70B quantized |
|
||||
|
||||
- Sources: [Hetzner GPU servers](https://www.hetzner.com/dedicated-root-server/matrix-gpu/), [Scaleway GPU pricing](https://www.scaleway.com/en/pricing/gpu/), [Hetzner GPU review 2026](https://gpuhosted.com/en/hetzner-gpu-review/)
|
||||
|
||||
**Tier 3: H100 for frontier models + future TEE migration**
|
||||
|
||||
For H100/H200 with Confidential Computing (TEE), needed when serving 70B+ models or implementing Tier 3 privacy:
|
||||
|
||||
| Provider | Renewable claim | GPU | Price | Location | TEE capable? |
|
||||
|----------|----------------|-----|-------|----------|-------------|
|
||||
| **Scaleway** | 100% renewable (DC5) | H100 SXM | ~$2.80/hr on-demand | Paris | Unknown |
|
||||
| **Lambda Labs** | Not renewable-specific | H100 | $2.21/hr reserved | US | H100 CC capable |
|
||||
| **IREN** | 100% renewable-powered DCs | H100, H200, B200 | Custom pricing | Global (AU, US, CA) | H100 CC capable |
|
||||
|
||||
- Sources: [Scaleway GPU pricing](https://www.scaleway.com/en/pricing/gpu/), [Lambda Labs pricing](https://lambda.ai/pricing)
|
||||
|
||||
### Key findings
|
||||
|
||||
1. **Hetzner is the greenest practical choice for the MVP.** 100% hydropower since 2008 (Germany) and 2018 (Finland). Now building own solar parks via HT Clean Energy GmbH. EMAS/ISO 14001 certified. GPU servers available from €199/month (L4) or €349/month (L40S). No H100/H200 — tops out at RTX 6000 Ada (96GB VRAM, can run 70B quantized).
|
||||
|
||||
2. **Scaleway DC5 is the greenest GPU cloud.** 100% renewable (wind + hydro, GO-certified), PUE 1.16, ultra-efficient cooling. Offers L4, L40S, and H100. On-demand pricing means you pay only when serving requests.
|
||||
|
||||
3. **Neither Hetzner nor Scaleway explicitly supports NVIDIA Confidential Computing (TEE).** For the TEE migration (Phase 3 privacy), IREN or Lambda Labs with H100 CC mode would be needed. This is a future concern, not an MVP requirement.
|
||||
|
||||
4. **Ollama on CPU is viable for the MVP.** For 5-10 user-testers using 7B models (Qwen 3.5 7B, Gemma 4 9B), Ollama runs on CPU — slowly (5-15 tokens/sec) but usable for a proof of concept. A Hetzner CPX31 (4 vCPU, 8GB RAM, ~€13/mo) can serve lightweight models without a GPU. When faster response is needed, add a GPU server or switch to Tinfoil/Ollama Cloud for inference.
|
||||
|
||||
### Recommended MVP hosting setup
|
||||
|
||||
**Option A: Minimal (CPU only, no GPU)**
|
||||
- Hetzner CPX31 (4 vCPU, 8GB RAM): €13/mo — runs Cloudron + Ollama (7B models on CPU, slow but functional)
|
||||
- Inference: Ollama self-hosted (lightweight models) + OpenRouter (frontier models)
|
||||
- Total infra: ~€28/mo (VPS + Cloudron)
|
||||
- Renewable: ✅ 100% hydropower
|
||||
|
||||
**Option B: With GPU (recommended for better UX)**
|
||||
- Hetzner CCX23 (dedicated vCPU, 8GB RAM) for Cloudron: ~€25/mo
|
||||
- Hetzner GEX44 (L4 GPU dedicated server): €199/mo
|
||||
- Inference: Ollama self-hosted on GPU (fast, 7B-13B models) + OpenRouter (frontier)
|
||||
- Total infra: ~€239/mo (VPS + Cloudron + GPU server)
|
||||
- Renewable: ✅ 100% hydropower
|
||||
- *Or:* Scaleway L4 on-demand (~$0.79/hr = ~$230/mo if running 24/7, but can shut down off-hours)
|
||||
|
||||
**Option C: Cloud inference, no GPU hosting**
|
||||
- Hetzner CPX31 for Cloudron only: €13/mo
|
||||
- Inference: Tinfoil API (TEE privacy, per-token) or Ollama Cloud ($25/seat/mo)
|
||||
- Total infra: ~€28/mo + per-token/seat costs
|
||||
- Renewable: ✅ VPS on hydropower; inference provider's energy mix unknown
|
||||
|
||||
### Sustainability tracking
|
||||
|
||||
GreenPT's real-time per-conversation CO₂ display is a feature the cooperative should replicate. This requires:
|
||||
- Measuring GPU/CPU energy per inference request
|
||||
- Multiplying by the grid carbon intensity of the data center's location
|
||||
- Displaying in the chat interface (Open WebUI supports custom UI elements via pipelines)
|
||||
|
||||
For Hetzner Germany/Finland (100% renewable), the carbon intensity is near-zero. For Scaleway DC5 Paris (100% renewable GO-certified), also near-zero. This is a genuine competitive differentiator: the cooperative can credibly claim near-zero carbon AI inference.
|
||||
|
||||
---
|
||||
|
||||
## Component 4: Unified account management (Cloudron)
|
||||
|
||||
### What Cloudron provides
|
||||
|
||||
Cloudron is a self-hosting platform that runs on your VPS and provides:
|
||||
- **One-click app installation** from an app store (100+ apps)
|
||||
- **Single Sign-On (SSO)** — users log in once, access all apps
|
||||
- **Centralized user management** — LDAP-based directory, admin/user roles
|
||||
- **Automatic SSL, DNS, backups, updates**
|
||||
- **Access control** — restrict which users can access which apps
|
||||
|
||||
- [Cloudron](https://www.cloudron.io/)
|
||||
- [Cloudron docs: user management](https://docs.cloudron.io/user-management/)
|
||||
- Cloudron pricing: ~€15/month for the Cloudron license + your VPS cost
|
||||
|
||||
### Key Cloudron apps for inference.coop
|
||||
|
||||
| App | Purpose | In Cloudron store? |
|
||||
|-----|---------|-------------------|
|
||||
| **Open WebUI** | Chat frontend | ✅ |
|
||||
| **Ollama** | Model serving | ✅ |
|
||||
| **LibreChat** | Alternative chat frontend | ✅ |
|
||||
| **Gitea** | Git hosting (for skills/agents) | ✅ |
|
||||
| **Nextcloud** | File sharing, collaboration | ✅ |
|
||||
| **Loomio** | Democratic decision-making / governance | ✅ |
|
||||
| **Discourse** | Forum / community discussion | ✅ |
|
||||
| **Matrix/Element** | Chat / communication | ✅ |
|
||||
|
||||
This means **the entire cooperative platform** — inference, chat, governance, communication, file sharing — can be deployed through Cloudron with unified SSO. Members log in once and access all services.
|
||||
|
||||
### Alternatives to Cloudron
|
||||
|
||||
| Platform | Open source? | SSO? | Notes |
|
||||
|----------|------------|------|-------|
|
||||
| **Coolify** | ✅ Yes (self-hostable) | Partial | Open-source Cloudron alternative. 280+ one-click services. No built-in SSO/LDAP. |
|
||||
| **YunoHost** | ✅ Yes | ✅ Yes | Free, Debian-based. Simpler than Cloudron. Fewer apps. Already used on Nathan's server. |
|
||||
| **Cosmo Cloud** | ✅ Yes | ✅ Yes | Newer, Cosmos Server based. |
|
||||
|
||||
### Recommendation
|
||||
|
||||
**Cloudron is the right choice for inference.coop.** The security model is materially stronger than YunoHost for running untrusted or semi-trusted software like model inference engines.
|
||||
|
||||
**Cloudron security model** (per [Cloudron security docs](https://docs.cloudron.io/security/)):
|
||||
- **Docker app isolation:** Each app runs in its own Docker container. One app cannot access another app's database or files.
|
||||
- **Read-only rootfs:** Apps cannot tamper with their own application code.
|
||||
- **AppArmor profiles:** System calls restricted, `/proc` and `/sys` filesystems blocked.
|
||||
- **Non-root execution:** All apps run as non-root user by default.
|
||||
- **Per-app subdomains:** Prevents XSS in one app from compromising others (unlike sub-path deployments).
|
||||
- **Dropped capabilities:** `CAP_SYS_ADMIN` and other dangerous capabilities are dropped.
|
||||
- **Automatic SSL/TLS:** All apps HTTPS-only, A+ SSL Labs rating, HSTS, wildcard certs to avoid certificate transparency leaks.
|
||||
- **Rate limiting:** Built-in protection against brute force on passwords, SSH, email, database.
|
||||
- **Signed updates:** Platform updates GPG-signed, keys maintained offline.
|
||||
- **No remote access:** No mechanism for cloudron.io to access your server.
|
||||
- **Automatic OS security updates:** Ubuntu unattended-upgrades enabled.
|
||||
|
||||
**YunoHost security model** — YunoHost does not provide Docker-based app isolation by default. Apps are installed directly on the system (not in containers), meaning a vulnerability in one app can potentially affect others. There is no AppArmor profiling, no read-only rootfs, no per-app sandboxing. This is acceptable for trusted, well-maintained apps on a personal server, but is a real concern for:
|
||||
- Running model inference engines that process untrusted user input
|
||||
- Running multiple services for multiple members (not just yourself)
|
||||
- A cooperative that has a duty of care to its members' data
|
||||
|
||||
**Cost comparison:**
|
||||
- Cloudron: €15/month license + VPS cost
|
||||
- YunoHost: free + VPS cost
|
||||
- The €15/month is the cost of containerized isolation, automatic security updates, and professional app packaging. For a cooperative handling member data and running inference workloads, this is a necessary expense, not a luxury.
|
||||
|
||||
**Alternative considered: Coolify** (open-source, self-hostable). Coolify deploys apps via Docker (better isolation than YunoHost) but lacks built-in SSO/LDAP user management and the polished app store. It's a middle ground but doesn't solve the unified account management problem as well as Cloudron.
|
||||
|
||||
- [Cloudron security docs](https://docs.cloudron.io/security/)
|
||||
- [Cloudron vs YunoHost comparison](https://sugggest.com/compare/cloudron-vs-yunohost)
|
||||
- [Self-hosted sandboxing platforms comparison](https://www.pistack.xyz/posts/2026-06-15-self-hosted-application-sandboxing-sandstorm-freedombox-cloudron/)
|
||||
|
||||
---
|
||||
|
||||
## Component 5: Payment and membership (Open Collective)
|
||||
|
||||
### What Open Collective provides
|
||||
|
||||
Open Collective is a platform for transparent funding of open projects:
|
||||
- **Recurring payments** — monthly/yearly tiers
|
||||
- **Fiscal sponsorship** — no need to incorporate immediately; a fiscal host handles legal/banking
|
||||
- **Transparent budget** — all income and expenses visible
|
||||
- **Tier system** — different membership levels with different benefits
|
||||
- **Expense submission** — members can submit expenses for reimbursement
|
||||
|
||||
- [Open Collective](https://opencollective.com/)
|
||||
- [Open Collective: fiscal hosts FAQ](https://docs.opencollective.com/help/fiscal-hosts/fiscal-hosts)
|
||||
|
||||
### How it fits inference.coop
|
||||
|
||||
1. Create an Open Collective for inference.coop
|
||||
2. Set up tiers with tiered budgets (see below)
|
||||
3. Connect to a fiscal host (e.g., Open Collective Foundation, or a cooperative-specific host)
|
||||
4. Members pay through Open Collective → funds held transparently
|
||||
5. Expenses (server costs, Tinfoil API costs) paid from the collective
|
||||
|
||||
### Membership tiers and budgets
|
||||
|
||||
The cooperative uses **tiered budgets** — each membership level maps to a LiteLLM budget cap, so "excessive use" is rate-limited natively by LiteLLM (no custom billing logic). A free tier lets people try the service before committing.
|
||||
|
||||
| Tier | Price | LiteLLM budget cap | Purpose |
|
||||
|------|-------|-------------------|---------|
|
||||
| **Free / Trial** | €0 | Small (e.g. €2/month of tokens) | Quick try — enough to test the service, not enough for sustained use |
|
||||
| **Member** | €10/month | Generous (e.g. €15/month of tokens) | Standard membership — day-to-day use |
|
||||
| **Supporter** | €20+/month | Higher (e.g. €30/month of tokens) | Heavy users + those who want to support the co-op |
|
||||
|
||||
The budget cap is just a number set on each member's LiteLLM virtual key. When a member hits their cap, LiteLLM returns a clean "budget exceeded" message — no surprise bills, no proration, no reconciliation logic. The exact cap values are a governance decision (set in Loomio), not a technical constraint.
|
||||
|
||||
**Key design decision: defer usage-based billing.** The MVP uses flat membership with tiered budgets. There is no per-token billing, no overage charges, no proration. If members later want metered billing, that's the point where the middleware grows in complexity — but it's not built speculatively.
|
||||
|
||||
### Gaps and additional needs
|
||||
|
||||
Open Collective handles **money in** (membership payments) but not:
|
||||
- **Access provisioning** — automatically granting/denying API access based on payment status
|
||||
- **Usage metering** — tracking token usage per member (LiteLLM does this)
|
||||
- **Rate limiting** — enforcing per-tier limits (LiteLLM does this via budget caps)
|
||||
|
||||
The only custom code needed is a lightweight middleware (webhook handler) between Open Collective and LiteLLM. LiteLLM handles metering and rate-limiting natively; the middleware just creates/enables/disables keys with the right budget.
|
||||
|
||||
### Alternatives
|
||||
|
||||
| Platform | Recurring payments | Fiscal sponsorship | Membership management | Cooperative-friendly? |
|
||||
|----------|-------------------|-------------------|----------------------|----------------------|
|
||||
| **Open Collective** | ✅ | ✅ | Partial (tiers, not access control) | ✅ Designed for collective governance |
|
||||
| **Patreon** | ✅ | ❌ | Partial | ❌ Extractive platform |
|
||||
| **Stripe + custom** | ✅ | ❌ | Custom build | Neutral |
|
||||
| **Memberstack** | ✅ | ❌ | ✅ Access control, member portals | ❌ SaaS, not cooperative |
|
||||
| **Outseta** | ✅ | ❌ | ✅ Full member management | ❌ SaaS |
|
||||
|
||||
### Recommendation
|
||||
|
||||
**Open Collective for payments + fiscal sponsorship, LiteLLM for metering/rate-limiting, custom middleware for access provisioning.** The middleware is a small webhook handler that:
|
||||
1. Receives Open Collective webhooks when payments are made/missed
|
||||
2. Creates/enables/disables the member's LiteLLM virtual key with the tier's budget cap
|
||||
3. On payment lapse, disables the key (LiteLLM then rejects requests)
|
||||
|
||||
This is maybe 200-500 lines of Python — the only custom code needed for the MVP. Rate-limiting excessive use is LiteLLM's native budget feature, not custom logic.
|
||||
|
||||
---
|
||||
|
||||
## Component 6: Governance (missing from original design goals)
|
||||
|
||||
A cooperative needs governance infrastructure. Cloudron has **Loomio** in its app store — an open-source platform for democratic decision-making used by cooperatives and community organizations worldwide.
|
||||
|
||||
### Recommendation
|
||||
|
||||
Install **Loomio via Cloudron** for member governance:
|
||||
- Proposals and voting
|
||||
- Discussion threads
|
||||
- Working groups
|
||||
- Transparent decision history
|
||||
|
||||
This gives members a space to decide on model selection, pricing, privacy policies, and platform direction — all within the same SSO ecosystem as the chat and API.
|
||||
|
||||
- [Loomio](https://loomio.org/)
|
||||
- [Loomio on Cloudron](https://www.cloudron.io/store/index.html)
|
||||
|
||||
---
|
||||
|
||||
## Component 7: API access gateway (LiteLLM)
|
||||
|
||||
The cooperative needs a layer between members and the inference backends (Tinfoil, Ollama, OpenRouter) that handles:
|
||||
|
||||
- **Per-user API keys** — each member gets their own key
|
||||
- **Usage metering** — track token usage per user for billing
|
||||
- **Rate limiting** — enforce per-tier limits (free, member, supporter)
|
||||
- **Model routing** — route requests to the right backend (Tinfoil for TEE, Ollama for free, OpenRouter for frontier)
|
||||
- **Cost tracking** — track spend per user, per model, per key
|
||||
|
||||
### Solution: LiteLLM Proxy
|
||||
|
||||
[LiteLLM](https://www.litellm.ai/) is an open-source AI gateway that does all of the above. It sits between Open WebUI / LobeChat and the inference backends:
|
||||
|
||||
```
|
||||
Members → Open WebUI / LobeChat → LiteLLM Proxy → Tinfoil (TEE)
|
||||
→ Ollama (self-hosted, free)
|
||||
→ OpenRouter (frontier fallback)
|
||||
```
|
||||
|
||||
**LiteLLM provides:**
|
||||
- **Virtual API keys** — per-user, per-team keys with individual budgets and rate limits
|
||||
- **Spend tracking** — per key, per user, per team, per model, across 140+ providers
|
||||
- **Model access control** — restrict which users can access which models
|
||||
- **Budget enforcement** — when a user hits their cap, requests stop automatically
|
||||
- **Admin dashboard** — usage analytics, spend tracking, key management
|
||||
- **OpenAI-compatible** — one endpoint for all backends; Open WebUI and LobeChat see a single API
|
||||
- **Fallback routing** — if Tinfoil is down, automatically fall back to OpenRouter
|
||||
- **Self-hostable** — runs as a Docker container, can be deployed via Cloudron or alongside it
|
||||
|
||||
Sources: [LiteLLM](https://www.litellm.ai/), [LiteLLM + Ollama gateway guide](https://everylocalai.com/stack/litellm-ollama-gateway), [LiteLLM proxy setup](https://markaicode.com/tutorial/litellm-configuration-guide/)
|
||||
|
||||
### How it fits the cooperative model
|
||||
|
||||
1. **Tinfoil API key** is stored only in LiteLLM's config — members never see it. The cooperative holds one Tinfoil account, LiteLLM distributes access via virtual keys.
|
||||
|
||||
2. **Per-member virtual keys** — each member gets an `sk-coop-xxxxx` key. LiteLLM tracks their usage and enforces their tier's limits. This is how the cooperative meters who uses what.
|
||||
|
||||
3. **Open WebUI integration** — Open WebUI supports per-user API keys and has built-in token tracking (via the [token tracking pipes feature](https://github.com/dartmouth/openwebui-token-tracking)). It can display per-conversation token counts and costs to users.
|
||||
|
||||
4. **Billing connection** — LiteLLM's spend tracking API can be queried by the billing middleware (Component 5) to reconcile usage with Open Collective payments. If a member stops paying, the middleware disables their LiteLLM virtual key.
|
||||
|
||||
5. **Model governance** — LiteLLM's model access control lets the cooperative decide which models each tier can access. Members vote on model selection in Loomio; admins update LiteLLM config to reflect decisions.
|
||||
|
||||
### Gap: self-serve portal and monetization
|
||||
|
||||
LiteLLM is the best AI-native gateway (model routing, token metering, fallback, virtual keys), but it is weak on two capabilities that are load-bearing for a cooperative where members are also API customers:
|
||||
|
||||
1. **Self-serve portal** — LiteLLM's admin UI is for *admins* creating keys. There is no clean way for a member to log in, see their own usage, and generate/rotate/revoke their own key.
|
||||
2. **Monetization** — LiteLLM tracks spend but does not *bill*. No payment, no invoice, no subscription lifecycle.
|
||||
|
||||
**Why not add a separate API gateway (Zuplo, Tyk, Kong)?**
|
||||
|
||||
| Option | Self-hostable? | AI-native routing | Self-serve portal | Monetization |
|
||||
|--------|---------------|-------------------|-------------------|-------------|
|
||||
| LiteLLM alone | ✅ | ✅ | ❌ (admin UI only) | ❌ (tracks, doesn't bill) |
|
||||
| LiteLLM + Tyk/Kong | ✅ | ✅ (LiteLLM) | ✅ | ✅ |
|
||||
| LiteLLM + Zuplo | ❌ (Zuplo is SaaS) | ✅ | ✅ | ✅ |
|
||||
| LiteLLM + custom portal | ✅ | ✅ | ✅ (build it) | ⚠️ (via Open Collective) |
|
||||
|
||||
- **Zuplo is out** — hosted SaaS, conflicts with the self-hostable, no-external-dependencies design. A cooperative should not route member API keys through a third-party SaaS.
|
||||
- **Tyk/Kong is a real option but adds a second gateway** — self-hostable with proper developer portals and monetization, but general-purpose (not AI-native) and oriented toward commercial API products, not cooperative membership. Two systems to operate.
|
||||
- **Custom portal is smaller than it sounds** — LiteLLM already does the hard part (metering, virtual keys, budgets). The portal is a thin web app: login (Cloudron SSO), "generate/rotate/revoke my key" (calls LiteLLM's key API), "my usage" (reads LiteLLM spend data), "my membership" (Open Collective status). ~300-500 lines, and it's the same middleware already planned for the Open Collective webhook — one build, not two.
|
||||
|
||||
**Monetization = Open Collective + LiteLLM spend data.** Open Collective handles the actual money (recurring membership, fiscal sponsorship). LiteLLM's spend tracking gives per-member usage. The middleware reconciles the two: if a member's usage exceeds their tier, or their payment lapses, the middleware adjusts their LiteLLM budget or disables their key.
|
||||
|
||||
**Decision:** Keep LiteLLM as the single AI gateway. Build the self-serve portal + billing reconciliation as part of the middleware (Component 5). Revisit Tyk/Kong only if the cooperative grows to need a full developer portal (docs, tiered self-serve plans, usage dashboards) — at which point the operational cost of a second gateway becomes justified.
|
||||
|
||||
### Architecture with LiteLLM
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────────────────────┐
|
||||
│ Member's browser │
|
||||
│ Chat: Open WebUI or LobeChat → api.inference.coop │
|
||||
│ API: Direct → api.inference.coop/v1/chat/completions │
|
||||
└────────────────────────┬─────────────────────────────────┘
|
||||
│ sk-coop-xxxxx (per-member key)
|
||||
┌────────────────────────▼─────────────────────────────────┐
|
||||
│ LiteLLM Proxy (Docker) │
|
||||
│ • Virtual keys per member │
|
||||
│ • Rate limits per tier │
|
||||
│ • Usage tracking per user/model │
|
||||
│ • Spend tracking for billing │
|
||||
│ • Model routing + fallback │
|
||||
└──────┬──────────────────┬──────────────────┬────────────┘
|
||||
│ │ │
|
||||
┌──────▼─────┐ ┌──────▼──────┐ ┌──────▼──────┐
|
||||
│ Tinfoil │ │ Ollama │ │ OpenRouter │
|
||||
│ (TEE) │ │ (self-host) │ │ (fallback) │
|
||||
│ Default │ │ Free tier │ │ Frontier │
|
||||
└────────────┘ └─────────────┘ └─────────────┘
|
||||
```
|
||||
|
||||
### Is LiteLLM in Cloudron?
|
||||
|
||||
LiteLLM is not in the Cloudron app store, but it can be **packaged as a custom Cloudron app**. This is the planned approach — it gives LiteLLM the full Cloudron treatment: SSO, sandboxing, SSL, backups, and access control, without a separate Docker container or reverse proxy.
|
||||
|
||||
**How custom Cloudron apps work:**
|
||||
1. Create a `CloudronManifest.json` (app ID, port, OIDC/LDAP addons, health check)
|
||||
2. Write a `Dockerfile` wrapping LiteLLM's image with Cloudron-specific configuration (read-only rootfs, `/app/data` for persistent DB, OIDC integration)
|
||||
3. `cloudron build && cloudron install` — deploys to Cloudron as a one-click app
|
||||
4. LiteLLM appears in the Cloudron dashboard with SSO, SSL, sandboxing, and backups
|
||||
|
||||
**Manifest (planned):**
|
||||
```json
|
||||
{
|
||||
"id": "ai.coop.litellm",
|
||||
"title": "LiteLLM Gateway",
|
||||
"httpPort": 4000,
|
||||
"addons": {
|
||||
"oidc": {},
|
||||
"ldap": {},
|
||||
"localstorage": {}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
The `oidc` addon gives LiteLLM access to Cloudron's SSO. The `ldap` addon lets it query the user directory. `localstorage` provides `/app/data` for the LiteLLM database.
|
||||
|
||||
**Result:** Members log in once via Cloudron SSO → access Open WebUI (chat), Loomio (governance), and LiteLLM admin dashboard (if admin) — all same credentials. API access uses per-member virtual keys managed by LiteLLM.
|
||||
|
||||
**Build task:** Package LiteLLM as a Cloudron app (~1 day). See [Cloudron packaging tutorial](https://docs.cloudron.io/packaging/tutorial/). Could also publish to the Cloudron community app store for others.
|
||||
|
||||
Sources: [Cloudron packaging docs](https://docs.cloudron.io/packaging/), [Cloudron manifest docs](https://docs.cloudron.io/packaging/manifest/), [Cloudron user directory (OIDC)](https://docs.cloudron.io/user-directory/)
|
||||
|
||||
### Alternatives to LiteLLM
|
||||
|
||||
| Tool | Per-user keys | Usage tracking | Rate limiting | Self-hosted? |
|
||||
|------|-------------|---------------|--------------|-------------|
|
||||
| **LiteLLM** | ✅ | ✅ Per-user, per-model | ✅ | ✅ Docker |
|
||||
| **Open WebUI built-in** | ✅ | ✅ Via token tracking pipes | Partial (in development, [#23323](https://github.com/open-webui/open-webui/issues/23323)) | ✅ (already in Cloudron) |
|
||||
| **Morph LLM Proxy** | ✅ | ✅ | ✅ | ✅ |
|
||||
| **Custom middleware** | Build it | Build it | Build it | ✅ |
|
||||
|
||||
Open WebUI's built-in token tracking is improving and may eventually handle per-user limits natively. For the MVP, LiteLLM is the more complete solution. As Open WebUI's features mature, the cooperative could simplify by dropping LiteLLM and using Open WebUI's native tracking.
|
||||
|
||||
- [LiteLLM](https://www.litellm.ai/) — open-source AI gateway
|
||||
- [LiteLLM docs](https://docs.litellm.ai/)
|
||||
- [Open WebUI token tracking](https://github.com/dartmouth/openwebui-token-tracking)
|
||||
- [Open WebUI per-user limits discussion](https://github.com/open-webui/open-webui/issues/23323)
|
||||
|
||||
## Component 8: Privacy and confidential inference
|
||||
|
||||
A cooperative handling member data has a stronger duty of care than a corporate provider — members own the cooperative, so the cooperative should not be able to surveil them. This is a structural advantage over corporate AI if we get it right.
|
||||
|
||||
### Privacy tiers (progressive)
|
||||
|
||||
The privacy model should be designed in layers, each stronger than the last:
|
||||
|
||||
**Tier 1: Self-hosted data sovereignty (MVP baseline)**
|
||||
- All data stays on cooperative infrastructure — no third-party access
|
||||
- Open WebUI stores conversation history in its own database (SQLite or PostgreSQL)
|
||||
- **Encryption at rest:** Use SQLCipher for SQLite or disk-level encryption for PostgreSQL
|
||||
- **Configurable retention:** Open WebUI supports automatic chat deletion after X days (feature request #21027, in development)
|
||||
- **No training on member data:** Models are pre-trained; member conversations are never used for fine-tuning
|
||||
- Sources: [Open WebUI security docs](https://docs.openwebui.com/enterprise/security/), [Open WebUI hardening guide](https://docs.openwebui.com/getting-started/advanced-topics/hardening/)
|
||||
|
||||
**Tier 2: Member-controlled retention + browser-local chat history**
|
||||
|
||||
Two complementary approaches:
|
||||
|
||||
*Option A: Open WebUI with server-side retention controls*
|
||||
- Members choose their own data retention: session-only, 7 days, 30 days, permanent
|
||||
- Session-only mode: conversations deleted when the session ends — no persistent storage at all
|
||||
- API access: no conversation logging for API users, only usage metering (token counts)
|
||||
- This requires a small configuration layer on top of Open WebUI's existing retention settings
|
||||
|
||||
*Option B: LobeChat in browser-local mode (Duck.ai-style)*
|
||||
|
||||
**LobeChat** is the better architecture for the Duck.ai approach. It has two deployment modes:
|
||||
- **Standalone/Lite mode:** All conversations stored in browser's IndexedDB. No server-side database. Chat history never leaves the user's device. This is exactly the Duck.ai model.
|
||||
- **Database mode:** PostgreSQL server-side storage with multi-user support for when persistence/sync is needed.
|
||||
|
||||
LobeChat features relevant to inference.coop:
|
||||
- **Browser-local storage:** IndexedDB — conversations stay on the user's device, not on cooperative servers
|
||||
- **CRDT-based sync:** Experimental multi-device sync using Conflict-Free Replicated Data Types — no central server needed
|
||||
- **Multi-provider support:** Connect to Ollama, vLLM, OpenAI, Anthropic, OpenRouter from one interface
|
||||
- **OpenAI-compatible:** Points at any OpenAI-compatible API endpoint — works with our inference backend
|
||||
- **Plugin system:** Extensible with web search, data tools, MCP (Model Context Protocol)
|
||||
- **PWA:** Works as a Progressive Web App on mobile and desktop
|
||||
- **MIT licensed, self-hostable, Docker deployment**
|
||||
- 55K+ GitHub stars
|
||||
|
||||
This means members can use LobeChat in standalone mode pointed at inference.coop's API endpoint — their chat history lives only in their browser, the cooperative never stores conversation content, and the cooperative's servers only see transient inference requests.
|
||||
|
||||
**Duck.ai's architecture** (for reference):
|
||||
- Chat history stored locally in browser by default — no server-side storage
|
||||
- Optional "Sync & Backup" uses end-to-end encrypted server storage (DuckDuckGo never holds decryption keys)
|
||||
- 18-month automatic deletion for synced chats not accessed
|
||||
- Up to 100 recent conversations stored locally
|
||||
- Notably, Duck.ai now uses **Tinfoil.sh** to host some models — the same TEE-based confidential inference platform relevant to our Tier 3 design
|
||||
- Sources: [DuckDuckGo help: recent chats](https://duckduckgo.com/duckduckgo-help-pages/duckai/recent-chats), [DuckDuckGo help: AI chat privacy](https://duckduckgo.com/duckduckgo-help-pages/duckai/ai-chat-privacy), [Factually.co: Duck.ai storage analysis](https://factually.co/fact-checks/technology/duck-ai-data-storage-fire-button-delete-chat-history-23ba86)
|
||||
|
||||
**Recommendation:** Offer both interfaces:
|
||||
- **Open WebUI** for members who want server-side chat history (with retention controls) and multi-user features
|
||||
- **LobeChat in standalone mode** for members who want Duck.ai-style browser-local privacy
|
||||
|
||||
Both point at the same inference API. Members choose their privacy posture by choosing their frontend.
|
||||
|
||||
- [LobeChat](https://github.com/lobehub/lobe-chat) — 55K+ stars, MIT license, browser-local mode
|
||||
- [LobeChat docs: local database](https://lobehub.com/docs/self-hosting/advanced/server-database)
|
||||
|
||||
**Tier 3: Encrypted inference via TEEs (Confer.to model)**
|
||||
|
||||
This is where it gets interesting — and where Confer.to's architecture becomes directly relevant.
|
||||
|
||||
### What Confer.to does
|
||||
|
||||
Confer.to (founded by Moxie Marlinspike, Signal founder) provides end-to-end encrypted AI chat:
|
||||
1. **Client-side encryption:** User's prompt is encrypted with WebAuthn passkeys *before* it leaves the browser
|
||||
2. **Trusted Execution Environment (TEE):** The encrypted prompt enters a hardware-enforced enclave (TEE) on the server. The server's OS cannot read the TEE's memory or execution state. The TEE decrypts the prompt, runs inference, encrypts the output, and sends it back.
|
||||
3. **Remote attestation:** Users can cryptographically verify that the TEE is running the claimed code — the cooperative *cannot* read conversations even if it wants to.
|
||||
4. **Statelessness:** LLMs are stateless (input in, output out) — making them ideal for TEE deployment since no state persists after the request.
|
||||
|
||||
- Source: [Confer.to blog: Private inference (Jan 2026)](https://confer.to/blog/2026/01/private-inference/)
|
||||
- Source: [Neurotechnus: AI privacy — Confer](https://neurotechnus.com/2026/01/19/ai-privacy-moxie-confer/)
|
||||
|
||||
### The TEE landscape in 2026
|
||||
|
||||
Confidential computing for AI inference matured significantly in 2026:
|
||||
|
||||
| Technology | Hardware | Status | Overhead |
|
||||
|-----------|----------|--------|---------|
|
||||
| **NVIDIA Confidential Computing** | H100, H200, B200 (GPU TEE) | Production-ready. CC-On mode encrypts GPU memory, protected PCIe between CPU and GPU. | <7% throughput |
|
||||
| **Intel TDX** | Intel CPUs (Trust Domain Extensions) | Mature. Creates encrypted VMs. | Low |
|
||||
| **AMD SEV-SNP** | AMD EPYC CPUs | Mature. Encrypted virtualization. | Low |
|
||||
| **AWS Nitro Enclaves** | AWS cloud | Available. Isolated compute environments. | Medium |
|
||||
|
||||
Key developments:
|
||||
- NVIDIA H100/H200 Confidential Computing went production in 2025-2026, enabling GPU-accelerated TEE inference
|
||||
- TEE overhead dropped below 7% — negligible for most workloads
|
||||
- Open-source frameworks emerging: **OpenPCC** (open-source confidential LLM serving on commodity TEEs, using vLLM, arXiv June 2026)
|
||||
- **Tinfoil** (tinfoil.sh) — open-source, OpenAI-compatible confidential inference API running in hardware enclaves. Collaborating with Red Hat. Production-ready.
|
||||
- Sources: [GeniusTechLab (Jul 2026)](https://geniustechlab.com/posts/2026-07-21-confidential-computing-tee-ai-inference-2026), [OpenPCC (arXiv Jun 2026)](https://arxiv.org/abs/2606.11145), [Tinfoil](https://tinfoil.sh/)
|
||||
|
||||
### How inference.coop could build on this
|
||||
|
||||
**Phase 1 (MVP):** Tier 1 + Tier 2. Self-hosted, encrypted at rest, member-controlled retention. The cooperative *could* technically read conversations, but has a policy commitment not to. This is already stronger than OpenAI/Anthropic, whose terms allow them to read your data.
|
||||
|
||||
**Phase 2 (growth):** Tier 3 with TEEs. When deploying on GPU hardware (for the vLLM migration), choose H100 GPUs with Confidential Computing enabled. Use a framework like OpenPCC or Tinfoil to serve models inside TEEs. At this point, the cooperative *cannot* read conversations even if it wants to — the privacy guarantee becomes architectural, not just policy-based.
|
||||
|
||||
**Phase 3 (long-term):** Client-side encryption + TEE (full Confer.to model). User's browser encrypts the prompt before transmission. Only the TEE can decrypt it. Even the network operator can't see the content. This requires custom frontend work on top of Open WebUI, but the architecture is proven by Confer.to.
|
||||
|
||||
### Tinfoil as a potential shortcut
|
||||
|
||||
[Tinfoil](https://tinfoil.sh/) is particularly interesting for the cooperative:
|
||||
- Open-source, OpenAI-compatible API
|
||||
- Runs models inside hardware enclaves
|
||||
- Verifiably private — users can attest the enclave
|
||||
- Collaborating with Red Hat on open-source confidential AI infrastructure
|
||||
- Could potentially be used as a backend instead of (or alongside) vLLM, providing TEE-protected inference without building the confidential computing layer from scratch
|
||||
|
||||
If Tinfoil's API is OpenAI-compatible, it could be a drop-in replacement for the vLLM backend, providing both inference *and* privacy in one package. Worth investigating for the Phase 2 migration.
|
||||
|
||||
### Privacy as cooperative advantage
|
||||
|
||||
This is the key strategic point: **privacy is where the cooperative can structurally outperform corporate AI providers.**
|
||||
|
||||
| | Corporate AI (OpenAI, Anthropic) | inference.coop Tier 1 | inference.coop Tier 3 (TEE) |
|
||||
|---|---|---|---|
|
||||
| Can the provider read your data? | Yes (per terms of service) | Technically yes, but policy prohibits | **No — architecturally impossible** |
|
||||
| Can the provider train on your data? | Yes (unless you pay for enterprise) | **No — never** | **No — never** |
|
||||
| Can you verify the privacy claim? | No (trust the company) | No (trust the cooperative) | **Yes — remote attestation** |
|
||||
| Data leaves your control? | Yes (sent to corporate servers) | No (stays on co-op infra) | No (encrypted before leaving your device) |
|
||||
|
||||
The cooperative's structural advantage is that it *can* deploy TEEs and offer verifiable privacy without a business model that depends on reading user data. OpenAI and Anthropic *could* deploy TEEs, but their business models (training on user data, targeting, enterprise upsells) create a conflict of interest. The cooperative has no such conflict.
|
||||
|
||||
- Source: [Confer.to blog](https://confer.to/blog/2026/01/private-inference/)
|
||||
- Source: [Confidential Inference Directory](https://confidentialinference.net/) — compares 7+ TEE-based AI providers
|
||||
|
||||
---
|
||||
|
||||
## Architecture summary (MVP)
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────┐
|
||||
│ Member's browser │
|
||||
│ Chat: https://chat.inference.coop (Open WebUI) │
|
||||
│ Chat: https://chat.inference.coop (LobeChat) │
|
||||
│ API: https://api.inference.coop/v1/... │
|
||||
│ Gov: https://gov.inference.coop (Loomio) │
|
||||
│ Pay: https://opencollective.com/inference-coop│
|
||||
└──────────────────────┬──────────────────────────┘
|
||||
│ HTTPS + Cloudron SSO (OIDC)
|
||||
┌──────────────────────▼──────────────────────────┐
|
||||
│ Cloudron (VPS, green energy) │
|
||||
│ ┌─────────┐ ┌──────────┐ ┌─────────────────┐ │
|
||||
│ │ SSO/OIDC│ │ Open WebUI│ │ Loomio │ │
|
||||
│ └─────────┘ └────┬─────┘ └─────────────────┘ │
|
||||
│ ┌────▼─────┐ │
|
||||
│ │ LiteLLM │ ← custom Cloudron app │
|
||||
│ │ Gateway │ (SSO, sandboxed) │
|
||||
│ └────┬─────┘ │
|
||||
│ ┌──────────┬───┴──────┬───────────┐ │
|
||||
│ │ │ │ │ │
|
||||
│ ┌──▼──┐ ┌──▼───┐ ┌──▼──┐ ┌──▼──────┐ │
|
||||
│ │Tinfoil│ │Ollama│ │Open │ │Billing │ │
|
||||
│ │(TEE) │ │(free)│ │Router│ │middleware│ │
|
||||
│ └───────┘└──────┘ └─────┘ └─────────┘ │
|
||||
│ (cloud) (Cloudron) (cloud) (custom webhook) │
|
||||
└──────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Green hosting:** Hetzner (100% hydropower) or Scaleway DC5 (100% renewable)
|
||||
**Privacy:** Tinfoil TEE for inference; LobeChat browser-local for chat history
|
||||
**SSO:** Cloudron OIDC — one login for chat, governance, API admin
|
||||
**Billing:** Open Collective → webhook middleware → LiteLLM key enable/disable
|
||||
|
||||
---
|
||||
|
||||
## Cost comparison: self-hosting vs cloud inference (10 users, ~12M tokens/month)
|
||||
|
||||
Assumptions: 10 user-testers, ~20 queries/day each, ~2000 tokens per conversation = ~400K tokens/day = ~12M tokens/month.
|
||||
|
||||
| Option | Cost/mo | Green? | TEE? | Self-hosted? | Inference quality |
|
||||
|--------|---------|--------|------|-------------|-------------------|
|
||||
| **Tinfoil API (GPT-OSS 120B)** | **~$18** | Unknown | ✅ | No | Full open models, fast |
|
||||
| **Tinfoil API (DeepSeek V4 Flash)** | **~$20** | Unknown | ✅ | No | 304B MoE, 1M context |
|
||||
| Hetzner CPU + Ollama + OpenRouter | €38 | ✅ Hydro | No | ✅ | 7B only on CPU (slow, 5-15 tok/s) |
|
||||
| Ollama Cloud Team (5 seats min) | ~$140 | Unknown | No | No | Full model library |
|
||||
| Scaleway L4 on-demand (8hrs/day) | ~$214 | ✅ Wind+hydro | No | ✅ | 7B-13B on GPU (fast) |
|
||||
| Hetzner L4 GPU + Ollama | €222 | ✅ Hydro | No | ✅ | 7B-13B on GPU (fast) |
|
||||
| Ollama Cloud Pro (10× individual) | ~$215 | Unknown | No | No | Full model library |
|
||||
|
||||
**Tinfoil wins on price by 2-10x** — the cooperative pays only for actual tokens used, not for idle GPU time. At 12M tokens/month, Tinfoil costs $3.96 (GPT-OSS 120B) to $5.52 (DeepSeek V4 Flash) in inference charges, plus €15/month for Cloudron = ~$18-20/month total.
|
||||
|
||||
**Self-hosting on green energy costs 2-10x more** because you're paying for a GPU server 24/7 regardless of whether anyone is querying it. The cheapest green self-hosted option (Hetzner CPU, €38/month) is slow (7B models on CPU, 5-15 tokens/sec) and still costs more than Tinfoil.
|
||||
|
||||
Sources: [Tinfoil pricing](https://tinfoil.sh/models), [Hetzner GPU servers](https://www.hetzner.com/dedicated-root-server/matrix-gpu/), [Scaleway GPU pricing](https://www.scaleway.com/en/pricing/gpu/), [Ollama Cloud pricing](https://ollama.com/pricing)
|
||||
|
||||
---
|
||||
|
||||
## Green + TEE: the gap
|
||||
|
||||
There is **no option today that is both clearly green AND TEE-protected at MVP scale**:
|
||||
|
||||
| Approach | Green? | TEE? | Cost/mo | MVP viable? |
|
||||
|----------|--------|------|---------|-------------|
|
||||
| Hetzner (hydro) + Ollama | ✅ | ❌ | €38-222 | Yes |
|
||||
| Scaleway DC5 (wind+hydro) + Ollama | ✅ | ❌ | ~$214 | Yes |
|
||||
| Tinfoil API (TEE) | Unknown | ✅ | ~$18-20 | Yes |
|
||||
| IREN (100% renewable) + H100 TEE | ✅ | ✅ | ~$1,800 | No (overkill) |
|
||||
| Google Cloud A3 Confidential VM + H100 | ~67% carbon-free | ✅ | ~$1,900 | No (overkill) |
|
||||
|
||||
The closest to "both" at MVP scale is **Tinfoil** — verifiable TEE privacy at $18-20/month, but energy mix unknown. The cooperative can:
|
||||
|
||||
1. **Be transparent about the trade-off** — "our inference is TEE-private; we cannot yet verify it is green"
|
||||
2. **Offset the gap** — purchase renewable energy certificates (RECs) to cover estimated inference emissions
|
||||
3. **Migrate when scale justifies it** — at 50+ members, self-host on IREN (renewable + TEE) or deploy Tinfoil's open-source framework on H100 hardware at a green data center
|
||||
|
||||
The alternative — self-hosting on Hetzner (green, €38/month, no TEE) — costs only €18/month more but requires GPU management and gives weaker privacy (policy-based, not architectural). For a cooperative whose value proposition is member privacy, the TEE advantage is worth more than the green gap.
|
||||
|
||||
- [Tinfoil](https://tinfoil.sh/) — TEE inference, open-source
|
||||
- [IREN](https://www.iren.com/) — 100% renewable GPU cloud with H100 TEE capability
|
||||
- [Google Cloud Confidential Computing](https://cloud.google.com/security/products/confidential-computing) — H100 TEE, targeting 24/7 carbon-free by 2030
|
||||
- [Canonical: Ubuntu Confidential VMs on Google Cloud A3 with H100](https://canonical.com/blog/2025/03/27/ubuntu-confidential-vms-now-available-on-google-cloud-a3-with-nvidia-h100-gpus)
|
||||
|
||||
---
|
||||
|
||||
**Hybrid approach (Tinfoil + Cloudron + optional Ollama):**
|
||||
|
||||
| Component | Monthly cost | Notes |
|
||||
|-----------|-------------|-------|
|
||||
| VPS (Cloudron host) | €15-20 | 2-4 vCPU, 4-8GB RAM — no GPU needed |
|
||||
| Cloudron license | €15 | SSO, app management, security sandboxing |
|
||||
| Tinfoil Inference API | $5-20 | Pay-as-you-go, GPT-OSS 120B at $0.15/$0.60 per M tok |
|
||||
| OpenRouter (fallback) | $2-5 | For models Tinfoil doesn't offer |
|
||||
| Domain | ~€1 | inference.coop or similar |
|
||||
| Open Collective | 0-5% | Platform fee depends on fiscal host |
|
||||
| **Total** | **~€40-60/month** | No GPU server needed |
|
||||
|
||||
At 10 members paying €10/month = €100/month → **surplus of €40-60/month** for reinvestment.
|
||||
|
||||
*Note: If Tinfoil per-token costs grow with usage, the cooperative can add self-hosted Ollama (free) for lightweight models or rent GPU hardware for vLLM at the break-even point (~50 members).*
|
||||
|
||||
---
|
||||
|
||||
## What's missing / needs further design
|
||||
|
||||
1. **Legal structure** — Open Collective fiscal sponsorship for MVP; formal cooperative incorporation later
|
||||
2. **Data policy** — what's logged, retained, shared. FERPA/GDPR considerations. Member-controlled retention (Tier 2).
|
||||
3. **Model governance** — how members decide which models to add/remove
|
||||
4. **Sustainability metrics** — energy per query, CO₂ tracking (à la GreenPT)
|
||||
5. **API key management** — whether to use Open WebUI's built-in keys or a dedicated gateway (LiteLLM)
|
||||
6. **Migration path** — Ollama → vLLM when scale requires it
|
||||
7. **Agent support** — whether to offer agentic capabilities (like Hermes Agent) beyond chat
|
||||
8. **Confidential inference (Phase 2-3)** — TEE deployment via OpenPCC or Tinfoil when GPU hardware is available. Client-side encryption (Confer.to model) as long-term goal. This is the cooperative's structural privacy advantage.
|
||||
|
||||
---
|
||||
|
||||
## Connections
|
||||
|
||||
- [co/core](../organizations/cocore.md) — most architecturally similar project (ATProto, mutual credit)
|
||||
- [GreenPT](../organizations/greenpt.md) — renewable energy AI hosting precedent
|
||||
- [Public AI Switzerland](../organizations/public-ai-switzerland.md) — consumer-facing AI cooperative model
|
||||
- [Ollama](../solidarity-economies/assessments/ollama-e2c.md) — inference engine + potential E2C candidate
|
||||
- [OpenRouter](../organizations/openrouter.md) — model brokering for frontier access
|
||||
- [Campus AI Cooperative](../solidarity-economies/projects/campus-ai-cooperative.md) — larger-scale version of the same concept
|
||||
- [AI Potluck](../organizations/ai-potluck.md) — coalition-built public AI stack; see learnings below
|
||||
|
||||
## Learnings from AI Potluck
|
||||
|
||||
[AI Potluck](../organizations/ai-potluck.md) (aipotluck.org) is the closest large-scale analogue to inference.coop's goal — "AI made and owned by the public" — but built as a *coalition* rather than a *cooperative*. Its design choices are directly instructive:
|
||||
|
||||
1. **Provenance transparency as a product feature.** AI Potluck's commitment "every response shows its provenance: the model, the organization, the compute, the country" is a concrete, implementable norm. inference.coop could adopt a lighter version: surface which model served each response (already possible via LiteLLM) and which backend (Tinfoil vs. Ollama) — turning the co-op's transparency value into a visible feature rather than a policy statement.
|
||||
|
||||
2. **Anti-engagement design as a differentiator.** AI Potluck explicitly rejects the extractive engagement model ("AI that wants you to turn it off," "doesn't use your emotional state to extend the conversation," "doesn't sell you sycophancy"). This is a values stance inference.coop shares and could articulate more explicitly in its charter and product copy — a cooperative has no engagement-maximization incentive, which is a structural advantage worth naming.
|
||||
|
||||
3. **The "no single point of failure" resilience argument.** AI Potluck's core claim — distributed ownership means no single actor's withdrawal is fatal — is a *resilience* argument, not just a fairness one. inference.coop's cooperative model achieves a version of this (no single member owns the whole), but the framing is worth borrowing: resilience-through-distribution is a selling point for members, not just a governance nicety.
|
||||
|
||||
4. **Data ethics as a trust signal.** "We don't host models trained on stolen data... you can delete everything in two steps." inference.coop's privacy posture (TEE, no training on member data) is stronger than AI Potluck's, but the *communication* of it — simple, concrete, verifiable — is a model to follow.
|
||||
|
||||
5. **The ownership-model contrast.** AI Potluck is a *coalition* (no single owner, no formal entity); inference.coop is a *cooperative* (member-owned, formal bylaws). The key difference is **accountability**: a cooperative has a defined membership, voting, and a legal entity that can be held responsible; a coalition does not. This is inference.coop's structural advantage over the potluck model — and worth articulating when explaining why "cooperative" rather than "coalition."
|
||||
|
||||
6. **Government-as-contributor framing.** AI Potluck reframes sovereign AI as *contribution* ("a government can own the parts that matter most to them without building everything else from scratch"). For inference.coop, this suggests a path for institutional members (universities, municipalities) to contribute compute or data without needing to own the whole — a membership tier worth considering.
|
||||
Reference in new issue
Block a user