Files
docs/models.md

9.7 KiB

Models

The co-op serves models from multiple providers. This is the single list — both the chat and the API draw from it, so it's the one place to look (and the one place to update).

Providers

Models are named provider/model-name, so you can choose not just which model but where it runs. Providers differ along two independent axes: privacy (how protected your prompts and responses are) and green energy (how the compute is powered). These are gradations, not binary switches — and no provider today is best on both.

Provider Value Privacy Green energy What that means
Tinfoil private Strongest — architectural Not a green claim Runs in hardware enclaves (TEEs) with end-to-end encryption (EHBP), verifiable via remote attestation. Neither we nor Tinfoil can read your prompts or responses at inference time. Energy mix is not disclosed.
GreenPT green Policy-based Strongest — 100% renewable Standard (non-enclave) hosting in the EU; privacy rests on GreenPT's commitments and EU data rules, not on cryptography. Compute runs on 100% renewable energy.
PublicAI public Policy-based Undisclosed A gateway for publicly developed, sovereign models — open-weight models built by national/public programs rather than private labs. Privacy is policy-based (no enclave); energy mix is not disclosed.

Reading the gradations:

  • Privacy: architectural (Tinfoil) beats policy-based (GreenPT, PublicAI). An enclave protects you even if the provider is compromised; a policy protects you only as far as the provider keeps its word. Both are better than a provider with neither.
  • Green energy: 100% renewable (GreenPT) beats undisclosed (Tinfoil, PublicAI). GreenPT reports its energy mix; the others do not, so we can't claim greenness for them.
  • Sovereignty: publicly developed (PublicAI) is its own value — the models come from public/national AI programs, built as public goods rather than proprietary products. It's a third axis, independent of privacy and energy.

Trade-offs are real: the most private option (Tinfoil) isn't the greenest, and the greenest option (GreenPT) isn't the most private. The point of the co-op is that you choose which axis matters more for a given task.

How to pick, quick guide:

  • Everyday chat and drafting → a fast model: tinfoil/glm-5-3-flash, greenpt/glm-5.3-flash, or greenpt/green-l. These handle the bulk of what people use chat for, at a fraction of flagship cost.
  • Hard reasoning, coding, agentic work → a flagship: tinfoil/glm-5-3 or greenpt/glm-5.3 (same underlying model, different provider trade-off), or tinfoil/deepseek-v4-1-flash for long documents and tool-driven tasks.
  • Images in, text out (vision) → tinfoil/glm-5-3-flash, tinfoil/deepseek-v4-1-flash, or the Apertus models — see the ✅ in the table below.
  • Cheap and cheerful → tinfoil/gpt-oss-120b or greenpt/glm-5.3-flash.
  • Model origins matter to you → the Apertus models (publicly developed; see below).

Current models

Pricing is per 1 million tokens, input / output (what the co-op pays the provider — your own cost is your monthly allowance, not per-token billing). Vision means the model accepts images as input (screenshots, photos, documents to reason about) — all models output text. None of our current models generate images.

Model Provider Privacy Green Vision Context Pricing (in / out) Notes
tinfoil/glm-5-3-flash Tinfoil architectural undisclosed ✅ 1M $0.40 / $1.25 Fast multimodal MoE; great everyday default with image input.
tinfoil/glm-5-3 Tinfoil architectural undisclosed — 1M $1.80 / $5.75 GLM flagship; strongest agentic model, private tier.
tinfoil/deepseek-v4-1-flash Tinfoil architectural undisclosed ✅ 1M $0.65 / $1.45 Efficient MoE, image input, tool calling; great for long docs and agents.
tinfoil/gpt-oss-120b Tinfoil architectural undisclosed — 131K $0.15 / $0.60 Lightweight open-weight reasoning; cheapest private tier.
tinfoil/kimi-k3 Tinfoil architectural undisclosed ✅ (image + video) 262K $4.00 / $20.00 Moonshot's flagship multimodal MoE; premium tier.
greenpt/green-r GreenPT policy 100% renewable — 128K $0.35 / $0.95 Reasoning-focused, renewable energy.
greenpt/green-l GreenPT policy 100% renewable — 128K $0.25 / $0.80 Lightweight, renewable energy.
greenpt/glm-5.3-flash GreenPT policy 100% renewable — 1M $0.11 / $0.44 Fast GLM model, renewable energy.
greenpt/glm-5.3 GreenPT policy 100% renewable — 1M $1.10 / $4.40 GLM flagship — strongest agentic model in the green tier.
publicai/apertus-v1.5-8b PublicAI policy undisclosed ✅ 262K $0.10 / $0.20 Fully open sovereign model, small/fast, image understanding.
publicai/apertus-v1.5-70b PublicAI policy undisclosed ✅ 262K $0.82 / $2.92 Fully open sovereign model, large/capable, image understanding.

Vision is verified working through the co-op gateway (base64 image input via /v1/chat/completions, standard OpenAI format). In the chat UI, just attach an image to your message while using a ✅ model.

Same model, two providers: glm-5-3 (Tinfoil) and glm-5.3 (GreenPT) are the same Z.ai model offered by two providers — Tinfoil's is architecturally private at $1.80/$5.75; GreenPT's is 100% renewable at $1.10/$4.40. That price/privacy trade-off in a single comparison is the co-op's whole idea in miniature: you pick which axis matters for the task at hand.

Audio models

These are not chat models — they power the chat's voice mode (dictation and spoken replies) and are available through the API on /v1/audio/* routes for members building speech into their own tools. They're hidden from the chat model picker to avoid confusion.

Model Provider Use Pricing
tinfoil/whisper-large-v3-turbo Tinfoil Speech-to-text $0.05 / 1M input tokens
tinfoil/voxtral-tts Tinfoil Text-to-speech listed $0 (verify on invoice)

Both run inside Tinfoil's enclaves, so voice data gets the same architectural privacy as the private chat models. TTS voices available: neutral_female, neutral_male, casual_female/male, cheerful_female, plus French, German, Spanish, Italian, Portuguese, Dutch, and Hindi variants.

# Speech-to-text
curl https://gateway.inference.coop/v1/audio/transcriptions \
  -H "Authorization: Bearer ***" \
  -F file=@your-audio.mp3 \
  -F model=tinfoil/whisper-large-v3-turbo

# Text-to-speech (returns WAV audio)
curl https://gateway.inference.coop/v1/audio/speech \
  -H "Authorization: Bearer ***" \
  -H "Content-Type: application/json" \
  -d '{"model": "tinfoil/voxtral-tts", "input": "Hello from the co-op", "voice": "neutral_female"}' \
  --output speech.wav

# Vision (image input to a multimodal chat model)
curl https://gateway.inference.coop/v1/chat/completions \
  -H "Authorization: Bearer ***" \
  -H "Content-Type: application/json" \
  -d '{"model": "tinfoil/glm-5-3-flash", "messages": [{"role": "user", "content": [
        {"type": "text", "text": "What is in this image?"},
        {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}
      ]}]}'

About the Apertus models

Apertus is the Swiss AI Initiative's open foundation model, built by a public collaboration between EPFL, ETH Zurich, and the Swiss National Supercomputing Centre (CSCS). It's "fully open" in the strict sense — not just downloadable weights:

  • Open weights & code — released under the Apache-2.0 license, so anyone can use, modify, and redistribute it, commercially or not.
  • Open training data — trained only on fully open data (unlike most proprietary models, which keep their training set secret). The data, code, methods, and alignment principles are all published and reproducible.
  • Open values — governed by a public Apertus Charter that documents the model's values and principles rather than leaving them implicit.
  • Sovereign & public-good — designed to European data-protection norms (GDPR, EU AI Act, Swiss law) as an example of "AI as a public good," not a proprietary product. Trained across 1,500+ languages (40% non-English).

The v1.5 generation adds image understanding, a longer 262K-token context window, improved tool use, and instruction-following. It's the publicly developed counterpart to the private (Tinfoil) and renewable (GreenPT) offerings — sovereignty as a third axis of member choice.

Source: apertus-ai.org and the Swiss AI Hugging Face model cards.

Some models also have a discounted cached-input rate (when a prompt is reused and doesn't need to be re-read) — see the code/litellm config for the exact numbers. Pricing here mirrors the model_info in the LiteLLM config (code/litellm/config.yaml), which is the authoritative source — the two are kept in sync.

Limitations on privacy

Only Tinfoil offers architectural privacy. The other providers are chosen for their value (renewable energy, public models) but do not run in enclaves — prompts and responses pass through them in the ordinary way. Your chat history is also stored on our server so you can revisit it, and that stored history is not encrypted in a way that prevents us from technically reading it — we commit not to. The full distinction — what's architecturally private versus what's a policy commitment — is in the Privacy Policy.