Files

168 lines
9.7 KiB
Markdown

# Models
The co-op serves models from multiple providers. This is the single list — both
the chat and the API draw from it, so it's the one place to look (and the one
place to update).
## Providers
Models are named `provider/model-name`, so you can choose not just *which* model
but *where* it runs. Providers differ along two independent axes: **privacy**
(how protected your prompts and responses are) and **green energy** (how the
compute is powered). These are gradations, not binary switches — and no
provider today is best on both.
| Provider | Value | Privacy | Green energy | What that means |
|---|---|---|---|---|
| **Tinfoil** | *private* | **Strongest — architectural** | Not a green claim | Runs in hardware enclaves (TEEs) with end-to-end encryption (EHBP), verifiable via remote attestation. Neither we nor Tinfoil can read your prompts or responses at inference time. Energy mix is not disclosed. |
| **GreenPT** | *green* | **Policy-based** | **Strongest — 100% renewable** | Standard (non-enclave) hosting in the EU; privacy rests on GreenPT's commitments and EU data rules, not on cryptography. Compute runs on 100% renewable energy. |
| **PublicAI** | *public* | Policy-based | Undisclosed | A gateway for **publicly developed, sovereign models** — open-weight models built by national/public programs rather than private labs. Privacy is policy-based (no enclave); energy mix is not disclosed. |
**Reading the gradations:**
- **Privacy:** *architectural* (Tinfoil) beats *policy-based* (GreenPT, PublicAI). An
enclave protects you even if the provider is compromised; a policy protects
you only as far as the provider keeps its word. Both are better than a
provider with neither.
- **Green energy:** *100% renewable* (GreenPT) beats *undisclosed* (Tinfoil,
PublicAI). GreenPT reports its energy mix; the others do not, so we can't
claim greenness for them.
- **Sovereignty:** *publicly developed* (PublicAI) is its own value — the models
come from public/national AI programs, built as public goods rather than
proprietary products. It's a third axis, independent of privacy and energy.
Trade-offs are real: the most private option (Tinfoil) isn't the greenest, and
the greenest option (GreenPT) isn't the most private. The point of the co-op is
that **you choose** which axis matters more for a given task.
**How to pick, quick guide:**
- **Everyday chat and drafting** → a fast model: `tinfoil/glm-5-3-flash`,
`greenpt/glm-5.3-flash`, or `greenpt/green-l`. These handle the bulk of what
people use chat for, at a fraction of flagship cost.
- **Hard reasoning, coding, agentic work** → a flagship: `tinfoil/glm-5-3` or
`greenpt/glm-5.3` (same underlying model, different provider trade-off),
or `tinfoil/deepseek-v4-1-flash` for long documents and tool-driven tasks.
- **Images in, text out (vision)** → `tinfoil/glm-5-3-flash`,
`tinfoil/deepseek-v4-1-flash`, or the Apertus models — see the ✅ in the
table below.
- **Cheap and cheerful** → `tinfoil/gpt-oss-120b` or `greenpt/glm-5.3-flash`.
- **Model origins matter to you** → the Apertus models (publicly developed;
see below).
## Current models
Pricing is per 1 million tokens, input / output (what the co-op pays the
provider — your own cost is your monthly allowance, not per-token billing).
**Vision** means the model accepts images as input (screenshots, photos,
documents to reason about) — all models output text. None of our current
models generate images.
| Model | Provider | Privacy | Green | Vision | Context | Pricing (in / out) | Notes |
|---|---|---|---|---|---|---|---|
| `tinfoil/glm-5-3-flash` | Tinfoil | architectural | undisclosed | ✅ | 1M | $0.40 / $1.25 | Fast multimodal MoE; great everyday default with image input. |
| `tinfoil/glm-5-3` | Tinfoil | architectural | undisclosed | — | 1M | $1.80 / $5.75 | GLM flagship; strongest agentic model, private tier. |
| `tinfoil/deepseek-v4-1-flash` | Tinfoil | architectural | undisclosed | ✅ | 1M | $0.65 / $1.45 | Efficient MoE, image input, tool calling; great for long docs and agents. |
| `tinfoil/gpt-oss-120b` | Tinfoil | architectural | undisclosed | — | 131K | $0.15 / $0.60 | Lightweight open-weight reasoning; cheapest private tier. |
| `tinfoil/kimi-k3` | Tinfoil | architectural | undisclosed | ✅ (image + video) | 262K | $4.00 / $20.00 | Moonshot's flagship multimodal MoE; premium tier. |
| `greenpt/green-r` | GreenPT | policy | 100% renewable | — | 128K | $0.35 / $0.95 | Reasoning-focused, renewable energy. |
| `greenpt/green-l` | GreenPT | policy | 100% renewable | — | 128K | $0.25 / $0.80 | Lightweight, renewable energy. |
| `greenpt/glm-5.3-flash` | GreenPT | policy | 100% renewable | — | 1M | $0.11 / $0.44 | Fast GLM model, renewable energy. |
| `greenpt/glm-5.3` | GreenPT | policy | 100% renewable | — | 1M | $1.10 / $4.40 | GLM flagship — strongest agentic model in the green tier. |
| `publicai/apertus-v1.5-8b` | PublicAI | policy | undisclosed | ✅ | 262K | $0.10 / $0.20 | Fully open sovereign model, small/fast, image understanding. |
| `publicai/apertus-v1.5-70b` | PublicAI | policy | undisclosed | ✅ | 262K | $0.82 / $2.92 | Fully open sovereign model, large/capable, image understanding. |
Vision is verified working through the co-op gateway (base64 image input via
`/v1/chat/completions`, standard OpenAI format). In the chat UI, just attach
an image to your message while using a ✅ model.
**Same model, two providers:** `glm-5-3` (Tinfoil) and `glm-5.3` (GreenPT)
are the same Z.ai model offered by two providers — Tinfoil's is
architecturally private at $1.80/$5.75; GreenPT's is 100% renewable at
$1.10/$4.40. That price/privacy trade-off in a single comparison is the
co-op's whole idea in miniature: you pick which axis matters for the task
at hand.
### Audio models
These are **not chat models** — they power the chat's voice mode (dictation
and spoken replies) and are available through the API on `/v1/audio/*` routes
for members building speech into their own tools. They're hidden from the chat
model picker to avoid confusion.
| Model | Provider | Use | Pricing |
|---|---|---|---|
| `tinfoil/whisper-large-v3-turbo` | Tinfoil | Speech-to-text | $0.05 / 1M input tokens |
| `tinfoil/voxtral-tts` | Tinfoil | Text-to-speech | listed $0 (verify on invoice) |
Both run inside Tinfoil's enclaves, so voice data gets the same architectural
privacy as the private chat models. TTS voices available: `neutral_female`,
`neutral_male`, `casual_female/male`, `cheerful_female`, plus French, German,
Spanish, Italian, Portuguese, Dutch, and Hindi variants.
```bash
# Speech-to-text
curl https://gateway.inference.coop/v1/audio/transcriptions \
-H "Authorization: Bearer ***" \
-F file=@your-audio.mp3 \
-F model=tinfoil/whisper-large-v3-turbo
# Text-to-speech (returns WAV audio)
curl https://gateway.inference.coop/v1/audio/speech \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{"model": "tinfoil/voxtral-tts", "input": "Hello from the co-op", "voice": "neutral_female"}' \
--output speech.wav
# Vision (image input to a multimodal chat model)
curl https://gateway.inference.coop/v1/chat/completions \
-H "Authorization: Bearer ***" \
-H "Content-Type: application/json" \
-d '{"model": "tinfoil/glm-5-3-flash", "messages": [{"role": "user", "content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}
]}]}'
```
### About the Apertus models
Apertus is the **Swiss AI Initiative's** open foundation model, built by a
public collaboration between **EPFL, ETH Zurich, and the Swiss National
Supercomputing Centre (CSCS)**. It's "fully open" in the strict sense — not just
downloadable weights:
- **Open weights & code** — released under the **Apache-2.0** license, so anyone
can use, modify, and redistribute it, commercially or not.
- **Open training data** — trained only on *fully open* data (unlike most
proprietary models, which keep their training set secret). The data, code,
methods, and alignment principles are all published and reproducible.
- **Open values** — governed by a public **Apertus Charter** that documents the
model's values and principles rather than leaving them implicit.
- **Sovereign & public-good** — designed to European data-protection norms
(GDPR, EU AI Act, Swiss law) as an example of "AI as a public good," not a
proprietary product. Trained across 1,500+ languages (40% non-English).
The v1.5 generation adds image understanding, a longer 262K-token context
window, improved tool use, and instruction-following. It's the *publicly
developed* counterpart to the private (Tinfoil) and renewable (GreenPT)
offerings — sovereignty as a third axis of member choice.
*Source: [apertus-ai.org](https://www.apertus-ai.org/) and the
[Swiss AI Hugging Face](https://huggingface.co/swiss-ai) model cards.*
Some models also have a discounted cached-input rate (when a prompt is reused
and doesn't need to be re-read) — see the `code/litellm` config for the exact
numbers. Pricing here mirrors the `model_info` in the LiteLLM config
(`code/litellm/config.yaml`), which is the authoritative source — the two are
kept in sync.
## Limitations on privacy
Only **Tinfoil** offers architectural privacy. The other providers are chosen
for their value (renewable energy, public models) but do not run in enclaves —
prompts and responses pass through them in the ordinary way. Your **chat
history** is also stored on our server so you can revisit it, and that stored
history is not encrypted in a way that prevents us from technically reading it —
we commit not to. The full distinction — what's architecturally private versus
what's a policy commitment — is in the [Privacy Policy](privacy-policy.md).