Add API access docs; fix model name to DeepSeek V4.1 Flash

This commit is contained in:
inference-bot committed 2026-09-14 16:34:24 -06:00
1 parent 20e76b4edf
commit 18ad848ef1
1 file changed
+20 -1
+20 -1
View File
@@ -58,7 +58,7 @@ The cooperative's software is open source and lives in the [`code`](https://git.
The gateway currently exposes three models (all served through Tinfoil's TEE-protected enclaves): The gateway currently exposes three models (all served through Tinfoil's TEE-protected enclaves):
- **DeepSeek V4 Flash** — default model, best for agentic tasks (1M context, tool calling) - **DeepSeek V4.1 Flash** — default model, best for agentic tasks (1M context, tool calling)
- **GPT-OSS 120B** — lightweight fallback - **GPT-OSS 120B** — lightweight fallback
- **GLM-5.3 Flash** — fast, efficient MoE model - **GLM-5.3 Flash** — fast, efficient MoE model
@@ -67,6 +67,25 @@ The gateway currently exposes three models (all served through Tinfoil's TEE-pro
- **Web search** — enabled by default, backed by a self-hosted SearXNG instance (no external search API or scraper key required). - **Web search** — enabled by default, backed by a self-hosted SearXNG instance (no external search API or scraper key required).
- **File uploads (RAG)** — works out of the box via Open WebUI's bundled local embedding model (`sentence-transformers/all-MiniLM-L6-v2`), no external embedding API. - **File uploads (RAG)** — works out of the box via Open WebUI's bundled local embedding model (`sentence-transformers/all-MiniLM-L6-v2`), no external embedding API.
### API access
Members can use the co-op's models programmatically through an OpenAI-compatible API:
- **Endpoint:** `https://gateway.inference.coop/v1`
- **Auth:** a bearer API key (created in the [Member Dashboard](https://dashboard.inference.coop))
- **Format:** standard OpenAI chat completions (`POST /v1/chat/completions`)
Example:
```bash
curl https://gateway.inference.coop/v1/chat/completions \
-H "Authorization: Bearer sk-..." \
-H "Content-Type: application/json" \
-d '{"model": "deepseek-v4-1-flash", "messages": [{"role": "user", "content": "Hello"}]}'
```
API usage draws from the **same monthly balance as chat** — there is no separate quota. The gateway is publicly reachable (member keys authenticate it), but its admin UI and API documentation are disabled; admin access is via the internal dashboard or the gateway's admin endpoints.
### Privacy ### Privacy
Privacy is a core value. Inference now runs through [Tinfoil](https://tinfoil.sh/inference), which provides **architectural** privacy: models run inside hardware enclaves (TEEs), and request/response bodies are encrypted end-to-end with the Encrypted HTTP Body Protocol (EHBP), so even Tinfoil's own infrastructure cannot read them. This is verifiable via remote attestation — not just a policy promise. Privacy is a core value. Inference now runs through [Tinfoil](https://tinfoil.sh/inference), which provides **architectural** privacy: models run inside hardware enclaves (TEEs), and request/response bodies are encrypted end-to-end with the Encrypted HTTP Body Protocol (EHBP), so even Tinfoil's own infrastructure cannot read them. This is verifiable via remote attestation — not just a policy promise.