inference-bot
2369c3e848
Add greenpt/glm-5.3-flash (/usr/bin/bash.11//usr/bin/bash.44 per 1M)
2026-09-24 10:43:14 -06:00
inference-bot
f86e005062
Rename Tinfoil models to tinfoil/* provider prefix
2026-09-24 08:16:16 -06:00
inference-bot
67192de232
Document pricing sync with co-op/docs/models.md
2026-09-23 21:08:33 -06:00
inference-bot
09ee36014c
Add GreenPT models (greenpt/green-r, greenpt/green-l) via provider/model naming convention
2026-09-23 20:10:18 -06:00
inference-bot
d1f57c45d5
Retry connection errors but never re-bill timeouts
...
num_retries=2 (retries transient connection errors), retry_policy with
TimeoutErrorRetries=0 (a timed-out generation is already billed by Tinfoil, so
re-hitting it re-bills). Combined with timeout=180, this hedges connection
errors without the silent re-billing that caused the Tinfoil spend gap.
2026-09-22 18:18:07 -06:00
inference-bot
576885ed21
Fix Tinfoil spend gap: raise timeout 30->180s, num_retries 2->0
...
DeepSeek V4.1 Flash is a reasoning model; 30s timeout caused LiteLLM to
abandon in-flight generations (falling back to GPT-OSS) while Tinfoil still
completed and billed them. num_retries re-hit Tinfoil up to 3x. Result: Tinfoil
billed ~2x the tokens LiteLLM logged. Generous timeout + zero retries stops
the silent re-billing.
2026-09-22 18:05:27 -06:00
inference-bot
9cbbc33bc7
Fix CORS patch to target site-packages (the copy the running CLI imports)
2026-09-14 18:01:13 -06:00
inference-bot
0a797ba50e
Patch hardcoded origins=[*] to honour LITELLM_CORS_ALLOWED_ORIGINS (v1.74.0 ignores the env var)
2026-09-14 17:55:08 -06:00
inference-bot
dae9bf4f31
Lock CORS to chat origin (was wildcard * with allow-credentials)
2026-09-14 17:51:38 -06:00
inference-bot
fb921b6e09
Disable admin UI (/ui) and API docs (Swagger/ReDoc/openapi) on the public gateway
2026-09-14 16:21:34 -06:00
inference-bot
700154c3d4
Add custom per-token pricing (model_info) so spend and budget enforcement work
2026-09-14 14:28:13 -06:00
inference-bot
da50b0438f
Replace deprecated deepseek-v4-flash with deepseek-v4-1-flash (default + fallback)
2026-09-14 14:06:23 -06:00
inference-bot
878ab8fa8d
Add GLM-5.3 Flash model
2026-09-09 14:18:14 -06:00
inference-bot
8a54e75d4e
Download tinfoil-proxy at runtime (build sandbox has no GitHub access)
2026-09-09 13:23:53 -06:00
inference-bot
26b0edc6a8
Hardcode tinfoil-proxy version in Dockerfile
2026-09-09 13:22:21 -06:00
inference-bot
e05fec0ca8
Switch to Tinfoil (TEE): add tinfoil-proxy sidecar for EHBP encryption, point LiteLLM at local proxy
2026-09-09 13:21:01 -06:00
inference-bot
2ce5c66b9b
Switch to Ollama Cloud (OpenAI-compatible /v1 endpoint) for plumbing smoke test
2026-09-04 14:15:49 -06:00
inference-bot
41068a6264
Fix Dockerfile for Cloudron: run as root (base image default), override ENTRYPOINT, fix .dockerignore
2026-09-04 10:29:15 -06:00
inference-bot
b9b4620e52
Initial commit: LiteLLM Cloudron app package for Inference Cooperative
2026-09-03 21:55:54 -06:00