diff --git a/infrastructure.md b/infrastructure.md index fbd3d33..0aaeaec 100644 --- a/infrastructure.md +++ b/infrastructure.md @@ -158,3 +158,49 @@ through it; the other providers are called directly. Maintaining the co-op's infrastructure — rotating Cloudron tokens, editing the portal's environment, and troubleshooting the provisioning pipeline — is documented in the [operator runbook](operator-token-runbook.md). + +## Gateway version and security update policy + +The LiteLLM gateway is pinned to an **exact upstream tag** in +[`code/litellm`](https://git.inference.coop/code/litellm) (currently +`v1.84.0`). Two reasons: + +1. **CVE fix line.** A member security review (Oct 2026) found the prior pin + (`main-v1.74.0-stable`) sat below the fix line for two published + vulnerabilities: CVE-2026-35029 (authorization bypass on + `/config/update`, fixed in 1.83.0, with documented in-the-wild probing) + and CVE-2026-59822 (MCP session auth bypass, fixed in 1.84.0, listed in + the CISA Known Exploited Vulnerabilities catalog). `v1.84.0` is at or + above both fix lines. +2. **Floating-tag drift.** `main-v1.74.0-stable` is a *floating* tag — the + registry silently moved it to a newer build. A rebuild intended as a + no-op therefore pulled different bits and triggered a full schema + migration, which is how the 2026-10-02 outage happened (see below). + +Rule: **never pin a floating tag** (`main-*` / `latest`); always the exact +version tag. When upgrading, expect the first boot after the upgrade to take +10–15 minutes (prisma migration with internal retries) — plan the restart +window accordingly. + +CORS: the gateway answers browser preflights only for +`https://chat.inference.coop` (via LiteLLM's native +`LITELLM_CORS_ORIGINS`). The old `LITELLM_CORS_ALLOWED_ORIGINS` name +(our v1.74.0-era sed-patch variable) is dead — v1.84.0 reads the native +name only. + +## Incident log + +**2026-10-02 — gateway outage after security-upgrade rebuild.** A Cloudron +update triggered by the CVE report rebuilt the app image; the floating +`main-v1.74.0-stable` tag had moved to newer bits (v1.84.0-era), so the +rebuild changed the running version and kicked off a full prisma DB +migration that takes 10–15 minutes per boot with internal timeouts/retries. +The gateway was down ~50 minutes (04:50–05:34 UTC); no data loss. Repairs: +version pinned to the exact tag in the repo (PR +[code/litellm#1](https://git.inference.coop/code/litellm/pulls/1) — pending +merge), `LITELLM_CORS_ORIGINS` set as a Cloudron env var (the upgrade had +silently reverted CORS to `*`), packaging synced to mirror the live +deployment. Also during the incident: the domain's authoritative +nameservers (yoursrs.com) were intermittently unreachable, briefly failing +resolution of some hostnames (git, portal) — transient registrar-side +issue, self-recovered, no records were wrong.