num_retries=2 (retries transient connection errors), retry_policy with
TimeoutErrorRetries=0 (a timed-out generation is already billed by Tinfoil, so
re-hitting it re-bills). Combined with timeout=180, this hedges connection
errors without the silent re-billing that caused the Tinfoil spend gap.
DeepSeek V4.1 Flash is a reasoning model; 30s timeout caused LiteLLM to
abandon in-flight generations (falling back to GPT-OSS) while Tinfoil still
completed and billed them. num_retries re-hit Tinfoil up to 3x. Result: Tinfoil
billed ~2x the tokens LiteLLM logged. Generous timeout + zero retries stops
the silent re-billing.