Skip to main content
Prism applies per-key rate limits. When a key exceeds it, the inference API returns 429 Too Many Requests with a Retry-After header.

Handle a 429

  • Retry-After is in seconds. Wait that long before the next attempt.
  • Use retryable as the primary retry signal. 429 responses are retryable.
  • If Retry-After is missing, use bounded exponential backoff with jitter.
Immediate retries add load and usually fail again. Cap the number of attempts and surface the final error to the caller.

Separate the limits you can hit

A 429 can also mean the model is at capacity rather than the key being rate-limited. Both return 429; both are retryable after the Retry-After interval. See Errors for the full status-code table and retry policy, and Streaming for retrying interrupted streams.
Last modified on October 6, 2026