> ## Documentation Index
> Fetch the complete documentation index at: https://docs.prisminference.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

> Understand Prism per-key rate limits and how to handle 429 responses.

Prism applies per-key rate limits. When a key exceeds it, the inference API
returns `429 Too Many Requests` with a `Retry-After` header.

## Handle a 429

```http theme={"theme":{"light":"github-light","dark":"github-dark"}}
HTTP/1.1 429 Too Many Requests
Retry-After: 12
```

```json theme={"theme":{"light":"github-light","dark":"github-dark"}}
{
  "error": {
    "message": "API key rate limit exceeded.",
    "type": "rate_limit_error",
    "param": null,
    "code": "rate_limit_exceeded",
    "retryable": true,
    "fix": "Wait for the Retry-After interval, then retry with bounded exponential backoff.",
    "docs_url": "https://docs.prisminference.com/rate-limits"
  }
}
```

* `Retry-After` is in seconds. Wait that long before the next attempt.
* Use `retryable` as the primary retry signal. `429` responses are retryable.
* If `Retry-After` is missing, use bounded exponential backoff with jitter.

```typescript theme={"theme":{"light":"github-light","dark":"github-dark"}}
const delay = (attempt: number, retryAfter: string | null): number => {
  if (retryAfter != null) {
    const seconds = Number(retryAfter);
    if (Number.isFinite(seconds)) {
      return Math.max(0, seconds * 1000);
    }
  }

  const exponential = Math.min(30_000, 500 * 2 ** attempt);
  return Math.random() * exponential;
};
```

Immediate retries add load and usually fail again. Cap the number of attempts and
surface the final error to the caller.

## Separate the limits you can hit

A `429` can also mean the model is at capacity rather than the key being
rate-limited. Both return `429`; both are retryable after the `Retry-After`
interval. See [Errors](/errors#rate-limits) for the full status-code table and
retry policy, and [Streaming](/guides/streaming) for retrying interrupted
streams.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.