Rate limits
Rate limits protect shared capacity. They are per key and per account, per minute.
Quick path
- Stay under the limits for your tier.
- Watch
RateLimit-*headers and honorRetry-Afteron429. - Need more: top up history and Enterprise raises the ceiling. See Enterprise.
Details
| Topic | Decision |
|---|---|
| Dimensions | Requests per minute (RPM) and tokens per minute (TPM), enforced per key and per account. The lower of the two ceilings applies. |
| Default ceilings | Regular: 60 RPM / 200K TPM. Roleplay and Cybersecurity: same defaults; higher on request with usage history. Exact ceilings are returned in response headers. |
| Headers | RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset. On 429, Retry-After in seconds. |
| Over limit | 429 rate_limited. Back off; do not retry in a tight loop. See Errors. |
| SLA | Rate-limited requests are customer errors and are excluded from availability measurement. See the Service level agreement. |
Checklist
- [ ] Client reads
Retry-Afterand backs off with jitter. - [ ] Bursts are smoothed to stay under TPM.
- [ ] Sustained need is raised via Enterprise, not via extra accounts.
Next step
Handle failures cleanly: Errors.