Limits
Design your client to respect the selected model’s context and request constraints.
A request may be limited by the model, API key, account, wallet, request rate, or client configuration. When a limit is reached, first identify the layer that caused it.
Types of limits
| Layer | What to check |
|---|---|
| Model capability | Context, maximum output, input types, and media size or count |
| API key | Budget, expiration, model scope, per-request cost, and source restrictions |
| Account and wallet | Account status, available balance, and service availability |
| Request rate | Requests, tokens, or concurrency allowed in a time window |
| Client | Timeout, connections, retry count, and request queue |
Check before sending
- Get actual limit values from the model documentation.
- Count conversation history, current input, and reserved output together.
- Check media files and total request size.
- Confirm the API key budget, expiration, and model scope.
- Set reasonable client timeout, concurrency, and retry limits.
Do not apply one model's limits to every model.
Respond to a limit
| Problem | Recommended action |
|---|---|
| Input is too long | Remove unnecessary history, compress, or split the content |
| Output limit is invalid | Use a field and value supported by the model |
| Key expired or budget exhausted | Review the key settings and follow the console process |
| Too many requests | Queue work, lower the rate, and reduce concurrency |
| Model access is restricted | Review account, key, and model availability |
| Client timeout | Check the original request in Requests |
Adding funds does not necessarily increase rate limits and does not bypass model capabilities or API key policies.
Request a limit adjustment
When contacting support, provide the Model ID, expected request volume, concurrency, average input and output, business hours, and example Request IDs. A submitted request does not mean that the limit has changed; wait for an explicit result before relying on it.