How limits work
Each rate-limited model has a limit for your organization and a shared global limit. Each holds at most one minute of its configured token allowance and refills continuously. For each request subject to these limits, Sail counts input tokens before inference starts:- Cached input reduces the count using the larger of a known Supercache read and Sail’s prefix-cache estimate. The two values are not added together.
- A Supercache write counts its full input before any completion-window discount. It does not receive a cache-read deduction.
- A configured completion-window discount reduces the remaining count, rounded up to a whole token. Windows share the same limits for the model and your organization.
- Final inference usage does not revise the rate-limit count.
When a request is limited
Sail returns:429when your organization reaches its limit for the model.529when the model reaches its global limit.
Retry-After in seconds. If retrying, wait at least that
long. This does not guarantee that the next attempt will succeed.
A rejected request does not start inference. If you sent an idempotency key,
reuse it when retrying. See Idempotency for examples.
Requests larger than the token allowance
A request whose counted input exceeds either limit’s entire one-minute allowance cannot pass that limit, even when it is full. Sail returnsRetry-After: 60 for this case, but waiting alone does not make the request
eligible. Reduce its counted input or contact support about the limit.
Temporary model availability
Some models can be temporarily unavailable even when you have not reached a rate limit. For non-streaming requests and errors detected before a stream begins, Sail returns503 with a Retry-After header. OpenAI-compatible
endpoints also return error code model_capacity_unavailable.
This response has no reset window. Wait at least the stated time, then retry
with backoff. The rejected request does not start inference.