Skip to main content
Sail limits how quickly your organization can send inference requests and how many can run at once. API keys in the same organization share these limits. Available capacity can vary during periods of high demand. Pro and Enterprise customers receive increased access to model capacity and higher rate limits, with Enterprise offering the highest capacity.

429: Too many requests

Your organization has reached a request or concurrency limit. Send fewer requests or run fewer at once. Wait for the delay in Retry-After before retrying.

503: Service unavailable

Sail cannot serve the request right now. Unavailable model capacity is one possible cause. Some 503 errors can occur after Sail has accepted the request. Follow Retrying requests before submitting again. When a response includes Retry-After, its value is the minimum delay in seconds before retrying.

Request higher limits

Upgrade to Pro for higher limits and prioritized access to models. For the highest capacity, or if your workload has specific throughput needs, contact support to ask about Enterprise.