Sail does not requantize weights. Each row links the exact Hugging Face
checkpoint Sail currently serves.
Core models
| Model | Model ID | Image | LoRA | Reasoning |
|---|---|---|---|---|
moonshotai/Kimi-K3 | ||||
zai-org/GLM-5.3 | ||||
zai-org/GLM-5.3-Flash | ||||
deepseek-ai/DeepSeek-V4-Pro-0813 | ||||
deepseek-ai/DeepSeek-V4-Flash-0731 | ||||
moonshotai/Kimi-K2.6INT4View checkpoint | ||||
google/gemma-4-31B-itBF16View checkpoint | ||||
nvidia/Gemma-4-31B-IT-NVFP4 | ||||
google/gemma-4-12B-itBF16View checkpoint | ||||
openai/gpt-oss-120b |
Flex-only models
Models served with theflex completion window exclusively.
| Model | Model ID | Image | LoRA | Reasoning |
|---|---|---|---|---|
Qwen/Qwen3.6-35B-A3BBF16View checkpoint |
Notes
- If we offer multiple quantizations of a model, we list them as separate model IDs.
- Use
GET /v1/modelsto confirm runtime availability for your API key. - For per-model rates by completion window, see Pricing.