Skip to main content
Time to first token (TTFT) and output speed for each model and completion window, measured on live Sail traffic.
ModelWindowTTFT (median)TTFT (p95)Output speed
Kimi K3
moonshotai/Kimi-K3
Default (ASAP)1.3 s19 s120 tok/s
Balanced2.1 s6.9 s108 tok/s
Flex1.6 s5.7 s137 tok/s
GLM-5.3
zai-org/GLM-5.3
Default (ASAP)1.5 s58 s65 tok/s
Balanced1.9 s1.1 min188 tok/s
Flex1.9 s7.5 s261 tok/s
GLM-5.3-Flash
zai-org/GLM-5.3-Flash
Default (ASAP)1.7 s1.9 min25 tok/s
Balanced3.8 min14.5 min19 tok/s
Flex4.4 s2.4 h23 tok/s
DeepSeek V4.1 Flash
deepseek-ai/DeepSeek-V4.1-Flash
Default (ASAP)3.6 s31 s44 tok/s
Balanced2.9 s54 s55 tok/s
Flex
DeepSeek V4 Pro 0813
deepseek-ai/DeepSeek-V4-Pro-0813
Default (ASAP)1.6 s4.2 s21 tok/s
Balanced
Flex27 tok/s
DeepSeek V4 Flash 0731
deepseek-ai/DeepSeek-V4-Flash-0731
Default (ASAP)1.5 s6.1 s82 tok/s
Balanced2.2 s11 s54 tok/s
Flex3.0 s7.9 s48 tok/s
Kimi-K2.6
moonshotai/Kimi-K2.6
Default (ASAP)1.1 s3.0 s41 tok/s
Balanced
Flex
Gemma 4 31B IT
google/gemma-4-31B-it
Default (ASAP)0.6 s6.1 s67 tok/s
Balanced10 s1.9 min33 tok/s
Flex1.5 s1.9 min32 tok/s
Gemma 4 31B IT (NVFP4)
nvidia/Gemma-4-31B-IT-NVFP4
Default (ASAP)1.0 s2.8 s27 tok/s
Balanced0.9 s2.5 s34 tok/s
Flex
Gemma 4 12B IT
google/gemma-4-12B-it
Default (ASAP)0.5 s1.0 s47 tok/s
Balanced
Flex
gpt-oss-120b
openai/gpt-oss-120b
Default (ASAP)0.6 s1.1 s203 tok/s
Qwen3.6 35B A3B
Qwen/Qwen3.6-35B-A3B
Flex0.6 s1.5 min47 tok/s

Notes

  • Time to first token grows with the length of your input, because the model processes the whole prompt before it produces the first token.
  • TTFT for balanced and flex includes time spent waiting in the queue.
  • These figures are typical rather than guaranteed, and vary with load and request size.