Rankings
Model rankings
Which models people actually run on Sothe, ranked by tokens processed through the API. Figures cover every account in aggregate — no request, prompt or customer is identifiable here.
Usage data through 2026-09-21 · last 7 days
Top models
Weekly tokens processed, stacked by model over the last 52 weeks. Hover a column for its breakdown; the final column is the week still in progress.
- Nova 2 Lite
- Claude Haiku 4.5
- Nova Micro
- Kimi K3
- gpt-oss-20b
- GPT OSS Safeguard 120B
- gpt-oss-120b
- GPT OSS Safeguard 20B
- Llama 4 Maverick 17B Instruct
- Others
Leaderboard
Ranked by tokens over the last 7 days, against the 7 days before that.
| # | Model | Tokens | Requests | Change |
|---|---|---|---|---|
| 1 | Nova 2 Liteamazon/nova-2-lite-v1:0 | 81K | 9 | new |
| 2 | Claude Haiku 4.5anthropic/claude-haiku-4-5-20251001-v1:0 | 23K | 16 | new |
| 3 | Nova Microamazon/nova-micro-v1:0 | 485 | 4 | new |
| 4 | Kimi K3moonshotai/kimi-k3 | 151 | 2 | new |
| 5 | GPT OSS Safeguard 120Bopenai/gpt-oss-safeguard-120b | 74 | 1 | new |
| 6 | gpt-oss-120bopenai/gpt-oss-120b-1:0 | 74 | 1 | new |
| 7 | gpt-oss-20bopenai/gpt-oss-20b-1:0 | 74 | 1 | new |
| 8 | GPT OSS Safeguard 20Bopenai/gpt-oss-safeguard-20b | 74 | 1 | new |
| 9 | Llama 4 Maverick 17B Instructmeta/llama4-maverick-17b-instruct-v1:0 | 46 | 3 | new |
| 10 | MiniMax M2.1minimax/minimax-m2.1 | 45 | 1 | new |
| 11 | MiniMax M2.5minimax/minimax-m2.5 | 45 | 1 | new |
| 12 | Llama 3.3 70B Instructmeta/llama3-3-70b-instruct-v1:0 | 42 | 1 | new |
| 13 | Kimi K2.5moonshotai/kimi-k2.5 | 34 | 1 | new |
| 14 | MiniMax M2minimax/minimax-m2 | 29 | 1 | new |
| 15 | NVIDIA Nemotron Nano 12B v2 VL BF16nvidia/nemotron-nano-12b-v2 | 23 | 1 | new |
| 16 | NVIDIA Nemotron 3 Super 120B A12Bnvidia/nemotron-super-3-120b | 23 | 1 | new |
| 17 | Nemotron Nano 3 30Bnvidia/nemotron-nano-3-30b | 23 | 1 | new |
| 18 | Llama 3.1 70B Instructmeta/llama3-1-70b-instruct-v1:0 | 22 | 1 | new |
| 19 | Llama 3.1 8B Instructmeta/llama3-1-8b-instruct-v1:0 | 22 | 1 | new |
| 20 | Llama 3 8B Instructmeta/llama3-8b-instruct-v1:0 | 21 | 1 | new |
| 21 | Llama 3 70B Instructmeta/llama3-70b-instruct-v1:0 | 20 | 1 | new |
| 22 | NVIDIA Nemotron Nano 9B v2nvidia/nemotron-nano-9b-v2 | 20 | 1 | new |
| 23 | Qwen3 32B (dense)qwen/qwen3-32b-v1:0 | 19 | 1 | new |
| 24 | Mistral 7B Instructmistral/mistral-7b-instruct-v0:2 | 17 | 1 | new |
| 25 | Mixtral 8x7B Instructmistral/mixtral-8x7b-instruct-v0:1 | 17 | 1 | new |
| 26 | Gemma 3 12B ITgoogle/gemma-3-12b-it | 16 | 1 | new |
| 27 | Gemma 3 27B PTgoogle/gemma-3-27b-it | 16 | 1 | new |
| 28 | Gemma 3 4B ITgoogle/gemma-3-4b-it | 16 | 1 | new |
| 29 | Qwen3-Coder-30B-A3B-Instructqwen/qwen3-coder-30b-a3b-v1:0 | 15 | 1 | new |
| 30 | Qwen3 Coder Nextqwen/qwen3-coder-next | 15 | 1 | new |
| 31 | Qwen3 Next 80B A3Bqwen/qwen3-next-80b-a3b | 15 | 1 | new |
| 32 | Qwen3 VL 235B A22Bqwen/qwen3-vl-235b-a22b | 15 | 1 | new |
| 33 | Kimi K2 Thinkingmoonshot/kimi-k2-thinking | 14 | 1 | new |
| 34 | GLM 5zai/glm-5 | 12 | 1 | new |
| 35 | GLM 4.7zai/glm-4.7 | 12 | 1 | new |
| 36 | GLM 4.7 Flashzai/glm-4.7-flash | 12 | 1 | new |
| 37 | DeepSeek-R1deepseek/r1-v1:0 | 12 | 1 | new |
| 38 | DeepSeek V3.2deepseek/v3.2 | 11 | 1 | new |
| 39 | Voxtral Mini 3B 2507mistral/voxtral-mini-3b-2507 | 10 | 1 | new |
| 40 | Ministral 3 8Bmistral/ministral-3-8b-instruct | 10 | 1 | new |
| 41 | Mistral Large (24.02)mistral/mistral-large-2402-v1:0 | 10 | 1 | new |
| 42 | Mistral Large 3mistral/mistral-large-3-675b-instruct | 10 | 1 | new |
| 43 | Ministral 3Bmistral/ministral-3-3b-instruct | 10 | 1 | new |
| 44 | Voxtral Small 24B 2507mistral/voxtral-small-24b-2507 | 10 | 1 | new |
| 45 | Writer Palmyra Vision 7Bwriter/palmyra-vision-7b | 10 | 1 | new |
| 46 | Pixtral Large (25.02)mistral/pixtral-large-2502-v1:0 | 10 | 1 | new |
| 47 | Devstral 2 123Bmistral/devstral-2-123b | 10 | 1 | new |
| 48 | Ministral 14B 3.0mistral/ministral-3-14b-instruct | 10 | 1 | new |
| 49 | Magistral Small 2509mistral/magistral-small-2509 | 10 | 1 | new |
| 50 | Mistral Small (24.02)mistral/mistral-small-2402-v1:0 | 10 | 1 | new |
| 51 | Nova Proamazon/nova-pro-v1:0 | 7 | 1 | new |
| 52 | Nova Liteamazon/nova-lite-v1:0 | 7 | 1 | new |
| 53 | Grok 4.6xai/grok-4.6 | 0 | 2 | new |
| 54 | Llama 4 Scout 17B Instructmeta/llama4-scout-17b-instruct-v1:0 | 0 | 2 | new |
Response time
Average end-to-end latency per request, measured at our gateway. It includes the model's own generation time, so a model that writes longer answers will look slower.
| Model | Average | Relative |
|---|---|---|
| Llama 4 Scout 17B Instruct | 0.35s | |
| Grok 4.6 | 0.36s | |
| gpt-oss-120b | 0.43s | |
| GLM 4.7 Flash | 0.45s | |
| GPT OSS Safeguard 120B | 0.45s | |
| Writer Palmyra Vision 7B | 0.45s | |
| Voxtral Mini 3B 2507 | 0.45s | |
| Mixtral 8x7B Instruct | 0.47s | |
| Kimi K2 Thinking | 0.47s | |
| Ministral 3B | 0.47s | |
| NVIDIA Nemotron Nano 9B v2 | 0.48s | |
| Gemma 3 4B IT | 0.48s | |
| NVIDIA Nemotron Nano 12B v2 VL BF16 | 0.48s | |
| Mistral Large 3 | 0.48s | |
| Qwen3 Coder Next | 0.49s | |
| Mistral 7B Instruct | 0.49s | |
| Llama 3 8B Instruct | 0.49s | |
| Devstral 2 123B | 0.50s | |
| Mistral Small (24.02) | 0.51s | |
| Llama 3.3 70B Instruct | 0.51s | |
| Gemma 3 27B PT | 0.52s | |
| Nemotron Nano 3 30B | 0.52s | |
| MiniMax M2 | 0.52s | |
| Qwen3 Next 80B A3B | 0.52s | |
| Llama 4 Maverick 17B Instruct | 0.53s | |
| Gemma 3 12B IT | 0.54s | |
| Llama 3.1 8B Instruct | 0.54s | |
| NVIDIA Nemotron 3 Super 120B A12B | 0.54s | |
| Voxtral Small 24B 2507 | 0.54s | |
| DeepSeek-R1 | 0.55s | |
| Mistral Large (24.02) | 0.57s | |
| Qwen3 32B (dense) | 0.57s | |
| Nova Lite | 0.57s | |
| Llama 3 70B Instruct | 0.58s | |
| Ministral 14B 3.0 | 0.58s | |
| Qwen3 VL 235B A22B | 0.58s | |
| DeepSeek V3.2 | 0.59s | |
| Ministral 3 8B | 0.59s | |
| GPT OSS Safeguard 20B | 0.60s | |
| Llama 3.1 70B Instruct | 0.62s | |
| Qwen3-Coder-30B-A3B-Instruct | 0.65s | |
| MiniMax M2.1 | 0.65s | |
| Pixtral Large (25.02) | 0.65s | |
| GLM 4.7 | 0.67s | |
| Nova Pro | 0.70s | |
| Magistral Small 2509 | 0.72s | |
| Kimi K3 | 0.78s | |
| Nova Micro | 0.98s | |
| MiniMax M2.5 | 1.40s | |
| Nova 2 Lite | 1.84s | |
| gpt-oss-20b | 2.00s | |
| GLM 5 | 2.58s | |
| Claude Haiku 4.5 | 3.19s | |
| Kimi K2.5 | 6.48s |
Prices for every model are on the Models page, and the API that produced these numbers is documented in the docs.