Rankings

Model rankings

Which models people actually run on Sothe, ranked by tokens processed through the API. Figures cover every account in aggregate — no request, prompt or customer is identifiable here.

Usage data through 2026-09-21 · last 7 days

Top models

Weekly tokens processed, stacked by model over the last 52 weeks. Hover a column for its breakdown; the final column is the week still in progress.

  • Nova 2 Lite
  • Claude Haiku 4.5
  • Nova Micro
  • Kimi K3
  • gpt-oss-20b
  • GPT OSS Safeguard 120B
  • gpt-oss-120b
  • GPT OSS Safeguard 20B
  • Llama 4 Maverick 17B Instruct
  • Others

Leaderboard

Ranked by tokens over the last 7 days, against the 7 days before that.

#ModelTokensRequestsChange
1Nova 2 Liteamazon/nova-2-lite-v1:081K9new
2Claude Haiku 4.5anthropic/claude-haiku-4-5-20251001-v1:023K16new
3Nova Microamazon/nova-micro-v1:04854new
4Kimi K3moonshotai/kimi-k31512new
5GPT OSS Safeguard 120Bopenai/gpt-oss-safeguard-120b741new
6gpt-oss-120bopenai/gpt-oss-120b-1:0741new
7gpt-oss-20bopenai/gpt-oss-20b-1:0741new
8GPT OSS Safeguard 20Bopenai/gpt-oss-safeguard-20b741new
9Llama 4 Maverick 17B Instructmeta/llama4-maverick-17b-instruct-v1:0463new
10MiniMax M2.1minimax/minimax-m2.1451new
11MiniMax M2.5minimax/minimax-m2.5451new
12Llama 3.3 70B Instructmeta/llama3-3-70b-instruct-v1:0421new
13Kimi K2.5moonshotai/kimi-k2.5341new
14MiniMax M2minimax/minimax-m2291new
15NVIDIA Nemotron Nano 12B v2 VL BF16nvidia/nemotron-nano-12b-v2231new
16NVIDIA Nemotron 3 Super 120B A12Bnvidia/nemotron-super-3-120b231new
17Nemotron Nano 3 30Bnvidia/nemotron-nano-3-30b231new
18Llama 3.1 70B Instructmeta/llama3-1-70b-instruct-v1:0221new
19Llama 3.1 8B Instructmeta/llama3-1-8b-instruct-v1:0221new
20Llama 3 8B Instructmeta/llama3-8b-instruct-v1:0211new
21Llama 3 70B Instructmeta/llama3-70b-instruct-v1:0201new
22NVIDIA Nemotron Nano 9B v2nvidia/nemotron-nano-9b-v2201new
23Qwen3 32B (dense)qwen/qwen3-32b-v1:0191new
24Mistral 7B Instructmistral/mistral-7b-instruct-v0:2171new
25Mixtral 8x7B Instructmistral/mixtral-8x7b-instruct-v0:1171new
26Gemma 3 12B ITgoogle/gemma-3-12b-it161new
27Gemma 3 27B PTgoogle/gemma-3-27b-it161new
28Gemma 3 4B ITgoogle/gemma-3-4b-it161new
29Qwen3-Coder-30B-A3B-Instructqwen/qwen3-coder-30b-a3b-v1:0151new
30Qwen3 Coder Nextqwen/qwen3-coder-next151new
31Qwen3 Next 80B A3Bqwen/qwen3-next-80b-a3b151new
32Qwen3 VL 235B A22Bqwen/qwen3-vl-235b-a22b151new
33Kimi K2 Thinkingmoonshot/kimi-k2-thinking141new
34GLM 5zai/glm-5121new
35GLM 4.7zai/glm-4.7121new
36GLM 4.7 Flashzai/glm-4.7-flash121new
37DeepSeek-R1deepseek/r1-v1:0121new
38DeepSeek V3.2deepseek/v3.2111new
39Voxtral Mini 3B 2507mistral/voxtral-mini-3b-2507101new
40Ministral 3 8Bmistral/ministral-3-8b-instruct101new
41Mistral Large (24.02)mistral/mistral-large-2402-v1:0101new
42Mistral Large 3mistral/mistral-large-3-675b-instruct101new
43Ministral 3Bmistral/ministral-3-3b-instruct101new
44Voxtral Small 24B 2507mistral/voxtral-small-24b-2507101new
45Writer Palmyra Vision 7Bwriter/palmyra-vision-7b101new
46Pixtral Large (25.02)mistral/pixtral-large-2502-v1:0101new
47Devstral 2 123Bmistral/devstral-2-123b101new
48Ministral 14B 3.0mistral/ministral-3-14b-instruct101new
49Magistral Small 2509mistral/magistral-small-2509101new
50Mistral Small (24.02)mistral/mistral-small-2402-v1:0101new
51Nova Proamazon/nova-pro-v1:071new
52Nova Liteamazon/nova-lite-v1:071new
53Grok 4.6xai/grok-4.602new
54Llama 4 Scout 17B Instructmeta/llama4-scout-17b-instruct-v1:002new

Market share

Each model's share of the tokens processed in the last 7 days.

Response time

Average end-to-end latency per request, measured at our gateway. It includes the model's own generation time, so a model that writes longer answers will look slower.

ModelAverageRelative
Llama 4 Scout 17B Instruct0.35s
Grok 4.60.36s
gpt-oss-120b0.43s
GLM 4.7 Flash0.45s
GPT OSS Safeguard 120B0.45s
Writer Palmyra Vision 7B0.45s
Voxtral Mini 3B 25070.45s
Mixtral 8x7B Instruct0.47s
Kimi K2 Thinking0.47s
Ministral 3B0.47s
NVIDIA Nemotron Nano 9B v20.48s
Gemma 3 4B IT0.48s
NVIDIA Nemotron Nano 12B v2 VL BF160.48s
Mistral Large 30.48s
Qwen3 Coder Next0.49s
Mistral 7B Instruct0.49s
Llama 3 8B Instruct0.49s
Devstral 2 123B0.50s
Mistral Small (24.02)0.51s
Llama 3.3 70B Instruct0.51s
Gemma 3 27B PT0.52s
Nemotron Nano 3 30B0.52s
MiniMax M20.52s
Qwen3 Next 80B A3B0.52s
Llama 4 Maverick 17B Instruct0.53s
Gemma 3 12B IT0.54s
Llama 3.1 8B Instruct0.54s
NVIDIA Nemotron 3 Super 120B A12B0.54s
Voxtral Small 24B 25070.54s
DeepSeek-R10.55s
Mistral Large (24.02)0.57s
Qwen3 32B (dense)0.57s
Nova Lite0.57s
Llama 3 70B Instruct0.58s
Ministral 14B 3.00.58s
Qwen3 VL 235B A22B0.58s
DeepSeek V3.20.59s
Ministral 3 8B0.59s
GPT OSS Safeguard 20B0.60s
Llama 3.1 70B Instruct0.62s
Qwen3-Coder-30B-A3B-Instruct0.65s
MiniMax M2.10.65s
Pixtral Large (25.02)0.65s
GLM 4.70.67s
Nova Pro0.70s
Magistral Small 25090.72s
Kimi K30.78s
Nova Micro0.98s
MiniMax M2.51.40s
Nova 2 Lite1.84s
gpt-oss-20b2.00s
GLM 52.58s
Claude Haiku 4.53.19s
Kimi K2.56.48s

Prices for every model are on the Models page, and the API that produced these numbers is documented in the docs.