AI 大模型排行榜 (Artificial Analysis LLM Ranking)

信息查询
3.9k 次浏览
100% 有帮助 · 1 人反馈

值品工具箱提供的 AI 大模型排行榜聚合了来自 Artificial Analysis 的权威数据,实时追踪并排名超过 100 个主流大语言模型。

AI 大模型排行榜数据中心

重置
排名 模型名称 综合指数 ▼ 编程 价格 ($/1M)
1 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 53.4 81.6 $20
2 Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) 53.2 80.7 $20
3 GPT-6 Astra (max) 52.7 76.9 $20
4 GPT-6 Astra (xhigh) 52.4 75.9 $20
5 Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback) 51.2 79.1 $20
6 GPT-6 Astra (high) 50.9 77.1 $20
7 Claude Opus 5 (Adaptive Reasoning, Max Effort) 50.8 78 $10
8 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) 49.7 77 $10
9 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 49.6 76.5 $20
10 GPT-6 Astra (medium) 49.6 76.7 $20
11 Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback) 48.9 77.1 $20
12 Claude Opus 5 (Adaptive Reasoning, High Effort) 48.1 76.5 $10
13 Muse Spark 1.3 (max) 48.1 75.8 $2
14 GPT-5.6 Sol (max) 47 77.4 $8
15 Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback) 46.8 75.2 $20
16 GPT-6 Astra (low) 45.8 75.7 $20
17 Qwen3.8 Max (0902) 45.4 76.2 $3
18 Muse Spark 1.3 (xhigh) 45.1 76.5 $2
19 Claude Opus 5 (Adaptive Reasoning, Medium Effort) 44.8 74.3 $10
20 GLM-5.3 (max) 44.8 74.8 $2.15
21 Grok 4.6 (high) 44.3 76.8 $3
22 Grok 4.6 (xhigh) 44.2 75.9 $3
23 GPT-5.6 Sol (xhigh) 44 78.3 $8
24 Step 5 Preview 43.7 - $1.425
25 Kimi K3 (max) 43.6 76.2 $6
26 Grok 4.6 (medium) 42.8 74.4 $3
27 GPT-5.6 Sol (high) 42.3 77.2 $8
28 GPT-5.6 Terra (max) 42.1 76.7 $4.5
29 GLM 5.3 Flash 41.8 71.5 $0.237
30 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 41.8 74.3 $10
31 Gemini 3.8 Flash (high) 40.9 76.3 $1.5
32 Claude Opus 4.7 (Adaptive Reasoning, Max Effort) 40.7 73.6 $10
33 Qwen3.8 Max 40.2 71.8 $3
34 Qwen3.8 2.4T A95B 39.9 71.9 $3
35 Qwen3.8-Flash-Next 39.8 73.1 $0.23
36 Gemini 3.8 Flash (medium) 39.8 74.1 $1.5
37 Gemini 3.7 Flash (medium) 39.6 71.5 $1.5
38 Muse Spark 1.2 (xhigh) 39.6 72.2 $2
39 DeepSeek V4.1 Flash (Reasoning, Max Effort) 39.5 - $0.525
40 Claude Opus 5 (Adaptive Reasoning, Low Effort) 39.4 66.9 $10
41 GPT-5.6 Sol (medium) 39.2 76.3 $8
42 Gemini 3.7 Flash (high) 39.1 76.1 $1.5
43 GPT-5.4 (xhigh) 39 71.1 $5.625
44 Grok 4.5 (high) 38.8 72.4 $3
45 GPT-5.5 (xhigh) 38.4 74.9 $11.25
46 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) 38.2 71.5 $4
47 GPT-5.6 Terra (xhigh) 38 70.6 $4.5
48 GPT-5.6 Luna (max) 37.3 71.4 $0.45
49 GPT-5.5 (high) 37 71.6 $11.25
50 Gemini 3.7 Flash (low) 36.9 71 $1.5
51 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) 36 68.8 $1.98
52 Agnes 3.0 Flash 35.5 - $0.075
53 Agnes 2.5 Pro Beta 35.2 62.3 $0.15
54 Grok 4.6 (low) 35.1 66.3 $3
55 DeepSeek V4 Flash Vision (Reasoning, Max Effort) 34.8 65 $0.66
56 GPT-5.6 Luna (xhigh) 34.6 68.6 $0.45
57 Kimi K3 (low) 34.5 72 $6
58 Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) 34.4 - $4
59 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) 34.3 69.1 $0.66
60 GPT-5.6 Terra (high) 34.2 67.1 $4.5
61 Gemini 3.6 Flash (high) 34 69.2 $1.5
62 GPT-5.5 (medium) 33.8 71.5 $11.25
63 Muse Spark 1.1 (xhigh) 33.7 71.3 $2
64 GLM-5.2 (max) 33.7 68.8 $2.15
65 Qwen3.8 27B (xhigh) 33.7 68.1 $1.125
66 Gemini 3.5 Flash (medium) 33.6 - $3.375
67 Motif 3 33.6 63.5 $0
68 GPT-5.6 Sol (low) 33.5 69.7 $8
69 Gemini 3.8 Flash (low) 33.5 73.5 $1.5
70 Gemini 3.5 Flash (high) 32.6 70.1 $3.375
71 GPT-5.3 Codex (xhigh) 32.5 - $4.813
72 Motif 3 (Beta) 32.3 62 $0
73 GPT-5.6 Luna (high) 32.1 63.3 $0.45
74 Claude Opus 4.6 (Adaptive Reasoning, Max Effort) 31.9 - $10
75 Claude Sonnet 5 (Adaptive Reasoning, High Effort) 31.7 - $4
76 Muse Spark 31.3 58.6 $0
77 Claude Opus 4.7 (Non-reasoning, High Effort) 30.9 - $10
78 GPT-5.5 (low) 30.7 60.9 $11.25
79 K2 Horizon 375B A23B 30.5 61.5 $0
80 DeepSeek V4 Pro 0424 (Reasoning, Max Effort) 30.4 59.4 $0.544
81 GPT-5.2 (xhigh) 30.4 - $4.813
82 Apodex 1.1 30.4 60.8 $0.975
83 DeepSeek V4 Pro 0424 (Reasoning, High Effort) 30.1 58.7 $0.544
84 GPT-5.6 Terra (medium) 30.1 64.7 $4.5
85 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) 30.1 63 $6
86 Gemini 3.1 Pro Preview 29.7 68.8 $4.5
87 Qwen3.7 Max 29.5 66 $3.75
88 MiniMax-M3 29.2 58.6 $0.525
89 Claude Opus 4.5 (Reasoning) 29.1 - $10
90 MiMo-V2-Pro 28.6 - $0
91 GPT-5.2 Codex (xhigh) 28.5 - $4.813
92 Qwen3.6 Max Preview 28.4 - $2.925
93 GPT-5.6 Sol (Non-reasoning) 28.3 65.1 $8
94 Nex-N2-Pro (based on Qwen3.5-397B-A17B) 28.2 59.1 $0
95 Solar Pro 4 28.2 52.7 $0.525
96 Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) 28.1 - $4
97 Gemini 3 Pro Preview (high) 28 - $4.5
98 GLM-5 (Reasoning) 27.9 - $1.55
99 Inkling Small 27.8 52.9 $0.525
100 GPT-5.4 (low) 27.6 - $5.625
101 Qwen3.8 27B (medium) 27.6 56.1 $1.125
102 GPT-5.6 Terra (low) 27.5 58.1 $4.5
103 JT-4.1 Flash 236B A21B 27.3 52.4 $0
104 Grok Build 0.1 0616 27.2 51.5 $1.25
105 Qwen3.6 Plus 27 54.5 $1.125
106 Kimi K2.6 27 61.8 $1.712
107 Agnes 2.5 Pro Alpha 26.8 58.8 $0.563
108 Quasar 438B (max, based on GLM-5.2) 26.7 61.2 $0.9
109 GLM-5-Turbo 26.6 - $0
110 GPT-5.2 (medium) 26.5 - $4.813
111 Claude Opus 4.6 (Non-reasoning, High Effort) 26.4 - $10
112 Gemini 3 Flash Preview (Reasoning) 26.3 - $1.125
113 Qwen3.8 27B (low) 26.2 58.2 $1.125
114 GLM-5.1 (Reasoning) 26.1 55.8 $2
115 GPT-5.5 Instant (June 2026) 26 39.4 $11.25
116 DeepSeek V4 Flash 0420 (Reasoning, High Effort) 26 52 $0.168
117 MiMo-V2.5-Pro 26 60.2 $0.544
118 Kimi K2.7 Code 25.8 60.8 $1.712
119 Grok 4.20 0309 v2 (Reasoning) 25.7 - $1.563
120 K2 Horizon MoVA 36B A4B 25.3 - $0
121 Hy3 25.3 58.8 $0.25
122 Grok 4.20 0309 (Reasoning) 25.2 - $3
123 MiMo-V2.5 25.2 56.8 $0.175
124 Qwen3.7 Plus 25.2 55.9 $0.7
125 MiMo-V2-Omni-0327 25.1 - $0
126 GPT-5.6 Luna (medium) 25 50.7 $0.45
127 Inkling (xhigh) 25 52.1 $1.762
128 Ling 3.0 Flash 24.9 50.6 $0.111
129 GPT-5 Codex (high) 24.9 - $3.438
130 Grok 4.3 (high) 24.9 42.2 $1.563
131 Grok 4.3 (medium) 24.8 - $1.563
132 Solar Open2 250B 24.7 45 $0
133 GPT-5.1 (high) 24.7 49.4 $3.438
134 Claude Sonnet 4.6 (Non-reasoning, High Effort) 24.7 - $6
135 Ling-3.0-flash-VL 24.6 57 $0
136 Grok 4.3 (low) 24.3 - $1.563
137 Claude Sonnet 5 (Adaptive Reasoning, Low Effort) 24.3 - $4
138 GLM-5.1 (Non-reasoning) 24.2 - $2.135
139 DeepSeek V4 Flash 0420 (Reasoning, Max Effort) 24.2 56.2 $0.168
140 GPT-5.4 mini (xhigh) 24.1 56.1 $1.688
141 MiMo-V2-Omni 23.9 - $0
142 Gemini 3.5 Flash (minimal) 23.8 - $3.375
143 GPT-5.1 Codex (high) 23.7 - $3.438
144 Claude Opus 4.5 (Non-reasoning) 23.7 - $10
145 Kimi K2.6 (Non-reasoning) 23.6 - $1.712
146 GLM 5V Turbo (Reasoning) 23.5 - $0
147 Kimi K2.5 (Reasoning) 23.5 46.8 $1.137
148 Claude Sonnet 4.6 (Non-reasoning, Low Effort) 23.3 - $6
149 Claude Sonnet 5 (Non-reasoning, High Effort) 23.2 66.4 $4
150 GPT-5.5 (Non-reasoning) 23.2 56.5 $11.25
151 GPT-5 (high) 23 37.8 $3.438
152 Nemotron 3 Ultra 550B A55B (Reasoning) 22.9 49.3 $1.05
153 Qwen3.5 27B (Reasoning) 22.9 - $0.825
154 GPT-5 (medium) 22.9 - $3.438
155 Claude 4.1 Opus (Reasoning) 22.8 - $30
156 MiniMax-M2.5 22.8 - $0.525
157 MiniMax-M2.7 22.8 52.6 $0.525
158 Hy3-preview (Reasoning) 22.7 - $0.1
159 A.X-K2 22.7 38.8 $0
160 GPT-5.5 Instant (May 2026) 22.7 - $11.25
161 Ling-3.0-flash-Fin 22.6 55.6 $0
162 Grok 4 22.5 - $6
163 MiMo-V2-Flash (Feb 2026) 22.4 - $0
164 GLM-5.2 (Non-reasoning) 22.4 46.5 $2.15
165 Qwen3.8 27B (Non-reasoning) 22.4 44.6 $1.125
166 Gemini 3 Pro Preview (low) 22.3 - $4.5
167 GLM-4.7 (Reasoning) 22.2 45.3 $1
168 Gemini 3.5 Flash-Lite 22.2 49.3 $0.85
169 Kimi K2 Thinking 22 - $1.075
170 o3-pro 21.9 - $35
171 G9v3-39A5B 21.8 33.1 $0
172 GLM-5 (Non-reasoning) 21.8 - $1.55
173 KAT Coder Pro V2 21.7 59.5 $0.525
174 DeepSeek V3.2 (Reasoning) 21.5 44.2 $0.315
175 Qwen3.5 397B A17B (Non-reasoning) 21.4 - $1.35
176 Qwen3.6 27B (Reasoning) 21.4 53.7 $1.35
177 Qwen3 Max Thinking 21.3 - $0
178 GPT-5.6 Luna (low) 21 44.2 $0.45
179 MiniMax-M2.1 20.9 - $0.525
180 DeepSeek V4 Pro 0424 (Non-reasoning) 20.8 - $0.544
181 MiMo-V2-Flash (Reasoning) 20.8 - $0.15
182 GPT-5 (low) 20.8 - $3.438
183 GPT-5.6 Terra (Non-reasoning) 20.8 52.3 $4.5
184 GPT-5.4 nano (xhigh) 20.7 56.1 $0.463
185 Claude 4.5 Sonnet (Reasoning) 20.7 52.1 $6
186 Claude 4 Opus (Reasoning) 20.6 - $30
187 GPT-5 mini (medium) 20.6 - $0.688
188 K2 Horizon 7B 20.6 38.6 $0
189 Qwen3.5 Omni Plus 20.4 - $1.5
190 GPT-5.1 Codex mini (high) 20.4 - $0.688
191 Grok 4.1 Fast (Reasoning) 20.4 - $0
192 o3 20.2 - $3.5
193 GPT-5.4 nano (medium) 20 - $0.463
194 Qwen3.6 27B (Non-reasoning) 19.8 46.6 $1.35
195 GPT-5.4 mini (medium) 19.7 - $1.688
196 K-EXAONE 2.0 0803 19.7 40.6 $0
197 Step 3.7 Flash 19.5 39.6 $0.438
198 Kimi K2.5 (Non-reasoning) 19.4 - $1.2
199 Qwen3.5 27B (Non-reasoning) 19.4 - $0.825
200 Claude 4.5 Sonnet (Non-reasoning) 19.3 - $6
201 Qwen3.5 35B A3B (Reasoning) 19.3 - $0.688
202 LongCat 2.0 19.1 45.3 $0.525
203 Gemma 4 31B (Reasoning) 19 43.4 $0
204 Claude 4 Sonnet (Reasoning) 18.9 37.6 $0
205 DeepSeek V4 Flash 0420 (Non-reasoning) 18.9 - $0.119
206 JT-35B-Flash 18.7 - $0
207 MiniMax-M2 18.6 - $0.525
208 KAT-Coder-Pro V1 18.6 - $0
209 Claude 4.1 Opus (Non-reasoning) 18.6 - $30
210 GLM-4.6 (Reasoning) 18.5 45.8 $0.963
211 Qwen3.5 397B A17B (Reasoning) 18.4 48.2 $1.35
212 MiMo-V2.5-Pro (Non-reasoning) 18.3 - $0.544
213 Qwen3.6 35B A3B (Reasoning) 18.2 41.9 $0.844
214 GPT-5.4 (Non-reasoning) 18.2 - $5.625
215 Grok 4 Fast (Reasoning) 17.9 - $0.275
216 Gemini 3 Flash Preview (Non-reasoning) 17.9 - $1.125
217 Qwen3.5 122B A10B (Non-reasoning) 17.7 43.3 $1.1
218 Claude 3.7 Sonnet (Reasoning) 17.7 36.4 $0
219 Muse Glimmer (high) 17.5 49 $0.637
220 GLM-4.7 (Non-reasoning) 17.4 - $1
221 Hy3-preview (Non-reasoning) 17 - $0.1
222 Ling-2.6-1T 17 - $0.85
223 GPT-5.2 (Non-reasoning) 17 - $4.813
224 Step 3.5 Flash 2603 17 - $0.15
225 Doubao Seed Code 16.9 - $0
226 Claude 4.5 Haiku (Reasoning) 16.9 43.9 $2
227 GPT-5 mini (high) 16.8 15.6 $0.688
228 Gemma 4 26B A4B (Reasoning) 16.7 39.3 $0.179
229 o4-mini (high) 16.7 - $1.925
230 Ring-2.6-1T 16.6 42.8 $0.85
231 Step 3.5 Flash 16.6 - $0.15
232 Claude 4 Opus (Non-reasoning) 16.6 - $30
233 Claude 4 Sonnet (Non-reasoning) 16.6 - $0
234 DeepSeek V3.2 Exp (Reasoning) 16.6 - $0.315
235 Qwen3 Max Thinking (Preview) 16.3 - $2.4
236 Gemini 2.5 Pro 16.1 33.3 $3.438
237 MiMo-V2-Flash (Non-reasoning) 16 49.8 $0
238 DeepSeek V3.2 (Non-reasoning) 16 - $0.315
239 K2 Horizon 3.7B 15.6 26.1 $0
240 Qwen3 Max 15.6 - $2.4
241 Qwen3.5 122B A10B (Reasoning) 15.6 45.7 $1.1
242 Gemini 3.1 Flash-Lite 15.6 34.7 $0.563
243 GPT-5.6 Luna (Non-reasoning) 15.5 39.3 $0.45
244 Gemini 2.5 Flash Preview (Sep '25) (Reasoning) 15.5 - $0
245 Claude 4.5 Haiku (Non-reasoning) 15.4 - $2
246 Kimi K2 0905 15.3 - $1.075
247 Ling 3.0 Tiny 15.3 26.5 $0
248 Claude 3.7 Sonnet (Non-reasoning) 15.3 - $6
249 Qwen3.6 35B A3B (Non-reasoning) 15.2 28.1 $0.844
250 o1 15.2 39.7 $26.25
251 Qwen3.5 35B A3B (Non-reasoning) 15.1 37 $0.688
252 Gemini 2.5 Pro Preview (Mar' 25) 15 46.7 $0
253 GLM-4.6 (Non-reasoning) 14.9 - $0.981
254 GLM-4.7-Flash (Reasoning) 14.9 - $0.153
255 Granite 4.2 30B 14.8 29.9 $0.282
256 DeepSeek V3.1 Terminus (Reasoning) 14.8 43.5 $1.914
257 Grok 3 mini Reasoning (high) 14.6 - $0.35
258 Grok 4.20 0309 (Non-reasoning) 14.6 - $3
259 Gemini 2.5 Pro Preview (May' 25) 14.5 - $3.438
260 DeepSeek V3.2 Speciale 14.5 - $0
261 K-EXAONE (Reasoning) 14.4 32.1 $0
262 ERNIE 5.0 Thinking Preview 14.3 - $0
263 Grok 4.20 0309 v2 (Non-reasoning) 14.2 - $1.563
264 Mistral Medium 3.5 14.2 46.9 $3
265 Gemma 4 12B (Reasoning) 14.2 31 $0.15
266 Nova 2.0 Pro Preview (medium) 14.2 34 $3.438
267 Grok Code Fast 1 14.1 - $0
268 Grok 4.3 (Non-reasoning) 14 35.2 $1.563
269 DeepSeek V3.1 Terminus (Non-reasoning) 13.9 - $0.453
270 Gemma 4 31B (Non-reasoning) 13.9 33.2 $0.205
271 DeepSeek V3.2 Exp (Non-reasoning) 13.9 - $0.315
272 Apriel-v1.5-15B-Thinker 13.8 - $0
273 Mercury 2 13.8 31.1 $0.375
274 DeepSeek V3.1 (Non-reasoning) 13.7 - $0.848
275 Qwen3.5 9B (Reasoning) 13.7 28.7 $0.151
276 Nova 2.0 Omni (medium) 13.6 - $0.85
277 DeepSeek V3.1 (Reasoning) 13.5 - $0.865
278 Qwen3 VL 235B A22B (Reasoning) 13.4 - $1.3
279 Apriel-v1.6-15B-Thinker 13.4 - $0
280 Nova 2.0 Lite (high) 13.4 23 $0.85
281 GPT-5.1 (Non-reasoning) 13.3 - $3.438
282 Qwen3.5 9B (Non-reasoning) 13.3 23.5 $0.19
283 EXAONE 4.5 33B 13.2 23.6 $0
284 Command A+ 13.1 27.8 $0
285 Gemma 4 26B A4B (Non-reasoning) 13.1 - $0.198
286 Qwen3.5 4B (Reasoning) 13.1 22.6 $0.06
287 DeepSeek R1 0528 (May '25) 13.1 - $1.763
288 Gemini 2.5 Flash (Reasoning) 13.1 - $0.85
289 GPT-5 nano (high) 13 - $0.138
290 Nemotron 3.5 Lightning 12.9 26.8 $0.095
291 Nemotron 3 Super 120B A12B (Reasoning) 12.8 37.7 $0.307
292 Nova 2.0 Pro Preview (low) 12.8 25.9 $3.438
293 GLM-4.5 (Reasoning) 12.8 - $0
294 Kimi K2 12.7 - $1.002
295 Qwen3 235B A22B 2507 (Reasoning) 12.7 22.1 $0.747
296 GPT-4.1 12.7 - $3.5
297 Qwen3 Max (Preview) 12.6 - $2.4
298 Nova 2.0 Lite (medium) 12.5 - $0.85
299 GPT-5 nano (medium) 12.5 - $0.138
300 Qwen3.5 Omni Flash 12.5 - $0.275
301 o3-mini 12.5 - $1.925
302 MiniCPM5-2B 12.5 14.5 $0
303 o1-pro 12.4 - $262.5
304 Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning) 12.4 - $0
305 JT-MINI 12.2 - $0
306 Grok 3 12.1 - $8
307 Seed-OSS-36B-Instruct 12.1 - $0.3
308 Qwen3 235B A22B 2507 Instruct 12 - $0.403
309 Qwen3 Coder 480B A35B Instruct 11.9 - $3
310 Qwen3 VL 32B (Reasoning) 11.9 - $0.28
311 Magistral Medium 1.2 11.8 21.3 $0
312 Sonar Reasoning Pro 11.8 - $0
313 Nova 2.0 Lite (low) 11.8 - $0.85
314 HyperNova 60B 2605 (high, based on gpt-oss-120b) 11.7 23.2 $0
315 MiniMax M1 80k 11.7 - $0.963
316 GPT-5.4 nano (Non-Reasoning) 11.7 - $0.463
317 Nemotron Cascade 2 30B A3B 11.7 25.3 $0
318 Gemini 2.5 Flash Preview (Reasoning) 11.7 - $0
319 gpt-oss-120b (high) 11.6 30.4 $0.261
320 K2 Think V2 11.5 21 $0
321 LongCat Flash Lite 11.5 - $0
322 GPT-5 (minimal) 11.4 - $3.438
323 DeepSeek R1 (Jan '25) 11.4 24.6 $2.5
324 o1-preview 11.4 34 $28.875
325 HyperCLOVA X SEED Think (32B) 11.4 - $0
326 Grok 4.1 Fast (Non-reasoning) 11.3 - $0
327 Mistral Small 4 (Reasoning) 11.3 26.6 $0.262
328 GLM-4.6V (Reasoning) 11.2 - $0.45
329 K-EXAONE (Non-reasoning) 11.2 - $0
330 Qwen3 Next 80B A3B (Reasoning) 11.2 17.4 $0.412
331 GPT-5.4 mini (Non-Reasoning) 11.1 - $1.688
332 Nova 2.0 Omni (low) 11.1 - $0.85
333 GLM-4.5-Air 11.1 - $0.372
334 Granite 4.2 8B 11.1 22.4 $0.107
335 Grok 4 Fast (Non-reasoning) 11.1 - $0.275
336 Mi:dm K 2.5 Pro 11 - $0
337 o3-mini (high) 11 16.3 $1.925
338 Ring-1T 10.9 - $0
339 G9v3-3B 10.8 9.9 $0
340 Trinity Large Thinking 10.8 25.8 $0.412
341 Qwen3.5 4B (Non-reasoning) 10.8 20.3 $0.06
342 INTELLECT-3 10.6 - $0
343 GLM-4.7-Flash (Non-reasoning) 10.6 - $0.153
344 GPT-5 (ChatGPT) 10.4 - $0
345 Solar Open 100B (Reasoning) 10.4 - $0
346 Grok 3 Reasoning Beta 10.4 - $0
347 Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning) 10.4 - $0.175
348 Nemotron 3 Nano Omni 30B A3B Reasoning 10.3 13.8 $0.158
349 gpt-oss-120b (low) 10.2 21.2 $0.249
350 GPT-4.1 mini 10.2 20.2 $0.7
351 Llama 4 Maverick 10 16.3 $0.422
352 MiniMax M1 40k 10 - $0
353 Nova 2.0 Pro Preview (Non-reasoning) 10 20.9 $3.438
354 gpt-oss-20b (low) 10 - $0.103
355 Qwen3 VL 235B A22B Instruct 9.9 - $0.7
356 North Mini Code 9.9 36.5 $0
357 GPT-5 mini (minimal) 9.9 - $0.688
358 K2-V2 (high) 9.9 - $0
359 Gemini 2.5 Flash (Non-reasoning) 9.9 - $0.85
360 Qwen3 30B A3B 2507 (Reasoning) 9.8 12.1 $0.75
361 o1-mini 9.8 - $0
362 DeepSeek V3 0324 9.7 21.2 $0.927
363 Ling 2.6 Flash 9.7 25.3 $0
364 Qwen3 Next 80B A3B Instruct 9.6 - $0.412
365 Tri-21B-think Preview 9.6 - $0
366 Qwen3 Coder 30B A3B Instruct 9.6 - $0.9
367 GPT-4.5 (Preview) 9.6 - $0
368 DiffusionGemma 26B A4B 9.5 19.7 $0
369 Qwen3 235B A22B (Reasoning) 9.5 - $2.625
370 QwQ 32B 9.5 - $0.745
371 Qwen3 VL 30B A3B (Reasoning) 9.5 - $0.75
372 Gemini 2.0 Flash Thinking Experimental (Jan '25) 9.4 24.1 $0
373 Gemma 4 12B (Non-reasoning) 9.4 - $0.15
374 Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning) 9.3 - $0.175
375 Mistral Large 3 9.3 20.1 $0.75
376 Qwen3 Coder Next 9.2 36.2 $0.563
377 Motif-2-12.7B-Reasoning 9.2 - $0
378 Mistral Medium 3.1 9.2 20.5 $0.8
379 Ling-1T 9.2 - $0
380 Nova Premier 9.2 - $5
381 Solar Pro 2 (Preview) (Reasoning) 9.1 - $0
382 Granite 4.2 3B 9.1 17.5 $0.052
383 Magistral Medium 1 9.1 - $0
384 Mistral Medium 3 9 - $0.8
385 K2-V2 (medium) 9 - $0
386 Llama Nemotron Super 49B v1.5 (Reasoning) 9 - $0.4
387 Devstral Medium 9 - $0
388 Mistral Small 4 (Non-reasoning) 9 - $0.262
389 Tri-21B-Think 9 - $0
390 gpt-oss-20b (high) 9 20.7 $0.092
391 GPT-4o (March 2025, chatgpt-4o-latest) 9 - $0
392 Gemini 2.0 Flash (Feb '25) 8.9 - $0
393 Claude 3.5 Haiku 8.9 15.9 $0
394 Llama 3.3 Nemotron Super 49B v1 (Reasoning) 8.9 - $0
395 Gemma 4 E4B (Reasoning) 8.9 9.4 $0.04
396 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) 8.9 14.4 $0.088
397 Qwen3 4B 2507 (Reasoning) 8.8 - $0
398 MiniCPM5-1B (Reasoning) 8.8 - $0
399 Sarvam 105B (high) 8.8 - $0.074
400 Gemini 2.0 Pro Experimental (Feb '25) 8.7 25.5 $0
401 Nova 2.0 Lite (Non-reasoning) 8.7 - $0.85
402 Devstral Small (May '25) 8.7 - $0
403 Claude 3 Opus 8.7 19.5 $30
404 MiniCPM5-1B (Non-reasoning) 8.7 - $0
405 Sonar Reasoning 8.7 - $0
406 Gemini 2.5 Flash Preview (Non-reasoning) 8.7 - $0
407 Devstral 2 8.6 31.3 $0
408 Magistral Small 1.2 8.6 14.7 $0.75
409 Qwen3 32B (Reasoning) 8.6 15.3 $0.28
410 Gemini 2.5 Flash-Lite (Reasoning) 8.5 - $0.175
411 DeepSeek V3 (Dec '24) 8.5 23 $0.463
412 GPT-4o (Nov '24) 8.4 - $4.375
413 Nanbeige4.1-3B 8.4 9.6 $0
414 LFM2.5-2.6B 8.4 7.7 $0
415 Qwen3 VL 32B Instruct 8.4 - $0.28
416 DeepSeek R1 Distill Qwen 32B 8.4 - $0
417 GLM-4.6V (Non-reasoning) 8.4 - $0.45
418 Qwen3 235B A22B (Non-reasoning) 8.3 - $1.225
419 Mistral Small 3.2 8.2 12.5 $0.15
420 Magistral Small 1 8.2 - $0
421 Gemini 2.0 Flash (experimental) 8.2 - $0
422 EXAONE 4.0 32B (Reasoning) 8.2 - $0
423 Qwen3 VL 8B (Reasoning) 8.2 - $0.66
424 Qwen3 14B (Reasoning) 8.2 13.8 $1.313
425 Nova 2.0 Omni (Non-reasoning) 8.2 - $0.85
426 DeepSeek R1 0528 Qwen3 8B 8.1 - $0
427 Llama 4 Scout 8.1 8.2 $0.313
428 Qwen2.5 Max 8 - $0
429 Qwen3 VL 30B A3B Instruct 7.9 - $0.35
430 Hermes 4 - Llama-3.1 70B (Reasoning) 7.9 - $0
431 Gemini 1.5 Pro (Sep '24) 7.9 23.6 $0
432 Solar Pro 2 (Preview) (Non-reasoning) 7.9 - $0
433 DeepSeek R1 Distill Llama 70B 7.9 - $0.8
434 Claude 3.5 Sonnet (Oct '24) 7.9 30.2 $6
435 DeepSeek R1 Distill Qwen 14B 7.8 - $0
436 Falcon-H1R-7B 7.8 - $0
437 GPT-4.1 nano 7.8 11.1 $0.175
438 Solar Pro 3 7.8 16.2 $0.262
439 Ling-flash-2.0 7.8 - $0.247
440 Gemma 4 E2B (Reasoning) 7.8 7.2 $0
441 Qwen3 Omni 30B A3B (Reasoning) 7.8 - $0.43
442 GPT-4o (Aug '24) 7.7 - $4.375
443 Qwen2.5 Instruct 72B 7.7 - $0.48
444 Sonar 7.7 - $0
445 Step3 VL 10B 7.7 - $0
446 Llama 3.3 Instruct 70B 7.7 11.9 $0.712
447 Qwen3 30B A3B (Reasoning) 7.6 - $0.75
448 Sonar Pro 7.6 - $0
449 Devstral Small (Jul '25) 7.6 - $0
450 QwQ 32B-Preview 7.6 - $0
451 GLM-4.5V (Reasoning) 7.6 - $0.9
452 Mistral Large 2 (Nov '24) 7.6 - $0
453 Devstral Small 2 7.5 29.3 $0
454 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) 7.5 - $0
455 Qwen3 30B A3B 2507 Instruct 7.5 - $0.35
456 ERNIE 4.5 300B A47B 7.5 - $0.485
457 Hermes 4 - Llama-3.1 405B (Reasoning) 7.5 - $1.5
458 Solar Pro 2 (Reasoning) 7.5 - $0
459 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) 7.5 - $0.3
460 Gemma 4 E4B (Non-reasoning) 7.5 - $0.04
461 Granite 4.1 30B 7.4 10.4 $0
462 NVIDIA Nemotron Nano 9B V2 (Reasoning) 7.4 - $0.07
463 Hermes 4 - Llama-3.1 405B (Non-reasoning) 7.4 - $1.5
464 Gemini 2.0 Flash-Lite (Feb '25) 7.4 - $0
465 NVIDIA Nemotron 3 Nano 4B 7.4 8 $0
466 Llama Nemotron Super 49B v1.5 (Non-reasoning) 7.4 - $0.4
467 Qwen3 32B (Non-reasoning) 7.3 - $0.28
468 GPT-4o (May '24) 7.3 24.2 $7.5
469 Gemini 2.0 Flash-Lite (Preview) 7.3 - $0
470 K2-V2 (low) 7.3 - $0
471 Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning) 7.3 - $0
472 Kimi Linear 48B A3B Instruct 7.3 - $0
473 Llama 3.1 Instruct 405B 7.3 - $0
474 Qwen3 8B (Reasoning) 7.3 9 $0.66
475 Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) 7.3 - $0
476 Qwen3 VL 8B Instruct 7.3 - $0.31
477 Qwen3 4B (Reasoning) 7.2 - $0
478 LFM2.5-8B-A1B 7.2 - $0
479 Claude 3.5 Sonnet (June '24) 7.2 26 $6
480 Llama 3.1 Tulu3 405B 7.2 - $0
481 GPT-4o (ChatGPT) 7.2 - $0
482 Ring-flash-2.0 7.2 - $0.247
483 Pixtral Large 7.1 - $0
484 Olmo 3.1 32B Think 7.1 - $0
485 Mistral Small 3.1 7.1 26.3 $0.15
486 Grok 2 (Dec '24) 7.1 - $0
487 GPT-5 nano (minimal) 7.1 - $0.138
488 Gemini 1.5 Flash (Sep '24) 7.1 - $0
489 Qwen3 VL 4B (Reasoning) 7 - $0
490 GPT-4 Turbo 7 21.5 $15
491 Solar Pro 2 (Non-reasoning) 7 - $0
492 Nova Pro 7 - $1.4
493 Command A 7 - $4.375
494 Qwen3.5 2B (Reasoning) 6.9 2.9 $0
495 Llama 3.1 Nemotron Instruct 70B 6.9 - $1.2
496 Llama 3.1 Instruct 8B 6.9 5.4 $0.028
497 Grok Beta 6.9 - $0
498 Qwen2.5 Instruct 32B 6.9 - $0
499 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) 6.8 - $0.088
500 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) 6.8 - $0.086
501 Mistral Large 2 (Jul '24) 6.8 - $3
502 Qwen3 4B 2507 Instruct 6.7 - $0
503 Qwen2.5 Coder Instruct 32B 6.7 - $0
504 Qwen3 14B (Non-reasoning) 6.7 - $0.612
505 GPT-4 6.7 13.1 $37.5
506 GLM-4.5V (Non-reasoning) 6.7 - $0.9
507 Mistral Small 3 6.7 - $0.15
508 Gemini 2.5 Flash-Lite (Non-reasoning) 6.7 - $0.175
509 Nova Lite 6.7 - $0.105
510 GPT-4o mini 6.7 11.4 $0.262
511 Hermes 4 - Llama-3.1 70B (Non-reasoning) 6.7 - $0
512 Qwen3 30B A3B (Non-reasoning) 6.6 - $0.35
513 DeepSeek-V2.5 (Dec '24) 6.6 - $0
514 Qwen3 4B (Non-reasoning) 6.6 - $0
515 Llama 3.1 Instruct 70B 6.6 - $0.56
516 Granite 4.1 8B 6.6 9.5 $0.063
517 Sarvam 30B (high) 6.6 - $0.047
518 Gemini 2.0 Flash Thinking Experimental (Dec '24) 6.6 - $0
519 DeepSeek-V2.5 6.6 - $0
520 Olmo 3.1 32B Instruct 6.5 - $0
521 Mistral Saba 6.5 - $0
522 DeepSeek R1 Distill Llama 8B 6.5 - $0
523 Gemma 4 E2B (Non-reasoning) 6.5 - $0
524 Olmo 3 32B Think 6.5 - $0
525 Gemini 1.5 Pro (May '24) 6.4 19.8 $0
526 R1 1776 6.4 - $0
527 Qwen2.5 Turbo 6.4 - $0.088
528 Reka Flash (Sep '24) 6.4 - $0.35
529 Llama 3.2 Instruct 90B (Vision) 6.4 - $0
530 Solar Mini 6.4 - $0.15
531 Celeris-1 6.3 14.4 $0.325
532 Grok-1 6.3 - $0
533 Qwen2 Instruct 72B 6.3 - $0
534 Phi-4 Mini Instruct 6.3 3.8 $0
535 EXAONE 4.0 32B (Non-reasoning) 6.3 - $0
536 Qwen3.5 2B (Non-reasoning) 6.2 2.4 $0
537 Gemini 1.5 Flash-8B 6.2 - $0
538 Qwen3.5 0.8B (Reasoning) 6.1 0 $0
539 DeepHermes 3 - Mistral 24B Preview (Non-reasoning) 6.1 - $0
540 Jamba 1.7 Large 6.1 - $0
541 Granite 4.0 H Small 6 - $0.107
542 Ministral 3 14B 6 14.4 $0.2
543 Jamba 1.5 Large 6 - $3.5
544 Qwen3 Omni 30B A3B Instruct 6 - $0.43
545 Hermes 3 - Llama-3.1 70B 6 - $0.7
546 Qwen3 8B (Non-reasoning) 6 - $0.31
547 DeepSeek-Coder-V2 6 - $0
548 OLMo 2 32B 6 - $0
549 Jamba 1.6 Large 6 - $0
550 LFM2 24B A2B 5.9 - $0
551 Gemini 1.5 Flash (May '24) 5.9 - $0
552 Phi-4 5.9 - $0.219
553 Claude 3 Sonnet 5.9 - $0
554 Nova Micro 5.9 - $0.061
555 Granite 4.1 3B 5.9 4.7 $0
556 Mistral Small (Sep '24) 5.8 - $0.3
557 Gemini 1.0 Ultra 5.8 17.6 $0
558 Phi-3 Mini Instruct 3.8B 5.8 - $0
559 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) 5.8 - $0.3
560 Gemma 3n E4B Instruct Preview (May '25) 5.8 - $0
561 Phi-4 Multimodal Instruct 5.8 - $0
562 Qwen2.5 Coder Instruct 7B 5.8 - $0
563 Mistral Large (Feb '24) 5.8 - $6
564 Mixtral 8x22B Instruct 5.7 - $0
565 Llama 2 Chat 7B 5.7 - $0.1
566 Llama 3.2 Instruct 3B 5.7 - $0
567 MiniCPM-V 4.6 1.3B 5.7 0.7 $0
568 Jamba Reasoning 3B 5.7 - $0
569 Qwen3 VL 4B Instruct 5.7 - $0
570 Qwen1.5 Chat 110B 5.7 - $0
571 Reka Flash 3 5.6 - $0.35
572 Olmo 3 7B Think 5.6 - $0
573 Claude 2.1 5.6 14 $0
574 Claude 3 Haiku 5.6 - $0.5
575 OLMo 2 7B 5.6 - $0
576 Molmo 7B-D 5.6 - $0
577 Ling-mini-2.0 5.5 - $0
578 DeepSeek R1 Distill Qwen 1.5B 5.5 - $0
579 Claude 2.0 5.5 12.9 $0
580 DeepSeek-V2-Chat 5.5 - $0
581 Mistral Small (Feb '24) 5.5 - $0.262
582 Mistral Medium 5.5 - $3
583 GPT-3.5 Turbo 5.5 10.7 $0.75
584 Ministral 3 8B 5.5 9.7 $0.15
585 Llama 3 Instruct 70B 5.5 - $1.175
586 Arctic Instruct 5.4 - $0
587 Qwen Chat 72B 5.4 - $0
588 LFM 40B 5.4 - $0
589 Llama 3.2 Instruct 11B (Vision) 5.4 - $0.345
590 Qwen3.5 0.8B (Non-reasoning) 5.4 1.2 $0
591 PALM-2 5.4 4.6 $0
592 Gemini 1.0 Pro 5.3 - $0
593 DeepSeek Coder V2 Lite Instruct 5.3 - $0
594 Sarvam M (Reasoning) 5.3 - $0
595 DeepSeek LLM 67B Chat (V1) 5.3 - $0
596 Llama 2 Chat 70B 5.3 - $0
597 Llama 2 Chat 13B 5.3 - $0
598 Command-R+ (Apr '24) 5.3 - $0
599 OpenChat 3.5 (1210) 5.3 - $0
600 DBRX Instruct 5.3 - $0
601 Exaone 4.0 1.2B (Reasoning) 5.3 - $0
602 Olmo 3 7B Instruct 5.2 - $0.125
603 Exaone 4.0 1.2B (Non-reasoning) 5.2 - $0
604 LFM2.5-1.2B-Thinking 5.2 - $0
605 Jamba 1.7 Mini 5.2 - $0
606 LFM2 2.6B 5.2 - $0
607 LFM2.5-1.2B-Instruct 5.2 - $0
608 Jamba 1.5 Mini 5.2 - $0.25
609 Granite 4.0 H 1B 5.2 - $0
610 Qwen3 1.7B (Reasoning) 5.2 - $0
611 Jamba 1.6 Mini 5.2 - $0
612 Mixtral 8x7B Instruct 5.1 - $0.512
613 Gemma 3 270M 5.1 - $0
614 Apertus 70B Instruct 5.1 - $1.345
615 Granite 4.0 Micro 5.1 - $0
616 DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) 5.1 - $0
617 Claude Instant 5 7.8 $0
618 Command-R (Mar '24) 5 - $0
619 Llama 65B 5 - $0
620 Mistral 7B Instruct 5 - $0.25
621 Qwen Chat 14B 5 - $0
622 Granite 4.0 1B 5 - $0
623 Molmo2-8B 5 - $0
624 LFM2 8B A1B 4.9 - $0
625 Granite 3.3 8B (Non-reasoning) 4.9 - $0.085
626 Qwen3 1.7B (Non-reasoning) 4.9 - $0
627 Gemma 3 27B Instruct 4.9 10.1 $0
628 Ministral 3 3B 4.8 4.8 $0.1
629 Apertus 8B Instruct 4.8 - $0.125
630 Gemma 3 1B Instruct 4.8 - $0
631 Gemma 3 4B Instruct 4.8 2.7 $0
632 Gemma 3n E2B Instruct 4.8 - $0
633 Gemma 3n E4B Instruct 4.8 3.2 $0
634 Granite 4.0 350M 4.8 - $0
635 Granite 4.0 H 350M 4.8 - $0
636 LFM2 1.2B 4.8 - $0
637 LFM2.5-VL-1.6B 4.8 - $0
638 Llama 3 Instruct 8B 4.8 - $0.07
639 Llama 3.2 Instruct 1B 4.8 - $0
640 Qwen3 0.6B (Non-reasoning) 4.8 - $0
641 Qwen3 0.6B (Reasoning) 4.8 - $0
642 Tiny Aya Global 4.8 - $0
643 Gemma 3 12B Instruct 3.8 5.8 $0
644 K2 Horizon 0.9B 3 3.4 $0
645 Cogito v2.1 (Reasoning) - - $1.25
646 EXAONE 4.5 33B (Non-reasoning) - - $0
647 Gemini 3 Deep Think - - $0
648 GPT-3.5 Turbo (0613) - - $0
649 GPT-4o mini Realtime (Dec '24) - - $0
650 GPT-4o Realtime (Dec '24) - - $0
651 GPT-5.4 Pro (xhigh) - - $67.5
652 GPT-5.5 Pro (xhigh) - - $0
653 Mi:dm K 2.5 Pro Preview - - $0

榜单解读建议

参考 AI 大模型排行榜 时,应综合考虑“综合指数”与“成本价格”。如果您是开发者,编程能力 (Coding) 是更核心的指标。

值品工具箱同步的 AI 大模型排行榜 数据每 24 小时更新,确保您获取到最新的模型性能对比。

指标说明

  • 综合指数:评估通用理解与逻辑。
  • 价格 $/1M:混合 3:1 输入输出比的平均成本。
  • 编程能力:衡量代码生成的准确性。

AI 大模型排行榜 常见问题 (FAQ)

Q1: AI 大模型排行榜 的数据多久更新?

AI 大模型排行榜 数据每 24 小时自动抓取一次,确保最新模型加入列表。

Q2: 这个 AI 大模型排行榜 包含国产模型吗?

是的,只要国产模型通过了 Artificial Analysis 的全球测评,就会出现在 AI 大模型排行榜 中。

Q3: 综合指数在 AI 大模型排行榜 中代表什么?

它代表模型的全能表现。AI 大模型排行榜 通过加权算法给出这个综合评分。

Q4: 如何在 AI 大模型排行榜 中查找性价比最高的游戏?

在 AI 大模型排行榜 页面中,您可以点击“价格”标题进行排序,寻找低价高分的模型。

Q5: AI 大模型排行榜 的编程能力测试准吗?

AI 大模型排行榜 参考了 LiveCodeBench 等权威基准测试,具有极高的参考价值。

Q6: 为什么有的新模型没进入 AI 大模型排行榜?

模型进入 AI 大模型排行榜 需要经过一系列测试,通常在新模型发布后数日内会完成更新。

Q7: AI 大模型排行榜 中的价格计算标准是什么?

价格是基于百万 Token 的调用成本,由 AI 大模型排行榜 统一混合计算得出。

Q8: 手机上能查看 AI 大模型排行榜 吗?

当然可以。AI 大模型排行榜 进行了移动端响应式深度优化。

Q9: AI 大模型排行榜 这个工具免费吗?

是的,由值品工具箱免费提供 AI 大模型排行榜 信息查询服务。

Q10: 我该怎么利用 AI 大模型排行榜 做选型?

如果您需要智能客服,参考 AI 大模型排行榜 的综合指数;如果做翻译,参考编程外的语言指标。

发表评论

请友善文明留言