Melhor IA para Código em 2026Atualizado por AA Coding Index

Qual IA programa melhor em 2026? Ranking de 486 modelos usando o AA Coding Index como sinal principal, com fallback para LiveCodeBench e SciCode. A lógica prioriza modelos atuais e evita distorções por picos isolados.

Sincronizado: 29 de julho de 2026 486 modelos com benchmarks de código

Casos de Uso

Autocompletar Código

Sugestões inline enquanto você digita. Ideal para IDEs como Cursor e VS Code.

Top modelos: GPT-5.6 Sol (xhigh), Claude Opus 5, GPT-5.6 Sol (max)

Geração de Código

Criar funções, classes e projetos completos a partir de descrições em linguagem natural.

Top modelos: GPT-5.6 Sol (xhigh), Claude Opus 5, GPT-5.6 Sol (max)

Debug e Code Review

Identificar bugs, sugerir correções e revisar pull requests automaticamente.

Top modelos: GPT-5.6 Sol (xhigh), Claude Opus 5, GPT-5.6 Sol (max)

Ranking de Coding — Top Modelos

#ModeloEmpresaScore de CodingBenchmarkContextPreço InputOpen Source
🥇GPT-5.6 Sol (xhigh)OpenAIOpenAI
78.3
AA Coding Index$5.00
🥈Claude Opus 5anthropicanthropic
78.0
AA Coding Index$5.00
🥉GPT-5.6 Sol (max)OpenAIOpenAI
77.4
AA Coding Index1.1M tokens$5.00
4GPT-5.6 Sol (high)OpenAIOpenAI
77.2
AA Coding Index$5.00
5GPT-5.6 Terra (max)OpenAIOpenAI
76.7
AA Coding Index1.1M tokens$2.50
6Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)AnthropicAnthropic
76.5
AA Coding Index1.0M tokens$10.00
7GPT-5.6 Sol (medium)OpenAIOpenAI
76.3
AA Coding Index$5.00
8Kimi K3Kimi
76.2
AA Coding Index$3.00
9GPT-5.5OpenAIOpenAI
74.9
AA Coding Index1.1M tokens$5.00
10Claude Opus 4.8 (Adaptive Reasoning, Max Effort)AnthropicAnthropic
74.3
AA Coding Index1.0M tokens$5.00
11Claude Opus 4.7AnthropicAnthropic
73.6
AA Coding Index1.0M tokens$5.00
12Claude Opus 4.7 (Fast)AnthropicAnthropic
73.6
AA Coding Index1.0M tokens$30.00
13Grok 4.5xaixai
72.4
AA Coding Index500K tokens$2.00
14GPT-5.5 ProOpenAIOpenAI
71.6
AA Coding Index1.1M tokens
15Claude Sonnet 5anthropicanthropic
71.5
AA Coding Index1.0M tokens$2.00
16GPT-5.6 Luna (max)OpenAIOpenAI
71.4
AA Coding Index1.1M tokens$1.00
17Muse Spark 1.1 (xhigh)MetaMeta
71.3
AA Coding Index$1.25
18GPT-5.4OpenAIOpenAI
71.1
AA Coding Index1.1M tokens$2.50
19GPT-5.4 ProOpenAIOpenAI
71.1
AA Coding Index1.1M tokens$30.00
20GPT-5.6 Terra (xhigh)OpenAIOpenAI
70.6
AA Coding Index$2.50
21Google: Gemini 3.5 FlashGoogleGoogle
70.1
AA Coding Index1.0M tokens$1.50
22Gemini 3.5 Flash (minimal)GoogleGoogle
70.1
AA Coding Index$1.50
23GPT-5.6 Sol (low)OpenAIOpenAI
69.7
AA Coding Index$5.00
24Gemini 3.6 Flash (high)GoogleGoogle
69.2
AA Coding Index1.0M tokens$1.50
25Gemini 3.1 Pro PreviewGoogleGoogle
68.8
AA Coding Index1.0M tokens$2.00
26GLM-5.2 (max)Z.aiZ.ai
68.8
AA Coding Index$1.40
27Gemini 3.1googlegoogle
68.8
AA Coding Index
28GPT-5.6 Luna (xhigh)OpenAIOpenAI
68.6
AA Coding Index$1.00
29GPT-5.6 Terra (high)OpenAIOpenAI
67.1
AA Coding Index$2.50
30Qwen3.7 MaxAlibabaAlibaba
66.0
AA Coding Index$2.50
31GPT-5.6 Sol (Non-reasoning)OpenAIOpenAI
65.1
AA Coding Index$5.00
32GPT-5.6 Terra (medium)OpenAIOpenAI
64.7
AA Coding Index$2.50
33GPT-5.6 Luna (high)OpenAIOpenAI
63.3
AA Coding Index$1.00
34Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)AnthropicAnthropic
63.0
AA Coding Index$3.00
35Claude Sonnet 4.6 (Non-reasoning, Low Effort)AnthropicAnthropic
63.0
AA Coding Index$3.00
36Motif 3 (Beta)Motif Technologies
62.0
AA Coding Index
37MoonshotAI: Kimi K2.6MoonshotAIMoonshotAI
61.8
AA Coding Index262K tokens$0.95
38Kimi K2.7 CodeKimi
60.8
AA Coding Index$0.95
39Xiaomi: MiMo-V2.5-ProXiaomi
60.2
AA Coding Index1.0M tokens$0.43
40Kwaipilot: KAT-Coder-Pro V2Kwaipilot
59.5
AA Coding Index256K tokens$0.30
41DeepSeek V4 ProDeepSeekDeepSeek
59.4
AA Coding Index1.0M tokens$0.43
42Nex-N2-ProNex AGI
59.1
AA Coding Index262K tokens$0.50
43Hy3-preview (Reasoning)Tencent
58.8
AA Coding Index262K tokens$0.14
44Agnes 2.5 Pro AlphaSapiens AI
58.8
AA Coding Index$0.45
45Muse SparkMetaMeta
58.6
AA Coding Index
46MiniMax-M3MiniMax
58.6
AA Coding Index1.0M tokens$0.30
47GPT-5.6 Terra (low)OpenAIOpenAI
58.1
AA Coding Index$2.50
48MiMo-V2.5Xiaomi
56.8
AA Coding Index$0.14
49DeepSeek V4 FlashDeepSeekDeepSeek
56.2
AA Coding Index1.0M tokens$0.14
50GPT-5.4 MiniOpenAIOpenAI
56.1
AA Coding Index400K tokens$0.75
51GPT-5.4 NanoOpenAIOpenAI
56.1
AA Coding Index400K tokens$0.20
52Qwen3.7 PlusAlibabaAlibaba
55.9
AA Coding Index$0.40
53GLM-5.1 (Reasoning)Z.aiZ.ai
55.8
AA Coding Index$1.38
54GLM-5.1 (Non-reasoning)Z.aiZ.ai
55.8
AA Coding Index$1.38
55MiniMax: MiniMax M2.7MiniMax
52.6
AA Coding Index197K tokens$0.30
56JT-4.1 Flash 236B A21BChina Mobile
52.4
AA Coding Index
57GPT-5.6 Terra (Non-reasoning)OpenAIOpenAI
52.3
AA Coding Index$2.50
58Claude 4.5 Sonnet (Reasoning)AnthropicAnthropic
52.1
AA Coding Index$3.00
59Claude 4.5 Sonnet (Non-reasoning)AnthropicAnthropic
52.1
AA Coding Index$3.00
60InklingThinking Machines
52.1
AA Coding Index$1.87
61Grok Build 0.1 0616xAIxAI
51.5
AA Coding Index$1.00
62GPT-5.6 Luna (medium)OpenAIOpenAI
50.7
AA Coding Index$1.00
63MiMo-V2-Flash (Reasoning)Xiaomi
49.8
AA Coding Index262K tokens$0.10
64MiMo-V2-Flash (Feb 2026)Xiaomi
49.8
AA Coding Index
65GPT-5.1OpenAIOpenAI
49.4
AA Coding Index400K tokens$1.25
66GPT-5.1 ChatOpenAIOpenAI
49.4
AA Coding Index128K tokens$1.25
67Gemini 3.5 Flash-LiteGoogleGoogle
49.3
AA Coding Index1.0M tokens$0.30
68Nemotron 3 Ultra 550B A55B (Reasoning)NVIDIANVIDIA
49.3
AA Coding Index1.0M tokens$0.68
69Mistral: Mistral Medium 3.5Mistral AIMistral AI
46.9
AA Coding Index262K tokens$1.50
70Kimi K2 ThinkingKimi
46.8
AA Coding Index262K tokens$0.60
71MoonshotAI: Kimi K2.5MoonshotAIMoonshotAI
46.8
AA Coding Index262K tokens$0.60
72Gemini 2.5 Pro Preview (Mar' 25)GoogleGoogle
46.7
AA Coding Index
73Gemini 2.5 Pro Preview (May' 25)GoogleGoogle
46.7
AA Coding Index$1.25
74GLM-4.6 (Reasoning)Z.aiZ.ai
45.8
AA Coding Index$0.55
75GLM-4.6 (Non-reasoning)Z.aiZ.ai
45.8
AA Coding Index$0.57
76GLM-4.7 (Reasoning)Z.aiZ.ai
45.3
AA Coding Index$0.60
77LongCat 2.0LongCat
45.3
AA Coding Index
78DeepSeek V3.2 Exp (Reasoning)DeepSeekDeepSeek
44.2
AA Coding Index$0.28
79DeepSeek V3.2DeepSeekDeepSeek
44.2
AA Coding Index164K tokens$0.28
80GPT-5.6 Luna (low)OpenAIOpenAI
44.2
AA Coding Index$1.00
81Claude 4.5 Haiku (Reasoning)AnthropicAnthropic
43.9
AA Coding Index$1.00
82DeepSeek V3.1 TerminusDeepSeekDeepSeek
43.5
AA Coding Index164K tokens$1.64
83Gemma 4 31BGoogleGoogle
43.4
AA Coding Index262K tokens$0.14
84Ring-2.6-1TInclusionAI
42.8
AA Coding Index$0.30
85Grok 4.3xAIxAI
42.2
AA Coding Index1.0M tokens$1.25
86o1OpenAIOpenAI
39.7
AA Coding Index200K tokens$15.00
87Step 3.7 FlashStepFun
39.6
AA Coding Index$0.20
88GPT-5.5 Instant (May 2026)OpenAIOpenAI
39.4
AA Coding Index$5.00
89GPT-5.5 Instant (June 2026)OpenAIOpenAI
39.4
AA Coding Index$5.00
90GPT-5.6 Luna (Non-reasoning)OpenAIOpenAI
39.3
AA Coding Index$1.00
91Gemma 4 26B A4B GoogleGoogle
39.3
AA Coding Index262K tokens$0.13
92GPT-5OpenAIOpenAI
37.8
AA Coding Index400K tokens$1.25
93GPT-5 (ChatGPT)OpenAIOpenAI
37.8
AA Coding Index$1.25
94NVIDIA Nemotron 3 Super 120B A12B (Reasoning)NVIDIANVIDIA
37.7
AA Coding Index1.0M tokens$0.25
95Claude 4 Sonnet (Reasoning)AnthropicAnthropic
37.6
AA Coding Index$3.00
96North Mini CodeCohereCohere
36.5
AA Coding Index
97Claude 3.7 Sonnet (thinking)AnthropicAnthropic
36.4
AA Coding Index200K tokens
98Claude 3.7 SonnetAnthropicAnthropic
36.4
AA Coding Index200K tokens$3.00
99Gemini 3.1 Flash Lite PreviewGoogleGoogle
34.7
AA Coding Index1.0M tokens$0.25
100Gemini 3.1 Flash LiteGoogleGoogle
34.7
AA Coding Index1.0M tokens$0.25
101Nova 2.0 Pro Preview (medium)AmazonAmazon
34.0
AA Coding Index$1.25
102o1-previewOpenAIOpenAI
34.0
AA Coding Index$16.50
103Gemini 2.5GoogleGoogle
33.3
AA Coding Index
104Gemini 2.5 ProGoogleGoogle
33.3
AA Coding Index1.0M tokens$1.25
105K-EXAONE (Reasoning)LG AI
32.1
AA Coding Index
106Devstral 2MistralMistral
31.3
AA Coding Index
107Inception: Mercury 2Inception
31.1
AA Coding Index128K tokens$0.25
108Gemma 4 12B (Reasoning)GoogleGoogle
31.0
AA Coding Index$0.10
109gpt-oss-120bOpenAIOpenAI
30.4
AA Coding Index131K tokens$0.15
110Claude 3.5 Sonnet (Oct '24)AnthropicAnthropic
30.2
AA Coding Index$3.00
111Devstral Small 2MistralMistral
29.3
AA Coding Index
112Qwen3.5 9B (Reasoning)AlibabaAlibaba
28.7
AA Coding Index
113Command A+CohereCohere
27.8
AA Coding Index
114Mistral: Mistral Small 4Mistral AIMistral AI
26.6
AA Coding Index262K tokens$0.15
115Mistral Small 3.1MistralMistral
26.3
AA Coding Index$0.10
116Claude 3.5 Sonnet (June '24)AnthropicAnthropic
26.0
AA Coding Index$3.00
117Trinity Large ThinkingArcee AI
25.8
AA Coding Index$0.23
118Gemini 2.0 Pro Experimental (Feb '25)GoogleGoogle
25.5
AA Coding Index
119Nemotron Cascade 2 30B A3BNVIDIANVIDIA
25.3
AA Coding Index
120Ling 2.6 FlashInclusion AI
25.3
AA Coding Index$0.10
121DeepSeek R1 (Jan '25)DeepSeekDeepSeek
24.6
AA Coding Index$1.68
122DeepSeek: R1DeepSeekDeepSeek
24.6
AA Coding Index164K tokens$0.70
123GPT-4o (March 2025, chatgpt-4o-latest)OpenAIOpenAI
24.2
AA Coding Index
124GPT-4o (2024-11-20)OpenAIOpenAI
24.2
AA Coding Index128K tokens$2.50
125OpenAI: GPT-4o (2024-05-13)OpenAIOpenAI
24.2
AA Coding Index128K tokens$5.00
126Gemini 2.0 FlashGoogleGoogle
24.1
AA Coding Index1.0M tokens
127Gemini 2.0 Flash Thinking Experimental (Jan '25)GoogleGoogle
24.1
AA Coding Index
128Gemini 1.5 Pro (Sep '24)GoogleGoogle
23.6
AA Coding Index
129EXAONE 4.5 33BLG AI
23.6
AA Coding Index
130HyperNova 60B 2605Multiverse Computing
23.2
AA Coding Index$0.04
131Nova 2.0 Lite (high)AmazonAmazon
23.0
AA Coding Index$0.30
132DeepSeek V3DeepSeekDeepSeek
23.0
AA Coding Index164K tokens$0.36
133Qwen3.5 4B (Reasoning)AlibabaAlibaba
22.6
AA Coding Index$0.03
134Qwen: Qwen3 235B A22B Instruct 2507AlibabaAlibaba
22.1
AA Coding Index262K tokens$0.70
135Qwen: Qwen3 235B A22B Thinking 2507AlibabaAlibaba
22.1
AA Coding Index131K tokens$0.15
136GPT-4 TurboOpenAIOpenAI
21.5
AA Coding Index128K tokens$10.00
137Magistral Medium 1.2Mistral AIMistral AI
21.3
AA Coding Index$2.00
138DeepSeek V3 0324DeepSeekDeepSeek
21.2
AA Coding Index$0.27
139K2 Think V2MBZUAI Institute of Foundation Models
21.0
AA Coding Index
140gpt-oss-20bOpenAIOpenAI
20.7
AA Coding Index131K tokens$0.07
141Mistral: Mistral Medium 3.1Mistral AIMistral AI
20.5
AA Coding Index131K tokens$0.40
142Qwen3.5 4B (Non-reasoning)AlibabaAlibaba
20.3
AA Coding Index$0.03
143GPT-4.1 MiniOpenAIOpenAI
20.2
AA Coding Index1.0M tokens$0.40
144Mistral Large 3MistralMistral
20.1
AA Coding Index$4.00
145Gemini 1.5 Pro (May '24)GoogleGoogle
19.8
AA Coding Index
146DiffusionGemma 26B A4BGoogleGoogle
19.7
AA Coding Index
147Claude 3 OpusAnthropicAnthropic
19.5
AA Coding Index$15.00
148Gemini 1.0 UltraGoogleGoogle
17.6
AA Coding Index
149Qwen3 Next 80B A3B (Reasoning)AlibabaAlibaba
17.4
AA Coding Index$0.50
150o3 Mini HighOpenAIOpenAI
16.3
AA Coding Index200K tokens$1.10
151Llama 4 MaverickMetaMeta
16.3
AA Coding Index1.0M tokens$0.27
152Solar Pro 3Upstage
16.2
AA Coding Index128K tokens
153Claude 3.5 HaikuAnthropicAnthropic
15.9
AA Coding Index200K tokens
154GPT-5 MiniOpenAIOpenAI
15.6
AA Coding Index400K tokens$0.25
155GPT-5 mini (minimal)OpenAIOpenAI
15.6
AA Coding Index$0.25
156Qwen3 32B (Reasoning)AlibabaAlibaba
15.3
AA Coding Index$0.70
157Magistral Small 1.2MistralMistral
14.7
AA Coding Index$0.50
158NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)NVIDIANVIDIA
14.4
AA Coding Index$0.05
159Ministral 3 14BMistralMistral
14.4
AA Coding Index$0.20
160Claude 2.1AnthropicAnthropic
14.0
AA Coding Index
161Qwen3 14B (Reasoning)AlibabaAlibaba
13.8
AA Coding Index$0.35
162Qwen3 14B (Non-reasoning)AlibabaAlibaba
13.8
AA Coding Index$0.35
163Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIANVIDIA
13.8
AA Coding Index$0.07
164OpenAI: GPT-4OpenAIOpenAI
13.1
AA Coding Index8K tokens$30.00
165Claude 2.0AnthropicAnthropic
12.9
AA Coding Index
166Mistral Small 3.2MistralMistral
12.5
AA Coding Index$0.10
167Qwen3 30B A3B 2507 (Reasoning)AlibabaAlibaba
12.1
AA Coding Index$0.20
168Llama 3.3 70B InstructMetaMeta
11.9
AA Coding Index131K tokens$0.59
169OpenAI: GPT-4o-miniOpenAIOpenAI
11.4
AA Coding Index128K tokens$0.15
170GPT-4.1 NanoOpenAIOpenAI
11.1
AA Coding Index1.0M tokens$0.10
171GPT-3.5 TurboOpenAIOpenAI
10.7
AA Coding Index$0.50
172Granite 4.1 30BIBM
10.4
AA Coding Index
173Gemma 3 27BGoogleGoogle
10.1
AA Coding Index262K tokens
174G9v3-3BAI9Stars
9.9
AA Coding Index
175Ministral 3 8BMistralMistral
9.7
AA Coding Index$0.15
176Nanbeige4.1-3BNanbeige
9.6
AA Coding Index
177Granite 4.1 8BIBM
9.5
AA Coding Index$0.05
178Gemma 4 E4B (Reasoning)GoogleGoogle
9.4
AA Coding Index$0.02
179Gemma 4 E4B (Non-reasoning)GoogleGoogle
9.4
AA Coding Index$0.02
180Qwen3 8B (Reasoning)AlibabaAlibaba
9.0
AA Coding Index$0.18
181Llama 4 ScoutMetaMeta
8.2
AA Coding Index1.3M tokens$0.18
182NVIDIA Nemotron 3 Nano 4BNVIDIANVIDIA
8.0
AA Coding Index
183Claude InstantAnthropicAnthropic
7.8
AA Coding Index
184Gemma 4 E2B (Reasoning)GoogleGoogle
7.2
AA Coding Index
185Gemma 3 12BGoogleGoogle
5.8
AA Coding Index131K tokens
186Llama 3.1 8B InstructMetaMeta
5.4
AA Coding Index16K tokens$0.07
187Ministral 3 3BMistralMistral
4.8
AA Coding Index$0.10
188Granite 4.1 3BIBM
4.7
AA Coding Index
189PALM-2GoogleGoogle
4.6
AA Coding Index
190Phi-4 Mini InstructMicrosoftMicrosoft
3.8
AA Coding Index
191Gemma 3n E4B InstructGoogleGoogle
3.2
AA Coding Index$0.06
192Qwen3.5 2B (Reasoning)AlibabaAlibaba
2.9
AA Coding Index
193Gemma 3 4BGoogleGoogle
2.7
AA Coding Index131K tokens
194Qwen3.5 0.8B (Non-reasoning)AlibabaAlibaba
1.2
AA Coding Index
195MiniCPM-V 4.6 1.3BOpenBMB
0.7
AA Coding Index
196Qwen3.5 0.8B (Reasoning)AlibabaAlibaba
0.0
AA Coding Index
197Gemini 3 Pro Preview (high)GoogleGoogle
92.0
LiveCodeBench$2.00
198Gemini 3 Flash Preview (Reasoning)GoogleGoogle
91.0
LiveCodeBench$0.50
199DeepSeek V3.2 SpecialeDeepSeekDeepSeek
90.0
LiveCodeBench164K tokens
200GPT-5.2OpenAIOpenAI
89.0
LiveCodeBench400K tokens$1.75
201Claude Opus 4.5 (Reasoning)AnthropicAnthropic
87.0
LiveCodeBench$5.00
202Gemini 3 Pro Preview (low)GoogleGoogle
86.0
LiveCodeBench$2.00
203o4 MiniOpenAIOpenAI
86.0
LiveCodeBench200K tokens$1.10
204o4 Mini HighOpenAIOpenAI
85.9
LiveCodeBench200K tokens$1.10
205GPT-5.1-CodexOpenAIOpenAI
85.0
LiveCodeBench400K tokens$1.25
206GPT-5.1-Codex-MaxOpenAIOpenAI
84.9
LiveCodeBench400K tokens$1.25
207GPT-5.1-Codex-MiniOpenAIOpenAI
84.0
LiveCodeBench400K tokens$0.25
208GPT-5 CodexOpenAIOpenAI
84.0
LiveCodeBench400K tokens$1.25
209Grok 4 FastxAIxAI
83.0
LiveCodeBench2.0M tokens$0.20
210MiniMax-M2MiniMax
83.0
LiveCodeBench205K tokens$0.30
211Grok 4xAIxAI
82.0
LiveCodeBench256K tokens$3.00
212Grok 4.1 FastxAIxAI
82.0
LiveCodeBench2.0M tokens
213MiniMax: MiniMax M2.1MiniMax
81.0
LiveCodeBench197K tokens$0.30
214o3OpenAIOpenAI
81.0
LiveCodeBench200K tokens$2.00
215ERNIE 5.0 Thinking PreviewBaidu
81.0
LiveCodeBench
216Apriel-v1.6-15B-ThinkerServiceNow
81.0
LiveCodeBench
217o3 ProOpenAIOpenAI
80.8
LiveCodeBench200K tokens$20.00
218Gemini 3 Flash Preview (Non-reasoning)GoogleGoogle
80.0
LiveCodeBench$0.50
219GPT-5 NanoOpenAIOpenAI
79.0
LiveCodeBench400K tokens$0.05
220INTELLECT-3Prime Intellect
78.0
LiveCodeBench131K tokens
221DeepSeek V3.1DeepSeekDeepSeek
78.0
LiveCodeBench164K tokens$0.56
222Doubao Seed CodeByteDance Seed
77.0
LiveCodeBench
223Seed-OSS-36B-InstructByteDance Seed
77.0
LiveCodeBench$0.21
224GPT-5.2 ProOpenAIOpenAI
76.5
LiveBench Coding400K tokens$21.00
225Claude Sonnet 4.5AnthropicAnthropic
76.1
LiveBench Coding1.0M tokens$3.00
226KAT-Coder-Pro V1KwaiKAT
75.0
LiveCodeBench
227EXAONE 4.0 32B (Reasoning)LG AI Research
75.0
LiveCodeBench
228Gemini 3.5GoogleGoogle
74.6
LiveBench Coding
229Claude Opus 4.5AnthropicAnthropic
74.0
LiveCodeBench200K tokens$5.00
230GLM-4.5 (Reasoning)Z.aiZ.ai
74.0
LiveCodeBench131K tokens
231Llama Nemotron Super 49B v1.5 (Reasoning)NVIDIANVIDIA
74.0
LiveCodeBench$0.40
232Qwen3 VL 32B (Reasoning)AlibabaAlibaba
74.0
LiveCodeBench$0.70
233Apriel-v1.5-15B-ThinkerServiceNow
73.0
LiveCodeBench
234o3 MiniOpenAIOpenAI
72.0
LiveCodeBench200K tokens$1.10
235Falcon-H1R-7BTII UAE
72.0
LiveCodeBench
236NVIDIA Nemotron Nano 9B V2 (Reasoning)NVIDIANVIDIA
72.0
LiveCodeBench$0.04
237Gemini 2.5 Flash Preview (Sep '25) (Reasoning)GoogleGoogle
71.0
LiveCodeBench
238MiniMax M1 80kMiniMax
71.0
LiveCodeBench$0.55
239Grok 3 MinixAIxAI
70.0
LiveCodeBench131K tokens$0.30
240Gemini 2.5 Flash Preview (Reasoning)GoogleGoogle
70.0
LiveCodeBench
241Olmo 3.1 32B ThinkAllen Institute for AI
70.0
LiveCodeBench
242Qwen3 VL 30B A3B (Reasoning)AlibabaAlibaba
70.0
LiveCodeBench$0.20
243NVIDIA Nemotron Nano 9B V2 (Non-reasoning)NVIDIANVIDIA
70.0
LiveCodeBench131K tokens$0.05
244Cogito v2.1 (Reasoning)Deep Cogito
69.0
LiveCodeBench$1.25
245K2-V2 (high)MBZUAI Institute of Foundation Models
69.0
LiveCodeBench
246Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning)GoogleGoogle
69.0
LiveCodeBench$0.10
247NVIDIA Nemotron Nano 12B v2 VL (Reasoning)NVIDIANVIDIA
69.0
LiveCodeBench$0.20
248Hermes 4 - Llama-3.1 405B (Reasoning)Nous Research
69.0
LiveCodeBench$1.00
249Ling-1TInclusionAI
68.0
LiveCodeBench
250Qwen3 Omni 30B A3B (Reasoning)AlibabaAlibaba
68.0
LiveCodeBench$0.25
251Qwen: Qwen3 Next 80B A3B InstructAlibabaAlibaba
68.0
LiveCodeBench262K tokens$0.50
252GLM-4.5-AirZ.aiZ.ai
68.0
LiveCodeBench$0.17
253Olmo 3 32B ThinkAllenAI
67.0
LiveCodeBench66K tokens
254GPT-5.2-CodexOpenAIOpenAI
66.9
LiveCodeBench400K tokens$1.75
255MiniMax M1 40kMiniMax
66.0
LiveCodeBench
256Nova 2.0 Omni (medium)AmazonAmazon
66.0
LiveCodeBench$0.30
257Grok Code Fast 1xAIxAI
66.0
LiveCodeBench256K tokens
258Mi:dm K 2.5 ProKorea Telecom
66.0
LiveCodeBench
259Claude 4.1 Opus (Non-reasoning)AnthropicAnthropic
65.4
LiveCodeBench$15.00
260xAI: Grok Build 0.1xAIxAI
65.4
LiveBench Coding256K tokens$1.00
261Claude 4.1 Opus (Reasoning)AnthropicAnthropic
65.0
LiveCodeBench$15.00
262Qwen3 VL 235B A22B (Reasoning)AlibabaAlibaba
65.0
LiveCodeBench$0.70
263Qwen3 Max (Preview)AlibabaAlibaba
65.0
LiveCodeBench$1.20
264Hermes 4 - Llama-3.1 70B (Reasoning)Nous Research
65.0
LiveCodeBench$0.13
265Motif-2-12.7B-ReasoningMotif Technologies
65.0
LiveCodeBench
266Claude 4 Opus (Reasoning)AnthropicAnthropic
64.0
LiveCodeBench$15.00
267Ring-1TInclusionAI
64.0
LiveCodeBench
268Llama 3.1 Nemotron Ultra 253B v1 (Reasoning)NVIDIANVIDIA
64.0
LiveCodeBench$0.60
269Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning)GoogleGoogle
64.0
LiveCodeBench$0.10
270Qwen3 4B 2507 (Reasoning)AlibabaAlibaba
64.0
LiveCodeBench
271QwQ 32BAlibabaAlibaba
63.0
LiveCodeBench$0.66
272HyperCLOVA X SEED Think (32B)Naver
63.0
LiveCodeBench
273Ring-flash-2.0InclusionAI
63.0
LiveCodeBench$0.14
274Qwen3 235B A22B (Reasoning)AlibabaAlibaba
62.0
LiveCodeBench$0.70
275Solar Pro 2 (Non-reasoning)Upstage
62.0
LiveCodeBench
276Olmo 3 7B ThinkAllen Institute for AI
62.0
LiveCodeBench
277MoonshotAI: Kimi K2 0905MoonshotAIMoonshotAI
61.0
LiveCodeBench262K tokens$0.60
278GLM-4.5V (Reasoning)Z.aiZ.ai
60.0
LiveCodeBench$0.60
279Qwen: Qwen3 VL 235B A22B InstructAlibabaAlibaba
59.0
LiveCodeBench262K tokens$0.70
280Qwen3 Coder 480B A35B InstructAlibabaAlibaba
59.0
LiveCodeBench$1.50
281Nova 2.0 Omni (low)AmazonAmazon
59.0
LiveCodeBench$0.30
282Ling-flash-2.0InclusionAI
59.0
LiveCodeBench$0.14
283Gemini 2.5 Flash LiteGoogleGoogle
59.0
LiveCodeBench1.0M tokens$0.10
284o1-miniOpenAIOpenAI
58.0
LiveCodeBench
285Mi:dm K 2.5 Pro PreviewKorea Telecom
58.0
LiveCodeBench
286GPT-5 (minimal)OpenAIOpenAI
56.0
LiveCodeBench$1.25
287Kimi K2Moonshot AIMoonshot AI
56.0
LiveCodeBench131K tokens$0.57
288DeepSeek V3.2 Exp (Non-reasoning)DeepSeekDeepSeek
55.0
LiveCodeBench$0.28
289Hermes 4 - Llama-3.1 405B (Non-reasoning)Nous Research
55.0
LiveCodeBench$1.00
290Claude Opus 4AnthropicAnthropic
54.0
LiveCodeBench200K tokens$15.00
291Qwen3 Max Thinking (Preview)AlibabaAlibaba
54.0
LiveCodeBench$1.20
292K2-V2 (medium)MBZUAI Institute of Foundation Models
54.0
LiveCodeBench
293Magistral Medium 1MistralMistral
53.0
LiveCodeBench
294GPT-5.3-CodexOpenAIOpenAI
53.0
SciCode400K tokens$1.75
295Qwen3 30B A3B 2507 InstructAlibabaAlibaba
52.0
LiveCodeBench$0.20
296Exaone 4.0 1.2B (Non-reasoning)LG AI Research
52.0
LiveCodeBench
297Claude Opus 4.6 (Adaptive Reasoning, Max Effort)AnthropicAnthropic
52.0
SciCode$5.00
298Claude Haiku 4.5AnthropicAnthropic
51.0
LiveCodeBench200K tokens$1.00
299Qwen: Qwen3 VL 32B InstructAlibabaAlibaba
51.0
LiveCodeBench131K tokens$0.70
300Qwen3 30B A3B (Reasoning)AlibabaAlibaba
51.0
LiveCodeBench$0.20
301Magistral Small 1MistralMistral
51.0
LiveCodeBench
302DeepSeek R1 0528 Qwen3 8BDeepSeekDeepSeek
51.0
LiveCodeBench
303Gemini 2.5 FlashGoogleGoogle
50.0
LiveCodeBench1.0M tokens$0.30
304Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning)NVIDIANVIDIA
49.0
LiveCodeBench
305Qwen: Qwen3 VL 30B A3B InstructAlibabaAlibaba
48.0
LiveCodeBench131K tokens$0.20
306Baidu: ERNIE 4.5 300B A47B Baidu
47.0
LiveCodeBench123K tokens$0.28
307GPT-5 nano (minimal)OpenAIOpenAI
47.0
LiveCodeBench$0.05
308EXAONE 4.0 32B (Non-reasoning)LG AI Research
47.0
LiveCodeBench
309Qwen3 4B (Reasoning)AlibabaAlibaba
47.0
LiveCodeBench
310Qwen3.6 Max PreviewAlibabaAlibaba
47.0
SciCode$1.30
311Claude Sonnet 4.6AnthropicAnthropic
47.0
SciCode1.0M tokens$3.00
312GPT-4.1OpenAIOpenAI
46.0
LiveCodeBench1.0M tokens$2.00
313Solar Pro 2 (Preview) (Reasoning)Upstage
46.0
LiveCodeBench
314Grok 4.20xAIxAI
46.0
SciCode2.0M tokens$2.00
315GLM-5 (Reasoning)Z.aiZ.ai
46.0
SciCode203K tokens$1.00
316Claude Opus 4.6AnthropicAnthropic
46.0
SciCode1.0M tokens$5.00
317Claude Sonnet 4AnthropicAnthropic
45.0
LiveCodeBench1.0M tokens$3.00
318Grok 4.20 0309 (Reasoning)xAIxAI
45.0
SciCode$2.00
319Reka Flash 3Reka Flash 3
44.0
LiveCodeBench66K tokens$0.20
320GLM 5V Turbo (Reasoning)Z.aiZ.ai
44.0
SciCode203K tokens
321GLM-5-TurboZ.aiZ.ai
44.0
SciCode203K tokens
322Grok 3xAIxAI
43.0
LiveCodeBench131K tokens
323Ling-mini-2.0InclusionAI
43.0
LiveCodeBench
324Xiaomi: MiMo-V2-ProXiaomi
43.0
SciCode1.0M tokens
325MiniMax: MiniMax M2.5MiniMax
43.0
SciCode197K tokens$0.30
326Qwen3 Omni 30B A3B InstructAlibabaAlibaba
42.0
LiveCodeBench$0.25
327GLM-4.6V (Non-reasoning)Z.aiZ.ai
41.0
LiveCodeBench$0.30
328Gemini 2.5 Flash Preview (Non-reasoning)GoogleGoogle
41.0
LiveCodeBench
329Hy3-preview (Reasoning)Tencent
41.0
SciCode262K tokens$0.07
330Qwen3.5 Omni PlusAlibabaAlibaba
41.0
SciCode$0.40
331Mistral: Mistral Medium 3Mistral AIMistral AI
40.0
LiveCodeBench131K tokens$0.40
332Qwen: Qwen3 Coder 30B A3B InstructAlibabaAlibaba
40.0
LiveCodeBench160K tokens$0.45
333MiMo-V2-Omni-0327Xiaomi
40.0
SciCode
334Step 3.5 FlashStepFun
40.0
SciCode$0.10
335Solar Pro 2 (Preview) (Non-reasoning)Upstage
39.0
LiveCodeBench
336Step 3.5 FlashStepFun
39.0
SciCode262K tokens$0.10
337DeepSeek R1 Distill Qwen 14BDeepSeekDeepSeek
38.0
LiveCodeBench
338Kimi Linear 48B A3B InstructKimi
38.0
LiveCodeBench
339Qwen3 4B 2507 InstructAlibabaAlibaba
38.0
LiveCodeBench
340GLM-5 (Non-reasoning)Z.aiZ.ai
38.0
SciCode$1.00
341Ling-2.6-1TInclusion AI
37.0
SciCode$0.30
342Xiaomi: MiMo-V2-OmniXiaomi
37.0
SciCode262K tokens
343Qwen2.5 MaxAlibabaAlibaba
36.0
LiveCodeBench
344NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning)NVIDIANVIDIA
36.0
LiveCodeBench262K tokens$0.05
345Qwen3 VL 8B (Reasoning)AlibabaAlibaba
35.0
LiveCodeBench$0.18
346GLM-4.5V (Non-reasoning)Z.aiZ.ai
35.0
LiveCodeBench$0.60
347NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning)NVIDIANVIDIA
35.0
LiveCodeBench$0.20
348Mistral: Devstral MediumMistral AIMistral AI
34.0
LiveCodeBench131K tokens
349QwQ 32B-PreviewAlibabaAlibaba
34.0
LiveCodeBench
350GLM-4.7-Flash (Reasoning)Z.aiZ.ai
34.0
SciCode$0.07
351Qwen: Qwen3 VL 8B InstructAlibabaAlibaba
33.0
LiveCodeBench131K tokens$0.18
352GPT-4o (ChatGPT)OpenAIOpenAI
33.0
SciCode
353Qwen: Qwen3 30B A3B Thinking 2507AlibabaAlibaba
32.2
LiveCodeBench131K tokens$0.08
354GPT-4o (2024-08-06)OpenAIOpenAI
32.0
LiveCodeBench128K tokens$2.50
355Amazon: Nova Premier 1.0AmazonAmazon
32.0
LiveCodeBench1.0M tokens$2.50
356Qwen: Qwen3 30B A3B Instruct 2507AlibabaAlibaba
32.0
LiveCodeBench262K tokens$0.20
357Qwen3 VL 4B (Reasoning)AlibabaAlibaba
32.0
LiveCodeBench
358OpenAI: GPT-4oOpenAIOpenAI
31.0
LiveCodeBench128K tokens$2.50
359Llama 3.1 Instruct 405BMetaMeta
31.0
LiveCodeBench$2.50
360Nova 2.0 Omni (Non-reasoning)AmazonAmazon
31.0
LiveCodeBench$0.30
361Qwen3 1.7B (Reasoning)AlibabaAlibaba
31.0
LiveCodeBench
362Step3 VL 10BStepFun
31.0
SciCode
363Qwen2.5 Coder 32B InstructAlibabaAlibaba
30.0
LiveCodeBench33K tokens
364SonarPerplexityPerplexity
30.0
LiveCodeBench127K tokens
365Sarvam M (Reasoning)Sarvam
30.0
LiveCodeBench
366Llama 3.1 Tulu3 405BAllen Institute for AI
29.0
LiveCodeBench
367Mistral Large 2 (Nov '24)MistralMistral
29.0
LiveCodeBench
368Qwen3 32B (Non-reasoning)AlibabaAlibaba
29.0
LiveCodeBench$0.70
369Llama Nemotron Super 49B v1.5 (Non-reasoning)NVIDIANVIDIA
29.0
LiveCodeBench$0.40
370Qwen3 VL 4B InstructAlibabaAlibaba
29.0
LiveCodeBench
371JT-35B-FlashChina Mobile
29.0
SciCode
372Llama 3.3 Nemotron Super 49B v1 (Reasoning)NVIDIANVIDIA
28.0
LiveCodeBench
373Qwen2.5 72B InstructAlibabaAlibaba
28.0
LiveCodeBench33K tokens$0.47
374Llama 3.3 Nemotron Super 49B v1 (Non-reasoning)NVIDIANVIDIA
28.0
LiveCodeBench
375Sonar Reasoning ProPerplexityPerplexity
28.0
LiveCodeBench128K tokens
376LongCat Flash LiteLongCat
28.0
SciCode
377DeepSeek: R1 Distill Qwen 32BDeepSeekDeepSeek
27.0
LiveCodeBench128K tokens
378R1 Distill Llama 70BDeepSeekDeepSeek
27.0
LiveCodeBench8K tokens$0.70
379Hermes 4 - Llama-3.1 70B (Non-reasoning)Nous Research
27.0
LiveCodeBench$0.13
380Grok 2 (Dec '24)xAIxAI
27.0
LiveCodeBench
381Gemini 1.5 Flash (Sep '24)GoogleGoogle
27.0
LiveCodeBench
382Mistral Large 2 (Jul '24)MistralMistral
27.0
LiveCodeBench131K tokens$2.00
383Olmo 3 7B InstructAllen Institute for AI
27.0
LiveCodeBench$0.10
384JT-MINIChina Mobile
27.0
SciCode
385Solar Open 100B (Reasoning)Upstage
27.0
SciCode
386Mistral: Pixtral Large 2411Mistral AIMistral AI
26.0
LiveCodeBench131K tokens
387Devstral Small (May '25)MistralMistral
26.0
LiveCodeBench
388Qwen3.5 Omni FlashAlibabaAlibaba
26.0
SciCode$0.10
389Sarvam 105B (high)Sarvam
26.0
SciCode$0.04
390Devstral Small (Jul '25)MistralMistral
25.0
LiveCodeBench131K tokens
391Mistral Small 3MistralMistral
25.0
LiveCodeBench$0.10
392Qwen2.5 Instruct 32BAlibabaAlibaba
25.0
LiveCodeBench
393Granite 4.0 H SmallIBM
25.0
LiveCodeBench$0.06
394Grok BetaxAIxAI
24.0
LiveCodeBench
395Mistral: SabaMistral AIMistral AI
24.0
SciCode33K tokens
396Llama 3.1 70B InstructMetaMeta
23.0
LiveCodeBench131K tokens$0.56
397Microsoft: Phi 4MicrosoftMicrosoft
23.0
LiveCodeBench16K tokens$0.13
398Nova ProAmazonAmazon
23.0
LiveCodeBench$0.80
399Qwen3 4B (Non-reasoning)AlibabaAlibaba
23.0
LiveCodeBench
400DeepSeek R1 Distill Llama 8BDeepSeekDeepSeek
23.0
LiveCodeBench
401Gemini 1.5 Flash-8BGoogleGoogle
22.0
LiveCodeBench
402Gemini 2.0 Flash (experimental)GoogleGoogle
21.0
LiveCodeBench
403Llama 3.2 Instruct 90B (Vision)MetaMeta
21.0
LiveCodeBench$2.04
404Jamba Reasoning 3BAI21 Labs
21.0
LiveCodeBench
405DeepHermes 3 - Mistral 24B Preview (Non-reasoning)Nous Research
20.0
LiveCodeBench
406Llama 3 70B InstructMetaMeta
20.0
LiveCodeBench8K tokens$0.65
407Gemini 1.5 Flash (May '24)GoogleGoogle
20.0
LiveCodeBench
408Qwen3 8B (Non-reasoning)AlibabaAlibaba
20.0
LiveCodeBench$0.18
409Gemma 4 E2B (Non-reasoning)GoogleGoogle
20.0
SciCode
410Gemini 2.0 Flash-Lite (Feb '25)GoogleGoogle
19.0
LiveCodeBench
411Hermes 3 - Llama-3.1 70BNous Research
19.0
LiveCodeBench$0.70
412Sarvam 30BSarvam
19.0
SciCode$0.03
413Gemini 2.0 Flash-Lite (Preview)GoogleGoogle
18.0
LiveCodeBench
414Claude 3 SonnetAnthropicAnthropic
18.0
LiveCodeBench$3.00
415AI21: Jamba Large 1.7AI21 Labs
18.0
LiveCodeBench256K tokens$2.00
416Granite 4.0 MicroIBM
18.0
LiveCodeBench131K tokens
417Tri-21B-think PreviewTrillion Labs
18.0
SciCode
418Mistral LargeMistral AIMistral AI
17.8
LiveCodeBench128K tokens$2.00
419Llama 3.1 Nemotron 70B InstructNVIDIANVIDIA
17.0
LiveCodeBench131K tokens$1.20
420Jamba 1.6 LargeAI21 Labs
17.0
LiveCodeBench$2.00
421Nova LiteAmazonAmazon
17.0
LiveCodeBench$0.06
422Tri-21B-ThinkTrillion Labs
17.0
SciCode
423Olmo 3.1 32B InstructAllenAI
17.0
SciCode66K tokens
424GLM-4.6V (Reasoning)Z.aiZ.ai
16.0
LiveCodeBench$0.30
425Qwen2 Instruct 72BAlibabaAlibaba
16.0
LiveCodeBench
426DeepSeek Coder V2 Lite InstructDeepSeekDeepSeek
16.0
LiveCodeBench
427Mixtral 8x22B InstructMistralMistral
15.0
LiveCodeBench
428Anthropic: Claude 3 HaikuAnthropicAnthropic
15.0
LiveCodeBench200K tokens$0.25
429LFM2 8B A1BLiquid AI
15.0
LiveCodeBench
430Mistral: Mixtral 8x22B InstructMistral AIMistral AI
14.8
LiveCodeBench66K tokens$2.00
431Mistral Small (Sep '24)MistralMistral
14.0
LiveCodeBench$0.20
432Jamba 1.5 LargeAI21 Labs
14.0
LiveCodeBench$2.00
433Gemma 3n E4B Instruct Preview (May '25)GoogleGoogle
14.0
LiveCodeBench
434Nova MicroAmazonAmazon
14.0
LiveCodeBench$0.04
435Qwen2.5 Coder Instruct 7B AlibabaAlibaba
13.0
LiveCodeBench
436Phi-4 Multimodal InstructMicrosoftMicrosoft
13.0
LiveCodeBench
437Granite 3.3 8B (Non-reasoning)IBM
13.0
LiveCodeBench$0.03
438Qwen3 1.7B (Non-reasoning)AlibabaAlibaba
13.0
LiveCodeBench
439Molmo2-8BAllen Institute for AI
13.0
SciCode
440Cohere: Command R+ (08-2024)CohereCohere
12.2
LiveCodeBench128K tokens$2.50
441Command-R+ (Apr '24)CohereCohere
12.0
LiveCodeBench$3.00
442Gemini 1.0 ProGoogleGoogle
12.0
LiveCodeBench
443Phi-3 Mini Instruct 3.8BMicrosoftMicrosoft
12.0
LiveCodeBench
444Granite 4.0 H 1BIBM
12.0
LiveCodeBench
445Qwen3 0.6B (Reasoning)AlibabaAlibaba
12.0
LiveCodeBench
446OpenChat 3.5 (1210)OpenChat
12.0
LiveCodeBench
447Mistral Small (Feb '24)MistralMistral
11.0
LiveCodeBench$0.15
448Llama 3.2 11B Vision InstructMetaMeta
11.0
LiveCodeBench131K tokens$0.36
449LFM2-24B-A2BLiquidAI
11.0
SciCode33K tokens
450Llama 3 8B InstructMetaMeta
10.0
LiveCodeBench8K tokens$0.04
451Mistral MediumMistralMistral
10.0
LiveCodeBench$1.50
452Llama 2 Chat 13BMetaMeta
10.0
LiveCodeBench
453LFM 40BLiquid AI
10.0
LiveCodeBench
454Gemma 3n E2B InstructGoogleGoogle
10.0
LiveCodeBench
455Llama 2 Chat 70BMetaMeta
10.0
LiveCodeBench
456DBRX InstructDatabricks
9.0
LiveCodeBench
457DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)Nous Research
9.0
LiveCodeBench
458Llama 3.2 3B InstructMetaMeta
8.0
LiveCodeBench80K tokens
459LFM2 2.6BLiquid AI
8.0
LiveCodeBench
460LFM2.5-8B-A1BLiquid AI
8.0
SciCode
461Jamba 1.6 MiniAI21 Labs
7.0
LiveCodeBench$0.20
462OLMo 2 32BAllen Institute for AI
7.0
LiveCodeBench
463DeepSeek R1 Distill Qwen 1.5BDeepSeekDeepSeek
7.0
LiveCodeBench
464Qwen3 0.6B (Non-reasoning)AlibabaAlibaba
7.0
LiveCodeBench
465Mistral: Mixtral 8x7B InstructMistral AIMistral AI
7.0
LiveCodeBench33K tokens$0.45
466Jamba 1.7 MiniAI21 Labs
6.0
LiveCodeBench
467Jamba 1.5 MiniAI21 Labs
6.0
LiveCodeBench$0.20
468Apertus 70B InstructSwiss AI Initiative
6.0
SciCode$0.82
469Granite 4.0 1BIBM
5.0
LiveCodeBench
470Command-R (Mar '24)CohereCohere
5.0
LiveCodeBench$0.50
471Mistral 7B InstructMistralMistral
5.0
LiveCodeBench$0.25
472OLMo 2 7BAllen Institute for AI
4.0
LiveCodeBench
473Molmo 7B-DAllen Institute for AI
4.0
LiveCodeBench
474MiniCPM5-1B (Non-reasoning)OpenBMB
4.0
SciCode
475Tiny Aya GlobalCohereCohere
4.0
SciCode
476LFM2.5-1.2B-ThinkingLiquid AI
4.0
SciCode
477Apertus 8B InstructSwiss AI Initiative
4.0
SciCode$0.10
478LFM2.5-VL-1.6BLiquid AI
3.0
SciCode
479LFM2 1.2BLiquid AI
2.0
LiveCodeBench
480Granite 4.0 H 350MIBM
2.0
LiveCodeBench
481Llama 3.2 1B InstructMetaMeta
2.0
LiveCodeBench60K tokens
482Granite 4.0 350MIBM
2.0
LiveCodeBench
483Gemma 3 1B InstructGoogleGoogle
2.0
LiveCodeBench
484LFM2.5-1.2B-InstructLiquid AI
2.0
SciCode
485Gemma 3 270MGoogleGoogle
0.0
LiveCodeBench
486Llama 2 Chat 7BMetaMeta
0.0
LiveCodeBench$0.05

+ 182 modelos sem benchmark de coding disponível.Ver todos os modelos

Guia Completo: IA para Programação em 2026

O Estado da IA para Código em 2026

A inteligência artificial transformou radicalmente o desenvolvimento de software nos últimos anos. Em 2026, modelos de linguagem (LLMs) são capazes de gerar código funcional em dezenas de linguagens, resolver bugs em projetos reais e até criar aplicações completas a partir de descrições em linguagem natural. O SWE-bench — o benchmark mais rigoroso para coding — avalia modelos em tarefas reais de engenharia de software extraídas de issues do GitHub.

SWE-bench: O Benchmark de Referência

O SWE-bench (Software Engineering Benchmark) é considerado o padrão ouro para avaliar capacidade de coding de LLMs. Diferente de benchmarks acadêmicos como HumanEval (que testa funções isoladas), o SWE-bench apresenta issues reais de repositórios populares como Django, Flask, scikit-learn e requests. O modelo precisa entender o contexto do projeto, localizar os arquivos relevantes e gerar um patch que resolva o bug — simulando o trabalho real de um desenvolvedor.

A versão “Verified” do SWE-bench (SWE-bench Verified) é curada por engenheiros humanos para garantir que cada tarefa tem uma solução clara e verificável. Os scores neste benchmark são particularmente informativos porque correlacionam fortemente com a experiência real de uso para coding.

HumanEval e LiveCodeBench

HumanEval, criado pela OpenAI, testa a capacidade de gerar funções Python a partir de docstrings. É um benchmark mais simples que o SWE-bench, mas útil para avaliar fluência básica em código. LiveCodeBench adiciona uma camada de complexidade ao testar com problemas que são atualizados regularmente, reduzindo o risco de contaminação (quando o modelo já viu as respostas durante o treinamento).

Como Escolher o Melhor Modelo para Código

A escolha do modelo ideal depende do caso de uso específico. Para autocompletar código em tempo real (Cursor, Copilot), velocidade e latência são mais importantes que score máximo — modelos das classes mini e flash costumam oferecer a melhor relação velocidade/qualidade. Para geração de projetos completos ou debug complexo, modelos frontier como Claude Fable 5, GPT-5.5 e Gemini 3.5 tendem a ser mais adequados, apesar do custo maior.

Para equipes que precisam de controle sobre os dados (compliance, segurança), modelos open source como DeepSeek Coder, Code Llama e StarCoder permitem deploy on-premises com performance competitiva. A decisão entre proprietário e open source envolve tradeoffs de custo, latência, privacidade e qualidade.

Ferramentas de Coding com IA

As principais ferramentas de desenvolvimento assistido por IA em 2026 incluem Cursor, GitHub Copilot, Windsurf e fluxos agentic integrados ao editor. Cada ferramenta usa diferentes modelos por baixo, e a qualidade do código gerado depende diretamente da capacidade do LLM utilizado.

Para desenvolvedores brasileiros, um fator importante é a capacidade do modelo de entender comentários, nomes de variáveis e documentação em português — algo que varia significativamente entre modelos e que não é capturado pelos benchmarks tradicionais em inglês.

Tendências para 2026 e Além

As tendências mais relevantes em IA para código incluem: agentes autônomos de engenharia (que resolvem tarefas complexas sem supervisão), geração de testes automatizados, refatoração inteligente, e integração nativa com pipelines de CI/CD. A fronteira está se movendo de “assistente de código” para “engenheiro autônomo”, com modelos cada vez mais capazes de navegar codebases grandes e tomar decisões arquiteturais.

Perguntas Frequentes

Qual é a melhor IA para programar?

Em 2026, os modelos que lideram em benchmarks de código são GPT-5.6 Sol (xhigh), Claude Opus 5, GPT-5.6 Sol (max). No entanto, a melhor escolha depende do caso de uso: autocompletar código, geração de projetos completos, debug ou code review.

ChatGPT ou Claude para código?

Hoje as opções mais fortes costumam orbitar Claude Fable 5 / Opus 4.8, GPT-5.5 e modelos da classe Gemini 3.5. Claude tende a brilhar em refactors longos; GPT é forte em geração rápida e tool use. O teste final precisa ser feito no seu próprio repositório.

O que é o SWE-bench?

SWE-bench (Software Engineering Benchmark) avalia a capacidade de modelos de resolver issues reais de repositórios open source no GitHub. É considerado o benchmark mais realista para coding, pois testa resolução de bugs em projetos reais, não exercícios acadêmicos.

Quais LLMs gratuitas são boas para código?

Modelos open source como DeepSeek Coder, Qwen Coder e Code Llama oferecem excelente performance em coding sem custo de API. Podem ser rodados localmente via Ollama ou acessados gratuitamente em plataformas como Together AI e Groq.

Cursor ou GitHub Copilot?

Cursor e Copilot são IDEs/extensões que usam LLMs por baixo. Cursor permite escolher o modelo (Claude, GPT, etc.), enquanto Copilot usa modelos da OpenAI. A qualidade do código gerado depende mais do modelo escolhido do que da ferramenta em si.

Explorar Outras Categorias