Overview

评估多步骤软推理能力,包含逻辑、数学和高效推理三个领域各250个问题

Metrics

MetricUnitDirection
accuracy%↑ Higher is better

Sources

Model Score Ranking

#ModelVendorScore
1Claude 3.5 Sonnet (2024-10-22)anthropic71.8
2o1openai71.8
3Gemini 1.5 Pro 002google71.6
4Gemini 2.0 Flash Thinkinggoogle70.8
5Sonar Reasoningother69.4
6Gemini 2.0 Flashgoogle69.3
7Qwen2 72Balibaba68.9
8GPT-4 0125 Previewopenai65.8
9Sonar Largeother65.6
10GPT-4 Vision Previewopenai64.5
11Gemini 1.0 Ultragoogle64.3
12Mistral Mediummistral64.3
13Claude 3 Sonnetanthropic64.1
14o1 Previewopenai64.1
15Qwen1.5 110Balibaba63.9
16Yi Large Turboother63.7
17Claude 3 Sonnet (2024-02-29)anthropic63.5
18GPT-4o (2024-08-06)openai62.9
19GPT-4o (2024-05-13)openai62.8
20Llama 3.3 70Bmeta62.3
21Yi Visionother61.4
22GPT-4 Turboopenai61.1
23Gemini 1.5 Flash 002google59.8
24Orca 2 13Bother59.8
25Yi Largeother59.4
26Falcon 180Bother59.1
27Qwen1.5 14Balibaba58.7
28GPT-4 1106 Previewopenai58.4
29Claude 3 Haikuanthropic58.1
30StableLM 2 12Bother58.1
31Microsoft WizardLM 2 8x22Bother58
32Gemini 1.0 Progoogle57.8
33Grok-2 Visionxai57.7
34Hermes 3 Llama 3.1 405Bother57.7
35Zephyr ORPO 141B Alphaother57.7
36Command Nightlycohere57.5
37Nous Hermes 2 Mixtral 8x7Bother57.4
38Claude 3 Opus (2024-02-29)anthropic57.3
39Jamba 1.5other57.1
40Command Rcohere57
41Jamba Instructother56.9
42WizardLM Team WizardLM 2 8x22Bother56.8
43GLM-4 Flashother56
44Gemini 1.0 Flashgoogle55.9
45Hermes 3 Llama 3.1 70Bother55.4
46Sonar Hugeother55.4
47GPT-4 32Kopenai54.9
48Qwen2.5 32Balibaba54.9
49DBRX Instructother54.7
50Gemini 1.5 Flash-8B 002google54.7
51GPT-4openai54.6
52Qwen1.5 32Balibaba54.5
53Llama 3.1 Nemotron 70Bmeta54.1
54o1 miniopenai54
55DBRX Baseother53.8
56Llama 3.2 90B Visionmeta53.5
57Mixtral 8x7Bmistral53.5
58DeepSeek V2deepseek53.4
59Microsoft WizardMath 7B v1other53.4
60Claude 3 Opusanthropic53.3
61GPT-4o miniopenai53.1
62Microsoft WizardCoder Python 34Bother53.1
63Gemma 2 27Bgoogle53
64Claude 3.5 Haikuanthropic52.3
65Mistral Largemistral52.1
66DeepSeek V2 Chatdeepseek52
67DeepSeek LLM 67Bdeepseek51.9
68StableCode 3Bother51.8
69StarCoder2 15Bother51.7
70Grok-2 Minixai51.4
71GLM-4 Plusother51.3
72GPT-4 Visionopenai51.2
73Llama 3 70Bmeta51.2
74Command R+ (08-2024)cohere51.1
75Jamba 1.5 Largeother51.1
76Claude 3 Haiku (2024-03-07)anthropic51
77Yi 1.5 34Bother51
78Qwen1.5 72Balibaba50.8
79DeepSeek Coder 7Bdeepseek50.5
80Qwen2.5 14Balibaba50.5
81Qwen2 57Balibaba50.5
82Phi-3.5 MoEother50.3
83Code Bisongoogle49.8
84DeepSeek Math 7Bdeepseek49.8
85GLM-4 9B Chatother49.5
86Microsoft WizardLM 2 7Bother49.3
87Mistral Smallmistral49.3
88Command R (08-2024)cohere49.1
89Phi-1other49
90Nous Hermes 2 Yi 34Bother48.8
91Jamba 1.5 Miniother48.1
92Llama 3.1 70Bmeta48.1
93Mistral 7B v0.3mistral48.1
94NVIDIA Llama 3.1 Nemotron 70Bother47.7
95Sonar Smallother47.6
96Code Llama 13Bmeta47.3
97Codestralmistral47.2
98Codestral Mambamistral47
99Command R7Bcohere47
100Qwen2.5 7Balibaba46.9
101OLMo 7Bother46.8
102GLM-4 Airother46.7
103Yi 1.5 6Bother46.6
104Yi 1.5 9Bother46.4
105Gemini 1.5 Flashgoogle46.1
106Llama 3 8Bmeta46.1
107Mistral Small 3mistral46
108DeepSeek Coder 33Bdeepseek45.8
109Gemini 1.5 Flash-8Bgoogle45.6
110Code Llama 34Bmeta44.9
111Grok Vision Betaxai44.5
112StarChat2 15B v0.1other44.4
113OLMo 7B SFTother44.3
114Gemma 2 9Bgoogle44.1
115Phi-3 Mediumother43.9
116Qwen2 7Balibaba43.9
117Grok Betaxai43.7
118ChatGLM3 6Bother42.9
119Code Llama 7Bmeta42.9
120Jurassic-2 Ultraother42.9
121Llemma 7Bother42.6
122OLMo 1.7 7Bother42.6
123Nous Capybara 7Bother42.5
124OLMo 7B Instructother42.5
125MPT 7Bother42.4
126Mathstral 7Bmistral42.3
127Flan-T5 XXLother42.2
128Llama 3.1 8Bmeta41.7
129Llama 3.2 11B Visionmeta41.7
130Mistral Tinymistral41.4
131Text Bisongoogle40.8
132Command Lightcohere40.7
133Zephyr 7B Alphaother40.5
134GPT-3.5openai40.4
135DeepSeek Coder V2deepseek40.3
136Flan-T5 XLother40.2
137Code Llama 70Bmeta40.1
138Open-Platypusother40
139Nous Hermes 2 Solar 10.7Bother39.6
140TigerBot 70B Chatother39.3
141OLMo 2 1124 7Bother39.2
142Phi-4other38.6
143GLM-4V 9Bother38.5
144Mistral Nemomistral38.5
145StarCoder2 3Bother38.4
146StarCoder2 7Bother38.3
147OpenChat 3.6 8Bother38.2
148Falcon 40Bother38.1
149Hermes 3 Llama 3.1 8Bother38
150Baichuan2 13B Chatother37.9
151Qwen2.5 1.5Balibaba37.8
152Claude 2.1anthropic37.7
153Mistral 7B v0.1mistral37.4
154Flan-UL2other37.2
155Ministral 8Bmistral37
156Yi 34Bother36.8
157Mistral 7B v0.2mistral36.7
158StableLM 3 4Bother36.6
159Phi-3 Smallother36.5
160Gemma 7Bgoogle36.2
161Llama 2 70Bmeta36.1
162Llama 2 7Bmeta35.3
163Claude 2anthropic35.2
164Llama Guard 2 8Bmeta34.9
165Capybara 1.5Bother34
166Llama 2 13Bmeta33.8
167Claude Instant 1anthropic33.5
168Phi-3 Miniother33.5
169Baichuan2 7B Chatother33.1
170Chat Bisongoogle32.8
171OpenChat 3.5 1210other32.6
172Zephyr 7B Betaother32.6
173Jurassic-2 Midother32.3
174Phi-3 Visionother32.1
175Nous Capybara 34Bother31.7
176Phi-3.5 Miniother31.2
177Gemma 2Bgoogle30.6
178Phi-1.5other30.3
179Argilla Notus 7B v1other30.1
180Yi 6Bother29.5
181Ministral 3Bmistral29.4
182Llama 3.2 3Bmeta29.3
183PaLM 2google29.3
184GPT-3.5 Turboopenai29.2
185GPT-3.5 Turbo 16Kopenai29.1
186MPT 30Bother28.9
187Qwen2.5 0.5Balibaba28.9
188Qwen2 1.5Balibaba27.8
189Llama 3.2 1Bmeta26.6
190StableLM 2 1.6Bother26.4
191Phi-2other23.7
192StableLM Zephyr 3Bother23.5
193Qwen2.5 3Balibaba23.3
194Llama Guard 3 8Bmeta22.8
195Embed English v3cohere0

MUSR

Описание

评估多步骤软推理能力,包含逻辑、数学和高效推理三个领域各250个问题

Основные характеристики

КатегорияЛицензияПоследнее обновление
reasoningMIT2024-02-01

Бенчмарк

Единица
%

Источники

Официальный URL