Gradio

Powered by Inspect and Inspect Evals, the Vector Evaluation Leaderboard presents an evaluation of leading frontier models across a comprehensive suite of benchmarks. Go beyond the summary metrics: click through to interactive reporting for each model and benchmark to explore sample-level performance and detailed traces.


Mistral-Large-Instruct-2407	66.89	78.66	60.04	94.69	94.45	86.66	98.23	43.18	74.24	90.48	86.05	80.34	64.82	41.51


DeepSeek-R1	83.83	95.68	92.72	95.45	96.67	91.79	98.74	70.45	--	--	--	--	--	--
Meta-Llama-3.1-70B-Instruct	66.89	78.66	60.04	94.69	94.45	86.66	98.23	43.18	88.11	86.99	86.05	80.34	--	--
Mistral-Large-Instruct-2407	69.42	86.59	65.74	93.78	94.37	85.48	98.53	47.35	74.24	90.48	82.85	80.36	--	--
Command R+	44.12	62.2	26.26	78.17	85.07	74.9	93.77	31.94	74.36	79.55	77.8	69.51	--	--
Claude-3.5-Sonnet	77.63	94.51	79.42	96.21	96.93	90.21	99.16	60.98	89.78	92.28	89.58	86.65	64.82	41.51
Gemini-1.5-Flash	59.93	74.39	45.2	85.82	93.09	78.85	98.4	40.4	75.1	85.57	76.81	77.15	57.02	16.98
Gemini-1.5-Pro	75.64	87.2	85.2	96.13	96.33	87.69	98.78	57.83	88.01	91.24	89.82	84.67	63.05	35.85
GPT-4o	74.51	90.85	70.54	94.47	96.33	90.13	99.16	51.01	75.12	92.43	87.8	84.35	59.03	35.85
GPT-4o-mini	63.96	85.98	63.3	91.81	92.49	75.3	97.94	38.38	80.66	87.5	84.19	76.98	53.96	18.87
o1	84.47	96.95	95.9	94.16	97.87	93.92	99.12	75.51	--	--	--	--	80.64	69.81
o3-mini	79.25	98.17	96.91	94.54	96.42	84.93	97.56	73.65	--	--	--	--	--	--


Gemini-1.5-Pro	Basic Agent	33.82	85.57	61.54	14.77	80.07	6.72


Claude-3.5-Sonnet	Basic Agent	33.82	85.57	61.54	14.77	80.07	6.72
Gemini-1.5-Pro	Basic Agent	13.82	52.91	23.08	28.99	59.61	0.4
GPT-4o	Basic Agent	16.61	63.8	23.08	49.95	82.49	1.2
o1	Basic Agent	41.09	84.81	46.15	8.78	72.35	0.36
o3-mini	Basic Agent	27.03	82.78	38.46	12.42	54.29	0.24

Vector Institute

The Vector Institute is dedicated to advancing the field of artificial intelligence through cutting-edge research and application. Our mission is to drive excellence and innovation in AI, fostering a community of researchers, developers, and industry partners.

🎯 Benchmarks

This leaderboard showcases performance across a comprehensive suite of benchmarks, designed to rigorously evaluate different aspects of AI model capabilities. Let's explore the benchmarks we use:

Inspect Evals

This leaderboard leverages Inspect Evals to power evaluation. Inspect Evals is an open-source repository built upon the Inspect AI framework. Developed in collaboration between the Vector Institute, Arcadia Impact and the UK AI Security Institute, Inspect Evals provides a comprehensive suite of high-quality benchmarks spanning diverse domains like coding, mathematics, cybersecurity, reasoning, and general knowledge.

Transparent and Detailed Insights

All evaluations presented on this leaderboard are run using Inspect Evals. To facilitate in-depth analysis and promote transparency, we provide Inspect Logs for every benchmark run. These logs offer sample and trace level reporting, allowing the community to explore the granular details of model performance.

⚙️ Base Benchmarks

These benchmarks assess fundamental reasoning and knowledge capabilities of models.

Benchmark	Description
ARC-Easy / ARC-Challenge	Multiple-choice science questions.
DROP	Comprehension benchmark evaluating advanced reasoning capability.
WinoGrande	Commonsense reasoning challenge.
GSM8K	Grade-school math word problems testing math capability & multi-step reasoning.
HellaSwag	Commonsense reasoning task.
HumanEval	Evaluates code generation and reasoning in a programming context.
IFEval	Specialized benchmark for instruction following.
MATH	Challenging questions sourced from math competitions.
MMLU / MMLU-Pro	Multi-subject multiple-choice tests of advanced knowledge.
GPQA-Diamond	Question-answering benchmark assessing deeper reasoning.
MMMU (Multi-Choice / Open-Ended)	Multi-modal tasks testing structured & open responses.

🚀 Agentic Benchmarks

These benchmarks go beyond basic reasoning and evaluate more advanced, autonomous, or "agentic" capabilities of models, such as planning and interaction.

Benchmark	Description
GAIA	Evaluates autonomous reasoning, planning, problem-solving for question answering.
InterCode-CTF	Capture-the-flag challenge testing cyber-security skills.
In-House-CTF	Capture-the-flag challenge testing cyber-security skills.
AgentHarm / AgentHarm-Benign	Measures harmfulness of LLM agents (and benign behavior baseline).
SWE-Bench-Verified	Tests AI agent ability to solve software engineering tasks.