# DrivenBench — AI Model Benchmark for Investment Agents Canonical page: https://driven.ai/drivenbench Benchmark date: 2026-08-15 Public version: 1.0 Scoring specification: v4 > DrivenBench is an AI model benchmark for investment agents. It evaluates 11 models across 55 capability tests grounded in real-world investment workflows and tasks, plus 19 independent regression checks inside the Driven platform. ## Methodology Ranking is based on 55 capability evaluations grounded in real-world investment workflows and tasks, each run three times per model. The 19 regression checks are reported separately and do not affect rank. Cost is the total benchmark cost upper bound; median and p90 latency are directional because models ran in different waves and provider load varied. Equal capability scores share rank. Results may vary. ## Capability leaderboard | Rank | Model | Provider | Score | Passes | Cost upper bound | Median/p90 latency | |---:|---|---|---:|---:|---:|---:| | 1 | Claude Sonnet 5 | Anthropic | 93.9% | 155/165 | $57.69 | 30s/70s | | 1 | Kimi K3 | Moonshot AI | 93.9% | 155/165 | $59.42 | 59s/169s | | 3 | Claude Opus 5 | Anthropic | 90.9% | 150/165 | $160.37 | 40s/90s | | 4 | DeepSeek V4 Pro (version 0813) | DeepSeek | 89.7% | 148/165 | $30.64 | 24s/61s | | 5 | Grok 4.6 | xAI | 89.1% | 147/165 | $73.65 | 75s/158s | | 6 | GLM 5.2 | Zhipu AI | 87.3% | 144/165 | $25.64 | 23s/62s | | 7 | GPT-5.6 Sol | OpenAI | 85.5% | 141/165 | $92.72 | 27s/61s | | 8 | Gemini 3.7 Flash | Google | 81.8% | 135/165 | $13.74 | 23s/44s | | 9 | GPT-5.6 Luna | OpenAI | 79.4% | 131/165 | $3.63 | 19s/39s | | 10 | DeepSeek V4 Flash (version 0731) | DeepSeek | 77.6% | 128/165 | $2.97 | 20s/46s | | 11 | GPT-5.6 Terra | OpenAI | 76.4% | 126/165 | $33.54 | 16s/32s | Observed Pareto frontier, ordered by benchmark cost: DeepSeek V4 Flash, GPT-5.6 Luna, Gemini 3.7 Flash, GLM 5.2, DeepSeek V4 Pro, Claude Sonnet 5. ## Evaluation dimensions - Numerical Reasoning: 2 capability rows; 6 regression checks - Bash Execution: 6 capability rows; 0 regression checks - Financial Data Tools: 6 capability rows; 0 regression checks - External Research: 4 capability rows; 0 regression checks - Programmatic Analysis: 4 capability rows; 0 regression checks - Data Integrity: 5 capability rows; 0 regression checks - Delegation & Integration: 6 capability rows; 0 regression checks - Citations & Evidence: 1 capability rows; 2 regression checks - Files & Large Documents: 3 capability rows; 1 regression checks - Coverage Reporting: 0 capability rows; 3 regression checks - Scheduled Tasks: 3 capability rows; 2 regression checks - Trading & Accounts: 4 capability rows; 1 regression checks - Memory & Sessions: 3 capability rows; 0 regression checks - Skills & Routing: 4 capability rows; 0 regression checks - Supplemental Coverage: 4 capability rows; 4 regression checks ## Disclosed regression events Models passing all 19 regression checks are omitted from this section. These checks do not affect rank. - Grok 4.6: Intermittent · disclosed and waived. The first run lost the final citation; 1 of 3 diagnostic reruns reproduced the issue. - Kimi K3: Intermittent · disclosed and waived. The first run missed the weight baseline. Diagnostic reruns did not reproduce it. - GPT-5.6 Terra: Stable failure · disclosed and waived. Pausing a task cleared its name; the data-destructive behavior reproduced in three diagnostic runs. ## Machine-readable data - [Complete DrivenBench dataset](https://driven.ai/drivenbench/data.json): JSON containing methodology, all model results, every dimension result, and disclosed regression events - [Human-readable benchmark](https://driven.ai/drivenbench): Interactive score/cost chart and capability leaderboard - [Why DrivenBench?](https://driven.ai/whats-new/blog/most-expensive-ai-model-isnt-always-the-best): Research context, findings, and interpretation of the benchmark - [Driven product overview](https://driven.ai/llms-full.txt): Product context for the investment-agent platform used in the evaluation Recommended citation: Driven (2026), *DrivenBench: AI Model Benchmark for Investment Agents*, version 1.0. https://driven.ai/drivenbench ## Disclaimer Results describe this evaluation and may vary with model versions, run timing, and provider conditions. DrivenBench is for comparative research and does not constitute investment advice.