August 25, 2026
DrivenBench
1.0
Compare AI models for investment agents across capability, cost, and latency, using real-world investment workflows and tasks.
Why DrivenBench?DrivenBench score
Capability score vs. total benchmark cost upper bound (log scale). The solid line shows the observed Pareto frontier.
Leaderboard
| Rank | Model | ScoreCapability score | Capability passes | Cost upper bound | Latency median / p90 | Details |
|---|---|---|---|---|---|---|
| =1=1 | 93.9%93.9% | 155/165 | $57.69 | 30s/70s | ||
| =1=1 | 93.9%93.9% | 155/165 | $59.42 | 59s/169s | ||
| 33 | 90.9%90.9% | 150/165 | $160.37 | 40s/90s | ||
| 44 | 89.7%89.7% | 148/165 | $30.64 | 24s/61s | ||
| 55 | 89.1%89.1% | 147/165 | $73.65 | 75s/158s | ||
| 66 | 87.3%87.3% | 144/165 | $25.64 | 23s/62s | ||
| 77 | 85.5%85.5% | 141/165 | $92.72 | 27s/61s | ||
| 88 | 81.8%81.8% | 135/165 | $13.74 | 23s/44s | ||
| 99 | 79.4%79.4% | 131/165 | $3.63 | 19s/39s | ||
| 1010 | 77.6%77.6% | 128/165 | $2.97 | 20s/46s | ||
| 1111 | 76.4%76.4% | 126/165 | $33.54 | 16s/32s |
Choose the model that fits your investment workflow, then put it to work in Driven.
Try Driven Free