ModelVerdict

Independent benchmarks of AI models on business work

Which model can do your paperwork, in your language, and what will it really cost?

Every week we run commercial and open models through the same business tasks: invoices, customer emails, personal data and contracts, with summaries and questions about documents next. Every number links to the actual model output and what it really costs, including fixing its mistakes.

Last run September 28, 2026 · 10 models · documents in: English (US), Czech, German

Leaderboards by task

Sorted by accuracy. Costs already include the language tax on tokens. Bar: 95%.

#ModelAccuracyPer 1,000Language taxData in the EU
1GPT-6 Sol
OpenAI
99.9%meets the bar
$9.65×1.00Yes
2Claude Sonnet 5
Anthropic
99.8%meets the bar
$8.84×1.00Yes
3Gemini 3.8 Flash
Google
99.8%meets the bar
$3.84×1.00No
4GPT-6 Luna
OpenAI
99.7%meets the bar
$0.51×1.00Yes
5DeepSeek V4.1 Flash
DeepSeek · open weights
99.4%meets the bar
$0.52×1.00Self-host only
6Gemini 3.5 Flash Lite
Google
99.2%meets the bar
$0.87×1.00Yes
7Mistral Medium 3.5
Mistral
97.8%meets the bar
$4.37×1.00Yes
8Gemma 4 31B
Google · open weights
97.6%meets the bar
$0.17×1.00Self-host only
9Mistral Small 4
Mistral · open weights
95.6%meets the bar
$0.41—Yes
10Qwen3.8 Flash
Qwen
82.4%below the bar
$1.15×1.00No

Full leaderboard and filters →