ModelVerdict

Leaderboard

Who does your task best

Sorted by the share of correct fields. Cost is the real token spend scaled to 1,000 documents. Bar: 95% correct fields; below it, manual fixes cost more than a cheaper model saves.

Compare two models side by side →

Invoice extraction · Czech · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$9.4103.6 sec100.0%meets the bar—
2Gemini 3.8 Flash
Google
100.0%99%$4.2673.5 sec100.0%meets the barwrong value
3DeepSeek V4.1 Flash
DeepSeek · open weights
99.9%99%$0.50135.0 sec100.0%meets the barwrong value
4GPT-6 Sol
OpenAI
99.9%99%$10.12133.9 sec100.0%meets the barwrong value
5Gemini 3.5 Flash Lite
Google
99.5%93%$1.03671.8 sec100.0%meets the barinvented value
6GPT-6 Luna
OpenAI
99.3%93%$0.51673.2 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
96.1%67%$0.2133310.4 sec100.0%meets the barwrong value
8Mistral Small 4
Mistral · open weights
93.6%56%$0.444402.0 sec100.0%below the barwrong value
9Mistral Medium 3.5
Mistral
93.5%63%$4.693672.6 sec100.0%below the barwrong value
10Qwen3.8 Flash
Qwen
91.8%77%$1.0223317.0 sec96.7%below the barinvalid JSON

Invoice extraction · Czech · PDF text (45 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$4.8702.7 sec100.0%meets the bar—
2DeepSeek V4.1 Flash
DeepSeek · open weights
100.0%100%$0.5006.0 sec100.0%meets the bar—
3Gemini 3.5 Flash Lite
Google
100.0%100%$0.9101.1 sec100.0%meets the bar—
4Gemini 3.8 Flash
Google
100.0%100%$3.4602.8 sec100.0%meets the bar—
5GPT-6 Luna
OpenAI
100.0%100%$0.2102.6 sec100.0%meets the bar—
6GPT-6 Sol
OpenAI
100.0%100%$4.0202.8 sec100.0%meets the bar—
7Mistral Medium 3.5
Mistral
99.8%98%$2.97221.5 sec100.0%meets the barmissing field
8Gemma 4 31B
Google · open weights
99.5%93%$0.24679.0 sec100.0%meets the barmissing field
9Mistral Small 4
Mistral · open weights
96.8%76%$0.262441.4 sec100.0%meets the barinvented value
10Qwen3.8 Flash
Qwen
92.0%53%$0.7446719.9 sec100.0%below the barwrong value

Invoice extraction · Czech · Clean image (45 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$13.2003.8 sec100.0%meets the bar—
2DeepSeek V4.1 Flash
DeepSeek · open weights
100.0%100%$0.4904.4 sec100.0%meets the bar—
3Gemini 3.8 Flash
Google
100.0%100%$4.6104.0 sec100.0%meets the bar—
4GPT-6 Sol
OpenAI
100.0%100%$14.5703.9 sec100.0%meets the bar—
5GPT-6 Luna
OpenAI
99.7%96%$0.73443.6 sec100.0%meets the barwrong value
6Gemini 3.5 Flash Lite
Google
99.5%93%$1.07672.0 sec100.0%meets the barinvented value
7Gemma 4 31B
Google · open weights
95.4%62%$0.2037811.0 sec100.0%meets the barwrong value
8Qwen3.8 Flash
Qwen
92.7%84%$1.1115615.9 sec95.6%below the barinvalid JSON
9Mistral Small 4
Mistral · open weights
92.3%44%$0.515561.9 sec100.0%below the barinvented value
10Mistral Medium 3.5
Mistral
92.1%60%$5.374002.5 sec100.0%below the barwrong value

Invoice extraction · Czech · Scan (45 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$10.4903.7 sec100.0%meets the bar—
2Gemini 3.8 Flash
Google
99.8%98%$4.61223.5 sec100.0%meets the barwrong value
3GPT-6 Sol
OpenAI
99.8%98%$11.89224.2 sec100.0%meets the barwrong value
4DeepSeek V4.1 Flash
DeepSeek · open weights
99.7%96%$0.49444.8 sec100.0%meets the barwrong value
5Gemini 3.5 Flash Lite
Google
99.0%87%$1.081331.8 sec100.0%meets the barwrong value
6GPT-6 Luna
OpenAI
98.3%87%$0.591333.4 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
94.9%56%$0.2044411.9 sec100.0%below the barwrong value
8Mistral Small 4
Mistral · open weights
92.1%47%$0.515332.6 sec100.0%below the barwrong value
9Qwen3.8 Flash
Qwen
90.9%89%$1.1711115.4 sec93.3%below the barinvalid JSON
10Mistral Medium 3.5
Mistral
89.1%36%$5.416442.9 sec100.0%below the barwrong value

Invoice extraction · Czech · Photo (15 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$8.4503.9 sec100.0%meets the bar—
2DeepSeek V4.1 Flash
DeepSeek · open weights
100.0%100%$0.5405.4 sec100.0%meets the bar—
3Gemini 3.8 Flash
Google
100.0%100%$4.5303.4 sec100.0%meets the bar—
4Gemini 3.5 Flash Lite
Google
99.5%93%$1.11671.8 sec100.0%meets the barinvented value
5GPT-6 Sol
OpenAI
99.5%93%$9.76674.9 sec100.0%meets the barwrong value
6GPT-6 Luna
OpenAI
99.0%87%$0.471333.3 sec100.0%meets the barwrong value
7Mistral Small 4
Mistral · open weights
92.8%60%$0.544002.9 sec100.0%below the barwrong value
8Gemma 4 31B
Google · open weights
91.8%33%$0.2066711.0 sec100.0%below the barwrong value
9Mistral Medium 3.5
Mistral
91.8%53%$5.654673.0 sec100.0%below the barwrong value
10Qwen3.8 Flash
Qwen
91.3%87%$1.1113315.4 sec100.0%below the barmissing field

Invoice extraction · German · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemini 3.5 Flash Lite
Google
100.0%100%$1.0301.5 sec100.0%meets the bar—
2GPT-6 Sol
OpenAI
100.0%99%$10.6574.9 sec100.0%meets the barwrong value
3DeepSeek V4.1 Flash
DeepSeek · open weights
99.5%94%$0.57606.4 sec100.0%meets the barwrong value
4GPT-6 Luna
OpenAI
99.2%91%$0.54873.7 sec100.0%meets the barwrong value
5Claude Sonnet 5
Anthropic
99.0%87%$10.421333.8 sec100.0%meets the barmissing field
6Mistral Medium 3.5
Mistral
97.4%75%$4.722532.5 sec100.0%meets the barinvented value
7Gemini 3.8 Flash
Google
97.2%96%$7.15404.7 sec97.3%meets the barinvalid JSON
8Gemma 4 31B
Google · open weights
94.9%64%$0.203607.7 sec100.0%below the barwrong value
9Mistral Small 4
Mistral · open weights
94.5%49%$0.445082.0 sec100.0%below the barinvented value
10Qwen3.8 Flash
Qwen
93.5%87%$1.1212719.0 sec94.0%below the barinvalid JSON

Invoice extraction · German · PDF text (45 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1DeepSeek V4.1 Flash
DeepSeek · open weights
100.0%100%$0.4906.7 sec100.0%meets the bar—
2Gemini 3.5 Flash Lite
Google
100.0%100%$0.9101.1 sec100.0%meets the bar—
3GPT-6 Luna
OpenAI
100.0%100%$0.2203.0 sec100.0%meets the bar—
4GPT-6 Sol
OpenAI
100.0%100%$3.8503.1 sec100.0%meets the bar—
5Qwen3.8 Flash
Qwen
100.0%100%$0.99027.8 sec100.0%meets the bar—
6Gemma 4 31B
Google · open weights
99.5%93%$0.23677.0 sec100.0%meets the barmissing field
7Mistral Medium 3.5
Mistral
99.5%93%$2.96671.4 sec100.0%meets the barmissing field
8Claude Sonnet 5
Anthropic
99.3%91%$5.89893.3 sec100.0%meets the barmissing field
9Mistral Small 4
Mistral · open weights
97.9%80%$0.242001.3 sec100.0%meets the barinvented value
10Gemini 3.8 Flash
Google
97.8%98%$6.24223.9 sec97.8%meets the barinvalid JSON

Invoice extraction · German · Clean image (45 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemini 3.5 Flash Lite
Google
100.0%100%$1.0701.7 sec100.0%meets the bar—
2GPT-6 Sol
OpenAI
100.0%100%$15.2605.6 sec100.0%meets the bar—
3DeepSeek V4.1 Flash
DeepSeek · open weights
99.7%96%$0.55445.7 sec100.0%meets the barwrong value
4GPT-6 Luna
OpenAI
99.7%96%$0.77444.1 sec100.0%meets the barwrong value
5Claude Sonnet 5
Anthropic
99.0%87%$13.271333.8 sec100.0%meets the barmissing field
6Mistral Medium 3.5
Mistral
96.6%69%$5.433112.5 sec100.0%meets the barinvented value
7Gemini 3.8 Flash
Google
95.2%91%$8.81895.1 sec95.6%meets the barinvalid JSON
8Gemma 4 31B
Google · open weights
93.9%51%$0.194897.4 sec100.0%below the barwrong value
9Mistral Small 4
Mistral · open weights
93.4%36%$0.516431.8 sec100.0%below the barinvented value
10Qwen3.8 Flash
Qwen
88.4%82%$1.2617816.7 sec88.9%below the barinvalid JSON

Invoice extraction · German · Scan (45 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemini 3.5 Flash Lite
Google
100.0%100%$1.0801.6 sec100.0%meets the bar—
2GPT-6 Sol
OpenAI
99.8%98%$13.15225.1 sec100.0%meets the barwrong value
3DeepSeek V4.1 Flash
DeepSeek · open weights
99.0%87%$0.611337.2 sec100.0%meets the barwrong value
4Claude Sonnet 5
Anthropic
98.6%82%$11.881783.8 sec100.0%meets the barmissing field
5GPT-6 Luna
OpenAI
98.3%82%$0.661784.1 sec100.0%meets the barwrong value
6Gemini 3.8 Flash
Google
97.8%98%$6.65225.0 sec97.8%meets the barinvalid JSON
7Mistral Medium 3.5
Mistral
96.9%67%$5.453332.8 sec100.0%meets the barinvented value
8Qwen3.8 Flash
Qwen
96.9%87%$1.1513316.6 sec97.8%meets the barinvalid JSON
9Mistral Small 4
Mistral · open weights
92.5%33%$0.526672.4 sec100.0%below the barinvented value
10Gemma 4 31B
Google · open weights
91.8%51%$0.194898.7 sec100.0%below the barwrong value

Invoice extraction · German · Photo (15 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemini 3.5 Flash Lite
Google
100.0%100%$1.1101.6 sec100.0%meets the bar—
2Gemini 3.8 Flash
Google
100.0%100%$6.4504.8 sec100.0%meets the bar—
3GPT-6 Sol
OpenAI
100.0%100%$9.7305.4 sec100.0%meets the bar—
4DeepSeek V4.1 Flash
DeepSeek · open weights
99.5%93%$0.70676.4 sec100.0%meets the barswapped dates
5Claude Sonnet 5
Anthropic
99.0%87%$11.051333.8 sec100.0%meets the barwrong value
6GPT-6 Luna
OpenAI
98.5%80%$0.502003.5 sec100.0%meets the barwrong value
7Mistral Medium 3.5
Mistral
94.9%60%$5.684002.7 sec100.0%below the barwrong value
8Mistral Small 4
Mistral · open weights
93.5%46%$0.535392.5 sec100.0%below the barinvented value
9Gemma 4 31B
Google · open weights
93.3%53%$0.194679.1 sec100.0%below the barwrong value
10Qwen3.8 Flash
Qwen
79.0%67%$0.9933315.6 sec80.0%below the barinvalid JSON

Invoice extraction · English (US) · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1GPT-6 Sol
OpenAI
99.9%99%$9.6574.1 sec100.0%meets the barwrong value
2Claude Sonnet 5
Anthropic
99.8%99%$8.84133.2 sec100.0%meets the barwrong value
3Gemini 3.8 Flash
Google
99.8%98%$3.84203.2 sec100.0%meets the barwrong value
4GPT-6 Luna
OpenAI
99.7%97%$0.51333.4 sec100.0%meets the barinvented value
5DeepSeek V4.1 Flash
DeepSeek · open weights
99.4%93%$0.52675.7 sec100.0%meets the barwrong value
6Gemini 3.5 Flash Lite
Google
99.2%92%$0.87801.7 sec100.0%meets the barwrong value
7Mistral Medium 3.5
Mistral
97.8%79%$4.372072.3 sec100.0%meets the barwrong value
8Gemma 4 31B
Google · open weights
97.6%81%$0.171878.9 sec100.0%meets the barwrong value
9Mistral Small 4
Mistral · open weights
95.6%58%$0.414201.8 sec100.0%meets the barwrong rate
10Qwen3.8 Flash
Qwen
82.4%57%$1.1542717.4 sec90.7%below the barinvalid JSON

Invoice extraction · English (US) · PDF text (45 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1DeepSeek V4.1 Flash
DeepSeek · open weights
100.0%100%$0.5207.4 sec100.0%meets the bar—
2Gemini 3.8 Flash
Google
100.0%100%$2.9202.3 sec100.0%meets the bar—
3Gemma 4 31B
Google · open weights
100.0%100%$0.1807.4 sec100.0%meets the bar—
4GPT-6 Luna
OpenAI
100.0%100%$0.1802.7 sec100.0%meets the bar—
5GPT-6 Sol
OpenAI
100.0%100%$3.2803.2 sec100.0%meets the bar—
6Claude Sonnet 5
Anthropic
99.8%98%$3.91222.4 sec100.0%meets the barmissing field
7Mistral Medium 3.5
Mistral
99.8%98%$2.24221.2 sec100.0%meets the barwrong value
8Gemini 3.5 Flash Lite
Google
99.4%96%$0.72441.0 sec100.0%meets the barwrong value
9Mistral Small 4
Mistral · open weights
98.4%84%$0.191561.2 sec100.0%meets the barwrong value
10Qwen3.8 Flash
Qwen
91.7%49%$0.9751118.7 sec100.0%below the barwrong rate

Invoice extraction · English (US) · Clean image (45 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$12.1903.3 sec100.0%meets the bar—
2GPT-6 Sol
OpenAI
100.0%100%$13.8104.2 sec100.0%meets the bar—
3Gemini 3.5 Flash Lite
Google
99.6%96%$0.93441.8 sec100.0%meets the barinvented value
4Gemini 3.8 Flash
Google
99.6%96%$4.13443.5 sec100.0%meets the barwrong value
5GPT-6 Luna
OpenAI
99.6%96%$0.73443.8 sec100.0%meets the barinvented value
6DeepSeek V4.1 Flash
DeepSeek · open weights
99.2%91%$0.50894.6 sec100.0%meets the barinvented value
7Gemma 4 31B
Google · open weights
98.2%84%$0.161569.2 sec100.0%meets the barwrong value
8Mistral Medium 3.5
Mistral
98.0%82%$5.251782.3 sec100.0%meets the barwrong value
9Mistral Small 4
Mistral · open weights
95.6%56%$0.504441.7 sec100.0%meets the barwrong rate
10Qwen3.8 Flash
Qwen
77.2%60%$1.1340015.2 sec86.7%below the barinvalid JSON

Invoice extraction · English (US) · Scan (45 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$10.6603.3 sec100.0%meets the bar—
2GPT-6 Sol
OpenAI
100.0%100%$12.0404.3 sec100.0%meets the bar—
3Gemini 3.8 Flash
Google
99.8%98%$4.36223.2 sec100.0%meets the barwrong value
4GPT-6 Luna
OpenAI
99.6%96%$0.63443.7 sec100.0%meets the barwrong value
5DeepSeek V4.1 Flash
DeepSeek · open weights
99.0%89%$0.551115.4 sec100.0%meets the barwrong value
6Gemini 3.5 Flash Lite
Google
98.8%87%$0.931331.8 sec100.0%meets the barinvented value
7Gemma 4 31B
Google · open weights
96.6%73%$0.162679.8 sec100.0%meets the barwrong value
8Mistral Medium 3.5
Mistral
96.0%62%$5.273782.8 sec100.0%meets the barwrong value
9Mistral Small 4
Mistral · open weights
93.5%38%$0.506222.3 sec100.0%below the barwrong rate
10Qwen3.8 Flash
Qwen
81.4%60%$1.3540017.4 sec91.1%below the barinvalid JSON

Invoice extraction · English (US) · Photo (15 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemini 3.8 Flash
Google
100.0%100%$4.1303.4 sec100.0%meets the bar—
2DeepSeek V4.1 Flash
DeepSeek · open weights
99.4%93%$0.52675.3 sec100.0%meets the barwrong value
3GPT-6 Luna
OpenAI
99.4%93%$0.48673.7 sec100.0%meets the barwrong value
4GPT-6 Sol
OpenAI
99.4%93%$9.08674.9 sec100.0%meets the barwrong value
5Claude Sonnet 5
Anthropic
98.8%93%$8.09673.2 sec100.0%meets the barwrong value
6Gemini 3.5 Flash Lite
Google
98.8%87%$0.941331.7 sec100.0%meets the barwrong value
7Mistral Medium 3.5
Mistral
96.4%67%$5.443332.8 sec100.0%meets the barwrong value
8Mistral Small 4
Mistral · open weights
93.3%47%$0.515332.6 sec100.0%below the barwrong value
9Gemma 4 31B
Google · open weights
92.1%40%$0.1660010.4 sec100.0%below the barwrong value
10Qwen3.8 Flash
Qwen
72.7%67%$1.2033319.6 sec73.3%below the barinvalid JSON

Email triage · Czech · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemini 3.8 Flash
Google
99.2%93%$3.06732.6 sec100.0%meets the barwrong value
2GPT-6 Sol
OpenAI
99.0%91%$3.19932.8 sec100.0%meets the barwrong value
3Qwen3.8 Flash
Qwen
98.5%87%$0.4513313.0 sec100.0%meets the barwrong value
4Claude Sonnet 5
Anthropic
98.2%83%$4.911674.7 sec100.0%meets the barwrong value
5GPT-6 Luna
OpenAI
97.9%81%$0.181872.6 sec100.0%meets the barwrong value
6DeepSeek V4.1 Flash
DeepSeek · open weights
97.4%77%$0.302274.2 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
97.1%75%$0.162532.5 sec100.0%meets the barwrong value
8Gemini 3.5 Flash Lite
Google
95.8%64%$0.513600.7 sec100.0%meets the barwrong value
9Mistral Medium 3.5
Mistral
94.5%55%$1.954470.8 sec100.0%below the barwrong value
10Mistral Small 4
Mistral · open weights
87.7%32%$0.126850.8 sec100.0%below the barwrong value

Email triage · Czech · Email (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemini 3.8 Flash
Google
99.2%93%$3.06732.6 sec100.0%meets the barwrong value
2GPT-6 Sol
OpenAI
99.0%91%$3.19932.8 sec100.0%meets the barwrong value
3Qwen3.8 Flash
Qwen
98.5%87%$0.4513313.0 sec100.0%meets the barwrong value
4Claude Sonnet 5
Anthropic
98.2%83%$4.911674.7 sec100.0%meets the barwrong value
5GPT-6 Luna
OpenAI
97.9%81%$0.181872.6 sec100.0%meets the barwrong value
6DeepSeek V4.1 Flash
DeepSeek · open weights
97.4%77%$0.302274.2 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
97.1%75%$0.162532.5 sec100.0%meets the barwrong value
8Gemini 3.5 Flash Lite
Google
95.8%64%$0.513600.7 sec100.0%meets the barwrong value
9Mistral Medium 3.5
Mistral
94.5%55%$1.954470.8 sec100.0%below the barwrong value
10Mistral Small 4
Mistral · open weights
87.7%32%$0.126850.8 sec100.0%below the barwrong value

Email triage · German · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemini 3.8 Flash
Google
98.4%85%$2.741472.9 sec100.0%meets the barwrong value
2GPT-6 Sol
OpenAI
98.2%84%$2.731602.9 sec100.0%meets the barwrong value
3Claude Sonnet 5
Anthropic
97.9%81%$5.111874.3 sec100.0%meets the barwrong value
4Qwen3.8 Flash
Qwen
97.9%81%$0.3918712.1 sec100.0%meets the barwrong value
5GPT-6 Luna
OpenAI
97.0%73%$0.152732.8 sec100.0%meets the barwrong value
6DeepSeek V4.1 Flash
DeepSeek · open weights
96.3%68%$0.303244.2 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
95.1%62%$0.153803.6 sec100.0%meets the barwrong value
8Gemini 3.5 Flash Lite
Google
94.7%57%$0.484330.8 sec100.0%below the barwrong value
9Mistral Medium 3.5
Mistral
92.7%48%$1.775200.8 sec100.0%below the barwrong value
10Mistral Small 4
Mistral · open weights
87.1%29%$0.117060.8 sec100.0%below the barwrong value

Email triage · German · Email (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemini 3.8 Flash
Google
98.4%85%$2.741472.9 sec100.0%meets the barwrong value
2GPT-6 Sol
OpenAI
98.2%84%$2.731602.9 sec100.0%meets the barwrong value
3Claude Sonnet 5
Anthropic
97.9%81%$5.111874.3 sec100.0%meets the barwrong value
4Qwen3.8 Flash
Qwen
97.9%81%$0.3918712.1 sec100.0%meets the barwrong value
5GPT-6 Luna
OpenAI
97.0%73%$0.152732.8 sec100.0%meets the barwrong value
6DeepSeek V4.1 Flash
DeepSeek · open weights
96.3%68%$0.303244.2 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
95.1%62%$0.153803.6 sec100.0%meets the barwrong value
8Gemini 3.5 Flash Lite
Google
94.7%57%$0.484330.8 sec100.0%below the barwrong value
9Mistral Medium 3.5
Mistral
92.7%48%$1.775200.8 sec100.0%below the barwrong value
10Mistral Small 4
Mistral · open weights
87.1%29%$0.117060.8 sec100.0%below the barwrong value

Email triage · English (US) · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1GPT-6 Sol
OpenAI
99.0%92%$2.44802.8 sec100.0%meets the barwrong value
2Gemini 3.8 Flash
Google
98.7%89%$2.391132.2 sec100.0%meets the barwrong value
3Claude Sonnet 5
Anthropic
98.2%84%$4.011603.8 sec100.0%meets the barwrong value
4DeepSeek V4.1 Flash
DeepSeek · open weights
97.6%79%$0.232133.8 sec100.0%meets the barwrong value
5GPT-6 Luna
OpenAI
97.5%77%$0.132272.5 sec100.0%meets the barwrong value
6Qwen3.8 Flash
Qwen
97.5%79%$0.302138.9 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
96.5%72%$0.122802.2 sec100.0%meets the barwrong value
8Gemini 3.5 Flash Lite
Google
96.3%69%$0.423130.7 sec100.0%meets the barwrong value
9Mistral Medium 3.5
Mistral
94.9%57%$1.474270.8 sec100.0%below the barwrong value
10Mistral Small 4
Mistral · open weights
89.7%37%$0.096270.8 sec100.0%below the barwrong value

Email triage · English (US) · Email (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1GPT-6 Sol
OpenAI
99.0%92%$2.44802.8 sec100.0%meets the barwrong value
2Gemini 3.8 Flash
Google
98.7%89%$2.391132.2 sec100.0%meets the barwrong value
3Claude Sonnet 5
Anthropic
98.2%84%$4.011603.8 sec100.0%meets the barwrong value
4DeepSeek V4.1 Flash
DeepSeek · open weights
97.6%79%$0.232133.8 sec100.0%meets the barwrong value
5GPT-6 Luna
OpenAI
97.5%77%$0.132272.5 sec100.0%meets the barwrong value
6Qwen3.8 Flash
Qwen
97.5%79%$0.302138.9 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
96.5%72%$0.122802.2 sec100.0%meets the barwrong value
8Gemini 3.5 Flash Lite
Google
96.3%69%$0.423130.7 sec100.0%meets the barwrong value
9Mistral Medium 3.5
Mistral
94.9%57%$1.474270.8 sec100.0%below the barwrong value
10Mistral Small 4
Mistral · open weights
89.7%37%$0.096270.8 sec100.0%below the barwrong value

Contract clauses · Czech · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$6.0503.6 sec100.0%meets the bar—
2DeepSeek V4.1 Flash
DeepSeek · open weights
100.0%100%$0.4204.3 sec100.0%meets the bar—
3Gemini 3.8 Flash
Google
100.0%100%$3.0902.2 sec100.0%meets the bar—
4Mistral Medium 3.5
Mistral
100.0%100%$3.0601.2 sec100.0%meets the bar—
5GPT-6 Luna
OpenAI
100.0%100%$0.2702.6 sec100.0%meets the bar—
6GPT-6 Sol
OpenAI
100.0%100%$4.6202.0 sec100.0%meets the bar—
7Gemma 4 31B
Google · open weights
100.0%99%$0.2576.0 sec100.0%meets the barwrong value
8Qwen3.8 Flash
Qwen
99.9%99%$0.631313.6 sec100.0%meets the barwrong value
9Gemini 3.5 Flash Lite
Google
99.9%98%$0.85200.9 sec100.0%meets the barwrong value
10Mistral Small 4
Mistral · open weights
99.0%92%$0.23801.2 sec100.0%meets the barwrong value

Contract clauses · Czech · Contract text (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$6.0503.6 sec100.0%meets the bar—
2DeepSeek V4.1 Flash
DeepSeek · open weights
100.0%100%$0.4204.3 sec100.0%meets the bar—
3Gemini 3.8 Flash
Google
100.0%100%$3.0902.2 sec100.0%meets the bar—
4Mistral Medium 3.5
Mistral
100.0%100%$3.0601.2 sec100.0%meets the bar—
5GPT-6 Luna
OpenAI
100.0%100%$0.2702.6 sec100.0%meets the bar—
6GPT-6 Sol
OpenAI
100.0%100%$4.6202.0 sec100.0%meets the bar—
7Gemma 4 31B
Google · open weights
100.0%99%$0.2576.0 sec100.0%meets the barwrong value
8Qwen3.8 Flash
Qwen
99.9%99%$0.631313.6 sec100.0%meets the barwrong value
9Gemini 3.5 Flash Lite
Google
99.9%98%$0.85200.9 sec100.0%meets the barwrong value
10Mistral Small 4
Mistral · open weights
99.0%92%$0.23801.2 sec100.0%meets the barwrong value

Contract clauses · German · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1GPT-6 Luna
OpenAI
100.0%100%$0.2703.0 sec100.0%meets the bar—
2GPT-6 Sol
OpenAI
100.0%100%$4.3602.9 sec100.0%meets the bar—
3Gemini 3.8 Flash
Google
99.9%99%$3.4272.6 sec100.0%meets the barinvented value
4Claude Sonnet 5
Anthropic
98.2%87%$7.111334.3 sec100.0%meets the barinvented value
5Gemini 3.5 Flash Lite
Google
98.0%84%$0.791601.0 sec100.0%meets the barinvented value
6Qwen3.8 Flash
Qwen
98.0%87%$0.7812719.7 sec99.3%meets the barinvented value
7DeepSeek V4.1 Flash
DeepSeek · open weights
97.2%77%$0.512284.6 sec100.0%meets the barinvented value
8Mistral Medium 3.5
Mistral
97.1%75%$2.762471.2 sec100.0%meets the barinvented value
9Gemma 4 31B
Google · open weights
94.4%27%$0.227272.6 sec100.0%below the barnumber format
10Mistral Small 4
Mistral · open weights
94.0%45%$0.215531.2 sec100.0%below the barinvented value

Contract clauses · German · Contract text (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1GPT-6 Luna
OpenAI
100.0%100%$0.2703.0 sec100.0%meets the bar—
2GPT-6 Sol
OpenAI
100.0%100%$4.3602.9 sec100.0%meets the bar—
3Gemini 3.8 Flash
Google
99.9%99%$3.4272.6 sec100.0%meets the barinvented value
4Claude Sonnet 5
Anthropic
98.2%87%$7.111334.3 sec100.0%meets the barinvented value
5Gemini 3.5 Flash Lite
Google
98.0%84%$0.791601.0 sec100.0%meets the barinvented value
6Qwen3.8 Flash
Qwen
98.0%87%$0.7812719.7 sec99.3%meets the barinvented value
7DeepSeek V4.1 Flash
DeepSeek · open weights
97.2%77%$0.512284.6 sec100.0%meets the barinvented value
8Mistral Medium 3.5
Mistral
97.1%75%$2.762471.2 sec100.0%meets the barinvented value
9Gemma 4 31B
Google · open weights
94.4%27%$0.227272.6 sec100.0%below the barnumber format
10Mistral Small 4
Mistral · open weights
94.0%45%$0.215531.2 sec100.0%below the barinvented value

Contract clauses · English (US) · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$5.0403.4 sec100.0%meets the bar—
2DeepSeek V4.1 Flash
DeepSeek · open weights
100.0%100%$0.3203.9 sec100.0%meets the bar—
3GPT-6 Sol
OpenAI
100.0%100%$3.3302.5 sec100.0%meets the bar—
4Gemini 3.5 Flash Lite
Google
100.0%99%$0.6971.0 sec100.0%meets the barwrong value
5Gemini 3.8 Flash
Google
99.9%99%$2.4772.0 sec100.0%meets the barwrong value
6GPT-6 Luna
OpenAI
99.8%99%$0.20132.7 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
99.6%97%$0.18272.5 sec100.0%meets the barwrong value
8Mistral Medium 3.5
Mistral
98.1%87%$2.351271.1 sec100.0%meets the barwrong value
9Qwen3.8 Flash
Qwen
95.1%69%$0.6230714.1 sec100.0%meets the barwrong value
10Mistral Small 4
Mistral · open weights
91.5%43%$0.175671.1 sec100.0%below the barwrong value

Contract clauses · English (US) · Contract text (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Claude Sonnet 5
Anthropic
100.0%100%$5.0403.4 sec100.0%meets the bar—
2DeepSeek V4.1 Flash
DeepSeek · open weights
100.0%100%$0.3203.9 sec100.0%meets the bar—
3GPT-6 Sol
OpenAI
100.0%100%$3.3302.5 sec100.0%meets the bar—
4Gemini 3.5 Flash Lite
Google
100.0%99%$0.6971.0 sec100.0%meets the barwrong value
5Gemini 3.8 Flash
Google
99.9%99%$2.4772.0 sec100.0%meets the barwrong value
6GPT-6 Luna
OpenAI
99.8%99%$0.20132.7 sec100.0%meets the barwrong value
7Gemma 4 31B
Google · open weights
99.6%97%$0.18272.5 sec100.0%meets the barwrong value
8Mistral Medium 3.5
Mistral
98.1%87%$2.351271.1 sec100.0%meets the barwrong value
9Qwen3.8 Flash
Qwen
95.1%69%$0.6230714.1 sec100.0%meets the barwrong value
10Mistral Small 4
Mistral · open weights
91.5%43%$0.175671.1 sec100.0%below the barwrong value

Personal data detection · Czech · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemma 4 31B
Google · open weights
99.8%99%$0.12132.9 sec100.0%meets the barmissing field
2Mistral Medium 3.5
Mistral
99.8%99%$1.40130.8 sec100.0%meets the barinvented value
3DeepSeek V4.1 Flash
DeepSeek · open weights
99.7%98%$0.31204.2 sec100.0%meets the barmissing field
4GPT-6 Sol
OpenAI
99.7%98%$2.22202.5 sec100.0%meets the barmissing field
5GPT-6 Luna
OpenAI
99.6%97%$0.12271.8 sec100.0%meets the barmissing field
6Gemini 3.8 Flash
Google
99.5%97%$2.19332.0 sec100.0%meets the barmissing field
7Claude Sonnet 5
Anthropic
98.9%92%$4.06803.2 sec100.0%meets the barinvented value
8Mistral Small 4
Mistral · open weights
98.8%92%$0.09810.9 sec100.0%meets the barinvented value
9Gemini 3.5 Flash Lite
Google
98.1%88%$0.451200.8 sec100.0%meets the barinvented value
10Qwen3.8 Flash
Qwen
97.0%89%$0.411139.6 sec99.3%meets the barwrong value

Personal data detection · Czech · Text (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemma 4 31B
Google · open weights
99.8%99%$0.12132.9 sec100.0%meets the barmissing field
2Mistral Medium 3.5
Mistral
99.8%99%$1.40130.8 sec100.0%meets the barinvented value
3DeepSeek V4.1 Flash
DeepSeek · open weights
99.7%98%$0.31204.2 sec100.0%meets the barmissing field
4GPT-6 Sol
OpenAI
99.7%98%$2.22202.5 sec100.0%meets the barmissing field
5GPT-6 Luna
OpenAI
99.6%97%$0.12271.8 sec100.0%meets the barmissing field
6Gemini 3.8 Flash
Google
99.5%97%$2.19332.0 sec100.0%meets the barmissing field
7Claude Sonnet 5
Anthropic
98.9%92%$4.06803.2 sec100.0%meets the barinvented value
8Mistral Small 4
Mistral · open weights
98.8%92%$0.09810.9 sec100.0%meets the barinvented value
9Gemini 3.5 Flash Lite
Google
98.1%88%$0.451200.8 sec100.0%meets the barinvented value
10Qwen3.8 Flash
Qwen
97.0%89%$0.411139.6 sec99.3%meets the barwrong value

Personal data detection · German · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Mistral Medium 3.5
Mistral
100.0%100%$1.3800.9 sec100.0%meets the bar—
2Gemma 4 31B
Google · open weights
99.9%99%$0.1071.9 sec100.0%meets the barmissing field
3DeepSeek V4.1 Flash
DeepSeek · open weights
99.8%99%$0.31133.7 sec100.0%meets the barmissing field
4Claude Sonnet 5
Anthropic
99.7%98%$4.40203.7 sec100.0%meets the barinvented value
5Gemini 3.8 Flash
Google
99.7%98%$2.12202.3 sec100.0%meets the barmissing field
6GPT-6 Sol
OpenAI
99.7%98%$2.13202.6 sec100.0%meets the barmissing field
7Gemini 3.5 Flash Lite
Google
99.6%97%$0.45270.8 sec100.0%meets the barinvented value
8GPT-6 Luna
OpenAI
99.6%97%$0.12272.2 sec100.0%meets the barmissing field
9Qwen3.8 Flash
Qwen
99.6%99%$0.36712.2 sec100.0%meets the barmissing field
10Mistral Small 4
Mistral · open weights
99.1%94%$0.09650.9 sec100.0%meets the barinvented value

Personal data detection · German · Text (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Mistral Medium 3.5
Mistral
100.0%100%$1.3800.9 sec100.0%meets the bar—
2Gemma 4 31B
Google · open weights
99.9%99%$0.1071.9 sec100.0%meets the barmissing field
3DeepSeek V4.1 Flash
DeepSeek · open weights
99.8%99%$0.31133.7 sec100.0%meets the barmissing field
4Claude Sonnet 5
Anthropic
99.7%98%$4.40203.7 sec100.0%meets the barinvented value
5Gemini 3.8 Flash
Google
99.7%98%$2.12202.3 sec100.0%meets the barmissing field
6GPT-6 Sol
OpenAI
99.7%98%$2.13202.6 sec100.0%meets the barmissing field
7Gemini 3.5 Flash Lite
Google
99.6%97%$0.45270.8 sec100.0%meets the barinvented value
8GPT-6 Luna
OpenAI
99.6%97%$0.12272.2 sec100.0%meets the barmissing field
9Qwen3.8 Flash
Qwen
99.6%99%$0.36712.2 sec100.0%meets the barmissing field
10Mistral Small 4
Mistral · open weights
99.1%94%$0.09650.9 sec100.0%meets the barinvented value

Personal data detection · English (US) · All levels (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemma 4 31B
Google · open weights
99.9%99%$0.0872.4 sec100.0%meets the barmissing field
2Claude Sonnet 5
Anthropic
99.7%98%$3.50202.5 sec100.0%meets the barmissing field
3Gemini 3.8 Flash
Google
99.7%98%$1.86201.9 sec100.0%meets the barmissing field
4Mistral Medium 3.5
Mistral
99.7%98%$1.15200.8 sec100.0%meets the barmissing field
5GPT-6 Sol
OpenAI
99.7%98%$2.09202.6 sec100.0%meets the barmissing field
6DeepSeek V4.1 Flash
DeepSeek · open weights
99.6%97%$0.29273.8 sec100.0%meets the barmissing field
7GPT-6 Luna
OpenAI
99.6%97%$0.11271.8 sec100.0%meets the barinvented value
8Gemini 3.5 Flash Lite
Google
99.5%97%$0.41270.8 sec100.0%meets the barmissing field
9Mistral Small 4
Mistral · open weights
99.4%96%$0.08400.8 sec100.0%meets the barinvented value
10Qwen3.8 Flash
Qwen
99.2%99%$0.28137.5 sec100.0%meets the barinvented value

Personal data detection · English (US) · Text (150 documents)

#ModelCorrect fieldsError-free documentsPer 1,000 documentsTo fix / 1,000Speed (median)Valid answersBarMost common error
1Gemma 4 31B
Google · open weights
99.9%99%$0.0872.4 sec100.0%meets the barmissing field
2Claude Sonnet 5
Anthropic
99.7%98%$3.50202.5 sec100.0%meets the barmissing field
3Gemini 3.8 Flash
Google
99.7%98%$1.86201.9 sec100.0%meets the barmissing field
4Mistral Medium 3.5
Mistral
99.7%98%$1.15200.8 sec100.0%meets the barmissing field
5GPT-6 Sol
OpenAI
99.7%98%$2.09202.6 sec100.0%meets the barmissing field
6DeepSeek V4.1 Flash
DeepSeek · open weights
99.6%97%$0.29273.8 sec100.0%meets the barmissing field
7GPT-6 Luna
OpenAI
99.6%97%$0.11271.8 sec100.0%meets the barinvented value
8Gemini 3.5 Flash Lite
Google
99.5%97%$0.41270.8 sec100.0%meets the barmissing field
9Mistral Small 4
Mistral · open weights
99.4%96%$0.08400.8 sec100.0%meets the barinvented value
10Qwen3.8 Flash
Qwen
99.2%99%$0.28137.5 sec100.0%meets the barinvented value