Leaderboard
Who does your task best
Sorted by the share of correct fields. Cost is the real token spend scaled to 1,000 documents. Bar: 95% correct fields; below it, manual fixes cost more than a cheaper model saves.
Compare two models side by side →
Invoice extraction · Czech · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $9.41 | 0 | 3.6 sec | 100.0% | meets the bar | — |
| 2 | Gemini 3.8 Flash | 100.0% | 99% | $4.26 | 7 | 3.5 sec | 100.0% | meets the bar | wrong value |
| 3 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.9% | 99% | $0.50 | 13 | 5.0 sec | 100.0% | meets the bar | wrong value |
| 4 | GPT-6 Sol OpenAI | 99.9% | 99% | $10.12 | 13 | 3.9 sec | 100.0% | meets the bar | wrong value |
| 5 | Gemini 3.5 Flash Lite | 99.5% | 93% | $1.03 | 67 | 1.8 sec | 100.0% | meets the bar | invented value |
| 6 | GPT-6 Luna OpenAI | 99.3% | 93% | $0.51 | 67 | 3.2 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 96.1% | 67% | $0.21 | 333 | 10.4 sec | 100.0% | meets the bar | wrong value |
| 8 | Mistral Small 4 Mistral · open weights | 93.6% | 56% | $0.44 | 440 | 2.0 sec | 100.0% | below the bar | wrong value |
| 9 | Mistral Medium 3.5 Mistral | 93.5% | 63% | $4.69 | 367 | 2.6 sec | 100.0% | below the bar | wrong value |
| 10 | Qwen3.8 Flash Qwen | 91.8% | 77% | $1.02 | 233 | 17.0 sec | 96.7% | below the bar | invalid JSON |
Invoice extraction · Czech · PDF text (45 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $4.87 | 0 | 2.7 sec | 100.0% | meets the bar | — |
| 2 | DeepSeek V4.1 Flash DeepSeek · open weights | 100.0% | 100% | $0.50 | 0 | 6.0 sec | 100.0% | meets the bar | — |
| 3 | Gemini 3.5 Flash Lite | 100.0% | 100% | $0.91 | 0 | 1.1 sec | 100.0% | meets the bar | — |
| 4 | Gemini 3.8 Flash | 100.0% | 100% | $3.46 | 0 | 2.8 sec | 100.0% | meets the bar | — |
| 5 | GPT-6 Luna OpenAI | 100.0% | 100% | $0.21 | 0 | 2.6 sec | 100.0% | meets the bar | — |
| 6 | GPT-6 Sol OpenAI | 100.0% | 100% | $4.02 | 0 | 2.8 sec | 100.0% | meets the bar | — |
| 7 | Mistral Medium 3.5 Mistral | 99.8% | 98% | $2.97 | 22 | 1.5 sec | 100.0% | meets the bar | missing field |
| 8 | Gemma 4 31B Google · open weights | 99.5% | 93% | $0.24 | 67 | 9.0 sec | 100.0% | meets the bar | missing field |
| 9 | Mistral Small 4 Mistral · open weights | 96.8% | 76% | $0.26 | 244 | 1.4 sec | 100.0% | meets the bar | invented value |
| 10 | Qwen3.8 Flash Qwen | 92.0% | 53% | $0.74 | 467 | 19.9 sec | 100.0% | below the bar | wrong value |
Invoice extraction · Czech · Clean image (45 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $13.20 | 0 | 3.8 sec | 100.0% | meets the bar | — |
| 2 | DeepSeek V4.1 Flash DeepSeek · open weights | 100.0% | 100% | $0.49 | 0 | 4.4 sec | 100.0% | meets the bar | — |
| 3 | Gemini 3.8 Flash | 100.0% | 100% | $4.61 | 0 | 4.0 sec | 100.0% | meets the bar | — |
| 4 | GPT-6 Sol OpenAI | 100.0% | 100% | $14.57 | 0 | 3.9 sec | 100.0% | meets the bar | — |
| 5 | GPT-6 Luna OpenAI | 99.7% | 96% | $0.73 | 44 | 3.6 sec | 100.0% | meets the bar | wrong value |
| 6 | Gemini 3.5 Flash Lite | 99.5% | 93% | $1.07 | 67 | 2.0 sec | 100.0% | meets the bar | invented value |
| 7 | Gemma 4 31B Google · open weights | 95.4% | 62% | $0.20 | 378 | 11.0 sec | 100.0% | meets the bar | wrong value |
| 8 | Qwen3.8 Flash Qwen | 92.7% | 84% | $1.11 | 156 | 15.9 sec | 95.6% | below the bar | invalid JSON |
| 9 | Mistral Small 4 Mistral · open weights | 92.3% | 44% | $0.51 | 556 | 1.9 sec | 100.0% | below the bar | invented value |
| 10 | Mistral Medium 3.5 Mistral | 92.1% | 60% | $5.37 | 400 | 2.5 sec | 100.0% | below the bar | wrong value |
Invoice extraction · Czech · Scan (45 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $10.49 | 0 | 3.7 sec | 100.0% | meets the bar | — |
| 2 | Gemini 3.8 Flash | 99.8% | 98% | $4.61 | 22 | 3.5 sec | 100.0% | meets the bar | wrong value |
| 3 | GPT-6 Sol OpenAI | 99.8% | 98% | $11.89 | 22 | 4.2 sec | 100.0% | meets the bar | wrong value |
| 4 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.7% | 96% | $0.49 | 44 | 4.8 sec | 100.0% | meets the bar | wrong value |
| 5 | Gemini 3.5 Flash Lite | 99.0% | 87% | $1.08 | 133 | 1.8 sec | 100.0% | meets the bar | wrong value |
| 6 | GPT-6 Luna OpenAI | 98.3% | 87% | $0.59 | 133 | 3.4 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 94.9% | 56% | $0.20 | 444 | 11.9 sec | 100.0% | below the bar | wrong value |
| 8 | Mistral Small 4 Mistral · open weights | 92.1% | 47% | $0.51 | 533 | 2.6 sec | 100.0% | below the bar | wrong value |
| 9 | Qwen3.8 Flash Qwen | 90.9% | 89% | $1.17 | 111 | 15.4 sec | 93.3% | below the bar | invalid JSON |
| 10 | Mistral Medium 3.5 Mistral | 89.1% | 36% | $5.41 | 644 | 2.9 sec | 100.0% | below the bar | wrong value |
Invoice extraction · Czech · Photo (15 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $8.45 | 0 | 3.9 sec | 100.0% | meets the bar | — |
| 2 | DeepSeek V4.1 Flash DeepSeek · open weights | 100.0% | 100% | $0.54 | 0 | 5.4 sec | 100.0% | meets the bar | — |
| 3 | Gemini 3.8 Flash | 100.0% | 100% | $4.53 | 0 | 3.4 sec | 100.0% | meets the bar | — |
| 4 | Gemini 3.5 Flash Lite | 99.5% | 93% | $1.11 | 67 | 1.8 sec | 100.0% | meets the bar | invented value |
| 5 | GPT-6 Sol OpenAI | 99.5% | 93% | $9.76 | 67 | 4.9 sec | 100.0% | meets the bar | wrong value |
| 6 | GPT-6 Luna OpenAI | 99.0% | 87% | $0.47 | 133 | 3.3 sec | 100.0% | meets the bar | wrong value |
| 7 | Mistral Small 4 Mistral · open weights | 92.8% | 60% | $0.54 | 400 | 2.9 sec | 100.0% | below the bar | wrong value |
| 8 | Gemma 4 31B Google · open weights | 91.8% | 33% | $0.20 | 667 | 11.0 sec | 100.0% | below the bar | wrong value |
| 9 | Mistral Medium 3.5 Mistral | 91.8% | 53% | $5.65 | 467 | 3.0 sec | 100.0% | below the bar | wrong value |
| 10 | Qwen3.8 Flash Qwen | 91.3% | 87% | $1.11 | 133 | 15.4 sec | 100.0% | below the bar | missing field |
Invoice extraction · German · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.5 Flash Lite | 100.0% | 100% | $1.03 | 0 | 1.5 sec | 100.0% | meets the bar | — |
| 2 | GPT-6 Sol OpenAI | 100.0% | 99% | $10.65 | 7 | 4.9 sec | 100.0% | meets the bar | wrong value |
| 3 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.5% | 94% | $0.57 | 60 | 6.4 sec | 100.0% | meets the bar | wrong value |
| 4 | GPT-6 Luna OpenAI | 99.2% | 91% | $0.54 | 87 | 3.7 sec | 100.0% | meets the bar | wrong value |
| 5 | Claude Sonnet 5 Anthropic | 99.0% | 87% | $10.42 | 133 | 3.8 sec | 100.0% | meets the bar | missing field |
| 6 | Mistral Medium 3.5 Mistral | 97.4% | 75% | $4.72 | 253 | 2.5 sec | 100.0% | meets the bar | invented value |
| 7 | Gemini 3.8 Flash | 97.2% | 96% | $7.15 | 40 | 4.7 sec | 97.3% | meets the bar | invalid JSON |
| 8 | Gemma 4 31B Google · open weights | 94.9% | 64% | $0.20 | 360 | 7.7 sec | 100.0% | below the bar | wrong value |
| 9 | Mistral Small 4 Mistral · open weights | 94.5% | 49% | $0.44 | 508 | 2.0 sec | 100.0% | below the bar | invented value |
| 10 | Qwen3.8 Flash Qwen | 93.5% | 87% | $1.12 | 127 | 19.0 sec | 94.0% | below the bar | invalid JSON |
Invoice extraction · German · PDF text (45 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash DeepSeek · open weights | 100.0% | 100% | $0.49 | 0 | 6.7 sec | 100.0% | meets the bar | — |
| 2 | Gemini 3.5 Flash Lite | 100.0% | 100% | $0.91 | 0 | 1.1 sec | 100.0% | meets the bar | — |
| 3 | GPT-6 Luna OpenAI | 100.0% | 100% | $0.22 | 0 | 3.0 sec | 100.0% | meets the bar | — |
| 4 | GPT-6 Sol OpenAI | 100.0% | 100% | $3.85 | 0 | 3.1 sec | 100.0% | meets the bar | — |
| 5 | Qwen3.8 Flash Qwen | 100.0% | 100% | $0.99 | 0 | 27.8 sec | 100.0% | meets the bar | — |
| 6 | Gemma 4 31B Google · open weights | 99.5% | 93% | $0.23 | 67 | 7.0 sec | 100.0% | meets the bar | missing field |
| 7 | Mistral Medium 3.5 Mistral | 99.5% | 93% | $2.96 | 67 | 1.4 sec | 100.0% | meets the bar | missing field |
| 8 | Claude Sonnet 5 Anthropic | 99.3% | 91% | $5.89 | 89 | 3.3 sec | 100.0% | meets the bar | missing field |
| 9 | Mistral Small 4 Mistral · open weights | 97.9% | 80% | $0.24 | 200 | 1.3 sec | 100.0% | meets the bar | invented value |
| 10 | Gemini 3.8 Flash | 97.8% | 98% | $6.24 | 22 | 3.9 sec | 97.8% | meets the bar | invalid JSON |
Invoice extraction · German · Clean image (45 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.5 Flash Lite | 100.0% | 100% | $1.07 | 0 | 1.7 sec | 100.0% | meets the bar | — |
| 2 | GPT-6 Sol OpenAI | 100.0% | 100% | $15.26 | 0 | 5.6 sec | 100.0% | meets the bar | — |
| 3 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.7% | 96% | $0.55 | 44 | 5.7 sec | 100.0% | meets the bar | wrong value |
| 4 | GPT-6 Luna OpenAI | 99.7% | 96% | $0.77 | 44 | 4.1 sec | 100.0% | meets the bar | wrong value |
| 5 | Claude Sonnet 5 Anthropic | 99.0% | 87% | $13.27 | 133 | 3.8 sec | 100.0% | meets the bar | missing field |
| 6 | Mistral Medium 3.5 Mistral | 96.6% | 69% | $5.43 | 311 | 2.5 sec | 100.0% | meets the bar | invented value |
| 7 | Gemini 3.8 Flash | 95.2% | 91% | $8.81 | 89 | 5.1 sec | 95.6% | meets the bar | invalid JSON |
| 8 | Gemma 4 31B Google · open weights | 93.9% | 51% | $0.19 | 489 | 7.4 sec | 100.0% | below the bar | wrong value |
| 9 | Mistral Small 4 Mistral · open weights | 93.4% | 36% | $0.51 | 643 | 1.8 sec | 100.0% | below the bar | invented value |
| 10 | Qwen3.8 Flash Qwen | 88.4% | 82% | $1.26 | 178 | 16.7 sec | 88.9% | below the bar | invalid JSON |
Invoice extraction · German · Scan (45 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.5 Flash Lite | 100.0% | 100% | $1.08 | 0 | 1.6 sec | 100.0% | meets the bar | — |
| 2 | GPT-6 Sol OpenAI | 99.8% | 98% | $13.15 | 22 | 5.1 sec | 100.0% | meets the bar | wrong value |
| 3 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.0% | 87% | $0.61 | 133 | 7.2 sec | 100.0% | meets the bar | wrong value |
| 4 | Claude Sonnet 5 Anthropic | 98.6% | 82% | $11.88 | 178 | 3.8 sec | 100.0% | meets the bar | missing field |
| 5 | GPT-6 Luna OpenAI | 98.3% | 82% | $0.66 | 178 | 4.1 sec | 100.0% | meets the bar | wrong value |
| 6 | Gemini 3.8 Flash | 97.8% | 98% | $6.65 | 22 | 5.0 sec | 97.8% | meets the bar | invalid JSON |
| 7 | Mistral Medium 3.5 Mistral | 96.9% | 67% | $5.45 | 333 | 2.8 sec | 100.0% | meets the bar | invented value |
| 8 | Qwen3.8 Flash Qwen | 96.9% | 87% | $1.15 | 133 | 16.6 sec | 97.8% | meets the bar | invalid JSON |
| 9 | Mistral Small 4 Mistral · open weights | 92.5% | 33% | $0.52 | 667 | 2.4 sec | 100.0% | below the bar | invented value |
| 10 | Gemma 4 31B Google · open weights | 91.8% | 51% | $0.19 | 489 | 8.7 sec | 100.0% | below the bar | wrong value |
Invoice extraction · German · Photo (15 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.5 Flash Lite | 100.0% | 100% | $1.11 | 0 | 1.6 sec | 100.0% | meets the bar | — |
| 2 | Gemini 3.8 Flash | 100.0% | 100% | $6.45 | 0 | 4.8 sec | 100.0% | meets the bar | — |
| 3 | GPT-6 Sol OpenAI | 100.0% | 100% | $9.73 | 0 | 5.4 sec | 100.0% | meets the bar | — |
| 4 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.5% | 93% | $0.70 | 67 | 6.4 sec | 100.0% | meets the bar | swapped dates |
| 5 | Claude Sonnet 5 Anthropic | 99.0% | 87% | $11.05 | 133 | 3.8 sec | 100.0% | meets the bar | wrong value |
| 6 | GPT-6 Luna OpenAI | 98.5% | 80% | $0.50 | 200 | 3.5 sec | 100.0% | meets the bar | wrong value |
| 7 | Mistral Medium 3.5 Mistral | 94.9% | 60% | $5.68 | 400 | 2.7 sec | 100.0% | below the bar | wrong value |
| 8 | Mistral Small 4 Mistral · open weights | 93.5% | 46% | $0.53 | 539 | 2.5 sec | 100.0% | below the bar | invented value |
| 9 | Gemma 4 31B Google · open weights | 93.3% | 53% | $0.19 | 467 | 9.1 sec | 100.0% | below the bar | wrong value |
| 10 | Qwen3.8 Flash Qwen | 79.0% | 67% | $0.99 | 333 | 15.6 sec | 80.0% | below the bar | invalid JSON |
Invoice extraction · English (US) · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Sol OpenAI | 99.9% | 99% | $9.65 | 7 | 4.1 sec | 100.0% | meets the bar | wrong value |
| 2 | Claude Sonnet 5 Anthropic | 99.8% | 99% | $8.84 | 13 | 3.2 sec | 100.0% | meets the bar | wrong value |
| 3 | Gemini 3.8 Flash | 99.8% | 98% | $3.84 | 20 | 3.2 sec | 100.0% | meets the bar | wrong value |
| 4 | GPT-6 Luna OpenAI | 99.7% | 97% | $0.51 | 33 | 3.4 sec | 100.0% | meets the bar | invented value |
| 5 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.4% | 93% | $0.52 | 67 | 5.7 sec | 100.0% | meets the bar | wrong value |
| 6 | Gemini 3.5 Flash Lite | 99.2% | 92% | $0.87 | 80 | 1.7 sec | 100.0% | meets the bar | wrong value |
| 7 | Mistral Medium 3.5 Mistral | 97.8% | 79% | $4.37 | 207 | 2.3 sec | 100.0% | meets the bar | wrong value |
| 8 | Gemma 4 31B Google · open weights | 97.6% | 81% | $0.17 | 187 | 8.9 sec | 100.0% | meets the bar | wrong value |
| 9 | Mistral Small 4 Mistral · open weights | 95.6% | 58% | $0.41 | 420 | 1.8 sec | 100.0% | meets the bar | wrong rate |
| 10 | Qwen3.8 Flash Qwen | 82.4% | 57% | $1.15 | 427 | 17.4 sec | 90.7% | below the bar | invalid JSON |
Invoice extraction · English (US) · PDF text (45 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash DeepSeek · open weights | 100.0% | 100% | $0.52 | 0 | 7.4 sec | 100.0% | meets the bar | — |
| 2 | Gemini 3.8 Flash | 100.0% | 100% | $2.92 | 0 | 2.3 sec | 100.0% | meets the bar | — |
| 3 | Gemma 4 31B Google · open weights | 100.0% | 100% | $0.18 | 0 | 7.4 sec | 100.0% | meets the bar | — |
| 4 | GPT-6 Luna OpenAI | 100.0% | 100% | $0.18 | 0 | 2.7 sec | 100.0% | meets the bar | — |
| 5 | GPT-6 Sol OpenAI | 100.0% | 100% | $3.28 | 0 | 3.2 sec | 100.0% | meets the bar | — |
| 6 | Claude Sonnet 5 Anthropic | 99.8% | 98% | $3.91 | 22 | 2.4 sec | 100.0% | meets the bar | missing field |
| 7 | Mistral Medium 3.5 Mistral | 99.8% | 98% | $2.24 | 22 | 1.2 sec | 100.0% | meets the bar | wrong value |
| 8 | Gemini 3.5 Flash Lite | 99.4% | 96% | $0.72 | 44 | 1.0 sec | 100.0% | meets the bar | wrong value |
| 9 | Mistral Small 4 Mistral · open weights | 98.4% | 84% | $0.19 | 156 | 1.2 sec | 100.0% | meets the bar | wrong value |
| 10 | Qwen3.8 Flash Qwen | 91.7% | 49% | $0.97 | 511 | 18.7 sec | 100.0% | below the bar | wrong rate |
Invoice extraction · English (US) · Clean image (45 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $12.19 | 0 | 3.3 sec | 100.0% | meets the bar | — |
| 2 | GPT-6 Sol OpenAI | 100.0% | 100% | $13.81 | 0 | 4.2 sec | 100.0% | meets the bar | — |
| 3 | Gemini 3.5 Flash Lite | 99.6% | 96% | $0.93 | 44 | 1.8 sec | 100.0% | meets the bar | invented value |
| 4 | Gemini 3.8 Flash | 99.6% | 96% | $4.13 | 44 | 3.5 sec | 100.0% | meets the bar | wrong value |
| 5 | GPT-6 Luna OpenAI | 99.6% | 96% | $0.73 | 44 | 3.8 sec | 100.0% | meets the bar | invented value |
| 6 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.2% | 91% | $0.50 | 89 | 4.6 sec | 100.0% | meets the bar | invented value |
| 7 | Gemma 4 31B Google · open weights | 98.2% | 84% | $0.16 | 156 | 9.2 sec | 100.0% | meets the bar | wrong value |
| 8 | Mistral Medium 3.5 Mistral | 98.0% | 82% | $5.25 | 178 | 2.3 sec | 100.0% | meets the bar | wrong value |
| 9 | Mistral Small 4 Mistral · open weights | 95.6% | 56% | $0.50 | 444 | 1.7 sec | 100.0% | meets the bar | wrong rate |
| 10 | Qwen3.8 Flash Qwen | 77.2% | 60% | $1.13 | 400 | 15.2 sec | 86.7% | below the bar | invalid JSON |
Invoice extraction · English (US) · Scan (45 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $10.66 | 0 | 3.3 sec | 100.0% | meets the bar | — |
| 2 | GPT-6 Sol OpenAI | 100.0% | 100% | $12.04 | 0 | 4.3 sec | 100.0% | meets the bar | — |
| 3 | Gemini 3.8 Flash | 99.8% | 98% | $4.36 | 22 | 3.2 sec | 100.0% | meets the bar | wrong value |
| 4 | GPT-6 Luna OpenAI | 99.6% | 96% | $0.63 | 44 | 3.7 sec | 100.0% | meets the bar | wrong value |
| 5 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.0% | 89% | $0.55 | 111 | 5.4 sec | 100.0% | meets the bar | wrong value |
| 6 | Gemini 3.5 Flash Lite | 98.8% | 87% | $0.93 | 133 | 1.8 sec | 100.0% | meets the bar | invented value |
| 7 | Gemma 4 31B Google · open weights | 96.6% | 73% | $0.16 | 267 | 9.8 sec | 100.0% | meets the bar | wrong value |
| 8 | Mistral Medium 3.5 Mistral | 96.0% | 62% | $5.27 | 378 | 2.8 sec | 100.0% | meets the bar | wrong value |
| 9 | Mistral Small 4 Mistral · open weights | 93.5% | 38% | $0.50 | 622 | 2.3 sec | 100.0% | below the bar | wrong rate |
| 10 | Qwen3.8 Flash Qwen | 81.4% | 60% | $1.35 | 400 | 17.4 sec | 91.1% | below the bar | invalid JSON |
Invoice extraction · English (US) · Photo (15 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash | 100.0% | 100% | $4.13 | 0 | 3.4 sec | 100.0% | meets the bar | — |
| 2 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.4% | 93% | $0.52 | 67 | 5.3 sec | 100.0% | meets the bar | wrong value |
| 3 | GPT-6 Luna OpenAI | 99.4% | 93% | $0.48 | 67 | 3.7 sec | 100.0% | meets the bar | wrong value |
| 4 | GPT-6 Sol OpenAI | 99.4% | 93% | $9.08 | 67 | 4.9 sec | 100.0% | meets the bar | wrong value |
| 5 | Claude Sonnet 5 Anthropic | 98.8% | 93% | $8.09 | 67 | 3.2 sec | 100.0% | meets the bar | wrong value |
| 6 | Gemini 3.5 Flash Lite | 98.8% | 87% | $0.94 | 133 | 1.7 sec | 100.0% | meets the bar | wrong value |
| 7 | Mistral Medium 3.5 Mistral | 96.4% | 67% | $5.44 | 333 | 2.8 sec | 100.0% | meets the bar | wrong value |
| 8 | Mistral Small 4 Mistral · open weights | 93.3% | 47% | $0.51 | 533 | 2.6 sec | 100.0% | below the bar | wrong value |
| 9 | Gemma 4 31B Google · open weights | 92.1% | 40% | $0.16 | 600 | 10.4 sec | 100.0% | below the bar | wrong value |
| 10 | Qwen3.8 Flash Qwen | 72.7% | 67% | $1.20 | 333 | 19.6 sec | 73.3% | below the bar | invalid JSON |
Email triage · Czech · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash | 99.2% | 93% | $3.06 | 73 | 2.6 sec | 100.0% | meets the bar | wrong value |
| 2 | GPT-6 Sol OpenAI | 99.0% | 91% | $3.19 | 93 | 2.8 sec | 100.0% | meets the bar | wrong value |
| 3 | Qwen3.8 Flash Qwen | 98.5% | 87% | $0.45 | 133 | 13.0 sec | 100.0% | meets the bar | wrong value |
| 4 | Claude Sonnet 5 Anthropic | 98.2% | 83% | $4.91 | 167 | 4.7 sec | 100.0% | meets the bar | wrong value |
| 5 | GPT-6 Luna OpenAI | 97.9% | 81% | $0.18 | 187 | 2.6 sec | 100.0% | meets the bar | wrong value |
| 6 | DeepSeek V4.1 Flash DeepSeek · open weights | 97.4% | 77% | $0.30 | 227 | 4.2 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 97.1% | 75% | $0.16 | 253 | 2.5 sec | 100.0% | meets the bar | wrong value |
| 8 | Gemini 3.5 Flash Lite | 95.8% | 64% | $0.51 | 360 | 0.7 sec | 100.0% | meets the bar | wrong value |
| 9 | Mistral Medium 3.5 Mistral | 94.5% | 55% | $1.95 | 447 | 0.8 sec | 100.0% | below the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 87.7% | 32% | $0.12 | 685 | 0.8 sec | 100.0% | below the bar | wrong value |
Email triage · Czech · Email (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash | 99.2% | 93% | $3.06 | 73 | 2.6 sec | 100.0% | meets the bar | wrong value |
| 2 | GPT-6 Sol OpenAI | 99.0% | 91% | $3.19 | 93 | 2.8 sec | 100.0% | meets the bar | wrong value |
| 3 | Qwen3.8 Flash Qwen | 98.5% | 87% | $0.45 | 133 | 13.0 sec | 100.0% | meets the bar | wrong value |
| 4 | Claude Sonnet 5 Anthropic | 98.2% | 83% | $4.91 | 167 | 4.7 sec | 100.0% | meets the bar | wrong value |
| 5 | GPT-6 Luna OpenAI | 97.9% | 81% | $0.18 | 187 | 2.6 sec | 100.0% | meets the bar | wrong value |
| 6 | DeepSeek V4.1 Flash DeepSeek · open weights | 97.4% | 77% | $0.30 | 227 | 4.2 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 97.1% | 75% | $0.16 | 253 | 2.5 sec | 100.0% | meets the bar | wrong value |
| 8 | Gemini 3.5 Flash Lite | 95.8% | 64% | $0.51 | 360 | 0.7 sec | 100.0% | meets the bar | wrong value |
| 9 | Mistral Medium 3.5 Mistral | 94.5% | 55% | $1.95 | 447 | 0.8 sec | 100.0% | below the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 87.7% | 32% | $0.12 | 685 | 0.8 sec | 100.0% | below the bar | wrong value |
Email triage · German · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash | 98.4% | 85% | $2.74 | 147 | 2.9 sec | 100.0% | meets the bar | wrong value |
| 2 | GPT-6 Sol OpenAI | 98.2% | 84% | $2.73 | 160 | 2.9 sec | 100.0% | meets the bar | wrong value |
| 3 | Claude Sonnet 5 Anthropic | 97.9% | 81% | $5.11 | 187 | 4.3 sec | 100.0% | meets the bar | wrong value |
| 4 | Qwen3.8 Flash Qwen | 97.9% | 81% | $0.39 | 187 | 12.1 sec | 100.0% | meets the bar | wrong value |
| 5 | GPT-6 Luna OpenAI | 97.0% | 73% | $0.15 | 273 | 2.8 sec | 100.0% | meets the bar | wrong value |
| 6 | DeepSeek V4.1 Flash DeepSeek · open weights | 96.3% | 68% | $0.30 | 324 | 4.2 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 95.1% | 62% | $0.15 | 380 | 3.6 sec | 100.0% | meets the bar | wrong value |
| 8 | Gemini 3.5 Flash Lite | 94.7% | 57% | $0.48 | 433 | 0.8 sec | 100.0% | below the bar | wrong value |
| 9 | Mistral Medium 3.5 Mistral | 92.7% | 48% | $1.77 | 520 | 0.8 sec | 100.0% | below the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 87.1% | 29% | $0.11 | 706 | 0.8 sec | 100.0% | below the bar | wrong value |
Email triage · German · Email (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash | 98.4% | 85% | $2.74 | 147 | 2.9 sec | 100.0% | meets the bar | wrong value |
| 2 | GPT-6 Sol OpenAI | 98.2% | 84% | $2.73 | 160 | 2.9 sec | 100.0% | meets the bar | wrong value |
| 3 | Claude Sonnet 5 Anthropic | 97.9% | 81% | $5.11 | 187 | 4.3 sec | 100.0% | meets the bar | wrong value |
| 4 | Qwen3.8 Flash Qwen | 97.9% | 81% | $0.39 | 187 | 12.1 sec | 100.0% | meets the bar | wrong value |
| 5 | GPT-6 Luna OpenAI | 97.0% | 73% | $0.15 | 273 | 2.8 sec | 100.0% | meets the bar | wrong value |
| 6 | DeepSeek V4.1 Flash DeepSeek · open weights | 96.3% | 68% | $0.30 | 324 | 4.2 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 95.1% | 62% | $0.15 | 380 | 3.6 sec | 100.0% | meets the bar | wrong value |
| 8 | Gemini 3.5 Flash Lite | 94.7% | 57% | $0.48 | 433 | 0.8 sec | 100.0% | below the bar | wrong value |
| 9 | Mistral Medium 3.5 Mistral | 92.7% | 48% | $1.77 | 520 | 0.8 sec | 100.0% | below the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 87.1% | 29% | $0.11 | 706 | 0.8 sec | 100.0% | below the bar | wrong value |
Email triage · English (US) · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Sol OpenAI | 99.0% | 92% | $2.44 | 80 | 2.8 sec | 100.0% | meets the bar | wrong value |
| 2 | Gemini 3.8 Flash | 98.7% | 89% | $2.39 | 113 | 2.2 sec | 100.0% | meets the bar | wrong value |
| 3 | Claude Sonnet 5 Anthropic | 98.2% | 84% | $4.01 | 160 | 3.8 sec | 100.0% | meets the bar | wrong value |
| 4 | DeepSeek V4.1 Flash DeepSeek · open weights | 97.6% | 79% | $0.23 | 213 | 3.8 sec | 100.0% | meets the bar | wrong value |
| 5 | GPT-6 Luna OpenAI | 97.5% | 77% | $0.13 | 227 | 2.5 sec | 100.0% | meets the bar | wrong value |
| 6 | Qwen3.8 Flash Qwen | 97.5% | 79% | $0.30 | 213 | 8.9 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 96.5% | 72% | $0.12 | 280 | 2.2 sec | 100.0% | meets the bar | wrong value |
| 8 | Gemini 3.5 Flash Lite | 96.3% | 69% | $0.42 | 313 | 0.7 sec | 100.0% | meets the bar | wrong value |
| 9 | Mistral Medium 3.5 Mistral | 94.9% | 57% | $1.47 | 427 | 0.8 sec | 100.0% | below the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 89.7% | 37% | $0.09 | 627 | 0.8 sec | 100.0% | below the bar | wrong value |
Email triage · English (US) · Email (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Sol OpenAI | 99.0% | 92% | $2.44 | 80 | 2.8 sec | 100.0% | meets the bar | wrong value |
| 2 | Gemini 3.8 Flash | 98.7% | 89% | $2.39 | 113 | 2.2 sec | 100.0% | meets the bar | wrong value |
| 3 | Claude Sonnet 5 Anthropic | 98.2% | 84% | $4.01 | 160 | 3.8 sec | 100.0% | meets the bar | wrong value |
| 4 | DeepSeek V4.1 Flash DeepSeek · open weights | 97.6% | 79% | $0.23 | 213 | 3.8 sec | 100.0% | meets the bar | wrong value |
| 5 | GPT-6 Luna OpenAI | 97.5% | 77% | $0.13 | 227 | 2.5 sec | 100.0% | meets the bar | wrong value |
| 6 | Qwen3.8 Flash Qwen | 97.5% | 79% | $0.30 | 213 | 8.9 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 96.5% | 72% | $0.12 | 280 | 2.2 sec | 100.0% | meets the bar | wrong value |
| 8 | Gemini 3.5 Flash Lite | 96.3% | 69% | $0.42 | 313 | 0.7 sec | 100.0% | meets the bar | wrong value |
| 9 | Mistral Medium 3.5 Mistral | 94.9% | 57% | $1.47 | 427 | 0.8 sec | 100.0% | below the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 89.7% | 37% | $0.09 | 627 | 0.8 sec | 100.0% | below the bar | wrong value |
Contract clauses · Czech · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $6.05 | 0 | 3.6 sec | 100.0% | meets the bar | — |
| 2 | DeepSeek V4.1 Flash DeepSeek · open weights | 100.0% | 100% | $0.42 | 0 | 4.3 sec | 100.0% | meets the bar | — |
| 3 | Gemini 3.8 Flash | 100.0% | 100% | $3.09 | 0 | 2.2 sec | 100.0% | meets the bar | — |
| 4 | Mistral Medium 3.5 Mistral | 100.0% | 100% | $3.06 | 0 | 1.2 sec | 100.0% | meets the bar | — |
| 5 | GPT-6 Luna OpenAI | 100.0% | 100% | $0.27 | 0 | 2.6 sec | 100.0% | meets the bar | — |
| 6 | GPT-6 Sol OpenAI | 100.0% | 100% | $4.62 | 0 | 2.0 sec | 100.0% | meets the bar | — |
| 7 | Gemma 4 31B Google · open weights | 100.0% | 99% | $0.25 | 7 | 6.0 sec | 100.0% | meets the bar | wrong value |
| 8 | Qwen3.8 Flash Qwen | 99.9% | 99% | $0.63 | 13 | 13.6 sec | 100.0% | meets the bar | wrong value |
| 9 | Gemini 3.5 Flash Lite | 99.9% | 98% | $0.85 | 20 | 0.9 sec | 100.0% | meets the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 99.0% | 92% | $0.23 | 80 | 1.2 sec | 100.0% | meets the bar | wrong value |
Contract clauses · Czech · Contract text (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $6.05 | 0 | 3.6 sec | 100.0% | meets the bar | — |
| 2 | DeepSeek V4.1 Flash DeepSeek · open weights | 100.0% | 100% | $0.42 | 0 | 4.3 sec | 100.0% | meets the bar | — |
| 3 | Gemini 3.8 Flash | 100.0% | 100% | $3.09 | 0 | 2.2 sec | 100.0% | meets the bar | — |
| 4 | Mistral Medium 3.5 Mistral | 100.0% | 100% | $3.06 | 0 | 1.2 sec | 100.0% | meets the bar | — |
| 5 | GPT-6 Luna OpenAI | 100.0% | 100% | $0.27 | 0 | 2.6 sec | 100.0% | meets the bar | — |
| 6 | GPT-6 Sol OpenAI | 100.0% | 100% | $4.62 | 0 | 2.0 sec | 100.0% | meets the bar | — |
| 7 | Gemma 4 31B Google · open weights | 100.0% | 99% | $0.25 | 7 | 6.0 sec | 100.0% | meets the bar | wrong value |
| 8 | Qwen3.8 Flash Qwen | 99.9% | 99% | $0.63 | 13 | 13.6 sec | 100.0% | meets the bar | wrong value |
| 9 | Gemini 3.5 Flash Lite | 99.9% | 98% | $0.85 | 20 | 0.9 sec | 100.0% | meets the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 99.0% | 92% | $0.23 | 80 | 1.2 sec | 100.0% | meets the bar | wrong value |
Contract clauses · German · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Luna OpenAI | 100.0% | 100% | $0.27 | 0 | 3.0 sec | 100.0% | meets the bar | — |
| 2 | GPT-6 Sol OpenAI | 100.0% | 100% | $4.36 | 0 | 2.9 sec | 100.0% | meets the bar | — |
| 3 | Gemini 3.8 Flash | 99.9% | 99% | $3.42 | 7 | 2.6 sec | 100.0% | meets the bar | invented value |
| 4 | Claude Sonnet 5 Anthropic | 98.2% | 87% | $7.11 | 133 | 4.3 sec | 100.0% | meets the bar | invented value |
| 5 | Gemini 3.5 Flash Lite | 98.0% | 84% | $0.79 | 160 | 1.0 sec | 100.0% | meets the bar | invented value |
| 6 | Qwen3.8 Flash Qwen | 98.0% | 87% | $0.78 | 127 | 19.7 sec | 99.3% | meets the bar | invented value |
| 7 | DeepSeek V4.1 Flash DeepSeek · open weights | 97.2% | 77% | $0.51 | 228 | 4.6 sec | 100.0% | meets the bar | invented value |
| 8 | Mistral Medium 3.5 Mistral | 97.1% | 75% | $2.76 | 247 | 1.2 sec | 100.0% | meets the bar | invented value |
| 9 | Gemma 4 31B Google · open weights | 94.4% | 27% | $0.22 | 727 | 2.6 sec | 100.0% | below the bar | number format |
| 10 | Mistral Small 4 Mistral · open weights | 94.0% | 45% | $0.21 | 553 | 1.2 sec | 100.0% | below the bar | invented value |
Contract clauses · German · Contract text (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Luna OpenAI | 100.0% | 100% | $0.27 | 0 | 3.0 sec | 100.0% | meets the bar | — |
| 2 | GPT-6 Sol OpenAI | 100.0% | 100% | $4.36 | 0 | 2.9 sec | 100.0% | meets the bar | — |
| 3 | Gemini 3.8 Flash | 99.9% | 99% | $3.42 | 7 | 2.6 sec | 100.0% | meets the bar | invented value |
| 4 | Claude Sonnet 5 Anthropic | 98.2% | 87% | $7.11 | 133 | 4.3 sec | 100.0% | meets the bar | invented value |
| 5 | Gemini 3.5 Flash Lite | 98.0% | 84% | $0.79 | 160 | 1.0 sec | 100.0% | meets the bar | invented value |
| 6 | Qwen3.8 Flash Qwen | 98.0% | 87% | $0.78 | 127 | 19.7 sec | 99.3% | meets the bar | invented value |
| 7 | DeepSeek V4.1 Flash DeepSeek · open weights | 97.2% | 77% | $0.51 | 228 | 4.6 sec | 100.0% | meets the bar | invented value |
| 8 | Mistral Medium 3.5 Mistral | 97.1% | 75% | $2.76 | 247 | 1.2 sec | 100.0% | meets the bar | invented value |
| 9 | Gemma 4 31B Google · open weights | 94.4% | 27% | $0.22 | 727 | 2.6 sec | 100.0% | below the bar | number format |
| 10 | Mistral Small 4 Mistral · open weights | 94.0% | 45% | $0.21 | 553 | 1.2 sec | 100.0% | below the bar | invented value |
Contract clauses · English (US) · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $5.04 | 0 | 3.4 sec | 100.0% | meets the bar | — |
| 2 | DeepSeek V4.1 Flash DeepSeek · open weights | 100.0% | 100% | $0.32 | 0 | 3.9 sec | 100.0% | meets the bar | — |
| 3 | GPT-6 Sol OpenAI | 100.0% | 100% | $3.33 | 0 | 2.5 sec | 100.0% | meets the bar | — |
| 4 | Gemini 3.5 Flash Lite | 100.0% | 99% | $0.69 | 7 | 1.0 sec | 100.0% | meets the bar | wrong value |
| 5 | Gemini 3.8 Flash | 99.9% | 99% | $2.47 | 7 | 2.0 sec | 100.0% | meets the bar | wrong value |
| 6 | GPT-6 Luna OpenAI | 99.8% | 99% | $0.20 | 13 | 2.7 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 99.6% | 97% | $0.18 | 27 | 2.5 sec | 100.0% | meets the bar | wrong value |
| 8 | Mistral Medium 3.5 Mistral | 98.1% | 87% | $2.35 | 127 | 1.1 sec | 100.0% | meets the bar | wrong value |
| 9 | Qwen3.8 Flash Qwen | 95.1% | 69% | $0.62 | 307 | 14.1 sec | 100.0% | meets the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 91.5% | 43% | $0.17 | 567 | 1.1 sec | 100.0% | below the bar | wrong value |
Contract clauses · English (US) · Contract text (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Sonnet 5 Anthropic | 100.0% | 100% | $5.04 | 0 | 3.4 sec | 100.0% | meets the bar | — |
| 2 | DeepSeek V4.1 Flash DeepSeek · open weights | 100.0% | 100% | $0.32 | 0 | 3.9 sec | 100.0% | meets the bar | — |
| 3 | GPT-6 Sol OpenAI | 100.0% | 100% | $3.33 | 0 | 2.5 sec | 100.0% | meets the bar | — |
| 4 | Gemini 3.5 Flash Lite | 100.0% | 99% | $0.69 | 7 | 1.0 sec | 100.0% | meets the bar | wrong value |
| 5 | Gemini 3.8 Flash | 99.9% | 99% | $2.47 | 7 | 2.0 sec | 100.0% | meets the bar | wrong value |
| 6 | GPT-6 Luna OpenAI | 99.8% | 99% | $0.20 | 13 | 2.7 sec | 100.0% | meets the bar | wrong value |
| 7 | Gemma 4 31B Google · open weights | 99.6% | 97% | $0.18 | 27 | 2.5 sec | 100.0% | meets the bar | wrong value |
| 8 | Mistral Medium 3.5 Mistral | 98.1% | 87% | $2.35 | 127 | 1.1 sec | 100.0% | meets the bar | wrong value |
| 9 | Qwen3.8 Flash Qwen | 95.1% | 69% | $0.62 | 307 | 14.1 sec | 100.0% | meets the bar | wrong value |
| 10 | Mistral Small 4 Mistral · open weights | 91.5% | 43% | $0.17 | 567 | 1.1 sec | 100.0% | below the bar | wrong value |
Personal data detection · Czech · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemma 4 31B Google · open weights | 99.8% | 99% | $0.12 | 13 | 2.9 sec | 100.0% | meets the bar | missing field |
| 2 | Mistral Medium 3.5 Mistral | 99.8% | 99% | $1.40 | 13 | 0.8 sec | 100.0% | meets the bar | invented value |
| 3 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.7% | 98% | $0.31 | 20 | 4.2 sec | 100.0% | meets the bar | missing field |
| 4 | GPT-6 Sol OpenAI | 99.7% | 98% | $2.22 | 20 | 2.5 sec | 100.0% | meets the bar | missing field |
| 5 | GPT-6 Luna OpenAI | 99.6% | 97% | $0.12 | 27 | 1.8 sec | 100.0% | meets the bar | missing field |
| 6 | Gemini 3.8 Flash | 99.5% | 97% | $2.19 | 33 | 2.0 sec | 100.0% | meets the bar | missing field |
| 7 | Claude Sonnet 5 Anthropic | 98.9% | 92% | $4.06 | 80 | 3.2 sec | 100.0% | meets the bar | invented value |
| 8 | Mistral Small 4 Mistral · open weights | 98.8% | 92% | $0.09 | 81 | 0.9 sec | 100.0% | meets the bar | invented value |
| 9 | Gemini 3.5 Flash Lite | 98.1% | 88% | $0.45 | 120 | 0.8 sec | 100.0% | meets the bar | invented value |
| 10 | Qwen3.8 Flash Qwen | 97.0% | 89% | $0.41 | 113 | 9.6 sec | 99.3% | meets the bar | wrong value |
Personal data detection · Czech · Text (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemma 4 31B Google · open weights | 99.8% | 99% | $0.12 | 13 | 2.9 sec | 100.0% | meets the bar | missing field |
| 2 | Mistral Medium 3.5 Mistral | 99.8% | 99% | $1.40 | 13 | 0.8 sec | 100.0% | meets the bar | invented value |
| 3 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.7% | 98% | $0.31 | 20 | 4.2 sec | 100.0% | meets the bar | missing field |
| 4 | GPT-6 Sol OpenAI | 99.7% | 98% | $2.22 | 20 | 2.5 sec | 100.0% | meets the bar | missing field |
| 5 | GPT-6 Luna OpenAI | 99.6% | 97% | $0.12 | 27 | 1.8 sec | 100.0% | meets the bar | missing field |
| 6 | Gemini 3.8 Flash | 99.5% | 97% | $2.19 | 33 | 2.0 sec | 100.0% | meets the bar | missing field |
| 7 | Claude Sonnet 5 Anthropic | 98.9% | 92% | $4.06 | 80 | 3.2 sec | 100.0% | meets the bar | invented value |
| 8 | Mistral Small 4 Mistral · open weights | 98.8% | 92% | $0.09 | 81 | 0.9 sec | 100.0% | meets the bar | invented value |
| 9 | Gemini 3.5 Flash Lite | 98.1% | 88% | $0.45 | 120 | 0.8 sec | 100.0% | meets the bar | invented value |
| 10 | Qwen3.8 Flash Qwen | 97.0% | 89% | $0.41 | 113 | 9.6 sec | 99.3% | meets the bar | wrong value |
Personal data detection · German · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Mistral Medium 3.5 Mistral | 100.0% | 100% | $1.38 | 0 | 0.9 sec | 100.0% | meets the bar | — |
| 2 | Gemma 4 31B Google · open weights | 99.9% | 99% | $0.10 | 7 | 1.9 sec | 100.0% | meets the bar | missing field |
| 3 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.8% | 99% | $0.31 | 13 | 3.7 sec | 100.0% | meets the bar | missing field |
| 4 | Claude Sonnet 5 Anthropic | 99.7% | 98% | $4.40 | 20 | 3.7 sec | 100.0% | meets the bar | invented value |
| 5 | Gemini 3.8 Flash | 99.7% | 98% | $2.12 | 20 | 2.3 sec | 100.0% | meets the bar | missing field |
| 6 | GPT-6 Sol OpenAI | 99.7% | 98% | $2.13 | 20 | 2.6 sec | 100.0% | meets the bar | missing field |
| 7 | Gemini 3.5 Flash Lite | 99.6% | 97% | $0.45 | 27 | 0.8 sec | 100.0% | meets the bar | invented value |
| 8 | GPT-6 Luna OpenAI | 99.6% | 97% | $0.12 | 27 | 2.2 sec | 100.0% | meets the bar | missing field |
| 9 | Qwen3.8 Flash Qwen | 99.6% | 99% | $0.36 | 7 | 12.2 sec | 100.0% | meets the bar | missing field |
| 10 | Mistral Small 4 Mistral · open weights | 99.1% | 94% | $0.09 | 65 | 0.9 sec | 100.0% | meets the bar | invented value |
Personal data detection · German · Text (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Mistral Medium 3.5 Mistral | 100.0% | 100% | $1.38 | 0 | 0.9 sec | 100.0% | meets the bar | — |
| 2 | Gemma 4 31B Google · open weights | 99.9% | 99% | $0.10 | 7 | 1.9 sec | 100.0% | meets the bar | missing field |
| 3 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.8% | 99% | $0.31 | 13 | 3.7 sec | 100.0% | meets the bar | missing field |
| 4 | Claude Sonnet 5 Anthropic | 99.7% | 98% | $4.40 | 20 | 3.7 sec | 100.0% | meets the bar | invented value |
| 5 | Gemini 3.8 Flash | 99.7% | 98% | $2.12 | 20 | 2.3 sec | 100.0% | meets the bar | missing field |
| 6 | GPT-6 Sol OpenAI | 99.7% | 98% | $2.13 | 20 | 2.6 sec | 100.0% | meets the bar | missing field |
| 7 | Gemini 3.5 Flash Lite | 99.6% | 97% | $0.45 | 27 | 0.8 sec | 100.0% | meets the bar | invented value |
| 8 | GPT-6 Luna OpenAI | 99.6% | 97% | $0.12 | 27 | 2.2 sec | 100.0% | meets the bar | missing field |
| 9 | Qwen3.8 Flash Qwen | 99.6% | 99% | $0.36 | 7 | 12.2 sec | 100.0% | meets the bar | missing field |
| 10 | Mistral Small 4 Mistral · open weights | 99.1% | 94% | $0.09 | 65 | 0.9 sec | 100.0% | meets the bar | invented value |
Personal data detection · English (US) · All levels (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemma 4 31B Google · open weights | 99.9% | 99% | $0.08 | 7 | 2.4 sec | 100.0% | meets the bar | missing field |
| 2 | Claude Sonnet 5 Anthropic | 99.7% | 98% | $3.50 | 20 | 2.5 sec | 100.0% | meets the bar | missing field |
| 3 | Gemini 3.8 Flash | 99.7% | 98% | $1.86 | 20 | 1.9 sec | 100.0% | meets the bar | missing field |
| 4 | Mistral Medium 3.5 Mistral | 99.7% | 98% | $1.15 | 20 | 0.8 sec | 100.0% | meets the bar | missing field |
| 5 | GPT-6 Sol OpenAI | 99.7% | 98% | $2.09 | 20 | 2.6 sec | 100.0% | meets the bar | missing field |
| 6 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.6% | 97% | $0.29 | 27 | 3.8 sec | 100.0% | meets the bar | missing field |
| 7 | GPT-6 Luna OpenAI | 99.6% | 97% | $0.11 | 27 | 1.8 sec | 100.0% | meets the bar | invented value |
| 8 | Gemini 3.5 Flash Lite | 99.5% | 97% | $0.41 | 27 | 0.8 sec | 100.0% | meets the bar | missing field |
| 9 | Mistral Small 4 Mistral · open weights | 99.4% | 96% | $0.08 | 40 | 0.8 sec | 100.0% | meets the bar | invented value |
| 10 | Qwen3.8 Flash Qwen | 99.2% | 99% | $0.28 | 13 | 7.5 sec | 100.0% | meets the bar | invented value |
Personal data detection · English (US) · Text (150 documents)
| # | Model | Correct fields | Error-free documents | Per 1,000 documents | To fix / 1,000 | Speed (median) | Valid answers | Bar | Most common error |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemma 4 31B Google · open weights | 99.9% | 99% | $0.08 | 7 | 2.4 sec | 100.0% | meets the bar | missing field |
| 2 | Claude Sonnet 5 Anthropic | 99.7% | 98% | $3.50 | 20 | 2.5 sec | 100.0% | meets the bar | missing field |
| 3 | Gemini 3.8 Flash | 99.7% | 98% | $1.86 | 20 | 1.9 sec | 100.0% | meets the bar | missing field |
| 4 | Mistral Medium 3.5 Mistral | 99.7% | 98% | $1.15 | 20 | 0.8 sec | 100.0% | meets the bar | missing field |
| 5 | GPT-6 Sol OpenAI | 99.7% | 98% | $2.09 | 20 | 2.6 sec | 100.0% | meets the bar | missing field |
| 6 | DeepSeek V4.1 Flash DeepSeek · open weights | 99.6% | 97% | $0.29 | 27 | 3.8 sec | 100.0% | meets the bar | missing field |
| 7 | GPT-6 Luna OpenAI | 99.6% | 97% | $0.11 | 27 | 1.8 sec | 100.0% | meets the bar | invented value |
| 8 | Gemini 3.5 Flash Lite | 99.5% | 97% | $0.41 | 27 | 0.8 sec | 100.0% | meets the bar | missing field |
| 9 | Mistral Small 4 Mistral · open weights | 99.4% | 96% | $0.08 | 40 | 0.8 sec | 100.0% | meets the bar | invented value |
| 10 | Qwen3.8 Flash Qwen | 99.2% | 99% | $0.28 | 13 | 7.5 sec | 100.0% | meets the bar | invented value |