ModelVerdict

Methodology

Methodology: how we measure

Every number on this site can be traced to a specific document, model output and run cost. This is how they are produced.

What we measure

Three everyday business tasks: extracting 13 fields from an invoice (supplier, identifiers, document number, payment reference, dates, net and tax by rate, total, currency, bank account); triaging customer emails (category, priority, sentiment, customer, order and invoice number, amount, currency, deadline); and finding personal data in business texts, e.g. before they are sent to an AI service (names, emails, phones, addresses, dates of birth, national IDs, bank accounts). We report the share of correct fields and the share of documents without a single error, because for a business the second number often matters more. Documents are in Czech and US English.

The test set

All documents are synthetic: generated from random but valid data (Czech company IDs and birth numbers with checksums, bank accounts per Czech National Bank rules, EINs, ABA routing numbers), so the correct answer is the source data, not a manual transcription, and has no typos. No real personal or company data: email addresses use .example domains, US phone numbers the fictional 555-01xx range, SSNs a range that is never issued. A Czech and an English document form a pair with the same structure, which lets us measure the language gap.

Invoices come in eight layouts with realistic variants (not VAT-registered, deposit invoice, credit note, foreign currency) and at four levels: 30% PDF text, 30% clean image, 30% scan (rotation, noise, JPEG), 10% photo (perspective, shadow). One third of every set is non-public; from it we publish only aggregate scores, so a model cannot score well by having seen the answers.

Emails come as a single message, with a long signature, as a thread with quoted earlier messages and as an email forwarded by a colleague; deadlines are also relative (“by Friday”). Priority follows fixed rules written in the prompt, so it is not a matter of opinion. Texts with personal data (HR records, support tickets, meeting minutes, emails) deliberately contain data that is not personal: company names made of surnames, company IDs, generic addresses such as info@, order numbers and dates that are not dates of birth.

How a run works

All models are called through OpenRouter with the same prompt in the document's language, at temperature 0, each model on one fixed provider so the results are reproducible. If a model supports structured output (JSON schema), we use it; if the provider rejects the schema, we ask again without it and record that. If the answer is not valid JSON, we ask once more and record that too. For every call we store the model, prompt and dataset versions, tokens, cost, latency and provider. Outputs are cached, so each week we only pay for new models, documents and prompts.

How we score

A field is correct when it matches the correct value exactly after normalising the format: dates to YYYY-MM-DD, amounts to a number with two decimals (both separator conventions), whitespace in identifiers and account numbers ignored. Diacritics are not normalised: “Horak” instead of “Horák” is an error. Errors are categorised (missing field, wrong value, lost diacritics, swapped dates, net instead of total, wrong rate, number format, invalid JSON). The bar for invoices is 95% correct fields.

Emails: category, priority and sentiment must match exactly; relative deadlines must be resolved to the right date. Personal data: each field is a list and is correct only when it is complete and has nothing extra; a company number reported as personal data counts as an invented value. The bar is 95% correct fields for every task.

Costs and the language tax

Cost is the real token spend as returned by the API in USD, scaled to 1,000 documents. Conversion to other currencies happens only for display, at the rate shown in the footer. The language tax is measured on a parallel corpus of 50 paragraphs in Czech, German, French, Spanish, Polish and English: we send the text with a minimal output limit, subtract the overhead of an empty prompt and compare input tokens with English.

True cost adds the work of a person who has to find and fix every document with an error: 3 minutes per document at 450 CZK / 35 EUR / 40 USD per hour. Speed is the median answer time (90th percentile in the tooltip); valid answers is the share of documents with valid JSON. Rate-limit errors of the API are not counted against a model: they are retried until every document has an answer.

The real cost calculator uses measured costs where we have documents in the language. For other languages it estimates: the surcharge measured on real documents in another language, scaled by the language tax of the chosen language (an invoice image costs the same in every language, so the plain token ratio would overstate it). Estimates are marked ≈ and have no accuracy.

The language gap

The accuracy difference between Czech and English documents is computed only on fields both versions have. US invoices have no VAT ID or tax point, so those fields are left out of the comparison.

Limitations

Synthetic documents do not cover everything a real company sees (handwritten notes, multi-page invoices, unusual layouts). With about 150 documents per language, accuracy is uncertain by roughly ±1 percentage point; do not treat smaller differences between models as decisive. Models and prices change, so every number shows the run date.

Independence

Rankings and recommendations cannot be bought. No model vendor pays us for placement. If you find an error in a correct answer or in the scoring, let us know; we fix it and note the change.

Current run: 20260928T085806Z-5d2ece of September 28, 2026, dataset version 0.1.0.