Facts & Figures
Bank statement extraction accuracy: facts and figures
Benchmark of bank statement transaction extraction on 47 real European banks: statement-level zero-error rate, fabricated-row rate, recall and precision for Holofin, GPT-5.5, Gemini 3.1 Pro and Claude Opus 4.8. CSV included.
- Holofin statements with zero errors
- 98%
- Holofin errored row across 44 statements
- 1
- errored rows per frontier model
- 70 to 115
- banks, one statement each, gold hand-verified
- 47
On this page
Summary
On a corpus of 47 real, anonymised statements from 47 different European banks (93 pages), Holofin returns 98% of statements with zero errors, one errored row across the 44 scored documents. The best frontier model, Gemini 3.1 Pro, returns 80% of statements with zero errors; GPT-5.5 77%; Claude Opus 4.8 75%.
The gap is fabrication, not reading. Recall is 0.93 to 1.00 for every system, but 8.3% to 10.0% of the rows a frontier model returns are not on the page. Holofin's fabricated-row rate is 0.1%. Giving the frontier models two pages or the whole document per call moves recall by about one point and does not close the gap.
Data
Statement-level results, 44 scored statements, per-page window
| System | Statements with zero errors | Errored rows (all 44 docs) | Fabricated-row rate | Recall | Precision |
|---|---|---|---|---|---|
| Holofin (production pipeline) | 98% | 1 | 0.1% | 1.000 | 0.999 |
| Gemini 3.1 Pro | 80% | 115 | 10.0% | 0.931 | 0.900 |
| GPT-5.5 | 77% | 84 | 8.3% | 0.939 | 0.917 |
| Claude Opus 4.8 | 75% | 70 | 9.2% | 0.929 | 0.908 |
A statement counts as correct only if every transaction row matches the hand-verified gold on (date, signed amount) at cent precision: no dropped rows, no fabricated rows. Fabricated-row rate is the share of returned rows whose (date, amount) is not on the page. Frontier models shown at their best (per-page) setting.
Download CSVRecall by context window, frontier models
| Model | Per page | Two pages | Whole document |
|---|---|---|---|
| GPT-5.5 | 0.939 | 0.942 | 0.932 |
| Gemini 3.1 Pro | 0.931 | 0.953 | 0.932 |
| Claude Opus 4.8 | 0.929 | 0.948 | 0.940 |
Feeding more pages per call is a wash: recall moves by about one point in either direction. Holofin runs one page at a time and is not shown because it scored 1.000 in its only configuration.
Download CSVBenchmark corpus: 47 banks, 93 pages, one statement per bank
| Bank | Country | Pages |
|---|---|---|
| BAMI Banque Michel Inchauspé | FR | 4 |
| Banque Dupuy de Parseval | FR | 1 |
| Banque Transatlantique | FR | 2 |
| Berliner Sparkasse | DE | 1 |
| Berliner Volksbank | DE | 1 |
| BNP Paribas | FR | 1 |
| BoursoBank | FR | 1 |
| BRED Banque Populaire | FR | 2 |
| bunq | NL | 2 |
| BW-Bank | DE | 2 |
| Caisse d'Epargne | FR | 2 |
| Commerzbank | DE | 2 |
| Crédit Agricole Brie Picardie | FR | 1 |
| Crédit Coopératif | FR | 2 |
| CIC | FR | 2 |
| Crédit Mutuel | FR | 1 |
| Deutsche Bank | DE | 2 |
| Deutsche Skatbank | DE | 2 |
| DKB Deutsche Kreditbank | DE | 3 |
| Fiducial Banque | FR | 1 |
| Finom | NL | 1 |
| Grenke Bank | DE | 3 |
| HSBC | FR | 1 |
| HypoVereinsbank | DE | 2 |
| iBanFirst | FR | 3 |
| Kontist | DE | 2 |
| La Banque Postale | FR | 3 |
| LCL | FR | 1 |
| Manager.one | FR | 2 |
| Mein ELBA (Raiffeisen) | AT | 3 |
| Memo Bank | FR | 1 |
| Monabanq | FR | 2 |
| Oberbank | AT | 1 |
| PayPal | LU | 4 |
| Postbank | DE | 1 |
| Qonto | FR | 1 |
| Raiffeisenbank Südstormarn Mölln | DE | 8 |
| Revolut Business | LT | 1 |
| SG Crédit du Nord | FR | 2 |
| Société Générale | FR | 1 |
| Shine | FR | 1 |
| Sparda-Bank | DE | 3 |
| SumUp | GB | 4 |
| Targobank | DE | 4 |
| UniCredit | DE | 1 |
| Viva Wallet | GR | 1 |
| Wise | BE | 1 |
Every statement is real, then anonymised so layout, tables and totals survive but names and numbers are synthetic. Gold labels were hand-verified line by line against the source PDFs. 44 of the 47 statements are scored in the results tables. Country is the issuing entity's home market.
Download CSVMethodology
Frontier candidates receive page images with a generic extraction prompt at three context sizes (per page, two pages, whole document). Holofin is the production pipeline (classify, OCR, per-page extract) driven over HTTP, one page at a time.
Every metric is document-macro: computed per statement, then averaged. A row matches gold when (transaction_date, signed amount) is exact at cent precision. A statement has zero errors when every gold row is returned and no returned row is missing from gold.
The corpus is one statement per distinct bank, picked for layout diversity rather than weighted by traffic. It over-represents rare and awkward layouts on purpose, so read the numbers as a worst-case probe of reliability, not a forecast of average production accuracy.
Balance reconciliation (opening balance plus transactions equals closing balance) was measured and is necessary but not sufficient: GPT-5.5 reconciles 42 of 45 statements yet still fabricates about 8% of rows, and Gemini 3.1 Pro left balances blank on 12 documents, which cannot be reconciled at all.
Spotted an error? Write to [email protected] and we will fix the row and note the change.
Sources
- Holofin, The Bank Statement Extraction Benchmark (full article, charts and corpus thumbnails)(accessed Sept. 3, 2026)
Frequently asked questions
Why score per statement rather than per row?
A statement is boolean for the team that receives it: either every transaction is right or the document is a liability. A statement with 27 of 30 rows right is 90% accurate per row and still wrong. In this benchmark, Gemini 3.1 Pro, GPT-5.5 and Claude Opus 4.8 returned a fully correct statement 75 to 80% of the time, even though they found most rows.
What counts as a fabricated row?
A returned row whose (date, signed amount) matches no transaction on the page. Traced by hand, 68% to 93% of these (by model) have no counterpart at all; the rest are a real row read with a wrong amount or date.
Does a larger context window fix it?
No. Across the three frontier models, two-page and whole-document windows move recall by roughly one point in either direction. The failure is per-layout, not per-window.
Is balance reconciliation enough to validate an extraction?
No. It is a necessary check, not a sufficient one. A fabricated row offset by another error still ties out, and a model that omits balances cannot be checked at all. Gold must be verified against the page.
Can I reuse these figures?
Yes, under CC BY 4.0 with attribution to Holofin. Each table has a CSV download; the full methodology is in the linked article.
Bank statements you can lend on.
Send Holofin your own statements: every row reconciled, every PDF checked for tampering, every value linked to the page it came from.