Facts & Figures

Bank statement extraction accuracy: facts and figures

Benchmark of bank statement transaction extraction on 47 real European banks: statement-level zero-error rate, fabricated-row rate, recall and precision for Holofin, GPT-5.5, Gemini 3.1 Pro and Claude Opus 4.8. CSV included.

By Holofin Research Published Updated
Holofin statements with zero errors
98%
Holofin errored row across 44 statements
1
errored rows per frontier model
70 to 115
banks, one statement each, gold hand-verified
47
On this page

Summary

On a corpus of 47 real, anonymised statements from 47 different European banks (93 pages), Holofin returns 98% of statements with zero errors, one errored row across the 44 scored documents. The best frontier model, Gemini 3.1 Pro, returns 80% of statements with zero errors; GPT-5.5 77%; Claude Opus 4.8 75%.

The gap is fabrication, not reading. Recall is 0.93 to 1.00 for every system, but 8.3% to 10.0% of the rows a frontier model returns are not on the page. Holofin's fabricated-row rate is 0.1%. Giving the frontier models two pages or the whole document per call moves recall by about one point and does not close the gap.

Data

Statement-level results, 44 scored statements, per-page window

Statement-level results, 44 scored statements, per-page window
SystemStatements with zero errorsErrored rows (all 44 docs)Fabricated-row rateRecallPrecision
Holofin (production pipeline)98%10.1%1.0000.999
Gemini 3.1 Pro80%11510.0%0.9310.900
GPT-5.577%848.3%0.9390.917
Claude Opus 4.875%709.2%0.9290.908

A statement counts as correct only if every transaction row matches the hand-verified gold on (date, signed amount) at cent precision: no dropped rows, no fabricated rows. Fabricated-row rate is the share of returned rows whose (date, amount) is not on the page. Frontier models shown at their best (per-page) setting.

Download CSV

Recall by context window, frontier models

Recall by context window, frontier models
ModelPer pageTwo pagesWhole document
GPT-5.50.9390.9420.932
Gemini 3.1 Pro0.9310.9530.932
Claude Opus 4.80.9290.9480.940

Feeding more pages per call is a wash: recall moves by about one point in either direction. Holofin runs one page at a time and is not shown because it scored 1.000 in its only configuration.

Download CSV

Benchmark corpus: 47 banks, 93 pages, one statement per bank

Benchmark corpus: 47 banks, 93 pages, one statement per bank
BankCountryPages
BAMI Banque Michel InchauspéFR4
Banque Dupuy de ParsevalFR1
Banque TransatlantiqueFR2
Berliner SparkasseDE1
Berliner VolksbankDE1
BNP ParibasFR1
BoursoBankFR1
BRED Banque PopulaireFR2
bunqNL2
BW-BankDE2
Caisse d'EpargneFR2
CommerzbankDE2
Crédit Agricole Brie PicardieFR1
Crédit CoopératifFR2
CICFR2
Crédit MutuelFR1
Deutsche BankDE2
Deutsche SkatbankDE2
DKB Deutsche KreditbankDE3
Fiducial BanqueFR1
FinomNL1
Grenke BankDE3
HSBCFR1
HypoVereinsbankDE2
iBanFirstFR3
KontistDE2
La Banque PostaleFR3
LCLFR1
Manager.oneFR2
Mein ELBA (Raiffeisen)AT3
Memo BankFR1
MonabanqFR2
OberbankAT1
PayPalLU4
PostbankDE1
QontoFR1
Raiffeisenbank Südstormarn MöllnDE8
Revolut BusinessLT1
SG Crédit du NordFR2
Société GénéraleFR1
ShineFR1
Sparda-BankDE3
SumUpGB4
TargobankDE4
UniCreditDE1
Viva WalletGR1
WiseBE1

Every statement is real, then anonymised so layout, tables and totals survive but names and numbers are synthetic. Gold labels were hand-verified line by line against the source PDFs. 44 of the 47 statements are scored in the results tables. Country is the issuing entity's home market.

Download CSV

Methodology

Frontier candidates receive page images with a generic extraction prompt at three context sizes (per page, two pages, whole document). Holofin is the production pipeline (classify, OCR, per-page extract) driven over HTTP, one page at a time.

Every metric is document-macro: computed per statement, then averaged. A row matches gold when (transaction_date, signed amount) is exact at cent precision. A statement has zero errors when every gold row is returned and no returned row is missing from gold.

The corpus is one statement per distinct bank, picked for layout diversity rather than weighted by traffic. It over-represents rare and awkward layouts on purpose, so read the numbers as a worst-case probe of reliability, not a forecast of average production accuracy.

Balance reconciliation (opening balance plus transactions equals closing balance) was measured and is necessary but not sufficient: GPT-5.5 reconciles 42 of 45 statements yet still fabricates about 8% of rows, and Gemini 3.1 Pro left balances blank on 12 documents, which cannot be reconciled at all.

Spotted an error? Write to [email protected] and we will fix the row and note the change.

Sources

  1. Holofin, The Bank Statement Extraction Benchmark (full article, charts and corpus thumbnails)(accessed Sept. 3, 2026)

License: CC BY 4.0. Reuse with attribution to Holofin.

Frequently asked questions

Why score per statement rather than per row?

A statement is boolean for the team that receives it: either every transaction is right or the document is a liability. A statement with 27 of 30 rows right is 90% accurate per row and still wrong. In this benchmark, Gemini 3.1 Pro, GPT-5.5 and Claude Opus 4.8 returned a fully correct statement 75 to 80% of the time, even though they found most rows.

What counts as a fabricated row?

A returned row whose (date, signed amount) matches no transaction on the page. Traced by hand, 68% to 93% of these (by model) have no counterpart at all; the rest are a real row read with a wrong amount or date.

Does a larger context window fix it?

No. Across the three frontier models, two-page and whole-document windows move recall by roughly one point in either direction. The failure is per-layout, not per-window.

Is balance reconciliation enough to validate an extraction?

No. It is a necessary check, not a sufficient one. A fabricated row offset by another error still ties out, and a model that omits balances cannot be checked at all. Gold must be verified against the page.

Can I reuse these figures?

Yes, under CC BY 4.0 with attribution to Holofin. Each table has a CSV download; the full methodology is in the linked article.

Bank statements you can lend on.

Send Holofin your own statements: every row reconciled, every PDF checked for tampering, every value linked to the page it came from.

Holofin