DEMO 01 / DOCUMENT PROCESSING
Receipt Reader
Drop a receipt or invoice — PDF or photo — and get clean spreadsheet rows back: merchant, date, VAT and every line item, checked against the total.
No file handy? Try a sample
Your file never leaves your browser. Text is read on this device by pdf.js and a WebAssembly build of Tesseract. This server has no upload endpoint — open your browser's network tab and watch.
Nothing read yet
Pick a sample above, or drop your own receipt. The first photo takes a few seconds longer: the OCR engine and its English + Latvian models (about 7 MB) download once and are then cached.
How it works
-
01
Read
A PDF with a text layer is read directly by pdf.js: exact characters, no guessing. Scans and photos are deskewed and read by Tesseract's LSTM engine (English + Latvian), compiled to WebAssembly and run in a worker thread in your browser.
-
02
Rebuild rows
OCR returns words with coordinates, not lines. Rows are rebuilt from geometry, so a table that page segmentation cut into a names column and a prices column comes back together, and two header columns on one baseline are split apart.
-
03
Parse
A deterministic parser, with no AI model, finds the totals block first (KOPĀ, Total due, PVN 21%), then reads the item lines above it: 2 x 1,50, 3 gab, 0,845 kg x 1,49 EUR/kg, comma decimals, discounts, dd.mm.yyyy dates.
-
04
Cross-check
Items are summed against the total, subtotal plus VAT against the total, and VAT against its rate. When the numbers agree, confidence goes up. When they don't, the check says by how much and the doubtful cells are highlighted for you to fix.