OCR invoice scanning software turns a picture of an invoice into fields a system can use: supplier, invoice number, date, totals and, if you are lucky, the lines. It is genuinely useful and it is routinely oversold, because the accuracy figure a vendor quotes is an average across fields that behave completely differently. Knowing which fields are the easy ones tells you how much human review you will still be paying for after you buy it.
The fields it gets right almost always
Totals, dates and invoice numbers are printed in predictable places in predictable formats, and modern extraction handles them well enough that reviewing them is usually wasted effort. Supplier identity is nearly as good when the supplier is already in your system, because the software is matching against a known list rather than reading freely. These are the fields worth trusting and spot-checking rather than checking one by one.
The fields it guesses at
Line items are the hard part, and they are the part the three-way match needs. Descriptions vary between what you ordered and what the supplier calls it, units of measure differ, and a single ordered line often arrives as several invoiced lines or the reverse. Tax treatment and general ledger coding are worse still, because they are not printed on the document at all: they are a judgement about the transaction, and no amount of reading the image produces them.
What that means for how you buy
Ask a vendor for accuracy by field rather than in aggregate, and ask specifically about line level extraction on your own invoices rather than on their samples. Then decide what review you will do on the fields that are guessed. A capture step that produces confident wrong data and no review queue is worse than keying, because a person keying knows they are the control and a person glancing at a screen full of green ticks does not.
Questions people ask about ocr invoice scanning software
Does OCR remove the need to check invoices?
No. It removes most of the keying. The check that matters is the match against the order and the receipt, which is a comparison of documents rather than a reading of one, and OCR does not perform it.
Do I need scanning if my invoices arrive as PDFs?
A PDF generated by a supplier's system usually carries a text layer, so extraction is easier and more accurate than from a scanned image. Emailed PDFs are the best case; photographs of paper are the worst.
What accuracy should I expect?
Ask for it by field and test it on your own documents, which is the only number that means anything. Aggregate accuracy figures blend easy fields with hard ones and are not comparable between vendors.