Organisations comparing automated invoice scanning products usually find them close, and then find their results vary anyway. The variation is nearly always upstream of the software: what the documents look like when they arrive, how consistently they are captured, and whether anybody scans the same document twice. Those three explain more than any engine difference.
What good input looks like
A flat, straight page at a sensible resolution with good contrast, captured once, at arrival. A supplier's own generated PDF is better still, because it carries a text layer and nothing has to be recognised at all. Where you can ask suppliers to email rather than post, that single change outperforms most software decisions.
What bad input looks like
A photograph taken at an angle in poor light, a scan of a photocopy, a page with a stamp over the total, or the same invoice captured twice at different settings by two people. None of these is fixable downstream, and all of them are cheap to fix at the point of capture with a short standard and an app that rejects a poor image.
Consistency beats cleverness
One resolution, one orientation, both sides where relevant, captured by the same route. Consistency produces better results than a superior engine fed a mixture, reliably, and it costs nothing but a written standard and an occasional check. This is the least interesting improvement available and usually the largest.
Questions people ask about automated invoice scanning
How do we know if input is our problem?
Look at where corrections cluster for a month. If they follow poor images rather than particular fields, no engine change will help.
Should we scan at arrival?
Yes. It gives a reliable received date and makes the invoice visible immediately, and it prevents annotated copies existing before the canonical one.
What resolution is enough?
High enough that small print and stamps stay legible when zoomed, tested on your worst supplier document.