Automatic data capture is often evaluated as a technology and is better evaluated as an economic proposition. It pays where three conditions hold at once, and it disappoints wherever one of them is missing. Checking your candidate documents against the three, before comparing any products, prevents most of the disappointment in this area.
Frequency
Enough documents that setup and maintenance amortise. A document type arriving twice a year cannot repay configuration however structured it is. This is why invoices are the classic case and why contracts, however important, are almost never worth automating extraction from in a finance team.
Structure
A defined set of fields, in broadly predictable places, meaning the same thing each time. Invoices qualify; correspondence does not, because its content is prose and the interesting part is meaning rather than fields. Documents that are structured but inconsistent, such as delivery notes across many suppliers, sit awkwardly in between.
A legible sender
Somebody with a reason to be readable. A supplier billing you wants to be paid and prints clearly. A handwritten note from a site has no such incentive and no consistency. Where the sender has no reason to be legible, capture at the point of creation, with a form or an app, beats extraction from whatever they produced.
Questions people ask about automatic data capture
What if a document type fails one condition?
Handle it manually or capture it at source with a form instead. Extraction is not the only way to get structured data, and it is the wrong way when the input is unstructured by nature.
Does capture need validation?
Always. Extracted values checked against what you already know catch a large share of errors before they reach a ledger.
How do we measure a trial?
By the human minutes still needed per document afterwards, on your own worst inputs, rather than by an accuracy figure.