Automate data extraction where volume, structure and legibility coincide

Updated

Automating data extraction is a general technique with a narrow economic sweet spot. It repays where three conditions hold together, and it disappoints wherever one is missing. Checking your candidate documents against the three before comparing any products prevents most of the disappointment, and it usually shortens the list of documents worth automating considerably.

Volume, so setup amortises

Enough documents that configuration and maintenance pay for themselves. A document type arriving twice a year cannot repay setup however structured it is, which is why contracts are almost never worth extracting from in a finance team even though they are important documents full of consequential terms.

Structure, so there are fields

A defined set of fields, in broadly predictable places, meaning the same thing each time. Invoices qualify; correspondence does not, because its content is prose and its value is meaning rather than fields. Delivery notes across many suppliers sit awkwardly between the two and are usually better captured at source.

A legible sender

Somebody with a reason to be readable. A supplier billing you wants to be paid and prints clearly. A handwritten note from a site has no such incentive and no consistency, and where the sender has no reason to be legible, capturing at the point of creation with a short form beats extracting afterwards.

Then validate whatever comes out

Supplier exists, reference is new, arithmetic works, date is plausible, bank details match. Extraction without validation is a faster route to a wrong number in a ledger, and the validation rules are usually the caller's responsibility rather than the extraction product's, which is easy to overlook.

Questions people ask about automate data extraction

What if a document type fails one condition?

Capture it at source with a form, or handle it manually. Extraction is not the only way to obtain structured data and it is the wrong way when the input is unstructured by nature.

How do we measure a trial?

By the human minutes still needed per document afterwards, on your own worst inputs, rather than by an accuracy figure that averages easy fields with hard ones.

Does it learn?

Often per supplier. Ask what a correction changes, whether it is local or global, and how soon the effect appears in your own results.

Sources

Related answers

Start Threewayly ProKeep the match, not the spreadsheet