Document extraction
Extract invoice data, line by line
Most tools stop at the header: supplier, date, total. Extractify goes down to the detail, every product with its quantity, unit price and line total, because that is where price rises and billing errors hide.
What gets extracted
Document header
- Supplier name, company registration and VAT numbers
- Invoice or delivery note number
- Document date
- Totals excluding and including tax
Every product line
- Wording exactly as the supplier writes it
- Reference or article code
- Quantity and packaging
- Unit price and line total
How a document is read
Four methods, tried in this order. Each result is checked before it is accepted: if it does not hold up, the next method takes over. AI is not the first reflex, it is the last resort.
- 1
Electronic invoice
If the PDF carries a Factur-X invoice, the data is already there, structured and exact. Nothing to guess, only to read.
- 2
Native PDF text
A PDF produced by billing software carries its own text. We split it into columns and rebuild the lines, checking that quantity times unit price really gives the stated total.
- 3
Character recognition
For scanned or photographed documents, where there are only pixels. The same arithmetic check applies.
- 4
AI model
As a last resort, on layouts the earlier steps could not untangle: nested columns, several notes in one file, a photo taken at an angle.
That order has a direct effect on your bill: reading a document from its native text costs nothing, while a call to an AI model is paid for. The earlier the cascade stops, the cheaper the reading.
What we accept
Factur-X and electronic invoices
The PDF and its structured data, read on arrival. This is the format becoming standard in France with the obligation in force since September 2026.
Ordinary PDFs
Your suppliers', with or without embedded text, on one page or many, including files that hold several delivery notes.
Photos and scans
JPEG, PNG, a delivery note photographed on a phone at the site or in the kitchen, with the image straightened.
Email attachments
A dedicated address receives documents, and your suppliers can write to it directly. ZIP archives are accepted and opened.
An extraction has to be checked
No automatic reading is 100 % reliable, and claiming otherwise would be a lie. What matters is making the mistake visible and quick to fix.
- Every line carries a confidence score, and the ones worth a look are flagged rather than lost in the pile.
- The original document stays on screen next to the data, with the line pinpointed on it: checking a price takes seconds.
- Arithmetic inconsistencies are caught before saving; a line whose amounts do not add up does not slip through quietly.
- The same document imported twice is recognised and refused, even if it changes supplier along the way.
- Your corrections are kept and serve later imports from the same supplier.
What we do not extract
- Contracts. Extractify reads invoices, delivery notes, quotes and purchase orders, not contractual clauses.
- Handwriting. Handwritten documents are not read reliably, and we would rather say so than produce wrong figures.
- Screens from other software. We read commercial documents, not screenshots of your ERP.
Frequently asked questions
Do we have to set up a template per supplier?
No. There is no layout to draw and no zone to mark out. Unusual formats are handled inside the product, not by you.
What happens with an unknown supplier?
Its name and legal identifiers are read from the document, with checksum validation. When in doubt, the document waits for your decision instead of being attached at random.
Are credit notes handled?
Yes, and properly: a credit note is recognised by its wording, not by the sign of its amounts which reading often loses, then deducted from your purchases instead of added to them.
How many documents can be processed?
That depends on your plan. Documents are processed in a queue, so you can drop dozens at once and close the tab.
