Glossary
What is invoice OCR, and why is it not enough?
By Tim White · Last updated
Invoice OCR (optical character recognition) converts a scanned or photographed invoice into machine-readable text. It is necessary for paper and photos, but on its own it only produces raw text: it does not know which number is the total or whether it read that number correctly. That is why OCR is only a first step.
What OCR does
OCR reads the characters off an image. For a digital PDF the text is already there; for a scan or a phone photo, OCR is what makes the document readable at all. InvoiceJet uses an OCR fallback so scans and photographs work alongside clean PDFs, and it handles European number and date formats.
What OCR cannot do
OCR gives you text, not understanding. It will not tell you which line is the vendor, which figure is tax versus total, or whether glare over the tax line caused a misread. It has no notion of confidence, so a wrong character looks exactly like a right one.
That is the gap InvoiceJet closes: two models cross-check the reading, math checks confirm the totals are internally consistent, and every field carries a confidence level with its source highlighted on the document.
What trips people up
OCR quality is decided before the software runs: by the capture. A photo taken at an angle, in a shadow, or at low resolution gives the reader warped or broken characters, and a total read off a glare-washed line can be off by a digit. If you photograph invoices, lay the page flat, fill the frame, and skip the hard shadows. When a capture is still poor, the per-field confidence level is what tells you which figures to check rather than trust.
Common questions
Does InvoiceJet work on phone photos of invoices?
Yes. Photographed and scanned documents are handled through an OCR fallback, and each field still carries a confidence level so you can see where a photo caused uncertainty.
Is OCR enough to automate invoice processing on its own?
No. OCR only produces raw text with no sense of which number is the total or whether it read correctly, so it has to be paired with extraction that structures and verifies each field.
Do digital PDFs need OCR?
Usually not. A digital PDF already contains selectable text, so OCR mainly matters for scans and phone photos where the text exists only as an image.
Sources
Keep reading
Turn your next invoice into verified data
Free for 10 invoices a month, no card. Every field carries a confidence level and cites its source.