Insights

AI & automationJuly 20266 min read

Extracting your invoices and contracts: why vision models replaced OCR

For years, automating the reading of a document meant stacking character recognition and brittle rules, which broke the first time a supplier changed its layout. That is no longer the case. Here is what changed, and how to tell if it pays off for you.

By Nathan · guinat6 min read

Take any company that receives invoices: somewhere, someone is still re-keying amounts and reference numbers into accounting software, field by field. It looks harmless enough, yet it costs a small fortune over a year, and for a long time the only credible alternative was almost worse: an OCR system loaded with rules that had to be repaired with every new supplier. Vision models have changed that equation. They read a page the way you read it, without being told in advance where each piece of information sits. Here is what that changes in practice, where it genuinely pays off, and where you should keep a cool head.

What does a document read by hand really cost you?

The numbers are obvious once you write them down. An invoice re-keyed by hand is a few minutes per document, multiplied by your volume. Over hundreds of documents a month, that adds up to whole days of work that show up on no budget line. But the keying time is only the visible part.

The real cost is what comes after. An amount copied wrong that only surfaces much later, a supplier you have to call back, a late payment, VAT booked to the wrong account. And a quieter cost: no one wants to spend the day typing numbers, and it is often your best people who end up doing it for lack of anything better. It is one of the most common cases among the AI automations that pay for themselves fast, precisely because the cost is permanent and consistently underestimated.

  • Direct keying time: a few minutes per document, which add up to days over a month.
  • Copy errors, paid for long after the moment they are made.
  • Delay: a batch processed on Friday is information that arrives several days late.
  • Human cost: qualified people tied up on a task no one wants to own.

Why did OCR and its rules always end up breaking?

The classic approach came in two steps: an OCR engine read the characters, then a layer of rules decided what to do with them. The invoice number is top right, the total is the last amount in the column, the due date comes right after its label. As long as the layout stayed put, it held.

The trouble is that layouts move all the time. Every supplier has its template, and every template has its exceptions. A new vendor, a credit note, a foreign quote, and the rule that worked yesterday returns a wrong value or nothing at all. You ended up writing one rule per special case, then a rule to catch the rule. In the end you spent more time patching the system than processing the documents it was meant to free you from.

What does a vision model read that OCR could not see?

A vision model does not just turn pixels into characters: it reads the page like a human, with its structure and meaning. It knows a total stays a total even if it has moved to another corner, that a due date stays a due date however it is phrased, that a table row belongs to its column. You no longer describe a position to it, you ask it for a piece of information.

In practice, that means a document it has never seen does not throw it off. Where OCR plus rules demanded a template per format, a vision model absorbs the variety without you maintaining hundreds of special cases.

  • A PDF scanned at an angle, stamped or annotated by hand.
  • An invoice from a supplier you had never received before, in a brand-new layout.
  • A clause buried in the middle of a twenty-page contract.
  • A table of line items to reconstruct, with its quantities and amounts.

Can you trust what a model extracts?

Not blindly, and it matters to say so. A vision model can read a 3 as an 8, or worse, return a plausible but wrong value with the same confidence as a correct one. It is the same mechanism that makes a language model hallucinate: it produces what looks like the right answer, not necessarily the true one.

The difference between a demo that impresses and a system in production is exactly what you put around the extraction. A confidence score on every field. Simple checks that catch the essentials: the sum of the lines must equal the total, an IBAN has a verifiable structure, a due date should fall after the invoice date. And a clear rule: below a certain confidence level, the document goes back to a human. You no longer check everything by hand, but you do not trust everything blindly either.

Does it pay off for you?

The starting math is simple: your volume of documents multiplied by the time actually spent on each. If the product runs into days per month, the question is no longer whether to automate, but when. Three triggers make it urgent.

  • A large volume handled by hand, where every document saved adds up quickly.
  • Electronic invoicing, mandatory since September 2026 and rolling out in phases, which pushes you to structure your flows anyway.
  • Sovereignty, when your documents are sensitive enough that their processing has to stay in-house.

One trap to watch, though: extracting is only half the job. A correct value stuck in a file has saved you nothing. The value appears when the extracted information lands directly in your accounting, your ERP or your line-of-business tool, without a human re-keying it a second time. That is where wiring extraction into your existing tools makes all the difference between a demo and a real gain.

Can your sensitive documents stay in-house?

Yes, and that is often what unblocks the project. Contracts, payslips, ID documents, client files: many companies hesitate to send that kind of material to an outside service whose location and reuse they do not control. The good news is that vision-model extraction does not require that trade-off.

Depending on your constraints, the processing runs on a European cloud or entirely within your own infrastructure, and your documents are never used to train a public model. That is the heart of what sovereign AI means in practice: the same extraction capability, without your data leaving the perimeter you have set.

OCR deciphered characters. A vision model reads a document. That is the whole difference.

The real shift is not that AI reads better: it is that you can finally automate document reading without building a rule factory that breaks at the first unexpected format. What remains is to know whether, in your case, the volume and the time spent justify it. Do you have a pile of documents still going through manual entry? Send me a few, anonymized, and in one call I will tell you what automating it would really save, with no commitment: let's talk.

Read next

Contact

Ready to go from demo to production?

Reply within 24 hours · first conversation free, no strings attached.