Documents go in. Structured data comes out.
Somewhere in your business there is work that consists of reading a document and typing what it says somewhere else. We make that stop needing hands: we build the thing, wire it to the systems you already use, and keep it running every month.
A process, not a tool
We are not selling an application where someone uploads files and hopes for the best. What we build is a closed process: documents arrive the way they already arrive — a mailbox, a shared folder, a supplier portal — and what comes out goes straight to where the data already lives.
In between there is a language model reading the document, rules validating what it returned, and a decision about what to do when confidence is not high enough. That last part is the one that is almost never built, and it is what separates a project that lasts from one that dies after three weeks.
The documents that usually cause this work
This is not a closed list. It is what comes up most often, and it is here so you can recognise your own case.
-
Supplier invoices
Every supplier has their own layout and none of them matches the last. Out of this comes the number, the date, the total, the tax broken down by rate, and the line items when the detail is needed.
-
Delivery and dispatch notes
They arrive as scanned paper, often crooked and with stamps over the top. Out comes the reference, the quantities, and what was delivered against what was ordered.
-
Emails from suppliers and customers
Orders, confirmations, changes to an order — written by people, in no format at all. Out comes what matters, in fields, ready to go into a system.
-
Forms filled in by hand
Job sheets, service records, timesheets. The things someone types up into a spreadsheet today.
-
Bank statements and financial documents
To reconcile against what is already recorded, instead of someone working down two lists side by side with a finger on the screen.
-
Reports assembled every month
The same files, the same joining up, the same format at the end. It is the simplest case of all and the one most often left undone.
It goes where the data already lives
You do not have to change system, and we would rather you did not. The work is connecting to what is already there.
We write to the accounting software, to the ERP, to the database, to the shared spreadsheet that is in practice the company's system, or to a file in whatever format your accountant asks for. Where there is an API, we use the API; where there is not, there is almost always a file import nobody knew existed.
If there is no way at all to write into it, we say so during the pilot and we do not build on top of a hope. It is one of the things the first week is for.
From the first file to a process that runs on its own
Five stages, and none of them starts before the one before it is closed. The first is the pilot, and it is the only one you pay for with no commitment to anything after it.
-
Pilot
You send real files. We process them and hand back a prototype running on your own data, with numbers: how much it gets right, what it costs per document, how many hours it saves. One week. How the pilot works.
-
Quote for the build
With the pilot's numbers in hand, the cost of building it is stated in full before anything starts. Fixed scope, fixed date, fixed price.
-
Build
The real version: extraction, validation, what happens when the model does not know, and the connections to the systems you already use. You also get a dashboard showing what went through, what failed, and what is waiting for a person.
-
Going live
It runs alongside what is already there, on real documents, until the numbers line up. Only then does the manual work get switched off — and switching it off is your decision, not ours.
-
Operations
Every month we watch the quality, fix what breaks and adapt when the systems around it change. That is the monthly fee, and it is not optional: abandoned software does not stay as it was, it stops working.
Why most of these projects die at the demo
Building a demo with a language model is easy, which is why everyone has one. Getting it to run on real documents, every day, with nobody watching, is another thing entirely.
-
Measuring accuracy instead of assuming it
You build a set of documents with the right answer written out by hand, and you measure against it. Without that, “it seems to be working well” is an opinion — and it is the opinion of the person who built it.
-
Deciding where a person is needed
Not everything has to be right 100 % of the time, and almost nothing manages it. What decides is the cost of the error: a wrong invoice total costs money, a wrong notes field costs nothing. Where it costs, it goes past a person.
-
Knowing what to do when the model does not know
A model that invents an answer is worse than one that says it could not manage. The process has to have somewhere for the doubtful case to go — a review queue, an alert, a document set aside — and it cannot be silence.
-
Watching quality over time
A supplier changes their invoice layout, a format goes from PDF to a scan, the model underneath is updated. Accuracy falls without anyone touching anything, and someone has to notice before the client does.
That is the work we do. The artificial intelligence is the easy part — what costs is everything that has to be built around it before you can trust the result without someone checking it.
When this is not worth doing
An administrator costs a business somewhere between €28 and €34 an hour once employer contributions and cover are counted. That is the figure that decides everything else — not the wish to have artificial intelligence in the company.
It pays off when the process eats more than 40 hours a month, or when more than 200 documents a month pass through someone's hands. Past that, the software pays for itself inside the first year.
Below it, usually it does not. We say so in the first conversation rather than selling you a project that never pays for itself — and often that you need less than you thought.
Where the documents end up
Processing happens inside the European Union and your documents are not used to train anything. We sign a data processing agreement — what Article 28 of the GDPR requires — before the first file arrives, not afterwards.
If there is personal data in there, and there almost always is, that gets written down: which data, how long it is kept, and what happens when the contract ends. The paperwork is in English.
Which process eats the most hours?
Tell us which one it is and send half a dozen example files. In a week you have a prototype running on them and the numbers that say whether it is worth it.