AI for back office and operations
Automating an administrative process is easy until the odd document arrives. We design for the exception from the start: the system recognises what it cannot process, escalates it with context, and that rate gets measured.
Key takeaways
Administrative automation does not fail on the normal case: it fails on the odd document, and that is where the project either works or creates new work.
The right metric is not "what share got automated" but what share of exceptions was detected and escalated. A flow that silently mis-processes 3% is worse than one that escalates 15%.
Document extraction gets measured field by field against a hand-validated batch, not overall. The average hides that the field that matters is the one failing most.
It pays where there is volume and stable rules. Where judgement changes case by case, automation creates more review than it removes.
The five processes that hold up
All five share a shape: high volume, stable rules and an outcome verifiable against a system.
- Supplier invoices, purchase orders, delivery notes, receipts. They arrive as PDFs, photos or emails and somebody retypes them.
- Issue, match against payment received, flag the difference. Mechanical, repeated and verifiable.
- Approvals, letters, access, small purchases. High volume, documented answers, handled by email today.
- What gets approved in one place has to be recorded in another, and today a person moves it by copying.
- The same report every week from three systems. That is collection work, not analysis.
The exception is not an edge case: it is the design
An automated administrative flow works on 85% of documents from week one. The project is decided by the remaining 15%: the invoice with a different format, the supplier who changed their layout, the crooked scan, the case the rule did not anticipate. A system that processes those like the normal ones introduces silent errors into accounting systems, and the cost of finding them later far exceeds the saving of having automated them.
So the design starts there: the system has to recognise when it is not sure — low confidence on a field, a total that does not match the line items, an unknown supplier — and escalate with the document and the reason, instead of filling the gap with the most likely value.
And extraction gets measured field by field against a hand-validated batch. The overall average is useless: a system can be 96% accurate overall and be failing on the tax ID or the total, which are exactly the two fields where an error costs money. It is reported per field, and per field you decide what gets automated and what gets reviewed.
The numbers behind an administrative flow
| What is measured | What it means | Why the other one is not enough |
|---|---|---|
| Per-field accuracy | Against a hand-validated batch, field by field: tax ID, date, total, line items. | Against overall accuracy, which hides that the expensive field is the one failing most. |
| Exception detection rate | Of the documents the system should not have processed, how many it actually escalated. | Against "percentage automated", which goes up precisely when the system stops escalating. |
| Cost of residual error | What happens to what got processed wrong and went undetected, and what it costs to find later. | It is the number that decides whether the process gets fully automated or keeps review. |
| Cycle time | From document arriving to record landing in the destination system. | It is the real saving, and it is what gets compared against the cost of running the flow. |
If the process changes every month or the judgement depends on who is looking, automating it produces a flow that has to be rebuilt constantly. Stabilise the process first, automate second, and that order is not negotiable.
And if volume is low, the saving does not pay for the build or the maintenance. We say so on the first call with the numbers on the table, even when it means a smaller project or none.
Related services
AI automation
Applied where there is volume and stable rules, not where there is expectation.
View serviceIntegrations and APIs
MCP servers against the systems already in operation.
View serviceAI agents
LangGraph orchestration, durable state and human approval on steps with consequences.
View serviceSoftware development
The system around the model, not just the model.
View serviceMore from the blog
Frequently asked questions
What happens with documents the system does not understand?
It escalates them, and that is the design rather than a failure. The system recognises signals that it is not sure — low confidence on a field, a total that does not match the line items, a supplier it has never seen — and sends the document to a person with the reason. The metric we report is what share of exceptions it detected, because a flow that silently mis-processes 3% costs more than one that escalates 15%.
Does it work with scans or phone photos?
It does, and the cost has to be stated. Image quality caps everything downstream: a crooked photo with a shadow drops extraction accuracy below acceptable on the numeric fields. We measure batch quality before promising anything and report what share falls below the threshold. Sometimes the right conclusion is to change how documents arrive before automating how they are processed.
Does it integrate with our ERP?
Yes, over API, an MCP server or whatever import mechanisms it exposes, against SAP, Dynamics or the system you run. When the ERP does not allow automated writes, the flow goes as far as leaving the record ready to approve, which already removes most of the manual work and keeps control where the team wants it.
Is this RPA or AI?
It is the combination, and the difference matters. RPA works when the process is deterministic: if the document always arrives the same way, you do not need a model. AI comes in when the document varies — every supplier invoices differently — and you have to interpret rather than read fixed positions. A good flow uses rules where rules suffice, because they are cheaper and more predictable, and a model only where it is needed.
How many people does it take to maintain?
Fewer than run it today, but not zero, and that is worth saying before signing. An administrative flow needs somebody to review the exceptions and flag when a business rule changes. We hand over the flow documented and with dashboards so the internal team can do that follow-up, and maintenance with us is optional and quoted separately.
How long does it take and how is it priced?
Two to six weeks per flow depending on complexity and integrations, with something running from week one. It is quoted with fixed scope, price and date in a proposal 48 hours after the first call. We recommend starting with one process, measuring it, and choosing the next with that number in hand rather than committing five at once.
Book 15 minutes
Tell us what you are building, or what stopped working. You leave the call with a concrete answer: it can be fixed, it can be built, or it isn't worth it.