Document Processing
Automation
Documents arrive in different formats and several teams need different parts of them. We build a controlled pipeline that extracts the data once, validates it and records every exception before posting to company systems.
Read the document
Extract the facts
Send for review
You get checked data with the source attached.
Checked
EXTRACTION AGAINST SOURCE DOCUMENTS
Measured
COST INCLUDING HUMAN REVIEW
One document type
IN THE FIRST PHASE
Document automation needs more than OCR
OCR, document-processing platforms and vision models offer different ways to extract information. The useful comparison is how accurately each handles your files and how much review remains.
We test mixed document packs, tables, poor scans and unusual layouts alongside routine examples. Extracted values then need business-rule checks before they reach another system.
The first pilot measures extraction errors, review effort and the cost of processing a complete record.
DOCUMENTS TO TEST
- Changing supplier templates
- Fields that appear reliable but are wrong
- Tables, stamps and handwriting
- Records requiring human approval
- Mixed packs and missing pages
VLM-BASED PIPELINE
- One pipeline, tested layouts, fewer templates
- Field-level checks before anything posts
- Tables, signatures and stamps checked
- Review based on agreed risk and quality rules
- Document-level cost and accuracy tracking
How a document pipeline works
An invoice, referral letter and lease need different fields and review rules. The pipeline gives each document a defined route from receipt to the target system.
Intake
We assess the inboxes, shared folders and upload routes already in use, then connect the agreed channels with access and file-handling controls.
Read
Extract the required totals, dates, parties and line items into a defined schema. Where confidence scores are available, we test their relationship to actual errors before using them for routing.
Check
Totals add up. VAT number checks pass. Supplier exists. PO is open. Date is in range. The rules your team already runs in their head, written down once.
Review
Records can post automatically only under agreed rules. Exceptions and higher-risk documents enter a review queue with the source and extraction side by side.
Post
Approved records pass through supported interfaces into the systems in scope. The audit history records the source, checks, approval and posting result.
Documents we process
Anything that arrives as a PDF, scan, photo or email attachment and ends up in someone's inbox to be typed up. These are common starting points.
Supplier invoices & credit notes
PO matching, VAT validation, GL coding, posting to Xero or Sage. Usually the first finance workflow worth inspecting.
Contracts & MSAs
Parties, term, renewal dates, liability caps, payment terms. Pulled into a register your legal and ops teams can search.
KYC, KYB & right-to-work
Passports, driving licences, share registers, Companies House filings. Captured, checked and retained against your policy and sector rules.
Insurance claims & FNOL
First notice of loss forms, repair quotes, medical reports, photos. Triaged, summarised, attached to the right claim file.
Customs, freight & PODs
Commercial invoices, packing lists, C88 / CDS entries, proof of delivery. Tied back to shipments in your WMS.
Clinical letters & referrals
Clinical documents need a separately agreed review process and appropriate clinical safety assessment. Extracted information must remain traceable to the source for the responsible clinician.
How we take it into live use
We assess the document flow, prove one class against representative examples and connect the approved result to the systems that need it.
You own the code, the prompts, the model choice and the data. We can host it on your stack or ours. If your AP team is on Xero today, they're still on Xero tomorrow.
BOOK A DOCUMENT AUDITDocument audit
We review a representative set of the documents causing the most work. The output records who touches them, where they go, which fields are retyped, the exceptions and the systems that need updating.
Pick the model, write the schema
We choose the model that wins on your document type (Claude, Gemini, GPT or open-weight), define the fields you need, and write the rules a reviewer would use to accept or reject. You see real extractions on your real documents before we wire anything up.
Launch the first pipeline
Intake, read, check, review queue, post. Running on your real volume, with your team in the loop. We run it alongside the manual process until the numbers match.
Watch, learn, add the next one
We turn useful corrections into test cases and assess proposed prompt or rule changes before release. Measurements include field errors, review time and cost per document.
Audit and validation controls
Documents may contain identity details, bank information or medical records. We identify the data, permitted uses and access requirements before choosing a processing service.
The design covers provider terms, hosting and transfer routes, access, retention and deletion. Logs retain the evidence needed to investigate an extraction or support a data request.
UK GDPR & DPA 2018
DPIAs where processing is likely to be high risk, lawful basis recorded per source, and Articles 22A to 22D safeguards where solely automated decisions have legal or similarly significant effects. We follow the ICO's AI and data protection guidance.
Making Tax Digital
MTD for Income Tax started on 6 April 2026 for eligible sole traders and landlords with qualifying income over £50,000. Document intake can help prepare digital records for compatible software.
Consumer Duty & accounts rules
For financial services firms and law firms: review queues for customer-facing decisions, evidence for Consumer Duty outcome monitoring, and records that match the SRA's accounts-rule expectation for accurate chronological records.
Signatures & seals
Detecting a signature image is different from validating an electronic signature. Where verification is needed, we scope the signature format, certificate checks and evidence required with your legal adviser.
Document processing already sits inside our client systems
Leo Associates extracts supplier quotes into project and finance records. Crystal keeps compliance findings and evidence with each deal. The extraction only matters because the next team receives controlled data.
When this is worth discussing
We work best when there is a real operating problem, enough volume to measure and people from the affected teams who can make decisions.
Usually a good fit
- An established UK business, usually with annual revenue above £10m
- A repeated process with a known cost, delay, error rate or capacity problem
- A senior sponsor and a day-to-day owner who understand the work
- Access to the relevant staff, systems, sample records and security requirements
We may point you elsewhere
- A standard product already covers the process well
- The requirement is a one-off small build with no wider operating case
- There is no owner or access to the people and data needed to test the result
- The plan relies on AI making high-impact decisions with nobody responsible for review
Questions before connecting the systems
What about hallucinations? I can't have made-up totals on an invoice.
We compare extracted values with source documents and apply rules such as reconciled totals, supplier matching and purchase-order checks. Confidence scores need calibration on your files. These controls reduce errors but do not guarantee every incorrect value will be detected.
Would Xero, Dext, Hubdoc, Rossum or Klippa cover this?
For standard supplier invoices entering a supported accounting package, an established product may be enough. Custom work is justified when document types, required fields, review rules or downstream systems fall outside the product's supported flow.
Where does the document data go? Who sees it?
You decide, document type by document type. Sensitive ones run through model endpoints in the UK or EU under a data processing agreement, or against a self-hosted open-weight model. Less sensitive ones can use a hosted frontier model under agreed processing terms. We write it into the DPIA before we touch your data.
How long until it's live and earning back?
Timing depends on document variation, validation rules and downstream systems. The pipeline runs alongside the manual process until the operational owner accepts the extraction and exception results.
What does it cost per document?
We benchmark your own documents, including model calls, storage, review time and failed processing. The business case compares that total with your current cost, rather than assuming a published industry average applies.
What about handwriting? Stamps? Foreign-language invoices?
Results depend on the language, format and scan quality. We benchmark each relevant category and agree when documents require review before posting.
Does my team need to use a new app?
We aim to retain existing intake channels and destination systems where they support the required controls. Reviewers may need a new screen for comparing the extraction with its source. Training and process changes are included in the scope.
Who owns it once it's built?
The commercial proposal states ownership of source code, prompts, evaluations and extracted data. Model, document-processing, storage and support costs are shown separately.
Talk to us about the documents
Send us representative documents, volumes, exceptions and the systems that receive the data. We will assess which document class has the strongest case for automation and what a safe pilot requires.