Document & intake automation
Invoices, purchase orders, claims, applications and intake forms — read, validated against your business rules, written into your system of record, with the awkward cases routed to a human. The highest-volume genuine automation in the mid-market.
Two people.
One keyboard's worth of work.
Most mid-market companies have a desk where documents arrive and a person retypes them into a system. Accounts payable, claims intake, order entry, patient or client onboarding — the shape is always the same: read a document, decide what it means, put the fields somewhere, escalate the odd one.
At 5,000 documents a month and four minutes each, that is roughly 333 hours — about CAD $12,000–13,300 a month in fully-loaded clerk time. It is the single most common process where automation pays back inside a year.
It is also the process most often automated badly. A system that extracts fields but cannot say how often it is wrong just moves the work into a review queue nobody measures.
What we actually build
Five stages, all of them yours at handover — source, infrastructure definitions and runbook included.
Ingestion
Documents arrive from wherever they already arrive: a shared mailbox, an SFTP drop, a scanner output folder, a web upload, or an API. We connect to the existing path rather than asking your senders to change.
Extraction
Structured fields pulled from PDFs, scans, spreadsheets and images, with schema validation on every output so a malformed response is caught rather than written. Handles multi-page documents and mixed formats.
Business-rule validation
Your rules, in deterministic code — not left to the model. Totals reconcile, tax is checked, the PO exists, the supplier is known, the date is in range, duplicates are caught. This layer is where correctness actually lives.
Write to the system of record
Posted into your accounting, ERP or line-of-business system via its API — or produced as a validated import file where no API exists. Every write is logged with the inputs that produced it.
Human exception queue
A small review interface for the cases the rules flagged: low confidence, failed validation, unknown supplier, or anything above a threshold you set. Approve, correct, or reject — and corrections feed the next evaluation round.
Fixed price. Written exclusions.
Prices in Canadian dollars, verified 23 August 2026. An advisory hour is CAD $195; fixed-scope advisory engagements start at CAD $4,500; minimum build engagement is CAD $9,500.
Single document type
$26,000CAD · One document type, two intake sourcesIncluded
- All five stages above, in production
- Held-out labelled evaluation set
- Exception review interface
- Per-document cost accounting
- Runbook and full source handover
Not included
- Additional document types
- Ongoing hosting — see run tiers
Multi-type & multi-source
$48,000CAD · Three document types, multiple sources, full review queueIncluded
- Everything in the single-type tier
- Three document types with separate rule sets
- Multiple intake sources, routed and deduplicated
- Role-based review queue with audit trail
- Per-stratum accuracy reporting
Not included
- Changes to your upstream systems
- Data migration or historical backfill
Hosting & care
$2,400 / $4,800CAD / month · Standard · ExtendedIncluded
- Hosting in a Canadian or U.S. region, or your own infrastructure
- Monitoring, alerting and incident response
- Monthly re-run of your evaluation set
- Model and cost optimisation as pricing shifts
Not included
- LLM inference — billed at cost, or run on your own API account
- New document types or rule changes
Accuracy is a cost line, not a slogan
On 5,000 documents a month, the difference between 80% and 95% accuracy is roughly CAD $1,950 a month of review labour — several times what the model API costs. That is why we measure it properly instead of quoting a vendor number.
Before we build, you help label a held-out set of real documents from your own operation, including the awkward ones. The system never sees them during development. Acceptance is scored against that set, at a threshold agreed in the contract.
What acceptance looks like
The set size and pass threshold are written into the statement of work before development starts. If the threshold is not met at acceptance, you choose: pay 60% and keep the code, or we run one full remediation cycle at no charge.
Accuracy is reported per document category, not as one blended number — it is entirely normal to see 99% on your dozen regular suppliers and materially less on the long tail. That split tells you where the human queue belongs.
Under roughly 500 documents a month. The manual baseline is around CAD $1,200/month, which will not repay a build in the tens of thousands before the process changes. Measure your real volume first — most people overestimate it.
When every document is a special case. Automation pays on a dominant pattern with a manageable tail. Take 50 recent documents; if you get 50 different stories, it is judgement work, not a process, and an agent will produce a queue that costs more than the original task.
When the upstream process changes monthly. You would be buying a maintenance obligation, not an asset. Stabilise and document it first.
Where these extraction patterns came from
Structured extraction from long technical documents, plus citation-grounded retrieval over a regulatory corpus. The schema validation, exception-queue and per-document cost accounting described above are the mechanisms we built here first and then productised.
Our own product, not a client engagement. We have no client case studies yet and will not invent any.
Questions worth asking first
Each answer restates what is set out above — how documents come in, how accuracy is judged, and where the system runs.
How do you know it is accurate enough before we accept it?
Before we build, you help label a held-out set of real documents from your own operation, including the awkward ones. The system never sees them during development. The set size and pass threshold are written into the statement of work before development starts, and acceptance is scored against that set.
What happens if the threshold is not met at acceptance?
You choose: pay 60% and keep the code, or we run one full remediation cycle at no charge. Accuracy is also reported per document category rather than as one blended number — it is entirely normal to see 99% on your dozen regular suppliers and materially less on the long tail, and that split tells you where the human queue belongs.
Do our senders have to change how they send documents?
No. Documents arrive from wherever they already arrive: a shared mailbox, an SFTP drop, a scanner output folder, a web upload, or an API. We connect to the existing path rather than asking your senders to change.
What if our system of record has no API?
Then we produce a validated import file instead. Where an API does exist, results are posted into your accounting, ERP or line-of-business system through it. Either way, every write is logged with the inputs that produced it.
Where does it run, and where do the documents live?
In a Canadian or U.S. region, or on your own infrastructure. The Canadian regions are AWS ca-central-1 (Montreal) or Azure Canada Central (Toronto); U.S. work runs on AWS or Azure in a U.S. region. We can also build inside your own cloud tenancy, where you hold the keys, the logs and the bill, or fully on-premise on your own hardware. The region is selected with you and named in the contract.
Start with the numbers
The readiness assessment measures your real volume, your current cost per document, and what accuracy is achievable on your actual paperwork — before you commit to a build.
See the assessment