Private & on-premise LLM
Open-weight models running on your own hardware or in a private cloud tenancy you control — for work where no data may leave your walls, and for nothing else.
Buy this for
the boundary.
There is exactly one good reason to run your own model: a requirement — legal, contractual, or from a customer's security review — that certain data never crosses into a third party's infrastructure. Clinical records under a hospital's own policy. Defence or classified-adjacent work. A privacy commitment you have already made in writing to your own customers.
If that requirement exists, this is the only architecture that satisfies it, and the cost is simply what compliance costs.
If it does not exist, self-hosting is usually the more expensive option — and we will show you the arithmetic rather than sell you the more impressive-sounding build.
Where the crossover actually is
Self-hosting replaces a variable per-token bill with a fixed monthly one: hardware or reserved GPU capacity, power and cooling, and the engineering time to keep it running. That fixed cost does not fall when your usage does.
On our own numbers, the crossover sits somewhere around 100 million tokens a month. Below that, a commercial API deployed in your own region — with the residency terms in the contract — is usually cheaper and always less work. Above it, and especially with steady round-the-clock load, owning the capacity starts to win.
Your crossover is not ours. It moves with your utilisation, your power cost, and whether you already have datacentre space. We calculate it with your numbers before quoting a build.
Where your data actually goesThe capability trade, stated plainly
Open-weight models have closed much of the gap, and for structured extraction, classification, summarisation and retrieval-grounded answering they are frequently indistinguishable in production.
On the hardest reasoning, long-context and coding work, the leading commercial models are still ahead. If your workload lives there, self-hosting costs you capability as well as money — and that trade should be a decision, not a surprise.
Benchmark positions move every few months. We re-test against your actual tasks at deployment rather than quoting a leaderboard.
What the deployment includes
A serving platform your own team can operate — not a model file on a server.
Model selection
Candidate open-weight models tested against your tasks on a labelled sample, at the quantisation you will actually run. Selection is evidence from your data, not a benchmark table.
Serving stack
An inference server with batching, quantisation and memory management tuned to your hardware, behind an OpenAI-compatible API so your applications need no special client.
Capacity sizing
Measured throughput and latency at your expected concurrency — before hardware is purchased. The most expensive mistake in this work is buying the wrong GPU.
Network isolation
Deployed inside your boundary with egress rules that make the isolation demonstrable, not merely intended. Documented well enough to hand to an auditor.
Observability
Utilisation, latency percentiles, queue depth and per-team usage accounting — so you can see whether the capacity you bought is the capacity you needed.
Upgrade path
A documented, rehearsed procedure for swapping in a newer model, because a materially better open-weight release lands every few months and you should be able to take it.
Build, run, and the part we don't mark up
Prices in Canadian dollars, verified 23 August 2026.
Deployment
from $34,000CAD · Serving platform, sized, tested and documentedIncluded
- Model selection tested on your data
- Serving stack with OpenAI-compatible API
- Capacity sizing before hardware purchase
- Network isolation and auditor documentation
- Operations runbook and team handover
Not included
- Hardware, colocation and network
- Fine-tuning or continued pre-training
Managed operation
from $3,200CAD / month · Or operate it yourself with the runbookIncluded
- Monitoring, alerting and incident response
- Model and serving-stack upgrades, tested first
- Capacity review as your load changes
- Quarterly re-benchmark against commercial APIs
Not included
- Hardware replacement or warranty
- Datacentre and facilities management
Hardware
At costClient-procured · No margin, no referral feeHow it works
- We specify exactly what to buy, and why
- You purchase it directly and own it outright
- Private-cloud GPU capacity is an alternative
- Sizing is done before you spend anything
Why
- We take no margin on hardware, so the recommendation is not a sales incentive
When you should not buy this
When "we want control" is the whole requirement. Control is a feeling; a residency clause is a contract. Most commercial providers will now commit in writing to a processing region and to not training on your data. If that satisfies your obligation, take it — it is cheaper and your team does not become a GPU operations team.
Below roughly 100 million tokens a month. The fixed cost does not amortise. We will run your numbers and tell you if you are on the wrong side of the line.
When nobody will own it after we leave. A self-hosted model is infrastructure. Somebody has to patch it, monitor it and eventually replace the hardware. If that person does not exist and you are not buying managed operation, the deployment will quietly rot.
When you need the frontier. If your workload genuinely requires the strongest available reasoning, an open-weight deployment will disappoint you, and no amount of tuning closes that gap.
Questions worth asking first
These are the ones that decide it. We will show you the arithmetic rather than sell you the more impressive-sounding build.
When is running our own model actually the right call?
There is exactly one good reason: a requirement — legal, contractual, or from a customer's security review — that certain data never crosses into a third party's infrastructure. If that requirement exists, this is the only architecture that satisfies it, and the cost is simply what compliance costs. If it does not exist, self-hosting is usually the more expensive option, and we will show you the arithmetic rather than sell you the more impressive-sounding build.
Where does the cost crossover sit?
On our own numbers, somewhere around 100 million tokens a month. Below that, a commercial API deployed in your own region — with the residency terms in the contract — is usually cheaper and always less work. Above it, and especially with steady round-the-clock load, owning the capacity starts to win. Your crossover is not ours: it moves with your utilisation, your power cost, and whether you already have datacentre space, so we calculate it with your numbers before quoting a build.
Will an open-weight model be as capable as a commercial API?
For structured extraction, classification, summarisation and retrieval-grounded answering, open-weight models are frequently indistinguishable in production. On the hardest reasoning, long-context and coding work, the leading commercial models are still ahead — if your workload lives there, self-hosting costs you capability as well as money, and that trade should be a decision, not a surprise. Benchmark positions move every few months, so we re-test against your actual tasks at deployment rather than quoting a leaderboard.
Do we buy the hardware through you?
No — hardware is a pass-through, client-procured at cost. We specify exactly what to buy and why, you purchase it directly and own it outright, and we take no margin and no referral fee, so the recommendation is not a sales incentive. Private-cloud GPU capacity is an alternative, and sizing is done before you spend anything.
Are you SOC 2 or ISO 27001 certified?
No. Quintessentia Network Inc. is not SOC 2 certified and is not ISO 27001 certified — we are a small, owner-operated Canadian firm, and if your procurement process requires either certification from every vendor, we will not pass it today. What we offer instead is documented practice, data residency in the country you choose, contractual commitments, and the option to run everything inside your own infrastructure so the trust boundary never leaves your organisation. It is all set out on the trust page.
Get the crossover calculated
Bring your expected volume, your residency requirement and your existing infrastructure. We will produce the arithmetic — including the case for not doing this.
Book a scoping call