Private & Local LLMs

Private LLM deployment: local AI that never leaves your business.

Open-weight models have closed most of the gap for the work a business actually does — drafting, summarising, answering questions over its own documents — and they run on hardware you own. We size the machine, load your files, build the evaluation, and keep it current. Confidentiality becomes a property of the architecture, not a clause in a vendor contract.

Book an AI Opportunity Sprint
Does this sound familiar?
  • A policy that forbids pasting client files into a public chatbot, and staff doing it anyway.
  • A data-processing agreement, a regulator, or a client contract that rules out a cloud model outright.
  • A per-token bill that grows with every document you would like the model to read.
  • A local LLM someone installed on a laptop that answers well in a demo and has no owner, no evaluation, and no backup.
What’s included

What’s included

  • Deployment design

    Local workstation, on-premise server, or a private cloud tenancy — chosen against your data-residency rules, load, and budget.

  • Model selection

    The open-weight model — Llama, Qwen, Gemma, Mistral or another — that clears your quality bar at a size your hardware can run.

  • Hardware sizing

    GPU memory, quantisation, and concurrency worked out before purchase, so the machine is neither undersized nor idle.

  • Retrieval over your documents

    Answers grounded in your own files, citing the source, with access control that follows the document rather than the login.

  • Evaluation harness

    A test set drawn from real questions, so a model update or a new document source is checked before it goes live.

  • Maintenance and updates

    Model upgrades, security patches, new sources, and monitoring on a schedule — an AI department without the headcount.

How we work

How we work

  1. Map~1 week · fixed fee

    AI Opportunity Sprint

    Find where AI creates real value — and what it’s worth.

    • Workshop with your team
    • Prioritized use-case backlog
    • ROI model for the top opportunities
  2. Prove3–4 weeks

    Production PoC

    Build one use case for real — and measure it.

    • One use case built production-grade
    • Evaluated against a real KPI
    • Evidence-based go / no-go
  3. Ship6–12 weeks

    Build & Integrate

    Put the system into production — integrated and compliant.

    • Full build, integrated with your stack
    • Security & compliance built in
    • Live, monitored, documented
  4. ScaleOngoing

    Scale & Enable

    Expand across workflows — and level up your team.

    • Rollout across teams and workflows
    • Evals & monitoring in place
    • Team enablement + fractional AI leadership

Open weights made local viable; quantisation made it affordable

Two years ago a local LLM was a compromise: a small open model on a workstation, giving answers a cloud model would have got right. That gap has narrowed for the tasks most businesses have — summarising, drafting, extracting fields, answering questions over internal documents. Open-weight families such as Llama, Qwen, Gemma and Mistral publish models from a few billion parameters up to frontier scale, and the mid-sized ones are the ones that matter here.

Quantisation is what makes the economics work. A model in the tens of billions of parameters, quantised to four bits, fits in the memory of one workstation-class GPU; the largest open models fit on a small server with a handful of them. That is a one-off purchase, not a per-token bill that scales with every document the model reads.

The trade is real and worth stating plainly. A self-hosted model is a step behind the largest cloud models on open-ended reasoning, and it does not improve on its own. What it offers instead is a fixed cost, a known location for every byte, and a model that does not change under you between Tuesday and Wednesday.

Confidentiality by architecture, not by policy

A policy says staff must not paste client files into a public chatbot. An architecture makes it impossible for the file to leave. The difference shows when a client, a regulator or an auditor asks where the data went. “Nowhere — here is the network diagram” is an answer no data-processing agreement can give.

Local does not automatically mean private. The model server, the retrieval index, the embedding step and the chat interface each hold a copy of your documents, and each needs the same access control the file server has. A retrieval layer that can surface a document to someone who was never allowed to read it is a breach with a friendlier interface. We treat the permission model as part of the deployment, not a hardening step for later.

Air-gapped deployment is the strict case. No outbound network from the inference host means updates arrive on media, telemetry is off, and the evaluation set has to travel with the model. It costs more to operate. For a firm whose business is other people’s secrets, it is the honest version of the promise.

A model nobody maintains is a demo

The failure mode of local AI in a business is rarely the model. It is the laptop it lives on. Someone technical installs it, it answers well, three people come to rely on it, and then a model update changes its behaviour, the index goes stale, or the person leaves. There is no test that says whether it still works, because there never was one.

The maintenance is the service. Model upgrades are run against the same question set before they replace the old version. New document sources are added with their permissions. The index is rebuilt on a schedule. Someone is on the hook when an answer is wrong. That is what an internal AI department would do, and what most firms cannot justify hiring for.

Before buying hardware, write down the fifty questions the system must answer well and who is allowed to see each answer. If that list cannot be written, the deployment is not ready, whatever the hardware.

Questions
Is a local LLM good enough compared with a cloud model?
For summarising, drafting, extraction and answering questions over your own documents, a well-chosen open-weight model is usually good enough, and grounding in your documents matters more to answer quality than the last few points of model capability. For open-ended reasoning at the frontier, cloud models still lead. The evaluation harness settles the question for your workload rather than in general.
What hardware does a private LLM need?
It depends on the model size and how many people ask at once. A mid-sized model serves a small team from one workstation-class GPU; larger models or heavier concurrent use call for a server with several. We size it from your question set and usage pattern before anything is bought.
Can it run entirely offline?
Yes. An air-gapped deployment has no outbound network from the inference host, with model updates delivered on media and telemetry off. It costs more to operate, and we recommend it only where the confidentiality requirement genuinely calls for it.
Which businesses is private AI for?
Any firm whose work is other people’s confidential information — law firms, accountancies, clinics, financial advisers, engineering firms holding client IP — and any company whose data-processing agreements or regulator rule out a cloud model. Size matters less than the confidentiality requirement.
Who owns the system afterwards?
You do. The hardware, the model weights, the index and the evaluation set are yours, and nothing routes through us. The maintenance contract keeps it current; it is not a dependency you cannot leave.
Do we still need the AI Opportunity Sprint?
If the use case is already clear — private question-answering over case files, say — we go straight to a scoped Production PoC on your hardware. The Sprint is for teams that also want to know which other workflows a local model could carry.
Ready to start?

Start with an Opportunity Sprint.

Book an AI Opportunity Sprint