AI & Innovation

Private AI for work that cannot leave the building

Financial records are among the least appropriate data to hand to a public cloud model. Our flagship research program is an attempt to make that trade-off unnecessary.

In development — research & development initiative

The problem

The privacy problem is not a policy problem

Ask an accountant why they have not put AI to work on their document pile and the answer is rarely about capability. It is that using a cloud model means sending client financial records — payroll, banking detail, personal tax information — to a third party's servers to be processed.

For many businesses, and for professionals holding confidentiality obligations, that is not something a terms-of-service page can settle. The data protection people actually want is architectural: the records do not leave the premises at all.

That is the constraint we are designing to. Not "we handle your data carefully", but "the processing happens on your hardware, in your office".

Where the documents go

Cloud model Your office Third-party servers On-premises model Documents SLM Nothing leaves

The flagship initiative

The On-Premises SLM Program

SIF is developing compact, task-specific models trained on vetted, niche domain data, designed to run locally on client premises. The goal is explicitly not to train a general-purpose model, and none of this is generative AI — the models classify, extract and match; they do not chat or generate content. The method is teacher–student distillation: larger teacher models supervise training, and only the small student ships.

This is a research and development initiative and is labelled as one throughout this site. It is being pursued because it is the honest answer to a constraint our own clients have — not because it makes for a good slide.

  • Curated, vetted data

    The training material is narrow and checked: financial and tax domain documents selected and reviewed, rather than scraped at scale. A model that only has to be right about a small, well-defined domain does not need to be large.

  • Distilled, compact models

    Teacher models supervise training; only the compact student ships, sized for hardware a client can reasonably own or lease. Getting useful, auditable accuracy inside that constraint is the hard part, and it is the research question.

  • Local deployment

    The model is intended to be deployed on the client premises and to run there. No document leaves the building to be processed. Privacy comes from where the computation happens, not from a contractual promise about it.

Where this applies

Built for sectors where the data cannot leave

Accounting and taxation — the domain this company grew out of — is the first proving ground. But the constraint this program is designed around — processing that must happen inside the client's own boundary — shows up across regulated and confidentiality-bound sectors. These are the deployment targets the research is scoped against.

  • Finance & accounting

    First proving ground

    Client financial records, payroll and banking detail carry confidentiality obligations a cloud terms-of-service page cannot settle.

    Candidate tasks: Transaction categorization, receipt and invoice extraction, reconciliation matching.

  • Taxation

    Personal and corporate tax files are among the most sensitive documents a business holds.

    Candidate tasks: Document classification, deduction and write-off discovery support.

  • Banking & finance

    Regulated institutions face the strictest deployment bar of all — often a fully disconnected environment.

    Candidate tasks: Document processing and extraction inside the institution’s own boundary.

  • Legal

    Solicitor-client privilege makes third-party processing of matter documents a genuine problem, not a preference.

    Candidate tasks: Clause and entity extraction, document triage.

  • Healthcare

    Patient data is governed by health-privacy law; for many workloads it simply cannot leave the premises.

    Candidate tasks: Structured intake and records extraction from unstructured documents.

On-premises is not always air-gapped — and we treat the difference seriously

Some deployments run locally but stay connected — for example, ingesting live bank feeds while all model processing happens on the client's hardware. Others must be fully air-gapped: no external calls at all, with documents ingested in batch inside the boundary. These are different deployment profiles with different guarantees, and part of the research program is supporting both without blurring the line between them.

A research direction

Continuous ledger automation

The clearest application of the SLM program in our own domain is what we are calling continuous ledger automation: bank, card and payment feeds flowing into a ledger that categorizes, reconciles and stays audit-ready as transactions post — instead of being caught up at month-end. The compact models this program trains are the layer that does the categorization, extraction and matching, locally.

This is a direction under active research and scoping, not a product. The industry literature around this pattern quotes striking automation and accuracy figures; we will publish our own numbers when we have measured them, and not before.

First research project

Learning how banks describe transactions

Every institution encodes statement transactions in its own way — the same purchase reads differently across banks, and each bank's own categorization logic is opaque and inconsistent. The program's first project trains compact models on how different banks structure, describe and categorize transactions, and builds our own classification framework on top of that: raw statement lines in, standardized, tax-ready categories out.

The intended outputs are financials generated from statements that are ready to flow into tax filing, and a model that existing platforms can embed and leverage — running locally where the data demands it. Whether a compact model can normalize descriptors reliably across banks it has never seen is precisely the open research question.

A rule of this research: statement data used for training is de-identified, consented, or synthetic. The work needs banks' descriptor conventions — not anyone's actual transactions.

  • 01

    Autonomous categorization

    Language models matching transaction descriptors against vendor history and a chart of accounts; OCR pulling data from receipts and invoices into the ledger.

  • 02

    Real-time financials

    A continuously maintained P&L and cost tracking, replacing the month-end reconciliation rush with books that are already current.

  • 03

    Tax-aware bookkeeping

    Liability tracking and deduction discovery running against the ledger as it updates, rather than as a year-end scramble.

Roadmap

Research, then pilot, then clients

Three stages, in order, with no dates attached. Each one has to earn the next. We will update this page as stages actually complete rather than as they are scheduled.

  1. 01 Current stage

    Research

    Data curation methodology, model selection and fine-tuning experiments, and evaluation design — establishing whether a compact model can hit the reliability bar this domain requires, and documenting where it does not.

  2. 02 Planned

    Pilot

    Controlled trials on real document workloads, running alongside the existing manual process rather than replacing it, so output quality can be measured against a known-good baseline before anything depends on it.

  3. 03 Planned

    Client deployments

    On-premises deployment for clients whose data sensitivity makes cloud processing unacceptable, with the support and review process around it that professional financial work requires.

Plain terms

What we are not claiming

There is a great deal of AI marketing that describes intentions in the past tense. This section exists so there is no ambiguity about which parts of the above are real work in progress and which parts are finished.

  • We are not claiming a finished product. There is nothing here to buy yet.
  • We are not publishing benchmark results, accuracy figures or client outcomes, because the work has not produced results we would stand behind publicly.
  • We are not naming clients or presenting case studies for this program.
  • We are not committing to dates. Research either reaches the reliability bar or it does not, and we would rather say so than hold a schedule.

Engaging today

The models are in development. The discovery is not.

While the research runs, engagements start with use-case discovery: finding the narrow, high-volume, checkable tasks in your operation that a compact model could hold, and designing the pilot that would prove it against a baseline.

Have a document workload that cannot go to the cloud?

That is exactly the constraint this program exists for. We are interested in talking to businesses who have it, even while the research is still running.