Book a free consultation
What we do Who we areIndustries Case studiesPortfolioContactBook a free consultation

The right model for each task, not one model for everything.

We use Claude, GPT, Gemini and open models side by side, routing each request to the model that fits its difficulty, sensitivity and cost.

What it does for a business

Model routing

Simple classification goes to a small, cheap model; complex reasoning goes to a frontier model.

Sensitivity routing

Anything containing personal or confidential data stays on a private model.

Provider independence

Swap or add providers behind one interface, with no change to your application.

Spend control

Budgets and caps per team, per feature and per month.

Who it's for

  • Teams paying frontier prices for simple tasks
  • Businesses that need some work to stay on private models
  • Product teams that want to switch providers without a rewrite

How we keep it safe

Private and local first for critical and sensitive data. Partnered cloud models only when a task needs them, behind a policy gate, with audit logs and monitoring you can see.

Technology stack

What we use for LLMs & multi-model, and what each piece is for.

Offer

MVP in one week

A working AI product in your users' hands in seven days, built on our proven components.

Offer

Free MVP for startups

No-obligation free MVP for startups. Scope agreed in the free consultation; you keep it either way.

Offer

Free fix-up for AI-built apps

Built your app with AI coding tools but it isn't secure or won't scale? We review it free and fix the critical security and scaling issues free.

Frontier models
Cl

Claude (Anthropic)

Long documents, careful reasoning, drafting; via the Claude Partner Network

GP

GPT (OpenAI)

General reasoning, tools and structured output; OpenAI Select Partner

Ge

Gemini (Google)

Multimodal and long-context work

Open and private models
Ll

Llama

Strong general open model for private deployment

Mi

Mistral / Mixtral

Efficient open models for high-volume tasks

Qw

Qwen and others

Multilingual and specialised open models

Platforms
Az

Azure OpenAI

Enterprise hosting with regional data residency

AW

AWS Bedrock

Managed access to several model families in your AWS account

Ol

Ollama / vLLM

Serving open models on your own hardware

Control
AI

AI gateway

One endpoint: auth, policy, redaction, routing, logs

Ev

Evaluation harness

Golden task sets scored for accuracy, cost, latency

Bu

Budgets

Caps per team, feature and month

How we deliver it

Typical phases and timelines; your plan is agreed after the free consultation.

  1. 1

    Usage audit

    Where models are used today and what it costs

    1 week
  2. 2

    Routing design

    Tasks mapped to models, with evaluation sets

    1–2 weeks
  3. 3

    Gateway live

    Policy, redaction, routing and dashboards

    2–4 weeks
  4. 4

    Optimise

    Monthly review of cost, quality and new models

    Ongoing

Real examples

From products we built and run, and engagements we measured.

Real example

LexEdge, legal practices

Each matter chooses its model: on-device via Ollama for privileged work, Claude, OpenAI or Gemini where allowed, or an air-gapped Private AI Box.

See LexEdge
Real example

NeuralGate, AI governance

An OpenAI-compatible gateway tracking 31 models across 242 tenants, with budgets and routing per team.

See NeuralGate

Under the hood

The technical detail, for your engineers.

Gateway pattern

Applications call one OpenAI-compatible endpoint; the gateway authenticates, applies policy, redacts, routes and logs.

Routing signals

Task type, input length, detected sensitivity, required latency, and per-team budget.

Evaluation

Each route is tested on a golden set of your real tasks for accuracy, cost and latency before it goes live.

Providers we work with

Claude (Claude Partner Network), OpenAI (Select Partner), Gemini, Azure OpenAI, AWS Bedrock, and open models such as Llama and Mistral run privately.

Questions we're asked

Will switching models break our app?

No. Your app talks to one gateway endpoint; models change behind it.

Is our data used for training?

No. Enterprise agreements and private models only; nothing trains anyone else's model.

How much can routing save?

It depends on your task mix; we measure it on your real traffic before you commit.

Which model is best?

The one that passes your evaluation set at the lowest cost and risk; it's usually more than one.

Can we use our existing OpenAI or Azure account?

Yes. The gateway works with your own accounts and keys.

What about response speed?

Routing takes latency into account; fast models handle interactive tasks, larger ones handle background work.

How do you compare models fairly?

On a golden set of your real tasks, scored for accuracy, cost and latency, repeated whenever a new model arrives.

Can we see what each team spends?

Yes. Spend and usage are broken down by team, feature and model.

Is our data used to train AI models?

No. We use private models or enterprise agreements that forbid training on your data.

How do we get started?

Book the free two-hour consultation. We look at one real workflow and tell you whether this technology fits.

Is there a guarantee?

Yes. Engagements we take on carry our 10× productivity guarantee on the agreed workflow, or the fee comes back.

Products built with LLMs & multi-model

Other technology

Where would LLMs & multi-model help your business?

Two free hours on a real workflow. We'll tell you whether it fits, and what it would take.

Book the free consultation