Start a conversation
AI and machine learning

Models trained on your data, evaluated honestly, shipped into something someone uses.

A model that lives in a notebook has not helped anyone. We take AI work from the decision it is meant to improve through to production, behind access controls, with monitoring and an explanation a person can inspect.

9AI systems delivered
4modalities: text, speech, vision, sensor
100%of judgements traceable to evidence
live
What we build

Six areas, and exactly what is included in each.

Every bullet below is work we have shipped. Where a line looks like filler on another company's site, it is here because a client needed it and we built it.

Model training and fine-tuning

Custom models trained on your own data, from dataset construction and labelling through training, evaluation and deployment — including fine-tuning open models where a general one will not do.

What this includes
  • Dataset construction, cleaning and labelling strategy
  • Baseline first, always, so you know what sophistication bought
  • Held-out evaluation with the metric that matches your decision
  • Fine-tuning open weights where a general model underperforms

Prediction and risk scoring

Classification, ranking, forecasting and risk scoring.

What this includes
  • Feature engineering from your operational data
  • Calibrated probabilities, not just labels
  • Threshold selection tied to the cost of each error type
  • Tutor, agent or operator-facing dashboards

Computer vision

Detection, segmentation and automated quality inspection, including on-device inference for places with no reliable connection.

What this includes
  • Instance segmentation with Mask R-CNN and successors
  • On-device inference through TensorFlow Lite
  • Annotation strategy and inter-annotator agreement
  • Confidence handling and human review routing

LLM applications and agents

Retrieval-augmented assistants, document understanding and task agents grounded in your own content, so an answer can be checked against a source rather than trusted on tone.

What this includes
  • Retrieval grounded in your documents, with citations
  • Local models through Ollama where data cannot leave
  • Evaluation harnesses that catch hallucination and drift
  • Tool use and agent orchestration with guardrails

Data engineering

The pipelines underneath.

What this includes
  • ETL pipelines with reconciliation and alerting
  • Schema design for analytics as well as transactions
  • Dashboards built for a decision, not for decoration
  • Data quality monitoring

Explainability and governance

Where a model touches a decision about a person, the explanation and the contest route are part of the deliverable, not a later phase that never gets funded.

What this includes
  • Attribution that reflects the computation, not a plausible story
  • A defined route for a person to challenge an outcome
  • Model cards and decision documentation
  • Bias assessment on the axes that matter for your use
What you receive

The deliverables, listed before you ask.

The same artefacts every time, regardless of project size.

  • Baseline model with measured performance
  • Trained model with held-out evaluation report
  • Inference service with monitoring and alerting
  • Drift detection and retraining schedule
  • Model card and decision documentation
  • Explanation surface for the people affected
Stack

What we build this with.

Deliberately narrow. A team our size claiming twenty frameworks is claiming one thing badly, twenty times over.

PythonPyTorchTensorFlowTensorFlow Litescikit-learnMask R-CNNEfficientNetLSTM and attentionRAGOllamaOpenAI-compatible APIsPandasFastAPI
Proof

Three things we have already shipped in this practice.

Open the work page for the full case note on any of these, including what did not go to plan.

Delivered

Dropout risk agent

50% less manual monitoring and 35% more consistent identification.

Delivered

Scratch and dent detection

Pixel-level instance segmentation separating scratch from dent on the CarDD benchmark.

Delivered

Multimodal emotion engine

Speech and text fused at decision level, beating both single-modality branches.

9AI systems delivered
4modalities: text, speech, vision, sensor
100%of judgements traceable to evidence
22systems shipped across the company
Questions

What clients ask about this practice.

Answered as we would answer them on a call, including the ones with awkward answers.

Usually less than you fear and more than you think, because operational systems hold more than people remember. The data audit in week two answers this properly, and it is the cheapest way to find out.

No. Client data trains client models. If a project benefits from a shared base model, that base is a public or licensed one, never another client's data.

It will. Performance decays as the world moves away from the training data. We build drift monitoring and a retraining schedule into the original scope so this is routine rather than a crisis.

Yes. We work with local models through Ollama and self-hosted inference where data cannot leave your estate.

Other practices

Most projects use more than one.

The reason a product is not working is rarely where the client first thought it was, so these three sit alongside this one more often than not.

Tell us what you are trying to build.

Describe the problem rather than the specification. If a smaller piece of work would answer it, we will propose that instead.