AI Engineering

AI that survives contact with production

Most AI projects die somewhere between the impressive demo and the Tuesday in month six when it quietly starts giving wrong answers. We build the unglamorous half — evaluation, guardrails, monitoring, retraining — so the thing you launch keeps working.

  • Custom models
  • LLM & RAG integration
  • Evaluation & guardrails
  • MLOps
NeuNov AI Engineering — custom models, LLM integration and production AI systems

Sound familiar?

“The pilot was great. Then it stalled.”

A notebook that works on your laptop is not a system. Getting from there to something your team can depend on is a different discipline — and it is the one that usually gets skipped.

“It’s right until it isn’t — and we can’t tell when.”

Without a test set and a scoring harness, quality is a feeling. We put numbers on it, so a change that makes things worse gets caught before your customers find it.

“Our best data is trapped in PDFs and spreadsheets.”

The proprietary knowledge that would make AI genuinely useful to you is sitting in invoices, contracts, email threads and folders nobody has opened in a year.

“Nobody owns it once it ships.”

Models drift, APIs change, costs creep. We hand over a system with monitoring, alerts and a runbook — and documentation written for the person who inherits it.

What we build

Everything below is scoped, priced and delivered as working software you own outright — not a strategy deck or a proof of concept that stops at the demo.

Custom model development

Classification, forecasting, scoring, ranking and anomaly detection trained on your own operational data — where a general-purpose model can’t see what makes your business different.

LLM integration & RAG

Retrieval-augmented systems that answer from your documents, policies and records — with citations back to the source, so answers can be checked instead of trusted blindly.

Assistants & chatbots

Narrowly scoped assistants for support, internal knowledge and onboarding, grounded in approved content, with clean handoff to a human the moment they are out of their depth.

Document intelligence

Extraction and structuring from invoices, contracts, forms, statements and scanned records — turning paperwork into fields your systems can act on automatically.

Evaluation & guardrails

A labelled test set, a regression suite that runs on every change, refusal behaviour for out-of-scope questions, PII handling, and rate and cost controls before anything goes live.

MLOps & lifecycle

Versioned prompts and models, reproducible deployments, drift and quality monitoring, alerting, and a retraining path — so improving it later doesn’t mean rebuilding it.

Deliverables

What you actually get

Every AI Engineering engagement ends with a working system in your environment and everything needed to run it without us. No lock-in, no per-seat licence on your own software.

Scope this with us
  • A deployed, working system running in your cloud account
  • An evaluation harness with a baseline score you can hold us to
  • Versioned prompts, model artifacts and configuration
  • Monitoring, alerting and cost controls wired up on day one
  • A runbook plus a live walkthrough for your team
  • Full source code and IP handover — you own all of it

From first call to handover

  1. 01

    Frame the decision

    We start from the business decision the model is meant to improve, not the technology. What action changes? Who acts on it? What does being wrong cost? That defines the accuracy bar everything else is measured against.

  2. 02

    Prove it on your data

    A focused build against a real sample of your data with a labelled test set, scored honestly — including the option to tell you it isn’t worth doing. Better to learn that in week two than month six.

  3. 03

    Harden it

    Guardrails, failure paths, PII handling, cost ceilings, and a regression suite so future changes can’t silently degrade quality. This is the step most projects skip, and the reason most projects don’t last.

  4. 04

    Ship, watch, hand over

    Deployment into your environment with monitoring and alerts, a walkthrough for your team, documentation, and full handover. Ongoing iteration is available on retainer if you want it — never required.

AI Engineering in the real world

AI Engineering pays off wherever expert judgement is applied over and over to the same kind of input. A few places it lands well:

Professional services

Contract and proposal review that flags unusual clauses and missing terms before a human reads page one.

Accounting & finance

Invoice and receipt coding to the right account, with confidence scores and low-confidence items routed for review.

Logistics & distribution

Extracting structured data from bills of lading, customs paperwork and supplier PDFs that arrive in fifty different formats.

Healthcare admin

Intake triage and referral routing on structured and free-text notes, keeping clinicians out of the paperwork queue.

Retail & e-commerce

Demand forecasting, product categorisation and review analysis on your own catalogue and sales history.

Field services

Turning technician voice notes and photos into clean job reports and follow-up tasks without a second pass.

AnalyseThisWC26 multi-model match predictor with out-of-sample model track record

A predictor we backtested before we published it

AnalyseThisWC26 runs a multi-model match predictor — Dixon-Coles, Poisson and Elo — with an out-of-sample track record published in the product itself, next to the predictions. That is the same discipline we bring to a client model: score it honestly on data it has never seen, then show the score.

Read the case study

Tools & platforms we build AI Engineering on

  • Python
  • PyTorch
  • scikit-learn
  • Claude
  • OpenAI
  • Open-weight models
  • LangChain
  • pgvector
  • AWS Bedrock
  • AWS SageMaker
  • Docker
  • GitHub Actions

AI Engineering — questions we get asked

Do we need a huge dataset before this is worth doing?
Usually not. Retrieval and LLM-based systems work from the documents you already have, and many classic models are useful with a few thousand well-labelled examples. The first thing we do is look at what you have and tell you honestly whether it is enough — that assessment is part of the Automation Audit.
Which model do you use — and are we locked into one vendor?
We pick per task and benchmark the options against your data: a frontier model where reasoning quality matters, a small or open-weight model where cost and latency matter. The integration layer is written so the model can be swapped without rewriting the application.
How do you stop it from making things up?
Grounding, scope limits and testing. Answers are retrieved from your approved sources with citations, the system is built to say “I don’t know” and hand off rather than guess, and a regression suite runs against a labelled test set on every change so quality drops are caught before release.
Does our data get used to train someone else’s model?
Not without your explicit decision. We use enterprise API terms that exclude training on your data by default, keep secrets in environment variables rather than code, apply least-privilege access, and can run entirely inside your own cloud account.
How long does an AI Engineering build take?
The Automation Audit runs about two weeks and tells you what is worth building. A first production build is fixed-scope and agreed before work starts — we scope it with you on a short call so there is no open-ended billing.
What happens after launch?
You get monitoring, alerts, documentation and a walkthrough, so your team can run it independently. If you would rather we keep iterating, monitoring and retraining, the Automation Retainer covers that at a monthly cadence.

Most engagements touch more than one

A good automation usually needs clean data behind it and a decent interface in front of it. Here’s the rest of what we do.

Have an AI idea that stalled at the demo?

Bring it to a free 30-minute call. We’ll tell you what it would take to make it production-grade — or tell you plainly if it isn’t worth building.

Free 30-min call · no commitment · you own everything we build