Pilots without a path
IDC found that 88% of AI pilots never reach production. The models usually perform; what is absent is the ownership, the controls and the unit economics that would make deployment a defensible decision.
Data, AI and decision systems
Most AI programmes stall between a working demo and a system someone will own. The reason is rarely technical — RAND found that the root causes are organisational in over 80% of cases. We work on that gap: the evaluation, the controls, the operating path and the cost model.

The current reality
McKinsey reports that only 12% of CEOs see both cost and revenue benefits from AI — despite near-universal adoption. The models work. What is missing is the operating layer around them: who owns the outcome, how it is evaluated and what happens when it is wrong.
IDC found that 88% of AI pilots never reach production. The models usually perform; what is absent is the ownership, the controls and the unit economics that would make deployment a defensible decision.
Gartner's CxO survey found that only 27% of executives have a comprehensive AI strategy. Without that frame, well-built systems go unused — teams do not adopt tools whose purpose has not been articulated from the top.
78% of enterprises are unprepared for EU AI Act obligations — not because controls are hard to build, but because most governance frameworks tell teams what they cannot do without showing them what they can.
What we do
We score competing initiatives on expected value, technical risk, dependencies and evidence quality — then turn the result into a sequence that leadership can defend when capacity or priorities shift.
We take one consequential workflow and build it to production: retrieval, agent orchestration, evaluation harness, human review path and the cost model that decides whether it should exist.
We design the data contracts, evaluation sets, logging standards and risk tiers that let a team ship a second and third system without renegotiating approval each time.
We redesign the roles, rituals and incentives around the new system, then train the people who will own it after we leave. If adoption depends on a change-comms deck, the design is wrong.
Selected work
A Fortune 500 analytics platform, a Lloyd’s-backed reinsurer’s risk workflow, and CTO portfolio governance across 35+ initiatives. Clients are anonymised; the numbers and the mechanics are not.

A Fortune 500 medical devices group had capable analysts in every market and no shared definition of a customer. A governed platform and a self-serve data catalogue grew organic adoption ~80% year on year.

At a Lloyd’s-backed reinsurer, scenario logic sat in actuarial models, exposure sat in portfolio systems and judgement sat in email. Discovery with underwriters and actuaries produced a single traceable workflow.

A Fortune 500 gaming and hospitality group ran 35+ concurrent technology initiatives with no common view of risk or return. The CTO and eight senior leaders adopted the resulting model as primary governance within six weeks.
Organisations we have worked with since 2016

How we work
We put the sponsor, the product lead and the engineers in the same room and force the disagreements out while they are still cheap. On the reinsurance work, running discovery with underwriters and actuaries together surfaced a definitional conflict that would otherwise have shipped.
Photo: Jo Szczepanska / UnsplashLatest thinking

A demo proves feasibility under conditions you chose. Five decisions stand between that and something an organisation can depend on.

Counting completed reviews tells you nothing about whether a system is trusted. Better signals exist, and they change what governance teams build.

An agent has no unit cost until you define a successful outcome. Retries, retrieval and human correction usually dominate the token bill.

A language model without memory restarts every conversation from zero. The architectures that close this gap are well understood — the hard part is deciding what a system should forget.
Start a conversation
Tell us what is changing, where the difficulty sits and what a good outcome would look like. We reply within two working days.