AI is most valuable when it improves a real workflow. The best starting point is not "where can we add AI?" but "where are people losing time, context or confidence?" That shift keeps investment tied to work that matters.
The evidence for that shift is now unusually clear. MIT's Project NANDA studied 300 public AI deployments alongside 150 leadership interviews and 350 employee surveys, and reported that around 95% of generative AI pilots delivered no measurable impact on profit and loss, despite an estimated $30–40 billion of enterprise spending. The authors attributed the gap not to model quality but to a learning gap — tools that failed to adapt to an organisation's actual workflows (MIT NANDA, *The GenAI Divide: State of AI in Business 2025*).
McKinsey's global survey the same year points in the same direction from the opposite end. Around 88% of organisations reported using AI in at least one function, but only about 6% qualified as high performers attributing more than 5% of EBIT to it — and the single organisational attribute most strongly linked to bottom-line impact was fundamental redesign of workflows, something only a minority of adopters had done (McKinsey, *The State of AI: Global Survey*, 2025).
Read together, those two findings say something specific: the constraint is not access to capable models. Everybody has that. The constraint is the willingness to change how work is done around them. That is the premise behind our AI development and automation practice.
Look for repeated judgement
Good candidates combine volume with a recognisable decision pattern: classifying enquiries, summarising complex records, drafting consistent responses or spotting exceptions. The goal is to support judgement while keeping accountability with people.
A practical filter is to score candidate workflows on four things:
- Frequency. Does this happen hundreds of times a month, or twice?
- Pattern. Could an experienced colleague explain the decision rule to a new starter in a few minutes?
- Tolerance. What is the cost of being wrong once, and can a person catch it before it reaches a customer?
- Evidence. Is there a record of past decisions to check performance against?
Workflows that score well on all four are where automation compounds. Workflows that score poorly on tolerance — irreversible decisions, regulated advice, anything touching someone's health, money or legal position — are places to build assistance rather than autonomy, if at all.
It is worth being direct about what this rules out. "Add a chatbot to the website" almost never scores well on this filter, which is one reason so many of them are quietly retired within a year.
Design the human handoff first
Automation needs a clear boundary. Define when the system can act, when it must ask for confirmation and how a person can understand what happened. Trust grows when uncertainty and provenance are visible.
In practice this means three design decisions made before any model is chosen:
- The confidence threshold. Below what level of certainty does the system stop and ask? This should be tunable, and it should start conservative.
- The audit trail. Every automated action needs a record of what was decided, on what basis, and which sources informed it. Without this, nobody can debug the system or defend a decision to a customer.
- The override. A person must be able to reverse an automated action easily and have that reversal fed back as a signal rather than lost.
MIT's research found that shortcomings in exactly this area — systems that could not retain feedback, adapt to context, or improve over time — were a common reason pilots looked impressive in a demonstration and collapsed in daily use.
Connect to reliable context
An assistant is only as useful as the information it can safely reach. Permissions, source quality, versioning and retrieval design matter more than a clever demonstration.
Build a narrow source set first and expand as accuracy is proven. A system grounded in twenty carefully curated documents will usually outperform one pointed at an entire shared drive, because the shared drive contains three versions of the same policy and nobody knows which is current.
This is where AI work turns into integration work. Access control has to mirror the permissions people already have — an assistant that surfaces a salary band to someone who should not see it is a data incident regardless of how good the answer was. Content needs structure and version control. Systems need to talk to each other reliably. That is ordinary custom development, and it is usually the larger share of the effort.
The same principle applies to content published on your own site: structured, well-maintained content is easier for both people and machines to retrieve accurately, which is one more reason content models and editorial workflow deserve attention before the AI layer is added on top.
Measure the workflow, not the model
Track time saved, resolution quality, adoption, escalation and customer impact. Model benchmarks can inform technical choices, but business value appears in the complete service.
One reason the MIT figure reads so badly is partly methodological: many pilots never established a pre-deployment baseline, so "no measurable impact" often means nobody measured. Before launch, record the current state — average handling time, error rate, backlog, escalation rate, customer satisfaction — so that the comparison exists later.
Adoption deserves particular scrutiny. MIT's researchers noted a pattern of employees using personal AI tools even where official deployments had stalled. If your sanctioned system has low usage while people are quietly pasting work into consumer tools, that is not an adoption problem to be solved with training. It is a signal that the sanctioned system is worse at the job, and a governance risk in its own right.
Move from pilot to product deliberately
Production AI needs monitoring, feedback, security and a plan for change. A focused pilot should test both usefulness and operating model so success can grow without creating unmanaged risk.
DORA's 2025 research, drawn from nearly 5,000 technology professionals, found that AI adoption behaves as an amplifier rather than a fix: it raised delivery throughput, but it also correlated with greater instability, more rework and longer recovery where testing, review and feedback loops were weak to begin with (DORA / Google Cloud, 2025). Teams with strong foundations got faster. Teams without them got faster at producing problems.
The implication for planning is straightforward. If your delivery pipeline, monitoring and review processes are already fragile, strengthening them is not a distraction from your AI programme — it is the part that determines whether the AI programme works. The same research found that a genuine user-centric focus was among the strongest predictors of teams seeing gains from AI at all.
A reasonable first step
If you are early in this, a sensible sequence looks like:
- Pick one workflow that scores well on frequency, pattern and tolerance.
- Record its current baseline for a month.
- Build a narrow assistant over a curated source set, with a conservative confidence threshold and a visible audit trail.
- Run it alongside the existing process, not instead of it.
- Compare against the baseline, then decide whether to widen the source set, raise the threshold, or stop.
Step five includes stopping. A pilot that produces a clear negative answer for a modest cost has done its job.
Where to go next
- To scope an automation opportunity properly, see AI development and automation.
- If the work involves connecting systems that do not currently talk to each other, that is custom development.
- For how this fits a broader product plan, read from idea to scalable digital product.
- For the platform foundations that determine whether AI helps or amplifies problems, read building platforms that stay useful as you scale.
Sources
- MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 — reported via Fortune, https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html
- McKinsey & Company / QuantumBlack, The State of AI: Global Survey (2025) — https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- DORA / Google Cloud, State of AI-assisted Software Development (2025) — https://dora.dev/
Useful digital work is a continuous loop: understand, design, build, learn and improve.
