Skip to main content
STRATHBERG

Market Perspective

28 August 2026 22 min read

If AI, Then

George Mudie

George Mudie

Senior Partner

Part 1 of 3

If you can keep your head about artificial intelligence while the boards around you are losing theirs.

If you are brave enough to resist the pressure to act before anyone has agreed what would be useful.

If you can insist on an answer that can be checked, while others act on one that merely reads well.

If you can recognise what your organisation needs for honest decisions, before the cost of misunderstanding becomes too much.

Then you may find yourself in the 5%.

5% is not a rhetorical flourish. BCG’s September 2025 study of 1,250 companies found that 60% were generating no material value from AI, 35% were scaling and starting to see returns, and 5% were creating substantial value at scale [1]. The distance between those groups is the subject of this article, and the explanation has surprisingly little to do with how clever the models are.

Consider the name for a moment. A large language model works by predicting, one piece at a time, which words are most likely to follow the words already written. It is extraordinarily good at this, and the results can be genuinely astonishing. But if the technology had been launched as Probably The Right Words, would the board have approved the budget quite so briskly? If the invoice had read Sounds Convincing Answer Generator, would anyone have signed it without first asking what the alternative cost?

This is not a criticism of the technology, which does things nothing else can. It is a question about the label. “Artificial intelligence” describes what the technology aspires to be. “Probably the right words” describes what it does. Much of the confusion in enterprise AI since 2023 sits in the gap between the two.

There is a second joke hiding in the word “if”. To a poet, it opens a conditional. To an engineer, it opens a rule. And a rule has a property a prediction does not: you can check it. Ask it the same question tomorrow and it answers the same way, for reasons someone can point to. This deterministic approach has been driving decisions in your organisation for as long as it has been computerised.

Perhaps it is time to pay more attention to the deterministic kind of intelligence, rather than the purely artificial.

What the adoption numbers actually say

Everyone is doing this already. That impression is widespread, and it is causing real damage in boardrooms.

The widely quoted figures come from global surveys that companies opt into, and they count whether an organisation has an AI initiative somewhere. On that basis, 88% of organisations report using AI regularly in at least one business function, and 72% report using generative AI, up from 33% a year earlier [2]. Stanford’s AI Index reports 78% of Fortune 500 companies with active generative AI initiatives [3].

National statistics count something different, and they count it more carefully. The Office for National Statistics found 35% of UK businesses with 10 or more employees using at least one AI technology in June 2026, up from around 12% in late 2023. The average adopter uses 1.6 technologies, barely moved from 1.4 three years earlier, and only 10% describe their use as extensive [4]. Eurostat and the OECD find much the same: 15.1% of people aged 16 to 74 across the EU use generative AI for work, and 20.2% of OECD companies use AI at all [5][6].

The ONS breakdown deserves a second look. Large language models lead at 18%, followed by visual content creation at 16%, data processing using machine learning at 12% and image processing using machine learning at 6% [4]. Those overlap heavily and cannot be added, which is what the average of 1.6 tells you. More interesting is what the list contains. Half of those categories are machine learning, which behaves quite differently from the other half, and the official statistics count them all as one thing called AI. That conflation is the subject of this article.

In summary, nearly every enterprise is playing with AI, but very few have successful production deployments. Only 39% report any EBIT impact, and where it exists it is most often below 5% [2]. Agentic systems have not changed this: around 66% of enterprises have experimented with them, and fewer than 10% have scaled them to quantifiable value [7].

The usual explanations are data quality, change management and skills. All three are real. Change management deserves the closest attention, because the ground beneath it has shifted.

Pew Research surveyed 3,488 US adults in June 2026 and found 52% more concerned than excited about the growing use of AI in daily life, up from 37% in 2021 [8]. The age breakdown is the part worth a CTO’s time. Among adults under 30, concern has climbed from 31% in 2021 to 55%, the first time a majority of that group has held that view, while the share more excited than concerned has fallen from 25% to 11% [8].

Under-30s use these tools more than any other age group. Their enthusiasm has cooled the fastest.

That is not unfamiliarity, and it will not respond to better communication. It is what happens when heavy users meet the gap between an answer that reads well and an answer that is right. The scepticism is earned, and it is one of the more valuable things in the building. An organisation that can show people precisely where the technology is trusted and where it is not gives that scepticism somewhere useful to go. The Pew figures are American, and the direction is unlikely to be unique to the US.

Better data or better change management would produce a spread, with well-run organisations doing somewhat better than the rest. What the numbers show is closer to a split: a large group getting nothing and a small group getting a great deal. Something is present in that small group and absent in the other, and it is worth working out what.

The demonstration problem

The pattern has now repeated in enough organisations to be worth setting out plainly.

A board reads the same research notes everyone else is reading and concludes, entirely reasonably, that the organisation is falling behind. Slow adoption is a genuine competitive risk and they have been told so by people they have good reason to trust. They apply pressure, which is what boards are for.

An AI task force is assembled, and within a few weeks there is something to show: a system that accepts a question in ordinary English and returns a serious, well-structured, confident answer. It cites sources. It explains its reasoning. In the demonstration it is close to magic, and the demonstration is not a deception, because the system genuinely does what it appears to do.

What got compressed in those weeks was the unglamorous work. Nobody settled what the terms meant. Ask that system for margin and it will give you a number. Which margin, though, against which cost base, in which currency, before or after returns, at which point in the season? The question was never defined, so it answered a question that did not exist, in a register that suggested otherwise. Nobody decided which questions it was permitted to answer, or what it should do when it did not know. The delivery plan rarely has as line labeled “agree the ontology”, and if it did, it would not survive a board asking for something to show by Thursday.

Then it goes into use, and the frustration curve begins. Wrong occasionally, then noticeably, then in a way that costs somebody something. Two colleagues ask the same question in slightly different words and receive different answers. Confidence collapses far faster than it was built, and it takes some confidence in the task force down with it.

This has happened in public, to firms whose entire product is verified analysis. One major professional services firm partially refunded a government fee in 2025 after AI-generated errors were found in a report, including references to papers that did not exist and a fabricated quotation from a court judgment [9]. An investigation into a report by another in 2026 found that most of its citations had been invented [10]. Both had mature review processes and every commercial incentive to catch exactly this.

Consider why the users were persuaded, because they were not being foolish. Fluency has been a dependable proxy for competence in every human interaction any of them has ever had. A colleague who answers precisely, in well-organised prose, with references, is usually a colleague who knows. That heuristic has served them their entire working lives. These systems are the first thing they have encountered that inverts it, producing the linguistic signature of expertise with none of the machinery underneath.

The illusion of intelligence blinds the user to the absence of reasoning. That is the failure, and it explains why the pattern recurs in organisations that have nothing else in common.

Take a moment to reflect. Has your organisation shown any of these signs?

Three kinds of AI, and what each can be trusted with

Most executives have met the standard taxonomy, and it is worth knowing why it does not help here.

CategoryWhat it meansWhere it stands today
Artificial narrow intelligenceCompetence within a limited domain or set of tasksNearly all deployed AI systems, including frontier language models, are generally classified here
Artificial general intelligenceHuman-level competence across a wide range of domainsNo consensus definition, no agreed test, and no widely accepted example exists today
Artificial superintelligenceCapability exceeding humans across essentially all domainsEntirely prospective

Each row places the system at a different point along the capability spectrum: how much the system can do. None says anything about whether the answer can be checked. Those are separate properties, and they have moved at very different speeds. System capability has expanded enormously over the past 3 years. The ability to verify an answer has barely moved, as nothing in the architecture that produced that capability was built to improve it.

For anyone responsible for a decision, it is more useful to sort the technology by what it can be trusted to hold.

CapabilityHow it behavesWhat it can be trusted to hold
Generative AI (large language models, image and code generation)Produces the most probable continuation. Same question, different answer. No mechanism to verify its own outputThe conversation around a decision: capturing intent, summarising, explaining, drafting, translating
Classical machine learning (forecasts, propensity, scoring, anomaly detection)Learns from data, but returns the same output for the same input once trained. Testable, monitorable, auditableInputs to a decision, and many decisions outright, subject to model governance
Rules, constraints and solvers (pricing engines, allocation optimisers, eligibility logic, screening)Executes logic someone wrote down. Same input, same output, always, with the reasoning inspectableThe decision itself, including anything that must be evidenced later

The second and third rows share something the first does not, and it has never had a shared name.

Deterministic intelligence

Most leaders already know the principle. You make your best decisions when you have information that is concise, relevant and accurate. Nothing about that changes because a technology is new. What has changed is how easy it has become to obtain something resembling accurate information without anyone checking whether it is.

Here is the working definition, in the plainest form.

Deterministic intelligence is any system that gives you the same answer every time you ask it the same question, and can show you how it got there.

Ask it twice, get the same result. Ask it why, and something can be produced: a rule, a calculation, a traceable path back to a source.

A promotion or redundancy selection makes the test concrete. Employment law generally requires an employer to explain the basis of the decision and defend it if discrimination is alleged. A system that scored candidates against stated criteria with stated weightings can support that. A system that generated a paragraph of persuasive reasoning cannot, however sound the reasoning appears, because there is no underlying decision record to inspect. The first qualifies as deterministic intelligence. The second does not.

For those who want the technical version: computation whose output is deterministic by construction and, where the stakes require it, verifiable against a specification.

Two distinct properties sit within that, and they are often conflated.

Deterministic means identical inputs produce identical outputs, by design rather than good fortune. You test it by running it again. This is the standard meaning of the term in computer science, and it is worth using precisely: the Association for Computing Machinery reserves “reproducible” for something different, meaning a separate team obtaining the same result using the same setup [16].

Verified means the output can be certified correct against a specification. You test it with a proof, a model checker, or by exhausting the possibilities.

That yields three assurance levels, measuring the thing the capability taxonomy ignores. The idea of grading a system by how much confidence its output can carry is well established: aviation software has used design assurance levels for decades, and functional safety engineering uses safety integrity levels in the same way.

LevelPropertyHow you test itWhere it lives
0. GenerativeSame input, different outputYou cannotLanguage models, generative retrieval, agentic chains
1. DeterministicSame input, same output, by constructionRun it twiceRules engines, decision tables, knowledge graphs, most classical ML at inference
2. VerifiedCertifiably correct against a specificationProof, model checking, exhaustive searchConstraint solvers, theorem provers, simulation with bounded error

Deterministic intelligence comprises levels one and two.

Three levels of assurance. Level 0 Generative gives a different output for the same input and cannot be tested. Level 1 Deterministic gives the same output for the same input by construction and is tested by running it twice. Level 2 Verified is certifiably correct against a specification, tested by proof or model checking. Deterministic intelligence is levels one and two.

Note where classical machine learning sits. Once trained, most of it is deterministic in use. A demand forecasting model built on decision trees, the standard approach in retail and supply chain, returns the same forecast on Tuesday as it did on Monday when given the same inputs. The same is true of a credit scorecard or a propensity model. What breaks determinism in a large language model is not the model itself but the sampling layer, together with whatever variability arrives through retrieval and tool use. This is not old technology versus new. It is architectures whose answers can be checked against architectures whose answers cannot.

Which means most organisations have been running a substantial deterministic intelligence estate for a decade or more. The pricing engine, whose promotional stacking and margin floors only work because the rules are written down. The allocation solver, which returns an answer that can be proved optimal against its constraints. The demand forecast. In financial services, the transaction monitoring thresholds that are calibrated and documented because a supervisor will eventually ask, the credit model that must explain itself to a customer, the sanctions screening where approximately right is not a category that exists.

None of that was ever named as a category. Which meant it never had a policy, an advocate, or a line in the strategy deck. And when the generative layer arrived, it was frequently wired into the same workflows and the same data, with no boundary drawn between the two.

The term is not entirely new. A short manifesto published in January 2025 set out deterministic intelligence as computation that is auditable, explainable and reproducible, under the neat formulations “AI guesses, DI knows” and “AI predicts, DI decides” [11]. The instinct was sound; what it lacked was any test for what belongs in the category, and any account of where machine learning fits. The three levels above are an attempt at both.

The underlying engineering is well established. The Alan Turing Institute runs an interest group on neuro-symbolic AI, and its framing supports the distinction directly: symbolic techniques built on rules, logic, and reasoning are far easier to inspect, explain and verify than methods that learn their behaviour from data alone [12]. IBM runs a comparable programme [13], and verified AI has serious literature behind it [14].

Those sources describe how to build systems. What has been missing is a way to describe the assurance of a decision in language a board can use. The closest alternative source is the UK AI Security Institute, whose published work is focused chiefly on frontier model safety and control rather than enterprise decision architecture [15]. It therefore does not address the broader question of what a business decision may properly rest on, which currently sits outside anyone’s formal remit. That is a question for organisations themselves to answer.

Where deterministic intelligence stops

There are three limits, and being clear about them is what makes the rest of it usable.

Deterministic is not the same as correct. A rules engine encoding a bad rule will give you the wrong answer reliably, on demand, forever. Determinism tells you the system is behaving consistently, not that it is behaving well. Only Level 2 speaks to correctness, and then only against the specification it was handed.

The specification is where the risk relocates. A provable system is provably correct with respect to something a human wrote down. Get that wrong and the proof remains valid while the outcome is wrong. Formalising a decision moves the judgment; it does not remove it, and it is worth knowing exactly where it has moved to.

And determinism cannot handle ambiguity at all. Intent. Disambiguation. Ordinary language. Summarising 400 pages. Explaining a declined credit application to someone who is upset about it. Turning a constraint model into something a board will actually interrogate. Drafting anything. No deterministic method touches any of this, and it represents an enormous proportion of the work an organisation does.

Where AI genuinely complements it

That third limit is not a gap to be apologised for. It is where AI earns, and it deserves to be said with some enthusiasm.

The reported gains point the same way. Where organisations describe measurable benefit from generative AI, it clusters in drafting, summarising, search, translation and coding assistance. Every one is a comprehension or communication task, and none is a decision that binds anybody to anything.

Language models belong at the interface. Let someone question the forecast in English rather than learn the reporting tool. Let the system explain to a store manager why the markdown rule fired, or to a relationship director why the credit model declined. Let it read 10,000 documents and summarise the compliance position, then draft the regulatory narrative once that position is established. Let it handle the customer conversation while entitlement is resolved by rules behind it. This is not a consolation prize. For most organisations there is more work of this kind than there is decision-making, and it is where the returns have shown up.

Classical machine learning is already inside the estate. The forecast, the propensity model, the anomaly score. Deterministic at inference, and it should be governed alongside the rest of the deterministic layer rather than swept into an AI policy drafted with chatbots in mind.

Generative AI outside evidenced decisions needs no permission from any of this. Trend and design ideation, personalisation, concept generation, marketing variants. The error cost is low, the feedback is immediate, and a human holds the judgment by design. Nothing in this argument applies, and nobody should be slowing that work down on the strength of it.

The trap was never using AI. The trap is letting the probabilistic layer hold the conclusion.

That is what explains the demonstration that worked. It was an interface with nothing underneath it, answering questions that had never been defined and using terms nobody had agreed upon. The users were not fooled by a bad system; they were fooled by a system that was excellent at the part they could see.

The neglect, and the reset that follows

There is a predictable consequence to all of this, and it is worth naming while there is still time to act on it.

The deterministic estate is being starved of attention at precisely the moment it is carrying the load. It is unfashionable. The people who maintain it are not on the AI steering committee. Its budget is defended, when it is defended at all, as legacy maintenance rather than as the only part of the estate whose answers can be checked. Meanwhile a generative layer is being built alongside it, with no boundary between them, by teams responding honestly to genuine pressure.

That will be corrected, and the correction will not be cheap. It will arrive in two or three years as a transformation programme with a proper name and a substantial budget, and it will consist largely of doing the unglamorous work that was compressed out of those first few weeks: agreeing what the terms mean, deciding which questions may be answered, and putting a boundary between the layer that guesses and the layer that decides.

Doing this work now is far cheaper than doing it twice, and the organisations that sequence it properly will not need the second programme at all.

Three things to take away

Your organisation already owns a deterministic intelligence estate, and it probably has no policy. Pricing, allocation, forecasting, monitoring, screening, decisioning: identify it, name it, assign an owner, and defend its budget on the right grounds. The case is not that it is legacy, but that it is the part of the stack whose answers can be evidenced.

Draw the boundary, then spend confidently on both sides of it. The question to ask of any AI-touched process is where an error becomes binding before a person sees it. Anything answering “immediately” belongs at level one or two. Everything else is fair game, and hesitating there costs more than it saves. Fund the interface layer properly rather than as a pilot; once the decision layer is secure, the constraint on AI value is usually the quality of the conversation around the decision rather than the model behind it.

When a system gives you an answer, ask what checked it. Not who reviewed the output, which is a weaker question. What in the system verified the claim, and against which source. If the answer is that it produced text resembling a checked answer, then a probabilistic system is holding the decision. It will eventually be exposed. The only variable is what it costs when it is.

A fair question to end on

There is an obvious challenge to all of this, which is that a consultancy has drawn a distinction it is well placed to sell into. It is a reasonable challenge, and the argument cannot answer it by asserting harder, having just spent several thousand words insisting that claims require evidence.

A better test is to look at what organisations are actually spending, because budgets record decisions more honestly than strategy documents do.

They record something worth knowing. AI has moved from 11% to 18% of the average enterprise IT budget in two years, and most of that has been funded by reallocating money rather than adding it. Tracing where it came from tells you a great deal about which systems an organisation was defending and which it was not, and it points to some genuinely useful conversations with suppliers. That is the subject of Part 2.


Acknowledgement

The opening lines are written with gratitude to Rudyard Kipling, whose “If” has survived rather more than a century of quotation and remains the finest four minutes of advice on keeping one’s composure while everyone else abandons theirs. For anyone who has not heard it read properly, Ralph Fiennes’ recording is worth the time: https://open.spotify.com/track/2xZpLP86l3a430BCNqDCXN

References

  1. BCG (2025). The Widening AI Value Gap: Build for the Future 2025.
  2. McKinsey (2025). The State of AI.
  3. Stanford HAI (2025). AI Index Report.
  4. Office for National Statistics (2026). Artificial Intelligence in UK Businesses: 2023 to 2026. Official statistics in development.
  5. Eurostat (2025). Use of Artificial Intelligence by Individuals.
  6. OECD (2025). AI in Business.
  7. McKinsey (2026). Building the Foundations for Agentic AI at Scale.
  8. Pew Research Center (2026). Young Adults in the US Are Increasingly Wary of AI, Concerned It Will Take Jobs.
  9. Deloitte Australia (2025). Report for the Australian Government; partial fee refund following AI-generated errors.
  10. GPTZero (2026). Investigation into citations in an EY Canada report.
  11. Deterministic Intelligence Initiative (2025). The Deterministic Intelligence Manifesto.
  12. The Alan Turing Institute. Neuro-symbolic AI interest group. Web resource, accessed August 2026.
  13. IBM Research. Neuro-Symbolic AI research programme. Web resource, accessed August 2026.
  14. Seshia, S. A., Sadigh, D. and Sastry, S. S. (2016). Towards Verified Artificial Intelligence. arXiv:1606.08514, revised 2020.
  15. UK AI Security Institute. Published research programme. Web resource, accessed August 2026.
  16. Association for Computing Machinery (2020). Artifact Review and Badging, version 1.1.

Part 2, Then Comes the Bill, traces where the AI budget came from and what it actually bought. Part 3, Then Prove It, applies the same test to transformation programmes.

Disclosure: Strathberg is developing capability in deterministic analysis. The definition above stands or falls on its own terms.

George Mudie

George Mudie

Senior Partner, Strathberg

Senior Partner at Strathberg. Three decades of experience as Global CTO and CISO across retail, telecoms, media, and intelligence.

Discuss this with us

If this resonates with a challenge you are facing, we are happy to have a direct conversation.

Get in touch