· Zenous Team · 8 min read
The Permanent Pilot, by the Numbers: What the 2026 Evidence Says About AI Programmes That Never Ship
Four independent 2026 publications, read together, give one answer to why AI programmes stall. Individual productivity is up, enterprise returns are flat, complexity is eating benefits, and the costs that produce the dip have names. The numbers, with sources, and what a sponsor does with them.
AI programmes stall because the organisation around them did not change. Workflows were not redesigned, decision rights did not move, and the work of checking agent output was never funded. The model did what the pilot promised. The operating model did not.
That is our reading of four independent publications from 2026, set out below with their numbers so a sponsor can check the reasoning rather than take it on trust. None of the figures are Zenous data. All of them are attributed, and the sources are at the foot of the page.
The four sources
| Publication | What it measures | Sample |
|---|---|---|
| McKinsey, The State of AI: Global Survey 2026 | Enterprise-level financial impact of AI, and what separates high performers | Global survey across regions, industries and company sizes |
| PMI, Pulse of the Profession 2026: Driving Success in Complex Projects | Project outcomes, complexity, and AI as a complexity driver | 16th edition, published 12 May 2026 |
| Gartner, press release of 25 June 2025 | Forecast for agentic AI project cancellations and their causes | Analyst forecast; a January 2025 poll of 3,412 webinar attendees on deployment status |
| DORA, ROI of AI-assisted Software Development, and the 2025 State of AI-assisted Software Development | Engineering return on AI, the shape of that return, and its costs | 2025 edition: nearly 5,000 technology professionals |
Number one: individual gains, flat enterprise returns
Eight in ten respondents to McKinsey’s 2026 survey say AI has improved their own productivity. The share of organisations reporting that AI contributed anything to earnings before interest and taxes is unchanged from a year earlier, at 37%. The share McKinsey calls high performers, organisations attributing at least 5% of EBIT to AI and describing the impact as significant, is about 6%.
What separates the high performers is not how much AI they bought. McKinsey reports that nearly three-quarters of them fundamentally redesigned workflows because of their AI use, up from 55% the year before, against about one quarter of everyone else, and that they pursue growth or innovation alongside efficiency rather than efficiency alone. One more figure from the same survey is worth a sponsor’s attention: 32% of respondents say their organisation has decided against buying software because it could be built internally with agentic coding tools. That is a procurement decision showing up in a survey about AI, which is what an operating-model change looks like from the outside.
Our reading: the distance between 80% and 37% is where the permanent pilot lives. The individual got faster. The organisation did not change what it does with the time.
Number two: complexity is eating the benefits
PMI’s 2026 Pulse of the Profession, the 16th edition, reports that 31% of complex projects fail to achieve the full scope of their intended benefits, more than twice the rate reported for projects overall. 81% of project professionals say projects have become more complex in recent years, with 37% describing a significant increase. Teams that navigate complexity effectively are five times more likely to succeed, at an 88% success rate against 14% for teams that manage it poorly.
AI appears in that report as a driver of the complexity, not as the cure. Asked what is driving it, respondents name AI and automation first at 72%, ahead of market and economic pressures at 48% and competition at 41%. PMI also finds the PMO matters more as complexity rises: respondents in PMO-equipped organisations were more likely to rate their complex project management very or extremely successful, 63% against 57%.
Our reading: an AI programme lands on top of a portfolio that is already losing benefits to complexity, and it adds to the load before it takes any away. A programme that has not moved decision rights first is adding acceleration to a system that could not decide at the old speed.
Number three: the cancellation forecast, and its causes
Gartner’s forecast, published in June 2025, is that more than 40% of agentic AI projects will be cancelled by the end of 2027. The three causes it names are escalating costs, unclear business value, and inadequate risk controls. Its analyst’s description of the field at the time: most agentic AI projects are early-stage experiments or proofs of concept, mostly driven by hype.
Our reading: all three named causes are governance failures. None of them is model capability. A project cancelled for unclear business value never had a value owner; a project cancelled for inadequate risk controls never had a governance page that said what the agent may touch.
Number four: the productivity dip, and what sits inside it
DORA’s report on the return on AI-assisted software development is built around what it calls the initial “productivity dip” of a new rollout, and it exists to help teams defend a budget through that dip. The central claim of its 2025 research is the one to put in front of a steering committee: AI’s primary role is that of an amplifier, magnifying the strengths of high-performing organisations and the dysfunctions of struggling ones.
Its companion AI Capabilities Model names seven foundational practices proven to amplify AI’s positive effect on organisational performance, which is the same argument from the other side. The returns come from the system around the tools, meaning the quality of the internal platform, the clarity of workflows and the alignment of teams, rather than from the tools themselves.
The 2025 DORA survey supplies the mechanism. AI adoption correlated positively with throughput and negatively with delivery stability, and roughly 30% of respondents reported little or no trust in AI-generated code while shipping it anyway.
Our reading: the dip has components, and each can be budgeted. The checking, the re-skilling and the plumbing. We set out how in Budget the Verification Tax.
Read together
Four bodies with different methods and different audiences land on one diagnosis. The gain is real at the level of the individual and the task. The return at the level of the enterprise depends on three things the pilot never tested:
- Workflow redesign. What distinguishes McKinsey’s high performers. The test is which step was removed from the process, not which step got faster.
- Decision rights that moved. PMI’s complexity finding. If the steering committee decides at the same cadence it did before the agents arrived, the programme has a faster pipeline into the same queue. Decision latency is the number that shows it.
- Funded verification. The productivity dip DORA describes. Agent output has to be checked to the standard the programme’s controls require, by someone who can still do the task by hand, and that work has to be on the plan.
A programme with none of the three is what we call pilot theatre: the model works, the demo works, and the operating model was never rebuilt around it. It does not fail. It stays a pilot.
What a sponsor does with this
Three questions for the next steering meeting, each answerable from evidence the programme should already have:
- Which workflow steps have been removed since the agents arrived, and who signed off the removal?
- What is the median decision latency on the programme this quarter, and what was it before the agents?
- Where on the plan is the line for checking agent output, who owns it, and what did it cost last month?
A programme that answers all three is past the pilot. A programme that cannot is in the 40%.
What the evidence does not say
It does not say AI fails to work. Throughput gains and individual productivity gains are consistent across all four sources. It does not say the technology is immature; the sources treat it as deployed. It says the return is organisational, and that the organisations collecting it did the unglamorous work first.
If you want the three questions put to your programme by someone outside it, a delivery audit does that in 7-14 days, or book a consultation and bring the last steering pack.
Sources: McKinsey & Company, “The State of AI: Global Survey 2026” (QuantumBlack); Project Management Institute, “Pulse of the Profession 2026: Driving Success in Complex Projects” and its accompanying press release; Gartner, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027” (press release, 25 June 2025), whose wording is quoted directly above and which also reports a January 2025 Gartner poll of 3,412 webinar attendees; DORA, “ROI of AI-assisted Software Development” and the 2025 “State of AI-assisted Software Development” (dora.dev). Every figure above was verified against the publisher’s own page on 11 September 2026: PMI’s press release of 12 May 2026 for the Pulse findings, McKinsey’s survey page for the EBIT, productivity, workflow-redesign and build-versus-buy figures, Gartner’s press release for the cancellation forecast, and DORA’s publications page for the productivity dip and the amplifier finding. All are attributed to those publications and are not proprietary Zenous data. Where we say “our reading,” that is Zenous interpretation, not a finding of the source.