8 min read

Why AI Pilots Fail: The Behavior Layer Underneath the 95% Number

MIT found 95% of generative AI pilots produce no P&L return. The reason sits under the tool, in the habit layer running the workday.
A worn footpath through tall grass
Photo by Courtney Smith on Unsplash

In 2025, MIT NANDA's State of AI in Business report put a number on something a lot of leaders already suspected: 95% of generative AI pilots produced no measurable return on the P&L. The number traveled fast, and the explanations attached to it traveled faster. Poor data and weak integration top the list. If either were the mechanism, the fix would be a procurement decision.

But the number hides a more ordinary story: the pilot demoed well, people used it for a few weeks, and then it faded. Nobody decided to stop. The workday resumed its old shape, and the tool became a tab nobody opened.

TL;DR: A pilot dies when a new tool lands on top of an unchanged habit, because roughly 40% of a workday runs as habit rather than decision. The fix is moving the feedback loop so the signal arrives before the old habit fires, and auditing the workday down to the keystrokes to find where that habit actually lives.

The Number Everyone Cites

Fortune's coverage of the MIT report captured how quickly the finding became a verdict about tools and talent.

Wendy Wood, a psychologist at USC, has spent her career measuring how much of daily behavior is deliberate. In a 2002 study of everyday behavior, she and her colleagues found that roughly 40% of what people do in a day repeats as habit, performed in stable contexts with little active decision-making. If 40% of a workday is a recording, then a pilot is asking a tool to interrupt a recording, and the recording plays on schedule without anyone paying attention to it.

The Habit That Runs the Workday

I wrote about this mechanism in 2016, nine years before the MIT number existed, in a piece about productivity apps. People eagerly buy the book, download the special app, and test out the system. They're all in on GTD, or bullet journaling, or four-hour workweeks. I know, because I've been there. A few weeks later, life happens and things get hectic. As a survival mechanism, it's easy to slip back into old habits and stop using the app, reading the book, or remembering the system. A few weeks after that, the cycle restarts. Frustration resurfaces, and most people buy another book or try a different app to see if they get better results.

Swap "GTD" for "the AI pilot" and every beat holds. The enthusiasm is real, and so are the first weeks. The post-mortem blames the app, so the organization buys the next app.

One thing I wrote in 2022 still shapes how I diagnose this: "'Good' and 'bad' habits are hard to drop, because we've assigned them these specific values in our life. But when you simply have a habit, you can change it whenever you want." A habit is running, and a habit carries no moral verdict, which means it can be changed without anyone being to blame for it.

The Adopted Tool and the Practiced Motion Still Running the Hands

I build AI systems and workflows for a living. I co-founded CTOx, a fractional CTO accelerator.

In a team meeting in the spring of 2025, I caught myself describing my own workflow out loud: "It's very easy, partially because I have this muscle memory. I'm prompting with it, and I'll export it as a Word doc. I noticed that you can still export it as a PDF. It looks funny. Maybe the large language models don't care what it looks like."

I had adopted the tool. I was actively prompting with it, in that moment. And my hands still reached for the Word export, because the Word export was the practiced motion, the one my muscle memory owned. I was even judging the output by the old reader's standards, flagging that the PDF "looks funny."

The First Error Writes the Story

Dietvorst, Simmons, and Massey published a 2015 study with a title that tells you the finding: Algorithm Aversion: People Erroneously Avoid Algorithms After Seeing Them Err. Participants watched an algorithm and a human forecaster both make mistakes, then chose whom to trust going forward. They abandoned the algorithm after a single visible error, even when the algorithm outperformed the human overall. MIT's Initiative on the Digital Economy has an explainer on the same body of work that makes the uncomfortable part plain: people forgive their own errors and the errors of other humans far more readily than a model's.

The tool produces one bad output in week two. The habit, which has produced thousands of acceptable outputs over years, is already practiced and immediately available to run. This is why the fade needs no decision.

Why More Training Misses the Broken Layer

The standard response to a faded pilot is more training, and I understand the instinct. The team already knew the tool existed. Knowing about the tool is one layer, doing the work by habit is another, and the pilot broke in the doing layer.

I said this in 2025: "AI is great at sitting beneath or behind the everyday processes you do. A trap that most people fall into is they add AI on top of their existing workflows, and then they rightfully complain that AI isn't saving them time. When I coach executives and entrepreneurs to help them get 5 to 10 hours a week back, we start by rebuilding the foundation of how they work. It's just as much unlearning as learning."

The unlearning has a granularity. "A lot of us do what we do because that's how we were taught. That's how we were trained. That's how we taught ourselves to do it. In some cases that could be months or years ago. First principles thinking is about identifying the assumptions you have about every step of your workday and how you actually do each task, right down to the keystrokes and the mouse clicks." That's where the habit lives. My Word export lived there.

Rebuilding my own workflows from the ground up after ChatGPT arrived took sustained daily effort over a long stretch. An hour of training against that kind of effort is a mismatch.

The environment around the habit matters too. In 2025 I talked about negative incentives: the moments in a team culture, or in your personal leadership, where you may penalize adaptability. If a team member uses ChatGPT and gets called out in front of the group, you've immediately set the expectation that no one is to use ChatGPT.

McKinsey's State of AI research keeps finding that redesigning workflows separates the organizations seeing returns from those running perpetual pilots. I'd hedge the exact proportions, since survey populations vary.

The Fix Is a Moved Feedback Loop

So what does changing a habit actually require? Psychology Today's explainer on how habits break points to context cues and friction. You change what the situation prompts, and you change how easy each path is. In other words, the signal has to arrive before the habit fires, at the moment of the cue.

I learned what that feels like in my own decision-making in a session in July 2025, and the parallel to a dying pilot is exact. "I feel like the internal volume knob has been turned up. When I'm making decisions that are not aligned, I'm hearing it more clearly when that happens. Do you want to do that? Hold it. I would do the pause. No, I don't want to do that thing."

The old loop: "In the old days, I'd do the thing, launch into action. Only then would I start to hear maybe that wasn't such a great idea." "The work has been in shortening and compressing the time between that intuitive signal and when I take corrective action. What I'm noticing is that the volume knob is turned up before I've made the decision out in the world, so the corrective action can happen before I've overcommitted to something."

The moved pilot loop: a standing 15-minute weekly check where the team hears the signal before the overcommitment, naming one moment the old habit fired and one moment the tool earned its place, while the week is still warm.

This is also why I've kept a practice since 2011, when I started Ridiculously Efficient, that I call Destroy to Create: every 90 days, I fundamentally reinvent one core workflow, a routine, or how I spend my weekend.

The Question to Ask About Your Pilot

Log your workday for one week so you can see where the hours actually go. I like doing this manually with a spreadsheet, because I want to capture the verbs and nouns involved. Which single task, handed to AI, would give you back an hour or more every day?

When you've found that task, design the loop around it. Take email as a generic worked example. The old way: manually write the email, manually edit it, overanalyze the wording, seek an external proofreader, send after multiple revisions. The new way: prompt AI with the desired outcome, tone, and recipient to generate a draft, then you, the human, edit for personal touches. The habit underneath it is a different object entirely, and building the second one is the actual work the 95% statistic is measuring the absence of.

The pilot that faded on your team's watch faded because a habit outlasted a tool.

Which habit is still running your workday underneath the tool you bought to change it?

Frequently Asked Questions

Why do AI pilots fail after a successful demo?

A new tool lands on top of an unchanged habit, and the habit outlasts it. The pilot demos well and gets used for a few weeks, then the workday resumes its old shape while the tool becomes a tab nobody opened. Nobody decides to stop, because roughly 40% of a workday runs as habit rather than decision.

What percentage of AI pilots fail, according to MIT?

MIT NANDA's 2025 State of AI in Business report found that 95% of generative AI pilots produced no measurable return on the P&L. The finding traveled fast, and the explanations attached to it traveled faster, with poor data and weak integration topping the list.

Why do people stop trusting AI after one mistake?

Dietvorst, Simmons, and Massey showed in 2015 that people abandon an algorithm after a single visible error, even when it outperforms a human forecaster overall. People forgive their own errors and the errors of other humans far more readily than a model's. The habit, already practiced and immediately available, runs again with no decision required.

How do you change a habit at work?

Change what the situation prompts and how easy each path is, so the signal arrives before the habit fires. A standing 15-minute weekly check, naming one moment the old habit fired and one moment the tool earned its place, moves the feedback loop ahead of the overcommitment. Start by logging your workday for one week, down to the keystrokes and mouse clicks, so you can find where the habit actually lives.

Why doesn't more AI training fix a failed pilot?

Knowing about the tool is one layer, and doing the work by habit is another, and the pilot broke in the doing layer. Rebuilding a workflow from the ground up takes sustained daily effort over a long stretch, and an hour of training is a mismatch against that. It is just as much unlearning as learning.

Stop Adding. Start Subtracting.

The world keeps accelerating. The Simplicity Protocol helps ambitious professionals do less to achieve more through weekly elimination strategies you can implement in 20 minutes or less.