Everyone wanted to know which tool would fix it. Nobody had asked why it was breaking.

The platform went down every day. Not weekly, not under peak load — every day. The first hours of every morning belonged to recovery, and had for long enough that the team had stopped calling it an outage. It was just early morning noise after months of this behavior.

This was a production environment running about a hundred devices after over two years of development before I got involved. The plan on the wall was to start to scale it with a specific customer in mind that would increase the number of devices by at least twenty times that.

The comfortable diagnosis was that the platform was the problem. Replace it. Buy bigger. Add capacity or start the development over.

Buying feels like diagnosis. It has a budget line, a timeline, and a name you can say out loud in a board meeting.

The one that was actually blocking plans to scale was a timing flaw in how two steps of the workflow synchronized. One flaw. In the process, not the platform — a process problem wearing a technology costume.

It took some time to find. The firefighting had been going on far longer.

It wasn't the only thing standing between us and scale. There were real technical problems underneath it. But this was the one taking a piece of every single day — and as long as it did, nobody had the hours to go after the rest. Every other fix that mattered came after this one.

Same platform. Flaw fixed. We landed that customer and scaled the platform 24x in 6 months.

I think about that one constantly now, because I keep watching a version of it happen with AI.

When you layer AI onto a workflow that has a flaw in it, the AI does not find the flaw. It executes it — faster, at volume, and with more confidence than the humans ever had. Automation is an amplifier. It amplifies whatever is actually there.

The tools have never been better than they are now, and stall stories have never been more common. Those two facts are related.

Last series I wrote about the frameworks we'll need for AI. This one is about what actually happens on the floor.

Over the coming Mondays I want to walk through the stall patterns I've seen across thirty years of operations, and what they predict about AI. Each one stands alone. Together they're closer to a map.

So, before the question of which AI tool: what does your process do when it breaks?

Because that's the thing you're about to automate at scale.