Every mid-size manufacturer I talk to has the same reason for not starting. The data isn't clean.

The ship dates in the ERP don't match the carrier portal. The same customer is spelled three ways. Half the item master hasn't been touched since the last migration. So the automation project waits — behind a data-cleanup initiative, an ERP re-implementation, a “let's get our house in order first.” The wait is the trap. The cleanup never finishes, and the automation never starts.

Data cleanup isn't a project with a finish line. Perfect data is a moving target you will never hit.


01Pilot Purgatory Has a Cause

Seventy-three percent of mid-market manufacturers are stuck in AI pilot mode, and the reason cited most often is the data. Only 27% have a data warehouse or data lake, less than half the rate of the broader mid-market. Forty-five percent still run on siloed data. Look at those numbers and the obvious move is to fix the data first, then automate. The firm that ran that survey recommends exactly that: audit the data, clean it, then bring in AI.

I read the same numbers the other way. A 73% pilot-stall rate isn't proof the data wasn't clean enough yet. It's proof that “clean it first” is the plan that keeps you in the pilot.

The manufacturers who wait for clean data don't get clean data. They get another year in the pilot. Data preparation is already the single largest time sink in an AI project — by most estimates half to three-quarters of the effort. Run as a standalone phase before the workflow, it has no natural finish line. There's always another field, another exception, another system that doesn't agree with the others. You can spend two years on it and still not be ready, because ready was never a real state.


02Clean Enough for What?

“Clean data” is a meaningless standard until you attach it to a job. Clean enough for what? Your entire item master doesn't need to be right to monitor on-time delivery. You need the open orders, their promise dates, and their ship status. That's it. The other 95% of the mess has nothing to do with that one workflow.

This is the distinction that gets lost. Reliable data matters — a workflow built on numbers nobody trusts will fail. But reliable enough to act on, for one specific process, is a far lower bar than enterprise-clean. Ninety-eight percent of manufacturers are exploring AI; 20% feel fully prepared to use it at scale. Some of that gap is real readiness — fragmented systems, data nobody trusts. But some of it is the bar itself: “at scale” is the wrong thing to clear first. The teams that get moving aren't the ones with cleaner data. They picked a job small enough that the data was already good enough.


03What This Looks Like on OTIF

Take delivery performance. You're tracking OTIF on a spreadsheet somebody updates Friday afternoon. You know it's costing you. Walmart alone charges 3% of the cost of goods on orders that miss the window, and every expedite to recover a slip runs two to three times normal freight, more in a crunch. So you want to automate the monitoring, and you're told the data's too dirty: the ERP ship dates lag, the carrier feed doesn't reconcile, the customer records don't map.

Here's the move. Build the monitor anyway, on the data you have. Pull the open orders, flag the ones trending late, put them in front of a person the same morning instead of the following Friday. It won't be perfect on day one. But inside two weeks it will have told you exactly which data problems actually cause missed deliveries, and it's a short list — a handful of customer mappings and one lagging status field, not the whole item master. You fix those, because now you know they're the ones that cost you. The automation didn't wait for clean data. It found the data worth cleaning.

Automation isn't what you do after the data is clean. It's the fastest way to find out which data was ever worth cleaning.


04The Cleanup Loop That Works

There's an honest caveat, and it cuts the other way. Dirty data does degrade AI, more at scale than in a pilot. A model reading across ten thousand records hits edge cases a demo never surfaces. So this isn't “ignore your data.” It's the opposite. Scope the workflow tight, validate the output, and keep a person on the exceptions. That's how you run on imperfect data without getting burned by it.

What kills you isn't the dirty data. It's fixing it by hand. When the correction loop runs through email threads and someone's follow-up, the mess grows faster than you close it. When the workflow itself flags the bad record, routes it to an owner, and tracks the fix, data quality climbs as a byproduct of running the process. The automation becomes the cleanup mechanism. That's the inversion. You were treating clean data as the entry ticket. It's the output.

The manufacturers pulling ahead in 2026 aren't the ones with the cleanest ERPs. They're the ones who stopped waiting to have one.


What is your operation waiting to fix before it starts — and what is the waiting already costing you?