Field Notes AI EXPLAINED August 2026

Do I need to clean up my data before starting an AI project?

No, and the belief that you do is probably the single most expensive misconception in business AI right now, because it keeps organisations in preparation mode for years while the problem they wanted to solve keeps costing them money every week.

Here is the thing the "data readiness" conversation misses. The businesses that get value from AI did not start with clean data. They started with a specific problem, and the AI was built to handle their data as it actually is, inconsistent formats, missing fields, information living in PDFs and email threads and one long-serving employee's memory. Handling that mess is not a barrier to the project. Handling that mess is the project.

Why modern AI is different from the automation that came before it

This is also where modern AI genuinely differs from the automation tools that came before it. Traditional rule-based automation needed clean, predictable inputs because a rule breaks the moment reality does not match it. Current AI is the opposite: it is specifically good at reading variable, messy, human-shaped information and making sensible decisions about it. Documents in different formats, customers who spell their own names three different ways, forms filled in with varying levels of care. A system built on rules chokes on that. An AI layer reads it, normalises it, and flags the genuinely ambiguous cases for a person to look at.

One of our clients processes building plans, which arrive as drawings produced by dozens of different architects with no standard format at all. There was no version of that project where the data got cleaned first, because the messiness was permanent and incoming. The platform was built to read the mess, and that is exactly why it was worth building.

What you actually need before starting

What you actually need before starting is much smaller than a data cleanup. You need someone who can describe how the process works today, including the exceptions and the unwritten rules, because scoping a build against a process nobody can articulate is where projects genuinely stall. And you need access to a reasonable sample of the real inputs, the actual documents and records the system will handle, warts included, because the AI should be built and tested against reality rather than an idealised version of it.

The one caveat

There is one caveat. If your core records are so fragmented that even a human cannot tell which version of a customer or job is correct, that specific confusion needs resolving for the affected process, though even then the fix is usually part of the build rather than a prerequisite project. What you should refuse to accept is the general-purpose "we need to sort out our data first" position, which I have written about before as one of the two great delay machines in business AI, because it is technically defensible, practically infinite, and mostly functions as a reason to never start.

Questions we get asked

How much sample data does a build actually need? Usually far less than people expect. A few dozen representative examples of the documents or records involved, including the ugly ones, is enough to scope and test most workflow builds.

What if our data is spread across systems that do not talk to each other? That is not a blocker, it is the most common starting condition we see, and connecting those systems is typically the core of the value.

Is there any data preparation worth doing before we talk to anyone? Gather examples and note where things live. Do not pay anyone to clean or migrate anything until the project that will use it is scoped, or you will clean the wrong things.

If messy data has been the reason you have not started, let's look at what is actually possible.