"\"We'll Clean the Data First\" Has No Stopping Condition"
"We'll clean the data first, then build the reporting." I've heard that sentence in most project kickoff meetings I've sat in, across e-commerce, automotive, marketing, and consumer brands. It sounds like the responsible path. Bad inputs make bad outputs. Get the foundation right before you build on it.
The logic is sound in isolation. In practice, I've almost never seen it survive contact with the second month.
Why the cleanup phase doesn't end
The problem is specific and structural: cleaning has no stopping condition.
Ask a team to clean the data and they will find work forever. Every dataset holds more defects than anyone will ever care about. Nulls in optional fields. Inconsistent casing in categorical columns. Duplicates that turn out to be legitimate on closer inspection. Date formats that vary across sources. Address fields with abbreviations in one system and full words in another.
Each of those is a real defect. Most of them change nothing. Without a downstream consumer telling you which defects alter an answer someone needs, you are sorting an infinite list by no criterion. The team works diligently through the backlog, and six months later the data is cleaner, no decision has been informed, and the project gets quietly defunded. Not because anyone made a bad decision. Because the scope was unbounded from day one and nobody noticed until the budget ran out.
I've seen this pattern three times at companies I worked for, and in each case the cleanup phase consumed between 40% and 60% of the total project budget before anyone built a report. In two of the three, the project was cancelled before a single dashboard shipped.
The stopping condition is downstream
Reporting supplies the stopping condition that the cleanup phase lacks.
Build the thinnest possible reporting for one decision somebody is currently making badly, and the defects that matter announce themselves immediately. A number comes out wrong. A chart breaks. A filter excludes rows it shouldn't. The person who reads the report says, "that's not right," and now you know exactly which defect to fix, because it produced a wrong answer in front of someone who cared.
The rest of the defects, the inconsistent casing and the duplicate addresses and the nulls in the optional fields, can wait. Some of them can wait indefinitely, which is the correct amount of attention for a defect that changes nothing.
This is not a new insight. Manufacturing figured it out decades ago: build the process to produce good output rather than inspecting every input before the line starts. W. Edwards Deming made the argument in Out of the Crisis in 1986, and Toyota's production system applied it through the 1950s and 1960s by pulling quality problems to the surface at the point they affected assembly.
A 2022 survey by dbt Labs found that analytics engineers spend roughly 44% of their time on data quality and pipeline maintenance. The cleanup is not a phase you finish. It is ongoing work, and the question is whether it is guided by a live consumer or by a backlog nobody prioritized.
The order that works
Pick one decision somebody is currently making badly, whether that means slowly, on gut feel, or with numbers nobody trusts. Fix only what breaks along the way. Write down every definition you resolve while you're in there, because those definitions will be invisible again in three months if they live only in the analyst's head.
Then take the next decision.
The first round usually surfaces three or four real problems: a join that drops rows, a field that maps two meanings to one code, a filter that silently excludes a segment. Those three or four problems are worth fixing. They're the defects that change an answer someone reads. Everything else on the infinite cleanup backlog ranks below them.
What it looks like when teams run it backwards
The backwards order, clean then build, creates a second problem beyond the budget waste. It trains the organization to expect a long wait before they see anything. Stakeholders learn that a data project means months of invisible work followed by a reveal. They disengage. By the time the cleaned data produces a report, the business question has moved, the sponsor has rotated, and the report answers a question nobody is asking anymore.
The forwards order, build then clean, inverts that dynamic. Stakeholders see something in the first two weeks. It's rough. Some numbers are wrong. But the conversation shifts from "when will we see something" to "this number looks off, can you fix it," which is exactly the conversation that produces a clean dataset, one real problem at a time.
I've run the backwards version and the forwards version. The backwards version always felt more professional at the start and always ran out of sponsorship before it delivered. The forwards version always felt scrappy at the start and always shipped.
The difference is finishing
This approach looks like cutting corners and it isn't. It's the difference between cleaning a house and cleaning the room you're about to have guests in. The first one is admirable and unbounded. The second one finishes, and the guests come, and the evening happens.
Every data project that ships has a stopping condition. Clean-first projects make that condition "when the data is clean," which is never. Build-first projects make that condition "when the decision is informed," which is Thursday.
Start with Thursday.