Innovation
Insights · Innovation

Cleaning up your data is the easy part

Auditing your data before you buy an AI tool is the easy part. Keeping it clean is the half that decides whether any of it survives the year.

The advice to audit your data before buying an AI tool has largely converged. Scope it tight, get one slice clean, let the rest wait. That advice is half right, and the missing half decides whether any of it survives the year.

Clean is a verb

Pick your slice. Reconcile the records. Retire the duplicate spreadsheet. Get the one number everybody quotes to actually mean one thing. That’s a real project, and it’s worth doing.

Now it’s ninety days later. Someone in operations started tracking a new thing in a private sheet because the system didn’t have a field for it. A process changed, and the documentation didn’t. Somebody’s manual export is the real source for a report that a dashboard claims to own. Your clean slice has quietly gone stale, and the AI sitting on top of it is still answering with total confidence. A system that’s confidently wrong is worse than no system at all, because people believe it.

This is the failure we see most, and it doesn’t look like failure at first. It looks like a successful pilot that gets less useful every quarter until people stop trusting it and drift back to asking each other.

The reason is simple. A cleanup is a one-time event. The mess is continuous. If you don’t change how work actually flows, the mess reasserts itself faster than any cleanup can keep up, because people generate it by doing their jobs the way the work is currently set up.

So there are two steps, and the second one is the one that gets skipped.

Connect the data. Every system you want the AI to reason over must be reachable, and changes must be trackable.

Then set the policy that keeps it current. If a department is still running a private spreadsheet, that work never comes back into the system. Whatever you clean today decays at exactly the rate your processes leak.

Neither step is an AI project. Both are things a company should do anyway. AI can help you enforce the policy once it exists, and it can flag the drift early, but it can’t decide where work is allowed to live. Do the first without the second, and you’ve bought yourself about two good quarters.

Finding all of it takes a scoping engagement that starts with working logins to real systems. This piece is about the second step, because it decides whether the first one holds.

A Tuesday at our company

We run on this, so here’s what the far side of those two steps looks like.

A designer signs in. Before they’re asked anything, they have their ticket list, the three Discord threads that moved overnight, and a note that the client said one specific item was urgent on last week’s sprint call. Not because someone briefed them. Because the system read the transcript. It suggests where to start, and why.

They work. As they go, their output gets checked against what the client actually agreed “done” means on that ticket, which is written down somewhere they don’t have to go find. When they finish, the documentation is written, the ticket moves, the artifacts are linked, and the developer who implements their design signs in tomorrow for the same briefing on the other side of the handoff.

We’re a 20-person software firm in Indianapolis. We’re not a lab. We built this because the alternative was a company where the answer to most questions was “ask Jason or me,” and that’s a single point of failure with a pulse.

None of it works without the two steps above. The briefing is only as good as the systems it can reach, and it only stays good because we changed how the work flows into them.

Humans send the messages

That system writes documentation, moves tickets, assembles context, and drafts communication. It does not send anything. It can write the Slack post or the client email, and then a person reads it and sends it.

That’s deliberate. The returns here come from the assembly, the recall, and the documentation that quietly eats your team’s week. Judgment stays with the people who are accountable for it. Anyone selling you a system that runs unattended is selling you the part that fails quietly and expensively.

The related question, and the one worth answering before anything gets built, is who gets access to the central knowledge base. Leadership only is simple. Opening it to the whole team is harder to design and usually worth more.

Then decide what the system is never allowed to hold. Ours loads its rules at the start of every session. If a meeting transcript comes in that contains a performance discussion about an employee, that content is rejected before it’s ever written. It doesn’t enter the repository at all. You want those rules enforced at the write, not in a policy document that assumes everyone remembers.

The part that compounds

Skip the workflows, and you’ll repeat the cleanup every eighteen months, paying full price each time and losing a little more of your team’s faith in it. By the third round, the people who have to live with the system have learned that it goes stale, and that lesson is expensive to unteach.

Fix the workflows once, and the cleanup holds, because nothing is quietly leaking back out.

The companies that get this right won’t be the ones with the best model. They’ll be the ones whose data was still true on the day they asked.


Drew Linn is CEO of Counterpart, a custom software firm in Indianapolis. We connect the systems businesses run on and stay the long-term owner of the complicated ones. We run our own company on what’s described above, so if you want to see it in action before you decide, ask, and we’ll show you.

Talk to us

Bring us the hard version of your challenge.

Thirty minutes with a senior engineer, no pitch and no pressure. Whatever you decide, you'll leave with a clearer map of what you're about to do.

Start the conversation Keep reading