Skip to content
Book a free consult
Interface sounds
By Kishan Thankey 6 min read StrategyDecision MakingAdoption

Is Your Data Ready for AI? The Audit to Run Before Any Build

More AI projects stall on data than on models, but 'get our data ready' does not mean what most teams fear. A six-question scorecard tells you whether the data behind your first workflow is ready, and what to fix if it is not.

Scattered data sources, a shared drive, an inbox, a spreadsheet, and a legacy system, funneling down into one small clean scope labeled one workflow's data.
Contents

Ask a room of founders why their AI project stalled and almost nobody says “the model was not smart enough.” The honest answers sound like: we could not agree which spreadsheet was the real one, the documents were three years stale, nobody knew who was allowed to see what. The model showed up ready. The data did not.

That failure has produced an equally expensive overreaction: teams that will not start until they have “fixed the data,” a project with no finish line. Both mistakes come from the same misunderstanding of what ready actually means.

”Ready” is scoped to a workflow, not a company

The question is never “is our data ready for AI.” It is “is the data behind this workflow ready,” and that changes everything about the size of the job.

An assistant that answers support questions from your policy documents needs those policy documents: current, deduplicated, and permissioned. It does not need your CRM cleaned, your legacy database migrated, or your data warehouse unified. Those might matter for workflow five. They are not the price of workflow one.

Company-wide data cleanup drawn as a huge mountain that takes years, next to one workflow's data drawn as a small step that takes weeks, with an arrow labeled scope it down.

This is the same discipline as picking your first workflow in the first place: narrow wins. A scoped readiness job takes weeks and ends. A company-wide one takes years and does not.

What “ready” actually means

Four properties do almost all of the work. For the one workflow you have in mind:

  • It exists digitally, and it is findable. Not “somewhere in the shared drive.” A named set of sources: these folders, this table, this system. If finding the data is itself a research project, that is the first fix.
  • It is fresh enough to trust. An AI assistant grounded in your 2023 pricing sheet will confidently quote 2023 prices. Staleness does not make AI fail loudly, it makes it wrong quietly, which is worse.
  • You know who is allowed to see it. If the source documents have access rules, the assistant needs to respect them from day one. Untangling permissions after launch is far more painful than mapping them before.
  • You have examples of right answers. A handful of real questions with answers you know are correct. This is your ground truth, and it is what turns “seems to work” into a quality number you can actually watch.

Notice what is not on the list: perfect structure. Modern AI reads unstructured data directly, the emails and PDFs and notes where most company knowledge actually lives. Messy format is fine. Contradictory copies are not: nine versions of the same policy with no marked winner will make any assistant contradict itself, and no model upgrade fixes that.

AI does not need your data to be tidy. It needs your data to be true, and it needs to know which copy to believe.

Score the workflow you have in mind, honestly:

Scorecard

How ready is the data behind your first AI workflow?

Answer for one specific workflow, not for the company.

0 of 6 checked

What the audit buys you

Findable sources A marked-authoritative version Known access rules Ten right answers

One workflow's data, audited

Weeks of scoped work, not a company-wide cleanup

An assistant that tells the truth about this corner of the business

The pieces converge as you scroll: four unglamorous properties, one workflow, and readiness stops being a mountain.

The fixes are smaller than they look

Each unchecked box has a boringly practical fix, and none of them is a platform project.

Data trapped on paper or in one person’s head gets captured by the workflow itself: a one-hour interview written up, a scanning pass, a template people fill from now on. Scattered sources get a named home: not moved, just listed, “these three folders and this table are the sources.” Duplicate versions get a decision: someone with authority spends an afternoon marking winners and archiving the rest. Staleness gets an owner: updating the source becomes part of a named person’s job, because data with no owner rots back within a quarter. And missing ground truth gets an hour of honesty: ten real questions, ten right answers, written down.

That list is the real “data preparation for AI.” Not a lake, not a warehouse migration, not a year of cleanup. Days to weeks of focused work on one corner, and every later workflow inherits part of it.

Where this goes next

Once one workflow’s data is ready and an assistant is answering from it with sources attached, something useful happens culturally: teams see what ready data buys, and the next corner gets cleaned because people want the next assistant. Readiness stops being an initiative and becomes a habit, funded by results instead of patience.

That is the honest sequencing: not “fix the data, then do AI,” and not “ignore the data, the model will cope.” Fix the corner you are building on. Build. Repeat.


Not sure whether your data can support the workflow you have in mind? That is a one-conversation diagnosis. Book a free consult and we will run this exact audit against your first workflow together, and map the shortest path from where your data is to where it needs to be.

Frequently asked questions

How do I know if my data is ready for AI?

Check four things for the one workflow you want to improve: the data exists digitally and is findable, it is fresh enough to trust, you know who is allowed to see it, and you have real examples of what a right answer looks like. If those hold for that single workflow, you are ready to start, whatever state the rest of the company's data is in.

Do we need to clean up all our data before starting with AI?

No, and trying to is the most expensive way to fail. Readiness is scoped per workflow, not per company. An assistant that answers policy questions needs the policy documents in order, not your entire data estate. Clean the corner you are building on, ship, and let each project pay for the next corner.

What kind of data does AI actually need?

Less than most teams assume, and messier than most teams assume is allowed. Modern AI reads unstructured data, emails, PDFs, notes, chat logs, directly. What it cannot survive is data that is stale, contradictory across copies, or so scattered that nobody can say which version is true.

What is the most common data problem that blocks AI projects?

Duplicates and staleness, not format. Nine versions of the same policy document with no marked winner will make any AI assistant contradict itself, no matter how good the model is. Picking the authoritative source for one workflow is usually days of work, and it is the highest-leverage data cleanup you can do.

Found this useful?

Share this with your network on LinkedIn, it helps more than you think.

Enjoyed this read? Get the next one in your inbox.

When we publish something worth your time, you will be first to know. No spam, unsubscribe anytime.

Keep reading

Several candidate workflows with one chosen as the highlighted first place to start with AI.

Where to Start With AI: How to Pick Your First Workflow

The hardest part of AI is not the model, it is choosing where to begin. A simple way to pick a first workflow that is high-value, low-risk, and proves itself fast.

Read article
A software box with its old workflow-automation label crossed out and a shiny AI AGENT sticker slapped on, while an inspection panel reveals the same fixed if-then rules inside.

Agent Washing: How to Tell a Real AI Agent From an Automation With a New Sticker

Every product renamed itself an agent this year, and the word stopped carrying information. A five-scenario quiz trains your eye, and five procurement questions expose what a vendor actually built, because the label decides the price, the failure modes, and the oversight you owe it.

Read article
A dial with three zones: automate it on the left for low-stakes reversible rules, put a copilot on it in the middle where AI drafts and a person approves, and keep the decision human on the right where stakes are high and the action is hard to undo.
Decision MakingTrust

What You Should Refuse to Automate

The hype says automate everything. The discipline is knowing where the line goes. A task belongs to a person, not a model, when it is hard to undo, needs real judgment, or puts money, health, a job, or safety at stake. Grade any task on those three axes here, and see where the boundary actually falls.

Read article

Have software that should be smarter?

Let’s map a free AI-transformation roadmap for your product.