Data Readiness Assessment for AI: Check the Fields, Not the Database
Avolis Research Group
·
·
22 min read
Data readiness assessment for AI: none of 17 ranking pages gave a sample size and pass mark for testing your records. Here's one, sized to one workflow.
A data readiness assessment for AI checks whether the information one specific workflow needs is recorded, reachable, accurate, consistent, and current enough for the kind of AI you plan to use. That last part is the one most assessments skip. An AI that drafts a proposal from inspection notes needs very different data from a model that predicts which bids you'll win, and the same records can be ready for the first and nowhere near ready for the second. So a useful assessment starts by naming the workflow, then checks only the fields that workflow touches, on a sample of real records, and ends with a list of fixes sized in hours, not a quote for cleaning the whole database.
In September 2026 we read the 17 pages ranking for "data readiness assessment services for ai initiatives," "ai data readiness assessment," and two close variants. Most of their publishers sell data platforms, data engineering, or cleanup work. Seven judged data against a specific use case, but only three set the use case before looking at the data, and three others did the reverse. Two separated what different kinds of AI need, and none gave a sample size and a pass mark you could apply to your own records.
This page fills those gaps for a business of 10 to 200 people. It's the data section of our guide to AI readiness assessment services, which covers the other five areas.
Key Takeaways
- Data is only "ready" for a defined task. Researcher Neil Lawrence put it plainly in 2017: a data set can reach the top band of readiness only "once a task is defined."
- Of 17 pages ranking for data readiness assessment services in September 2026, none gave a sample size and a pass mark you could run on your own records, and only 3 named the use case before assessing the data.
- What counts as ready depends on what the AI does. Drafting from documents needs readable, current inputs, while prediction needs years of outcomes recorded the same way.
- In a 2025 survey of 565 data professionals, only 12% said their data was of sufficient quality and accessibility for AI.
- In a 2026 survey of 505 data leaders, 88% said they had the data readiness AI needs, yet 43% named data readiness as a top challenge.
- In 75 measurements where managers checked the last 100 records their departments created, 47% of new records had at least one critical error, and only 3% of scores met even the loosest acceptable standard.
Table of Contents
- What a data readiness assessment is
- Why "clean your data first" is the wrong first step
- What the data needs depends on what the AI does
- The five checks, field by field
- Which data gaps block a workflow
- What a data readiness assessment service should deliver
- A worked example: a 55-person glazing contractor
- Where Avolis fits
- Frequently Asked Questions
- Continue Learning
What Is a Data Readiness Assessment for AI?
A data readiness assessment for AI is a check of whether the records a specific AI use would rely on exist, can be reached, and are good enough for that use. It isn't an audit of every system you own. Its output is a short list of the fields that fall short for that workflow, what it takes to fix each one, and whether the workflow can go ahead in the meantime.
The idea that readiness depends on the job has a long history in data science. In a 2017 paper, the machine-learning researcher Neil Lawrence proposed three bands of what he called data readiness levels. The lowest, Band C, "is about the accessibility of a data set," and it starts with what he called hearsay data: data someone believes exists, signaled by lines like "The sales department should have a record of that." Band B "is about the faithfulness and representation of the data," meaning whether what's recorded matches what it claims to record. Band A "is about data in context," and he was explicit that "a data set can only be considered in Band A once a task is defined" (Lawrence, "Data Readiness Levels", arXiv, May 2017).
The formal standards say much the same thing. The international data quality model, ISO/IEC 25012, defines data quality through 15 characteristics, among them accuracy, completeness, consistency, credibility, and currentness, and frames each one in terms of data "used under specified conditions" (ISO 25000 portal, ISO/IEC 25012). Even Gartner, whose figures are widely quoted to argue that data is the problem, says AI-ready data "must be representative of the use case" (Gartner, "Lack of AI-Ready Data Puts AI Projects at Risk", February 2025).
In other words, "is our data ready for AI?" has no answer until you say which AI, doing which job. The question an operations business can actually answer is narrower: is the data this one workflow needs good enough for the thing we want AI to do in it?
Why "Clean Your Data First" Is the Wrong First Step
Most businesses' data has problems, which is exactly why cleaning it first is a poor place to start. If every data set fails a general inspection, then a general inspection can't tell you where to spend, and the useful question becomes which problems sit in the fields your first workflow depends on.
The surveys agree that data is rarely in good shape. Gartner found that 63% of organizations "either do not have or are unsure if they have the right data management practices for AI," in a survey of 1,203 data management leaders in July 2024, and it predicts that through 2026 organizations will abandon 60% of AI projects unsupported by AI-ready data (Gartner, February 2025). In a survey of 565 data and analytics professionals by Drexel University's LeBow College of Business and Precisely, only 12% said their data was of sufficient quality and accessibility for AI. That sample wasn't only big companies: 30% of respondents worked at organizations with under 250 employees (LeBow and Precisely, 2025 Outlook: Data Integrity Trends and Insights, September 2024).
The most practical evidence comes from a much simpler exercise. Data quality consultant Thomas Redman and two co-authors had managers in their classes pull 10 to 15 critical data attributes for the last 100 units of work their department completed, mark the obvious errors, and count the records with none. Across 75 of those measurements, "47% of newly-created data records have at least one critical" error, a quarter of the scores fell below 30 out of 100, and "only 3%" met the loosest acceptable standard. The managers themselves had never considered anything below the "high nineties" acceptable (Nagle, Redman, and Sammon, "Only 3% of Companies' Data Meets Basic Quality Standards", Harvard Business Review, September 2017).
Here's the idea in simple terms: if nearly everyone's records fail, "our data isn't perfect" tells you nothing, because it's true of your competitors too. What it does tell you is that a company-wide cleanup has no natural end. The studies also show why ignoring data is expensive. In interviews with 53 AI practitioners, Google researchers found that 92% had run into at least one "data cascade," a data problem that went unnoticed early and caused compounding damage later (Sambasivan and others, "Everyone wants to do the model work, not the data work", CHI 2021). Those were high-stakes projects in areas such as health, conservation, and lending, so the lesson for a contractor isn't the 92% itself. It's that data problems stay hidden until someone checks the specific fields a system relies on, which is one of the recurring patterns in why AI projects fail.
The self-reported numbers can't be trusted on their own, either. The 2026 edition of the LeBow and Precisely survey, this time of 505 senior data leaders at companies with at least 1,000 employees or $250 million in revenue, found what its authors called a "confidence-reality gap": 88% said they had the data readiness they needed for AI, while 43% listed data readiness among their biggest AI challenges (LeBow and Precisely, 2026 State of Data Integrity and AI Readiness, January 2026).
Our reading of the data: these are big-company numbers, and they still show that the people closest to the data can't tell from the inside whether it's ready. At 10 to 200 people the fix is cheaper, not harder. You don't need a survey or a data team. You need someone to pull 20 real records for the one workflow you care about and check them against the paperwork, which is what the rest of this page sets out.
What the Data Needs Depends on What the AI Does
Different kinds of AI lean on different parts of your data, so the same records can be ready for one use and not another. A model that drafts from documents mostly needs its inputs to be readable and current. A model that predicts needs a long history of outcomes recorded the same way every time. Most assessments treat "data" as one thing, and of the 17 pages we read, only 2 drew this distinction at all.
Five kinds of use cover most of what an operations business would build, and each puts its weight on different checks:
| What the AI does | Example in an operations business | What its data has to be | Usually doesn't matter much | Typical blocker |
|---|---|---|---|---|
| Drafts from your documents | Repair proposals from inspection notes and photos; follow-up emails from a quote | Readable, current, and the right source (the latest price list, not last year's) | Spelling variants, free text, some missing fields a person fills in | An out-of-date reference file the AI will copy from faithfully |
| Pulls fields out of documents | Reading supplier invoices, work orders, or permits into the job system | Documents that arrive consistently and a clear list of fields to extract | History; the data doesn't need to exist in a system yet, since this creates it | No agreed place for the extracted fields to land |
| Sorts and routes | Tagging incoming service requests by trade and urgency | Categories that are defined and used the same way by everyone | Volume of history, once a few hundred examples are labeled | Categories that mean different things to different people |
| Answers questions from your records | A crew lead asks which warranty covers a roof, and the AI answers from the files | One current version of each document, and permission to share it with whoever's asking | Formatting | Three versions of the same policy, and nobody knows which is live |
| Predicts or forecasts | Which bids you'll win; which jobs will run over; next month's material needs | Past outcomes recorded every time, the same way, for long enough to see a pattern | How the records look | The outcome was never recorded, so there's nothing to learn from |
The table is our framework, not a published standard, but it's consistent with Lawrence's bands. Every use needs a defined task, which is what Band A asks for. What differs is how much faithful history each one needs. Drafting and extraction need the current inputs to be reachable (Band C) and a few fields to be accurate (Band B), while prediction needs months or years of outcomes recorded faithfully and the same way before the question it's meant to answer can be settled. That difference matters to a small operation because the first two kinds are where most of the early payoff is, and they're also the kinds whose data is usually closest to ready.
It's also why emails, PDFs, and photos aren't the problem many assessments treat them as. One ranking page listed "Data locked in PDFs, email threads, or systems with no API" as a common readiness failure. For a model that drafts or extracts, those documents are the input, and what matters is whether they're the right, current ones. For a model that predicts, they may well be a problem, because a forecast needs the outcome in a field it can count.
Citation-ready summary: Data readiness for AI depends on the task. AI that drafts or extracts from documents needs readable, current inputs and can work from emails, PDFs, and photos. AI that predicts needs a long history of outcomes recorded consistently. Neil Lawrence's 2017 data readiness levels put it this way: data reaches the top band only "once a task is defined."
The Five Checks, Field by Field
Once you've named the workflow and the kind of AI, list the handful of fields it needs, usually five to ten, and run five checks on each: is it captured, can it be reached and used, is it accurate, is it consistent, and is it current. Run them on a sample of real, recent records, not on anyone's description of the system.
- Captured. Is the information recorded in a system at all, or does it live in someone's inbox, notebook, or memory? This is Lawrence's hearsay test. Ask where the last record is, then go and look. Often the gap turns out to be a single field nobody has needed to record before, as the example in our data template shows.
- Reachable and permitted. Can the data be exported or read by another system, and are you allowed to use it this way? Check both the export (a CSV, an API, a report you can schedule) and the rules. Customer contracts, insurer requirements, and privacy law can all limit where personal or client data goes. Lawrence put legal and ethical constraints in his lowest band for a reason: if you can't use the data, nothing else about it matters.
- Accurate. Does the record match what actually happened? This is the check most assessments describe and few run. Take the same sample and compare each field against the source of truth, whether that's the signed contract, the invoice, the photo, or the measurement report. Redman's version counts a record as good only if every critical field is right, which is the standard to use here.
- Consistent. Is the same thing recorded the same way every time? "IGU," "insulated unit," and "dbl pane" are one glass type to a person and three to a report. This matters a lot for sorting and prediction, and much less for drafting, where the AI reads the words the way a person would.
- Current. Is it updated when reality changes? Price lists, labor rates, contact details, and policy documents go stale quietly. An AI that drafts from a stale price list will copy the stale price every time, with complete confidence.
Where the data lives decides the first two checks, and at this size the money side is usually in better shape than the job side. In the UK government's 2024 survey of 8,396 small and medium-sized employers, 88% of small businesses (10 to 49 employees) and 92% of medium-sized ones (50 to 249) that use business software had accounting software, but only 42% and 54% used a customer relationship management system, the kind of software that holds customer and job history. Among construction firms of all sizes up to 249 employees, only 58% used business software for anything (Department for Business and Trade, Longitudinal Small Business Survey 2024: SME employers, September 2025).
That's a UK survey, and we didn't find a US equivalent broken out by firm size, but the pattern is the one the first check is designed to catch. The invoice total is in the accounting system and exports cleanly. The reason the job ran over, or which customer prefers a call to an email, can easily be sitting in an inbox, a spreadsheet, or someone's head.
The simplest version of the first checks is to count how many of your last 20 records have each field filled in, and our data readiness template is set up for exactly that. Accuracy takes one more step. For each of the same 20 records, open the source document and mark any field that doesn't match. If you check six fields on 20 records, that's 120 comparisons, and at roughly half a minute each, by our estimate, it's an hour's work for the person who knows the paperwork best.
Set the pass mark before you count, and tie it to what happens when a field is wrong. For fields that reach a customer with no person checking first, such as a price, a date, or an address in an automatic email, hold to the standard the managers in Redman's study set for themselves, the high nineties, which on a 20-record sample means all 20. Where a person reviews every draft before it goes out, or where a field only steers the work internally, like a job-type tag used to route requests, a lower bar on the inputs can be fine, because the review catches what the sample misses. Either way, the number that matters is the error rate in the fields the AI will actually use, measured on your own records, against a bar you set in advance.
This is also where the cost of skipping the check shows up. Redman's "rule of ten" holds that work costs ten times as much to complete when its data is flawed as when it's right. His worked example: 100 tasks at $1 each cost $100 with perfect data, and $199 when 11 of them carry an error (Nagle, Redman, and Sammon). An AI that works from those records doesn't remove the rework. It produces it faster, which is why the accuracy check comes before the build, not after the first complaint.
Which Data Gaps Block a Workflow
Not every gap you find is a reason to wait. Sort each one by what it does to the workflow in question: some block it outright, some can be fixed while the build goes ahead, and some don't matter for this use. It's the same three-way sort our infrastructure assessment applies to systems, and our methodology page explains why a single blocker shouldn't be averaged away.
| What you find | Drafting or extraction | Prediction | What to do |
|---|---|---|---|
| A needed field is never recorded | Fix alongside if a person can supply it at review | Blocker until there's history | Add the field now; for prediction, set a date to revisit |
| Errors above your pass mark in a field that goes to customers | Blocker | Blocker | Fix the source process first, not the historical records |
| The same thing is written several ways | Not a blocker in most cases | Blocker | Add a dropdown going forward; map old variants only if prediction needs them |
| A reference file (prices, rates, policies) is out of date | Blocker, because the AI will copy it | Fix alongside | Update it and name an owner who updates it on a schedule |
| The data can't be exported | Fix alongside if documents can be read directly | Blocker | Check for a report or export in the software you already pay for |
| You aren't sure you're allowed to use it this way | Blocker | Blocker | Read the customer contract or policy, and ask before building |
Notice that the fix is almost always to the process that creates the data, going forward, and rarely to the backlog. Correcting three years of old records helps only if a model needs three years of history. For drafting and extraction it's usually wasted effort, because the AI works on the next record, not the last thousand.
What a Data Readiness Assessment Service Should Deliver
A data readiness assessment service should hand you a finding for each workflow in scope: the fields that workflow needs, where each one lives, the error rate measured on a sample, a blocker, fix-alongside, or not-a-blocker call on each gap, and the hours each fix will take. If what you get is a maturity score, a platform recommendation, or a proposal to clean everything, you've bought something else.
It helps to know who's selling. Of the 17 pages we read, about 14 were published by firms that sell data platforms, data engineering, or cleanup work, and several made the sequence plain. Three others went the opposite way and audited the data before any use case was chosen. One mid-market provider's page says "most clients move into a foundation build within 60 days" of the readout. None of that is wrong in itself. A 2,000-person company with dozens of AI projects may well need a data platform. But at 10 to 200 people, an assessment that ends in a foundation build deserves a hard look at whether your first workflow actually needs one.
Among the pages we read, published durations ran from two to eight weeks, and the page that scoped most narrowly, to one or two use cases, quoted two to four weeks. Who offers these assessments, what they tend to cost, and the free and subsidized routes are compared on our AI readiness assessment consulting firms page, so we won't repeat that here.
Before you sign, a few questions will tell you which kind of assessment you're buying:
- Which workflow will you assess the data for? If the answer is "all of it," the scope is the whole company, and the price and timeline will follow.
- Will you check records against the source documents, or review how the systems are set up? Only the first tells you whether the data is accurate.
- What sample size and pass mark will you use? A good answer names both. Twenty to a hundred records per workflow is a reasonable range.
- Will you tell us what doesn't need fixing? At this size, that list can be longer and more useful than the list of what does.
- Do you sell the fix? Many good firms do. It's worth knowing, so you can weigh the recommendations accordingly.
If a partner can't answer the first question, it can't answer the rest, and our guide to which AI consulting company to choose covers the wider partner decision.
A Worked Example: A 55-Person Glazing Contractor
This is an illustrative composite, not a client, and its figures are examples rather than benchmarks. Picture a 55-person commercial glazing contractor doing storefronts, curtain wall, and service repairs, with about 40 installers in the field, four estimators who also run site surveys, three project managers, and an office of eight. The owner has two AI ideas and has been told by a software vendor that the company needs a data cleanup before it can do either.
The first idea is drafting service repair quotes. The estimators run about 25 site surveys a week for broken glass, failed seals, and damaged frames, and each one turns into a quote that takes around 45 minutes to write from the survey notes, the photos, the measurement sheet, and the unit price list. The second is predicting which project bids are worth chasing. These are the new-construction and tenant-improvement bids, a separate stream from the service quotes, and the estimators spend hours on ones that go nowhere.
Checked field by field, the two ideas come out very differently. For the service quotes, the estimator who knew the paperwork best pulled the last 20 surveys and compared the key fields against the photos and measurement sheets, which took a little over an hour. They set the pass mark first. An estimator reviews every draft before it goes out, so for the survey fields a lower bar of 19 of 20 was enough, and the prices would be checked separately, against the supplier's current price sheet, since those go to the customer as written. The opening sizes matched on all 20, and the damage notes matched the photos on 19, with the one miss the kind a reviewer would catch. The glass type was typed six different ways, which didn't matter, because the AI reads it the way a person would.
The real finding was the price list. It hadn't been updated in 14 months, and 9 of its 60 line items no longer covered what the company's supplier now charges. Nine wrong prices out of 60 is nowhere near the high nineties, and those prices would reach the customer as written.
The bid-prediction idea failed on the first check. The customer database held about 1,400 project bids from three years, but won or lost was recorded on only about 520 of them, and the reason for a loss on almost none. There was nothing yet for a model to learn from.
| Workflow | Kind of AI | Key finding | Call | Owner and next step |
|---|---|---|---|---|
| Service repair quotes from site surveys | Drafts from documents | Survey fields at 20 and 19 of 20, above the reviewed-draft bar; price list 14 months stale | Go, once the price list is updated (about an afternoon) | Senior estimator updates prices and owns quarterly updates |
| Which project bids to chase | Predicts | Outcome recorded on about 37% of project bids; loss reason almost never | Wait. Make won/lost and a loss reason required fields now | Office manager; revisit in 12 months |
| The vendor's company-wide cleanup | None | Would correct three years of records no current workflow needs | No | Nobody |
The go, wait, and no calls follow the decision rules on our services page. The vendor wasn't entirely wrong about the data, which did have problems. But the cleanup it proposed would have fixed old records that no current workflow uses, and it would have left the one problem that mattered, the stale price list, untouched.
What the example shows: the fix for the prediction idea costs almost nothing, adding two required fields, but it takes a year to pay off, so it's worth starting now even though nothing gets built this year. And the honest answer on bids may turn out to be simpler than AI. With a few hundred outcomes a year, a monthly report of win rates by customer type might tell the estimators most of what a model would.
Where Avolis Fits
The data checks on this page are part of our two-week diagnostic, run alongside measuring the workflows themselves, rather than a separate data service. We work through each workflow with the people who run it, pull a sample of real records, check the fields that workflow needs against the source documents, and sort what we find into blockers, fixes to make alongside the build, and things that don't need fixing. Every engagement starts with that diagnostic, and the ranked result is yours whatever you decide to do next.
If you'd rather run the checks yourself, the data template in our readiness checklist covers completeness, and the accuracy check above adds an hour. If the question behind this one is whether it's worth paying for any assessment at your size, our AI readiness assessment for SMBs page covers that.
Frequently Asked Questions
What is a data readiness assessment for AI?
A data readiness assessment for AI checks whether the records a specific AI use would rely on are captured, reachable, accurate, consistent, and current enough for that use. A good one names the workflow first, tests a sample of real records against the source documents, and sorts each gap into a blocker, a fix to make alongside the build, or not a problem.
What does an AI data readiness assessment check?
It checks five things for each field a workflow needs: whether the information is recorded in a system, whether it can be exported and legally used, whether it matches the source documents, whether it's recorded the same way each time, and whether it's kept up to date. Which checks matter most depends on whether the AI drafts, extracts, sorts, answers, or predicts.
Do I need to clean my data before using AI?
Usually not all of it. Surveys show most companies' data has errors, so a company-wide cleanup has no natural end. Fix the fields your first workflow depends on, and fix the process that creates them going forward. Cleaning old records only pays off when a model needs that history, as prediction does and drafting usually doesn't.
How long does a data readiness assessment take?
Of the pages we reviewed in September 2026, published assessment durations ran from two to eight weeks, with narrower scopes at the short end. For one workflow at a business of 10 to 200 people, the checks themselves are smaller: by our estimate, about an hour to count and compare 20 records, plus the time to list the fields and agree on a pass mark.
Can AI use data from emails, PDFs, and photos?
Yes, for many uses. AI that drafts documents or extracts fields works directly from emails, PDFs, and photos, so those don't need converting into a database first. What matters is that they're the right, current versions. Prediction is different: a forecast needs the outcome recorded in a field it can count, consistently, over a long enough period.
Continue Learning
Data readiness has no answer for the company as a whole. It has an answer for each workflow, and it usually comes down to a few fields and an hour with the paperwork.
Assessing readiness:
- AI readiness assessment services
- AI readiness assessment methodology
- AI readiness assessment for SMBs
- AI readiness assessment checklist and PDF
- AI infrastructure readiness assessment
Choosing who does it:
- Which AI consulting company should I choose?
- AI readiness assessment consulting firms
- Why do AI projects fail?
Sources
All sources retrieved 2026-09-25.
- Neil D. Lawrence, "Data Readiness Levels," arXiv 1705.02245, submitted May 5, 2017 (full text via ar5iv), retrieved 2026-09-25: https://arxiv.org/abs/1705.02245
- ISO 25000 portal, "ISO/IEC 25012" (the ISO/IEC 25012:2008 data quality model), retrieved 2026-09-25: https://iso25000.com/en/iso-25000-standards/iso-25012
- Gartner, "Lack of AI-Ready Data Puts AI Projects at Risk," press release, February 26, 2025, retrieved 2026-09-25: https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk
- LeBow College of Business, Drexel University, and Precisely, 2025 Outlook: Data Integrity Trends and Insights, September 2024, retrieved 2026-09-25: https://www.lebow.drexel.edu/sites/default/files/2024-09/drexel-lebow-precisel-data-integrity-trends-insights-2025-outlook.pdf
- LeBow College of Business, Drexel University, and Precisely, 2026 State of Data Integrity and AI Readiness, January 2026, retrieved 2026-09-25: https://www.lebow.drexel.edu/sites/default/files/2026-01/lebow-precisely-state-data-integrity-ai-readiness-2026.pdf
- Tadhg Nagle, Thomas C. Redman, and David Sammon, "Only 3% of Companies' Data Meets Basic Quality Standards," Harvard Business Review, September 11, 2017, retrieved 2026-09-25: https://hbr.org/2017/09/only-3-of-companies-data-meets-basic-quality-standards
- Nithya Sambasivan and others, "'Everyone wants to do the model work, not the data work': Data Cascades in High-Stakes AI," Proceedings of CHI 2021, Google Research publication page, retrieved 2026-09-25: https://research.google/pubs/everyone-wants-to-do-the-model-work-not-the-data-work-data-cascades-in-high-stakes-ai/
- Department for Business and Trade, Longitudinal Small Business Survey 2024: SME employers (businesses with 1 to 249 employees), GOV.UK, September 25, 2025, retrieved 2026-09-25: https://www.gov.uk/government/statistics/small-business-survey-2024-businesses-with-employees/longitudinal-small-business-survey-2024-sme-employers-businesses-with-1-to-249-employees
On the page review. "The 17 pages" means the distinct organic results for "data readiness assessment services for ai initiatives," "ai data readiness assessment," "data readiness assessment for AI small business," and "data readiness assessment AI consulting what's included," fetched and read on 2026-09-25. Twenty distinct URLs came back; three returned "not found" errors and aren't counted. They included pages from EY, RSM, Deloitte, CBIZ, AWS, Actian, TDWI, Telus Digital, Infomineo, and several data and AI consultancies. "Judged against a use case" means the page said data readiness should be assessed for a specific use; "set the use case first" means the page's own sequence named use cases before the data was evaluated. "A test with a count and a pass mark" means a sample size and a threshold an owner could apply to their own records. "Did the reverse" means the page evaluated data before choosing use cases (Techment, AWS, and Cybic). AWS's small-business checklist came closest to a test, with an inventory of source systems, owners, and update frequency, but no error count or threshold. "About 14 sell data work" is our reading of each publisher's services. The page that called PDFs and email threads a failure and the page quoting a 60-day move to a foundation build are Phos AI Labs and Techment. The durations were RSM's four weeks, Techment's two, four, or six, Phos AI Labs' two to four weeks for one or two use cases, Cybic's three to six, and Infomineo's four to eight. It's a snapshot of one day's results, not a market survey.
On the surveys. The Gartner, LeBow, and Precisely surveys are of data management and analytics professionals, mostly at large organizations. The 2025 LeBow sample was broader (30% at organizations under 250 employees) than the 2026 one (companies with at least 1,000 employees or $250 million in revenue), and its authors note the shift, so the two editions aren't a trend. The 2026 "confidence-reality gap" figures come from separate questions. Gartner's 60% figure is a prediction, not an outcome.
On the Redman study. Harvard Business Review paywalls the article body. The figures quoted (75 measurements, 47%, the quarter below 30, "only 3%," the "high nineties," and the rule of ten with its $100 and $199 example) were checked against the article's opening and a full attributed reprint. The managers who ran the measurement were in the authors' classes, so it isn't a random sample of companies.
On the data-cascades study. The 92% is from interviews with 53 practitioners working on high-stakes AI in India, East and West Africa, and the US. It describes those practitioners' projects, not small businesses.
On the table of AI uses. The five kinds of use and what each needs are our own framework, built on Lawrence's bands. They're not a published standard.
On the glazing contractor. It's an illustrative composite, not a client. Its figures (25 site surveys a week, 45 minutes per quote, 20 of 20 and 19 of 20 matches, 9 of 60 stale prices, about 1,400 project bids with about 520 outcomes) are representative, not measured.
On first-party claims. Descriptions of how the Avolis diagnostic works, and the observation that the most common gap is a field nobody records, describe our own work. They are not independent research and are not offered as benchmarks.
About Avolis Research Group
Avolis Research Group is Avolis's in-house research practice, focused on how operations-heavy small and mid-sized businesses actually adopt AI. It synthesizes primary economic research, government survey data, and results from real implementations into practical, vendor-neutral guidance.
Ready to make AI work for you?
Book an AI readiness evaluation. If there’s nothing worth automating, we’ll tell you.
