Generative AI Readiness Assessment: Test the Output, Not the Company
Avolis Research Group
·
·
24 min read
A generative AI readiness assessment tests whether your team can check AI drafts faster than writing them. Legal AI tools erred on 17–33% of test questions.
A generative AI readiness assessment checks whether your business can use AI that drafts, summarizes, and answers in plain language on one specific piece of work, safely, and for less effort than doing that work by hand. It asks different questions from a general AI readiness assessment because you don't build or train a generative model; you rent one that already exists. That makes years of clean data and your own computing power matter much less. What matters more is whether the person who knows the work can check every draft faster than they could have written it, whether the documents the AI draws from are current, and whether your staff know what they may put into which tool. The most useful single test is to run the AI on ten jobs you've already finished and compare its drafts with what you actually sent.
In September 2026 we reviewed the 19 pages ranking for "generative ai readiness assessment" and three close variants. Most were general AI readiness quizzes and service pages that scored the company, not the task. Ten scored the organization as a whole, and only one explained what changes when the AI is generative. None of their readable text used the word "hallucination," only one counted the time a person spends fixing the cases the AI gets wrong, and none suggested testing the AI on work the business had already finished, where the right answers are known.
This page covers the generative side of our guide to AI readiness assessment services, for businesses of 10 to 200 people.
Key Takeaways
- Generative AI changes where the readiness risk sits. You rent the model, so the questions shift from data volume and infrastructure to checking the output, keeping source documents current, and setting rules for what goes in.
- NIST's Generative AI Profile lists 12 risks unique to or made worse by generative AI, including "confabulation," its term for confidently stated but false output.
- Even purpose-built tools make confident mistakes. In a peer-reviewed Stanford study, legal research tools built on retrieval gave false answers on 17% to 33% of test questions, against 43% for a general chatbot, and a 2026 test of eight systems found rates from under 10% to nearly half.
- Of 19 pages ranking for "generative ai readiness assessment" and three variants in September 2026, none suggested testing the AI on work the business had already finished, and none used the word "hallucination."
- The practical test is a back-test: run the AI on 10 finished jobs, have the person who knows the work review each draft, and time the review. At a 17% error rate, 10 cases show at least one error about 84% of the time.
- In Cisco's 2024 survey of 2,600 privacy and security professionals, 48% admitted entering non-public company information into generative AI tools.
Table of Contents
- What a generative AI readiness assessment is
- How it differs from a general AI readiness assessment
- Why checking the output is the real readiness test
- The back-test: run it on work you've already finished
- What goes in: source documents, accounts, and rules
- Which work is generative-shaped
- What the assessment should deliver
- A worked example: a 38-person building inspection firm
- Where Avolis fits
- Frequently Asked Questions
- Continue Learning
What Is a Generative AI Readiness Assessment?
A generative AI readiness assessment is a check of whether a specific piece of work can be handed, in part, to AI that produces text (drafts, summaries, answers, extracted fields) and whether the business can use that output without it causing harm. Its result should be a decision for each piece of work it examined, not a score for the company.
Generative AI is the kind that writes. It drafts a proposal from inspection notes, summarizes a 40-minute site call, pulls the line items off a supplier invoice, or answers a crew lead's question from the warranty file. It isn't the older kind of AI that learns a pattern from your history, such as a model that forecasts material usage from three years of jobs. That older kind is what most readiness frameworks were designed for, and the difference matters, because a generative model arrives already trained. You pay a monthly fee or a per-use charge, and the provider runs the computers, which is why our infrastructure assessment says to rent the compute and assess the connections instead.
The risk arrives with the model, too. In July 2024 the National Institute of Standards and Technology published a Generative AI Profile to sit alongside its AI Risk Management Framework, and it lists 12 risks that are "unique to or exacerbated by" generative AI. The second on the list is confabulation, which NIST defines as "the production of confidently stated but erroneous or false content," known colloquially as hallucinations (NIST, AI 600-1: Generative Artificial Intelligence Profile, July 2024). Other entries that matter to a small operation include data privacy, intellectual property, and what NIST calls human-AI configuration, which covers people trusting the output too much or too little.
In other words, a generative AI readiness assessment is less about whether your business can support AI and more about whether it can supervise it.
How It Differs From a General AI Readiness Assessment
A general AI readiness assessment asks whether you have the data, systems, skills, and rules to build and run AI. For generative AI, several of those questions shrink and one new one takes over: can the people who know the work catch a confident mistake before it leaves the building?
Our pillar guide covers six areas every readiness assessment should examine. Here's how each one shifts when the AI is generative:
| Area | What a classic AI assessment asks | What a generative AI assessment asks instead |
|---|---|---|
| The work | Is there a repeatable task with enough volume? | Same, plus: is the output text a person can read and judge, and how long does judging it take? |
| Data | Is there enough clean history to train or tune a model? | Are the few documents the AI draws from (price lists, specs, templates, policies) the current ones? |
| Systems | Can we host, connect, and feed a model? | Which accounts are people using, and can the business see and close them? |
| People | Do we have technical staff to build and maintain it? | Does the reviewer know the work well enough to spot a plausible error, and do they have time to? |
| Rules | Is there governance for models and data? | Do staff know what may be pasted into which tool, and have we read that tool's data terms? |
| Checking output | Measure the model's accuracy once, before launch | Check every output that matters, every time, because the same prompt can produce a different answer |
The table is our framework, not a published standard. The last row is the one most assessments miss, and the research on how people actually work with these tools explains why it matters.
Researchers at Microsoft and Carnegie Mellon surveyed 319 knowledge workers who shared 936 real examples of using generative AI at work. They found that the tools shift effort "toward information verification, response integration, and task stewardship," in other words, from doing the work to overseeing it. They also found that "higher confidence in GenAI is associated with less critical thinking," while higher confidence in one's own skill went with more of it (Lee and others, "The Impact of Generative AI on Critical Thinking", CHI 2025). Among the barriers to checking the output, 44 of the 319 mentioned a lack of time.
NIST names the same pattern automation bias, "excessive deference to automated systems," and warns that it can make the confabulation risk worse. Put plainly, the person reviewing an AI draft is doing a different job from the person who used to write it, and a readiness assessment has to ask whether that person has the knowledge and the minutes to do it well.
Why Checking the Output Is the Real Readiness Test
Generative AI makes mistakes that look exactly like correct answers, and the rate varies so much by tool and task that the only reliable number is one you measure on your own work. That's the core of generative readiness. If a person who knows the work can catch those mistakes quickly, the tool saves time. If they can't, it moves the work downstream and adds risk.
The clearest measurement comes from law, where a wrong answer is easy to prove. A Stanford team ran more than 200 preregistered legal research questions through three commercial tools built on retrieval, the method vendors use to ground answers in a library of real documents, and through GPT-4 as a general-purpose comparison. Vendors had described that method as "eliminating" hallucinations and promised "hallucination-free" citations. The study found otherwise. Lexis+ AI gave a false answer on 17% of questions, Ask Practical Law AI on 17%, and Westlaw's AI-Assisted Research on 33%, while GPT-4 on its own reached 43% (Magesh and others, "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools", Journal of Empirical Legal Studies, 2025).
Those tools have been updated since the spring of 2024, so the specific rates are history. The spread isn't. In an August 2026 preprint, researchers at the University of Luxembourg tested eight legal retrieval systems on European data protection law and a national civil code written in French. False answers ranged "from less than 10% of responses for the best-performing systems to nearly half in the worst case" (Das, Abualhaija, and Bianculli, "How Much Do Legal RAG Systems Still Hallucinate?", arXiv, August 2026). Questions built on a wrong assumption produced especially high rates, which is worth knowing, because an estimator asking "what's the markup on the Type B unit?" when there is no Type B unit is exactly that kind of question.
Law isn't a contracting business, and we cite these studies for their method, not their numbers. What they show is that a confident, well-formatted answer tells you nothing about whether it's right, that a tool's marketing tells you even less, and that the rate for your work can only come from checking your work.
The cost of not checking shows up in the office rather than in court. In a September 2025 survey of 1,150 full-time US desk workers by BetterUp Labs and the Stanford Social Media Lab, 40% said they'd received AI-generated "workslop" in the past month, meaning work that looked finished but wasn't. Respondents estimated each instance took about two hours to sort out (BetterUp Labs, "Workslop", September 2025). Those are self-reported estimates, but the mechanism is the one a small business should worry about: the time the AI saved the sender gets spent, with interest, by whoever receives the draft.
The arithmetic that decides readiness: generative AI pays off on a task only when the time to prompt, review, and fix a draft is clearly less than the time to do the task by hand. In our diagnostics, that comparison settles more generative AI ideas than any question about data or tools, and it's the one number most assessments never measure. A draft that takes 12 minutes to check and correct beats a 45-minute task. A draft that takes 35 minutes to check, because the reviewer has to re-derive every figure, doesn't, however impressive it looks. It's the second of the two added checks on our should we implement AI page, and the back-test below is how you measure it.
The Back-Test: Run It on Work You've Already Finished
The fastest reliable way to test generative AI readiness is a back-test: give the AI the original inputs from ten jobs you've already completed, and compare its drafts with the versions a person actually approved. You already know the right answers, so every error is visible, and the whole test takes an afternoon rather than a two-month rollout.
It's a different test from a trial. A trial tells you whether people keep using a tool, and our page on how businesses assess readiness sets out how to run one properly. A back-test tells you whether the output is right, and how long it takes to make it right, before anyone depends on it. For generative AI we'd run the back-test first, because there's no point measuring adoption of drafts nobody should send.
- Pick one task and pull ten finished examples. Choose recent jobs where you still have the inputs (the field notes, photos, call recording, or customer email) and the final output a person approved and sent. Include one or two awkward ones, not just the clean cases.
- Write the instructions once. Use the same prompt, template, and reference documents for all ten, the way the tool would really be used. Don't tune the instructions to each case after seeing the answer.
- Run the AI on the inputs only. It must not see the finished version. If it's a tool that reads your files, make sure the answer isn't sitting in a folder it can reach.
- Have the person who'd really review it check each draft. Not the owner, and not whoever is most excited about the tool, but the estimator, coordinator, or project manager who'd be accountable for what goes out. They compare each draft with the source inputs and the approved version.
- Time the review and fix, and grade every error. Record minutes per draft and sort each error into one of three grades: a wrong fact that would mislead the customer, something missing that should be there, or style only.
- Compare with the time the task takes today. Use a few timed recent examples rather than anyone's estimate, which is the baseline our DIY readiness guide walks through.
| Record for each of the 10 cases | Why it matters |
|---|---|
| Minutes to review and fix the draft | The number that decides whether it saves time |
| Wrong facts, counted | Anything that would mislead a customer, supplier, or inspector |
| Missing items, counted | Often caught at review, but they add minutes |
| Style-only edits | Usually fixable once in the instructions, not each time |
| Whether the reviewer caught every wrong fact | Have a second person check three of the ten; a miss here matters more than the error itself |
Ten cases is small, but it's enough to catch the error rates the studies above found. The arithmetic is simple: if a tool makes a serious error 17% of the time, the chance that ten independent cases show none is 0.83 to the tenth power, about 16%, so a back-test of ten catches it about 84% of the time. At twenty cases the chance of seeing no errors falls to about 2%. So ten is enough to decide whether an idea is worth pursuing, and twenty is our bar for anything that would reach a customer without a person reading it first. Those are our calculations from the binomial formula, not figures from the studies.
Set the pass mark before you start. Our default for a reviewed draft is that review and fix take no more than half the time of doing the task by hand, and that the reviewer catches every wrong fact in the cases the second person checks. For anything that would go out unreviewed, the bar is zero wrong facts in twenty. Even that can't rule out an error rate of around one in seven, which is one more reason most businesses shouldn't send generative AI output to customers unreviewed at all.
What Goes In: Source Documents, Accounts, and Rules
What a generative AI produces depends on what goes into it, which puts three things on the readiness list: the reference documents it draws from, the accounts people use to reach it, and the rules about what they may type or upload.
Reference documents. A generative tool doesn't need three years of clean records. It needs the handful of documents it will copy from to be the current ones: this year's price list, the latest spec sheet, the proposal template the owner actually likes, the warranty terms in force. An AI that drafts from last year's price list will quote last year's prices every time, fluently. Our data readiness assessment sets out why drafting needs readable, current inputs while prediction needs long, consistent history, so the generative check is short: list the reference documents, confirm each is current, and name who keeps it current.
Accounts and vendor terms. Staff are very likely using generative AI already, often on personal accounts, which our infrastructure assessment covers under shadow AI. The generative-specific question is what the terms of each account allow the provider to do with what goes in, and those terms often differ between a consumer plan and a business one on the same product. Google's consumer privacy notice for its Gemini apps, for example, says trained human reviewers read some conversations and that reviewed chats are kept for up to three years. It asks users not to enter "confidential information that you wouldn't want a reviewer to see" (Google, "Gemini Apps Privacy Hub", retrieved September 2026).
Its business counterpart says the opposite about the same model family: customer content "is not human reviewed or otherwise used for Generative AI model training outside your domain without permission" (Google, "Generative AI in Google Workspace Privacy Hub", retrieved September 2026). We're not recommending either product. The point is that the same brand can come with very different terms, and an assessment should read the terms for the accounts people are actually using.
Terms also change. In January 2024 the Federal Trade Commission reminded AI model providers that customers "may reveal sensitive or confidential information when using a company's models, such as internal documents and even their own users' data." It warned that quietly going back on privacy commitments, "for example, by surreptitiously changing its terms of service or privacy policy," can bring enforcement (FTC, "AI Companies: Uphold Your Privacy and Confidentiality Commitments", January 2024). That protects you after the fact. The cheaper protection is a named person who re-reads the terms once a year.
Rules for staff. Most organizations know this is a gap. In Cisco's 2024 Data Privacy Benchmark, which surveyed 2,600 privacy and security professionals across 12 geographies, 63% said their organizations limited what data could be entered into generative AI tools and 61% limited which tools employees could use. Yet 48% admitted entering non-public company information into those tools themselves, and 45% had entered information about employees (Cisco, "Organizations Ban Use of Generative AI Over Data Privacy Security", January 2024).
Those respondents were privacy and security professionals, the people most likely to know the rules. At a 30-person contractor with no written rule at all, it's safe to assume the estimators are pasting customer drawings into whatever chat tool they like best. The fix is a one-page rule covering which accounts, what stays out, what gets checked, and who owns it, and our page on whether to implement AI sets it out. An assessment should confirm the rule exists and that people know it, not rewrite it.
Which Work Is Generative-Shaped
Generative AI is strongest where the work is turning information that already exists into text a person can check, and weakest where the work is arithmetic, scheduling, or a judgment that commits money or safety. Sorting your candidate tasks by that pattern is the first step of the assessment, before any back-test.
| Task pattern | Example in an operations business | Generative fit | The readiness question |
|---|---|---|---|
| Drafting from notes and photos | Inspection report narratives, repair proposals, change-order letters | Strong | Can the reviewer check a draft in well under the time it takes to write one? |
| Summarizing | Site-call summaries, handoff notes between shifts, a digest of a long email thread | Strong | Will anyone notice if the summary leaves out the one line that mattered? |
| Pulling fields from documents | Supplier invoices, permits, and work orders into the job system | Good | Is there an agreed place for each field to land, and a spot check? |
| Answering from your own files | A crew lead asks which warranty covers a unit | Good internally | Is there one current version of each document, and a way to see the source? |
| Talking to customers directly | A website chatbot quoting service windows or prices | Risky | Would you stand behind every answer it could give? (See customer-facing AI.) |
| Calculating and pricing | Takeoffs, markups, repair cost estimates | Poor on its own | Keep the math in a spreadsheet or estimating tool, and let the AI write around it |
| Scheduling and routing | Assigning techs to jobs by skill, location, and availability | Not generative work | This is optimization, a different kind of software, and often a setting in a tool you already pay for |
The table is our own sorting, not a published taxonomy. The two bottom rows matter because they're where generative AI gets pitched most often and fits least well. A chat assistant can explain how to price a job and draft the cover letter for the estimate, but the numbers in it should come from the system that calculates them, not from the model's best guess at what they ought to be.
Once the list is sorted, the strong and good rows go forward to a back-test, and the others get a plain "not this tool" in the findings. Our methodology page explains why we treat the work as a gate rather than one score among several, and the same logic applies here: a task that fails the back-test isn't made ready by a strong result anywhere else.
What the Assessment Should Deliver
A generative AI readiness assessment should hand you a short list of tasks, each with a back-test result, a decision, and the one or two fixes that stand between it and use. If what you get is a maturity level, a generic acceptable-use policy template, or a recommendation for a particular product, you've bought something else.
For each task in scope, the finding should include:
- The task and who reviews the output. A named role, not "the team."
- The back-test numbers. Cases run, wrong facts found, minutes to review and fix, and the baseline time for doing the task by hand.
- The source documents it depends on, whether each is current, and who owns keeping it that way.
- The account and terms it would run under, checked against what the data involved allows.
- A decision: go, wait, or no, following the rules on our services page, with the reason in a sentence.
A few questions will tell you what kind of assessment you're buying before you sign:
- Will you test the AI on our own finished work, or score us on a questionnaire? Only the first tells you how the output will perform.
- Who reviews the drafts in your test? It should be the person who'd review them for real.
- Will you time the review, not just count the errors? A tool with few errors can still cost more time than it saves if every draft has to be re-derived.
- Will you read the data terms for the tools our staff already use? That's a half-hour job that's often skipped.
- Do you sell or resell a particular AI product? Many good firms do, and it's worth knowing so you can weigh the recommendation.
Who offers these assessments and what they tend to cost is compared on our AI readiness assessment consulting firms page, and the wider partner decision is covered in which AI consulting company to choose.
A Worked Example: A 38-Person Building Inspection Firm
This is an illustrative composite, not a client, and its figures are examples rather than benchmarks. Picture a 38-person commercial building inspection firm that writes property condition reports for buyers and lenders. It has 14 field inspectors, six report writers, four cost estimators, and an office of 14. The firm sends about 30 reports a month. A report runs 30 to 60 pages, and timed across five recent reports, the writers spent about three hours on the narrative sections of each one, turning the inspectors' notes and photos into findings a lender can rely on. The owner has four generative AI ideas and a partner who's worried about liability.
Idea 1: drafting report narratives. The senior report writer pulled ten reports sent in the last two months, including two awkward buildings with mixed-use additions. The writer ran the AI on the field notes and photos only, with the firm's report template and its standard condition definitions as reference. Reviewing and fixing each draft took between 35 and 80 minutes, a median of about 50, against roughly three hours to write the same sections by hand. Across the ten drafts, the AI made four wrong-fact errors, three of them on the awkward buildings (a roof age taken from the wrong note, a system described as original when the photos showed a replacement). The senior writer caught all four, and a second writer who checked three drafts, including both awkward buildings, found no additional misses.
Idea 2: a client chatbot for report questions. Lenders and buyers call with questions about findings, and the owner wanted a chatbot to answer from the report. It was never back-tested, because it failed the first question: would the firm stand behind every answer it could give? A report is a professional opinion that clients rely on for lending decisions, and an unsupervised answer that paraphrased a finding wrongly would carry the firm's name.
Idea 3: staff use of personal chat accounts. In working through the report workflow, the team mentioned that two writers already pasted inspectors' notes, including client names and addresses, into personal chat accounts to tidy the wording. Nobody had read those accounts' terms, and the firm's engagement letters promise clients confidentiality.
Idea 4: repair cost tables. The estimators wanted AI to produce the cost-to-cure tables. This is arithmetic from unit costs, not generative work, and the firm's estimating spreadsheet already does it; what they actually wanted was faster lookups of current unit costs, which is a data-entry problem.
| Idea | Kind of work | Key finding | Call | Owner and next step |
|---|---|---|---|---|
| Report narratives | Drafting from notes and photos | Median ~50 minutes to review and fix vs ~3 hours by hand; 4 wrong facts in 10, all caught | Go, with a senior writer reviewing every draft | Senior report writer owns the template and condition definitions |
| Client chatbot | Talking to customers | Can't stand behind every possible answer on a professional opinion | No | Owner; revisit only as an internal tool for staff taking calls |
| Personal chat accounts | Staff use | Client data in accounts with unread terms, against confidentiality promises | Not a task: rule fix, one company account and a one-page rule | Office manager, this month |
| Cost-to-cure tables | Calculating | Arithmetic already handled by the estimating spreadsheet | No (not generative work) | Lead estimator updates unit costs quarterly |
The go and no calls follow the decision rules on our services page. The personal accounts aren't a task to decide on; they're a gap in the one-page rule. The narrative drafts cleared our half-time bar comfortably, but the errors clustered on the awkward buildings, which is why the call keeps a senior writer on every draft rather than any writer.
What the example shows: the idea with the clearest payoff and the idea with the clearest risk used the same kind of AI. The difference was who checks the output before it matters. The report narratives go through a senior writer who knows what a wrong roof age looks like, while the chatbot's answers would go straight to a lender. And the most urgent finding wasn't a project at all. It was the two personal accounts, which cost nothing to fix and carried the most exposure.
Where Avolis Fits
For generative AI, the back-test is part of our two-week diagnostic rather than a separate service. We work through each workflow with the people who run it, measure how long the work takes today, run the AI on a sample of finished jobs with the person who'd review it, and time and grade what comes back. Every engagement starts with that diagnostic, and the ranked result is yours whatever you decide to do next.
If you'd rather run a back-test yourself, the six steps above are all you need, plus an afternoon with the person who knows the work best. If the question behind this one is whether any assessment is worth paying for at your size, our page on AI readiness assessment for SMBs covers that. Our coverage of why AI projects fail shows what happens when nobody checks the output before it counts.
Frequently Asked Questions
What is a generative AI readiness assessment?
A generative AI readiness assessment checks whether a specific task can be handed in part to AI that drafts, summarizes, or answers in plain language, and whether the business can use that output safely. A good one tests the AI on your own finished work, times how long a knowledgeable person takes to review and fix each draft, checks the source documents and account terms, and ends in a go, wait, or no decision for each task.
How is a gen AI readiness assessment different from a general AI readiness assessment?
A general AI readiness assessment focuses on whether you have the data, systems, and skills to build or train AI. A gen AI readiness assessment starts from the fact that you rent a model that's already trained, so data volume and infrastructure matter less. It focuses instead on checking output, keeping reference documents current, reading the data terms of the tools staff use, and setting rules for what goes in.
How do you test whether generative AI output is accurate enough?
Run a back-test. Give the AI the original inputs from ten jobs you've already finished, and have the person who'd really review the output compare each draft with the version you sent. Count wrong facts, time the review and fix, and compare that time with doing the task by hand. Use twenty cases for anything that would reach a customer without review.
Do we need clean data to use generative AI?
Usually not in the way older AI needed it. A generative tool doesn't learn from years of your records. It works from the documents you give it, so what matters is that those few reference documents, such as price lists, templates, specs, and policies, are the current ones and that someone owns keeping them current. Prediction and forecasting are different and do need consistent history.
Is it safe for employees to put company information into AI tools?
It depends on the account and its terms. Consumer and business plans of the same product can differ on whether people at the provider may read conversations and whether content is used to train models. Check the terms of every account in use, move work onto a company account where the terms fit your data, and write a one-page rule on what stays out.
Continue Learning
Generative AI readiness isn't a property of the company. It's a property of each task, and it comes down to whether the person who knows the work can check the AI's drafts faster than they could write them.
Assessing readiness:
- AI readiness assessment services
- AI readiness assessment methodology
- AI readiness assessment for SMBs
- Data readiness assessment for AI
- AI infrastructure readiness assessment
- Should we implement AI?
Choosing who does it:
- Which AI consulting company should I choose?
- AI readiness assessment consulting firms
- Why do AI projects fail?
Sources
All sources retrieved 2026-09-25 unless noted.
- National Institute of Standards and Technology, NIST AI 600-1: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, July 2024, retrieved 2026-09-25: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- Hao-Ping (Hank) Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson, "The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers," CHI 2025, Microsoft Research, retrieved 2026-09-25: https://www.microsoft.com/en-us/research/wp-content/uploads/2025/01/lee_2025_ai_critical_thinking_survey.pdf
- Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning, and Daniel E. Ho, "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools," Journal of Empirical Legal Studies 22(2), 216–242, 2025 (arXiv 2405.20362, version of June 2024), retrieved 2026-09-25: https://arxiv.org/abs/2405.20362
- Souvick Das, Sallam Abualhaija, and Domenico Bianculli, "How Much Do Legal RAG Systems Still Hallucinate?," arXiv 2608.14210, submitted August 14, 2026, retrieved 2026-09-25: https://arxiv.org/abs/2608.14210
- BetterUp Labs and Stanford Social Media Lab, "Workslop: The Hidden Cost of AI-Generated Busywork," September 2025, retrieved 2026-09-25: https://www.betterup.com/workslop
- Google, "Gemini Apps Privacy Hub," retrieved 2026-09-25: https://support.google.com/gemini/answer/13594961?hl=en
- Google, "Generative AI in Google Workspace Privacy Hub," retrieved 2026-09-25: https://support.google.com/a/answer/15706919?hl=en
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments," January 9, 2024, retrieved 2026-09-25: https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- Cisco, "Organizations Ban Use of Generative AI Over Data Privacy Security: Cisco Study," newsroom release on the 2024 Data Privacy Benchmark Study, January 2024, retrieved 2026-09-25: https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2024/m01/organizations-ban-use-of-generative-ai-over-data-privacy-security-cisco-study.html
On the page review. "The 19 pages" means the distinct organic results for "generative ai readiness assessment," "gen ai readiness assessment," "generative ai readiness assessment small business," and "generative ai readiness checklist," fetched and read on 2026-09-25. Twenty distinct URLs came back; one (Udemy Business) blocked retrieval and isn't counted, and two Microsoft quizzes and an image-only CoSN checklist were judged from their landing pages. They included pages from Microsoft, Cisco, RSM, Eide Bailly, EDUCAUSE, OvalEdge, Domo, several AWS and Microsoft marketplace listings, and a handful of consultancies. "Scored the organization as a whole" means the page assessed the company rather than specific use cases (10 pages; 2 assessed per use case, 6 mixed the two, and 1 tested an individual's skills). The page that explained what changes for generative AI is OvalEdge's, which puts the extra burden on governed inputs. The page that counts time spent on the cases the AI gets wrong is Tommaso Maria Ricci's guide ("time to resolve an exception"). "Used the word hallucination" means the term (or "hallucinate") appeared in the page's readable text; the two Microsoft quizzes and the CoSN checklist were checked on their landing pages only. "Testing on finished work" means running the AI on past jobs with known outcomes before a pilot; pages that recommended pilots on new work (Orases, Ricci) aren't counted. It's a snapshot of one day's results, not a market survey.
On the legal AI studies. The Stanford figures come from preregistered queries run between March and May 2024 on the named products, which have since been updated; we cite them for the method and the spread, not as current rates for those products. "False" means the study's "hallucination" label: a response that was incorrect or cited a source that didn't support it. The Luxembourg study is a preprint on the GDPR and a national civil code in French, and it hasn't yet been peer-reviewed.
On the surveys. The Microsoft and Carnegie Mellon study is a survey of 319 knowledge workers recruited through the Prolific platform, and its findings are associations from self-reports, not measured effects. The BetterUp figures (40%, about two hours per instance) come from an online survey of 1,150 full-time US desk workers and are self-reported estimates. Cisco's respondents are privacy and security professionals; we cite the figures as its newsroom release reports them, because the full report page blocked retrieval.
On the vendor terms. Google's two privacy hubs are quoted as an example of consumer and business terms differing on one product family. Terms change, and we don't recommend either product or any other.
On the back-test arithmetic. The 84% and 2% figures are our own calculations (1 − 0.83¹⁰ ≈ 0.84; 0.83²⁰ ≈ 0.024), assuming each case is independent and the tool's serious-error rate is 17%, the low end of the Stanford range. They're a guide to sample size, not a finding from any study.
On the tables. The comparison of classic and generative assessments and the sorting of task patterns are our own frameworks. They're not published standards.
On the building inspection firm. It's an illustrative composite, not a client. Its figures (38 people, three hours per narrative, 35 to 80 minutes to review, four wrong facts in ten drafts) are representative, not measured.
On first-party claims. Descriptions of how the Avolis diagnostic works, and the observation that the review-time comparison settles more generative AI ideas than any other question, describe our own work. They are not independent research and are not offered as benchmarks.
About Avolis Research Group
Avolis Research Group is Avolis's in-house research practice, focused on how operations-heavy small and mid-sized businesses actually adopt AI. It synthesizes primary economic research, government survey data, and results from real implementations into practical, vendor-neutral guidance.
Ready to make AI work for you?
Book an AI readiness evaluation. If there’s nothing worth automating, we’ll tell you.
