Which AI Consulting Company Should You Choose?

Avolis Research Group

·

·

18 min read

In 2025, 95% of enterprise AI pilots showed no measurable payoff. Here's how to choose an AI consulting company that ships working systems, not decks.

Choose the partner who can describe your operation back to you in plain language. Then make them prove it with a live demo on your own messy data. Neither test requires technical knowledge. In 2024, RAND interviewed 65 experienced data scientists and engineers about why AI projects fail. The leading cause wasn't weak technology. It was the people commissioning the work misunderstanding — or miscommunicating — the problem to be solved (RAND Corporation, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed). That finding changes your position in the room. The most common failure mode is a translation problem between the person who knows the business and the person who knows the technology. You already hold one side of it. Your job isn't to become technical. It's to be impossible to misunderstand, and to hire someone who can be understood back.

Worth saying plainly: we sell AI consulting. So does nearly every site ranking for this search. So this guide is built from things you can check without trusting us: federal enforcement records, government survey data, and questions whose answers you can judge yourself.

Key Takeaways

  • In 2024, RAND found the leading cause of AI project failure is miscommunicating the problem, not weak technology (RAND, 2024).
  • You can't audit a model. You can audit a live demo on your own messy records, and you should insist on one before signing.
  • "AI washing" — claiming a product runs on AI when people are doing the work — is now an enforcement category. The SEC and FTC have both charged it, and the records are public and free to read.
  • In 2026, 76% of small businesses Goldman Sachs surveyed used AI, but only 14% had it fully embedded in core operations (Goldman Sachs, 2026). That gap is what you're hiring for.

Table of Contents

As a Non-Technical Founder, You're Not Judging the Technology

Judge the fit between the work and the proposal, because that's where projects actually break. In 2024, RAND cited estimates that more than 80% of AI projects fail (RAND Corporation). That's roughly twice the failure rate of IT projects with no AI in them. RAND quotes that figure rather than measuring it, so treat it loosely. Its own findings are the useful part.

RAND's interviews surfaced five root causes of failure. Look at what they are:

  1. The people commissioning the work misunderstand or miscommunicate the problem to be solved.
  2. The organization lacks the data to train an effective model.
  3. The team chases the newest technology instead of the user's real problem.
  4. The organization can't manage its data or deploy what gets built.
  5. The problem is simply too hard for AI to solve.

Only the last one is a question about the technology. The other four are questions about your operation, your information, and how well two groups of people understand each other. Those you can assess.

Our finding: In our experience, being non-technical isn't the risk everyone tells you it is. The risk is the translation gap, and it runs in both directions. Consider two owners. One says "I don't know anything about AI" but can explain exactly where a job stalls. The other has read every article and can't name a single bottleneck. The first is in far better shape. Domain clarity is the scarce input here, and you already own it.

RAND's first recommendation says the same thing from the other side: make sure the technical staff understand the project's purpose and domain context. So when you sit in that first meeting, you're not there to be quizzed. You're there to find out whether they can learn your business. For the wider decision this sits inside, start with which AI consulting company you should choose.

What Are You Actually at Risk of Here?

The specific risk is buying a capability that doesn't exist yet, and it's documented. In April 2025, the SEC charged the founder of Nate, Inc. over its AI claims (SEC, Litigation Release No. 26282). According to the SEC's complaint, he raised over $42 million by telling investors the shopping app completed purchases without human involvement. The complaint alleges Nate relied in large part on contract employees entering orders by hand. These are allegations, not findings. And Nate is the extreme end of what regulators have charged — the everyday version looks far duller.

What Two Regulators Have Actually Charged

Presto Automation is the everyday version, and it settled. In January 2025, the SEC's order found Presto had described the AI behind its drive-thru product misleadingly (SEC, Administrative Proceeding File No. 3-22413). For a period, according to the order, it didn't disclose that a third party owned and operated the speech recognition in every deployed unit. Presto later shipped its own version, then claimed that product removed the need for human order-taking. The order found the vast majority of orders still needed a human. Presto settled without admitting or denying the findings.

Two federal regulators have now made this a named category. In March 2024, the SEC settled charges against two investment advisers over their AI claims, collecting $400,000 in penalties. The orders found the claims false and misleading, and neither firm admitted or denied those findings. "Such AI washing hurts investors," said then-Chair Gary Gensler (SEC, Press Release 2024-36). In September 2024, the FTC announced Operation AI Comply — five enforcement actions in one sweep (FTC, FTC Announces Crackdown on Deceptive AI Claims and Schemes). One was a $193,000 settlement with a company selling an "AI lawyer." As then-Chair Lina M. Khan put it: "There is no AI exemption from the laws on the books."

Choosing an AI Consulting Partner as a Non-Technical Founder - Avolis AI "AI washing" is an enforcement category now Selected US federal actions over false or misleading AI claims Mar 2024 SEC · 2 advisers $400,000 penalties Sep 2024 FTC · 5 actions Operation AI Comply Jan 2025 SEC · Presto Nasdaq-listed co. Apr 2025 SEC · Nate founder $42m raised Two recurring claims drew charges: presenting a third party's technology as your own, and overstating how often the system runs without a human. Sources: SEC press release 2024-36; FTC, Sept 2024; SEC AP 3-22413; SEC LR-26282.
Sources: SEC, Press Release 2024-36, March 2024; FTC, Operation AI Comply, September 2024; SEC, Administrative Proceeding 3-22413, January 2025; SEC, Litigation Release 26282, April 2025.

Be fair about what this does and doesn't mean. Every one of these cases involves selling to investors or consumers, not a consulting engagement, and the overwhelming majority of consultancies are nothing like them. Nobody should read this section as a reason to distrust the firm in front of you. What the records give you is narrower and more useful than suspicion: a pattern of exactly which claims proved hollow under scrutiny. Two came up repeatedly. Presenting another company's technology as your own, and overstating how often the system runs unattended. Both are questions a paying customer can ask in a first meeting, and both have answers a non-technical buyer can evaluate, which is the whole value of reading enforcement records instead of vendor blogs.

How to Check the Public Record Yourself

Search the SEC and FTC databases before you sign anything. The SEC's Action Lookup covers individuals only, so send it the names of the people who will actually sign your contract. Send the firm's own name to the FTC's enforcement actions and to the SEC's litigation releases. All three are free, and none of them needs a login. If you've been told who the subcontractors are, run those too.

Our finding: No buyer's guide we've read tells you to do this, and it takes about five minutes. Be careful how you read the result, though. A clean search proves very little, because most firms have never been charged with anything. A hit tells you plenty.

Ask Any AI Consulting Partner for a Demo on Your Own Data

This is the one test that needs no technical knowledge and can't easily be faked. In 2024, RAND's interviews found that models are too often optimized for the wrong metrics (RAND Corporation). Others get built so they don't fit the business workflow and context at all. A demo on someone else's clean data can't reveal either problem. A demo on yours will. Here's how to run it. Pull twenty real records from the workflow you want fixed — and choose the ugly ones. The invoice with handwriting on it. The voicemail from the customer with bad reception. The estimate that got revised four times.

Protect the records before they leave your building. Get a mutual NDA signed first, then black out customer names, addresses, and any payment or account numbers. In our experience, most workflows test fine on redacted records, because the mess you're testing lives in the format, not the identity. If a workflow genuinely needs real identifiers, ask for written data-processing terms covering where the data sits and when it's deleted, and a firm that handles regulated clients will have those terms ready. Then watch, live, while they run it. Not a recorded video, not a deck with screenshots, and not a sanitised sample file they prepared last week. Ask them to narrate what's happening as it happens.

You're listening for four things:

  • Does it break, and do they say so? Everything breaks on real data. Honesty about where is the signal.
  • What happens when it's wrong? There should be a named answer: a human checks it, a flag gets raised, a queue exists.
  • How often does it need a human? Ask for a number, then ask how they measured it.
  • Can they explain the failure in your language? Not the model's language. Yours.

That third question is the one the SEC's Presto order turned on, so it's worth asking plainly. In our experience, this is also the moment a sales process either becomes a working conversation or quietly stalls. A firm that has built the thing will reach for specifics without being pushed. They'll tell you which record types are hard and what share of records needed a human last month. They'll name the step that still needs a person to sign off. A firm that hasn't will keep the discussion at the level of capability rather than cases. You don't need to understand the model to hear that difference, and if your target is automating a specific process, the criteria for choosing a partner for AI-driven process automation go deeper on scoping the demo.

How Do You Tell a Diagnostic From a Sales Call?

Count who's talking. A real diagnostic spends most of its time on your operation, not their capabilities. In 2024, RAND recommended that leaders commit a team to one specific problem for at least a year (RAND Corporation). If a project isn't worth that commitment, RAND concluded, it probably isn't worth starting. A diagnostic exists to work out which problem earns it.

The difference shows up in the first thirty minutes.

Sales call Diagnostic
Opens with their platform, clients, and logos Opens with questions about how a job moves through your shop
Asks what you want built Asks what takes the longest and what comes back wrong
Wants your budget early Wants your volumes, your handoffs, and who touches what
Produces a proposal Produces a written description of your process you can correct
Scopes to what they already sell Scopes to one workflow, and sometimes to nothing

That last row matters most. If a firm has never told a prospect there's nothing worth automating yet, you're talking to a vendor with one answer. We turn down work when the numbers don't support it. Any firm worth hiring should be able to name a time they did the same. The deliverable test is simple. Ask for a written summary of your own process after the first meeting, then read it the way you'd read a quote from a subcontractor who has never seen the site. Can they describe it accurately enough that you'd correct only small details? That's the translation check, and it's what an AI readiness assessment is built to produce.

Which Questions Work When You Can't Check the Answer?

Ask questions whose answers you can judge on clarity and specificity rather than accuracy. The ranking below travels; the absolute rates don't, since this is EU enterprise data and you are not an EU enterprise. In 2025, Eurostat found that 70.89% of EU enterprises which considered AI and declined cited a lack of relevant expertise (Eurostat, Use of artificial intelligence in enterprises). Legal uncertainty came second at 52.52%, privacy third at 48.83%. Expertise leads by a wide margin. Adoption itself was still thin that year: 19.95% of EU enterprises used AI at all, and 17% of small ones.

Choosing an AI Consulting Partner as a Non-Technical Founder - Avolis AI The number one barrier isn't the technology Reasons cited by EU enterprises that considered AI and didn't adopt it, 2025 0% 25% 50% 75% Lack of expertise 70.9% Unclear legal risk 52.5% Privacy concerns 48.8% In the same year, 19.95% of EU enterprises used AI; among small enterprises it was 17%.
Source: Eurostat, Use of artificial intelligence in enterprises, 2025 reference year.

So the market knows most buyers can't evaluate the technology, and the good firms build for exactly that: they lead with the workflow and let you judge the fit. The questions below are designed so a vague answer is itself the finding.

Ask this A good answer sounds like Be careful if
"Which parts of this are yours, and which belong to someone else?" A plain list: the model is a third party's, the workflow and integrations are ours They call the whole stack "our platform" without naming a vendor
"What share of cases will still need a person?" A number, plus how it was measured, plus what happens to the rest "It's fully automated," or a shrug about measurement
"Explain this to me like I'll have to explain it to my foreman." A version with no jargon, in one minute The answer needs jargon to survive
"What's the smallest version we could ship first?" One workflow, one metric, a date A phased program covering the whole business
"What would make you tell us not to do this?" A specific condition they've actually hit before They treat it as an objection to handle
"Who on your team will I be talking to in month seven?" Names, and their role after go-live "Your account manager will coordinate"

Notice what none of these require. No jargon. You don't need to know what a model is. You need to notice whether clarity survives contact with a hard question.

Who's Actually Going to Do the Work?

Ask for names and roles, because people still run this. In 2026, Goldman Sachs surveyed 1,256 owners in its 10,000 Small Businesses program (Goldman Sachs, 10,000 Small Businesses Voices). It found 87% see AI as augmenting their workforce, not replacing it. That instinct is right, and it has a purchasing consequence. If humans stay in the loop, you're buying a working relationship, not a product.

So get specific about who:

  • Who builds it. A named person, not a pool.
  • Who trains your team. And how many hours, in your building or on your calls.
  • Who owns it after go-live. Someone on your side has to. Name them in the meeting.
  • Who's subcontracted. If part of the work goes elsewhere, you should know before you sign.
  • Who picks up the phone in month seven. The answer to this predicts most of your experience.

Our finding: Across our own client diagnostics we found the strongest predictor of a build that sticks isn't technical at all. It's whether one person inside the business — usually an ops manager or office lead — actually wants to own it. Where nobody claims it, the system quietly reverts to the old way within a quarter. We now ask who that person is before we scope anything, and if there isn't one, that's the first thing to fix. That's a pattern across our engagements, not a measured statistic — we haven't published a sample.

None of this is about distrust. It's that a system nobody owns isn't a system. It's a subscription. For the behavioral differences behind these answers, see what separates a good AI consulting partner from a typical vendor.

Why Does This Feel Harder Than It Should?

Because the gap between using AI and running on AI is enormous, and almost nobody has crossed it. In 2026, Goldman Sachs found 76% of the small businesses it surveyed use AI (Goldman Sachs, 10,000 Small Businesses Voices, n=1,256). Only 14% have it fully embedded in core operations. And 73% said they'd benefit from more training and implementation support.

Choosing an AI Consulting Partner as a Non-Technical Founder - Avolis AI Most use AI. Very few run on it. Goldman Sachs 10,000 Small Businesses participants, 2026 (n=1,256) 76% use AI in some form 14% fully embedded in core operations 0% 100% Both bars use one zero-based scale; the dashed frame is 100%. The 62-point gap is the work a partner is actually for. Separately, 73% wanted more implementation support, and 93% of those already using AI reported a positive impact.
Source: Goldman Sachs 10,000 Small Businesses Voices survey of 1,256 program participants, conducted by Babson College and David Binder Research, January 27–February 4, 2026. These are owners who had already signed up for a business-training program, so they are not a random sample of US small businesses and adoption among them runs ahead of the field. The 93% figure is a share of AI users, not of all respondents.

Read that 76% carefully, though, because a broader survey lands much lower. In 2026, the U.S. Census Bureau's Business Trends and Outlook Survey put AI use at 37% among firms with 250 or more employees (U.S. Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users). It was 32% at 100 to 249 employees, and under 20% among the smallest firms. Between December 2025 and May 2026, use grew among firms above 20 employees and stayed flat below that line.

Those two numbers measure different populations, and the difference matters. Goldman surveyed participants in its own business-education program — owners who already sought out training, so adoption among them runs far ahead of the field. Census surveys businesses at random, so read the Goldman figures as the shape of what happens once an owner engages, and the Census figures as where most operations your size actually sit today.

So if this decision feels hard, that's the norm at your size, not a personal shortfall. Very few operations your size have done it, which means there are fewer people around you to ask, and that's precisely why the demo and the plain-language test matter more than references.

What Belongs in the Agreement, in Plain English?

Anything you'd struggle to argue about later. In March 2024, the SEC's AI washing case turned up something else too (SEC, Press Release 2024-36). According to the SEC's order, Global Predictions' advisory contract also included a clause limiting what clients could hold the firm responsible for. Advisers aren't allowed to do that. Contract language is exactly where an expertise gap gets locked in. Read the agreement like it's the product, because for a year it is.

Five items to insist on, in words you can read without help:

  • What "done" means. One workflow, one number, one date. Not a list of deliverables.
  • Who owns what you paid for. The workflows, the configurations, the documentation. In writing.
  • Access to your own systems and data. You keep the keys, always.
  • Documentation a new hire could follow. Written for your team, not for engineers.
  • What happens if it misses, and how you exit. Including what you keep on the way out.

Two of RAND's five root causes were about data and infrastructure. Either the organization lacked the data to train on, or it couldn't deploy and manage what got built. Both are cost centers disguised as assumptions, and both usually land on you rather than the firm. So make the agreement name who does that work, in hours, before anyone signs. If a line item is described as "prep" or "discovery," ask whose calendar it comes out of. Then ask what happens if it takes twice as long. The answer reshapes the price more often than the headline fee does. A firm that has done this before will have a view on it already, and for how these terms interact with fee structures, see what AI consultants actually charge.

When Should You Walk Away?

Walk when jargon does the work that evidence should. Census data gives a useful reality check here. In 2026, AI use grew only among firms with 20 or more employees (U.S. Census Bureau, May 2026). Among the smallest firms it stayed flat. So a firm proposing a company-wide program to a 30-person operation is presuming a maturity that adoption data says almost nobody that size has reached.

The signals that should end a conversation:

  • They can't explain it without jargon. If a plain-English version doesn't exist, it's protecting something.
  • No demo on your data. Confidentiality is a real constraint, so offer an NDA and redacted records first. If they still won't, the constraint was convenient.
  • "Fully automated," with no number. Two SEC cases turned on exactly this claim.
  • A quote before a diagnostic. They're pricing a product, not your problem.
  • No named people. You're buying a relationship. Insist on knowing with whom.
  • Talking down to you about your own business. You're the domain expert. That's the input they need most.

Does one flag mean walk? No. Two means slow down and ask again. What you're testing across all of them is a single thing: does clarity hold up when the questions get harder? And for the ROI arithmetic sitting behind those conversations, see how to choose an AI consulting service for business ROI.

Your Next Step

Before you talk to anyone, write one page about a single workflow. Where it starts, who touches it, where it stalls, what it costs when it comes back wrong. No AI terms at all. That page is your half of the translation, and it's the thing RAND's research says most projects are missing.

Then take it to a conversation and see whether they can hand it back to you sharper than you wrote it. If you'd like us to run that with you, that's what our readiness diagnostic does. And if it turns out there's nothing worth automating yet, we'll say so. Talk to us about a readiness assessment.

Frequently Asked Questions

Do I need technical knowledge to choose an AI consulting partner?

No, but you do need domain clarity. In 2024, RAND's interviews with 65 practitioners found the leading cause of AI project failure was the people commissioning the work misunderstanding or miscommunicating the problem. Your job is describing your operation precisely and checking whether the firm can describe it back. That's a communication test, not a technical one.

How do I know if a firm's AI is real?

Ask two questions the SEC has already litigated, both of which a firm with nothing to hide will answer in a sentence each. First, which parts of the technology are theirs and which belong to a third party. Second, what share of cases still need a human, and how that was measured. In January 2025 the SEC charged Presto Automation over misleading claims on both points.

What's the single best test of an AI consultant?

A live demo on twenty of your own messy records, watched in real time. RAND found models are often optimized for the wrong metrics or built so they don't fit the actual workflow. A clean canned demo hides that. Your worst invoices and hardest voicemails will surface it in minutes.

Why do so many small businesses struggle to get AI into operations?

Because using AI and running on it are different problems. In 2026, Goldman Sachs found 76% of the small businesses it surveyed use AI, while only 14% had embedded it in core operations. Some 73% wanted more implementation support. Census data puts adoption lower still across the wider population. Tools are easy to buy. Workflow change and adoption are the hard part.

Is a bigger consulting firm safer if I'm not technical?

Not necessarily, and size doesn't fix the translation gap. In 2025, Eurostat found 70.89% of EU enterprises that considered AI and decided against it cited lack of expertise as their reason for not adopting. What closes that gap is someone embedded in your operation who explains their work in plain language, whatever their firm's size.

Sources


About Avolis Research Group

Avolis Research Group is Avolis's in-house research practice, focused on how operations-heavy small and mid-sized businesses actually adopt AI. It synthesizes primary economic research, government survey data, and results from real implementations into practical, vendor-neutral guidance.

More about Avolis and how we work · Get in touch

Continue Learning

Ready to make AI work for you?

Book an AI readiness evaluation. If there’s nothing worth automating, we’ll tell you.