Which AI Consulting Company Should You Choose?
Avolis Research Group
·
·
16 min read
In 2025, 95% of enterprise AI pilots showed no measurable payoff. Here's how to choose an AI consulting company that ships working systems, not decks.
Choose the partner who maps how your work actually runs before they name any technology. Then make them show you what happens on the cases the automation gets wrong, because in November 2025 Forrester predicted that process intelligence would rescue 30% of failed AI projects during 2026 (Forrester, Predictions 2026: Automation At The Crossroads).
Process intelligence means software that reads the logs in the systems you already run and reconstructs how work actually flows. It's the machine version of standing behind someone and watching.
It's a forecast, not a measurement. But look at what it assumes. A large share of failed AI work is recoverable by understanding the process, not by changing the technology.
That's the whole basis for choosing well here. AI-driven process automation is a process project wearing a technology project's clothes, which is why the firm that wins the demo so often loses the build. So the questions that predict the outcome are questions about your workflow, not their stack.
One disclosure before we start: Avolis does this work for a living, so we're not a neutral party. What follows is the evaluation we'd want to be put through, with the arithmetic shown so you can check it against anyone.
This guide gives you two things most buyer's guides skip. First, the arithmetic of what happens on the cases the system gets wrong, and second the questions that separate a partner who designs for that from one who demos the happy path and invoices you.
Key Takeaways
- Forrester predicted process intelligence would rescue 30% of failed AI projects in 2026 (Forrester, 2025). A forecast — but it puts the fault in the process, not the model.
- Nobody has solved unsupervised autonomy. The best agent on OSWorld 2.0 finished 20.6% of open-ended multi-hour tasks unaided (2026) — a ceiling on autonomy, not a forecast for your build.
- Workflow redesign beat 24 other attributes McKinsey tested for profit impact. Automating your process unchanged aims low.
- Scope for production, not a pilot. Only 11% of organizations had AI agents in production, against 38% still piloting (Deloitte, 2026).
Table of Contents
- What Actually Separates One Partner From Another?
- How Do They Learn Your Process?
- Will They Redesign the Work or Just Automate It?
- What Happens on the Cases That Go Wrong?
- Who Owns It on the Tuesday After Go-Live?
- Does the Scope Match a Business Your Size?
- Twelve Questions to Ask a Process Automation Partner
- When Should You Walk Away From a Process Automation Proposal?
- Map One Workflow First
- Frequently Asked Questions
- Sources
- Continue Learning
What Actually Separates One Partner From Another?
Not the tooling. In 2025, McKinsey's The State of AI tested 25 organizational attributes against EBIT impact from AI (McKinsey, The State of AI in 2025). EBIT is operating profit, before interest and tax — the line that tells you whether something paid. Workflow redesign had the biggest effect of any of the 25. Roughly 39% of organizations reported any EBIT impact at all. About 6% qualified as high performers. Almost every firm you talk to will have access to the same models and the same integration platforms. The models aren't the moat. What differs is how much of the engagement goes into understanding and reshaping the work before anything gets built.
So you're really assessing five things:
- Process discovery. How they figure out what actually happens, not what the written procedure says.
- Redesign nerve. Whether they'll change the process or just automate it as-is.
- Exception design. What the system does when it's unsure or wrong.
- Ownership. Who runs and repairs this in month six.
- Scope fit. Whether the proposal is sized for your operation.
Each one is testable in a conversation before you sign anything. The rest of this article is how to test them. If you're earlier in the decision and still comparing types of firm, start with which AI consulting company you should choose.
How Do They Learn Your Process?
By watching it, not by asking about it. In February 2026, Gartner published research titled Understand Your Processes Before Investing in Agentic Automation (Gartner). The title is the finding. Gartner's guidance is to sort out what your processes actually are first, then decide how much automation the work needs. A note on the word agentic, since you'll hear it constantly. An agent is software given a goal and left to work out the steps itself, rather than following rules someone wrote down. It's the most capable and the least predictable end of this market. Plenty of useful process automation isn't agentic at all.
A good discovery phase produces a map of the real process. Where work arrives from, who touches it, what they check, what they decide by feel, and where it stalls. That map is a deliverable you should be able to keep.
Here's the test. Ask how many hours they'll spend with the people who do the job. A firm that answers in days is doing discovery. A firm that answers with "a kickoff workshop" is doing sales.
Our finding: In one general-contracting build, a multimillion-dollar trades business runs two service lines with a team of three, on a custom system it owns outright at $0 per-seat licensing (Avolis, General contracting case study). We cost the piecemeal equivalent — a custom CRM build, AI training, and a back-office admin seat — at roughly $130K–$250K in year one. That's our estimate of the alternative, not a measured saving: it excludes what the build itself cost, which varies by scope. That came from redesigning how quoting and job flow work, not from buying more software. (Our own client work, not independent research.)
Also from that work: In our experience, the documented process and the running process are almost never the same thing. The written version covers the clean case. The real version includes the three customers who get invoiced differently. The supplier whose confirmations arrive by text. The step one person does from memory. Those exceptions aren't noise around the process. Very often they are the process, and they're what breaks an automation built from the written procedure.
That gap is why the Forrester number lands where it does. Process intelligence rescues projects because it finds the work nobody wrote down.
Will They Redesign the Work or Just Automate It?
Ask directly, then listen for whether the process changes shape. Only about 6% of organizations qualified as AI high performers in McKinsey's 2025 survey (McKinsey, The State of AI in 2025). That means attributing 5% or more of EBIT to AI and reporting significant value. Most of that group were redesigning workflows, not layering AI onto existing ones.
Deloitte's 2026 read on agentic strategy says the same thing in one line (Deloitte, Tech Trends 2026: Agentic AI strategy). Value comes from redesigning operations, not from putting agents on top of old workflows.
There's an old term for automating a process without changing it: paving the cowpath. You get a faster version of a route that was never sensible.
In practice, redesign usually means fewer steps rather than smarter ones. Approvals that existed because someone, once, was slow. Data re-entered because two systems never talked, a habit nobody chose and everybody inherited. And checks that only exist because an earlier step is unreliable. A partner worth hiring will propose deleting some of that. A vendor will automate all of it, because more steps mean more scope. One caution, since we're arguing for redesign: it has a cost your team pays in attention. It isn't free. Redesign means retraining people and rewriting how the job is done. If a firm proposes redesign without budgeting for that, they've only done half the thinking.
What Happens on the Cases That Go Wrong?
This is the question almost nobody asks, and it's the one that decides whether the thing survives contact with your business. In 2026, the OSWorld 2.0 benchmark tested agents on 108 realistic computer workflows, each taking a person a median of about 1.6 hours, and the best agent completed just 20.6% of them end to end (OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks, arXiv).
Read that number carefully, because it's easy to misuse. OSWorld measures an agent working unsupervised across hours of open-ended computer work, deciding for itself what to do next. A narrow, supervised automation is a different class of problem. Read the email, pull four fields, create the job — bounded, repeatable, and checked. So the benchmark tells you where unsupervised autonomy currently tops out. It doesn't predict your build's accuracy, in either direction.
What it tells you is that nobody has solved the unsupervised part, so somebody has to design for the cases the system gets wrong, and that somebody should be the firm you hire.
Compare it with the headline numbers. Stanford HAI's 2026 AI Index reported agent accuracy on the original OSWorld rising from roughly 12% in 2024 to 66.3% (Stanford HAI, 2026 AI Index Report). That's within six percentage points of the roughly 72% humans score. Real progress. And as the AI Index puts it, agents still fail about one attempt in three on structured benchmarks.
Here's why the distinction matters more than it looks. Older process automation was deterministic: a rules engine either ran or it visibly broke, and you found out the same day. AI-driven process automation is probabilistic, so it can be wrong quietly, on a small share of cases nobody notices for weeks. That changes what a partner owes you. The deliverable isn't "the automation works," because on some percentage of cases it won't. The deliverable is a designed path for those cases. That means a confidence threshold, which decides what the system won't attempt on its own. A review queue, where a person checks the doubtful ones. An audit trail, and a named fallback.
Now do the arithmetic the brochure skips. The figures below are illustrative, not a benchmark. Say the workflow runs 200 times a month and the build is right 95% of the time. That's 10 cases a month landing somewhere. If that somewhere is a queue one person clears in twenty minutes, you have a working system. If it's nowhere, you have ten silent errors a month compounding into a mess someone finds in the spring, so it's the same accuracy and a completely different outcome — and the difference is entirely design, not technology. So ask for the accuracy target and the exception volume it implies. Both numbers, not one. If a firm can't sketch that path on a whiteboard, they haven't designed it. From our work, this is the section most often missing from a proposal. It's also the one that decides whether a build survives its first odd month.
Two moves make this concrete during evaluation. First, ask them to demo a failure, not a success — any demo runs the happy path. Second, hand them five of your genuinely weird cases and ask what the system would do with each. For example, take a lead-intake build. What happens when a customer sends the site address as a photo of a business card? Or asks for two separate jobs in one message? Vague answers here are the single most reliable signal you'll get. For more on the criteria specific to automation scoping, see how to choose an AI provider for business automation.
Who Owns It on the Tuesday After Go-Live?
Name that person before you sign. In November 2025, Forrester predicted that in 2026 fewer than 15% of firms would switch on the agentic features already inside their automation suites (Forrester). The capability was bought and then left off. That's an ownership failure, not a technology one.
Automated processes drift. Your supplier changes a form. A rule changes. Volume doubles in spring. Someone has to notice, diagnose, and fix — and if that person isn't named, the answer defaults to whoever complains loudest.
Four ownership clauses belong in the agreement, on top of the scope itself:
- A named owner on your side, with the hours to do it acknowledged out loud.
- What the partner covers after go-live, and for how long, in writing.
- How a problem gets raised, and what response time you can expect.
- What you keep if you part ways — the process map, the configuration, the credentials, the documentation.
That last one deserves a hard look. Ask what happens if you end the relationship in month nine. A partner will tell you what transfers. A vendor will change the subject. Deloitte's Tech Trends 2026 reported that 35% of organizations had no formal strategy for AI agents at all (Deloitte, Tech Trends 2026: Agentic AI strategy). No strategy usually means no owner after launch.
Does the Scope Match a Business Your Size?
Most published guidance on this decision was written for enterprises, and your economics are different. In 2026, the U.S. Census Bureau's Business Trends and Outlook Survey found AI use rising steeply with headcount (U.S. Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users). It's about 37% among firms with 250 or more employees, and under 20% among firms with fewer than 20 employees, while Census data also shows 57% of adopting firms use AI in three or fewer business functions. There's a sharper finding underneath. In 2026, the U.S. Chamber of Commerce Foundation's Main Street AI Monitor looked at what small businesses use AI for (U.S. Chamber of Commerce Foundation, Main Street AI Monitor). Only 26% of that use touches recurring workflows. Some 64% is personal productivity.
That's the actual gap in the market. Lots of small businesses use AI — for drafting, for summarising, for the odd bit of research — but very few have it doing repeated work inside a process. Moving from the first to the second is the entire job you're hiring for.
One data point argues for specialists, and we'll flag our own interest in it. Buying from specialized vendors succeeded roughly 67% of the time in MIT Project NANDA's 2025 study (MIT Project NANDA, The GenAI Divide: State of AI in Business 2025). That's about three times the rate of internal builds. Treat it as suggestive rather than settled. The same report's headline claim about pilot failure rates drew serious methodological criticism. And we're a specialist firm citing a finding that flatters specialists. Either way, a company-wide program pitched at a 40-person operation is an enterprise engagement wearing a smaller label.
Twelve Questions to Ask a Process Automation Partner
Ask these before you see a proposal. Each one tests something specific, and the pattern of answers tells you more than any reference call. For the general version of this test, see what makes a good AI consulting partner versus a typical vendor.
| Ask this | What the answer reveals |
|---|---|
| How many hours will you spend watching the people who do this job? | Whether discovery is real work or a workshop |
| What will you give us that documents how the process runs today? | Whether you keep an asset or just get a build |
| Which steps would you delete rather than automate? | Redesign nerve versus paving the cowpath |
| What does the system do when it isn't confident? | Whether exception handling exists at all |
| Show me the failure case, not the demo. | Whether they've run this in anger before |
| Here are five odd cases from last month. What happens to each? | How specific their thinking gets under pressure |
| Who reviews the output, how often, and for how long? | Whether a human stays in the loop by design |
| What integration work is in scope, and what's discovered later? | Where the budget overrun will come from |
| What percentage of cases do you expect to route to human review in month one? | Whether they have modelled exceptions at all |
| What breaks this — and how would we find out? | Honesty, and whether they've thought about drift |
| If we end this in month nine, what do we keep? | Lock-in, plainly |
| Which of our systems will you read from, and which will you write to? | Whether they have looked at your stack |
Our finding: We've sat on both sides of these conversations. The questions that do the work are the ones about exceptions and ownership, not the ones about capability. Everyone can answer "can you do this?" Far fewer can answer "what happens when it's wrong, and who fixes it?" without going quiet. We'd rather be asked the second kind. This is an impression from our own engagements, not a tally.
Want the same evaluation without the jargon? The version for non-technical founders choosing an AI consulting partner covers this ground more slowly.
When Should You Walk Away From a Process Automation Proposal?
Walk when the proposal arrives before the process is understood. Gartner's 2025 forecast put more than 40% of agentic AI projects as canceled by the end of 2027 (Gartner). The reasons cited were escalating costs, unclear business value, and inadequate risk controls. It's a forecast, not a measurement. But all three are things you can check for in a sales conversation.
The signals worth walking on:
- A fixed price before any discovery. They're pricing a product, not your process.
- No answer on exceptions. The most expensive gap, and the easiest to spot.
- A named tool in the first meeting. The stack should follow the process, not lead it.
- Integration described as "straightforward." In 2024, RAND found data preparation and engineering consume the bulk of real AI effort (RAND Corporation, The Root Causes of Failure for AI Projects and How They Can Succeed). In 2026, Deloitte's Tech Trends 2026 found data searchability and reusability named as obstacles by 48% and 47% of organizations (Deloitte, Tech Trends 2026: Agentic AI strategy).
- A headcount-savings business case. If the payoff depends on letting people go, you've bought a redundancy plan, not a process improvement.
That last one matters to us, so we'll be plain about it. The point of taking paperwork off your crews is that they spend their hours on the work only they can do, so if a proposal's arithmetic only works by removing people, ask what it delivers if you keep them.
Map One Workflow First
Start narrow. Pick one workflow and write down what actually happens in it — including the exceptions people handle from memory. That document is the most useful thing you can bring to any conversation with any firm. It costs you a week of attention and it changes every meeting afterward.
If you'd rather have that mapped properly, that's what a diagnostic is for — a couple of weeks of watching the work, not a workshop. And if the process isn't worth automating yet, we'll tell you that. Talk to us about a readiness assessment.
Frequently Asked Questions
What's the difference between an AI automation partner and an RPA vendor?
RPA — robotic process automation — follows rules someone wrote down, so it runs or it visibly breaks. AI-driven automation judges each case and can be wrong quietly instead. In 2026 no agent passed 21% on OSWorld 2.0's unsupervised open-ended tasks. That's a ceiling on autonomy, not your error rate. Exception design matters more here.
How long should process discovery take before anything gets built?
Long enough to see the work happen — usually one to three weeks on a single workflow. Forrester predicted process intelligence, software that reconstructs workflows from system logs, would rescue 30% of failed AI projects in 2026. That's a forecast, not a measurement. A fixed price after one kickoff call means they haven't seen your exceptions.
Should we automate our existing process or redesign it first?
Redesign, in most cases. In 2025, McKinsey found workflow redesign had the biggest effect on EBIT impact of 25 attributes tested. Most high performers were redesigning rather than layering AI on. Automating an unchanged process gives you a faster version of a route that was never sensible.
What should be in the contract for an AI process automation project?
Six things. The process map as an appendix, one primary metric, and a defined exception path. Then a named owner on each side, post-launch support terms, and what transfers if you part ways. In 2026, Deloitte reported 35% of organizations had no formal strategy for AI agents — usually meaning nobody owned the system after launch.
Is AI process automation worth it for a business under 50 people?
Often, but only at narrow scope. In 2026, Census data showed 57% of AI-adopting firms use it in three or fewer functions. The U.S. Chamber Foundation found just 26% of small-business AI use touches recurring workflows. One repeated, measurable process beats a company-wide program at that size.
Sources
All sources retrieved 2026-07-28.
- Forrester, Predictions 2026: Automation At The Crossroads, Leslie Joseph, November 2025 — https://www.forrester.com/blogs/predictions-2026-automation-at-the-crossroads/
- Stanford HAI, 2026 AI Index Report, Technical Performance chapter, 2026 — https://hai.stanford.edu/ai-index/2026-ai-index-report/technical-performance
- OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks, arXiv:2606.29537, June 2026 — https://arxiv.org/abs/2606.29537
- McKinsey & Company (QuantumBlack), The State of AI in 2025: Agents, innovation, and transformation, November 2025 — https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- Deloitte, Tech Trends 2026: Agentic AI strategy (drawing on the 2025 Emerging Technology Trends study), 2026 — https://www.deloitte.com/us/en/insights/topics/technology-management/tech-trends/2026/agentic-ai-strategy.html
- Gartner, Understand Your Processes Before Investing in Agentic Automation, February 2026 — https://www.gartner.com/en/documents/7420462
- Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, June 2025 — https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- RAND Corporation, The Root Causes of Failure for AI Projects and How They Can Succeed, August 2024 — https://www.rand.org/pubs/research_reports/RRA2680-1.html
- U.S. Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users (Business Trends and Outlook Survey), May 2026 — https://www.census.gov/library/stories/2026/05/ai-use-businesses.html
- U.S. Chamber of Commerce Foundation, Main Street AI Monitor, June 2026 — https://www.uschamberfoundation.org/workforce/half-of-small-business-workers-use-ai-most-to-boost-productivity-not-automate-jobs
- MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, July 2025, as reported by Fortune — https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
- Avolis, General contracting: a multimillion-dollar trades business run by three people (first-party client work) — https://www.avolis.ai/resources/case-studies/general-contractor
About Avolis Research Group
Avolis Research Group is Avolis's in-house research practice, focused on how operations-heavy small and mid-sized businesses actually adopt AI. It synthesizes primary economic research, government survey data, and results from real implementations into practical, vendor-neutral guidance.
More about Avolis and how we work · Get in touch
Continue Learning
- Which AI consulting company should you choose? — the full decision guide this article sits under.
- What makes a good AI consulting partner vs. a typical vendor? — the six general signals, without the process-automation specifics.
- How to choose an AI provider for business automation — scoping criteria for automation work.
- Choosing an AI consulting partner for non-technical founders — the same decision without the jargon.
- What do AI consultants charge? — pricing models and what drives the number.
- AI readiness assessment services — how a diagnostic maps the process before anything gets built.
- How to choose an AI consulting service for business ROI — the arithmetic behind proving the payoff.
Ready to make AI work for you?
Book an AI readiness evaluation. If there’s nothing worth automating, we’ll tell you.
