How to Choose a Consulting Partner for AI-Driven Process Automation
Avolis Research Group
·
·
17 min read
Forrester predicted process intelligence would rescue 30% of failed AI projects in 2026. How to choose a process automation partner on the work, not the model.
Choose the partner who maps how your work actually runs before they name any technology, and then make them show you what happens on the cases the automation gets wrong, because in November 2025 Forrester predicted that process intelligence would rescue 30% of failed AI projects during 2026 (Forrester, Predictions 2026: Automation At The Crossroads).
Process intelligence means software that reads the logs in the systems you already run and reconstructs how work actually flows. It's the machine version of standing behind someone and watching.
It's a forecast, not a measurement. But look at what it assumes: a large share of failed AI work is recoverable by understanding the process, not by changing the technology.
That assumption is the whole basis for choosing well here. AI-driven process automation is a process project wearing a technology project's clothes, which is why the firm that wins the demo so often loses the build. Put another way, the questions that predict the outcome are questions about your workflow, not about their stack.
One disclosure before we start: Avolis does this work for a living, so we're not a neutral party. What follows is the evaluation we'd want to be put through ourselves, with the arithmetic shown so you can check it against anyone.
This guide gives you two things most buyer's guides skip: first, the arithmetic of what happens on the cases the system gets wrong, and second, the questions that separate a partner who designs for those cases from one who demos the happy path and then invoices you.
Key Takeaways
- Forrester predicted process intelligence would rescue 30% of failed AI projects in 2026 (Forrester, 2025). It's a forecast, but it puts the fault in the process rather than the model.
- Nobody has solved unsupervised autonomy: the best agent on OSWorld 2.0 finished 20.6% of open-ended multi-hour tasks unaided (2026), which sets a ceiling on autonomy rather than a forecast for your build.
- Workflow redesign beat 24 other attributes McKinsey tested for profit impact, so automating your process unchanged aims low.
- Scope for production, not a pilot, because only 11% of organizations had AI agents in production, against 38% still piloting (Deloitte, 2026).
Table of Contents
- What Actually Separates One Partner From Another?
- How Do They Learn Your Process?
- Will They Redesign the Work or Just Automate It?
- What Happens on the Cases That Go Wrong?
- Who Owns It on the Tuesday After Go-Live?
- Does the Scope Match a Business Your Size?
- Twelve Questions to Ask a Process Automation Partner
- When Should You Walk Away From a Process Automation Proposal?
- Map One Workflow First
- Frequently Asked Questions
- Sources
- Continue Learning
What Actually Separates One Partner From Another?
It isn't the tooling, because when McKinsey's The State of AI tested 25 organizational attributes against EBIT impact from AI in 2025 (McKinsey, The State of AI in 2025), workflow redesign had the biggest effect of any of the 25. EBIT is operating profit, before interest and tax, which is the line that tells you whether something actually paid. Roughly 39% of organizations reported any EBIT impact at all, and only about 6% qualified as high performers. Almost every firm you talk to will have access to the same models and the same integration platforms, so the models aren't the moat. What differs is how much of the engagement goes into understanding and reshaping the work before anything gets built.
So you're really assessing five things:
- Process discovery. How they figure out what actually happens, not what the written procedure says.
- Redesign nerve. Whether they'll change the process or just automate it as-is.
- Exception design. What the system does when it's unsure or wrong.
- Ownership. Who runs and repairs this in month six.
- Scope fit. Whether the proposal is sized for your operation.
Each one is testable in a conversation before you sign anything, and the rest of this article shows you how to test them. If you're earlier in the decision and still comparing types of firm, start with which AI consulting company you should choose.
How Do They Learn Your Process?
A good partner learns your process by watching how it actually runs, not by asking about it. In February 2026, Gartner published research titled Understand Your Processes Before Investing in Agentic Automation (Gartner) that makes the same point. The title is the finding: Gartner's guidance is to sort out what your processes actually are first, and only then decide how much automation the work needs. A note on the word agentic, since you'll hear it constantly: an agent is software given a goal and left to work out the steps itself, rather than following rules someone wrote down. It sits at the most capable and least predictable end of this market, and plenty of useful process automation isn't agentic at all.
A good discovery phase produces a map of the real process, showing where work arrives from, who touches it, what they check, what they decide by feel, and where it stalls. That map is a deliverable you should be able to keep, whatever you decide next.
Here's the test: ask how many hours they'll spend with the people who do the job, meaning the dispatcher who builds the day's schedule or the office manager who reconciles supplier invoices. A firm that answers in days is doing discovery, while a firm that answers with "a kickoff workshop" is doing sales.
Our finding: In one general-contracting build, a multimillion-dollar trades business runs two service lines with a team of three, on a custom system it owns outright at $0 per-seat licensing (Avolis, General contracting case study). We cost the piecemeal equivalent (a custom CRM build, AI training, and a back-office admin seat) at roughly $130K–$250K in year one. That's our estimate of the alternative rather than a measured saving, because it excludes what the build itself cost, which varies by scope. The result came from redesigning how quoting and job flow work rather than from buying more software (this is our own client work, not independent research).
Also from that work: In our experience, the documented process and the running process are almost never the same thing, because the written version covers the clean case while the real version includes the three customers who get invoiced differently, the supplier whose confirmations arrive by text, and the step one person does from memory. Those exceptions aren't noise around the process; very often they are the process, and they're what breaks an automation built from the written procedure.
That gap is why the Forrester number lands where it does, since process intelligence rescues projects by finding the work nobody wrote down.
Will They Redesign the Work or Just Automate It?
Ask directly, then listen for whether the process changes shape. Only about 6% of organizations qualified as AI high performers in McKinsey's 2025 survey (McKinsey, The State of AI in 2025), and most of that group were redesigning workflows rather than layering AI onto existing ones. Qualifying meant attributing 5% or more of EBIT to AI and reporting significant value.
Deloitte's 2026 read on agentic strategy says the same thing in one line (Deloitte, Tech Trends 2026: Agentic AI strategy): value comes from redesigning operations, not from putting agents on top of old workflows.
There's an old term for automating a process without changing it, paving the cowpath, and what you get is a faster version of a route that was never sensible.
In practice, redesign usually means fewer steps rather than smarter ones: approvals that exist because someone, once, was slow, data the estimator re-enters because two systems never talked (a habit nobody chose and everybody inherited), and checks that only exist because an earlier step is unreliable. A partner worth hiring will propose deleting some of that, whereas a vendor will automate all of it, because more steps mean more scope. One caution, since we're arguing for redesign: it isn't free. It has a cost your team pays in attention, because redesign means retraining people and rewriting how the job is done. If a firm proposes redesign without budgeting for that, they've only done half the thinking.
What Happens on the Cases That Go Wrong?
This is the question almost nobody asks, and it's the one that decides whether the thing survives contact with your business. In 2026, the OSWorld 2.0 benchmark tested agents on 108 realistic computer workflows, each taking a person a median of about 1.6 hours, and the best agent completed just 20.6% of them end to end (OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks, arXiv).
Read that number carefully, because it's easy to misuse. OSWorld measures an agent working unsupervised across hours of open-ended computer work, deciding for itself what to do next. A narrow, supervised automation is a different class of problem. Read the email, pull four fields, create the job: that kind of work is bounded, repeatable, and checked. So the benchmark tells you where unsupervised autonomy currently tops out, and it doesn't predict your build's accuracy in either direction.
What it tells you is that nobody has solved the unsupervised part, so somebody has to design for the cases the system gets wrong, and that somebody should be the firm you hire.
Compare it with the headline numbers. Stanford HAI's 2026 AI Index reported agent accuracy on the original OSWorld rising from roughly 12% in 2024 to 66.3% (Stanford HAI, 2026 AI Index Report), which is within six percentage points of the roughly 72% humans score. That's real progress, and yet, as the AI Index puts it, agents still fail about one attempt in three on structured benchmarks.
Here's why the distinction matters more than it looks: older process automation was deterministic, so a rules engine either ran or visibly broke, and you found out the same day. AI-driven process automation is probabilistic, which means it can be wrong quietly, on a small share of cases nobody notices for weeks. That changes what a partner owes you, because the deliverable isn't "the automation works" (on some percentage of cases it won't) but a designed path for those cases. In other words, you're owed a confidence threshold that decides what the system won't attempt on its own, a review queue where a person checks the doubtful ones, an audit trail, and a named fallback.
Now do the arithmetic the brochure skips, using figures that are illustrative rather than a benchmark. Say the workflow runs 200 times a month and the build is right 95% of the time, which leaves 10 cases a month landing somewhere. If that somewhere is a queue the office manager clears in twenty minutes, you have a working system. If it's nowhere, you have ten silent errors a month compounding into a mess someone finds in the spring. It's the same accuracy and a completely different outcome, and the difference is entirely design, not technology. So ask for the accuracy target and the exception volume it implies, because you need both numbers, not one. If a firm can't sketch that path on a whiteboard, they haven't designed it. From our work, this is the section most often missing from a proposal, and it's also the one that decides whether a build survives its first odd month.
Two moves make this concrete during evaluation. The first is to ask them to demo a failure rather than a success, because any demo runs the happy path. The second is to hand them five of your genuinely weird cases and ask what the system would do with each. Take a lead-intake build, for example: what happens when a customer sends the site address as a photo of a business card, or asks for two separate jobs in one message? Vague answers to cases like those are the single most reliable signal you'll get. For more on the criteria specific to automation scoping, see how to choose an AI provider for business automation.
Who Owns It on the Tuesday After Go-Live?
Name that person before you sign. In November 2025, Forrester predicted that in 2026 fewer than 15% of firms would switch on the agentic features already inside their automation suites (Forrester). The capability was bought and then left off, which makes it an ownership failure, not a technology one.
Automated processes drift, because your supplier changes a form, a pricing rule changes, or volume doubles in spring. Someone has to notice, diagnose, and fix the problem, and if that person isn't named, the answer defaults to whoever complains loudest.
Four ownership clauses belong in the agreement, on top of the scope itself:
- A named owner on your side, with the hours to do it acknowledged out loud.
- What the partner covers after go-live, and for how long, in writing.
- How a problem gets raised, and what response time you can expect.
- What you keep if you part ways: the process map, the configuration, the credentials, the documentation.
That last one deserves a hard look, so ask what happens if you end the relationship in month nine. A partner will tell you what transfers, and a vendor will change the subject. Deloitte's Tech Trends 2026 reported that 35% of organizations had no formal strategy for AI agents at all (Deloitte, Tech Trends 2026: Agentic AI strategy), and no strategy usually means no owner after launch.
Does the Scope Match a Business Your Size?
Most published guidance on this decision was written for enterprises, and your economics are different. In 2026, the U.S. Census Bureau's Business Trends and Outlook Survey found AI use rising steeply with headcount (U.S. Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users). It reaches about 37% among firms with 250 or more employees and stays under 20% among firms with fewer than 20 employees. Census data also shows 57% of adopting firms use AI in three or fewer business functions. There's a sharper finding underneath. In 2026, the U.S. Chamber of Commerce Foundation's Main Street AI Monitor looked at what small businesses use AI for (U.S. Chamber of Commerce Foundation, Main Street AI Monitor) and found that only 26% of that use touches recurring workflows, while some 64% is personal productivity.
That's the actual gap in the market, because lots of small businesses use AI (for drafting, for summarising, for the odd bit of research) but very few have it doing repeated work inside a process, like the service coordinator's daily job intake or the bookkeeper's month-end reconciliation. Moving from the first kind of use to the second is the entire job you're hiring for.
One data point argues for specialists, and we'll flag our own interest in it. Buying from specialized vendors succeeded roughly 67% of the time in MIT Project NANDA's 2025 study (MIT Project NANDA, The GenAI Divide: State of AI in Business 2025), which is about three times the rate of internal builds. Treat it as suggestive rather than settled, because the same report's headline claim about pilot failure rates drew serious methodological criticism. And we're a specialist firm citing a finding that flatters specialists. Either way, a company-wide program pitched at a 40-person operation is an enterprise engagement wearing a smaller label.
Twelve Questions to Ask a Process Automation Partner
Ask these before you see a proposal, because each one tests something specific, and the pattern of answers tells you more than any reference call. For the general version of this test, see what makes a good AI consulting partner versus a typical vendor.
| Ask this | What the answer reveals |
|---|---|
| How many hours will you spend watching the people who do this job? | Whether discovery is real work or a workshop |
| What will you give us that documents how the process runs today? | Whether you keep an asset or just get a build |
| Which steps would you delete rather than automate? | Redesign nerve versus paving the cowpath |
| What does the system do when it isn't confident? | Whether exception handling exists at all |
| Show me the failure case, not the demo. | Whether they've run this in anger before |
| Here are five odd cases from last month. What happens to each? | How specific their thinking gets under pressure |
| Who reviews the output, how often, and for how long? | Whether a human stays in the loop by design |
| What integration work is in scope, and what's discovered later? | Where the budget overrun will come from |
| What percentage of cases do you expect to route to human review in month one? | Whether they have modelled exceptions at all |
| What breaks this, and how would we find out? | Honesty, and whether they've thought about drift |
| If we end this in month nine, what do we keep? | Lock-in, plainly |
| Which of our systems will you read from, and which will you write to? | Whether they have looked at your stack |
Our finding: We've sat on both sides of these conversations, and the questions that do the work are the ones about exceptions and ownership, not the ones about capability. Everyone can answer whether they can do the job, but far fewer can say what happens when it's wrong, and who fixes it, without going quiet. We'd rather be asked the second kind, though this is an impression from our own engagements, not a tally.
If you want the same evaluation without the jargon, the version for non-technical founders choosing an AI consulting partner covers this ground more slowly.
When Should You Walk Away From a Process Automation Proposal?
Walk when the proposal arrives before the process is understood. Gartner's 2025 forecast put more than 40% of agentic AI projects as canceled by the end of 2027 (Gartner), citing escalating costs, unclear business value, and inadequate risk controls. It's a forecast, not a measurement, but all three are things you can check for in a sales conversation.
The signals worth walking on:
- A fixed price before any discovery. They're pricing a product, not your process.
- No answer on exceptions. It's the most expensive gap, and the easiest to spot.
- A named tool in the first meeting. The stack should follow the process, not lead it.
- Integration described as "straightforward." In 2024, RAND found data preparation and engineering consume the bulk of real AI effort (RAND Corporation, The Root Causes of Failure for AI Projects and How They Can Succeed). In 2026, Deloitte's Tech Trends 2026 found data searchability and reusability named as obstacles by 48% and 47% of organizations (Deloitte, Tech Trends 2026: Agentic AI strategy).
- A headcount-savings business case. If the payoff depends on letting people go, you've bought a redundancy plan, not a process improvement.
That last one matters to us, so we'll be plain about it. The point of taking paperwork off your crews is that they spend their hours on the work only they can do, so if a proposal's arithmetic only works by removing people, ask what it delivers if you keep them.
Map One Workflow First
Start narrow by picking one workflow and writing down what actually happens in it, including the exceptions that people like the dispatcher or the estimator handle from memory. That document is the most useful thing you can bring to any conversation with any firm. It costs you a week of attention, and it changes every meeting afterward.
If you'd rather have that mapped properly, that's what a diagnostic is for: a couple of weeks working through the workflow with the people who run it, measuring how it runs and mapping where it stalls, rather than a workshop. And if the process isn't worth automating yet, we'll tell you that. Talk to us about a readiness assessment.
Frequently Asked Questions
What's the difference between an AI automation partner and an RPA vendor?
RPA (robotic process automation) follows rules someone wrote down, so it runs or it visibly breaks, while AI-driven automation judges each case and can be wrong quietly instead. In 2026 no agent passed 21% on OSWorld 2.0's unsupervised open-ended tasks, which is a ceiling on autonomy rather than your error rate, and it's why exception design matters more here.
How long should process discovery take before anything gets built?
Long enough to see the work happen, which is usually one to three weeks on a single workflow. Forrester predicted process intelligence, software that reconstructs workflows from system logs, would rescue 30% of failed AI projects in 2026, though that's a forecast, not a measurement. A fixed price after one kickoff call means they haven't seen your exceptions.
Should we automate our existing process or redesign it first?
Redesign it first in most cases. In 2025, McKinsey found workflow redesign had the biggest effect on EBIT impact of 25 attributes tested, and most high performers were redesigning rather than layering AI on. Automating an unchanged process gives you a faster version of a route that was never sensible.
What should be in the contract for an AI process automation project?
Six things belong in it: the process map as an appendix, one primary metric, a defined exception path, a named owner on each side, post-launch support terms, and what transfers if you part ways. In 2026, Deloitte reported 35% of organizations had no formal strategy for AI agents, which usually means nobody owned the system after launch.
Is AI process automation worth it for a business under 50 people?
It often is, but only at narrow scope. In 2026, Census data showed 57% of AI-adopting firms use it in three or fewer functions, and the U.S. Chamber Foundation found just 26% of small-business AI use touches recurring workflows. At that size, one repeated, measurable process beats a company-wide program.
Sources
All sources retrieved 2026-07-28.
- Forrester, Predictions 2026: Automation At The Crossroads, Leslie Joseph, November 2025: https://www.forrester.com/blogs/predictions-2026-automation-at-the-crossroads/
- Stanford HAI, 2026 AI Index Report, Technical Performance chapter, 2026: https://hai.stanford.edu/ai-index/2026-ai-index-report/technical-performance
- OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks, arXiv:2606.29537, June 2026: https://arxiv.org/abs/2606.29537
- McKinsey & Company (QuantumBlack), The State of AI in 2025: Agents, innovation, and transformation, November 2025: https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- Deloitte, Tech Trends 2026: Agentic AI strategy (drawing on the 2025 Emerging Technology Trends study), 2026: https://www.deloitte.com/us/en/insights/topics/technology-management/tech-trends/2026/agentic-ai-strategy.html
- Gartner, Understand Your Processes Before Investing in Agentic Automation, February 2026: https://www.gartner.com/en/documents/7420462
- Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, June 2025: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- RAND Corporation, The Root Causes of Failure for AI Projects and How They Can Succeed, August 2024: https://www.rand.org/pubs/research_reports/RRA2680-1.html
- U.S. Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users (Business Trends and Outlook Survey), May 2026: https://www.census.gov/library/stories/2026/05/ai-use-businesses.html
- U.S. Chamber of Commerce Foundation, Main Street AI Monitor, June 2026: https://www.uschamberfoundation.org/workforce/half-of-small-business-workers-use-ai-most-to-boost-productivity-not-automate-jobs
- MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, July 2025, as reported by Fortune: https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
- Avolis, General contracting: a multimillion-dollar trades business run by three people (first-party client work): https://www.avolis.ai/resources/case-studies/general-contractor
About Avolis Research Group
Avolis Research Group is Avolis's in-house research practice, focused on how operations-heavy small and mid-sized businesses actually adopt AI. It synthesizes primary economic research, government survey data, and results from real implementations into practical, vendor-neutral guidance.
More about Avolis and how we work · Get in touch
Continue Learning
- Which AI consulting company should you choose?: the full decision guide this article sits under.
- What makes a good AI consulting partner vs. a typical vendor?: the six general signals, without the process-automation specifics.
- How to choose an AI provider for business automation: scoping criteria for automation work.
- Choosing an AI consulting partner for non-technical founders: the same decision without the jargon.
- What do AI consultants charge?: pricing models and what drives the number.
- AI readiness assessment services: how a diagnostic maps the process before anything gets built.
- How to choose an AI consulting service for business ROI: the arithmetic behind proving the payoff.
Ready to make AI work for you?
Book an AI readiness evaluation. If there’s nothing worth automating, we’ll tell you.
