Ninety-seven percent of midmarket leaders say they are satisfied with their AI return on investment, according to RSM’s survey of more than a thousand midmarket executives. Budgets are climbing, adoption is mainstream, and on the executive dashboard, AI appears to be working.
Look one layer deeper, and the operational reality tells a different story.
Only 36% of those same organizations have AI embedded across core processes. Kauffman Rossin’s study of midmarket C-levels puts a finer point on what that "return" actually consists of: 95% cite time savings as the primary benefit, while only 26% can point to revenue gained. Meanwhile, Grant Thornton finds that just 14% of midmarket firms have an AI strategy actually implemented in operations, and 70% report that at least half their core applications aren't AI-ready.
So which is it — is AI delivering, or isn't it? Both numbers are true. They are just measuring different things. The 97% represents program satisfaction reported by the executives who approved the budget. The time-not-revenue split and the implementation gaps represent the operational reality underneath it.
Satisfaction is a leading indicator of nothing. Only your foundation determines whether the P&L actually sees the impact.
Why the number isn't moving
Search "AI ROI" and you'll find dozens of frameworks from Deloitte, IBM, PwC, and the rest — all variations on the same advice: define your metrics, track usage, calculate cost savings, iterate. It's sound advice. But, if measurement was the problem, better measurement frameworks would have fixed it by now.
Measurement was never the primary constraint. You can measure a number precisely and watch it stay flat if nothing underneath changes. The midmarket data makes this plain: the value AI is actually producing shows up as time saved and smoother workflows — Kauffman Rossin finds 82% of companies measure AI by time saved, only 26% by revenue gained — while the structural readiness to turn that into financial outcomes is missing. Grant Thornton finds just 14% of midmarket firms have an AI strategy actually implemented in operations, and 70% say at least half their core applications aren't AI-ready. Frameworks assume AI is deployed onto a foundation ready to convert outputs into outcomes. For most companies, that foundation isn't there yet.
The three fractures keeping the number flat
Across the pilots that stall without moving anything, the same three gaps show up, in some combination, almost every time.
- It's less about dirty data than about disconnected context. You don't need a multi-year data-cleansing project. The AI doesn't need perfect data — it needs a way to reach the operational facts already sitting in your systems: the CRM notes, the ticket threads, the inbox nobody exports. When that context is scattered across three tools and a spreadsheet, the AI fills the gaps with a plausible-sounding but ultimately useless answer. No amount of measurement rescues an output built on a broken input.
- Adoption happened in isolation. One team gets real value; everyone else tried it twice, got a mediocre result, and went back to the old way. The ROI math then averages a genuine win against a dozen non-adoptions, and the blended number looks like nothing happened — because for most of the business, nothing did.
- There's no memory that carries over. Every session starts from zero: the same context re-explained, the same corrections re-made, the same institutional knowledge that never got captured anywhere the AI could use it. Without a curated library of what the business actually knows, the AI keeps re-solving the same problems, forgetting the corrections you made to your imperfect data, and burning the hours that were supposed to be the return.
None of these are model problems. They are foundation problems, which is why throwing a smarter model at a broken workflow never changes the P&L
Narrow scope, real foundation, then scale what worked
There's a clear pattern in the minority of programs that move a financial number — Kauffman Rossin's "Operators," the 2% where AI is genuinely how the business runs, or the firms NCMM finds outgrowing their peers. They pick one pain point, in one part of the business, and get the foundation right. That means a single workflow using the right engine for the task, and a human checkpoint on the output until trust is earned.
Measure outcomes, not usage
Measuring tokens used, seats, or sessions does absolutely nothing for your bottom line — unless you're an AI vendor. Instead, focus solely on metrics that reflect operational speed and cost.
| Stop measuring (usage) | Start measuring (outcome) |
|---|---|
| Seats activated, license utilization | Cycle time per output (e.g. hours per quote) |
| Prompt volume, login counts | Error / rework rate on completed tasks |
| Session duration | Cost per unit of work delivered |
| Self-reported "time saved" | The business metric the pilot was meant to move |
Baseline your chosen metric the way you already measure it, then track it over 90 days. If the number doesn't move, the spend is wasted.
Find your fractures. Fix your foundation.
Projects don't fail to deliver ROI because AI doesn't work. They fail because companies measure the wrong layer at scale before the foundation exists to support it.
Pick one pain point in one silo. Build a saved-answers layer, deploy the right engine for the workflow, keep a human checkpoint on outputs until trust is earned, and prove the number shifts in 90 days before touching anything else.
This is the work we do at PARALLAX: a fixed-fee diagnostic to find what's blocking your ROI, followed by outcome-aligned implementation in one silo. No platform rollouts or twelve-month transformation programs, just solid, provable value.
