Many CFOs and functional leaders share the frustration of viewing dashboards that report all sorts of AI-driven efficiencies while not seeing those efficiencies translate into real reductions in cost or headcount. In previous posts, I’ve discussed some of the issues with AI productivity gains and ROI, but I realized I hadn’t addressed a central assumption underlying much of the disconnect between business leaders’ expectations for AI-driven efficiencies and the reality: the AI Substitution Fallacy.

The AI Substitution Fallacy is the assumption that AI can replace human work in a business process while holding the quality of the resulting work constant. This assumption helps explain why many organizations measure impressive local efficiency gains from AI yet realize little or no actual business value in the aggregate.

Anyone who has deployed AI at scale has seen this firsthand. AI rarely produces work of identical quality to humans. Sometimes it’s better. Sometimes it’s worse. But it’s almost never the same.

Of course “quality” is not a single characteristic. AI may outperform humans on some dimensions while underperforming on others. It may be more consistent but less adaptable, more precise but less accurate, or faster but less capable of nuanced judgment or common sense. The importances of those differences depends entirely on the business process. Yet many organizations still describe their AI strategy primarily in terms of substituting AI for human labor and treating the resulting reduction in effort as an efficiency gain or cost savings. The math almost never that simple. Unless the underlying business process is largely insensitive to quality (which is uncommon), or AI consistently exceeds the quality threshold required by that process, differences between AI and human output inevitably affect the downstream value of the overall business process.

Quality-insensitive processes certainly exist — simple classification, basic data entry, first-drafts of documents nobody will read closely — but most business processes that organizations care enough to measure are quality-sensitive by definition. If not, there wouldn’t be reviews or approval gates. The presence of those checkpoints is itself evidence that someone, somewhere, is already paying attention to quality. And if there’s an end-customer as well, someone either inside or outside the process will notice when quality changes.

Many organizations are discovering this reality as they implement AI use cases that appear to raise productivity but ultimately deliver little additional business value. Consider the examples of agentic coding and ITSM I addressed in my previous essay. If AI coding assistants produce higher error rates than experienced developers, code reviews will be longer and costlier, and the resulting code will be less valuable. If AI-powered ITSM tools provide lower quality support, the result will be escalation to higher (and more expensive) tiers of support or lower customer satisfaction.

This isn’t an argument against general-purpose GenAI as a knowledge tool. In many personal productivity tasks, the difference in quality is negligible and the benefits of speed/reduction of friction easily outweigh quality differences. It’s about what happens when AI is brought into existing, high-volume processes I described in a separate essay on ROI. Those processes already have built in standards for quality – even if only implicitly – which means that quality differences will ultimately show up in the overall workflow and the underlying economics, for better or for worse.

In quality sensitive processes, local productivity gains can be misleading. AI doesn’t simply substitute human labor – it substitutes work of a different quality level. And because workflows and business processes convert both labor and quality into value, changes in quality affect the overall economics of the workflow. What appears locally as an efficiency can lead to a shifting of costs elsewhere – or an overall lowering of the value of the output.

Of course quality cuts both ways – sometimes AI or AI+human produces higher quality output. And when that happens, it can create opportunities. For example, if property harnessed, the quality differences of AI-enabled work can mitigate the bottlenecks that often show up in workflows where AI assistance leads to higher upstream outputs. Or harnessed to produce outputs of higher overall value.

Consider AI-assisted SW development. As AI assistance makes code faster to write – or takes over coding entirely – many SW organizations discover that the bottleneck simply moves downstream (https://about.gitlab.com/resources/ai-accountability-survey-2026/?utm_source=chatgpt.com). But as this report also shows, AI-driven code quality is also increasing, creating an opportunity (or perhaps a necessity) to rethink the review process itself. Higher-quality code may require less intensive reviews, freeing up reviewers to focus on high-risk/high-impact portions of code while automating more of the review of less important sections with independent tools. Those changes can ultimately drive higher overall output – and greater value – across the entire SDLC.

The practical corollary: before crediting AI with creating value through an efficiency gain, ask where and how the quality difference is affecting the overall process. If you can’t answer that, you don’t know whether value is actually being created, costs are being relocated to a place you’re not measuring, or perhaps additional value is being generated that you’re not even aware of.