Artificial intelligence is rapidly changing software development, but one of the industry’s most important questions remains surprisingly difficult to answer: is it actually making companies more productive? Developers increasingly use AI assistants to generate code, troubleshoot problems, write tests and accelerate routine engineering work. The experience can feel dramatically faster, yet improvements at the level of an individual programmer do not automatically translate into faster product launches, lower technology costs or better business performance.
That gap was the central argument of a presentation by TechBlocks at AI4 2026, where the technology services company challenged the way enterprises are measuring returns from AI-assisted software development. Its contention is that organisations risk repeating an error seen during previous technology transitions: adopting new tools while leaving the operating model surrounding them largely unchanged. The comparison with cloud computing is instructive. Many early cloud programmes moved existing applications and infrastructure onto external platforms without fundamentally redesigning how those systems operated. The lesson is not that cloud computing failed, but that simply moving an existing system onto a new technology does not necessarily capture the economic value of that technology.
AI may now be creating a similar problem in software engineering. Companies are buying coding assistants, testing tools, security platforms, AI agents, governance systems and analytics products. Each can improve a particular part of the development process, but optimising individual tasks is different from improving the performance of the entire software organisation. Coding represents only one part of software delivery. Engineers also spend considerable time understanding requirements, reviewing existing systems, designing architecture, testing, fixing defects, coordinating releases, dealing with security issues and maintaining software after deployment. Generating code faster can therefore create impressive local productivity gains while leaving many of the surrounding bottlenecks unchanged.
It can even create new ones. Research involving experienced software developers working on repositories they already knew well demonstrated how significant the perception gap can become. Before undertaking the work, developers expected AI to reduce completion time considerably. After using the tools, they still believed they had worked faster. The measured result, however, showed that they actually took longer to complete the assigned tasks. The experiment should not be interpreted as proof that AI coding tools generally reduce productivity, particularly as models continue to improve rapidly. The more important finding is that people can feel substantially more productive while objective measures tell a different story.
Generating code rapidly feels like progress, but time can subsequently be spent checking suggestions, correcting errors or understanding machine-generated changes. That perception gap represents a major management problem because companies have traditionally measured technology teams through activity indicators such as tickets closed, code commits, development hours and releases. AI can increase many of these numbers almost automatically, but more activity is not necessarily more value.
A development team could produce substantially more code while releasing products at the same speed. It could complete more tickets while introducing additional technical debt. It could increase the number of software changes while producing no measurable improvement in revenue, customer experience or operating costs. Recent industry research has pointed towards the same problem. Adoption of AI development tools has increased sharply across hundreds of companies, yet median improvements in software delivery have remained well below some of the dramatic productivity claims associated with generative AI.
The disparity does not necessarily mean the tools have little value. It suggests that the economic benefits may not materialise simply because companies distribute AI licences to developers. The real productivity unit is the software delivery system rather than the individual programmer. Instead of asking how many employees use AI, management needs to ask whether software reaches customers more quickly, defects are declining, engineering costs are improving and technology projects are producing better commercial outcomes.
Those questions also expose weaknesses in traditional technology outsourcing. For decades, much external software development has been sold through time-and-materials contracts. Customers effectively purchase engineering capacity, whether measured in people, hours or development teams. AI potentially disrupts that arrangement. If a software supplier can use AI to complete work much faster while continuing to charge for the same number of development hours, most of the productivity gain remains with the supplier rather than the customer.
The incentives become even more complicated when clients themselves pay for AI licences, cloud resources and model usage while still purchasing essentially the same labour-based delivery model. This creates a fundamental question for technology procurement: if AI genuinely increases productivity, who receives the economic benefit?
TechBlocks argues that software contracts will increasingly need to move from paying for effort towards paying for results. Traditional development contracts allocate much of the delivery risk to the customer. If projects require more work, encounter unexpected complexity or generate defects, the customer often pays for the additional engineering time required to resolve them. An outcome-based model shifts some of that risk back towards the technology provider, making the supplier financially responsible not simply for providing developers but for delivering defined results.
TechBlocks is positioning its AI-native delivery approach around this idea, combining AI-assisted software development, measurement, governance and human engineering oversight rather than presenting the system as simply another coding assistant. The strategy reflects a broader evolution taking place across enterprise technology. AI tools initially entered organisations at the individual level, with employees adopting assistants capable of writing documents, producing code or analysing information. Companies are now discovering that enterprise-scale productivity requires another layer.
Different AI systems need access to appropriate organisational knowledge. Their actions require governance. Work performed by agents and humans needs to be coordinated. Outputs require validation. Costs need to be measured and responsibility for mistakes needs to remain clear. In software engineering this is particularly important because applications rarely exist in isolation. Large organisations may operate systems that have evolved for decades, with decisions made years earlier embedded in databases, integration patterns and business logic that may never have been fully documented.
An AI model reading the code can see what exists without necessarily knowing why it exists. This helps explain why AI can appear extraordinarily capable when creating a new application from scratch while struggling with established enterprise environments. A mature application contains institutional memory. A seemingly unnecessary piece of logic may exist because of a regulatory requirement. An unusual integration may support a major customer. A database structure may accommodate an earlier acquisition. A software dependency may have survived because replacing it would disrupt another system.
Developers who have worked with an application for years may understand these relationships intuitively. AI does not automatically possess that context. Companies will therefore need systems capable of connecting software agents with architectural documentation, historical decisions, business requirements, security rules and operational information. Without that context, AI can generate technically plausible changes that are commercially or operationally wrong.
Governance presents another challenge. As AI generates an increasing share of software, organisations need to determine which changes can be automated and which require human approval. Not every piece of code carries equal risk. An internal reporting tool may tolerate a high degree of autonomous development, while software governing payments, medical information, infrastructure or regulated transactions requires considerably more control.
The future development organisation will probably contain varying levels of AI autonomy. Agents may write code, generate tests, analyse failures and prepare releases, while humans increasingly concentrate on architecture, security, unusual problems and approval of high-risk changes. The objective is not to remove engineers from software development but to redesign how human and machine capabilities are combined.
Measurement becomes equally important. Software organisations possess enormous amounts of operational data, yet connecting engineering work to business value remains difficult. A company can measure lines of code, tickets, deployments and incidents relatively easily. Determining whether a particular software investment increased revenue, reduced operating costs or improved customer retention is much harder.
AI makes resolving this problem increasingly important because traditional activity measures can become misleading. If AI allows an engineer to generate twice as much code, lines of code become an even less useful productivity measure. If an agent can automatically create hundreds of pull requests, the number of pull requests ceases to reveal much about organisational effectiveness. Software management therefore needs to move towards measures such as time from requirement to production, defect rates, system reliability, cost per delivered capability and commercial impact.
This could eventually transform how corporate technology departments are managed. Engineering has historically been treated primarily as a cost centre in many organisations, with budgets often defined by headcount, contractor numbers and infrastructure costs. An outcome-oriented model would treat software more like an investment portfolio. Management would evaluate how much capital is being committed to a particular capability and what business value that capability ultimately generates.
AI could make such measurement easier because digital agents can generate detailed information throughout the development process. Every requirement, AI interaction, software change, test and deployment can theoretically be traced. The challenge is converting that enormous amount of activity data into useful economic information.
This is where TechBlocks sees the opportunity for an AI-native software factory. Its model combines execution of the work, intelligence surrounding the work and governance of the work. AI agents can assist with coding, testing and releases. Measurement systems can follow activity through the development lifecycle, while governance determines which systems and people can make particular decisions.
The commercial component may prove just as important. TechBlocks says its approach can connect payment more directly to engineering outcomes rather than development hours. Whether this model becomes widely adopted remains to be seen, but the underlying direction is significant because AI is weakening the historical relationship between labour hours and software output.
If one engineer equipped with AI can eventually accomplish substantially more than one engineer could previously, charging customers by the hour becomes increasingly disconnected from value. The same disruption is likely to spread beyond software development. Consultancies, law firms, accounting companies and other professional-services businesses face a similar problem because their economics have traditionally depended partly on the number of professional hours required to deliver an assignment.
AI can reduce those hours. Clients will therefore increasingly ask why productivity improvements generated by technology should accrue entirely to the service provider. Outcome-based pricing could become one response. The transition will not be simple because outcomes are harder to define than hours. Software projects frequently change during development, business requirements evolve, external dependencies interfere with delivery and clients themselves can create delays.
Providers accepting greater outcome risk will therefore require more precise contracts and considerably better measurement. AI may also provide some of the infrastructure needed to make those arrangements possible by creating detailed records of who performed work, which tools were used, when requirements changed and what happened after software entered production. That creates the possibility of much more transparent delivery economics.
The implications extend to chief financial officers and boards. Enterprise AI spending is increasingly moving beyond experimentation, and management teams will be expected to demonstrate what financial returns those investments generate. Licence adoption is unlikely to remain an acceptable measure, nor will employee surveys saying that people feel more productive. Companies will need objective evidence that AI is improving operating performance.
This is where the productivity paradox becomes important. AI can make individual tasks dramatically easier without materially changing company-level results if the surrounding organisation remains the same. A developer may write code faster while still waiting days for approval. A customer-service agent may produce responses instantly while underlying customer problems remain unresolved. An analyst may prepare reports more quickly while executives continue making decisions through the same slow governance process.
Productivity consequently depends as much on organisational redesign as on model capability. Companies that merely add AI to existing processes may capture incremental efficiencies. Those willing to redesign workflows around human and machine collaboration could capture much larger gains.
The difference echoes earlier waves of enterprise technology. The internet created far more value when businesses stopped treating websites as digital brochures and began building entirely new digital business models. Cloud computing created greater value when companies stopped thinking of it simply as outsourced infrastructure and redesigned applications around scalable architectures. AI could require an equivalent change.
The important question is therefore not simply how existing work can be completed more quickly. It is how work should be organised if intelligent machines are available from the beginning. For software development, that may mean smaller engineering teams coordinating specialised agents. Testing could become continuous, documentation could update automatically, and systems could detect operational problems and prepare potential repairs before an engineer intervenes.
Human roles would shift towards architecture, judgement, governance and complex problem-solving. Procurement could move from purchasing engineering hours towards purchasing software outcomes. Management could shift from counting developers towards measuring what combined human-and-AI teams actually deliver.
The organisations that gain the most from AI may therefore not be those purchasing the largest number of tools. They may be the companies that rebuild their operating models around them. That is the real productivity challenge now facing enterprise technology. AI has already demonstrated that it can make individual tasks faster. The harder task is proving that the company itself has become more productive.
Source: CIJ.World Research & Analysis Team