Artificial intelligence was supposed to save employees time, but in many organisations it is creating a new category of work: checking whether the machine was right. The result is an emerging form of AI fatigue in which employees are not necessarily overwhelmed by the technology itself, but by the continuous judgement required to supervise it. That was the central argument presented at Ai4 2026 in Las Vegas by Jessica Davis, Vice President of AI Product at USA TODAY Co., who works across an organisation encompassing USA TODAY and more than 200 local publications. In journalism, where accuracy and public trust are fundamental to the product, the consequences of unreliable AI are particularly visible, but the experience offers a broader lesson for companies in other industries. When employees are expected to verify every AI output manually, human oversight can quickly become the bottleneck that prevents the technology from delivering the productivity gains businesses expected.
Many organisations entered the generative AI era expecting employees to complete existing tasks faster. Instead, the nature of the work itself has changed. AI can draft, analyse, search and recommend at considerable speed, but somebody still needs to decide whether the result is accurate, useful, safe and appropriate. When those standards have not been clearly defined by management, individual employees effectively become their own AI governance function. Human-in-the-loop oversight remains essential in many circumstances, particularly where mistakes can affect customers, financial decisions, legal obligations, safety or reputation, but asking people to scrutinise every output indefinitely is difficult to scale. The more AI an organisation deploys, the more employees can find themselves supervising machines rather than performing the higher-value work automation was supposed to create time for.
Davis argued that this is fundamentally a leadership problem rather than simply a technology problem. Organisations need to ask whether AI is working for employees or employees are increasingly working for AI. The gap between changing workflows and unchanged expectations is where much of the fatigue emerges. If someone is still expected to deliver the same volume of traditional work while also checking every automated output, AI may add responsibility without removing enough labour elsewhere. The problem is compounded by the fact that AI is not conventional deterministic software. Similar prompts can produce different results, behaviour can change when models are updated, and performance can vary considerably between situations. Governance structures designed around relatively stable software therefore do not always translate comfortably to AI.
One potential answer is greater use of structured evaluations. Instead of asking employees repeatedly whether an individual AI output looks acceptable, companies can first define what successful performance actually means and then systematically test systems against those criteria. Human judgement moves away from sitting inside every individual transaction towards designing and supervising the standards by which the system operates. If a company can demonstrate through repeated testing that a system performs reliably within a clearly defined class of tasks, employees do not necessarily need to inspect every routine output with the same intensity. Human intervention can instead concentrate on exceptions, failures, unusual cases and genuinely high-risk decisions.
USA TODAY Co.’s experience with public-records requests provides an example of how this can work. Journalists regularly use federal and state public-records laws to obtain information from government organisations, but requirements differ between jurisdictions and mistakes can result in delayed or rejected requests. The company developed an AI-supported system intended to help journalists prepare these requests, but Davis said the project initially struggled with hallucinations and incorrect information. Journalists worked with the development team over a period of months, repeatedly identifying problems and refining the system, yet its performance remained below the level the organisation considered suitable for newsroom use.
Progress accelerated when the company began converting journalists’ judgement into structured evaluations and introduced additional supervisory mechanisms. Davis told the Ai4 audience that the resulting system eventually achieved a very high level of accuracy across the use cases tested internally. The precise figures should be treated as company-reported results rather than independently audited performance data, but the more important lesson is how the development process changed. Instead of relying primarily on journalists to discover errors one interaction at a time, the organisation captured what they considered a successful public-records request and turned those requirements into repeatable tests. Once that framework existed, additional functionality that had previously taken much longer to develop could reportedly be introduced considerably faster.
This approach effectively converts expertise into infrastructure. Experienced employees still determine what good performance looks like, but their judgement can be reused repeatedly rather than manually applied to every output. That has potentially significant implications for companies struggling with AI adoption because it provides a way to maintain human oversight without requiring perpetual human supervision of routine tasks. A bank using AI to review loan documents could establish which financial and compliance conditions must always be identified. A logistics operator could test whether recommendations comply with route, service and safety requirements. A property investor could build evaluations around lease expiries, tenant concentration, financing assumptions, capital expenditure and investment mandates. The objective is not to remove professional judgement but to capture enough of it to determine when automated work can be trusted and when a person needs to intervene.
The same principle can improve the relationship between product teams and governance functions. Product teams are generally expected to create business value, improve adoption and deliver products quickly, while governance teams concentrate on risk, compliance and monitoring. If those functions operate independently, AI projects can become trapped between pressure to accelerate and pressure to minimise exposure. Davis described USA TODAY Co. using a risk-value framework to create a common basis for those conversations. Instead of debating whether an application simply feels too risky or sufficiently useful, management can examine measurable evidence about performance, failure rates and expected value. A low-risk, high-value application may move through approval relatively quickly, while a system with significant consequences and uncertain performance receives greater scrutiny.
That can also reduce governance-related fatigue. Employees may currently be required to complete repeated approvals, document every experiment and manually verify outputs because management lacks confidence in the systems being deployed. Structured evaluation does not eliminate governance, but it can make it more targeted. The meaning of human oversight consequently begins to change. The phrase human in the loop is often interpreted to mean that somebody must review everything an AI system produces. A more sustainable model may increasingly place people above the loop: specialists establish standards, determine escalation points, investigate failures and decide when systems can be allowed greater autonomy.
There will remain situations where every output requires human approval. High-stakes financial, healthcare, legal or safety decisions may justify extensive supervision, particularly when the consequences of an error are substantial. Applying exactly the same standard to low-risk routine tasks, however, can destroy much of the productivity value AI was intended to create. Governance therefore needs to distinguish between applications rather than imposing one level of human involvement across an entire organisation.
For commercial real estate, this distinction is likely to become increasingly important as AI enters lease abstraction, investment screening, due diligence, valuation support, property management, procurement, financing and development analysis. If analysts must manually repeat every calculation or verify every extracted lease clause, automation may simply create another layer of work. A stronger model would establish conditions under which particular outputs can be trusted. An organisation might, for example, test large numbers of lease provisions against expert-reviewed answers and establish confidence thresholds for particular document types. Common clauses that repeatedly meet the required standard could be processed with less intervention, while unusual provisions or low-confidence results would be escalated to legal or asset-management teams.
The same approach could apply to investment analysis. AI does not need to replace investment committees or professional judgement to generate substantial value. It can automate repeatable analytical work if the organisation knows how to determine whether those tasks are being performed correctly. Human attention can then move towards assumptions, negotiations, strategy and exceptions where experience matters most. In that model, AI handles more of the repeatable work while professionals concentrate on decisions that require contextual understanding and accountability.
The remaining challenge is continuous monitoring. Davis acknowledged that USA TODAY Co. is still working on how to extend evaluation across more tools and workflows. That matters because an AI system that performs satisfactorily today may behave differently after its underlying model, prompts, information sources or surrounding processes are changed. AI governance therefore increasingly resembles continuous quality control rather than a one-time approval process. Companies need mechanisms capable of identifying when performance deteriorates, when new categories of failure appear and whether systems remain within the standards originally established for them.
The broader lesson is that AI fatigue cannot be solved simply by providing employees with more training or deploying more powerful models. Companies need to redesign the work surrounding AI. Leadership must determine what people should continue doing themselves, which tasks machines can perform reliably, how performance will be measured and where human judgement creates the greatest value. Without that clarity, employees risk becoming permanent supervisors of technology that was supposed to make them more productive.
The objective is therefore not to remove humans from AI processes but to use their expertise more intelligently. People define what good performance looks like, evaluation systems measure whether AI continues to meet that standard, and human attention is concentrated where uncertainty or consequences justify it. Organisations that learn how to capture professional judgement in repeatable evaluation systems may be better positioned to achieve both the productivity improvements and the trust that enterprise AI adoption was intended to deliver.
In that sense, AI fatigue may be an early warning that an organisation’s operating model has not evolved as quickly as its technology. The companies that solve the problem will not necessarily be those deploying the most AI. They may be those that become best at deciding when machines can be trusted, when people need to intervene and how the relationship between the two should change as the technology improves.
Source: CIJ.World Research & Analysis Team