Companies trying to control artificial intelligence are discovering a difficult problem: writing rules for an AI system is considerably easier than ensuring those rules produce the right decision when the system encounters the complexity of the real world. As AI moves deeper into financial analysis, healthcare, customer service and corporate decision-making, businesses may therefore need to supplement policies and guardrails with something more demanding — continuous testing against the judgement of people who actually understand the relevant business domain. That was the central argument presented at Ai4 2026 in Las Vegas by Robbie Goldfarb, Co-Founder and CTO of Forum AI, who argued that many of the hardest AI problems cannot be reduced to simple instructions about what a system is or is not permitted to do.
The issue is becoming increasingly important as companies create formal AI policies. Major model developers already maintain their own behavioural frameworks, while enterprises are creating governance structures covering privacy, transparency, fairness, record keeping and employee access. Such frameworks are essential, but they are not sufficient by themselves because organisations frequently operate in situations where several legitimate objectives conflict. One policy might require an AI system to preserve an individual’s autonomy while another requires it to consider that person’s longer-term wellbeing. A company may require employees to collect only the minimum information necessary while simultaneously requiring sufficient records to reconstruct and audit an AI decision. Privacy may conflict with legal obligations, while transparency can sometimes conflict with security. These are not necessarily examples of badly written policies. They illustrate that the real world contains situations where several reasonable principles can apply simultaneously.
Goldfarb argued that humans often deal with such ambiguity by considering consequences rather than consulting rules mechanically. Experienced professionals think about what is likely to happen if a particular action is taken, who will be affected and whether an apparently satisfactory short-term result could create a larger problem later. Domain expertise matters because experienced people accumulate knowledge about how apparently similar situations can develop differently depending on context. Forum AI describes one approach to translating this thinking into AI evaluation as consequence mapping. Instead of simply classifying an AI response as acceptable or unacceptable, evaluators examine the likely effects of the decision and whether it actually produced the outcome the organisation wanted.
A customer-service chatbot provides a straightforward example. The system might respond politely to an unhappy customer and technically follow every instruction it has been given, yet leave the underlying problem unresolved. Viewed only as an individual interaction, the response might appear successful. Viewed through its consequences, it may result in another customer contact, additional work for human employees, higher support costs and a longer resolution time. The same principle can apply throughout business. A financial-analysis system could generate technically correct calculations while failing to identify a material risk. An HR assistant could respond accurately while revealing information that should remain restricted, while an automated procurement system might obtain a lower price but accept contractual conditions creating greater liabilities elsewhere. Evaluating only whether AI followed its instructions can therefore miss whether it produced the result the organisation actually needed.
Context creates another problem. Many AI evaluations still test systems through isolated questions and responses, while real business processes develop over time. The correct response to a customer today may depend on what happened yesterday, what another department promised last week or what the organisation already knows about the case. Goldfarb used the example of a passenger asking an airline chatbot about missing luggage and being instructed to complete a form. If the passenger returns later because nothing has changed, a system without appropriate memory might simply recommend completing the same form again. Each individual response could appear reasonable, but the overall customer experience would clearly be poor.
The opposite problem also exists. AI systems can rely too heavily on historical information and introduce irrelevant context into new situations. Effective memory therefore involves more than storing greater quantities of information. A system needs to identify which previous information matters to the current decision and which should be ignored. This becomes increasingly important as companies deploy agents across longer workflows. AI systems may communicate with customers, access corporate records and execute several stages of a process over hours or days. An agent unable to maintain appropriate context may repeatedly restart processes, contradict earlier actions or act on information that is no longer valid.
That leads to perhaps the most commercially significant argument in the Forum AI presentation: companies increasingly need their own evaluation systems defining what good AI performance actually means inside their organisation. Goldfarb argued that engineers should not be solely responsible for making that judgement. Technical teams can measure latency, model performance and system reliability, but determining whether financial analysis is genuinely useful requires financial expertise. Assessing clinical behaviour requires medical knowledge, evaluating legal workflows requires legal expertise, and determining whether an AI system performs well inside a particular company requires people who understand that company’s customers, processes and standards.
This changes AI evaluation from a predominantly technical exercise into a business capability. Forum AI itself is developing products around this concept, combining standardised evaluations created with external specialists with customised assessments reflecting organisations’ own knowledge and workflows. The broader implication extends beyond Forum AI’s platform. As companies become increasingly dependent on third-party foundation models, their competitive advantage may not come from owning the underlying model. Many organisations can access similar models from the same technology providers. What may differentiate them is their ability to define, measure and continually improve the outcomes they expect those models to produce.
Microsoft CEO Satya Nadella has made a related argument, suggesting that companies’ private evaluations could become particularly valuable intellectual property. The logic is important. If an organisation has a proprietary collection of tests accurately representing what excellent performance looks like, it can compare competing models against its own requirements rather than relying entirely on public benchmarks. That could also reduce dependence on individual AI suppliers. If another model becomes cheaper, faster or more capable, the company can test it against its existing evaluation framework before deciding whether to switch.
Private evaluations can also capture institutional knowledge that previously existed mainly inside experienced employees’ heads. An experienced property investment professional, for example, may recognise that an acquisition appearing attractive according to headline yield contains warning signs in lease expiries, tenant concentration, future capital expenditure, financing or competing supply. Turning those considerations into structured evaluation criteria could allow a company to test whether an AI system reproduces the judgement required by the business rather than merely generating plausible financial analysis.
The same principle could become increasingly relevant throughout commercial real estate. AI systems are beginning to analyse leases, valuation reports, financing documents, building information, tenant correspondence and potential acquisitions. Generic model accuracy matters, but organisations also need to know whether those systems follow their particular investment mandates, risk limits and operating procedures. A logistics developer evaluating development land has different priorities from a residential investor assessing an apartment portfolio. A bank underwriting a property loan requires different outputs from a property manager handling tenant requests. Each organisation consequently has its own definition of acceptable AI performance.
This is why standardised benchmarks, although useful, can only go so far. Public tests can identify broad differences between models, but they cannot fully reproduce a company’s proprietary workflows, customers, policies and risk appetite. Goldfarb described internal Forum AI testing in which automated AI judges were compared with domain experts and sometimes reached materially different conclusions. Those results should be treated as Forum AI’s own research rather than an established industry-wide finding, but the underlying warning is relevant. Using one AI system automatically to determine whether another AI system performed correctly can create false confidence if the evaluator itself does not reflect the judgement of people who understand the subject.
Human expertise therefore remains important even within highly automated evaluation systems. The objective is not necessarily to have specialists manually review every AI response, which would undermine much of the economic benefit of automation. Instead, expert judgement can be used to design and calibrate evaluation systems that subsequently operate at greater scale. Companies can effectively convert part of their institutional knowledge into a reusable mechanism for assessing machine decisions.
Organisations also need to recognise that AI performance is not static. A system that performs satisfactorily today may behave differently after the underlying model is updated, prompts are changed, new information sources are connected or additional tools are made available. Even apparently small modifications to the surrounding software can alter model behaviour. An AI application therefore cannot simply pass an acceptance test before launch and then be assumed to remain reliable indefinitely.
Continuous evaluation may consequently become analogous to the monitoring companies already apply to other critical infrastructure. Organisations track cybersecurity vulnerabilities, financial controls, networks and operational systems after deployment rather than testing them once and assuming nothing will change. AI may require a similar operating model in which performance is continually measured, failures are investigated and evaluation criteria evolve as the business changes. This becomes particularly important as AI gains the ability to act rather than simply answer questions. A system producing text can cause reputational or informational problems when it fails, while an agent capable of changing records, initiating transactions or communicating independently with customers can translate the same failure directly into an operational or financial consequence.
The wider lesson for businesses is that AI governance cannot consist entirely of restrictions written before deployment. Policies define boundaries, but evaluation determines whether the system is actually operating successfully within them. Companies therefore need both. They require clear rules covering privacy, security, access and accountability while simultaneously developing methods for testing whether AI produces the quality of decisions expected in ambiguous real-world situations.
As AI moves away from isolated assistants and becomes embedded inside operating processes, the ability to define what constitutes a good decision could itself become a competitive asset. The more decisions organisations delegate to machines, the more valuable their accumulated human judgement becomes. The advantage may ultimately belong not to the companies using the largest or newest models, but to those that can most accurately translate their expertise into measurable standards, continuously test AI against those standards and improve the systems when performance begins to drift.
In that sense, the next stage of enterprise AI governance may be less about writing another rulebook and more about building an institutional capability to determine whether the technology is actually making good decisions. Models, platforms and AI suppliers will continue to change. A company’s understanding of its customers, risks, workflows and definition of a successful outcome is far more proprietary. Capturing that knowledge and turning it into a continuous evaluation system may become one of the more valuable foundations for deploying AI reliably at scale.
Source: CIJ.World Research & Analysis Team