Category: Validation

  • AI Moves at the Speed the Business Can Absorb

    The pressure to “do AI” has become one of the loudest forces in business. Boards ask for a strategy, competitors announce pilots, and employees experiment with new tools before governance has caught up. In that environment, speed is often mistaken for progress. The more useful leadership question instead of How quickly can we deploy AI? is How quickly can our organisation responsibly absorb the change it creates?

    That distinction matters because AI does not merely add a new application to the technology stack. Used seriously, it changes workflows, decision rights, data practices, roles, controls, customer interactions and, in some cases, the economic logic of a business. If those changes arrive faster than the organisation can understand, govern and operationalise them, the outcome is not transformation or improvement, but accumulated work, inconsistent practices and a loss of confidence.

    Jeff Wilke, a senior executive who worked at Amazon for 25 years, made this point bluntly to Jeff Bezos early in the company’s life. Bezos was generating ideas faster than the organisation could act on them, and Wilke told him: “You have to release the work at the right rate that the organization can accept it.” Bezos later described it as a profound insight. Every idea released beyond the organisation’s capacity to absorb it did not accelerate progress. It created distraction.

    That principle applies directly to AI adoption. The objective is not to slow down for its own sake but to sequence change so that each step increases the organisation’s capacity for the next.

    Moving too fast creates invisible failure

    When leaders mandate AI adoption at a pace the business cannot sustain, the first failure is often invisible. A team may launch an impressive pilot, produce an executive demonstration, or buy licences at scale. The operating model around the tool, however, remains unfinished. Staff do not know when to rely on the output, where sensitive information may go, who owns errors, how exceptions are handled or what is expected of them once the old process is retired.

    The result is a familiar pattern: workarounds multiply, quality varies by team, risk teams intervene late, and employees quietly revert to older methods when the new approach causes friction. AI earns a reputation as a management fad rather than a source of practical value. In the most severe cases, an organisation disrupts a reliable service model before it has built a dependable replacement. The approach fails because the change exceeds the organisation’s ability to cope, not because the model used was insufficiently powerful.

    A useful way to think about this is as a mismatch between deployment velocity and absorption capacity.

    DimensionWhen deployment outpaces capacityWhen capacity sets the pace
    Process designAI is layered onto unclear or unstable work.The team redesigns one defined workflow and clarifies hand-offs.
    PeopleEmployees are told to adopt, but not taught how to exercise judgement.Training, job aids and feedback are built into the rollout.
    Data and controlsData use and accountability are resolved after launch.Guardrails, escalation paths and quality checks are established before broader use.
    Value measurementActivity is reported: licences, pilots and prompts.Outcomes are measured: cycle time, quality, cost, customer experience and risk.
    TrustEarly mistakes become evidence against the programme.Small, well-governed wins create permission for the next change.

    The difference is not caution versus ambition. It is operational discipline versus performative speed.

    Established businesses are complex

    Startups can often change rapidly because the organisation is smaller, its operating model is still forming and its technology carries little legacy integration. A founder can decide in the morning, change a product workflow by afternoon and observe the impact within days.

    Established businesses face a different reality. They may serve millions of customers, operate under regulatory obligations, rely on complex supplier relationships and carry years of interconnected systems and policies. Their scale is an advantage, but it also means that a change in one function creates consequences in many others. A new AI-assisted credit decision, for example, has implications for compliance, model governance, customer communication, front-line procedures, auditability and dispute resolution. The business must move deliberately because the cost of a poorly absorbed change is multiplied by its reach.

    This a design constraint to work within, not a weakness to apologise for. Large organisations should build a repeatable mechanism for safe acceleration. The goal is to increase the rate at which the business can absorb change, not to force a rate that breaks it.

    Build absorption capacity before demanding velocity

    The practical response is to treat AI adoption as a managed portfolio of changes rather than a single enterprise-wide instruction. Start with a workflow that is sufficiently valuable to matter, sufficiently bounded to govern and sufficiently measurable to learn from. Assign a business owner, a process owner and a risk or control partner. Define what a good outcome looks like before deployment, including the conditions under which the team will pause or reverse the change.

    A sensible sequence has four stages.

    StageLeadership questionEvidence required before moving on
    ProveDoes this solve a real operational or customer problem?A defined baseline and a demonstrable improvement in a controlled workflow.
    StabiliseCan people use it reliably under normal conditions?Clear ownership, training, controls and a functioning exception process.
    ReplicateDoes the operating pattern transfer to similar work?Consistent results across more than one team or business unit.
    ScaleCan the enterprise support this without degrading quality or trust?Resourcing, governance and technology capacity match the expanded scope.

    First, prove usefulness in a narrow setting where people can compare the AI-assisted outcome with the existing process. Then stabilise the new way of working by documenting roles, controls, exceptions and training. Then replicate the pattern in adjacent use cases that share similar data, risks or operating habits. Finally, scale the platform and governance only after the organisation has evidence that the operating model works repeatedly.

    This approach may appear slower at the beginning because it resists the temptation to announce universal adoption. In practice, it is usually faster over time. Each successful use case leaves behind reusable assets such as trusted data pathways, trained leaders, approved controls, implementation playbooks, measurement methods and internal advocates. The next deployment begins from a stronger base, drawing on accumulated organisational knowledge rather than repeating the same arguments from scratch.

    Leaders must protect the rate of learning, not the rate of adoption

    The leadership task is to sponsor AI usage and to protect the business’s learning rate. That means creating room for teams to identify flaws without being labelled resistant, refusing to measure progress solely by adoption volume and being explicit about what will not yet be automated. It also means resisting the false choice between reckless acceleration and organisational paralysis.

    A mature AI programme should be demanding, but its demands should be specific: improve a process, raise decision quality, reduce a known friction point, protect customers, and demonstrate results. “Use AI everywhere” is an instruction to create unmanaged variation at scale.

    Wilke’s observation to Bezos offers a better standard. Ideas have value only when an organisation can turn them into coherent action. Every AI deployment released beyond the organisation’s capacity to absorb it creates a backlog of unfinished transformation.

    Organisations that make change smooth – clear enough for people to adopt, controlled enough to trust and valuable enough to sustain – will accumulate capability faster than those that simply move first. Slow is smooth, smooth is fast.

  • Why AI Automation Fails Without Knowledge Engineering

    Most AI automation projects for business processes start the same way. A vendor demonstrates a compelling proof of concept and the demo works. The board approves the investment but six months later, the automated process is producing results that no one trusts, and the organisation is quietly routing work around it.

    The failure wasn’t in the AI. It was in what the implementation team did before they deployed the AI.

    Quick automation captures what systems do while knowledge engineering captures why they do it. These are not the same thing, and conflating them is the most common and costly mistake in enterprise AI implementation today.

    The Gap That Quick Automation Cannot See

    Process mining and rule extraction are useful tools. They read event logs from your ERP, CRM, and core systems, then reconstruct how your processes actually flow. They surface bottlenecks, flag deviations, and give your implementation team a working map of operations.

    For simple, well-documented processes, this is enough but for anything more complex, it is a structural failure. Event logs record what happened in the system, not the reasoning behind it. They cannot see the workarounds your experienced people have built over fifteen years. They cannot capture the judgement your senior underwriter applies when a claim is technically within policy but smells like fraud. They cannot document the exception a finance team made for a long-standing client that never made it into the system because everyone just knew.

    Research published in 2025 confirmed that event logs systematically underrepresent operational reality, with an estimated 70% of knowledge work occurring outside the systems that generate those logs. What gets automated by the AI is the documented surface of a process, not its full depth.

    When the AI encounters the undocumented exception, it either fails, escalates, or produces a wrong answer with complete confidence. All three outcomes erode trust, and once trust is gone, the automation sits idle while your people route around it.

    What Knowledge Engineering Actually Does

    Knowledge engineering is the discipline of making implicit knowledge explicit and verifiable before your team encodes it into an AI system.

    This work involves structured interviews with domain experts. It involves observing how work actually gets done, not how the procedure manual says it should be done. It involves encoding the results into knowledge bases that subject matter experts can review, challenge, and correct. It involves validating that the encoded logic produces the right outcomes before anything touches a production environment.

    This is the pattern that distinguishes durable AI automation from the kind that works in demos and fails in production. The AI surfaces candidate rules and logic and domain experts validate or correct them. The validated knowledge becomes the foundation your team builds automation on.

    This model is explicit about something that many vendors prefer not to emphasise: AI will confidently surface intent that sounds right and is not. Business-logic reconstruction can leverage AI but remains human-led by design. The AI drafts the interpretation while the expert confirms or corrects it.

    Where Systems Thinking Changes the Problem

    A systems thinking approach to AI automation asks a different question than most implementation teams ask. Instead of “what processes can we automate?”, it asks “what does the system as a whole need to keep working correctly after automation is introduced?”

    This reframing matters because automation does not operate in isolation. When you automate a claims processing workflow, you change the load on the humans who handle escalations. You change what data your compliance team can access and when. You change the feedback mechanisms that let your experienced people spot when something is going wrong. You may also change the incentives in ways you did not intend, particularly if the organisation measures the automated process on speed but measured the people it replaced on accuracy.

    Systems thinking requires mapping those interdependencies before implementation, not discovering them six months after go-live.

    Three principles from this approach are directly applicable to AI automation decisions.

    The first is that problems appear long before they are noticed. An AI model that is slowly drifting from its training distribution will produce subtly degrading outputs for weeks before anyone raises a flag. IBM’s 2025 Cost of a Data Breach Report found that 13% of organisations reported breaches involving AI models or applications, and 97% of those breached organisations lacked proper AI access controls. The failures were systemic, not sudden.

    The second is that you get what you incentivise. If your AI implementation is measured on the percentage of cases processed automatically, you will get high automation rates but you may also get poor outcomes on the cases that needed human review but were incorrectly classified as routine. Design the measurement framework before designing the automation.

    The third is that you fall to your level of preparation. When something goes wrong with an automated process at scale, your team’s ability to respond depends entirely on what they practised before it happened. Oversight protocols, escalation paths, and retraining procedures need to exist and be tested well before an incident requires them.

    The Real-Time Oversight Problem

    Traditional governance was built for a different tempo: annual audits, quarterly reviews, weekly metrics. These cadences made sense when processes ran at human speed and human scale.

    AI automation changes both variables simultaneously. A model processing thousands of claims per day can accumulate hundreds of errors before a weekly metrics review would detect the pattern. By the time a quarterly audit surfaces a systematic bias in lending decisions, the exposure may already be material.

    This is not a hypothetical concern. The EU AI Act’s Article 96 now requires organisations to demonstrate compliance through continuously updated, machine-readable evidence, not point-in-time assessments.

    The implication for senior leaders is direct. Governance frameworks designed for human-speed processes need to be redesigned before your organisation deploys AI at scale, not retrofitted after a problem surfaces. Real-time monitoring of model outputs, data quality, and prediction confidence is not optional infrastructure. It is the mechanism by which you maintain accountability for decisions your organisation is making at machine speed.

    What a Systemic Approach Actually Looks Like

    Implementing AI automation well follows a recognisable pattern.

    Begin with knowledge extraction, not process mapping. Before any automation is designed, invest in capturing the institutional knowledge that lives in people, not systems. This means structured elicitation sessions with domain experts, protocol analysis of how difficult cases are actually handled, and documentation of the exceptions and judgement calls that define quality outcomes.

    Validate before you deploy. The people who encoded the knowledge review it. Pilot programmes run in environments where your team checks outputs against known-good outcomes before the model is trusted with live decisions.

    Design oversight into the process architecture. Human-in-the-loop checkpoints are not emergency interventions but are designed features, placed where the model’s confidence is lowest or where the cost of an error is highest.

    Instrument for continuous monitoring. Model performance, data quality, and output confidence scores are tracked continuously. Thresholds trigger review before drift becomes visible to customers or regulators.

    Govern at the speed of your systems. Review cadences match the operational tempo of the automated process, not the legacy tempo of the governance process it replaced.

    The Question to Ask Your Implementation Team

    If your organisation is evaluating or currently implementing AI business process automation, one question will reveal more about the quality of the approach than any other: what knowledge engineering work did the team complete before the AI model was trained?

    If the answer is “we used automated process mining and rule extraction from our existing systems”, that is a starting point, not a complete answer. It describes what the system does, it does not describe what your experienced people know that the system does not record.

    If the answer is “we conducted structured knowledge capture with domain experts and validated the results before training”, you are working with a team that understands where AI automation actually fails.

    The technology for automating business processes is mature, accessible, and increasingly affordable. The discipline required to automate them well, preserving institutional knowledge, designing real-time oversight, and aligning governance to the speed of the systems being governed, remains genuinely rare.

    That gap is where durable competitive advantage is built.

  • Your AI Interface Is a Cognitive Design, Not a Cosmetic One

    Most organisations deploying AI spend months selecting the right model, agonising over accuracy rates, vendor contracts, and compliance implications. Then, in the final weeks before launch, someone asks: “What should the screen look like?”

    The interface is not decoration applied after the real decisions are made, instead it determines whether your people reason clearly alongside the AI or quietly work around it. Get it right, and your AI system makes better decisions than either the human or the system could alone. Get it wrong, and you have an expensive system that your team has learned to distrust, override, or ignore.

    This article gives you a structured method to audit any AI interface before deployment. Based on a field called Cognitive Systems Engineering (CSE), developed in the early 1980s by Erik Hollnagel and David Woods to address exactly the kind of high-stakes, human-machine decision environments that organisations across every industry are now building. The method is built around a five-level analysis called the Abstraction Hierarchy. You will walk through each level (we’ll use a loan decision AI as the working example), leaving with a set of questions you can apply to your own deployment.

    The Shift That Changes Everything

    Traditional software is deterministic. You press a button, you get a result. Designing the interface for that kind of system still requires some skill, but is largely a matter of clarity and efficiency because we have many examples of good (and bad) design.

    AI is different as it produces uncertain outputs and reasons probabilistically. It can be right most of the time and catastrophically wrong in ways that are difficult to anticipate. The interface for a traditional system needs to be usable. The interface for an AI system needs to do something different: it needs to make the AI’s reasoning visible so that the human can judge when to act on it, when to question it, and when to override it.

    Hollnagel and Woods called this a “joint cognitive system“: the human and the AI are not separate entities where one hands off to the other. They are a single thinking unit, and the interface is the connective tissue between them. Whether the decision involves a loan approval, a clinical recommendation, a fraud alert, or an operational plan, that connective tissue either holds or it tears.

    The Abstraction Hierarchy, developed by Jens Rasmussen and Kim Vicente at the Risø National Laboratory in Denmark, gives you a disciplined way to design that connective tissue. It asks you to understand your work domain at five levels, from purpose down to physical configuration. Each level reveals a different set of interface requirements that you might otherwise miss.

    The Abstraction Hierarchy: Five Questions for Any AI Interface

    Think of the Abstraction Hierarchy less as a taxonomy and more as a diagnostic interview. At each level, you ask a question about your work domain. The answers tell you what your interface must show, what it must prevent, and what it must make possible.

    Here is the framework applied to a loan decision AI to help understand the details.

    Level 1: What is this system ultimately for?

    This is the question of functional purpose: the values and goals the system exists to serve. Not the technical goals (“classify loan applications”), but the human and organisational values at stake.

    For a loan decision AI, the answer is not simply “approve or decline applications faster”. The real purposes are: extend credit to people who can repay it, protect the institution from credit losses, comply with responsible lending obligations, and treat applicants fairly across demographic groups. These purposes can conflict. A model optimised for speed may sacrifice fairness. A model optimised for loss minimisation may discriminate. These tensions exist whether you name them or not and exposing them allows informed decisions.

    What this means for your interface: The interface must make the system’s governing priorities visible. If the model has been calibrated to weight certain risk factors above others, the loan officer reviewing its recommendation should be able to see that calibration, not just the output. When the AI recommends declining an application, the interface should surface which of the system’s core purposes drove that recommendation: is this a credit risk concern, a compliance flag, or something the model cannot categorise cleanly?

    Questions to ask about your deployment:

    • What are the two or three values this system is genuinely optimised for?
    • What values are in tension, and does the interface make those tensions visible?
    • When the AI’s recommendation conflicts with a user’s instinct, does the interface help the user understand why?

    Level 2: What principles govern how the system operates?

    This is the question of abstract function: the rules, regulations, and governing principles that constrain what the system can and cannot do. In aviation, these are physics and safety regulations. In lending, they are responsible lending laws, anti-discrimination regulations, internal credit policy, and audit requirements.

    These constraints do not change with each transaction. They define the boundaries within which all decisions must fall. The problem is that most AI interfaces present outputs as though these constraints do not exist. The model returns a score. The interface shows the output metric. The loan officer is left to remember, from training, which regulatory constraints apply in this situation.

    What this means for your interface: Regulatory and policy constraints should be structurally present in the interface, not stored in the user’s head. If the AI’s recommendation would require a manual review under responsible lending obligations, the interface should flag that requirement automatically. If the model’s confidence falls below a threshold that your compliance team’s internal policy requires to be reviewed, show that threshold, and ideally the policy, visibly.

    This is the principle that Vicente and Rasmussen called Ecological Interface Design (EID): the constraints of the work domain should be visible in the interface itself, not stored in the user’s memory. An EID-informed loan interface does not require the loan officer to remember the regulatory rulebook. It makes the rulebook structurally visible in the decision flow.

    Questions to ask about your deployment:

    • Which regulatory obligations apply to the decisions this AI supports?
    • Are those obligations visible in the interface, or do users have to remember them independently?
    • When the AI’s recommendation sits in a regulatory grey zone, does the interface make that visible?

    Level 3: What processes does the system need to perform?

    This is the question of generalised function: the operational processes and workflows that need to happen for the system to achieve its purpose within its constraints.

    In a loan context, this includes the steps of gathering applicant data, running the credit model, checking for compliance flags, presenting a recommendation, capturing the loan officer’s decision and rationale, escalating edge cases, and creating an audit trail. These processes are not all equally visible in most AI deployments. Typically, what is visible is the output of the model. The processes that produced it, and the processes that need to follow from it, are hidden or scattered across different systems.

    What this means for your interface: The interface should reflect the full process, not just the model’s output. If the correct process requires a loan officer to review the AI’s recommendation alongside the applicant’s supporting documents before deciding, the interface should make that sequence natural and difficult to skip. If the process requires capturing the officer’s reasoning when overriding the AI, the interface should prompt for that reasoning at the point of override, not as a retrospective form filed later.

    This is also where structured input design matters. If you want the AI to produce consistent, auditable outputs, you need the interface to guide consistent, structured inputs. A free-text prompt box for a loan officer to query the AI is the wrong design. A structured form that constructs the query from validated fields, treating the prompt as a template and the user’s input as variables, produces far better consistent results and a far cleaner audit trail.

    Questions to ask about your deployment:

    • What is the full sequence of steps the process requires, including before and after the AI’s output?
    • Does the interface make the correct sequence the natural path, or can users shortcut it?
    • How does the interface capture the human’s reasoning, not just the AI’s recommendation?

    Level 4: What are the capabilities and limits of each component?

    This is the question of physical function: what each component of the system can and cannot do. For an AI, this means understanding the model’s actual capability boundaries. Where does it perform well? Where does it degrade? What kinds of inputs push it outside its training distribution?

    Loan officers who use AI tools daily develop intuitions about where the model is reliable and where it is not. New officers do not have those intuitions, and even experienced officers can be misled when the model presents its outputs with uniform visual confidence regardless of whether it is in familiar or unfamiliar territory.

    What this means for your interface: The interface must distinguish between high-confidence and low-confidence outputs, and it must do so in a way that reflects the model’s actual calibration, not a standardised disclaimer. If the AI is recommending approval on an application that combines features it has rarely seen together, that uncertainty should be visible. If the model’s confidence score is below a meaningful threshold, the interface should communicate that clearly and differently from high-confidence outputs, not with a footnote, but with a structural difference in how the recommendation is presented.

    This is what the CSE literature calls making the AI’s epistemics visible: the interface should show not just what the AI concluded, but how firmly it concluded it and on what basis.

    Questions to ask about your deployment:

    • Does the interface distinguish between high-confidence and low-confidence recommendations?
    • Can users tell when the AI is operating in territory close to the edge of its training?
    • What happens when the AI encounters an input type it was not trained on?

    Level 5: What is the actual configuration of the system?

    This is the question of physical form: the literal layout, controls, and information architecture of the interface as it exists on the screen.

    This is where most interface design effort is spent and also the level where most AI interface problems are most visible: the designer burying a recommendation at the bottom of a long screen, the confidence score presented in a font smaller than the surrounding data, or the override button placed three clicks away. These are not aesthetic problems but decision quality problems.

    If the most important signal the AI is sending is that it is uncertain about this application, and the interface makes that signal hard to find, the loan officer will miss it. That is design failure, not a training failure.

    What this means for your interface: The visual hierarchy of the interface should reflect the information hierarchy of the decision. The AI’s recommendation and its confidence level should be visually prominent. Flags and caveats should not be hidden in tooltips. The action the interface makes easiest should be the action the process intends to be most common. The action that requires more care, like an override of the AI’s recommendation, should require commensurate effort in the interface: not so much effort that it becomes a workaround, but enough that it cannot happen accidentally.

    Questions to ask about your deployment:

    • Does the visual hierarchy of the interface match the decision hierarchy?
    • Is the AI’s uncertainty as visible as the AI’s recommendation?
    • What is the path of least resistance in the interface, and is that the right path?

    What Happens When You Skip This Analysis

    The most common failure mode is not that the AI model is wrong but that the interface makes it impossible to know when the model is wrong.

    People in high-stakes roles learn quickly. If an AI interface presents confident-sounding recommendations without surfacing the model’s reasoning or uncertainty, they will test it against their own judgement for a few weeks, find cases where it was clearly wrong, and start treating all its outputs with blanket scepticism. The AI becomes a checkbox, not a collaborator. The organisation has paid for a decision-support system and deployed a bureaucratic step.

    The second failure mode is the reverse: people defer to the AI when they should not, because the interface presents its outputs with more authority than the model’s actual confidence warrants. This produces decisions that look considered but are indefensible when scrutinised: by regulators, auditors, customers, or a board asking why something went wrong.

    Both failures are interface failures. The model may be performing exactly as designed. The interface is simply not communicating what the model knows and does not know.

    Where to Start

    You do not need to redesign your entire interface before launch. You need to run through these five levels with the people who will use the system and the people responsible for the process it supports.

    Bring three groups into a room: the people who will use the AI day to day, the people who own the process it sits inside, and whoever is accountable for risk or compliance in that domain. Walk through the five levels as questions. At each level, ask: what does the interface currently show, and what does this analysis say it needs to show? The gaps between those two answers are your design priorities.

    Pay particular attention to Levels 1 and 4. Purpose misalignment (Level 1) and invisible uncertainty (Level 4) are the two most dangerous gaps, and they are the two most commonly overlooked in AI deployments focused on model performance rather than interface design.

    The Abstraction Hierarchy was originally developed to allow human operators to make complex, safety-critical decisions under pressure. Whether your AI is supporting credit decisions, clinical triage, fraud detection, or operational planning, the underlying challenge is the same: consequential, often regulated decisions where the interface either helps people reason well or quietly gets in the way.

    Your AI model does not make decisions. The human-AI system makes decisions. The interface is what makes that system work so design it accordingly.


    References

    [1] E. Hollnagel and D. D. Woods, “Cognitive systems engineering: New wine in new bottles,” International Journal of Man-Machine Studies, vol. 18, pp. 583–600, 1983.

    [2] K. J. Vicente and J. Rasmussen, “Ecological interface design: Theoretical foundations,” IEEE Transactions on Systems, Man, and Cybernetics, vol. 22, no. 4, pp. 589–606, 1992.

  • Modelling Genius: Encoding Human Expertise into AI

    Ask a master negotiator how they read a room and you will usually get a shrug. Ask a surgeon how they decide where to cut first, or a trader how they sense a market turn before the numbers confirm it, and the answer is the same kind of shrug. Expertise this deep runs on autopilot. The expert cannot see their own process because it stopped being conscious years ago.

    This is the real obstacle standing between human expertise and AI systems making the right decisions. It is not a data problem but a structure problem. Before an AI can inherit a decision-making pattern, someone has to find that pattern first.

    This article looks at what behavioural modelling is, the specific skills it demands, and how those skills can be used to produce frameworks such as the Galdren Lead Gen Decision and Judgement Framework for building AI systems that reason the way an expert reasons, not only the way a manual describes.

    What Expert Behavioural Modelling Means

    Practitioners use the word “modelling” to describe at least three different things, and the difference matters for what an AI system ends up inheriting.

    One approach requires the modeller to absorb an expert’s patterning unconsciously first, and only analyse it afterwards. The modeller suspends any conscious, analytic attempt to understand the expert’s patterning during that initial assimilation stage, and only claims success once they can reproduce the pattern in new situations and draw out responses of comparable quality and speed to the original expert.

    A second approach skips the immersive step and instead applies existing categories and labels to describe what an expert does. This can mean structured strategy elicitation that genuinely surfaces the expert’s underlying structure, or it can mean something far shallower: someone asks the expert a few questions, writes up a summary, and calls it a knowledge base. People use “modelling” for both. The results are not comparable. One reaches the structure underneath the expert’s behaviour. The other copies what is visible on the surface and stops there.

    The third approach, now the most common one in AI projects, does not involve the expert at all. It works from what the expert has already produced: emails, documents, and speeches, parsed into a retrieval system the AI searches when it needs an answer. The AI is left to infer the underlying decisions, choices, and behaviour from the artefacts alone.

    Most AI knowledge capture today uses this third approach, because it is faster and easier for everyone involved. It is also the approach least likely to produce a usable model. Feeding an AI what an expert wrote is not the same as feeding it how the expert decided what to write. The result is a system that can echo the expert’s language without inheriting the expert’s judgement.

    True behavioural modelling sets a higher bar than any of these three. It insists someone tests the extracted pattern against the real world before anyone calls the job complete.

    The TOTE: A Map of How a Decision Runs

    The core unit of any modelling is the TOTE (Test, Operate, Test, Exit), a concept the psychologists George Miller, Eugene Galanter, and Karl Pribram introduced in their 1960 book on cognitive psychology, drawing on ideas from cybernetics. A TOTE breaks a piece of behaviour into a trigger (test), an action taken in response (operate), a check on whether the action worked (test again), and the point at which the process stops (exit). If the outcome is not reached, the person loops back through testing and operating until it has. The model can describe something as small as hammering a nail, or be nested to build a map of far more complex behaviour.

    This is the piece most AI encoding projects skip. They capture what the expert decided. They do not capture the loop: what triggered attention in the first place, what the expert checked before acting, what told them the first attempt had not worked, and what specific signal told them to stop. Encode the decision without the loop and the AI inherits an answer, not a method. It tends to fail the moment conditions shift even slightly from the example it saw.

    Strategy Elicitation: The Skill of Finding a Pattern the Expert Cannot See

    Extracting a TOTE from an expert requires strategy elicitation, and it comes in two forms. Informal elicitation happens in conversation. People tend to run through their internal strategy as they describe a past experience, and simply asking “how did you know?” is often enough to draw the sequence out naturally. Formal elicitation is more deliberate. It works through a set sequence, because experts routinely leave out the step that matters most, not because they are hiding it, but because it happens too fast for them to notice.

    From Elicited Pattern to Working System

    Once the modeller has elicited a strategy, it needs a structure an AI can act on. Prompt engineering and knowledge graphs earn their place here. They translate the TOTE sequence and its trigger conditions into something the AI checks at each step, rather than a paragraph it reads once.

    The Galdren Lead Gen Decision and Judgement Framework is one working example of this translation in practice. It runs three parallel layers: an Execution Layer where the AI acts as generator and architect, a Judgement Layer where it acts as coach, challenger, and critic, and an AI Behaviour Layer where it acts as recorder and analyst, shaping thinking through the interaction itself. Underneath sits a seven-step decision sequence: frame the problem, gather the evidence, list the options, weigh the trade-offs, make the decision, state a confidence level from 0 to 100 percent, and apply a reversal test asking what would change the decision-maker’s mind.

    This sequence is a TOTE with the trigger and exit conditions made explicit. Frame and Evidence form the test that opens the loop. Options, Trade-offs, and Decision form the operate stage. Confidence and the Reversal Test form the second test, the check on whether the decision holds up, before the process exits. Structuring it this way is what lets the framework’s judgement training loop, Ask, Question, Push Back, Educate, Support, Iterate, work as a genuine feedback cycle rather than a checklist someone bolts onto the AI’s output after the fact.

    Validation: Proving the Model Works

    Any model is unproven until it produces the expert’s results, as confirmed by the expert. The modeller only claims success once they can reproduce the original patterning in new situations and get responses of comparable quality and timing to what the original expert would have produced. That is a demanding standard, and it should be. Extracted expertise that only works inside the interview room is not extracted expertise. It is a transcript.

    For AI systems, this means human-in-the-loop review cannot stop at “does the output look reasonable”. It has to ask whether the AI, given a new situation the expert never discussed, reaches the judgement the expert would have reached. This is where the Lead Gen Decision and Judgement framework’s emphasis on validation and correction earns its keep. Experts review outputs against unfamiliar cases, flag the gap between what the AI did and what they would have done, and the difference points straight back to which part of the TOTE the modeller captured wrongly or left out.

    The Difference Between a Coder and a Modeller

    Anyone can prompt an AI with instructions that someone already wrote down somewhere. That is transcription, and it is useful, but it is not modelling.

    A modeller does something else entirely. They sit with an expert who cannot explain their own genius, use structured elicitation to surface the trigger, the check, and the exit point the expert never noticed themselves, sort those pieces by the level they operate at, and only then hand the result to an AI system for testing against reality. A coder formats information that already exists in explicit form. A modeller extracts a pattern that has never been explicit, and proves it works before calling the job finished.

    Three Problems Behavioural Modelling Has Not Solved

    Several problems remain unresolved.

    Depth of elicitation. Formal strategy elicitation takes time and skill most AI projects do not budget for. Rushing it tends to produce a shallow model dressed up as a deep one.

    Drift. An expert’s TOTE for a given decision changes as their environment changes. A model captured once and never revisited tends to fall behind the person it was built from.

    Proof of transfer. The real test of encoded expertise is whether it performs on cases the expert never saw, not whether it matches the training examples. Too few AI projects test for this directly.

    The Discipline That Data Cannot Replace

    AI systems thinking delegation will not advance much further on data volume alone. What moves it forward is the discipline of behavioural modelling: find the trigger the expert cannot name, map the loop they run without noticing, sort what you find by the level it belongs to, and refuse to call the job done until the pattern produces the expert’s results in someone else’s hands. Frameworks such as Galdren LGD show what this looks like once someone builds it into a working system. Applied properly, behavioural modelling gives AI encoding projects something most currently lack: a defined way to know whether the expertise was ever captured at all.

  • The AI Parrot Problem: Why Your AI Sounds Smart and Isn’t

    A RAG system trained on an expert’s own documents can answer a question in the expert’s own words, sound completely authoritative, and still reach a conclusion the expert would never make. That gap, between sounding right and being right, is the real risk in most AI knowledge systems today. It should worry you more than a system that gets things obviously wrong. An obvious mistake at least tells you something is broken.

    The AI Parrot Problem: Why Your AI Sounds Smart and Isn't

    What RAG Actually Does

    Retrieval-Augmented Generation does two things, and neither one is judgement. First, it retrieves: given a question, it searches a knowledge base and pulls out passages that seem relevant, in much the same way an advanced search engine would. Second, it generates: it turns those passages into a fluent, coherent answer.

    That is the whole job. The system finds information and writes it up well. It does not weigh conflicting evidence, apply experience to an unfamiliar case, or reason about what an expert would decide. Retrieval and judgement are not two points on the same scale. They are different operations entirely, and a system built for one does not automatically gain the other.

    Why It Demos Well and Fails Quietly

    RAG systems look impressive in demonstrations, and for good reason. On questions close to their source material, they act as an efficient lookup tool: find the right passage, phrase it well, done. This is where most people form their impression of what the system can do.

    The trouble starts on the cases the documents never covered, the ones where an expert would normally lean on experience and judgement rather than a reference page. This is not a random weak spot but a structural failing. The system has no mechanism for reasoning through a situation it has not seen, only for retrieving and rephrasing what it has. The output can still sound confident and plausible even when it is wrong, which means the failure stays invisible until someone acts on it.

    Air Canada’s website chatbot is a well-documented instance of this pattern. In November 2022, a customer asked the chatbot about bereavement fares after the death of his grandmother. The chatbot told him he could apply for the discount after booking, within 90 days of ticket issue. He booked his flight on that basis and later submitted a claim. Air Canada refused it. The airline’s actual bereavement policy, published elsewhere on its own website, did not allow retroactive claims. The customer took the case to the BC Civil Resolution Tribunal, which found Air Canada liable for the chatbot’s inaccurate advice and rejected the airline’s argument that the chatbot was a separate entity not covered by its duty to customers. The tribunal ordered Air Canada to cover the fare difference and costs, a little over 800 Canadian dollars in total.

    Air Canada never disclosed exactly how the chatbot worked, so this was not confirmed as a RAG system specifically. What matters here is the pattern. An AI tool produced a fluent, specific, confident answer, using the airline’s own bereavement-fare language, that directly contradicted the airline’s own documented policy. That is the Parrot Problem, playing out with real financial and legal consequences.

    The Cost, at Two Levels

    The Parrot Problem is not only an inconvenience. It costs something at the individual level, and something different at the organisational level.

    For the individual, the cost is trust placed in the wrong direction. Someone asks a system a genuine judgement question, a recommendation, a diagnosis, advice on a decision that matters, and the system answers with the fluency and vocabulary of an expert. They act on it. Later they find out the “expert” behind the answer never reasoned through their specific case. Depending on the domain, that can mean a wasted booking, a financial loss, or worse.

    For the organisation, the cost compounds quietly. If nobody checks AI output against what an expert would decide, and only checks whether it sounds reasonable, the gap between the system’s answers and correct practice widens with every case it handles. A support AI that confidently recommends the wrong fix for an undocumented configuration is likely to fail the same way each time that configuration comes up, and nobody notices until the pattern of complaints does.

    Three Questions to Diagnose Your Own System

    You do not need a research team to find out whether your AI knowledge base has this problem. Ask these three questions.

    1. Has anyone tested it on a case the source documents never covered? If every test question can be answered directly from the training material, you have tested retrieval, not judgement.
    2. When it gets something wrong, can anyone say which piece of reasoning it skipped, or does it only look “off”? If nobody can point to the specific step the system missed, that is a sign it was never reasoning through the problem. It was matching patterns and hoping.
    3. Would the original expert sign off on this answer, or only recognise the words? An expert can recognise their own language in an AI’s response and still disagree completely with the conclusion. Recognising the vocabulary is not the same as agreeing with the judgement.

    It’s a Structure Problem, Not a Data Problem

    More documents will not fix this. A bigger knowledge base makes a RAG system a better parrot, not a better judge. The real work is capturing how an expert decides, not only what they have already written down, and that is a different kind of project entirely.

    That raises the obvious next question. If retrieval isn’t enough, what does it take to build a system that reasons the way an expert reasons? That is where we pick up next.

  • API vs. Local LLMs: The Cost, Privacy, and Security trade-off

    The AI landscape has handed enterprises a genuinely difficult architectural question: do you call a vendor API (OpenAI, Anthropic, Google), or do you self-host an open-weight model (Llama 3, Mistral, Qwen)?

    This is not a question of convenience. It touches data privacy, compliance obligations, long-term economics, and how much operational risk you’re willing to carry. The right answer varies by organisation. This guide maps the trade-offs so you can make the call with clarity.

    Data Privacy, Security, and Compliance: Where the Decision Often Starts

    For many enterprises, data sovereignty ends the conversation early.

    When you call a vendor API, your prompts and data travel to a third-party server. Even with enterprise agreements that prohibit training on your inputs, the data physically leaves your network. For healthcare, financial services, and government, that exposure creates real legal and compliance risk.

    A self-hosted model eliminates that exposure entirely. Prompts stay inside your infrastructure making compliance audits simpler and intellectual property does not pass through external systems.

    If your data classification policies, regulatory obligations, or legal counsel would flag external data transmission as a problem, self-hosting is the baseline requirement.

    Cost: The Maths Depends on How Much You Use It

    The cost comparison between vendor APIs and self-hosted models is heavily dependent on utilisation rates.

    Vendor APIs charge per token. That structure works well for prototyping, low-volume applications, and irregular usage. You pay for what you consume and carry no idle infrastructure costs.

    Self-hosting flips the model. You pay for compute uptime, not token generation. Costs are fixed and predictable, regardless of whether the model is processing requests or sitting idle.

    Cost FactorVendor API (e.g., OpenAI)Self-Hosted (e.g., Llama 3)
    Pricing ModelUsage-based (per token)Time-based (per hour of compute or power use)
    Upfront InvestmentZero to minimalHigh (cloud instance setup or hardware purchase)
    PredictabilityVariable, scales with usageFixed, predictable monthly infrastructure costs
    High Utilisation CostExpensive at scaleHighly cost-effective
    Low Utilisation CostVery cheapExpensive (paying for idle compute)

    The crossover point is utilisation. High, continuous workloads favour self-hosting because the fixed cost per token drops as throughput increases. Sporadic or unpredictable workloads favour vendor APIs as you avoid paying for idle capacity.

    Self-hosting also provides protection against vendor price changes and model deprecations, which add a different kind of cost: the engineering time to migrate when a vendor discontinues a model you’ve built on.

    Uptime, Maintenance, and Operational Burden

    Vendor APIs abstract away the infrastructure entirely. The provider manages scaling, load balancing, hardware failures, and security patching. Your team consumes a service.

    Self-hosting transfers that burden internally. Running a production-grade LLM requires people who can manage GPU memory constraints, configure inference servers (such as vLLM or Ollama), handle hardware failures, and maintain network security. If you don’t have that capability today, building it takes time and money.

    Speed is also a practical consideration. Vendor APIs are purpose-built for throughput, delivering hundreds of tokens per second. Self-hosted models struggle to match that performance without significant infrastructure investment. For latency-sensitive applications, that gap matters.

    Version Control, Model Updates, and Stability

    Vendor APIs give you access to the latest models the moment they release. New capabilities, longer context windows, and multimodal features arrive without any action on your part.

    The trade-off is stability. Vendors deprecate older models and sometimes alter model behaviour in ways that affect existing applications. An output that worked reliably on one version may behave differently after an update, and you may not get much notice.

    Self-hosting locks you to a specific version. Behaviour is consistent and reproducible for as long as you need it. Updating to a newer model is a deliberate choice, not something imposed on you. The open-source community releases capable models frequently, and quantisation techniques now allow large models to run on hardware that would have been insufficient even a year ago.

    Guardrails, Customisation, and Model Behaviour

    Vendor models ship with built-in safety guardrails. These are designed for general audiences and broad use cases. For most applications, they’re appropriate. For specialised enterprise use cases, they can be too restrictive and refusing prompts that are benign in context, or producing outputs that don’t match your organisation’s voice.

    The US government’s June 2026 export control directive suspending access to Anthropic’s Claude Fable 5 for all foreign nationals, citing a jailbreak vulnerability, illustrated a different kind of risk: vendor-side disruption that enterprises have no control over. Organisations that had built workflows on Fable 5 found access cut without warning. That event accelerated a shift in enterprise thinking toward models that no government directive can reach.

    Self-hosted models give you complete control over behaviour. You can fine-tune on proprietary data, define your own guardrails, and build an AI capability that reflects your specific domain. Competitors using generic vendor APIs cannot replicate that. The model becomes an asset, not a commodity.

    Conclusion: Match the Architecture to the Actual Risk

    The vendor vs. self-host decision comes down to where your risks sit and what you’re optimising for.

    Choose a Vendor API when:

    • Speed to market is the primary goal.
    • Usage is low or unpredictable.
    • Data leaving your network is not a compliance concern.
    • Your organisation lacks the MLOps capability to run production infrastructure.

    Choose to Self-Host when:

    • Data privacy or regulatory requirements prohibit external data transmission.
    • Utilisation is high and continuous, making per-token pricing prohibitive.
    • Deep customisation or proprietary fine-tuning is required.
    • The application must operate in offline or air-gapped environments.
    • Resilience against third-party disruption such as regulatory, commercial, or otherwise, is a business requirement.

    For many organisations, neither option alone is the right answer. A hybrid architecture routes sensitive, high-volume, or compliance-constrained work to self-hosted models, while vendor APIs handle complex reasoning tasks or public-facing applications where frontier capability matters more than data control. Simple, ad-hoc queries with no sensitive context go to the vendor. Anything carrying proprietary data stays internal.

    The architecture question is really a risk question. Map your actual risks first. The right infrastructure follows from that.

  • Designing AI Systems for Human-in-the-Loop

    Most systems are designed for a human who doesn’t exist. The assumption baked into policy, process, and AI system design is a person who is alert, rational, and fully attentive. Real operators are tired, stressed, and distracted. Closing that gap is not a matter of better training or stricter procedures. It requires designing systems that account for human cognition as it actually works, not as we wish it would.

    The Human-in-the-Loop Is Not What We Think

    The concept of the “human-in-the-loop” implies a vigilant, rational operator capable of overseeing and correcting automated systems. In high-stakes environments, that assumption collapses quickly.

    Humans get stressed, tired, and distracted. They are susceptible to automation bias, the tendency to over-rely on automated systems even when evidence contradicts them. They also engage in cognitive offloading, delegating mental tasks to machines and reducing their own engagement in the process. In AI systems, these patterns increase the likelihood of errors, misjudgements, and unintended consequences will pass right through the checks designed to stop them. The presence of a human in front of a screen does not automatically create meaningful oversight.

    As AI systems become more autonomous, the human role shifts from active control to passive monitoring. That shift is a problem. Passive monitoring over extended periods degrades situational awareness, erodes skills, and slows intervention when it matters most. Placing a human in the loop introduces its own risks when the design of that loop ignores how human attention actually works.

    Human Factors Engineering: Design for the Real Operator

    To build effective AI systems, you need an accurate picture of what humans can and cannot do. Human Factors Engineering (HFE) provides that picture. The field optimises the design of systems, tasks, equipment, and environments to enhance human performance and reduce error.

    James Reason, whose Swiss cheese model shaped modern safety thinking, captured this well. In his 1990 book Human Error, he wrote: “Rather than being the main instigators of an accident, operators tend to be the inheritors of system defects. Their part is that of adding the final garnish to a lethal brew whose ingredients have already been long in the cooking.”

    The implication is direct. When operators fail, the system usually failed first. HFE addresses that by examining three connected areas:

    The job. What does the task actually require? What is the workload, the environment, the design of controls and displays? Tasks should match human perceptual, attentional, and decision-making capabilities, not exceed them.

    The individual. What are the operator’s skills, attitudes, and physical capabilities? Some characteristics are fixed; others can be developed. Systems need to account for variation across the people who use them.

    The organisation. What work patterns, cultural norms, and communication structures are in place? These factors shape individual and group behaviour in ways that no individual training programme can override.

    Getting HFE right means designing systems where it is easier to do the right thing and harder to do the wrong thing, where errors are visible and recoverable before they become consequential.

    Decision Support Tools That Actually Support Decisions

    Cognitive limits become most dangerous under pressure. Well-designed decision support systems account for that by actively aiding interpretation and action, not just presenting data.

    Three elements matter most:

    Dashboards that prioritise. An effective dashboard does not display everything. It surfaces critical information, visualises trends, and highlights anomalies with clear hierarchies and minimal cognitive load. The goal is fast comprehension without overwhelm.

    Alerts that mean something. Alert fatigue is a genuine hazard. Alerts should be timely, relevant, and actionable. Each one should give the operator enough context to understand the problem and a clear indication of what to do next. Volume without prioritisation produces noise, not signal.

    Tools integrated into the workflow. Decision support works when it appears at the moment of need, not as a separate system the operator has to consult. Checklists, guided procedures, and automated pre-analysis reduce cognitive load by providing structure where it would otherwise be absent.

    The goal is not to replace human judgement. It is to give that judgement a better foundation.

    Decision Infrastructure: Shaping the Environment Around the Decision

    Individual tools are not enough. The broader decision infrastructure, the design choices, processes, and cultural norms that surround every decision, needs to be built with the same deliberateness.

    Behavioural science offers practical levers. Cognitive biases like confirmation bias and the availability heuristic are not character flaws. They are predictable features of how human minds work under pressure. Designers can work with them rather than against them by:

    • Setting safe defaults so that inaction does not create risk
    • Framing information to make risks and consequences legible
    • Structuring choices so the better option is also the easier one

    Feedback loops enable learning. A feedback loop feeds a system’s outputs back into its inputs to reveal cause-and-effect relationships. In human-system interaction, they make it possible to measure what is actually happening, identify where decisions are breaking down, and improve the system over time.

    Don Norman. author of “The Design of Everyday Things” has noted that the human mind struggles with interconnected systems where feedback is delayed and consequences are invisible. Good design shortens those delays and makes consequences visible before they become irreversible.

    Four principles guide effective feedback loop design. First, understand the operators’ world as they experience it, not as designers imagine it. Second, solve the right problem by tracing issues to their root causes rather than addressing symptoms. Third, treat every element as part of a larger system, because effects in complex environments are often distant from their causes. Fourth, make changes incrementally. Small, testable interventions reveal what works and allow for adjustment. Large-scale fixes rarely survive contact with reality.

    Design Systems with the Outcome in Mind

    Designing for the actual human requires accepting an uncomfortable premise: the person using your system will not always be at their best. Here is what that means in practice.

    Start with human factors. Integrate HFE from the beginning of system design. Analyse job demands, individual capabilities, and organisational influences before finalising any interface or workflow.

    Design for cognitive limits. Human attention, memory, and processing capacity are finite. Dashboards, alerts, and decision tools should reduce cognitive load, not add to it. Give operators what they need to decide, not everything you could show them.

    Apply behavioural science deliberately. Design choice architectures that make the right action the default. Use framing and feedback to guide behaviour toward outcomes that serve both the operator and the organisation.

    Build feedback loops into everything. Monitor human-system interaction continuously. Collect data on where decisions go wrong, identify patterns, and improve the system based on what you find. One failure is a signal. Two of the same failure means the system has not changed.

    Treat errors as system signals, not individual failures. Most errors reflect poor system design, not poor operators. An organisation that blames individuals for system-induced failures will keep producing those failures. One that investigates errors for systemic causes will keep reducing them.

    The rational human in the loop is a fiction worth abandoning. The stressed, tired, and distracted human is not a problem to be managed through compliance. That human is the person your system must be designed to support. Build for them, and you build something that actually works.


    References

    1. Reason, J. (1990). Human Error. Cambridge University Press.

  • Your AI is Confident. Can it Tell You Why?

    As AI moves from predictive pattern-matching to autonomous decision-making, the stakes for boards have changed. Directors don’t need to understand the technical mechanics of AI, but they do need to understand its reasoning frameworks.

    The next critical shift in AI governance moves beyond correlation toward causation and counterfactuals. This briefing explains why these concepts matter for risk management, strategic decision-making, and regulatory compliance.

    The Right Tool for the Right Question

    Most enterprise AI systems rely on statistical analysis and pattern recognition. This has been the standard analytical engine for decades and is deeply integrated across finance, operations, marketing, and risk functions. For the questions it was designed to answer, it performs well.

    The key word is designed. Correlation-based AI answers one question: “When A occurs, how often does B follow?”

    Causal AI asks something different: “Did A actually cause B?”

    That distinction matters more than most boards realise. A correlative AI has no model of why things happen, only what has happened before. Ask it to predict outcomes under novel conditions and it breaks down. It simply wasn’t built for those kinds of questions.

    This is not a criticism, but a design boundary.

    Counterfactuals: The “What If?” That Changes Everything

    To manage risk effectively, boards need to evaluate not just what did happen, but what could have happened under different circumstances. This is the domain of counterfactual reasoning.

    A counterfactual is a conditional rooted in a hypothetical:

    “Given that A happened and led to B, what would have happened if A had been different or absent?”

    In human decision-making, counterfactuals underpin accountability, ethics, and strategy. Boards ask them constantly: “If we hadn’t entered that market, would our margins have held?”

    When AI systems incorporate counterfactual reasoning, they can stress-test their own conclusions before acting on them. Rather than following a predictive model to its output, a causally aware AI can simulate alternative scenarios, testing what would change if one variable shifted, before recommending a course of action.

    This is the difference between a model that identifies what is likely and one that appreciates the drivers behind outcomes.

    What This Means for Board Governance

    The integration of causal and counterfactual AI has four concrete implications for directors exercising their fiduciary duties.

    A. Explainability and Regulatory Compliance

    Regulatory pressure toward explainable AI is accelerating globally. The EU AI Act, Australia’s AI Ethics Framework, and financial regulators across major markets are moving toward requiring that high-stakes AI decisions be justifiable, not just statistically defensible.

    Correlative models have difficulty meeting this standard. They rarely explain why a specific decision was made because they were never designed to understand cause, only pattern.

    Counterfactual reasoning provides a direct path forward. A causal AI can justify its output by stating: “The loan application was declined because the debt-to-income ratio was 45% combined with the credit score and payment history. Had that ratio been 40%, the application would have been approved.” That level of transparency supports compliance audits, withstands regulatory scrutiny, and creates the documented decision trail that protects the organisation.

    B. Systemic Bias and Legal Exposure

    Traditional AI models embed the biases present in historical data. When those models operate in credit, hiring, or pricing decisions, they can perpetuate outcomes that disadvantage protected groups, not through intent, but through the patterns they have learned.

    Counterfactual fairness testing offers a rigorous method for identifying this risk. By systematically modelling how changing a protected attribute, such as an applicant’s gender or postcode, alters the AI’s decision, organisations can determine whether that attribute is causally influencing outcomes it should not. This methodology is well established and increasingly expected by regulators as evidence of due diligence. It converts a compliance risk into a defensible position.

    C. Scenario Planning and Strategic Resilience

    Boards rely on stress-testing to navigate uncertainty. Causal AI makes that stress-testing materially more useful.

    Rather than forecasting future performance based strictly on historical data, management can model complex, multi-variable scenarios – “What if inflation rises by 2% while a key supplier faces a 30-day delay?” – with a system that maps the structural dependencies driving those outcomes, not just their historical co-occurrence.

    This is the difference between a model that has seen similar conditions before and one that identifies the mechanisms at work. In a volatile operating environment, the latter is a governance asset.

    D. Architectural Enhancements: Bridging the Reasoning Gap

    Standard LLMs are natively correlative; they predict what comes next based on patterns in training data. The frontier of AI deployment involves layering structured reasoning protocols over these models to address that limitation directly.

    Research into counterfactual inference frameworks shows measurable results. Studies applying structured causal reasoning algorithms to frontier models have achieved accuracy rates above 90% on complex causal logic tasks. This is a substantial improvement over unassisted LLMs, which show accuracy drops of 25–40 percentage points when tested on counterfactual reasoning compared to standard pattern-matching tasks. Separately, counterfactual probing approaches have demonstrated hallucination reductions in the range of 20–25% on established benchmarks.

    These are meaningful gains, but they also illustrate the scale of the gap that unassisted, correlative AI leaves open. For the board, the implication is clear: operational risk and hallucination are not inherent, unfixable flaws of AI. They are architectural challenges that respond to rigorous engineering.

    Questions the Board Should Be Asking

    To ensure the organisation is prepared, directors should test their current AI governance posture against three questions:

    1. Audit Capability: Does our AI risk framework require systems to provide counterfactual explanations for high-stakes decisions – credit, pricing, hiring – or do we accept outputs without traceable reasoning?
    2. Model Resilience: How exposed are our operational AI models to conditions they have not encountered before? Are we over-reliant on purely correlative systems in areas where novel risk is most likely?
    3. Governance Alignment: Is our AI governance policy keeping pace with regulatory demands for transparency and explainability, or are we managing to a standard that is already being superseded?

    Conclusion

    As AI becomes a strategic actor, its reasoning must be held to the same scrutiny as executive decision-making. Boards that champion causal clarity and counterfactual rigour are managing compliance and building the infrastructure for decisions that are defensible, resilient, and genuinely informed.

  • Encoding Expert Intuition: Cognitive Task Analysis for AI Agents

    You’ve given your AI agent thousands of examples. Historical records, past decisions, expert-authored reports. The data is comprehensive. So why is the agent still getting it wrong?

    The answer lies in what the data doesn’t contain. When an AI agent trains on an expert’s emails or historical logs, it learns the outputs of decisions, not the judgement, pattern recognition, or edge-case handling that produced them. The expert who wrote those records was drawing on years of accumulated intuition, scanning for cues that never made it into the document, and applying mental models that existed only in their head.

    To capture that invisible expertise, a growing number of AI development teams are turning to Cognitive Task Analysis (CTA) – a structured methodology for eliciting tacit knowledge from domain experts. By systematically mapping how experts perceive cues, weigh options, and navigate ambiguity, CTA translates human intuition into concrete engineering specifications.

    Why Historical Data Is Not Enough

    Large Language Models operate on statistical probabilities, predicting the next most likely token based on vast datasets. This is highly effective for language generation. It becomes unreliable when applied to complex, high-stakes decision-making.

    Two limitations become apparent quickly.

    The first is reasoning. LLMs do not possess an inherent understanding of physical-world constraints or domain-specific logic. When they encounter novel situations, such as edge cases not well-represented in training data, they tend to produce confident-sounding responses that are fundamentally wrong. This “hallucination” problem is particularly dangerous in expert domains, where a plausible-sounding error can be indistinguishable from the correct answer.

    The second is situational awareness. A human expert doesn’t passively receive information; they actively scan their environment for subtle cues, filtering out noise to focus on what matters. An LLM treats all information in its context window with roughly equal weight unless explicitly instructed otherwise. Without the ability to distinguish critical signals from irrelevant data, an AI agent can make decisions that appear logical on the surface but fail in practice.

    The solution isn’t more data. It’s a more rigorous approach to capturing what the expert actually does and the reasons why.

    Eliciting the Knowledge That Experts Can’t Easily Articulate

    Cognitive Task Analysis encompasses a range of techniques designed to surface the cognitive work that experts perform without consciously thinking about it. This is not a single interview. It is a structured investigation into how an expert perceives situations, what draws their attention, and how they make decisions when the stakes are high and the information is incomplete.

    The process falls roughly into five stages:

    1. Preparation: Understand the problem domain, the nature of the tasks being performed, and which analytical methods are best suited to uncovering the relevant expertise.
    2. Knowledge Elicitation: Apply specific methods to capture the key decisions and cognitively demanding tasks that define expert performance.
    3. Analysis and Representation: Decompose the elicited data and restructure it into a form that can be translated into system design.
    4. Implementation: Iteratively apply the identified decisions and strategies to the AI system being built.
    5. Evaluation: Establish clear performance measures, evaluate the results, and improve based on evidence.

    Applied Cognitive Task Analysis (ACTA)

    The most practical framework for AI development teams is Applied Cognitive Task Analysis. ACTA streamlines the knowledge elicitation process through three primary methods:

    1. Task Diagram: The analyst works with the expert to produce a high-level map of the task, specifically identifying which subtasks require the most cognitive effort or judgement.
    2. Knowledge Audit: Probes how the expert diagnoses problems, anticipates future states, and recognises anomalies. The focus is on the specific cues the expert relies on and the strategies they employ.
    3. Simulation Interview: The expert is presented with a challenging, realistic scenario and asked to walk through their decision-making process step by step. This reveals how they handle pressure, ambiguity, and shifting priorities in conditions that mirror real work.

    Translating Expert Judgement into Agent Architecture

    The practical value of CTA lies in what you do with the findings. The outputs of analysis map directly onto the components of modern AI agent architectures.

    CTA FindingAI Agent ComponentImplementation Strategy
    Critical CuesContext Engineering / MemoryEnsure the agent’s system prompt or Retrieval-Augmented Generation pipeline explicitly surfaces these specific data points before a decision is made.
    Expert StrategiesPrompt EngineeringTranslate the expert’s mental model into Chain-of-Thought instructions, directing the LLM to replicate the expert’s step-by-step reasoning process.
    Common ErrorsGuardrails / ConstraintsUse identified pitfalls to write explicit negative constraints in the system prompt (e.g., “Never approve a claim if X is missing”).
    Decision PointsAgentic Workflows / Human-in-the-loopDesign the agent’s workflow to pause at high-stakes decision points, either triggering a specific sub-agent for deeper analysis or requesting Human-in-the-Loop validation.

    This mapping also creates an evaluation framework. When an AI or RAG pipeline isn’t performing as expected, the CTA findings provide a structured lens for identifying precisely where the reasoning breaks down and what needs to change.

    Conclusion

    Giving an AI agent a large dataset of historical decisions is a necessary first step. It is not sufficient for building a system you can trust with consequential decisions.

    To move beyond brittle automation, you need to look beyond the data and examine how the expert actually thinks. Cognitive Task Analysis provides the methodology to extract that tacit knowledge; systematically mapping cues, strategies, and decision points so that your AI agent doesn’t merely mimic expert language, but genuinely replicates expert reasoning.

    The problem you’re trying to solve existed long before you noticed it. The expertise you need to encode has been sitting in people’s heads all along.


    Contact Galdren to explore how Cognitive Task Analysis can be applied to your AI development process.

  • Measuring AI Risk: A Practical Framework

    You can’t govern what you can’t measure. That’s the core problem facing organisations deploying AI today.

    Most AI risk conversations stay at the level of vague concern: “we need to manage bias”, “we should think about data privacy”. That’s not governance. It’s wishful thinking. Real AI governance requires specific metrics, clear thresholds, and defined accountability for each category of risk your systems create.

    The below framework covers ten AI risk categories that most consistently cause operational failures, regulatory penalties, and reputational damage. For each category, it provides the primary business impact, recommended mitigation strategies, and the specific metrics your teams should be tracking in production.

    How to use this framework: Not all ten categories apply equally to every deployment. Start by assessing which risks are most material to your specific AI systems and business context. Use the metrics as a baseline, adapt thresholds to your risk appetite and regulatory environment. Revisit your measurements as your systems evolve, because AI risk changes as models drift, data shifts, and threat actors adapt.


    1. Model Inaccuracy

    Primary Business Impact: Operational failure, loss of customer trust

    An AI model is only as valuable as its predictions are reliable. When a predictive maintenance model misses an impending equipment failure, the result is unplanned downtime. When a customer-facing AI provides incorrect information, the result is frustrated users and eroded trust. Model inaccuracy compounds over time through a process called concept drift, the model’s training data becomes less representative of the real world it’s operating in, and performance quietly degrades until something breaks visibly.

    Mitigation Strategies: Continuous monitoring, human-in-the-loop

    Continuous monitoring of production performance is the foundation. Pair this with human-in-the-loop (HITL) mechanisms that allow expert review of high-stakes decisions. HITL serves a dual purpose: it catches errors before they cause harm, and the correction data it generates feeds back into model improvement.

    Measurement Metrics for Model Inaccuracy:

    MetricDescriptionHow to Measure
    Drift RateThe rate at which confident incorrect predictions increase over time.Monitor the proportion of high-confidence predictions subsequently identified as incorrect through human review or ground truth data.
    Accuracy / Precision / Recall / F1-scoreStandard statistical measures of predictive performance.Calculate against a held-out test set or ground truth data. Accuracy measures overall correct predictions; Precision measures true positives among all positive predictions; Recall measures true positives among all actual positives; F1-score is the harmonic mean of precision and recall.
    Mean Absolute Error (MAE) / Root Mean Square Error (RMSE)Average magnitude of errors for regression tasks.Calculate the average absolute difference (MAE) or the square root of the average squared differences (RMSE) between predicted and actual values.
    HITL Intervention RateThe percentage of model outputs requiring human correction or override.Track the frequency of human interventions in the AI workflow. A high rate signals significant inaccuracy or the need for retraining.
    Edge Case Failure RateHow often the model fails on rare, unusual, or out-of-distribution inputs.Design specific test cases or monitor real-world performance on identified edge cases. Requires domain expertise to define what constitutes an edge case.

    2. Data Privacy (PII)

    Primary Business Impact: Regulatory fines, legal liability

    AI systems process vast quantities of data, including sensitive personal information. Without rigorous controls, that data can leak into model outputs, training datasets, or logs, exposing individuals and triggering regulatory action. The EU AI Act and GDPR both impose substantial penalties for PII mishandling, and the complexity of AI pipelines makes it easy to lose track of exactly where personal data flows and how it’s used.

    Mitigation Strategies: Advanced PII classifiers, differential privacy

    Automated PII classifiers identify and categorise sensitive data before it reaches models or outputs. Differential privacy adds carefully calibrated noise to data, preserving aggregate utility while making it mathematically difficult to reconstruct individual records. Together, these approaches reduce exposure without eliminating the data’s analytical value.

    Measurement Metrics for Data Privacy (PII):

    MetricDescriptionHow to Measure
    PII Leakage RateHow often sensitive personal information appears inadvertently in AI outputs or logs.Monitor system outputs, logs, and data flows using automated scanning tools configured to detect PII that should have been protected or anonymised.
    Differential Privacy EpsilonA quantitative measure of privacy loss in a differentially private algorithm. Lower epsilon means stronger privacy guarantees.Calculate the epsilon value for algorithms employing differential privacy. This requires detailed understanding of the algorithm’s mechanics and the data it processes.
    Re-identification Risk ScoreThe estimated probability that an individual can be uniquely identified from supposedly anonymised data.Use statistical methods and re-identification tests to assess the likelihood of linking anonymised data back to individuals. K-anonymity and l-diversity assessments can inform this score.
    Data Minimisation RatioThe proportion of sensitive data used relative to total available sensitive data. Lower is better.Audit data pipelines and storage to quantify PII collected and processed versus the minimum required for the AI system’s function.

    3. Shadow AI

    Primary Business Impact: Uncontrolled data egress, security gaps

    When employees adopt AI tools without IT or security approval, they create risk the organisation can’t see and therefore can’t manage. Sensitive corporate data gets fed into external models. Unapproved third-party services introduce unvetted security vulnerabilities. The absence of oversight can expose the organisation to data breaches, compliance violations, and intellectual property loss, all from tools that nobody officially sanctioned.

    Mitigation Strategies: Discovery tools, strict procurement policies

    Network monitoring and endpoint detection reveal unauthorised AI applications that are already in use. But discovery alone isn’t enough. Procurement policies need teeth: every AI tool adoption should require security, legal, and ethical review before deployment. The goal is centralised visibility without creating so much friction that employees route around the process anyway.

    Measurement Metrics for Shadow AI:

    MetricDescriptionHow to Measure
    Unsanctioned AI Usage CountThe number of unauthorised AI tools or services detected within the organisation’s network.Use network monitoring, endpoint detection and response (EDR) solutions, and cloud access security brokers to identify and count unapproved AI applications.
    Data Egress Volume to AI DomainsTotal data volume transferred from internal systems to external, unapproved AI service endpoints.Monitor network traffic and data loss prevention (DLP) systems for transfers to unsanctioned AI service domains.
    Procurement Bypass RateThe percentage of AI tool acquisitions occurring outside established procurement processes.Audit expense reports, software licences, and cloud service usage against approved vendor lists and procurement records.
    Discovery CoverageThe percentage of the enterprise network, endpoints, and cloud environments actively monitored for Shadow AI.Assess the scope and effectiveness of AI discovery tools across the full IT infrastructure.

    4. Agentic Risk

    Primary Business Impact: Autonomous system failure, goal hijacking

    Agentic AI systems – those that plan, execute, and adapt autonomously to achieve goals – operate with a degree of independence that creates new categories of failure. They can make compounding errors without human checkpoints. They can be manipulated through prompt injection into pursuing objectives entirely different from their original purpose. The more autonomous the system, the wider the gap between what it does and what a human would have approved in the moment.

    Mitigation Strategies: Sandboxing, restricted tool access

    Sandboxing allows safe observation of agent behaviour before deployment in consequential environments. Restricting which tools and resources an agent can access limits the blast radius when things go wrong, and in sufficiently complex systems, things will go wrong. These aren’t permanent constraints; as an agent demonstrates reliable behaviour in controlled conditions, access can expand incrementally.

    Measurement Metrics for Agentic Risk:

    MetricDescriptionHow to Measure
    Task Success RateThe percentage of times an AI agent successfully achieves its assigned objective.Define clear objectives and success criteria. Measure task completion rates in controlled environments and real-world deployments.
    Tool Call AccuracyThe precision with which an agent selects and correctly uses external tools or APIs.Monitor agent logs to assess whether correct tools are invoked with appropriate parameters. Compare against expert-defined optimal tool usage.
    Goal Hijacking RateHow often an agent deviates from its intended objective due to external manipulation or internal misalignment.Test with adversarial prompts designed to attempt goal hijacking. Track instances where agent behaviour deviates from intended outcomes.
    Autonomous Step CountThe number of actions an agent takes without human intervention or oversight.Log the sequence of agent actions and identify points where human review would typically occur. Higher counts indicate greater autonomy and require proportionally stronger controls.
    Reasoning Coherence ScoreAn assessment of the logical consistency of an agent’s decision-making process.Use human evaluators or an LLM-as-a-Judge approach to score the agent’s reasoning steps for logical soundness and alignment with expected behaviour.

    5. Adversarial Attacks

    Primary Business Impact: System compromise, data poisoning

    Adversarial attacks are deliberate attempts to exploit AI system vulnerabilities. Subtle perturbations in images can cause misclassification. Carefully crafted prompts can hijack the behaviour of large language models. Poisoned data introduced into training sets can degrade performance or create hidden backdoors. These attacks are designed to look like normal inputs while producing abnormal outcomes.

    Mitigation Strategies: Robustness testing, prompt filtering, HITL

    Robustness testing evaluates models against known adversarial techniques before attackers get the chance. Prompt filtering – through input validation, sanitisation, and anomaly detection – reduces the attack surface by catching malicious inputs before they reach the core model. Neither measure is sufficient alone; layering them together significantly raises the cost of a successful attack.

    Measurement Metrics for Adversarial Attacks:

    MetricDescriptionHow to Measure
    Attack Success RateThe percentage of adversarial inputs that successfully produce incorrect or manipulated outputs.Conduct controlled experiments with adversarial examples and record the rate of misclassification or undesired behaviour.
    Robustness ScoreA model’s measured resilience to various adversarial attack techniques.Evaluate model performance under different types and strengths of adversarial perturbations. Higher scores indicate greater resilience.
    Filter Bypass LatencyThe average time or effort required to craft an adversarial input that bypasses existing defences.Simulate attack scenarios and measure the time and number of attempts needed for a successful bypass.
    Prompt Injection SensitivityThe threshold at which a model begins to follow malicious instructions over its intended system instructions.Test models with a range of prompt injection techniques, varying strength and subtlety, and observe adherence to malicious instructions at each level.

    6. Bias & Discrimination

    Primary Business Impact: Reputational damage, legal action

    AI models don’t create bias from nothing. They learn it from training data that reflects existing societal inequities. A hiring algorithm trained on historical decisions inherits the patterns of who was hired historically. A loan approval model reflects the lending patterns of the past. The problem compounds when these biases aren’t surfaced before deployment and the model makes thousands of consequential decisions before anyone notices the pattern.

    Anti-discrimination law doesn’t care that the bias came from training data. The legal and reputational exposure is the same as deliberate discrimination.

    Mitigation Strategies: Fairness audits, diverse training sets

    Regular fairness audits systematically identify and quantify biases in model outputs across demographic groups. Diverse training sets that accurately represent target populations reduce the initial bias the model learns. Techniques including re-sampling, re-weighting, and adversarial debiasing can further reduce bias in models where the training data alone is insufficient.

    Measurement Metrics for Bias & Discrimination:

    MetricDescriptionHow to Measure
    Disparate Impact RatioCompares the selection or outcome rate for a protected group against a majority group. A ratio below 0.8 (the “80% rule”) indicates potential disparate impact.Calculate the ratio of selection rates (e.g., hiring rate, loan approval rate) for a protected group versus the most favoured group.
    Equalized Odds / Demographic ParityStatistical fairness metrics assessing whether a model performs equally across groups (Equalized Odds) or produces positive outcomes at equal rates across groups (Demographic Parity).For Equalized Odds, compare true positive and false positive rates across groups. For Demographic Parity, compare the proportion of positive predictions per group.
    Bias Amplification FactorThe extent to which a model amplifies biases already present in its training data.Compare bias observed in model outputs against bias in the input data. A factor greater than 1 indicates amplification.
    Stereotype Consistency ScoreHow often a model generates outputs that align with harmful stereotypes about specific demographic groups.Design targeted prompts to test for stereotypical associations and measure the rate at which the model reinforces them.

    7. Regulatory Fragmentation

    Primary Business Impact: Compliance overhead, market exit

    The global AI regulatory landscape is accelerating in every direction simultaneously. The EU AI Act imposes obligations based on risk classification. GDPR governs personal data. National AI strategies are emerging across jurisdictions with different requirements and enforcement approaches. An organisation operating internationally may face genuinely conflicting obligations; what one jurisdiction requires, another prohibits or leaves undefined.

    This isn’t a problem that resolves itself with time. The organisations that treat regulatory fragmentation as a compliance administration problem will find themselves perpetually reactive and expensive to operate.

    Mitigation Strategies: Global policy mapping, NIST AI RMF

    Systematic mapping of international AI regulations to internal governance frameworks, before those regulations come into force, is the only sustainable approach. The NIST AI Risk Management Framework provides a common structure that satisfies significant portions of requirements across multiple global regulations, giving organisations a stable foundation to build jurisdiction-specific controls on top of.

    Measurement Metrics for Regulatory Fragmentation:

    MetricDescriptionHow to Measure
    Compliance Coverage ScoreThe percentage of relevant global AI regulations for which the organisation has established corresponding internal controls and policies.Conduct a regulatory mapping exercise comparing internal policies against external requirements, and calculate the proportion of covered regulations.
    Audit Readiness ScoreThe time and resources required to produce comprehensive compliance evidence for a specific AI system or regulation.Track effort involved in preparing for and undergoing AI-related audits. Lower effort indicates higher readiness and better-maintained documentation.
    Policy Update LatencyThe average time from enactment of a new AI regulation to the corresponding adjustment of internal policies and controls.Monitor regulatory developments alongside internal policy revision cycles. Shorter latency indicates greater organisational agility.
    Market Exit Risk FactorThe organisation’s exposure to operational restrictions in specific jurisdictions due to inability to comply with local AI regulations.Evaluate the stringency of regulations in key operating markets and assess the organisation’s capacity to meet them.

    8. IP Infringement

    Primary Business Impact: Copyright litigation, loss of IP, loss of reputation

    AI models trained on large datasets frequently ingest copyrighted material. Generative AI can produce outputs substantially similar to works those models were trained on. The result is potential copyright infringement at scale, automatically, invisibly, and without any deliberate intent on the part of the organisation deploying the model.

    The legal landscape around AI-generated content and training data is still developing, but litigation is already underway in multiple jurisdictions. Organisations cannot afford to wait for case law to settle before managing this risk.

    Mitigation Strategies: Data provenance tracking, legal review

    Data provenance tracking ensures all training data is properly licenced, attributed, and traceable to its source. Legal review of AI-generated outputs, such as for generative AI applications, identifies potential IP issues before deployment or public release. These controls need to be built into the AI development lifecycle, not applied as an afterthought.

    Measurement Metrics for IP Infringement:

    MetricDescriptionHow to Measure
    Data Provenance ScoreThe percentage of training data for which origin, licencing, and usage rights are fully documented and verifiable.Audit training datasets for complete and accurate records of data sources, licences, and terms of use.
    Copyright Similarity IndexA quantitative measure of overlap between AI-generated content and known copyrighted works.Use computational tools and human review to compare AI outputs against databases of copyrighted material, identifying substantial similarity.
    Licence Violation CountThe number of instances where licensing terms for training data or model usage have been breached.Track and log deviations from licencing agreements, including unauthorised use of data or models.
    IP Indemnification CoverageThe percentage of AI outputs or components covered by legal indemnification clauses from third-party AI vendors.Review vendor contracts and SLAs to assess the extent of IP indemnification provided.

    9. Operational Complexity

    Primary Business Impact: Integration debt, high maintenance costs, security exposure

    AI systems rarely operate in isolation. They connect to data pipelines, monitoring infrastructure, deployment platforms, and downstream business processes. As these systems scale, the connections multiply. Each connection is a potential failure point, and the effort to maintain them compounds over time. Teams that spend the majority of their engineering capacity keeping existing systems running have little left for improvement or innovation.

    This isn’t an AI-specific problem, but AI accelerates it. Models need retraining. Data pipelines need updating. Infrastructure needs scaling. Without deliberate architectural choices, complexity grows faster than the capacity to manage it.

    Mitigation Strategies: Modular architecture, MLOps maturity

    Modular architecture allows independent development, deployment, and scaling of AI components, limiting the cascading impact of any single failure. MLOps maturity, applying DevOps and security principles to the full machine learning lifecycle, automates the most labour-intensive parts of model management, from data ingestion through to monitoring and retraining. The investment in MLOps pays back through reduced operational overhead and faster, safer iteration.

    Measurement Metrics for Operational Complexity:

    MetricDescriptionHow to Measure
    Integration Debt ScoreA measure of technical debt from complex, non-standard, or poorly documented integrations between AI system components.Assess the number of manual integrations, custom scripts, and non-standard APIs in use. Higher scores indicate greater fragility and maintenance burden.
    MLOps Maturity LevelAn assessment of adherence to MLOps best practices across automation, collaboration, and continuous delivery.Use established MLOps maturity models (e.g., Google’s MLOps maturity model) to score capabilities in areas including CI/CD for ML, automated testing, and model monitoring.
    Maintenance-to-Development RatioThe proportion of engineering resources allocated to maintaining existing AI systems versus building new capabilities.Track resource allocation between maintenance tasks (bug fixes, infrastructure updates, model retraining) and new development. A high ratio indicates unsustainable operational complexity.
    System Latency / ThroughputThe responsiveness of the AI system (latency) and the volume of requests it can handle within a given time (throughput).Monitor real-time performance metrics in production. High latency or low throughput under normal load indicates bottlenecks worth investigating.

    10. Third-Party AI Opacity

    Primary Business Impact: Supply chain vulnerability, PII and security exposure, reputation impact

    Most organisations don’t build their AI from scratch. They use cloud-based ML platforms, pre-trained models, and AI-powered APIs. Those third-party systems are often proprietary black boxes. The vendor’s internal workings, data sources, and risk profiles are not disclosed, and may not be disclosable. When something goes wrong with a third-party model, the organisation bearing the business consequence may have had no visibility into the risk that caused it, but are ultimatly responsible for it.

    This is supply chain risk applied to AI, and it compounds when third-party vendors themselves rely on sub-vendors, open-source components, and external data sources.

    Mitigation Strategies: Vendor risk assessments, SLA enforcement

    Comprehensive vendor risk assessments should require Model Cards, System Cards, and audit reports from suppliers, not as a compliance checkbox, but as genuine evaluation of their AI governance practices. SLA enforcement converts transparency requirements into contractual obligations with consequences, covering performance, security, ethical standards, and incident response. Organisations that exercise their audit rights regularly send a clear signal about the standard they expect.

    Measurement Metrics for Third-Party AI Opacity:

    MetricDescriptionHow to Measure
    Vendor Transparency ScoreA composite score reflecting the completeness and clarity of documentation provided by third-party AI vendors.Evaluate vendor-provided Model Cards, System Cards, audit reports, and security certifications against a predefined checklist of transparency requirements.
    SLA Compliance RateHow consistently third-party AI providers meet contractual SLAs related to performance, uptime, security, and data handling.Monitor and track vendor performance against agreed SLAs, including downtime, response times, and adherence to security protocols.
    Supply Chain Vulnerability IndexA measure of risks introduced by upstream dependencies in third-party AI services, including sub-vendors, open-source components, and data sources.Map the supply chain of third-party AI solutions, identifying and assessing the risk profile of each component and sub-vendor.
    Audit Rights Exercise RateHow often the organisation exercises its contractual rights to audit third-party AI vendors.Track audits conducted on third-party AI providers and findings from those audits. Regular exercise of audit rights is a leading indicator of proactive risk management.

    Conclusion

    The ten categories in this framework – Model Inaccuracy, Data Privacy (PII), Shadow AI, Agentic Risk, Adversarial Attacks, Bias & Discrimination, Regulatory Fragmentation, IP Infringement, Operational Complexity, and Third-Party AI Opacity – Each has produced documented business failures in organisations that treated AI risk as someone else’s problem, or as a problem for later.

    Measurement is where governance becomes real. The metrics here give your teams something concrete to track, report on, and improve over time. They make risk visible, and visible risk can be managed.

    Start with the categories most material to your current deployments. Build measurement into your operating rhythm, not as an audit exercise but as a production standard. And when your metrics show something unexpected, treat it as information: a signal that something in your system or environment has changed, and an invitation to understand why before it becomes a crisis.