Category: Organisational

  • Training People to Think With, Against, and Through the Machine

    The real AI skill is judgement

    Many organisations describe their AI training goals in deceptively simple terms: teach people how to use the tools. That usually means showing employees how to write a prompt, summarise a document, generate a presentation, or ask a chatbot for ideas.

    AI is capable of very good tactical responses. The day-to-day, line-by-line work a capable junior or mid-level employee would produce. Those capabilities matter, but they are only the entry point, because the more consequential skill is not producing an answer with AI. It is deciding whether the answer deserves confidence, identifying what it missed, exposing where it might fail, and improving it through deliberate interaction.

    There’s a long list of damaging failures where the AI produced a confident answer that turns out to be based on flawed assumptions or made up facts. Using some simple approaches can catch these mistakes before they reach a client or regulator while still allowing AI’s speed.

    The core idea is treating AI as a dynamic sparring partner: a system that can propose, challenge, reframe, simulate, and revise, but whose contributions must be examined by a capable human. This changes the purpose of AI training. The goal is not merely fluency with interfaces, as with traditional software. It is the development of disciplined judgement under conditions in which plausible language can conceal weak reasoning.

    Prompt literacy is not enough, judgement is the harder skill

    Prompt literacy asks, “How do I get the model to produce what I want?” Judgement literacy asks a harder sequence of questions: What should the model be asked to do? What would a good answer contain? Which assumptions are hidden in the response? What evidence would change my mind? How could this answer fail in practice?

    That distinction is central because AI systems are optimised to produce useful-looking responses, not to guarantee that every claim is correct, complete, current, or appropriate for the situation. A confident paragraph may combine sound reasoning with an unsupported assumption. A concise recommendation may omit a constraint that a subject-matter expert would regard as decisive. A polished analysis may be internally coherent while based on incorrect premises.

    Training must therefore teach employees to separate fluency from validity, because AI producing outputs at scale, a small error rate becomes large absolute numbers. A team processing 500 customer communications a week with a 3% undetected error rate has 15 problems a week compounding.

    Treat AI as an argument simulator, not an answer machine

    A productive human–AI interaction has several distinct moves. The user first asks the AI to generate a provisional response. Instead of accepting it, the user then assigns the system an adversarial role: critic, sceptical customer, hostile reviewer, domain expert, risk officer, opposing counsel, or operational implementer. The user asks it to identify weaknesses, test assumptions, and present alternatives. Finally, the user decides what to retain, verify, revise, or reject.

    This process is more reliable than asking for a single “best answer” because it creates structured friction. The AI is used to generate content and also to interrogate the content it produces. Its role shifts from answer machine to argument simulator.

    StageHuman responsibilityUseful AI roleOutput to preserve
    FrameDefine the objective, audience, constraints, and stakesAsk clarifying questions and expose ambiguityA precise problem statement
    GenerateRequest a provisional answer without treating it as finalProduce options, hypotheses, drafts, or modelsMultiple candidate approaches
    CritiqueInspect logic, evidence, omissions, and assumptionsAct as an adversarial reviewerA failure and risk register
    Stress-testCompare the idea against edge cases and real-world constraintsSimulate stakeholders, scenarios, and objectionsConditions under which the idea breaks
    RefineApply human expertise and verified evidenceRewrite, reorganise, and make uncertainty explicitA revised, decision-ready artifact
    ValidateOwn the final judgement and consequencesHelp create checklists or audit trailsA documented approval decision

    The critique protocol preventing AI’s fluency from masking weak reasoning

    A useful training programme should provide a repeatable critique protocol. Employees can begin by asking whether the response answered the actual question rather than a nearby, easier one. They should then examine the assumptions: What does the answer presume about the customer, market, data, timeline, resources, law, or operating environment? Which assumptions are explicit, and which are hidden?

    The next step is to inspect the evidence. Are factual claims traceable to reliable sources? Does the response distinguish observed facts from estimates, interpretations, and recommendations? Does it indicate uncertainty where uncertainty matters? A response that makes no distinction between “we know”, “we infer”, and “we might try” is difficult to govern.

    Employees should also look for omissions. What relevant stakeholder is absent? What downside is underdeveloped? What implementation cost has been ignored? What would a sceptical expert object to? In many professional settings, the most dangerous error is not a false statement but a missing consideration.

    A compact minimal critique sequence can be remembered as TRACE:

    LetterQuestion
    T — TargetDid the output address the real objective and intended audience?
    R — ReasoningAre the logic, assumptions, and causal links sound?
    A — AccuracyWhich claims require verification, and what evidence supports them?
    C — CoverageWhat perspectives, constraints, risks, or alternatives are missing?
    E — ExecutionCould a real person or team implement this, and what would fail first?

    The protocol matters less than the habit. Critique should become a normal stage of work, not an emergency response after an AI-generated mistake reaches a customer, executive, or regulator. Even one good adversarial question catches more than zero.

    The strongest exercises make people disagree with the AI, not just improve it

    The strongest exercises do not ask participants merely to improve an AI answer. They train the reflex to know when to disagree with it.

    A team might give an AI system a proposed product launch plan and ask it to produce three critiques: one from a cash-constrained finance lead, one from a sceptical customer, and one from an operations manager responsible for execution. Participants then rank the criticisms by importance, identify which are supported by evidence, and decide what additional information is needed.

    A second exercise is the assumption reversal. Participants take a central assumption in the AI’s response and invert it. If the plan assumes rapid adoption, they ask what happens if adoption is slow. If it assumes reliable data, they ask how the recommendation changes when the data is incomplete or biased. If it assumes users will follow a process, they ask what incentives would cause users to bypass it.

    A third is the minimum viable rebuttal. Each participant must name the single strongest reason not to accept the AI’s recommendation. The purpose is not to be contrarian. It is to prevent the tendency to equate a polished output with a persuasive one.

    A fourth is the pre-mortem. Participants assume the project has already failed completely, then work backward to identify what caused it. This removes the social pressure to stay quiet in a room full of consensus, because the failure is already stipulated.

    Accountability requires making human judgement visible in the work

    Organisations should not evaluate AI training by counting prompts or measuring how quickly employees produce drafts. Those metrics reward activity. Better measures reward judgement: the number of material assumptions identified, the proportion of important claims verified, the quality of alternatives considered, and the clarity with which uncertainty is communicated.

    A practical workflow is to use an AI work note for consequential outputs. It need not be long. It can record the task given to the system, the key assumptions detected, the main criticisms generated, the facts independently checked, the changes made by the human, and the person who approved the final version. This creates a lightweight audit trail without making every interaction bureaucratic. As a side effect, in high-pressure organisations when something goes wrong the person who approved the AI output without documented scrutiny is exposed. Where the person who has a note showing they checked the key assumptions is not.

    Managers should also distinguish between different levels of risk. A brainstorming exercise may require only a basic plausibility check. A customer communication, hiring recommendation, safety procedure, financial analysis, or policy decision requires a more demanding review. The higher the stakes, the more the process should emphasise source verification, domain expertise, independent reasoning, and explicit approval.

    Correction must be rewarded, not tolerated

    No training framework will work if employees believe that questioning AI marks them as inefficient or resistant to technology. Leaders must communicate that revision is not failure but the mechanism by which value is created.

    This cultural point applies to humans as well as AI outputs. People routinely accept outputs that confirm their preferences, especially when those outputs are articulate and fast. AI can amplify that tendency by presenting a conclusion before the user has fully examined the problem. The organisation must reward people who find a flaw early, surface an inconvenient alternative, or slow down a high-stakes decision long enough to validate its premises.

    People who use the sparring-partner approach understand their work better, because interrogating the AI forces them to articulate what they know and don’t know. Leaders can model the behaviour by asking “What would make this wrong?” and “Show me the strongest case against this recommendation” before asking whether the answer is useful. Over time, these questions become part of the institution’s decision vocabulary.

    The goal is better thinking

    Treating AI as a sparring partner does not mean distrusting every output or forcing every task through an elaborate review ritual. It means assigning the system the right role for the task. AI can be a rapid generator, tireless critic, perspective simulator, editor, tutor, and rehearsal partner. It cannot replace accountability for decisions whose consequences belong to people and institutions.

    The central discipline is simple: generate broadly, challenge deliberately, verify selectively, and decide consciously. When people are trained this way, AI becomes more than a productivity shortcut. While additive initially, it becomes a structured environment for thinking, one that helps users see alternatives, discover weaknesses, and improve the quality of their own judgement. People who develop this capability become the ones their organisation trusts with higher-stakes work, because they’re visibly producing better outputs with fewer errors.

    The organisations that benefit most from AI will not necessarily be those that use it the fastest. They will be those that learn how to build a culture where speed is measured over the long term.

    Start here: a 30-minute exercise any team can run this week

    Bring one ordinary work product such as a proposal, briefing, process note, or customer message. Ask an AI system to improve it. Then ask the system to criticise its own revision from three different perspectives. Have the team independently identify the two most serious weaknesses, verify the most consequential claims, and produce a final version with changes documented.

    At the end, ask three questions: What did the AI notice that we missed? What did we notice that the AI missed? Which part of the final judgement could not responsibly be delegated?

    The answers will reveal the real training agenda. AI competence is not the ability to obtain an answer. It is the ability to engage an answer critically enough to make it better, and to know when it should not be trusted at all.

  • AI Moves at the Speed the Business Can Absorb

    The pressure to “do AI” has become one of the loudest forces in business. Boards ask for a strategy, competitors announce pilots, and employees experiment with new tools before governance has caught up. In that environment, speed is often mistaken for progress. The more useful leadership question instead of How quickly can we deploy AI? is How quickly can our organisation responsibly absorb the change it creates?

    That distinction matters because AI does not merely add a new application to the technology stack. Used seriously, it changes workflows, decision rights, data practices, roles, controls, customer interactions and, in some cases, the economic logic of a business. If those changes arrive faster than the organisation can understand, govern and operationalise them, the outcome is not transformation or improvement, but accumulated work, inconsistent practices and a loss of confidence.

    Jeff Wilke, a senior executive who worked at Amazon for 25 years, made this point bluntly to Jeff Bezos early in the company’s life. Bezos was generating ideas faster than the organisation could act on them, and Wilke told him: “You have to release the work at the right rate that the organization can accept it.” Bezos later described it as a profound insight. Every idea released beyond the organisation’s capacity to absorb it did not accelerate progress. It created distraction.

    That principle applies directly to AI adoption. The objective is not to slow down for its own sake but to sequence change so that each step increases the organisation’s capacity for the next.

    Moving too fast creates invisible failure

    When leaders mandate AI adoption at a pace the business cannot sustain, the first failure is often invisible. A team may launch an impressive pilot, produce an executive demonstration, or buy licences at scale. The operating model around the tool, however, remains unfinished. Staff do not know when to rely on the output, where sensitive information may go, who owns errors, how exceptions are handled or what is expected of them once the old process is retired.

    The result is a familiar pattern: workarounds multiply, quality varies by team, risk teams intervene late, and employees quietly revert to older methods when the new approach causes friction. AI earns a reputation as a management fad rather than a source of practical value. In the most severe cases, an organisation disrupts a reliable service model before it has built a dependable replacement. The approach fails because the change exceeds the organisation’s ability to cope, not because the model used was insufficiently powerful.

    A useful way to think about this is as a mismatch between deployment velocity and absorption capacity.

    DimensionWhen deployment outpaces capacityWhen capacity sets the pace
    Process designAI is layered onto unclear or unstable work.The team redesigns one defined workflow and clarifies hand-offs.
    PeopleEmployees are told to adopt, but not taught how to exercise judgement.Training, job aids and feedback are built into the rollout.
    Data and controlsData use and accountability are resolved after launch.Guardrails, escalation paths and quality checks are established before broader use.
    Value measurementActivity is reported: licences, pilots and prompts.Outcomes are measured: cycle time, quality, cost, customer experience and risk.
    TrustEarly mistakes become evidence against the programme.Small, well-governed wins create permission for the next change.

    The difference is not caution versus ambition. It is operational discipline versus performative speed.

    Established businesses are complex

    Startups can often change rapidly because the organisation is smaller, its operating model is still forming and its technology carries little legacy integration. A founder can decide in the morning, change a product workflow by afternoon and observe the impact within days.

    Established businesses face a different reality. They may serve millions of customers, operate under regulatory obligations, rely on complex supplier relationships and carry years of interconnected systems and policies. Their scale is an advantage, but it also means that a change in one function creates consequences in many others. A new AI-assisted credit decision, for example, has implications for compliance, model governance, customer communication, front-line procedures, auditability and dispute resolution. The business must move deliberately because the cost of a poorly absorbed change is multiplied by its reach.

    This a design constraint to work within, not a weakness to apologise for. Large organisations should build a repeatable mechanism for safe acceleration. The goal is to increase the rate at which the business can absorb change, not to force a rate that breaks it.

    Build absorption capacity before demanding velocity

    The practical response is to treat AI adoption as a managed portfolio of changes rather than a single enterprise-wide instruction. Start with a workflow that is sufficiently valuable to matter, sufficiently bounded to govern and sufficiently measurable to learn from. Assign a business owner, a process owner and a risk or control partner. Define what a good outcome looks like before deployment, including the conditions under which the team will pause or reverse the change.

    A sensible sequence has four stages.

    StageLeadership questionEvidence required before moving on
    ProveDoes this solve a real operational or customer problem?A defined baseline and a demonstrable improvement in a controlled workflow.
    StabiliseCan people use it reliably under normal conditions?Clear ownership, training, controls and a functioning exception process.
    ReplicateDoes the operating pattern transfer to similar work?Consistent results across more than one team or business unit.
    ScaleCan the enterprise support this without degrading quality or trust?Resourcing, governance and technology capacity match the expanded scope.

    First, prove usefulness in a narrow setting where people can compare the AI-assisted outcome with the existing process. Then stabilise the new way of working by documenting roles, controls, exceptions and training. Then replicate the pattern in adjacent use cases that share similar data, risks or operating habits. Finally, scale the platform and governance only after the organisation has evidence that the operating model works repeatedly.

    This approach may appear slower at the beginning because it resists the temptation to announce universal adoption. In practice, it is usually faster over time. Each successful use case leaves behind reusable assets such as trusted data pathways, trained leaders, approved controls, implementation playbooks, measurement methods and internal advocates. The next deployment begins from a stronger base, drawing on accumulated organisational knowledge rather than repeating the same arguments from scratch.

    Leaders must protect the rate of learning, not the rate of adoption

    The leadership task is to sponsor AI usage and to protect the business’s learning rate. That means creating room for teams to identify flaws without being labelled resistant, refusing to measure progress solely by adoption volume and being explicit about what will not yet be automated. It also means resisting the false choice between reckless acceleration and organisational paralysis.

    A mature AI programme should be demanding, but its demands should be specific: improve a process, raise decision quality, reduce a known friction point, protect customers, and demonstrate results. “Use AI everywhere” is an instruction to create unmanaged variation at scale.

    Wilke’s observation to Bezos offers a better standard. Ideas have value only when an organisation can turn them into coherent action. Every AI deployment released beyond the organisation’s capacity to absorb it creates a backlog of unfinished transformation.

    Organisations that make change smooth – clear enough for people to adopt, controlled enough to trust and valuable enough to sustain – will accumulate capability faster than those that simply move first. Slow is smooth, smooth is fast.

  • What to Do with The Feelings When the Machine Can Do Your Job

    Something real is being taken from you, and it is not just your workload.

    Ask most people to introduce themselves, and they describe what they do. “I’m an architect” “I’m a nurse” “I’m a lawyer”. That is not modesty or habit but a statement of identity. The role names the person. It signals competence, years of effort, and a place in the world. When AI starts performing the core tasks of that role with speed and reasonable accuracy, the threat is not just professional. Losing a task is inconvenient, but losing the thing you built yourself around is something else entirely.

    Your Role Has Shifted, and That Shift Hurts More Than You Expected

    The primary change AI seems to bring to knowledge work is not elimination but demotion from creator to checker.

    When an AI system can draft a legal argument, generate a diagnosis, or produce a structural engineering report, the human professional moves from producing the work to reviewing it. The cognitive load shifts from high-level synthesis to what researchers call “AI managerial labour”: monitoring outputs, catching errors, and signing off on conclusions you did not reach yourself.

    What ChangedBefore AIAfter AI
    Where authority comes fromYour training and judgementThe algorithm’s prediction
    What you spend your time onSolving and creatingOverseeing and refining
    Your role in the decisionThe expert who decidesThe human buffer who approves

    This shift is widely documented and spreading across industries faster than most organisations are acknowledging. Software developers have already lived through it. A few years ago, AI-generated code was treated as a curiosity, too unreliable to trust. Now most developers use it routinely, and the ones who don’t are increasingly at a disadvantage. The code is not perfect, but it does not need to be. It just needs to be good enough to change the expectation of what a developer does, and that pattern is moving through law, medicine, finance, consulting, and design.

    Why This Feels Like a Betrayal

    The fear of this shift is rational. Researchers have identified a specific phenomenon as AI guilt, or AI shame. Professionals who built their identity around a craft feel genuine discomfort when they hand that craft to a machine. Using AI to do what you once did yourself can feel like a violation of your professional standards, or, as one research team put it, like “conspiring with the enemy” to make yourself obsolete. When your expertise is the thing you are proudest of, and a tool makes that expertise feel commoditised, it does not produce neutral feelings.

    The problem runs deeper when the AI cannot explain itself. Many systems produce outputs without revealing how they reached them. When you must approve a decision you cannot trace back through logic, when you are accountable for an outcome you did not reason your way to, your sense of professional autonomy erodes. You are responsible without being in control. That is a specific and corrosive kind of discomfort.

    The emotional toll is not evenly distributed. Those who have spent the longest building their expertise tend to feel this most acutely. A career built over decades on a particular kind of cognitive authority faces a different kind of threat than a career just starting out. There is more time and effort being attached to the individuals identity with less time to rebuild.

    The Way Through Is Changing How You See Yourself, Not What You Do

    Research and practical experience suggest that professionals who navigate this best are the ones who change how they think about themselves, not what they do.

    One concrete shift is in how you describe yourself. “I am a lawyer” places your identity in a function that AI is now performing. “I work as a lawyer” places it in a behaviour which is something you consciously practise, adapt, and develop. The first statement sets a boundary, but the second sets a direction. It sounds like a small distinction, but in a period of rapid change it is a significant one, because capabilities and behaviours can be learned while identities tend to be defended.

    Beyond language, the professionals who contiue to improve tend to do it by deliberately focusing on the elements of their work that remain genuinely human. Not in a defensive way, as if protecting territory, but in a generative way, actively building the part of their practice that AI cannot replicate. That means different things in different fields, but the pattern is consistent.

    The things only humans bring. Empathy, ethical reasoning, and the ability to read an unspoken concern in a client’s face are not tasks but capacities. AI can simulate some of them, but it cannot hold responsibility for the outcomes, and clients know the difference.

    The human relationship. Who decides, in the end? Who is accountable when something goes wrong? In almost every profession, the answer remains a person. Redefining your collaborations so that human judgement is clearly the final arbiter is not just good practice but also where your irreplaceability lives.

    The shift from doer to strategist. The professionals adapting best are not the ones doing less, but the ones doing differently. They use AI to accelerate the execution of ideas they are directing, rather than treating AI as a replacement for the thinking itself.

    You Still Own What the Machine Cannot Take

    AI can replicate the execution of a task, but it cannot replicate the responsibility for the consequences.

    When a diagnosis is wrong, a human clinician is accountable. If a structural report misses something and a building fails, a human engineer bears that weight. When a contract fails its client, a human lawyer answers for it. This is a legal fact and a fundamental distinction between a tool and a professional.

    The future of expertise is not in competing with AI on speed or scale, as that’s a losing proposition. It is in combining the fundamentals of your business, ethical responsibility, tacit knowledge, and relational intelligence that human professionals carry with the capability AI provides. The professionals who hold that combination clearly in mind, and invest in the human side of it deliberately, are the ones who will survive this transition and be more valuable because of it.

    The discomfort you are feeling right now is a signal worth listening to. While it might feel as if your identity is being challenged, you are much more than your role. That signal and challenge is a call to action and growth, if you act on it.


    This article draws on research by Sartirana & Salvatore (2025), Law & Varanasi (2025), Ziegelmayer & James (2024), Mirbabaie et al. (2021), among others.

  • Why AI Automation Fails Without Knowledge Engineering

    Most AI automation projects for business processes start the same way. A vendor demonstrates a compelling proof of concept and the demo works. The board approves the investment but six months later, the automated process is producing results that no one trusts, and the organisation is quietly routing work around it.

    The failure wasn’t in the AI. It was in what the implementation team did before they deployed the AI.

    Quick automation captures what systems do while knowledge engineering captures why they do it. These are not the same thing, and conflating them is the most common and costly mistake in enterprise AI implementation today.

    The Gap That Quick Automation Cannot See

    Process mining and rule extraction are useful tools. They read event logs from your ERP, CRM, and core systems, then reconstruct how your processes actually flow. They surface bottlenecks, flag deviations, and give your implementation team a working map of operations.

    For simple, well-documented processes, this is enough but for anything more complex, it is a structural failure. Event logs record what happened in the system, not the reasoning behind it. They cannot see the workarounds your experienced people have built over fifteen years. They cannot capture the judgement your senior underwriter applies when a claim is technically within policy but smells like fraud. They cannot document the exception a finance team made for a long-standing client that never made it into the system because everyone just knew.

    Research published in 2025 confirmed that event logs systematically underrepresent operational reality, with an estimated 70% of knowledge work occurring outside the systems that generate those logs. What gets automated by the AI is the documented surface of a process, not its full depth.

    When the AI encounters the undocumented exception, it either fails, escalates, or produces a wrong answer with complete confidence. All three outcomes erode trust, and once trust is gone, the automation sits idle while your people route around it.

    What Knowledge Engineering Actually Does

    Knowledge engineering is the discipline of making implicit knowledge explicit and verifiable before your team encodes it into an AI system.

    This work involves structured interviews with domain experts. It involves observing how work actually gets done, not how the procedure manual says it should be done. It involves encoding the results into knowledge bases that subject matter experts can review, challenge, and correct. It involves validating that the encoded logic produces the right outcomes before anything touches a production environment.

    This is the pattern that distinguishes durable AI automation from the kind that works in demos and fails in production. The AI surfaces candidate rules and logic and domain experts validate or correct them. The validated knowledge becomes the foundation your team builds automation on.

    This model is explicit about something that many vendors prefer not to emphasise: AI will confidently surface intent that sounds right and is not. Business-logic reconstruction can leverage AI but remains human-led by design. The AI drafts the interpretation while the expert confirms or corrects it.

    Where Systems Thinking Changes the Problem

    A systems thinking approach to AI automation asks a different question than most implementation teams ask. Instead of “what processes can we automate?”, it asks “what does the system as a whole need to keep working correctly after automation is introduced?”

    This reframing matters because automation does not operate in isolation. When you automate a claims processing workflow, you change the load on the humans who handle escalations. You change what data your compliance team can access and when. You change the feedback mechanisms that let your experienced people spot when something is going wrong. You may also change the incentives in ways you did not intend, particularly if the organisation measures the automated process on speed but measured the people it replaced on accuracy.

    Systems thinking requires mapping those interdependencies before implementation, not discovering them six months after go-live.

    Three principles from this approach are directly applicable to AI automation decisions.

    The first is that problems appear long before they are noticed. An AI model that is slowly drifting from its training distribution will produce subtly degrading outputs for weeks before anyone raises a flag. IBM’s 2025 Cost of a Data Breach Report found that 13% of organisations reported breaches involving AI models or applications, and 97% of those breached organisations lacked proper AI access controls. The failures were systemic, not sudden.

    The second is that you get what you incentivise. If your AI implementation is measured on the percentage of cases processed automatically, you will get high automation rates but you may also get poor outcomes on the cases that needed human review but were incorrectly classified as routine. Design the measurement framework before designing the automation.

    The third is that you fall to your level of preparation. When something goes wrong with an automated process at scale, your team’s ability to respond depends entirely on what they practised before it happened. Oversight protocols, escalation paths, and retraining procedures need to exist and be tested well before an incident requires them.

    The Real-Time Oversight Problem

    Traditional governance was built for a different tempo: annual audits, quarterly reviews, weekly metrics. These cadences made sense when processes ran at human speed and human scale.

    AI automation changes both variables simultaneously. A model processing thousands of claims per day can accumulate hundreds of errors before a weekly metrics review would detect the pattern. By the time a quarterly audit surfaces a systematic bias in lending decisions, the exposure may already be material.

    This is not a hypothetical concern. The EU AI Act’s Article 96 now requires organisations to demonstrate compliance through continuously updated, machine-readable evidence, not point-in-time assessments.

    The implication for senior leaders is direct. Governance frameworks designed for human-speed processes need to be redesigned before your organisation deploys AI at scale, not retrofitted after a problem surfaces. Real-time monitoring of model outputs, data quality, and prediction confidence is not optional infrastructure. It is the mechanism by which you maintain accountability for decisions your organisation is making at machine speed.

    What a Systemic Approach Actually Looks Like

    Implementing AI automation well follows a recognisable pattern.

    Begin with knowledge extraction, not process mapping. Before any automation is designed, invest in capturing the institutional knowledge that lives in people, not systems. This means structured elicitation sessions with domain experts, protocol analysis of how difficult cases are actually handled, and documentation of the exceptions and judgement calls that define quality outcomes.

    Validate before you deploy. The people who encoded the knowledge review it. Pilot programmes run in environments where your team checks outputs against known-good outcomes before the model is trusted with live decisions.

    Design oversight into the process architecture. Human-in-the-loop checkpoints are not emergency interventions but are designed features, placed where the model’s confidence is lowest or where the cost of an error is highest.

    Instrument for continuous monitoring. Model performance, data quality, and output confidence scores are tracked continuously. Thresholds trigger review before drift becomes visible to customers or regulators.

    Govern at the speed of your systems. Review cadences match the operational tempo of the automated process, not the legacy tempo of the governance process it replaced.

    The Question to Ask Your Implementation Team

    If your organisation is evaluating or currently implementing AI business process automation, one question will reveal more about the quality of the approach than any other: what knowledge engineering work did the team complete before the AI model was trained?

    If the answer is “we used automated process mining and rule extraction from our existing systems”, that is a starting point, not a complete answer. It describes what the system does, it does not describe what your experienced people know that the system does not record.

    If the answer is “we conducted structured knowledge capture with domain experts and validated the results before training”, you are working with a team that understands where AI automation actually fails.

    The technology for automating business processes is mature, accessible, and increasingly affordable. The discipline required to automate them well, preserving institutional knowledge, designing real-time oversight, and aligning governance to the speed of the systems being governed, remains genuinely rare.

    That gap is where durable competitive advantage is built.

  • Your AI Interface Is a Cognitive Design, Not a Cosmetic One

    Most organisations deploying AI spend months selecting the right model, agonising over accuracy rates, vendor contracts, and compliance implications. Then, in the final weeks before launch, someone asks: “What should the screen look like?”

    The interface is not decoration applied after the real decisions are made, instead it determines whether your people reason clearly alongside the AI or quietly work around it. Get it right, and your AI system makes better decisions than either the human or the system could alone. Get it wrong, and you have an expensive system that your team has learned to distrust, override, or ignore.

    This article gives you a structured method to audit any AI interface before deployment. Based on a field called Cognitive Systems Engineering (CSE), developed in the early 1980s by Erik Hollnagel and David Woods to address exactly the kind of high-stakes, human-machine decision environments that organisations across every industry are now building. The method is built around a five-level analysis called the Abstraction Hierarchy. You will walk through each level (we’ll use a loan decision AI as the working example), leaving with a set of questions you can apply to your own deployment.

    The Shift That Changes Everything

    Traditional software is deterministic. You press a button, you get a result. Designing the interface for that kind of system still requires some skill, but is largely a matter of clarity and efficiency because we have many examples of good (and bad) design.

    AI is different as it produces uncertain outputs and reasons probabilistically. It can be right most of the time and catastrophically wrong in ways that are difficult to anticipate. The interface for a traditional system needs to be usable. The interface for an AI system needs to do something different: it needs to make the AI’s reasoning visible so that the human can judge when to act on it, when to question it, and when to override it.

    Hollnagel and Woods called this a “joint cognitive system“: the human and the AI are not separate entities where one hands off to the other. They are a single thinking unit, and the interface is the connective tissue between them. Whether the decision involves a loan approval, a clinical recommendation, a fraud alert, or an operational plan, that connective tissue either holds or it tears.

    The Abstraction Hierarchy, developed by Jens Rasmussen and Kim Vicente at the Risø National Laboratory in Denmark, gives you a disciplined way to design that connective tissue. It asks you to understand your work domain at five levels, from purpose down to physical configuration. Each level reveals a different set of interface requirements that you might otherwise miss.

    The Abstraction Hierarchy: Five Questions for Any AI Interface

    Think of the Abstraction Hierarchy less as a taxonomy and more as a diagnostic interview. At each level, you ask a question about your work domain. The answers tell you what your interface must show, what it must prevent, and what it must make possible.

    Here is the framework applied to a loan decision AI to help understand the details.

    Level 1: What is this system ultimately for?

    This is the question of functional purpose: the values and goals the system exists to serve. Not the technical goals (“classify loan applications”), but the human and organisational values at stake.

    For a loan decision AI, the answer is not simply “approve or decline applications faster”. The real purposes are: extend credit to people who can repay it, protect the institution from credit losses, comply with responsible lending obligations, and treat applicants fairly across demographic groups. These purposes can conflict. A model optimised for speed may sacrifice fairness. A model optimised for loss minimisation may discriminate. These tensions exist whether you name them or not and exposing them allows informed decisions.

    What this means for your interface: The interface must make the system’s governing priorities visible. If the model has been calibrated to weight certain risk factors above others, the loan officer reviewing its recommendation should be able to see that calibration, not just the output. When the AI recommends declining an application, the interface should surface which of the system’s core purposes drove that recommendation: is this a credit risk concern, a compliance flag, or something the model cannot categorise cleanly?

    Questions to ask about your deployment:

    • What are the two or three values this system is genuinely optimised for?
    • What values are in tension, and does the interface make those tensions visible?
    • When the AI’s recommendation conflicts with a user’s instinct, does the interface help the user understand why?

    Level 2: What principles govern how the system operates?

    This is the question of abstract function: the rules, regulations, and governing principles that constrain what the system can and cannot do. In aviation, these are physics and safety regulations. In lending, they are responsible lending laws, anti-discrimination regulations, internal credit policy, and audit requirements.

    These constraints do not change with each transaction. They define the boundaries within which all decisions must fall. The problem is that most AI interfaces present outputs as though these constraints do not exist. The model returns a score. The interface shows the output metric. The loan officer is left to remember, from training, which regulatory constraints apply in this situation.

    What this means for your interface: Regulatory and policy constraints should be structurally present in the interface, not stored in the user’s head. If the AI’s recommendation would require a manual review under responsible lending obligations, the interface should flag that requirement automatically. If the model’s confidence falls below a threshold that your compliance team’s internal policy requires to be reviewed, show that threshold, and ideally the policy, visibly.

    This is the principle that Vicente and Rasmussen called Ecological Interface Design (EID): the constraints of the work domain should be visible in the interface itself, not stored in the user’s memory. An EID-informed loan interface does not require the loan officer to remember the regulatory rulebook. It makes the rulebook structurally visible in the decision flow.

    Questions to ask about your deployment:

    • Which regulatory obligations apply to the decisions this AI supports?
    • Are those obligations visible in the interface, or do users have to remember them independently?
    • When the AI’s recommendation sits in a regulatory grey zone, does the interface make that visible?

    Level 3: What processes does the system need to perform?

    This is the question of generalised function: the operational processes and workflows that need to happen for the system to achieve its purpose within its constraints.

    In a loan context, this includes the steps of gathering applicant data, running the credit model, checking for compliance flags, presenting a recommendation, capturing the loan officer’s decision and rationale, escalating edge cases, and creating an audit trail. These processes are not all equally visible in most AI deployments. Typically, what is visible is the output of the model. The processes that produced it, and the processes that need to follow from it, are hidden or scattered across different systems.

    What this means for your interface: The interface should reflect the full process, not just the model’s output. If the correct process requires a loan officer to review the AI’s recommendation alongside the applicant’s supporting documents before deciding, the interface should make that sequence natural and difficult to skip. If the process requires capturing the officer’s reasoning when overriding the AI, the interface should prompt for that reasoning at the point of override, not as a retrospective form filed later.

    This is also where structured input design matters. If you want the AI to produce consistent, auditable outputs, you need the interface to guide consistent, structured inputs. A free-text prompt box for a loan officer to query the AI is the wrong design. A structured form that constructs the query from validated fields, treating the prompt as a template and the user’s input as variables, produces far better consistent results and a far cleaner audit trail.

    Questions to ask about your deployment:

    • What is the full sequence of steps the process requires, including before and after the AI’s output?
    • Does the interface make the correct sequence the natural path, or can users shortcut it?
    • How does the interface capture the human’s reasoning, not just the AI’s recommendation?

    Level 4: What are the capabilities and limits of each component?

    This is the question of physical function: what each component of the system can and cannot do. For an AI, this means understanding the model’s actual capability boundaries. Where does it perform well? Where does it degrade? What kinds of inputs push it outside its training distribution?

    Loan officers who use AI tools daily develop intuitions about where the model is reliable and where it is not. New officers do not have those intuitions, and even experienced officers can be misled when the model presents its outputs with uniform visual confidence regardless of whether it is in familiar or unfamiliar territory.

    What this means for your interface: The interface must distinguish between high-confidence and low-confidence outputs, and it must do so in a way that reflects the model’s actual calibration, not a standardised disclaimer. If the AI is recommending approval on an application that combines features it has rarely seen together, that uncertainty should be visible. If the model’s confidence score is below a meaningful threshold, the interface should communicate that clearly and differently from high-confidence outputs, not with a footnote, but with a structural difference in how the recommendation is presented.

    This is what the CSE literature calls making the AI’s epistemics visible: the interface should show not just what the AI concluded, but how firmly it concluded it and on what basis.

    Questions to ask about your deployment:

    • Does the interface distinguish between high-confidence and low-confidence recommendations?
    • Can users tell when the AI is operating in territory close to the edge of its training?
    • What happens when the AI encounters an input type it was not trained on?

    Level 5: What is the actual configuration of the system?

    This is the question of physical form: the literal layout, controls, and information architecture of the interface as it exists on the screen.

    This is where most interface design effort is spent and also the level where most AI interface problems are most visible: the designer burying a recommendation at the bottom of a long screen, the confidence score presented in a font smaller than the surrounding data, or the override button placed three clicks away. These are not aesthetic problems but decision quality problems.

    If the most important signal the AI is sending is that it is uncertain about this application, and the interface makes that signal hard to find, the loan officer will miss it. That is design failure, not a training failure.

    What this means for your interface: The visual hierarchy of the interface should reflect the information hierarchy of the decision. The AI’s recommendation and its confidence level should be visually prominent. Flags and caveats should not be hidden in tooltips. The action the interface makes easiest should be the action the process intends to be most common. The action that requires more care, like an override of the AI’s recommendation, should require commensurate effort in the interface: not so much effort that it becomes a workaround, but enough that it cannot happen accidentally.

    Questions to ask about your deployment:

    • Does the visual hierarchy of the interface match the decision hierarchy?
    • Is the AI’s uncertainty as visible as the AI’s recommendation?
    • What is the path of least resistance in the interface, and is that the right path?

    What Happens When You Skip This Analysis

    The most common failure mode is not that the AI model is wrong but that the interface makes it impossible to know when the model is wrong.

    People in high-stakes roles learn quickly. If an AI interface presents confident-sounding recommendations without surfacing the model’s reasoning or uncertainty, they will test it against their own judgement for a few weeks, find cases where it was clearly wrong, and start treating all its outputs with blanket scepticism. The AI becomes a checkbox, not a collaborator. The organisation has paid for a decision-support system and deployed a bureaucratic step.

    The second failure mode is the reverse: people defer to the AI when they should not, because the interface presents its outputs with more authority than the model’s actual confidence warrants. This produces decisions that look considered but are indefensible when scrutinised: by regulators, auditors, customers, or a board asking why something went wrong.

    Both failures are interface failures. The model may be performing exactly as designed. The interface is simply not communicating what the model knows and does not know.

    Where to Start

    You do not need to redesign your entire interface before launch. You need to run through these five levels with the people who will use the system and the people responsible for the process it supports.

    Bring three groups into a room: the people who will use the AI day to day, the people who own the process it sits inside, and whoever is accountable for risk or compliance in that domain. Walk through the five levels as questions. At each level, ask: what does the interface currently show, and what does this analysis say it needs to show? The gaps between those two answers are your design priorities.

    Pay particular attention to Levels 1 and 4. Purpose misalignment (Level 1) and invisible uncertainty (Level 4) are the two most dangerous gaps, and they are the two most commonly overlooked in AI deployments focused on model performance rather than interface design.

    The Abstraction Hierarchy was originally developed to allow human operators to make complex, safety-critical decisions under pressure. Whether your AI is supporting credit decisions, clinical triage, fraud detection, or operational planning, the underlying challenge is the same: consequential, often regulated decisions where the interface either helps people reason well or quietly gets in the way.

    Your AI model does not make decisions. The human-AI system makes decisions. The interface is what makes that system work so design it accordingly.


    References

    [1] E. Hollnagel and D. D. Woods, “Cognitive systems engineering: New wine in new bottles,” International Journal of Man-Machine Studies, vol. 18, pp. 583–600, 1983.

    [2] K. J. Vicente and J. Rasmussen, “Ecological interface design: Theoretical foundations,” IEEE Transactions on Systems, Man, and Cybernetics, vol. 22, no. 4, pp. 589–606, 1992.

  • What Bunnings and a Face Scanner Taught Australian Businesses About Privacy

    Most organisations deploying AI-powered technology assume the hard part is the technology. The Bunnings case shows the hard part is the governance.

    Between November 2018 and November 2021, Bunnings installed facial recognition technology (FRT) in a number of its Australian and New Zealand stores. The system captured the face of every person who walked through the door, compared it against a database of individuals previously involved in criminal conduct or incidents involving staff, and alerted store employees when it found a match. The facial data itself was deleted in milliseconds. Nobody was told it was happening.

    That decision triggered a regulatory investigation that took six years to resolve, a determination by the Privacy Commissioner, an appeal to the Administrative Review Tribunal, and a still-open question about whether the matter ends there. The legal outcome was mixed, but the compliance lessons are not.

    The Regulator Found Bunnings Collected Biometric Data, Whether It Knew It or Not

    On 29 October 2024, Privacy Commissioner Carly Kind found that Bunnings had breached several Australian Privacy Principles (APPs) under the Privacy Act 1988 (Cth). The breaches related to transparency, collection, and notification. In the Commissioner’s view, FRT was a highly privacy-invasive tool that interfered disproportionately with the privacy of all store entrants, not just the small number of people it was designed to identify.

    Bunnings appealed and on 4 February 2026, the Administrative Review Tribunal (ART) partially set aside the determination and upheld the findings that Bunnings breached APP 1 (transparent management of personal information) and APP 5 (notification of collection). It overturned the finding on APP 3 (collection of sensitive information without consent).

    The APP 3 decision turned on a specific provision. Under section 16A of the Privacy Act, an organisation can collect sensitive information without consent when obtaining that consent would be unreasonable or impracticable, and when the collection is reasonably necessary to prevent a serious threat to the life, health, or safety of any individual or to public health or safety. The Tribunal accepted that Bunnings faced genuine and serious threats, noting the scale of retail crime, the nature of incidents involving staff and customers, and the fact that many products in a Bunnings store can be used as a weapon. On those facts, the Tribunal found FRT to be a reasonable and proportionate response.

    The Privacy Commissioner subsequently confirmed she would not appeal the Tribunal’s decision.

    The Tribunal was explicit on that this outcome is not a green light for biometric deployment, but is a fact-specific finding about a specific organisation facing specific threats, with specific controls in place. Other organisations attempting to rely on the same reasoning without equivalent circumstances will likely find the exception does not apply to them.

    The Breaches Upheld Tell the More Useful Story

    Bunnings succeeded on the consent question but did not succeed on governance, and that distinction matters more for most organisations thinking about where their own exposure sits.

    “Milliseconds” Is Not a Defence Against Collection

    Bunnings argued that because the facial matching process occurred in RAM and the data was deleted in milliseconds, it had not technically “collected” personal information. The Tribunal rejected this.

    The Tribunal found that Bunnings collected facial images captured by CCTV cameras, that those images constituted biometric information, and that biometric information is sensitive information under the Privacy Act, regardless of how briefly it was held. The legal threshold for collection is low. If your system processes personal information, even in real time, even transiently, you have collected it, and privacy obligations follow.

    This matters beyond FRT. Any organisation using automated systems that touch personal information, including AI tools processing voice, images, location, or behavioural data, should assume collection is occurring and govern accordingly.

    A Sign on the Door Is Not Enough When You Are Taking Someone’s Face

    Bunnings had posted privacy notices at store entry points. The Tribunal found this insufficient.

    The Tribunal’s reasoning was that given the sensitive nature of biometric data, the size of Bunnings as an organisation, and the resources available to it, more was required. Customers needed to know the specific type of information being collected, the technology being used, the purposes behind collection, and what happened if they chose not to provide it.

    Generic privacy policies are written to cover everything. Biometric collection requires notices written to cover that specific thing. When the information being collected is sensitive, specificity is required.

    The Absence of a Documented Risk Assessment is Treated as a Governance Failure

    This is the finding with the broadest application.

    APP 1.2 requires organisations to take reasonable steps to implement practices, procedures, and systems that ensure compliance with the APPs. The Tribunal found that Bunnings failed to meet this standard. The steps taken prior to deployment were described as “random enquiries and actions”. No formal, structured, and documented privacy risk assessment had been conducted before the FRT system went live.

    The Tribunal’s language is deliberate. When an organisation collects sensitive information, it faces a serious intrusion of privacy. Serious intrusions require formal responses, not ad hoc ones. A Privacy Impact Assessment (PIA) conducted before deployment, documented, and retained, is what reasonable steps look like in this context.

    The compliance lesson here goes beyond privacy. Regulators across multiple domains, including privacy, cybersecurity, and financial services, are increasingly focused on whether organisations can demonstrate that their governance preceded their deployment decisions, not followed them. If the PIA exists only after a complaint is made, it does not count.

    The Regulatory Environment Around This Decision Has Changed

    The Bunnings case was decided under the existing Privacy Act 1988 framework. That framework is now amended.

    The Privacy and Other Legislation Amendment Act 2024 (Cth), which received royal assent in December 2024, introduced changes that take effect progressively through 2026. Three are particularly relevant to organisations using AI or data-intensive technologies.

    The definition of personal information now explicitly covers inferred or generated data, such as the outputs of machine learning models. AI generating risk scores, predicted preferences, and behavioural classifications are now personal information for the purposes of the APPs. Organisations that assumed model outputs sat outside the Act’s reach need to revisit that assumption.

    From December 2026, organisations must disclose in their privacy policies when personal information is used in substantially automated decision-making that has significant effects on individuals. This requires organisations to audit existing systems now, identify where automated decisions are being made, what professional function they perform, and build disclosure workflows before the obligation takes effect.

    A Children’s Online Privacy Code is being developed, with registration expected by December 2026. Organisations whose services are likely to be accessed by minors face heightened obligations and need to begin preparing ahead of the code’s commencement.

    A further reform expected in the second tranche of amendments is a general “fair and reasonable” test for the collection, use, and disclosure of personal information. If introduced, this would shift the compliance question from whether a notice is given to whether the underlying practice can withstand objective scrutiny. An AI system that targets individuals based on inferred sensitive characteristics, or that generates unexpected outcomes from opaque models, may fail that test regardless of what the privacy policy says.

    What Organisations Should Do Now

    The Bunnings case illustrates a familiar pattern. The organisation deployed the technology first and worked out the governance as concerns arose. That sequence creates risk, and the regulatory direction of travel is toward holding organisations accountable for the order in which they do things.

    The following steps address the specific failure modes the Tribunal identified.

    Conduct a Privacy Impact Assessment before deployment, not after. For any system that collects sensitive information, a formal, structured, and documented PIA is what APP 1.2 compliance looks like. This assessment should identify the risks, document the controls, and record who made the decisions and why, and should exist before the system goes live.

    Write specific notices for specific technologies. If your organisation uses FRT, AI-powered monitoring, or any other tool that collects biometric, health, or other sensitive information, the privacy notice for that collection needs to name the technology, describe what is collected, state the purpose, and explain the consequences of not providing the information. Generic notices covering all data collection are not adequate for high-risk processing.

    Audit your systems for the amended definition of personal information. AI model outputs, including risk scores, classifications, and generated assessments, now sit within the Act’s scope. Organisations using AI tools that generate outputs about individuals should review how those outputs are managed, retained, disclosed, and corrected.

    Map automated decisions and prepare for disclosure. Identify where your organisation uses personal information in automated or substantially automated decision-making with significant effects. Build the workflows and policy language needed to disclose this before December 2026.

    Maintain records of AI-assisted decisions. Accountability under the APPs requires being able to demonstrate what decisions were made, on what basis, and with what human oversight. If that record does not exist, accountability cannot be demonstrated.

    The Bunnings outcome is sometimes framed as a partial win for business. In a narrow legal sense, that is accurate. On the question of governance, Bunnings did not win and was found to have deployed an invasive technology without adequate transparency, without specific notification, and without a documented risk assessment. Those findings were not overturned.

    The compliance clock started the moment the technology was deployed. For organisations considering similar decisions, it starts now.

  • You’re Not Getting AI Answers. You’re Getting Your Own Assumptions

    When you type a question into an AI system, you assume you are asking something neutral. Every sentence you write carries silent commitments, background assumptions so deeply embedded in language that you produce them without noticing. These assumptions shape what the AI returns before it processes a single fact.

    Linguists call them presuppositions, and understanding how they work is the difference between using AI as a precision tool and using it as a confirmation machine.

    Presuppositions steer AI before the question is answered

    A presupposition is an implicit assumption that must be true for a sentence to make sense. Consider a mundane example: “I need to go to the store and buy milk.” The sentence takes for granted that a store selling milk exists. That assumption is never stated. It does not need to be, because language operates on shared background conditions.

    Add one word and the sentence changes: “I need to go to the store and buy more milk.” Now the sentence presupposes you already have some milk. The word “more” does quiet work.

    This is how presuppositions function in everyday language. In AI search, they do the same thing, but the consequences are less forgiving.

    Research shows that all tested models display sensitivity to presuppositions, but instruction-tuned models (common descendants you would recognise are ChatGPT and Gemini) were particularly vulnerable to surface-level cues and prompt phrasing. Instruction tuning trains models to be “helpful and responsive to the user’s framing”, and how you write the question changes what the model treats as true.

    The presuppositions you already use

    There are three types worth recognising, because each one shows up differently in the prompts people write to AI.

    The existence assumption. This is the most common. When you refer to something as though it exists, you are presupposing it does. “What are the benefits of this approach?” presupposes there are benefits. “Who is responsible for this problem?” presupposes someone is. You probably use this pattern dozens of times a day in conversation, and it works fine there because the other person can push back. AI systems trained to be helpful are less likely to.

    Notice it in your own language when you hear phrases like “the reason why”, “the best option”, or “the impact of”. Each one smuggles in an assumed fact. There is a reason, a best option, an impact. The question is built on ground the AI did not get to inspect.

    The change assumption. Certain words signal that a shift has occurred. “Still”, “again”, “yet”, “anymore”, “stopped”, and “started” all carry the assumption that something was previously different. “Why is this approach still popular?” presupposes it has been popular for a while. “When did the company start losing money?” presupposes it is losing money. “Is he still the right person for this?” presupposes he once clearly was.

    These words feel precise and factual, which is part of why they slip past unexamined. They are not neutral. Each one commits the AI to a version of history before it evaluates any evidence.

    The framing assumption. This one is subtler. The words around your question set a context that the AI carries into its response. Ask about a topic in a critical frame and the AI weights critical evidence. Ask in an optimistic frame and it finds reasons for optimism. “Given how disruptive this technology has been, what should organisations do next?” has already decided the technology is disruptive. The question is just asking about next steps.

    The framing assumption is particularly hard to catch because it does not sit in a single word but in the setup, the tone, the choice of examples you include before asking. But it is the most powerful of the three, because it can prime the entire direction of an AI’s response before the actual question begins.

    How a single word loads a false assumption

    The practical risk sits in queries that presuppose a relationship or fact the user has not yet verified.

    Consider two ways to ask about this (intentionally weird) topic:

    1. “How are strawberries related to pine trees?”
    2. “Are strawberries related to pine trees?”

    The first query contains an existential presupposition. It assumes a relationship exists and asks only for the description. An AI seeking to satisfy the prompt may overstate minor biological similarities rather than evaluate whether the relationship is real. The second query allows the model to assess the claim directly. The same logic applies to every existence, change, and framing assumption in the previous section. If you are trying to determine whether something is true, the presupposition inside your question is working against you.

    Four practices for cleaner AI interactions

    Presupposition awareness is not about being precious with language, or that using them is incorrect. It is about noticing when your phrasing has already decided what the answer should be.

    Audit your hidden assumptions before submitting. Ask: what must be true for my question to make sense? If that assumed truth is exactly what you are trying to find out, your prompt is circular. Restructure it to raise the question openly rather than bury it as a background condition.

    Test the inverse. If you search for the benefits of a policy, also search for its costs. If your initial phrasing produces strong positive results, check whether the framing was doing the work. Results that evaporate when the question is reversed were probably produced by the question, not the evidence.

    State your assumptions explicitly. Rather than allowing the AI to carry a background assumption forward, name it. “Assuming X is true, what follows?” separates the factual check from the logical inference. You can then run the same structure with “Assuming X is false” to map the counterfactual.

    Refuse the premise when the AI acts on a false one. If an AI returns results built on an assumption you did not intend, do not just rephrase. Explicitly deny the assumption first. “There is no established relationship between X and Y. Given that, what does the evidence show?” Resetting the context is often more effective than softening the original question.

    The research confirms that providing explicit context against a presupposition significantly reduces the model’s tendency to treat it as given. The AI did not add the bias, your prompt did. That means fixing it is entirely within your control.


    Reference: Wörgötter, M. L., Lai, S., & Schuster, S. (2026). There is No Spoon: Existential Presupposition in Large Language Models. University of Vienna & University College London.

  • Modelling Genius: Encoding Human Expertise into AI

    Ask a master negotiator how they read a room and you will usually get a shrug. Ask a surgeon how they decide where to cut first, or a trader how they sense a market turn before the numbers confirm it, and the answer is the same kind of shrug. Expertise this deep runs on autopilot. The expert cannot see their own process because it stopped being conscious years ago.

    This is the real obstacle standing between human expertise and AI systems making the right decisions. It is not a data problem but a structure problem. Before an AI can inherit a decision-making pattern, someone has to find that pattern first.

    This article looks at what behavioural modelling is, the specific skills it demands, and how those skills can be used to produce frameworks such as the Galdren Lead Gen Decision and Judgement Framework for building AI systems that reason the way an expert reasons, not only the way a manual describes.

    What Expert Behavioural Modelling Means

    Practitioners use the word “modelling” to describe at least three different things, and the difference matters for what an AI system ends up inheriting.

    One approach requires the modeller to absorb an expert’s patterning unconsciously first, and only analyse it afterwards. The modeller suspends any conscious, analytic attempt to understand the expert’s patterning during that initial assimilation stage, and only claims success once they can reproduce the pattern in new situations and draw out responses of comparable quality and speed to the original expert.

    A second approach skips the immersive step and instead applies existing categories and labels to describe what an expert does. This can mean structured strategy elicitation that genuinely surfaces the expert’s underlying structure, or it can mean something far shallower: someone asks the expert a few questions, writes up a summary, and calls it a knowledge base. People use “modelling” for both. The results are not comparable. One reaches the structure underneath the expert’s behaviour. The other copies what is visible on the surface and stops there.

    The third approach, now the most common one in AI projects, does not involve the expert at all. It works from what the expert has already produced: emails, documents, and speeches, parsed into a retrieval system the AI searches when it needs an answer. The AI is left to infer the underlying decisions, choices, and behaviour from the artefacts alone.

    Most AI knowledge capture today uses this third approach, because it is faster and easier for everyone involved. It is also the approach least likely to produce a usable model. Feeding an AI what an expert wrote is not the same as feeding it how the expert decided what to write. The result is a system that can echo the expert’s language without inheriting the expert’s judgement.

    True behavioural modelling sets a higher bar than any of these three. It insists someone tests the extracted pattern against the real world before anyone calls the job complete.

    The TOTE: A Map of How a Decision Runs

    The core unit of any modelling is the TOTE (Test, Operate, Test, Exit), a concept the psychologists George Miller, Eugene Galanter, and Karl Pribram introduced in their 1960 book on cognitive psychology, drawing on ideas from cybernetics. A TOTE breaks a piece of behaviour into a trigger (test), an action taken in response (operate), a check on whether the action worked (test again), and the point at which the process stops (exit). If the outcome is not reached, the person loops back through testing and operating until it has. The model can describe something as small as hammering a nail, or be nested to build a map of far more complex behaviour.

    This is the piece most AI encoding projects skip. They capture what the expert decided. They do not capture the loop: what triggered attention in the first place, what the expert checked before acting, what told them the first attempt had not worked, and what specific signal told them to stop. Encode the decision without the loop and the AI inherits an answer, not a method. It tends to fail the moment conditions shift even slightly from the example it saw.

    Strategy Elicitation: The Skill of Finding a Pattern the Expert Cannot See

    Extracting a TOTE from an expert requires strategy elicitation, and it comes in two forms. Informal elicitation happens in conversation. People tend to run through their internal strategy as they describe a past experience, and simply asking “how did you know?” is often enough to draw the sequence out naturally. Formal elicitation is more deliberate. It works through a set sequence, because experts routinely leave out the step that matters most, not because they are hiding it, but because it happens too fast for them to notice.

    From Elicited Pattern to Working System

    Once the modeller has elicited a strategy, it needs a structure an AI can act on. Prompt engineering and knowledge graphs earn their place here. They translate the TOTE sequence and its trigger conditions into something the AI checks at each step, rather than a paragraph it reads once.

    The Galdren Lead Gen Decision and Judgement Framework is one working example of this translation in practice. It runs three parallel layers: an Execution Layer where the AI acts as generator and architect, a Judgement Layer where it acts as coach, challenger, and critic, and an AI Behaviour Layer where it acts as recorder and analyst, shaping thinking through the interaction itself. Underneath sits a seven-step decision sequence: frame the problem, gather the evidence, list the options, weigh the trade-offs, make the decision, state a confidence level from 0 to 100 percent, and apply a reversal test asking what would change the decision-maker’s mind.

    This sequence is a TOTE with the trigger and exit conditions made explicit. Frame and Evidence form the test that opens the loop. Options, Trade-offs, and Decision form the operate stage. Confidence and the Reversal Test form the second test, the check on whether the decision holds up, before the process exits. Structuring it this way is what lets the framework’s judgement training loop, Ask, Question, Push Back, Educate, Support, Iterate, work as a genuine feedback cycle rather than a checklist someone bolts onto the AI’s output after the fact.

    Validation: Proving the Model Works

    Any model is unproven until it produces the expert’s results, as confirmed by the expert. The modeller only claims success once they can reproduce the original patterning in new situations and get responses of comparable quality and timing to what the original expert would have produced. That is a demanding standard, and it should be. Extracted expertise that only works inside the interview room is not extracted expertise. It is a transcript.

    For AI systems, this means human-in-the-loop review cannot stop at “does the output look reasonable”. It has to ask whether the AI, given a new situation the expert never discussed, reaches the judgement the expert would have reached. This is where the Lead Gen Decision and Judgement framework’s emphasis on validation and correction earns its keep. Experts review outputs against unfamiliar cases, flag the gap between what the AI did and what they would have done, and the difference points straight back to which part of the TOTE the modeller captured wrongly or left out.

    The Difference Between a Coder and a Modeller

    Anyone can prompt an AI with instructions that someone already wrote down somewhere. That is transcription, and it is useful, but it is not modelling.

    A modeller does something else entirely. They sit with an expert who cannot explain their own genius, use structured elicitation to surface the trigger, the check, and the exit point the expert never noticed themselves, sort those pieces by the level they operate at, and only then hand the result to an AI system for testing against reality. A coder formats information that already exists in explicit form. A modeller extracts a pattern that has never been explicit, and proves it works before calling the job finished.

    Three Problems Behavioural Modelling Has Not Solved

    Several problems remain unresolved.

    Depth of elicitation. Formal strategy elicitation takes time and skill most AI projects do not budget for. Rushing it tends to produce a shallow model dressed up as a deep one.

    Drift. An expert’s TOTE for a given decision changes as their environment changes. A model captured once and never revisited tends to fall behind the person it was built from.

    Proof of transfer. The real test of encoded expertise is whether it performs on cases the expert never saw, not whether it matches the training examples. Too few AI projects test for this directly.

    The Discipline That Data Cannot Replace

    AI systems thinking delegation will not advance much further on data volume alone. What moves it forward is the discipline of behavioural modelling: find the trigger the expert cannot name, map the loop they run without noticing, sort what you find by the level it belongs to, and refuse to call the job done until the pattern produces the expert’s results in someone else’s hands. Frameworks such as Galdren LGD show what this looks like once someone builds it into a working system. Applied properly, behavioural modelling gives AI encoding projects something most currently lack: a defined way to know whether the expertise was ever captured at all.

  • The AI Parrot Problem: Why Your AI Sounds Smart and Isn’t

    A RAG system trained on an expert’s own documents can answer a question in the expert’s own words, sound completely authoritative, and still reach a conclusion the expert would never make. That gap, between sounding right and being right, is the real risk in most AI knowledge systems today. It should worry you more than a system that gets things obviously wrong. An obvious mistake at least tells you something is broken.

    The AI Parrot Problem: Why Your AI Sounds Smart and Isn't

    What RAG Actually Does

    Retrieval-Augmented Generation does two things, and neither one is judgement. First, it retrieves: given a question, it searches a knowledge base and pulls out passages that seem relevant, in much the same way an advanced search engine would. Second, it generates: it turns those passages into a fluent, coherent answer.

    That is the whole job. The system finds information and writes it up well. It does not weigh conflicting evidence, apply experience to an unfamiliar case, or reason about what an expert would decide. Retrieval and judgement are not two points on the same scale. They are different operations entirely, and a system built for one does not automatically gain the other.

    Why It Demos Well and Fails Quietly

    RAG systems look impressive in demonstrations, and for good reason. On questions close to their source material, they act as an efficient lookup tool: find the right passage, phrase it well, done. This is where most people form their impression of what the system can do.

    The trouble starts on the cases the documents never covered, the ones where an expert would normally lean on experience and judgement rather than a reference page. This is not a random weak spot but a structural failing. The system has no mechanism for reasoning through a situation it has not seen, only for retrieving and rephrasing what it has. The output can still sound confident and plausible even when it is wrong, which means the failure stays invisible until someone acts on it.

    Air Canada’s website chatbot is a well-documented instance of this pattern. In November 2022, a customer asked the chatbot about bereavement fares after the death of his grandmother. The chatbot told him he could apply for the discount after booking, within 90 days of ticket issue. He booked his flight on that basis and later submitted a claim. Air Canada refused it. The airline’s actual bereavement policy, published elsewhere on its own website, did not allow retroactive claims. The customer took the case to the BC Civil Resolution Tribunal, which found Air Canada liable for the chatbot’s inaccurate advice and rejected the airline’s argument that the chatbot was a separate entity not covered by its duty to customers. The tribunal ordered Air Canada to cover the fare difference and costs, a little over 800 Canadian dollars in total.

    Air Canada never disclosed exactly how the chatbot worked, so this was not confirmed as a RAG system specifically. What matters here is the pattern. An AI tool produced a fluent, specific, confident answer, using the airline’s own bereavement-fare language, that directly contradicted the airline’s own documented policy. That is the Parrot Problem, playing out with real financial and legal consequences.

    The Cost, at Two Levels

    The Parrot Problem is not only an inconvenience. It costs something at the individual level, and something different at the organisational level.

    For the individual, the cost is trust placed in the wrong direction. Someone asks a system a genuine judgement question, a recommendation, a diagnosis, advice on a decision that matters, and the system answers with the fluency and vocabulary of an expert. They act on it. Later they find out the “expert” behind the answer never reasoned through their specific case. Depending on the domain, that can mean a wasted booking, a financial loss, or worse.

    For the organisation, the cost compounds quietly. If nobody checks AI output against what an expert would decide, and only checks whether it sounds reasonable, the gap between the system’s answers and correct practice widens with every case it handles. A support AI that confidently recommends the wrong fix for an undocumented configuration is likely to fail the same way each time that configuration comes up, and nobody notices until the pattern of complaints does.

    Three Questions to Diagnose Your Own System

    You do not need a research team to find out whether your AI knowledge base has this problem. Ask these three questions.

    1. Has anyone tested it on a case the source documents never covered? If every test question can be answered directly from the training material, you have tested retrieval, not judgement.
    2. When it gets something wrong, can anyone say which piece of reasoning it skipped, or does it only look “off”? If nobody can point to the specific step the system missed, that is a sign it was never reasoning through the problem. It was matching patterns and hoping.
    3. Would the original expert sign off on this answer, or only recognise the words? An expert can recognise their own language in an AI’s response and still disagree completely with the conclusion. Recognising the vocabulary is not the same as agreeing with the judgement.

    It’s a Structure Problem, Not a Data Problem

    More documents will not fix this. A bigger knowledge base makes a RAG system a better parrot, not a better judge. The real work is capturing how an expert decides, not only what they have already written down, and that is a different kind of project entirely.

    That raises the obvious next question. If retrieval isn’t enough, what does it take to build a system that reasons the way an expert reasons? That is where we pick up next.

  • Beyond the Assembly Line: Redesigning Knowledge Work

    Why the current approach to AI adoption is repeating the costly mistakes of the offshoring era, and what organisations can do differently.

    The Pattern We Have Seen Before

    Artificial intelligence is entering the enterprise the same way offshoring did twenty years ago. Both promised the same thing: lower costs and the freedom for onshore teams to focus on “high-value strategy”. Both are driven by an industrial-era assembly line mindset, one that treats knowledge work as a series of discrete tasks to be optimised rather than a connected system to be understood.

    This mindset is the belief that cognitive labour can be broken into interchangeable parts, the same way a car is built from interchangeable components. The flaw is that knowledge work carries tacit, contextual knowledge that cannot be stripped out without losing what makes the work valuable in the first place.

    Offshoring proved this the hard way. Senior managers spent half their week managing vendors, fixing broken handoffs, and rewriting deliverables that missed the context only a tenured employee would have understood. Today, the same pattern is repeating in digital form. Managers and developers are drowning in AI-generated output that takes longer to check and correct than it would have taken to produce from scratch.

    This article sets out why that is happening, what it costs organisations long-term, and three strategic shifts that break the cycle.

    Outsourcing’s Hidden Tax, and AI’s Version of It

    FeatureIndustrial Assembly LineKnowledge Work (Outsourcing or AI)
    Primary unitPhysical componentCognitive task
    LogicModular and standardisedContextual and tacit
    GoalLower cost per unitFaster output generation
    Failure modeMechanical breakdownContextual “slop”

    The table above captures the core problem. An assembly line works because every component is interchangeable and every step is independent of context. Knowledge work does not follow that logic. A report, a piece of code, or a client strategy only has value once it reflects the specific history, relationships, and politics of the organisation that needs it.

    That is the context AI does not have unless specifically designed for. Large language models lack what we might call “home office” knowledge, namely the unwritten skill of the individual and understanding of a company’s history and its culture. Without it, AI produces generic solutions to specific problems. The output looks complete but is often lacking.

    The Rise of AI Slop and the Auditing Tax

    We are entering the era of AI slop: content, code, and reports that look flawless on the surface but are hollow underneath. If AI cannot draw on the specific context of a business, it fills the gaps with plausible generalities.

    Outsourcing was meant to free teams for strategic work. Instead, it shifted effort from production to auditing. AI is creating the same shift. Organisations are spending more time checking the machine’s work than they would have spent doing the work themselves.

    This is the auditing tax, and it explains why AI adoption so often fails to show up in the numbers. According to MIT Media Lab’s 2025 study, “The GenAI Divide: State of AI in Business 2025”, 95 per cent of organisations have yet to see a measurable return on their generative AI investment, despite tens of billions of dollars in enterprise spending. The researchers found the gap was driven by implementation, not by model quality. Most deployments cannot retain context or learn from correction, so every interaction starts from zero. The auditing tax consumes the time AI was meant to save.

    The Broken Talent Pipeline

    There is a second cost that takes longer to show up: the erosion of the talent pipeline.

    When entry-level tasks moved offshore, the home office lost its training ground. Junior employees no longer did the “grunt work” that once built the foundation for senior expertise. AI threatens to repeat this at a faster pace. Summarising a report, writing a first draft of code, and conducting initial research used to be where junior staff built the instincts that, over a decade, turned into senior judgement.

    When AI takes over that work, organisations are removing time from the calendar and removing the training ground itself. The struggle of synthesising a report or debugging a simple script is exactly where the mental models of a future expert are honed. Without it, the next generation of knowledge workers will lack the intuition needed to do the very auditing and strategic oversight an AI-heavy workplace demands. Left unaddressed, this creates a leadership vacuum for the next decade.

    Three Pillars for Redesigning Knowledge Work

    Breaking the cycle requires more than better prompts or faster tools. It requires a different architecture for how knowledge work gets done.

    1. Build the Coordination Layer Beneath the AI

    The real bottleneck in most organisations is not intelligence but coordination. Employees spend a significant share of their week acting as human connectors for computers: copying a Slack message into a Jira ticket, then summarising it again for a Notion page.

    A coordination layer automates these handoffs and keeps context flowing between tools. Think of it as the electricity grid of knowledge work, the invisible infrastructure that lets the silos talk to each other. Without it, AI stays organisationally blind, forced to start every conversation from zero. With it, AI can draw on the same tacit knowledge as a tenured employee, becoming an operator that understands the flow of work rather than a conversationalist that only understands the task in front of it.

    2. Replace Task Speed With Outcome Velocity

    Counting prompts sent or emails generated is an industrial-era metric dressed up in AI language. The metric that matters is outcome velocity: how fast an organisation moves from identifying a problem to delivering a validated solution.

    • Task speed: “We generated 100 reports today.” This measures activity.
    • Outcome velocity: “We spotted a market shift and adjusted strategy within 48 hours.” This measures results.

    An organisation can increase task speed and still slow down, because every fast output adds to the queue of work that someone else has to audit, follow up, or fix.

    3. Treat AI as Cognitive Offloading, Not Cognitive Replacement

    Cognitive replacement removes the human from the loop to cut costs. Cognitive offloading uses AI to handle the mental drudgery, such as data synthesis, formatting, and first drafts, while keeping human judgement at the centre of the work.

    This distinction determines whether AI strengthens an individuals or organisation’s expertise or quietly hollows it out. Used as a lever for human judgement, AI increases what good people can do. Used as a replacement for judgement, it produces faster slop.

    Reclaiming the Knowledge in Knowledge Work

    The future of work is not an assembly line of bots producing slop at scale. It is a coordinated system where AI handles the logistics of information, freeing people for deep thought, contextual judgement, and genuine innovation.

    Organisations that build the coordination layer and measure outcome velocity instead of task speed can finally deliver on what offshoring and early AI adoption both failed to provide: technology that makes work better, not only faster.