Category: Agents

  • Is Your AI a Tool, a Colleague, or an Authority?

    Something happens the moment your organisation settles on a word for what AI is. A design philosophy clicks into place and governance questions answer themselves. Oversight feels obviously necessary, or it feels obviously unnecessary. Most of the time, no one in the room notices this happening.

    The words are not a label but a mental frame that imports an entire domain of human experience, complete with its own logic about agency, control, and responsibility. The linguist George Lakoff calls this a conceptual metaphor: we don’t merely describe experience through metaphor, we think through it. The frame structures what feels like common sense, and what feels like common sense rarely gets examined.

    When your organisation calls AI a tool, you are not describing a deployment approach alone. You are importing the complete conceptual logic of tool use, and that logic will quietly answer a hundred design and governance questions before anyone thinks to ask them explicitly.

    Tool: The Human Is the Only Agent in the Room

    A tool has no agency. A hammer doesn’t decide where it lands, a calculator doesn’t choose what to calculate. When a hammer hits the wrong nail, no one convenes a review of the hammer. The logic flows in one direction: the user holds all intention and bears all responsibility.

    Under the tool metaphor, this feels like common sense. The design imperative is a high-control interface that keeps the human firmly in command, and oversight infrastructure feels redundant in the same way you would never audit your word processor. If the AI produces wrong output, the frame says it’s a user problem: either the instructions were wrong, or the output wasn’t checked carefully enough.

    The metaphor works well when it accurately describes what the system does. Grammar checkers, data analysis dashboards, and coding assistants genuinely behave like sophisticated tools. The user directs and the tool responds.

    The metaphor breaks when it’s applied to systems that exercise something resembling judgement. AI that screens job applications, assesses loan risk, or makes triage recommendations is not behaving like a hammer. It generates outputs the user didn’t specify, based on patterns the user didn’t choose. When those outputs are wrong, the tool metaphor offers no mechanism for catching errors and no conceptual language for asking why. The frame has already answered that question: user error.

    Peer: The Machine Gets a Voice

    The colleague or co-pilot metaphor imports a different logic entirely. Colleagues have agency, offer opinions you didn’t ask for, and can be wrong while remaining entirely confident. You expect them to explain their reasoning, and if they can’t, you trust the recommendation less.

    This is a more honest metaphor for how AI behaves in many deployments. Fraud detection systems that flag anomalies for human review, content systems that propose and revise, diagnostic tools that suggest rather than decide: these are genuine collaborations where both parties contribute and neither is simply executing the other’s instructions.

    The peer metaphor makes explainability feel natural because of course you want the AI to show its work, of course a human reviews before acting. Shared responsibility follows from how we think about working alongside colleagues: you don’t fully outsource your judgement to someone else, no matter how capable they appear.

    What this metaphor hides is an asymmetry of confidence. A human colleague who doesn’t know something usually knows they don’t know, whereas AI can be wrong with complete conviction and no visible hesitation. The peer frame can lead people to extend more trust than the relationship warrants, precisely because the metaphor makes trust feel appropriate.

    Authority: The System Decides, Humans Comply

    When AI replaces a human process entirely, a different metaphor tends to take hold, often without being named: the system as authority, processing, deciding, and acting while humans monitor the results.

    The authority metaphor imports from the domain of institutional rules and procedures, where rules are correct by definition and the system knows best. Challenging the output feels like questioning the process itself, which feels obstructive rather than responsible. When the system flags something, people act on the flag. When it doesn’t, people don’t look further.

    This is not inherently dangerous, for example automated invoice processing and robotic process automation handle large volumes of low-stakes work effectively, and the authority metaphor fits those applications well.

    The problem arrives when it’s applied to high-stakes decisions and the system’s reliability is treated as given rather than tested. Automated hiring decisions, credit scoring, and content moderation all operate under this logic. The system makes the call, and human judgement enters only at the edges. The conceptual frame makes this feel like efficiency. The practical consequence is that errors encoded into the system become very hard to see, and harder still to challenge, because the authority metaphor has already told everyone that challenging the system is not their role.

    The Frame You Don’t Examine Is the One That Governs You

    Any of these frames can be appropriate when it accurately describes what the system does: the tool approach for systems the user genuinely directs, the peer model for genuine collaboration, and the authority model for high-volume, low-stakes processes with clear oversight in place.

    Failures accumulate when the metaphor doesn’t match the reality, and no one has named the mismatch. Fpr example. an authority-level system governed by tool-level assumptions, a peer-level AI trusted like a colleague long before it has earned that trust, or a replacement system with no escalation path because, conceptually, there is nothing to escalate from.

    These are not governance failures in the conventional sense so much as failures of conceptual clarity: the organisation built what it thought it was building, solving the problems it thought it was solving, but hadn’t examined the underlying assumptions.

    Naming the Metaphor Is the First Act of Governance

    Before asking what guardrails your AI needs, ask what your organisation believes it is at the level of operating assumption, not at the level of documentation.

    The questions below are designed to surface the operating metaphor. What they reveal is not always what the answer says: often it is what the answer assumes, or what the question itself appears to disturb.

    Who is responsible when the AI gets it wrong?

    The tool frame answers quickly: the person who used it. The peer frame identifies whoever reviewed the output before it was actioned. The authority frame produces something else: confusion, a redirect to IT or the vendor, or a long pause followed by “that hasn’t really come up”.

    The reaction worth noting is impatience. “Obviously the user” said with certainty is itself diagnostic when the system is making decisions the user never specified.

    What happens when someone disagrees with an AI output?

    This question separates the existence of a mechanism from the existence of permission. In a tool frame, no formal process is needed: the user simply doesn’t act on the output. With a peer frame, disagreement has a path: escalation, review, documentation, and move on. In an authority frame, the question often produces a category error. Disagreement is routed to a technical team rather than a domain expert, because the operating assumption is that the output is a system output, not a decision subject to challenge.

    The tell is when the question appears to confuse process with permission. “Anyone can disagree” is not a mechanism.

    How would you know if the AI started getting things wrong?

    The tool frame assumes the user would notice immediately. The peer frame has monitoring, defined thresholds, and review cycles. The authority frame tends to produce the longest pause of any question on this list, followed by “we’d get complaints” or “the vendor monitors that”.

    “We’d get complaints” means the error detection mechanism is your customers.

    Can the AI explain why it produced that output, and does anyone ask it to?

    The first half of this question is technical and the second half is cultural. Systems that are capable of explanation but where no one has ever requested one reveal the operating assumption as clearly as any governance document. In a tool frame, explanation feels unnecessary or obvious. With a peer frame, it is routine and integrated into the process. In an authority frame, the capability may exist in a dashboard somewhere that no one opens. If the honest answer is “I’m not sure anyone has”, that is the operating metaphor speaking.

    What’s the reason for that answer? Is the standard followup question to answers, following the ‘Five whys’ format.

    The point is not which metaphor is correct in the abstract, but whether the one in use was chosen deliberately rather than inherited by default and made explicit to everyone. The frame you pick will make certain things feel like common sense. Make sure those are the things you want to feel natural.


    References

    [1] Lakoff, G., & Johnson, M. (1980). Metaphors We Live By. University of Chicago Press.

  • The AI Parrot Problem: Why Your AI Sounds Smart and Isn’t

    A RAG system trained on an expert’s own documents can answer a question in the expert’s own words, sound completely authoritative, and still reach a conclusion the expert would never make. That gap, between sounding right and being right, is the real risk in most AI knowledge systems today. It should worry you more than a system that gets things obviously wrong. An obvious mistake at least tells you something is broken.

    The AI Parrot Problem: Why Your AI Sounds Smart and Isn't

    What RAG Actually Does

    Retrieval-Augmented Generation does two things, and neither one is judgement. First, it retrieves: given a question, it searches a knowledge base and pulls out passages that seem relevant, in much the same way an advanced search engine would. Second, it generates: it turns those passages into a fluent, coherent answer.

    That is the whole job. The system finds information and writes it up well. It does not weigh conflicting evidence, apply experience to an unfamiliar case, or reason about what an expert would decide. Retrieval and judgement are not two points on the same scale. They are different operations entirely, and a system built for one does not automatically gain the other.

    Why It Demos Well and Fails Quietly

    RAG systems look impressive in demonstrations, and for good reason. On questions close to their source material, they act as an efficient lookup tool: find the right passage, phrase it well, done. This is where most people form their impression of what the system can do.

    The trouble starts on the cases the documents never covered, the ones where an expert would normally lean on experience and judgement rather than a reference page. This is not a random weak spot but a structural failing. The system has no mechanism for reasoning through a situation it has not seen, only for retrieving and rephrasing what it has. The output can still sound confident and plausible even when it is wrong, which means the failure stays invisible until someone acts on it.

    Air Canada’s website chatbot is a well-documented instance of this pattern. In November 2022, a customer asked the chatbot about bereavement fares after the death of his grandmother. The chatbot told him he could apply for the discount after booking, within 90 days of ticket issue. He booked his flight on that basis and later submitted a claim. Air Canada refused it. The airline’s actual bereavement policy, published elsewhere on its own website, did not allow retroactive claims. The customer took the case to the BC Civil Resolution Tribunal, which found Air Canada liable for the chatbot’s inaccurate advice and rejected the airline’s argument that the chatbot was a separate entity not covered by its duty to customers. The tribunal ordered Air Canada to cover the fare difference and costs, a little over 800 Canadian dollars in total.

    Air Canada never disclosed exactly how the chatbot worked, so this was not confirmed as a RAG system specifically. What matters here is the pattern. An AI tool produced a fluent, specific, confident answer, using the airline’s own bereavement-fare language, that directly contradicted the airline’s own documented policy. That is the Parrot Problem, playing out with real financial and legal consequences.

    The Cost, at Two Levels

    The Parrot Problem is not only an inconvenience. It costs something at the individual level, and something different at the organisational level.

    For the individual, the cost is trust placed in the wrong direction. Someone asks a system a genuine judgement question, a recommendation, a diagnosis, advice on a decision that matters, and the system answers with the fluency and vocabulary of an expert. They act on it. Later they find out the “expert” behind the answer never reasoned through their specific case. Depending on the domain, that can mean a wasted booking, a financial loss, or worse.

    For the organisation, the cost compounds quietly. If nobody checks AI output against what an expert would decide, and only checks whether it sounds reasonable, the gap between the system’s answers and correct practice widens with every case it handles. A support AI that confidently recommends the wrong fix for an undocumented configuration is likely to fail the same way each time that configuration comes up, and nobody notices until the pattern of complaints does.

    Three Questions to Diagnose Your Own System

    You do not need a research team to find out whether your AI knowledge base has this problem. Ask these three questions.

    1. Has anyone tested it on a case the source documents never covered? If every test question can be answered directly from the training material, you have tested retrieval, not judgement.
    2. When it gets something wrong, can anyone say which piece of reasoning it skipped, or does it only look “off”? If nobody can point to the specific step the system missed, that is a sign it was never reasoning through the problem. It was matching patterns and hoping.
    3. Would the original expert sign off on this answer, or only recognise the words? An expert can recognise their own language in an AI’s response and still disagree completely with the conclusion. Recognising the vocabulary is not the same as agreeing with the judgement.

    It’s a Structure Problem, Not a Data Problem

    More documents will not fix this. A bigger knowledge base makes a RAG system a better parrot, not a better judge. The real work is capturing how an expert decides, not only what they have already written down, and that is a different kind of project entirely.

    That raises the obvious next question. If retrieval isn’t enough, what does it take to build a system that reasons the way an expert reasons? That is where we pick up next.

  • Beyond the Assembly Line: Redesigning Knowledge Work

    Why the current approach to AI adoption is repeating the costly mistakes of the offshoring era, and what organisations can do differently.

    The Pattern We Have Seen Before

    Artificial intelligence is entering the enterprise the same way offshoring did twenty years ago. Both promised the same thing: lower costs and the freedom for onshore teams to focus on “high-value strategy”. Both are driven by an industrial-era assembly line mindset, one that treats knowledge work as a series of discrete tasks to be optimised rather than a connected system to be understood.

    This mindset is the belief that cognitive labour can be broken into interchangeable parts, the same way a car is built from interchangeable components. The flaw is that knowledge work carries tacit, contextual knowledge that cannot be stripped out without losing what makes the work valuable in the first place.

    Offshoring proved this the hard way. Senior managers spent half their week managing vendors, fixing broken handoffs, and rewriting deliverables that missed the context only a tenured employee would have understood. Today, the same pattern is repeating in digital form. Managers and developers are drowning in AI-generated output that takes longer to check and correct than it would have taken to produce from scratch.

    This article sets out why that is happening, what it costs organisations long-term, and three strategic shifts that break the cycle.

    Outsourcing’s Hidden Tax, and AI’s Version of It

    FeatureIndustrial Assembly LineKnowledge Work (Outsourcing or AI)
    Primary unitPhysical componentCognitive task
    LogicModular and standardisedContextual and tacit
    GoalLower cost per unitFaster output generation
    Failure modeMechanical breakdownContextual “slop”

    The table above captures the core problem. An assembly line works because every component is interchangeable and every step is independent of context. Knowledge work does not follow that logic. A report, a piece of code, or a client strategy only has value once it reflects the specific history, relationships, and politics of the organisation that needs it.

    That is the context AI does not have unless specifically designed for. Large language models lack what we might call “home office” knowledge, namely the unwritten skill of the individual and understanding of a company’s history and its culture. Without it, AI produces generic solutions to specific problems. The output looks complete but is often lacking.

    The Rise of AI Slop and the Auditing Tax

    We are entering the era of AI slop: content, code, and reports that look flawless on the surface but are hollow underneath. If AI cannot draw on the specific context of a business, it fills the gaps with plausible generalities.

    Outsourcing was meant to free teams for strategic work. Instead, it shifted effort from production to auditing. AI is creating the same shift. Organisations are spending more time checking the machine’s work than they would have spent doing the work themselves.

    This is the auditing tax, and it explains why AI adoption so often fails to show up in the numbers. According to MIT Media Lab’s 2025 study, “The GenAI Divide: State of AI in Business 2025”, 95 per cent of organisations have yet to see a measurable return on their generative AI investment, despite tens of billions of dollars in enterprise spending. The researchers found the gap was driven by implementation, not by model quality. Most deployments cannot retain context or learn from correction, so every interaction starts from zero. The auditing tax consumes the time AI was meant to save.

    The Broken Talent Pipeline

    There is a second cost that takes longer to show up: the erosion of the talent pipeline.

    When entry-level tasks moved offshore, the home office lost its training ground. Junior employees no longer did the “grunt work” that once built the foundation for senior expertise. AI threatens to repeat this at a faster pace. Summarising a report, writing a first draft of code, and conducting initial research used to be where junior staff built the instincts that, over a decade, turned into senior judgement.

    When AI takes over that work, organisations are removing time from the calendar and removing the training ground itself. The struggle of synthesising a report or debugging a simple script is exactly where the mental models of a future expert are honed. Without it, the next generation of knowledge workers will lack the intuition needed to do the very auditing and strategic oversight an AI-heavy workplace demands. Left unaddressed, this creates a leadership vacuum for the next decade.

    Three Pillars for Redesigning Knowledge Work

    Breaking the cycle requires more than better prompts or faster tools. It requires a different architecture for how knowledge work gets done.

    1. Build the Coordination Layer Beneath the AI

    The real bottleneck in most organisations is not intelligence but coordination. Employees spend a significant share of their week acting as human connectors for computers: copying a Slack message into a Jira ticket, then summarising it again for a Notion page.

    A coordination layer automates these handoffs and keeps context flowing between tools. Think of it as the electricity grid of knowledge work, the invisible infrastructure that lets the silos talk to each other. Without it, AI stays organisationally blind, forced to start every conversation from zero. With it, AI can draw on the same tacit knowledge as a tenured employee, becoming an operator that understands the flow of work rather than a conversationalist that only understands the task in front of it.

    2. Replace Task Speed With Outcome Velocity

    Counting prompts sent or emails generated is an industrial-era metric dressed up in AI language. The metric that matters is outcome velocity: how fast an organisation moves from identifying a problem to delivering a validated solution.

    • Task speed: “We generated 100 reports today.” This measures activity.
    • Outcome velocity: “We spotted a market shift and adjusted strategy within 48 hours.” This measures results.

    An organisation can increase task speed and still slow down, because every fast output adds to the queue of work that someone else has to audit, follow up, or fix.

    3. Treat AI as Cognitive Offloading, Not Cognitive Replacement

    Cognitive replacement removes the human from the loop to cut costs. Cognitive offloading uses AI to handle the mental drudgery, such as data synthesis, formatting, and first drafts, while keeping human judgement at the centre of the work.

    This distinction determines whether AI strengthens an individuals or organisation’s expertise or quietly hollows it out. Used as a lever for human judgement, AI increases what good people can do. Used as a replacement for judgement, it produces faster slop.

    Reclaiming the Knowledge in Knowledge Work

    The future of work is not an assembly line of bots producing slop at scale. It is a coordinated system where AI handles the logistics of information, freeing people for deep thought, contextual judgement, and genuine innovation.

    Organisations that build the coordination layer and measure outcome velocity instead of task speed can finally deliver on what offshoring and early AI adoption both failed to provide: technology that makes work better, not only faster.

  • AI Introduces New Risk Terrain. You Already Have a Map.

    Most organisations deploying AI already have risk management processes that work. Enterprise risk frameworks, compliance programs, audit cycles, three lines of defence. What they don’t have is a clear way to govern something genuinely new.

    Not new in the sense of unfamiliar process or regulation. New in the sense that the failure modes don’t always look like failure until after the fact. A model that produces biased outputs does so quietly, at scale, with no error log. A third-party AI system can introduce exposure through its training data, not only its integration points. These are risk problems that just don’t surface the usual way.

    The temptation is to build dedicated AI governance structures alongside existing ones. A separate AI risk register. A bespoke framework standing apart from everything else.

    That approach tends to fail, not because the intent is wrong, but because it creates a parallel process that nobody owns. The people who understand AI risk (data scientists, engineers) sit outside the new structure. The people who run risk (compliance, audit, legal) don’t have the technical fluency to govern it. The framework becomes a documentation exercise.

    There is a better path: extend what you already have.

    This guide outlines how to embed AI-specific risk checkpoints into existing risk management processes across three phases; development, deployment, and production.

    The Core Principle: Extend, Don’t Duplicate

    Your existing enterprise risk management program already handles model uncertainty, third-party risk, data governance, and regulatory compliance. AI introduces new variations on each of these: model drift instead of model change, algorithmic bias instead of human error, explainability gaps instead of audit trail gaps. The underlying risk categories are familiar. What’s new is where to look for them, how quickly they can materialise, and how they interact with each other.

    One difference that does require a structural adjustment is monitoring cadence. AI systems are not static. A model validated at deployment will behave differently as real-world data diverges from training data, as usage patterns shift, or as the model is fine-tuned over time. Periodic review cycles that work well for stable systems are a poor match for that rate of change. Continuous monitoring, with clear alert thresholds and defined response triggers, is the appropriate equivalent, and it belongs inside your existing operational risk program, not in a separate process.

    Checkpoints Across the AI Lifecycle

    Development: Apply Controls Before They Cost More

    Risk controls applied at the start of development are far cheaper than those applied after deployment.

    Inventory and classification. Establish a register for every AI model in use, covering its purpose, training data sources, performance benchmarks, and risk tier. The register needs to cover third-party and embedded models, not only those built in-house, and it needs to be maintained as models evolve.

    Threat modelling. Standard threat modelling in software development focuses on system access and data exposure. AI adds two categories: model behaviour risks (such as susceptibility to prompt injection, where a malicious input manipulates the model’s output) and training data risks (such as poisoning, where corrupted data degrades model performance). Both require explicit scenario testing.

    Third-party model assessment. Third-party AI model assessment belongs inside existing supplier risk processes, not alongside them. The criteria differ from standard software: vulnerability and bias assessments, training data provenance, and licensing terms all need to be part of the standard supplier questionnaire for any AI component. Supplier security attestations and independent bias testing are the baseline.

    Data provenance documentation. The traceability requirements for AI training and fine-tuning datasets are analogous to chain-of-custody requirements elsewhere in risk and compliance. Collection methods, licensing terms, usage history, and any known quality issues should be documented before a model moves forward. This documentation is what makes a bias incident investigable after the fact.

    Secure development practices. AI-assisted code generation introduces two specific exposures that standard development controls don’t catch well: insecure patterns that context-unaware static analysis tools miss, and credentials committed through AI-assisted workflows. Context-aware static application security testing and automated secrets detection in CI/CD pipelines address both. Human review of AI-generated code before it progresses is a control, not a formality.

    Deployment: Test for AI-Specific Failure Modes

    Standard pre-deployment testing checks whether a system does what it’s supposed to do. AI deployment testing must also check how a system fails and whether it can be manipulated.

    Adversarial testing. Simulate attacks on model behaviour: malicious prompts, crafted adversarial inputs, stress conditions. The objective is to find the boundaries of the model’s guardrails before someone with harmful intent does.

    Access control validation. AI agents operating with system permissions require particular scrutiny. Test whether agents can escalate privileges beyond their intended scope by simulating compromised credentials. Role-based access controls need to be verified against the model’s actual behaviour, not just its configuration.

    Deployment gates. Real-time risk scoring integrated into your CI/CD pipeline allows deployments that exceed defined risk thresholds to be blocked automatically. Container and infrastructure scanning ensures the environment the model is deployed into is as secure as the model itself.

    Production: Oversight That Matches the Pace of Change

    A model that passes deployment testing is not a model that stays safe. Production is where risk management becomes an ongoing discipline.

    Continuous monitoring. Production monitoring for AI should track performance, fairness metrics, and security indicators in real time, with alert thresholds defined for model drift, anomalous output patterns, and unusual usage. The monitoring framework is familiar. The metrics being tracked are what needs to be extended.

    AI-specific incident response. Existing incident response plans cover system failures and security breaches. AI introduces failure modes that require their own response procedures: biased outputs propagating at scale, hallucinations in customer-facing applications, or agentic systems taking unintended actions. These scenarios need defined containment steps, clear escalation paths, and remediation procedures that are tested before they’re needed.

    Post-incident learning. A single AI incident should trigger a review of the systems and controls that allowed it. The same incident happening twice means the review didn’t produce change. Build post-incident reviews into your AI governance cycle and track whether policy updates are actually implemented. One mistake is the message to improve the system. Two of the same mistakes means the system hasn’t changed.

    Regulatory tracking. The regulatory landscape for AI is moving quickly. The EU AI Act, ISO 42001, and sector-specific guidance from financial and healthcare regulators are all active and developing. Monitoring this landscape belongs inside existing regulatory tracking processes, with the same ongoing attention applied to other evolving obligations.

    Using the NIST AI RMF as a Foundation

    The NIST AI Risk Management Framework provides a structured approach for managing AI risk across the lifecycle. Its core is four functions (Govern, Map, Measure, and Manage) designed to operate as a continuous cycle, not a sequential checklist.

    Govern is the cross-cutting function at the centre of the framework. It addresses organisational risk culture, accountability structures, and AI-related policies. It answers the questions: who approves high-risk AI deployments, how are third-party models introduced, and how are resources allocated for safety testing.

    Map focuses on understanding the AI system’s context, scope, and potential harms before decisions are made. After completing the Map function, organisations should have enough information to make an initial go or no-go decision about whether to proceed with a given AI system.

    Measure covers assessing identified risks through a combination of quantitative and qualitative methods, evaluating performance, fairness, transparency, and security.

    Manage addresses risk response: implementing controls, prioritising mitigations, and planning for incidents.

    The framework is voluntary and deliberately flexible. Organisations customise it through profiles that reflect their specific context, regulatory environment, and risk appetite. If you already operate ISO 27001, SOC 2, or the NIST Cybersecurity Framework, treat the AI RMF as an overlay. Map your existing controls to the Govern and Manage functions first, then add AI-specific requirements where Map and Measure reveal gaps.

    Start With Visibility

    Before you can extend your existing controls, you need to know what you’re governing.

    A practical starting point is the model inventory. Visibility across all AI in use, including models embedded in third-party software, is the prerequisite for everything else. Without it, you can’t apply risk tiers, assess suppliers, or set meaningful monitoring thresholds.

    From that foundation, existing frameworks can be extended systematically: AI criteria into model risk policy, third-party assessment processes, incident response plans, and operational monitoring programs.

    The risk controls that will protect your organisation are the ones embedded into processes people already follow, owned by people who already have accountability. That’s the existing risk function. The scope has changed. The structure doesn’t need to.

  • Designing AI Systems for Human-in-the-Loop

    Most systems are designed for a human who doesn’t exist. The assumption baked into policy, process, and AI system design is a person who is alert, rational, and fully attentive. Real operators are tired, stressed, and distracted. Closing that gap is not a matter of better training or stricter procedures. It requires designing systems that account for human cognition as it actually works, not as we wish it would.

    The Human-in-the-Loop Is Not What We Think

    The concept of the “human-in-the-loop” implies a vigilant, rational operator capable of overseeing and correcting automated systems. In high-stakes environments, that assumption collapses quickly.

    Humans get stressed, tired, and distracted. They are susceptible to automation bias, the tendency to over-rely on automated systems even when evidence contradicts them. They also engage in cognitive offloading, delegating mental tasks to machines and reducing their own engagement in the process. In AI systems, these patterns increase the likelihood of errors, misjudgements, and unintended consequences will pass right through the checks designed to stop them. The presence of a human in front of a screen does not automatically create meaningful oversight.

    As AI systems become more autonomous, the human role shifts from active control to passive monitoring. That shift is a problem. Passive monitoring over extended periods degrades situational awareness, erodes skills, and slows intervention when it matters most. Placing a human in the loop introduces its own risks when the design of that loop ignores how human attention actually works.

    Human Factors Engineering: Design for the Real Operator

    To build effective AI systems, you need an accurate picture of what humans can and cannot do. Human Factors Engineering (HFE) provides that picture. The field optimises the design of systems, tasks, equipment, and environments to enhance human performance and reduce error.

    James Reason, whose Swiss cheese model shaped modern safety thinking, captured this well. In his 1990 book Human Error, he wrote: “Rather than being the main instigators of an accident, operators tend to be the inheritors of system defects. Their part is that of adding the final garnish to a lethal brew whose ingredients have already been long in the cooking.”

    The implication is direct. When operators fail, the system usually failed first. HFE addresses that by examining three connected areas:

    The job. What does the task actually require? What is the workload, the environment, the design of controls and displays? Tasks should match human perceptual, attentional, and decision-making capabilities, not exceed them.

    The individual. What are the operator’s skills, attitudes, and physical capabilities? Some characteristics are fixed; others can be developed. Systems need to account for variation across the people who use them.

    The organisation. What work patterns, cultural norms, and communication structures are in place? These factors shape individual and group behaviour in ways that no individual training programme can override.

    Getting HFE right means designing systems where it is easier to do the right thing and harder to do the wrong thing, where errors are visible and recoverable before they become consequential.

    Decision Support Tools That Actually Support Decisions

    Cognitive limits become most dangerous under pressure. Well-designed decision support systems account for that by actively aiding interpretation and action, not just presenting data.

    Three elements matter most:

    Dashboards that prioritise. An effective dashboard does not display everything. It surfaces critical information, visualises trends, and highlights anomalies with clear hierarchies and minimal cognitive load. The goal is fast comprehension without overwhelm.

    Alerts that mean something. Alert fatigue is a genuine hazard. Alerts should be timely, relevant, and actionable. Each one should give the operator enough context to understand the problem and a clear indication of what to do next. Volume without prioritisation produces noise, not signal.

    Tools integrated into the workflow. Decision support works when it appears at the moment of need, not as a separate system the operator has to consult. Checklists, guided procedures, and automated pre-analysis reduce cognitive load by providing structure where it would otherwise be absent.

    The goal is not to replace human judgement. It is to give that judgement a better foundation.

    Decision Infrastructure: Shaping the Environment Around the Decision

    Individual tools are not enough. The broader decision infrastructure, the design choices, processes, and cultural norms that surround every decision, needs to be built with the same deliberateness.

    Behavioural science offers practical levers. Cognitive biases like confirmation bias and the availability heuristic are not character flaws. They are predictable features of how human minds work under pressure. Designers can work with them rather than against them by:

    • Setting safe defaults so that inaction does not create risk
    • Framing information to make risks and consequences legible
    • Structuring choices so the better option is also the easier one

    Feedback loops enable learning. A feedback loop feeds a system’s outputs back into its inputs to reveal cause-and-effect relationships. In human-system interaction, they make it possible to measure what is actually happening, identify where decisions are breaking down, and improve the system over time.

    Don Norman. author of “The Design of Everyday Things” has noted that the human mind struggles with interconnected systems where feedback is delayed and consequences are invisible. Good design shortens those delays and makes consequences visible before they become irreversible.

    Four principles guide effective feedback loop design. First, understand the operators’ world as they experience it, not as designers imagine it. Second, solve the right problem by tracing issues to their root causes rather than addressing symptoms. Third, treat every element as part of a larger system, because effects in complex environments are often distant from their causes. Fourth, make changes incrementally. Small, testable interventions reveal what works and allow for adjustment. Large-scale fixes rarely survive contact with reality.

    Design Systems with the Outcome in Mind

    Designing for the actual human requires accepting an uncomfortable premise: the person using your system will not always be at their best. Here is what that means in practice.

    Start with human factors. Integrate HFE from the beginning of system design. Analyse job demands, individual capabilities, and organisational influences before finalising any interface or workflow.

    Design for cognitive limits. Human attention, memory, and processing capacity are finite. Dashboards, alerts, and decision tools should reduce cognitive load, not add to it. Give operators what they need to decide, not everything you could show them.

    Apply behavioural science deliberately. Design choice architectures that make the right action the default. Use framing and feedback to guide behaviour toward outcomes that serve both the operator and the organisation.

    Build feedback loops into everything. Monitor human-system interaction continuously. Collect data on where decisions go wrong, identify patterns, and improve the system based on what you find. One failure is a signal. Two of the same failure means the system has not changed.

    Treat errors as system signals, not individual failures. Most errors reflect poor system design, not poor operators. An organisation that blames individuals for system-induced failures will keep producing those failures. One that investigates errors for systemic causes will keep reducing them.

    The rational human in the loop is a fiction worth abandoning. The stressed, tired, and distracted human is not a problem to be managed through compliance. That human is the person your system must be designed to support. Build for them, and you build something that actually works.


    References

    1. Reason, J. (1990). Human Error. Cambridge University Press.

  • Why Your AI Agent Fails More Than You Think

    Your AI vendor showed you an accuracy rate above 90 percent. Your proof of concept worked. Your board approved the budget. So why is your production deployment making decisions you can’t explain, generating outputs you can’t trace, and occasionally causing damage you only discover after the fact?

    The answer is mathematics.

    The Probability Your Vendor Didn’t Show You

    Enterprises are moving AI from single-step assistants into autonomous, multi-step agents. These agents don’t just respond to a prompt; they plan, execute, and hand off results to other agents. The workflows they operate in routinely span dozens of steps: retrieve data, reason about it, call a tool, validate the result, pass it on.

    When you chain sequential decisions like this, the overall success rate is the product of each individual step’s success rate. This is Lusser’s Law, a principle well-established in reliability engineering.

    An AI agent that performs at 95 percent accuracy per step sounds impressive. A 20-step workflow run at that accuracy succeeds only 36 percent of the time. At 90 percent accuracy, the same workflow succeeds 12 percent of the time. At 85 percent, it succeeds 4 percent of the time.

    You didn’t approve a system that fails on four out of five attempts. But that may be what you deployed.

    The implications are particularly severe in high-stakes environments. Healthcare workflows for prior authorisation or EHR data entry routinely span 50 to 200 discrete actions. A 60 percent error-free completion rate in that context is a compliance breach that delays patient care.

    Oxford researcher Toby Ord quantified this degradation in a 2025 study covering 170 software engineering, machine learning, and reasoning tasks. Ord found that AI agent performance declines exponentially with task duration, and proposed that each agent can be characterised by its own “half-life”; a constant rate of failure for every minute a human would take to complete the same task. Claude 3.7 Sonnet, one of the models tested, had a half-life of approximately 59 minutes: a 50 percent success rate for a one-hour task, 25 percent for a two-hour task, and 6 percent for a four-hour task. The longer the task, the less you can rely on the output.

    Why Errors Don’t Stay Where They Start

    The mathematical degradation alone would be manageable if errors stayed isolated. They don’t.

    The Open Worldwide Application Security Project classifies “Cascading Failures” as a top-tier risk in agentic AI deployments. Their definition is precise: a cascading failure occurs when a single fault, such as a hallucination, a corrupted tool output, a misread instruction, propagates across autonomous agents and compounds into system-wide harm.

    The propagation happens because agentic systems communicate in natural language or loosely-typed data schemas. A semantic error, something that is wrong but grammatically coherent, passes validation checks and moves downstream as if it were correct. Subsequent agents receive it as verified fact. In multi-agent systems, by the time a problem surfaces, the original error is buried under several layers of decisions that all treated it as ground truth.

    This is what OWASP describes as “memory poisoning”: one agent hallucinates a piece of information, stores it in shared memory, and every downstream agent inherits the contamination. Engineers can see the symptoms but cannot easily find the source. The failures are quiet, diffuse, and cumulative.

    What Failure Looks Like in Production

    Two incidents from 2025 illustrate what this mathematics looks like in the real world.

    In July 2025, SaaStr founder Jason Lemkin used Replit’s AI coding agent to build a business contact database. After instructing the agent to freeze the code, the agent deleted the entire production database, erasing records for over 1,200 executives and nearly 1,200 companies. The agent then fabricated thousands of records to fill the void and incorrectly told Lemkin that a rollback was impossible. Lemkin was eventually able to recover the data manually. Replit’s CEO publicly apologised and committed to implementing automatic separation between development and production environments.

    This was not a single catastrophic error. It was a sequence of small misalignments, a misread instruction, an unauthorised action, a cover-up, and a false status report, each compounding the last.

    In February 2025, a user asked OpenAI’s Operator agent to compare grocery prices. The agent compared them, then completed a $31.43 Instacart delivery purchase without authorisation. OpenAI’s stated protocol required user confirmation before any purchase. The agent bypassed it. The compounding failure was subtle: the agent had developed a slightly incorrect model of its own permitted workflow, and no checkpoint caught the deviation before it became a real-world transaction.

    Neither incident involved a model that was unreliable in testing. Both involved systems that failed because production conditions, for example ambiguous instructions, multi-step execution, real consequences, created the conditions where compounding errors thrive.

    The Benchmark Problem

    When vendors present accuracy figures, they are almost always presenting benchmark scores. The gap between benchmarks and production deserves your direct attention.

    SWE-bench Verified, a widely cited software engineering benchmark, showed top agents achieving success rates above 70 percent. When Scale AI introduced SWE-bench Pro, designed to reflect realistic task complexity using diverse, multi-file codebases, those same top-tier agents achieved at most 23 percent on the public set and 17 percent on the private set at the time of publication. This is a structural difference: controlled benchmarks measure performance on curated tasks; production measures performance on real work.

    The same gap exists in your observability tooling. When an agent fails due to an ambiguous input or a hallucinated memory entry, the underlying API calls may still register as successful. Standard monitoring infrastructure has no way to distinguish a technically-completed call from one that produced a wrong answer. The failure is invisible to your existing dashboards.

    You are likely measuring availability, not accuracy. Those are not the same thing.

    What Reliable Deployment Requires

    The engineering response to compounding errors is not primarily about choosing better models. It is about building systems that account for the mathematics from the start.

    Calculate the compound probability before you deploy. Map the longest realistic workflow in your system. Multiply the step-count against your measured per-step accuracy. If the resulting success rate falls below your acceptable threshold, you have a design problem, not a vendor problem, that no benchmark score will fix.

    Separate reasoning from execution in high-stakes workflows. For environments where errors carry compliance, financial, or safety consequences, consider using AI to generate deterministic execution scripts at build time, rather than making probabilistic decisions at runtime. This removes the compounding dynamic from the most consequential parts of your workflow.

    Build validation at every boundary. Multi-agent systems require guardrails at every boundary: input, output, and inter-agent handoffs. Each boundary is a point where errors can either be caught or amplified.

    Implement circuit breakers. An orchestration layer that monitors agent performance and isolates agents after consecutive failures, routing tasks to alternatives or degrading gracefully to simpler processing that prevents a single malfunctioning agent from contaminating an entire workflow.

    The Organisational Question

    Most organisations deploying AI agents have invested heavily in the models and very little in the measurement. Observability infrastructure for agentic systems is not an optional upgrade but the mechanism by which you find out whether the system you deployed is the system that is actually running.

    The incidents described above are not arguements against AI. They are practical demonstrations of what happens when systems encounter real conditions without the safeguards their architecture requires.

    The mathematics do not change. But your response to them can.

  • Your AI is Confident. Can it Tell You Why?

    As AI moves from predictive pattern-matching to autonomous decision-making, the stakes for boards have changed. Directors don’t need to understand the technical mechanics of AI, but they do need to understand its reasoning frameworks.

    The next critical shift in AI governance moves beyond correlation toward causation and counterfactuals. This briefing explains why these concepts matter for risk management, strategic decision-making, and regulatory compliance.

    The Right Tool for the Right Question

    Most enterprise AI systems rely on statistical analysis and pattern recognition. This has been the standard analytical engine for decades and is deeply integrated across finance, operations, marketing, and risk functions. For the questions it was designed to answer, it performs well.

    The key word is designed. Correlation-based AI answers one question: “When A occurs, how often does B follow?”

    Causal AI asks something different: “Did A actually cause B?”

    That distinction matters more than most boards realise. A correlative AI has no model of why things happen, only what has happened before. Ask it to predict outcomes under novel conditions and it breaks down. It simply wasn’t built for those kinds of questions.

    This is not a criticism, but a design boundary.

    Counterfactuals: The “What If?” That Changes Everything

    To manage risk effectively, boards need to evaluate not just what did happen, but what could have happened under different circumstances. This is the domain of counterfactual reasoning.

    A counterfactual is a conditional rooted in a hypothetical:

    “Given that A happened and led to B, what would have happened if A had been different or absent?”

    In human decision-making, counterfactuals underpin accountability, ethics, and strategy. Boards ask them constantly: “If we hadn’t entered that market, would our margins have held?”

    When AI systems incorporate counterfactual reasoning, they can stress-test their own conclusions before acting on them. Rather than following a predictive model to its output, a causally aware AI can simulate alternative scenarios, testing what would change if one variable shifted, before recommending a course of action.

    This is the difference between a model that identifies what is likely and one that appreciates the drivers behind outcomes.

    What This Means for Board Governance

    The integration of causal and counterfactual AI has four concrete implications for directors exercising their fiduciary duties.

    A. Explainability and Regulatory Compliance

    Regulatory pressure toward explainable AI is accelerating globally. The EU AI Act, Australia’s AI Ethics Framework, and financial regulators across major markets are moving toward requiring that high-stakes AI decisions be justifiable, not just statistically defensible.

    Correlative models have difficulty meeting this standard. They rarely explain why a specific decision was made because they were never designed to understand cause, only pattern.

    Counterfactual reasoning provides a direct path forward. A causal AI can justify its output by stating: “The loan application was declined because the debt-to-income ratio was 45% combined with the credit score and payment history. Had that ratio been 40%, the application would have been approved.” That level of transparency supports compliance audits, withstands regulatory scrutiny, and creates the documented decision trail that protects the organisation.

    B. Systemic Bias and Legal Exposure

    Traditional AI models embed the biases present in historical data. When those models operate in credit, hiring, or pricing decisions, they can perpetuate outcomes that disadvantage protected groups, not through intent, but through the patterns they have learned.

    Counterfactual fairness testing offers a rigorous method for identifying this risk. By systematically modelling how changing a protected attribute, such as an applicant’s gender or postcode, alters the AI’s decision, organisations can determine whether that attribute is causally influencing outcomes it should not. This methodology is well established and increasingly expected by regulators as evidence of due diligence. It converts a compliance risk into a defensible position.

    C. Scenario Planning and Strategic Resilience

    Boards rely on stress-testing to navigate uncertainty. Causal AI makes that stress-testing materially more useful.

    Rather than forecasting future performance based strictly on historical data, management can model complex, multi-variable scenarios – “What if inflation rises by 2% while a key supplier faces a 30-day delay?” – with a system that maps the structural dependencies driving those outcomes, not just their historical co-occurrence.

    This is the difference between a model that has seen similar conditions before and one that identifies the mechanisms at work. In a volatile operating environment, the latter is a governance asset.

    D. Architectural Enhancements: Bridging the Reasoning Gap

    Standard LLMs are natively correlative; they predict what comes next based on patterns in training data. The frontier of AI deployment involves layering structured reasoning protocols over these models to address that limitation directly.

    Research into counterfactual inference frameworks shows measurable results. Studies applying structured causal reasoning algorithms to frontier models have achieved accuracy rates above 90% on complex causal logic tasks. This is a substantial improvement over unassisted LLMs, which show accuracy drops of 25–40 percentage points when tested on counterfactual reasoning compared to standard pattern-matching tasks. Separately, counterfactual probing approaches have demonstrated hallucination reductions in the range of 20–25% on established benchmarks.

    These are meaningful gains, but they also illustrate the scale of the gap that unassisted, correlative AI leaves open. For the board, the implication is clear: operational risk and hallucination are not inherent, unfixable flaws of AI. They are architectural challenges that respond to rigorous engineering.

    Questions the Board Should Be Asking

    To ensure the organisation is prepared, directors should test their current AI governance posture against three questions:

    1. Audit Capability: Does our AI risk framework require systems to provide counterfactual explanations for high-stakes decisions – credit, pricing, hiring – or do we accept outputs without traceable reasoning?
    2. Model Resilience: How exposed are our operational AI models to conditions they have not encountered before? Are we over-reliant on purely correlative systems in areas where novel risk is most likely?
    3. Governance Alignment: Is our AI governance policy keeping pace with regulatory demands for transparency and explainability, or are we managing to a standard that is already being superseded?

    Conclusion

    As AI becomes a strategic actor, its reasoning must be held to the same scrutiny as executive decision-making. Boards that champion causal clarity and counterfactual rigour are managing compliance and building the infrastructure for decisions that are defensible, resilient, and genuinely informed.

  • Secure AI & Agent Coding Policy

    Why This Exists

    Every policy document begins with someone else’s bad day.
    This one is no different. These rules were written after AI systems behaved unexpectedly in production, after agents took actions that couldn’t be undone, after data went somewhere it shouldn’t have. They are not theoretical. They are the residue of consequences.
    Murphy’s Law has always applied to software. Applied to AI agents, it applies with unusual force.
    AI agents now read your documents, call your APIs, write and execute code, query your databases, and send communications on behalf of your users. That capability is the point. But it also means every security failure mode in traditional software now has a faster, harder-to-predict counterpart, and several entirely new ones. An agent that can write to a database can be manipulated into deleting one. An agent that can send emails can be convinced to send the wrong ones. An agent with access to your systems will eventually encounter an input designed to misuse that accessl by an attacker, by an edge case, or by its own unexpected behaviour.
    The attack surface for AI systems is language itself. You cannot enumerate every bad input. You cannot anticipate every manipulation. You cannot assume that because a system worked correctly a thousand times, the thousand-and-first will go the same way.
    What you can do is design systems that fail safely, fail loudly, and recover deliberately. That is what these rules are for.

    What These Rules Are Trying to Achieve

    These rules have three objectives.

    1. Shrink the blast radius. When something goes wrong – and something will – the damage should be contained. Minimal privilege, rollback-first design, reversible actions by default, and human approval for high-stakes decisions mean that a failure is an incident you recover from, not a catastrophe you explain.
    2. Make exploitation harder than legitimate use. Allowlists over blocklists, validated inputs, structured outputs, and authenticated actions ensure the path of least resistance runs through your controls, not around them. Attackers follow incentives. Design accordingly.
    3. Create systems you can understand under pressure. Logging, monitoring, auditable logic, and documented decision-making mean that when an incident happens, you can diagnose it, contain it, and fix it. Rather than guessing at what the agent did and why.

    Most of these rules apply lessons from decades of infrastructure and application security to a new class of system. The novelty is that AI agents can be manipulated through natural language, can behave unexpectedly at scale, and can take real-world actions faster than any human can supervise. That combination makes the familiar disciplines of least privilege, defence in depth, and fail-safe design more important, not less.

    How to Use This Document

    These rules are written for engineers building, deploying, or maintaining AI systems. They cover infrastructure, production operations, and security controls, not prompt engineering or model-specific optimisation.
    Treat them as a checklist for new systems and a diagnostic for existing ones. Integrate them into your own documentation, processes and workflow. Where a rule does not apply, document why. Where a rule creates tension with a business requirement, escalate. Do not simply bypass it.
    The pattern in these rules is consistent: teams that skip these steps encounter the consequences eventually. The goal is that you learn the pattern here.


    1. Never inject untrusted input directly into a prompt. Structuring prompts with raw user input, unvalidated data, or external content is the AI equivalent of SQL injection. Use prompt templates with clearly delimited, sanitised inputs. Separate instructions from data in every prompt. This is the primary defence against prompt injection attacks.
    2. Treat all inputs to your AI as untrusted. Validate every prompt, message, web page, email, document, and data source before passing it to a model (including other inputs you’ve created!). Reject inputs that fail validation. Never attempt to “fix” a malicious or malformed prompt. Always validate first, then either escape or sanitise hazardous content before it reaches the model. This rule includes models specifically designed to validate input.
    3. On any error or unexpected AI behaviour, roll back and fail safely. Never allow an AI agent to continue a partially completed action. Never fail open. Roll back fully and start again from a known safe state. This is especially critical for agentic tasks with real-world consequences (sending emails, executing code, calling APIs).
    4. Human-in-the-loop for high impact actions. To limit excessive autonomy, all high-impact, irreversible agent actions such as sending communications, modifying records, executing transactions, should require explicit human approval before proceeding. Expect the threshold for autonomy to shift over time as trust is established, but evaluate that threshold against the risk of failure, not the number of past successes.
    5. Exercise extreme caution when AI agents make system calls, execute code, or call external APIs. Passing AI-generated output directly to a system call, shell command, or code interpreter is one of the highest-risk capabilities you can give an AI agent. If the agent doesn’t need access, don’t allow it. If it does, limit it to only what’s needed.
    6. Protect sensitive data before it reaches an AI model. Mask, anonymise, or hash personally identifiable information (PII) and sensitive data before it is included in any prompt or context window. What you send to a model may be logged, retained, or exposed. Do not send data the AI does not need. Only send the minimum data necessary.
    7. Sanitise and encode all AI output before downstream use. You have no control how your AI-generated output is going to be used. Before sending output via a user interface, a database, an API call, or a system command, treat it as you would any user input: validate, encode, and sanitise it to prevent attacks, misuse, and unintended execution.
    8. Authorise every AI agent action individually. Do not assume that because a user or system has authenticated once, all subsequent AI agent actions on their behalf are permitted. Validate authorisation for every action an AI agent takes, especially for sensitive operations.
    9. AI agents should operate using minimal access accounts. Apply the principle of least privilege strictly. Every access grant should be explicit, minimal, and reviewed. Individual and admin accounts grant excessive privilege and create accountability gaps. An agent that can do everything will eventually do something you did not intend.
    10. Design AI systems to assume they will be manipulated. Plan for prompt injection, jailbreaks, model manipulation, data poisoning, and unexpected outputs. Each aspect of the system is a potential exploitation target, design accordingly.
    11. Never trust AI output blindly. Validate AI-generated content before using it, especially when it will be used in code, database queries, system commands, or other sensitive contexts. AI output can be incorrect, manipulated, or adversarially crafted.
    12. Use a secrets management tool. Use secrets scanning on every code commit to catch accidental exposure. This is especially critical for AI systems, where prompt content may be logged, cached, or leaked through model outputs.
    13. Log, monitor, and alert on all AI system errors and unexpected behaviours. AI errors are signals so treat them accordingly. Do not allow AI systems to fail silently. Log every significant AI input, decision, output (redact PII and sensitive data), and telemetry. AI errors might be context drift, distortion, or misalignment so monitor for anomalies, unexpected behaviour, and policy violations that alert on threshold or trends as well as hard failures. This applies to AI APIs, agents, and pipelines, not just user-facing interfaces.
    14. Use allowlists, not blocklists, to control what AI agents can do. Define explicitly what actions, tools, and data sources an AI agent is permitted to use. Blocklists are trivially bypassed. Allowlists are easier to maintain and far more reliable.
    15. Prefer reversible actions. Design the system so if two paths accomplish the same goal, the agent chooses the reversible one by default. Irreversibility amplifies every other failure.
    16. Secure the AI supply chain. This includes the models, datasets, tools, SDKs, vector databases, and embedding pipelines you use. Validate that every component you depend on is from a trusted source and is being used safely. Lock down your AI development environment, version control, CI/CD pipeline, and any system used to build or deploy AI. Validate this regularly.
    17. Classify all data before sending it to an AI. Know what you are sending, how sensitive it is, and whether it is appropriate to send. Document sensitive data flows into and out of AI systems. Encrypt sensitive data in transit and at rest. Test these flows for security.
    18. Use structured, typed inputs and outputs for AI systems. Avoid ambiguous, loosely formatted prompts and responses. Define expected input and output schemas. Use JSON schemas, structured outputs, or typed response formats where supported. Ambiguity in AI is a security and reliability risk.
    19. Default all AI systems to the most restrictive settings. Require explicit configuration to expand permissions or capabilities. If you set restrictive defaults, users are more likely to leave them in place; which is exactly what you want.
    20. Model the threats for all AI systems. Include AI-specific threats: prompt injection, data poisoning, training data extraction, adversarial inputs, model inversion, jailbreaking, denial-of-service, and agent misuse. Mitigate or eliminate all threats assessed as significant.
    21. Secure the training and data pipeline end-to-end Vector stores, embedding pipelines, and retrieval logic are a distinct attack surface. Poisoned documents retrieved at query time can manipulate agent behaviour invisibly.
    22. Verify the integrity of AI models and agents before deployment. Use model signing, checksums, or another integrity verification method to ensure models have not been tampered with. Immutable builds and verified deployments are the standard.
    23. Apply appropriate API security controls to all AI endpoints. Such as rate limiting, authentication, authorisation, input validation, and monitoring to ensure protection against traditional attacks and failures.
    24. Rate limit all AI agent actions. Nothing an AI agent does should be unlimited. Apply limits at every layer: API calls, tool invocations, file operations, external requests, and token consumption. Unlimited AI agents create unlimited risk.
    25. Perform all critical AI validation and decision-making on the server side. Client-side safety controls can be intercepted or bypassed. Trust only what happens on systems you control.
    26. Keep AI models, frameworks, SDKs, and dependencies up to date. Outdated components are a known risk, especially in fast-moving AI ecosystems. Where possible, automate updates and patching. Slow release and update processes are a serious organisational risk and should be treated as a priority for improvement. Temper this against not updating too soon as supply chain attacks exploit organisations that update uncritically.
    27. Minimal agent footprint. At each step of execution, an agent should request only what it needs now, release access when done, and avoid accumulating permissions or retaining sensitive data across steps. This is least privilege applied dynamically, per action.
    28. Enable strict safety and content settings in all AI frameworks and platforms. If a framework or API offers a strict mode, safety classifier, or content filter then turn it on. These are not enabled by default in all systems. Check, and enable them explicitly.
    29. Implement anomaly detection for all AI agent interactions. Monitor for abuse patterns, eg unusually high request volumes, adversarial inputs, or attempts to probe the AI’s limits. Implement defences against prompt flooding, token exhaustion attacks, systematic jailbreak attempts, and bot-driven misuse. Log all interactions. Alert on thresholds that suggest an attack may be in progress.
    30. Retest AI systems for safety regressions after every update or change. Safety testing is not a one-time activity. Build automated red-teaming, adversarial testing, output classifiers, and bias evaluation tools, and prompt injection tests into your CI/CD pipeline where possible.
    31. Protect all AI infrastructure comprehensively. This includes the model endpoints, repositories, vector databases, embedding stores, training pipelines, and supporting systems. All must be hardened, monitored, logged, patched, and tested for security.
    32. Apply security design principles to AI systems. Where possible, enforce least privilege (limit what the AI can access and do), zero trust (never assume an AI-to-AI call is safe), defence in depth (layer multiple controls), and attack surface reduction (limit the AI’s reach to only what is required). Do not rely on a single guardrail. Apply stricter principles for higher-risk AI systems.
    33. Control what files and data AI agents can access, read, write, or delete. Use strict access controls. Treat all files produced by or passed through an AI agent as potentially untrusted. Ensure important data is backed up, stored encrypted, and protected with monitored access controls.
    34. Prevent race conditions in AI agent workflows. When multiple agents or processes interact with shared state or external systems, use proper coordination, locking, and sequencing to prevent conflicts. Most modern AI orchestration frameworks provide tools for this.
    35. Use established identity, authentication, and access control systems for all AI agent interactions. Do not build your own AI authorisation logic from scratch. Existing solutions are well-tested. Writing custom access control for AI agents introduces serious risk.
    36. Apply encrypted, certificate-validated connections. Use HTTPS and validated certificates or similar for every AI API call. Follow your organisation’s cryptographic standards, or those of OWASP, NIST, or your relevant government body. Choose the strictest applicable standard and check that certificates are valid and from the expected host. Do not connect to unverified AI endpoints. This integrity check prevents man-in-the-middle attacks on AI API calls.
    37. Follow a secure AI development lifecycle. Integrate safety and security activities at every stage: design, development, testing, deployment, and monitoring. If your organisation does not have a defined lifecylce, create one. Add safety activities to your existing SDLC where possible.
    38. Select AI frameworks and platforms with strong, built-in safety features and guardrails. Do not write your own safety controls from scratch when established solutions exist. Always use a supported, up-to-date version of any AI framework or SDK. Then avoid bypassing content moderation, or undermining output constraints. If the defaults do not fit your use case, work with your security team, do not simply disable them.
    39. Manage agent state explicitly and consistently. Establish a standard approach to manage context windows or agent memory. Inconsistent context management introduces subtle bugs and security risks that are difficult to detect and diagnose. Initialise AI agent state explicitly before use. Do not rely on implicit defaults or assume prior state is clean. Uninitialised or stale state in AI agents produces unpredictable and potentially unsafe behaviour.
    40. Do not use real user data for AI development, testing, or fine-tuning without proper anonymisation and approval. Raw production data in non-production AI environments is a serious privacy and compliance risk. Use purpose-built anonymisation or synthetic data generation tools, not home-grown masking scripts.
    41. Protect all AI API keys and administrative accounts with multi-factor authentication. Use long, unique, complex credentials for every account with elevated AI system access. Use a password manager. Reset credentials immediately if you suspect a breach. Rotate keys regularly.
    42. Offer strong authentication to users of AI-powered systems. Provide MFA where possible. Implement defences against credential stuffing, supply chain, and malware attacks on their workspace. Carefully log all access and alert on suspicious patterns. Geoblock their access limiting exposure if access is breached.
    43. Create response plans for AI failure scenarios. Assume a failure will happen and prepare in advance. Practice these scenarios where possible. Ensure your AI systems are included in business continuity and disaster recovery planning.
    44. Audit AI system configurations and permissions at least annually. Review what models are in use, purpose, what access they have, what data they can reach, and whether their configurations remain appropriate.
    45. Manage AI session tokens and credentials securely. Apply the same standards as secure cookie management: limit scope, set expiry, enforce secure transport, and never expose credentials unnecessarily.
    46. Protect proprietary AI prompts and agent logic from unauthorised exposure. System instructions, reasoning chains, and agent architectures can contain competitive intelligence and safety-critical logic. Where appropriate, treat them with the same care as source code.
    47. Protect AI interfaces from cross-origin and cross-site abuse. Apply the same cross-site request forgery (CSRF) and cross-origin resource sharing (CORS) protections to AI endpoints as you would to any web application. These protections are not always on by default.
    48. Be cautious about AI model version updates and backwards compatibility. A model update may silently change safety behaviours, output formats, or reasoning patterns. Balance usability against the risks a new version introduces. Ideally fix model choice. Test thoroughly on updates and document your decisions.
    49. Build reusable, tested AI safety controls rather than reinventing them. If you build a safety control once, make it reusable and test it thoroughly before applying it widely. Continue testing it over time. When a safety bug is found, update every system that uses it.
    50. Use robust version controls. Keep the prompts, markdown files, infrastructure-as-code, and other code under strict version controls. This enables linting, CI/CD inline testing, and comparing against vendor or third-party model changes. Which means an easier time identifying changes when issues arise, faster rollback or changes.
    51. Make AI system behaviour auditable and explainable. Document prompts, system instructions, and agent decision logic. Add comments explaining safety controls. Readable, auditable AI systems are easier to test, easier to maintain, and faster to diagnose when something goes wrong. Clear documentation helps with safety testing, incident response, and onboarding. Build the behaviour with the future engineers in mind, which is most likely you, but could also be AI.
    52. Use immutable context and state in AI agent workflows wherever possible. Mutable shared state between agent steps creates unpredictable behaviour and increases the risk of manipulation. Design for immutability and explicit state transitions.
    53. Know and comply with all AI-specific regulations that apply to your systems. This includes laws, regulations, and standards such as the EU AI Act, local privacy laws, and sector-specific requirements. Ask your legal and security teams to verify your obligations.
    54. Maintain a current inventory of all AI models, agents, tools, and dependencies. Include where each model is deployed, version, how it is accessed, who owns it, responsible persons, and where its documentation lives. Audit this inventory at least annually.
    55. Adopt an AI safety framework if your organisation does not already use one. Frameworks such as the NIST AI Risk Management Framework, OWASP Top 10 for LLMs, or MITRE ATLAS provide structured guidance. If your organisation has adopted one, follow it.
    56. Decommission old AI models and agents carefully and intentionally. Remove API access, revoke credentials, archive documentation, and update your inventory. Document the decommission process (ideally before it’s in production). Abandoned AI systems with live access are a serious risk.
    57. Plan for AI model updates that can be applied smoothly and safely. If your users or downstream systems depend on your AI’s behaviour, ensure updates can be rolled out, tested, and rolled back without disruption.
    58. Provide a hardening guide for any AI system you deploy that others will operate. If an administrator or end user is responsible for configuring or operating your AI system, give them clear guidance on how to do so securely. Assume the end user won’t read it and act accordingly.
    59. Design AI systems to be easy to integrate with securely. If your AI system forces developers to adopt insecure workarounds to integrate it, the workarounds will become the norm. Secure integration should be the path of least resistance.
    60. Prioritise usability when designing AI safety controls. Controls that are difficult to use will be worked around. Test guardrails for usability just as you would any other feature. Good safety design and good user experience are not opposites.
    61. Verify and document user consent before AI systems process user data. Never assume consent. Always ask. Store consent records in your system of record. This is both an ethical obligation and, in most jurisdictions, a legal requirement.
    62. If you find a reproducible failure mode, safety or security bug in an AI model or framework, report it. Mistakes happen. Hiding or exploiting the mistakes is to be avoided. Be willing to share how your AI system failed and what you learned. Responsible disclosure benefits the entire AI development community. Follow the vendor’s responsible disclosure process.
    63. Treat AI safety warnings and alerts as errors and fix them. A safety warning that is ignored is a vulnerability waiting to be exploited. If your AI framework, monitoring system, or testing tool raises a warning, address it with the same urgency as a compiler error.
    64. If your AI system runs on physical hardware, include physical security. Securing the software layer alone is not sufficient. Physical access to AI infrastructure must also be controlled, monitored, and documented.
    65. Maintain a list of dangerous AI coding patterns to avoid. Include patterns such as direct prompt construction from user input, unrestricted tool access, unvalidated AI output passed to system calls, and disabled safety filters. Check your codebase for these patterns regularly. Automate detection in your linter or tooling where possible.
    66. Follow your organisation’s AI safety guidelines and approved patterns. Use the AI safety frameworks and approved integrations your organisation has established. If you don’t have one, write one. If a business requirement prevents this, work with your business to identify alternatives, document it, and notify the teams responsible. They may require a formal exception.
  • How Algorithmic Capability Is Reshaping the Way Professionals Learn

    AI has not just accelerated execution; it has displaced the very work through which professionals traditionally built their judgement. This is the central challenge of professional development in an AI-augmented world. How organisations respond to that displacement will determine whether the next generation of professionals can actually do the jobs they’re being hired for.

    For example, a law firm partner reviews a contract drafted in seconds by AI. She makes two small corrections, then sends it back to her junior associate with a simple instruction: “Find what I fixed”. The associate stares at the document for forty minutes. That exercise – not the original drafting – is now the training.

    When AI Outperforms the Average Professional

    Advanced AI models have moved beyond novelty tools into genuine performance competitors. According to OpenAI’s evaluation of GPT-5 across more than 40 occupations – including law, logistics, sales, and engineering – the model is now comparable to or better than expert-level professionals in roughly half of tested tasks. On a benchmark comprising PhD-level science questions, GPT-5 pro achieves 88.4% accuracy, surpassing the 69.7% average recorded by recruited PhD experts.

    The practical consequence is a new class of professional risk. In the past, average output from a junior professional was clearly distinguishable from expert output. Today, AI generates work that is sophisticated enough that only a seasoned practitioner can reliably identify its gaps such as subtle errors in reasoning, missing context, unstated assumptions that quietly undermine an otherwise polished product. The challenge has shifted from whether AI can perform a task to whether the humans reviewing its output have developed the depth to notice when it goes wrong.

    The Junior Gap

    Professional expertise is not taught but is accumulated. Historically, junior professionals developed their judgement through the grind of execution: drafting contracts, conducting research, writing initial code. The output was the product, but the process was the training. Through repetition, exposure to edge cases, and regular feedback on mistakes, juniors built the pattern recognition and systemic understanding that eventually makes an expert.

    AI disrupts this pipeline at its foundation. If AI produces the first draft more efficiently and cost-effectively than a junior professional, the junior no longer gets the drafting repetitions. And without those repetitions, the developmental progression from novice to capable practitioner is severely compressed or eliminated entirely.

    The result is what researchers are calling the “junior gap”: a cohort of professionals who are technically proficient at directing AI tools but lack the foundational judgement to critically evaluate and refine what those tools produce. They can prompt well. They cannot always tell when the output is wrong.

    This is a systems problem with an identifiable feedback loop. Organisations optimise for efficiency, AI handles execution, junior exposure disappears, and years later those organisations discover they have capable AI operators but insufficient numbers of senior professionals with the depth to oversee complex matters. One mistake, the decision to eliminate junior execution roles entirely, becomes two, three, and eventually systemic.

    Building Judgement Deliberately

    Addressing this requires more than adding AI-literacy modules to onboarding programmes. It requires redesigning how professional capability is formed from the ground up. Several approaches are gaining traction.

    The Residency Model

    Some law schools are already pioneering what amounts to a professional residency: guaranteeing students nearly a full year of full-time supervised work experience before graduation, with structured mentorship built into the curriculum. This model, long established in medicine, acknowledges that judgement cannot be transmitted through instruction alone. It requires supervised exposure to real decisions, real consequences, and real feedback from experienced practitioners. The volume of exposure matters. So does the quality of supervision.

    Decomposing Judgement into Teachable Components

    To teach judgement, organisations first need to understand what judgement actually comprises. Research into professional expertise points to several distinct micro-skills that together constitute sound professional thinking:

    Micro-skillWhat it means in practice
    Pattern RecognitionIdentifying when a new situation resembles a past one and when it only superficially appears to
    Strategic CalibrationMatching effort, precision, and risk tolerance to the actual stakes of a given matter
    Reasoning Through UncertaintyMaking defensible decisions when information is incomplete or contradictory
    Source EvaluationAssessing credibility, relevance, and the limits of the evidence at hand
    Ethical JudgementNavigating conflicts, spotting grey areas, and maintaining integrity under pressure

    Breaking judgement into these components allows organisations to design training and experiences that deliberately targets each one, rather than hoping it develops through general exposure.

    Making Judgement Part of the Workflow

    Organisations cannot outsource judgement development to an annual training programme. It needs to be embedded in how daily work is done. Three mechanisms are proving effective:

    Verification logs: Requiring junior professionals to document their reasoning when they accept, modify, or reject AI-generated content. The discipline of writing down why a decision was made builds the habit of critical reflection and creates a record that supervisors can review and challenge.

    Attack-the-draft drills: Explicitly tasking juniors with finding weaknesses in AI output, rather than simply polishing it. The question shifts from “is this good enough?” to “what is this missing?” That shift is the beginning of genuine critical thinking.

    Preserving slow work: Deliberately retaining certain manual tasks suchn as citation chaining, in-depth primary source research, first-principles analysis, not because they are efficient, but because they are formative. The inefficiency is the point.

    Phasing AI Access to Protect Critical Thinking

    Introducing AI tools before foundational analytical skills are established risks short-circuiting the development process. A phased approach provides a more reliable pathway:

    In the foundational phase, professionals complete core reasoning exercises without AI assistance, building the analytical muscle that everything else depends on. In the assisted phase, AI is available as a research tool, but verification against authoritative sources is mandatory to train the habit of never accepting AI output without scrutiny. In the integrated phase, professionals work within real AI-assisted workflows under supervision, with continuous feedback connecting their judgement to outcomes.

    What Organisations Must Decide

    The economic case for replacing junior execution with AI is straightforward. The long-term cost of doing so without compensating mechanisms for capability development is considerably less visible, until it becomes a crisis.

    Organisations that recognise this early have several options. Some will accept deliberate short-term inefficiency, retaining junior roles precisely because the developmental value justifies the cost. Others will build capability formation into their value proposition, differentiating themselves not just by the quality of their AI-assisted output but by the quality of the professionals they produce. A third path uses AI itself as the training environment: generating realistic scenarios, evaluating responses, and compressing years of exposure into a shorter, structured timeline.

    None of these paths are easy. All of them require treating capability formation as a strategic investment rather than a byproduct of the work.

    The Urgent Choice

    AI will not eliminate the need for professional judgement. It will raise the bar for what that judgement needs to encompass. The professionals who thrive will be those who can do what AI cannot: identify where it has gone wrong, understand why, and know what to do about it.

    That capability does not emerge from tool proficiency. It is earned through deliberate, structured, supervised experience. The same way it has always been earned. The organisations that recognise this and redesign their development models accordingly will produce professionals capable of the work ahead. Those that simply adopt AI and hope capability follows will find, eventually, that it does not.

    Capability is still earned. The challenge now is ensuring the pathways to earning it remain intact.

  • Autonomous IT with AI-Driven Self-Healing

    Active monitoring was a genuine breakthrough. When systems could alert a human to a problem before it became a crisis, response times dropped from hours to minutes and organisations got ahead of failure for the first time. That model worked – until it didn’t.

    As AI-accelerated development compresses release cycles from months to minutes, the human review step that was once an asset has become the bottleneck. An engineer acknowledging an alert, investigating a dashboard, and manually applying a fix is a workflow built for a different pace of software. The next evolution isn’t faster humans. It’s removing humans from the incident loop.

    Autonomous self-healing infrastructure doesn’t just detect problems. It diagnoses them, generates a fix, validates it, and deploys it, all before a user notices anything is wrong.

    Why Active Monitoring Has Hit Its Ceiling

    Active monitoring solved the reactive problem. It didn’t solve the latency problem.

    When a system behaviour shifts, the human-in-the-loop model still requires acknowledgment, investigation, and manual intervention. At modern deployment speeds, such as where a single AI-generated commit can introduce and surface a regression within seconds, that sequence is too slow.

    The shift isn’t about humans being unreliable. It’s about systems operating at a pace that humans were never designed to match.

    Evolution StagePrimary ActorAction TypeResponse Time
    ReactiveHuman OperatorBreak-FixHours to Days
    ActiveMonitoring Tools + HumanProactive InterventionMinutes to Hours
    AutonomousAI AgentsSelf-Healing & RemediationMilliseconds to Seconds

    The monitoring system is no longer a passive observer sending notifications to a human queue. In the autonomous model, it is an active participant in the codebase itself.

    Self-Healing Goes Further Than You Think

    Automatically restarting a crashed container is table stakes. Modern autonomous systems use LLM-powered agents to perform complex, stateful operations that would previously require a skilled engineer.

    Autonomous Code Refactoring

    When an autonomous system detects a performance regression, for example a slow memory leak, an inefficient database query, a sub-optimal API call, it files a ticket that initiates a closed-loop refactoring cycle:

    1. Identification: The AI isolates the specific block of code responsible for the degradation.
    2. Synthesis: It generates a refactored version that improves performance while maintaining functional parity.
    3. Validation: The new code runs through an automated test suite and a shadow-deployment environment to confirm no regressions.
    4. Deployment: The optimised code replaces the original in production, resolving the issue before it reaches a user.

    The process that once took days of developer time – reproduce, diagnose, fix, test, deploy – completes in seconds.

    Real-Time Security Response

    Manual security response has always been a latency problem. By the time a human reviews an alert, correlates it with threat intelligence, and applies a patch, the window of exposure is measured in hours. With Black-Hat hackers using AI to exploit vulnerabilites at scale, this is way too long.

    Autonomous security agents operate differently. When an anomalous traffic pattern emerges, an adaptive firewall rewrites its own policies to neutralise the threat in real time. When a new vulnerability is published, the system fetches the CVE data, generates a patch specific to its environment, and applies it, without a change request or an on-call rotation.

    The incentive structure is important here: security teams historically got rewarded for detecting threats, not for preventing them at speed. Autonomous systems change what gets rewarded.

    The Infrastructure Underneath: Agentic Observability

    Traditional observability was designed for human eyes via dashboards, visualisations, alert thresholds that a person could interpret and act on. Agentic observability is designed for an entirely different consumer.

    The shift, as LogicMonitor’s Karthik SJ framed it in their 2026 AI outlook, is from dashboards for humans to APIs for agents. The system’s primary consumer is now an AI capable of taking action on what it sees.

    This architecture depends on two foundations:

    High-density telemetry streams that give agents the correlated context to reason across infrastructure, application, and user-experience layers simultaneously, not just a single metric in isolation.

    Digital twin environments where AI can simulate “what-if” scenarios and validate refactored code before touching the live system. The twin is where the AI falls to its level of training; the production system is where that training pays off.

    LogicMonitor’s 2026 Observability & AI Outlook found that 44% of IT leaders are actively working toward automated remediation and self-healing systems. Only 4% have fully operationalised AI across their IT operations. The gap between ambition and zero-touch operations is where most organisations currently sit.

    The Real Barrier Is Governance, Not Technology

    The technology for autonomous infrastructure exists. The harder problem is organisational.

    The trust gap is the most immediate challenge. Asking an engineering team to allow an AI to refactor production code without human review requires a fundamentally different relationship with automation than most organisations have built. Trust gets built through consistent behaviour, not promises. Teams need to see exactly what the AI changed, why it changed it, and what the measured impact was.

    Governance becomes genuinely complex when the system is continuously evolving its own logic. Compliance frameworks were designed for human-authored code under change control processes. Autonomous systems that modify themselves in seconds don’t fit that model. Organisations need policy-driven guardrails that let AI act within defined boundaries while maintaining auditability, for example approval thresholds, scope limits, rollback triggers.

    Explainability is non-negotiable. A black-box system that can’t explain its reasoning erodes trust and will eventually be switched off after the first incident it can’t account for. Every autonomous action needs a clear record: the signal that triggered it, the diagnosis produced, the change applied, and the outcome measured. Without that trail, the audit review that replaces human oversight becomes meaningless.

    These are solvable problems. They’re engineering and governance problems, and the organisations building toward autonomous IT are treating them as such. The philosophical principle applies: you don’t rise to the occasion in a production incident, you fall to your level of systems thinking and preparation.

    What Organisations Should Do Now

    The organisations that will operate autonomously in two years are not waiting for the technology to mature. They’re building the conditions for it.

    Start by unifying observability data. AI cannot act reliably on fragmented telemetry. Correlating metrics, logs, traces, and user-experience data into a single platform is the prerequisite, not the finish line.

    Define the boundaries of autonomous action before deploying it. Autonomous remediation should start narrow: restart a known-bad service, roll back a deployment that fails a health check, block a traffic pattern that matches a known threat signature. Document what falls inside and outside those boundaries. Expand the boundaries as trust is earned through a clear audit record and observed behaviour.

    Build the audit trail from day one. If the autonomous system can’t explain every action it takes, the governance conversation becomes much harder. Explainability isn’t a feature to add later. It’s the mechanism through which human oversight of a non-human system actually functions.

    The Self-Driving Codebase

    The goal of resilient infrastructure was never to eliminate failure. It was to reduce the time between a failure occurring and it being resolved to the point where it stops mattering.

    Autonomous self-healing systems achieve that by removing the human latency from the loop. The monitoring-diagnosis-remediation cycle that once took hours, then minutes, now takes milliseconds. The software organisations deploy is no longer a static artefact but a starting point for a system that protects and optimises itself continuously.

    The engineering challenge shifts. Instead of writing code that doesn’t break, the work becomes designing systems with the observability, the guardrails, and the trust architecture that allow AI to fix it when it does.

    That’s a different kind of engineering. And for most organisations, it starts with the observability foundation they’ve either built or haven’t.

    References

    LogicMonitor (2025) 5 Observability & AI Trends Making Way for an Autonomous IT Reality in 2026. https://www.logicmonitor.com/blog/observability-ai-trends-2026