Author: admin

  • Your Marketing Copy as Conduct Control

    A financial services firm’s website is much more than a brand asset that happens to mention products. It is a conduct record. Every headline, testimonial, calculator, performance chart and call to action can create a representation about the firm, its services, internal controls, a financial product, a likely outcome, or the level of risk involved. The relevant regulatory question is not whether each sentence is technically defensible in isolation, but whether the communication, viewed as a whole and received by the audience likely to see it, creates a false, misleading or deceptive impression.

    ASIC’s updated Regulatory Guide 234, Advertising financial products and services (including credit), sets the standard. RG 234 is directed to promoters of financial products and financial advice services, and to publishers of advertising. Its core requirements are straightforward: statements must be true, accurate and capable of being substantiated; predictions about future returns or risk require reasonable grounds; and an intention not to mislead does not prevent a communication from being misleading. The overall impression matters as much as the individual words. The audience reached matters as much as the audience intended.

    That framing has a direct compliance implication and marketing review should be treated as a front-end conduct control, integrated into the AFSL holder’s broader system of supervision, monitoring, record-keeping and accountability. A polished approval workflow is not enough if published claims remain unbalanced, unsupported or impossible to reconstruct after the fact.

    ASIC’s June 2026 Update Changes What Firms Must Review

    On 9 June 2026, ASIC reissued RG 234, the first substantial rework of the document since it was introduced in 2012. The revision drew on more than a decade of enforcement and regulatory action and followed a public consultation that ran from November 2025 to January 2026. The biggest structural change is the absorption of Regulatory Guide 53, The use of past performance in promotional material, into RG 234. RG 234 is now the single reference point for ASIC’s advertising expectations, with no stated transition period. Which means the revised guidance is now in effect.

    The update also added AI-generated content, AI specific claims, and a greenwashing enforcement example, expanded guidance on the distinction between required warnings and general disclaimers, clarified the performance-period comparison rules for products with less than five years of history, and added further commentary on what constitutes an “ordinary and reasonable person” when assessing how an advertisement lands.

    The practical message is that the firm must ensure that any communication is fair, balanced, clear, accurate and supported by evidence, with qualifications presented in a way that an ordinary member of the intended audience can notice and understand. A disclaimer cannot reliably fix a dominant headline that communicates an inconsistent promise.

    A Claim-Level Benchmark Makes Regulatory Exposure Visible

    A useful approach to marketing review does not attempt to declare a breach or confirm that copy is free from risk. It identifies individual claims, maps them to relevant risk themes and assigns a triage status. A report of this type might present a result such as 79 out of 100 across 42 claims, with 23 green, 18 yellow and 1 red. The value lies less in the numbers than in the audit trail behind it detailing the exact wording, the page location, the issue classification, the evidence required, the accountable owner and the remediation status.

    As a concrete example, a sentence such as “It is this level of total care that ensures expert advice to get you where you need to be” may warrant a red classification because “ensures” may convey an absolute guarantee about both the quality of advice and the outcome.

    A report of this type should be treated as an internal screening tool, not as an ASIC-issued finding or an independently verified legal conclusion. Its governance value is that it converts a large body of website text into a prioritised remediation queue.

    Benchmark resultGovernance interpretationTypical next step
    GreenNo issue identified against the benchmark criteria, subject to supporting evidence and ongoing change control.Retain evidence and include in periodic re-review.
    YellowWording, context, qualification, substantiation or audience fit may create elevated risk.Assign an owner, obtain evidence, revise wording or add prominent context, then re-approve.
    RedAn absolute, guarantee, unsupported prediction, material omission or other high-risk element may require urgent attention.Escalate to compliance or legal; consider pausing or removing the content; document the decision and check related channels.

    The thresholds for green, yellow and red should be written down and calibrated to the firm’s risk appetite. A firm that treats a yellow finding as an acceptable residual is making a deliberate decision. A firm that discovers its yellow findings are proliferating across channels is detecting a systemic pattern. The score is a diagnostic, not a verdict and shows where the lines sit and how often the firm crosses them. What the firm does next is the real test of its governance.

    Three Failure Patterns Appear Repeatedly in AFSL Marketing

    Misleading promotional claims understate what the firm is promising

    Absolute words such as “guaranteed”, “ensures”, “always”, “never”, “risk-free”, “best” or “no downside” deserve immediate scrutiny. They may be acceptable in a narrow factual context. In financial services marketing, however, they can imply certainty about an investment outcome, the quality of advice, the suitability of a strategy or the absence of risk.

    The review should test the net impression the claim creates. “Helping clients pursue their objectives” is materially different from “ensuring clients achieve their objectives”. “We provide tailored advice” is different from “our advice will get you where you need to be”. The second formulation in each pair may require proof the firm cannot realistically provide and may overstate what advice can achieve.

    A sound review examines adjacent text, page design, imagery, testimonials and calls to action together. A modest qualification buried below a bold guarantee may not change the overall impression. A claim can also mislead through omission if a reasonable consumer would need material information to understand its significance.

    Disclaimers that are present but ineffective provide false comfort

    Disclaimers are not a universal safe harbour. They can be important where a communication needs context, but they must be specific, readable, proximate and consistent with the main message to carry any protective weight. A generic statement such as “past performance is not indicative of future performance” does not, by itself, explain fees, volatility, relevant time periods, portfolio composition, assumptions, conflicts or the possibility of loss where those matters are material to the impression created.

    The review should distinguish between a missing disclaimer and an ineffective one. A disclaimer may be present yet ineffective if it is hidden behind a link, placed far from the claim it qualifies, displayed in unreadable type, contradicted by the headline, or drafted in language the audience is unlikely to understand. The firm should record not only the disclaimer text but also its location, formatting, presentation across device types and relationship to the specific claim. Any advice business copy should also be checked against the boundary between factual information, general advice, personal advice and a promotional invitation to engage.

    Performance statements create an incomplete picture

    Performance content carries a high risk of presenting a partial account. A favourable return figure may be technically correct yet misleading if the period, benchmark, fees, tax treatment, volatility, drawdowns, assumptions or risk of loss are unclear. Cherry-picked periods and isolated success stories can produce a stronger impression than the underlying evidence supports.

    The June 2026 update consolidates all past-performance advertising guidance into RG 234. A practical review should ask: what exactly is being measured, and over what period? Is the result gross or net of fees? Is the comparison like-for-like? Are negative or less favourable periods relevant to understanding the claim? Is the benchmark appropriate? Are forecasts or projections clearly distinguished from historical results? Can the firm reproduce the data and methodology that generated the number?

    The same discipline applies to testimonials and case studies. A statement such as “we helped this client retire early” may imply a typical or repeatable outcome even if it describes one client’s experience. The firm should assess whether the example is representative, whether material qualifications are needed, and whether the audience could mistake an individual result for a promise.

    A Defensible Claim-Review Process Starts With a Complete Inventory

    A firm that reviews its marketing reactively, one page at a time, when someone raises a concern, is not running a control but responding to incidents. A defensible review process starts with a complete inventory of every public-facing channel: the main website, landing pages, calculators, downloadable guides, newsletters, paid search copy, social channels, adviser biographies, webinars, podcasts, third-party profiles and any influencer content. ASIC monitors for misleading or deceptive representations and unlicensed financial services, which supports treating digital channels as part of the controlled perimeter rather than as informal exceptions.

    Each piece of content should then be decomposed into claims, that might not be limited to a sentence. It may be a number, superlative, visual comparison, implied promise, omission, testimonial or a combination of headline and design. The reviewer should capture the precise wording, URL or channel, screenshot or archived version, date observed, intended audience, product or service involved, risk theme, evidence required, decision and owner.

    Review questionEvidence to retain
    What does the audience reasonably take away?The full page or post, including headline, imagery, buttons and nearby qualifications.
    Is the claim factual, predictive, comparative or opinion-based?Source documents, calculations, benchmark definitions, assumptions and approval records.
    Can the firm substantiate it now?Dated evidence, data lineage, research, client-file samples where appropriate and sign-off.
    Is important context prominent and understandable?Screenshots across desktop and mobile, disclaimer placement and readability checks.
    Does the claim remain accurate over time?Review date, expiry trigger, owner, change log and monitoring result.
    Does the communication create advice, distribution or target-market issues?Product scope, audience analysis, advice classification and relevant disclosure and DDO assessment.

    Classification rules should be written down before review begins. A red rule might include an absolute guarantee, an unsupported future-return prediction, a material misstatement, a claim contradicted by available evidence, or a missing qualification that would change the audience’s likely decision. Yellow might include unclear scope, weak substantiation, poor disclaimer prominence, stale data, ambiguous comparisons or an outcome-oriented testimonial. Green should mean “no issue identified under the tested criteria”, not “approved forever”.

    The Pattern of Findings Tells You More Than the Score

    The most valuable question after a review is not “What was our score?” It is “What does the pattern of findings say about our controls?” If most yellow items relate to missing context, the problem may sit in the copy template or brand guidelines. If red items are concentrated in adviser biographies, the training and approval process may be underperforming. If stale performance numbers recur, ownership and review triggers may be unclear. If the same issue appears on the website and across social channels, the content management process may lack a single source of truth.

    A governance committee or responsible manager should review aggregate results alongside remediation ageing, repeat findings, approval exceptions, complaints, incidents and regulatory changes. Management information should distinguish between newly detected issues, accepted residual risk, items awaiting evidence and items closed after independent verification. A numerical score should never conceal a serious single issue: one unqualified guarantee can warrant more attention than many low-risk wording observations.

    Control areaWhat the firm should be able to demonstrate
    OwnershipEvery public claim has a business owner and a compliance escalation path.
    Pre-publication reviewHigher-risk claims receive documented compliance or legal review before release.
    SubstantiationEvidence is current, traceable and sufficient for the specific wording and audience.
    Disclosures and qualificationsMaterial context appears prominently and is tested in the actual publication format.
    Change managementCopy is re-reviewed when products, fees, performance, law, guidance or audience changes.
    SurveillancePublished content is periodically scanned, sampled and compared with the approved version.
    RemediationRed and yellow issues have deadlines, accountable owners, decisions and closure evidence.
    LearningRepeat findings feed back into templates, training, controls and risk appetite.

    Red and Yellow Findings Require Concrete Action, Not Just Documentation

    A red finding should trigger prompt containment. The firm should preserve the relevant version of the content, confirm whether it is still live, assess the channels and audiences affected, and escalate under its incident and breach-reporting framework where appropriate. It should then decide whether to remove, pause, correct or qualify the communication, documenting the rationale and approval.

    For yellow findings, the response should be proportionate and concrete. The firm may need to obtain evidence, narrow the claim, replace certainty language, add context, correct a comparison, improve disclaimer prominence or introduce an expiry date. The revised copy should be tested as a whole, because adding words at the bottom of a page may not fix the impression created at the top.

    In both cases, remediation should include a look-back. Search for similar wording, related claims and syndicated versions across web pages, social posts, PDFs, email journeys and third-party channels. A review is most useful when it reveals a pattern that can be corrected systematically, not a single sentence that is edited in isolation.

    Marketing Review Is a Control System, Not a One-Time Exercise

    ASIC’s advertising guidance places the emphasis on substance and overall impression. For AFSL firms, that means marketing-copy review must be integrated with evidence management, approvals, disclosure controls, monitoring and remediation. It cannot be treated as a one-off brand exercise or an annual tidy-up.

    A benchmark result, whether 78 out of 100 or any other score, provides a useful starting point for management discussion. Its real value is the claim-level detail underneath: which words create risk, what evidence is missing, how prominent the qualification is, who owns the fix and whether the control environment learns from the result. Used that way, marketing review becomes more than a surveillance exercise. It becomes a practical test of whether the firm’s governance and review processes are producing the outcomes the firm expects, before a regulator asks the same question.

    Contact us to explore how to automate your marketing review process.

    References

    [1] ASIC, RG 234 Advertising financial products and services (including credit), issued 9 June 2026. https://www.asic.gov.au/regulatory-resources/find-a-document/regulatory-guides/rg-234-advertising-financial-products-and-services-including-credit/

    [2] ASIC, INFO 269 Discussing financial products and services online. https://www.asic.gov.au/regulatory-resources/financial-services/giving-financial-product-advice/discussing-financial-products-and-services-online

  • Building AI in a Financial Advice Business: Design the System First

    The most reliable path to AI is to build a clean operating system for the business, then use AI and automation to make that system faster, clearer, and easier to govern.

    For most financial advice businesses, AI looks like a fast productivity win. It can read client documents, summarise meetings, prepare drafts, answer internal questions, and cut repetitive administration. The temptation is to pick a tool, connect some files, and let the team experiment.

    That is rarely the right starting point. If client information is spread across spreadsheets, email, task boards, shared drives, and specialist platforms, AI will not remove the complexity but increase it. It may generate polished language or save time on a specific task, but it cannot determine which record is authoritative, whether a document is current, who owns the next decision, or whether the required evidence has been captured and with each added tool complexity, mistakes, and gaps increases.

    The more durable approach treats AI implementation as a systems-design programme. First, design the operating environment: its workflows, data, ownership, permissions, and records. Then introduce AI on top of that clean foundation. The result is an implementation that improves efficiency without sacrificing the consistency, traceability, and professional accountability that advice businesses require.

    ASIC’s October 2024 report, Beware the Gap, reviewed 624 AI use cases across 23 Australian financial services licensees and found that AI adoption is outpacing the governance frameworks meant to manage it. Nearly half of licensees had no policies addressing consumer fairness or bias. ASIC’s conclusion was direct: build the governance before you build the capability.

    This six-step sequence provides a practical roadmap to build the governance, as you build the . The order is deliberate and prevents the practice from automating disconnected work or giving AI access to poorly governed data.

    StageWhat it createsWhat it prevents
    1. Map every workflowA shared picture of how work actually happens.Digitising an imagined process rather than the real one.
    2. Consolidate the stackA purposeful technology backbone with fewer duplicate tools.Fragmented client records and repeated manual entry.
    3. Design the source of truthOne governed model of clients, entities, advice, documents, and controls.Competing versions of the same information.
    4. Move the team into one homeGenuine adoption of the central system before automation.Building automations around poor data and shadow processes.
    5. Add AI agentsUseful AI grounded in controlled, permissioned business data.Scaling stale, incomplete, or unauthorised information.
    6. Automate the backgroundContinuous handling of routine handoffs, reminders, and filing.Staff spending time chasing status updates and administrative tasks.

    Step 1: Map every workflow before touching anything

    Before writing code, buying software, or selecting an AI vendor, map the firm’s full operation. The aim is to understand the advice lifecycle as one connected system: how a prospect becomes a client, how data is collected, how advice is prepared and reviewed, how documents are issued, how client actions are implemented, and how ongoing review obligations are managed.

    The map should follow work from the initial trigger to the final outcome. It needs to show every handoff, decision, system touchpoint, data re-entry, approval, wait state, and exceptions. It should also show where important evidence is created, stored, amended, or lost.

    The most important practical rule: walk each process with the person actually doing the work. The practice principal may describe onboarding as a clear, linear sequence. The client services officer may reveal a different operational reality: information copied from a fact find into a spreadsheet, then entered into planning software, then checked against documents in a shared drive, with missing information chased by email. Both perspectives matter but the detailed operational view is the one that should shape the future system.

    Workflow questionWhat to identifyWhy it matters for later AI use
    What starts the work?Referral, client enquiry, review date, life event, adviser request, or compliance finding.Defines when and how a workflow, automation, or AI assistant may be invoked.
    Who performs the work?Adviser, client services officer, paraplanner, compliance manager, client, or external provider.Establishes accountability, permissions, and escalation paths.
    Which data is required?Client details, entities, tax information, superannuation data, portfolios, documents, and instructions.Defines the data scope and quality needed for reliable AI assistance.
    Where does data move?Planning systems, email, forms, spreadsheets, shared drives, and manual exports.Exposes duplicate entry points and data-reconciliation risks.
    Where is judgement applied?Suitability checks, document review, advice preparation, exception handling, and compliance sign-off.Identifies decisions where human responsibility must remain explicit.
    What proves completion?Approved documents, meeting records, client acceptance, completed tasks, or review evidence.Defines the records the central system must retain.

    The output is a prioritised workflow inventory. You will find issues but do not attempt to fix everything immediately. Start with the department or journey losing the most manual hours, generating the most rework, or creating the greatest operational risk. In many practices, client onboarding, review preparation, advice-document production, and post-meeting administration are sensible early candidates.

    Step 2: Consolidate the stack, absorb, keep, and kill

    Workflow mapping nearly always reveals an overlapping technology stack. Team members may initiate work through email, track it in Monday.com or Trello, maintain client information in spreadsheets, store documents in a shared drive, and then re-enter the same information into planning or accounting software. Individually, each tool may appear useful. Together, they obscure ownership and create significant manual coordination overhead.

    The next step is to sort the stack into three categories: absorb, keep, and kill.

    CategoryTypical examplesPractical action
    AbsorbSpreadsheets, Monday.com, Trello, form tools, internal wikis, disconnected task lists, and lightweight databases.Bring their work-management function and relevant records into the central system wherever practical.
    KeepCore financial-planning software, accounting packages, portfolio or custody systems, and other specialist products with material operational value.Retain these as specialist systems and connect them through controlled integrations.
    KillUnused subscriptions, duplicate platforms, unofficial tools, and applications holding isolated records.Preserve any necessary records, define an archive approach, and decommission them promptly.

    Consolidation does not require replacing every specialist system with one giant platform. Advice businesses will continue to need purpose-built tools for planning, portfolio management, accounting, and related functions. The aim is to create a central operating environment that coordinates work, holds the core relationships, manages permissions, and gives the team a reliable view of each client and process.

    A simple design question provides a useful test: where should a team member go to understand the current state of a client, task, document, review, or exception? If the answer is “it depends”, the stack is not ready to support AI effectively.

    Step 3: Build a single source of truth

    A central system is valuable only when its data model is clear. The third step is to define the core nouns of the practice and the relationships between them. This is the foundation of the single source of truth.

    For an advice business, those core nouns include clients, households, entities, advisers, portfolios, accounts, advice engagements, Statements of Advice, invoices, meeting notes, documents, approvals, obligations, exceptions, and remediation actions. Each item needs a unique identity, a clear owner, a defined lifecycle, appropriate access rules, and a history of material changes.

    The relationships matter as much as the nouns.

    RelationshipWhy it matters operationally
    Client → Household / EntityPrevents an incomplete view of the client relationship and associated structures.
    Client or Entity → Account / PortfolioConnects advice activity to the relevant financial record.
    Advice engagement → SOA → Approval → Client acceptanceCreates a traceable document and decision lifecycle.
    Task → Owner → Status → Due dateShows who is responsible for an action and whether it is progressing.
    Compliance review → Finding → RemediationTurns review work into owned, visible corrective action.
    Document → Client / Entity → Version → SourcePreserves context and reduces doubt about which document is current.

    The governing rule is this: if client information belongs in the central system, it must not also be maintained in an uncontrolled side spreadsheet. This does not prohibit reporting exports or specialist-system records. It means the practice must be able to identify the authoritative record, secure it, update it in one place, and trace material changes.

    This structure also makes governance easier, producing exactly what ASIC expects. When core relationships are explicit, the practice can show what is associated with a client, which document version was approved, who made a decision, and what follow-up work remains outstanding. Without this structure, AI will retrieve only fragments of the story.

    Step 4: Move the team into one home before adding anything else

    Only after the central system and its data model are ready should the practice migrate teams. Move one department at a time, starting with the function carrying the greatest manual load or experiencing the most rework.

    This is not simply a data-migration project but a change-management exercise. Each team needs views that match how it thinks about work. A client services officer may need a queue of missing information, documents awaiting filing, forthcoming review activity, and tasks blocked by a client response. An adviser may need a client-centric view of meetings, outstanding actions, portfolios, advice documents, and review dates. A compliance manager may need visibility of higher-risk cases, incomplete evidence, overdue reviews, and unresolved remediation work.

    The immediate objective is adoption. People should be able to complete their normal work from the central environment without maintaining a parallel process elsewhere. This requires clean data, useful views, clear ownership, sensible access permissions, and active support during the transition.

    Do not automate anything yet. Clean the data and make the central system the team’s home first.

    Premature automation is understandable when a team is already overloaded, but automation built on inconsistent records or unclear responsibility simply repeats poor process more quickly. Establish a stable operating rhythm first: staff trust the data, leaders can see the work, and exceptions are recorded rather than hidden. Only then should automation be introduced.

    Step 5: Add AI agents once the data is ready

    Once the practice has one working home, reliable data, defined records, and appropriate permissions, AI can become useful by augmenting and expanding the decision making. It can retrieve information from the governed environment, prepare drafts, identify missing information, and support routine work without requiring staff to search across disconnected systems.

    Early AI use cases should be narrow, high-value, and easy to supervise. The focus should be on reducing repetitive administrative effort and improving access to existing information, not on replacing accountable professional judgement.

    AI agentPractical roleRequired control
    Document-ingestion agentReads client PDFs and extracts relevant tax, superannuation, entity, or account data into a review queue.Retain the source file, show extracted fields for verification, and require authorised approval before changing records.
    Generation agentDrafts scopes of work, meeting minutes, client correspondence, and first-pass advice-document sections from approved templates and data.Treat output as a draft and preserve review, approval, versioning, and record-keeping steps.
    Operational QA agentLets authorised users query the practice database in plain English, identifying overdue reviews or incomplete client records.Return traceable results, respect permissions, and avoid unsupported conclusions.
    Workflow-assistance agentIdentifies missing evidence, suggests the next task, and prepares process-stage checklists.Recommend and route work. Never silently override workflow or accountability.

    Every agent should have a defined purpose, approved data scope, bounded actions, visible source traceability, an audit history, and a named human owner. Users should understand which sources informed an answer, what may be incomplete, and how to correct a record or escalate a concern.

    The line between assistance and accountability must remain clear. An agent can retrieve, draft, classify, summarise, check, and route. Decisions with material client, professional, legal, or compliance consequences must remain within appropriately authorised and reviewable human processes.

    Step 6: Automate the background so the foreground stays human

    Once the central workflows are established and AI is operating inside clear boundaries, automate the background work that keeps the practice organised. This is where the operating system eliminates low-value coordination without removing visibility or control.

    Examples include status updates, compliance reminder triggers, review scheduling, document filing, requests for missing information, task creation, approval routing, exception-expiry notices, and data-quality checks. These processes should run quietly in the background while making the need for human attention clear when a decision or exception arises.

    Background automationPractical outcome
    Status updates and routingCompletion of one task updates the relevant record and creates the next owned action.
    Compliance remindersResponsible staff are notified before a review, document, evidence item, or exception becomes overdue.
    Review schedulingReview work is created according to agreed service cadence, client circumstances, or engagement status.
    Document filingApproved records are linked to the correct client, entity, engagement, and process stage using structured rules.
    Data-quality checksMissing fields, duplicate identifiers, or inconsistent relationships enter a managed exception queue.
    Management visibilityLeaders see current workload, client pipeline, outstanding obligations, and operational bottlenecks.

    Each automation should be observable and governed. The practice should be able to see what ran, what changed, what failed, and who owns the resulting exception. Background processes reduce administrative burden. They must not become invisible black boxes.

    A useful by-product of working through this sequence is that AI governance develops naturally alongside it. Identification of the use cases, classifying data clearly, identifying the owner and where a human-in-the-loop is required, understanding what third-party vendor tools actually do, and establishing logging, validation, and audit trails are not separate governance tasks. They emerge from building the system correctly in the first place and allows producing effective Governance guidelines, policies, and procedures that regulators expect easy.

    The sequence works because daily utility drives adoption

    Systems change and software adoption is rarely a training problem. It is more often a daily-utility problem. Staff adopt systems that simplify their work, help them find the right information, and remove unnecessary follow-up. They resist systems that add another place to log activity or create obligations without making the work easier.

    This sequence changes the experience of work before it adds intelligence. First, it maps reality. Then it removes duplicate tools, establishes one source of truth, and moves teams into a practical home. Only after the operating environment is stable does it add AI and background automation.

    For a financial advice and other regulated businesses, this is the practical path to AI implementation: design the operating system, establish the data foundation, move the team into one home, then automate and augment with control. The result is not only a more modern technology stack. It is a more coherent, usable, and scalable practice. If your practice is ready to build this foundation, get in touch and we can help you design and implement each step.

  • Training People to Think With, Against, and Through the Machine

    The real AI skill is judgement

    Many organisations describe their AI training goals in deceptively simple terms: teach people how to use the tools. That usually means showing employees how to write a prompt, summarise a document, generate a presentation, or ask a chatbot for ideas.

    AI is capable of very good tactical responses. The day-to-day, line-by-line work a capable junior or mid-level employee would produce. Those capabilities matter, but they are only the entry point, because the more consequential skill is not producing an answer with AI. It is deciding whether the answer deserves confidence, identifying what it missed, exposing where it might fail, and improving it through deliberate interaction.

    There’s a long list of damaging failures where the AI produced a confident answer that turns out to be based on flawed assumptions or made up facts. Using some simple approaches can catch these mistakes before they reach a client or regulator while still allowing AI’s speed.

    The core idea is treating AI as a dynamic sparring partner: a system that can propose, challenge, reframe, simulate, and revise, but whose contributions must be examined by a capable human. This changes the purpose of AI training. The goal is not merely fluency with interfaces, as with traditional software. It is the development of disciplined judgement under conditions in which plausible language can conceal weak reasoning.

    Prompt literacy is not enough, judgement is the harder skill

    Prompt literacy asks, “How do I get the model to produce what I want?” Judgement literacy asks a harder sequence of questions: What should the model be asked to do? What would a good answer contain? Which assumptions are hidden in the response? What evidence would change my mind? How could this answer fail in practice?

    That distinction is central because AI systems are optimised to produce useful-looking responses, not to guarantee that every claim is correct, complete, current, or appropriate for the situation. A confident paragraph may combine sound reasoning with an unsupported assumption. A concise recommendation may omit a constraint that a subject-matter expert would regard as decisive. A polished analysis may be internally coherent while based on incorrect premises.

    Training must therefore teach employees to separate fluency from validity, because AI producing outputs at scale, a small error rate becomes large absolute numbers. A team processing 500 customer communications a week with a 3% undetected error rate has 15 problems a week compounding.

    Treat AI as an argument simulator, not an answer machine

    A productive human–AI interaction has several distinct moves. The user first asks the AI to generate a provisional response. Instead of accepting it, the user then assigns the system an adversarial role: critic, sceptical customer, hostile reviewer, domain expert, risk officer, opposing counsel, or operational implementer. The user asks it to identify weaknesses, test assumptions, and present alternatives. Finally, the user decides what to retain, verify, revise, or reject.

    This process is more reliable than asking for a single “best answer” because it creates structured friction. The AI is used to generate content and also to interrogate the content it produces. Its role shifts from answer machine to argument simulator.

    StageHuman responsibilityUseful AI roleOutput to preserve
    FrameDefine the objective, audience, constraints, and stakesAsk clarifying questions and expose ambiguityA precise problem statement
    GenerateRequest a provisional answer without treating it as finalProduce options, hypotheses, drafts, or modelsMultiple candidate approaches
    CritiqueInspect logic, evidence, omissions, and assumptionsAct as an adversarial reviewerA failure and risk register
    Stress-testCompare the idea against edge cases and real-world constraintsSimulate stakeholders, scenarios, and objectionsConditions under which the idea breaks
    RefineApply human expertise and verified evidenceRewrite, reorganise, and make uncertainty explicitA revised, decision-ready artifact
    ValidateOwn the final judgement and consequencesHelp create checklists or audit trailsA documented approval decision

    The critique protocol preventing AI’s fluency from masking weak reasoning

    A useful training programme should provide a repeatable critique protocol. Employees can begin by asking whether the response answered the actual question rather than a nearby, easier one. They should then examine the assumptions: What does the answer presume about the customer, market, data, timeline, resources, law, or operating environment? Which assumptions are explicit, and which are hidden?

    The next step is to inspect the evidence. Are factual claims traceable to reliable sources? Does the response distinguish observed facts from estimates, interpretations, and recommendations? Does it indicate uncertainty where uncertainty matters? A response that makes no distinction between “we know”, “we infer”, and “we might try” is difficult to govern.

    Employees should also look for omissions. What relevant stakeholder is absent? What downside is underdeveloped? What implementation cost has been ignored? What would a sceptical expert object to? In many professional settings, the most dangerous error is not a false statement but a missing consideration.

    A compact minimal critique sequence can be remembered as TRACE:

    LetterQuestion
    T — TargetDid the output address the real objective and intended audience?
    R — ReasoningAre the logic, assumptions, and causal links sound?
    A — AccuracyWhich claims require verification, and what evidence supports them?
    C — CoverageWhat perspectives, constraints, risks, or alternatives are missing?
    E — ExecutionCould a real person or team implement this, and what would fail first?

    The protocol matters less than the habit. Critique should become a normal stage of work, not an emergency response after an AI-generated mistake reaches a customer, executive, or regulator. Even one good adversarial question catches more than zero.

    The strongest exercises make people disagree with the AI, not just improve it

    The strongest exercises do not ask participants merely to improve an AI answer. They train the reflex to know when to disagree with it.

    A team might give an AI system a proposed product launch plan and ask it to produce three critiques: one from a cash-constrained finance lead, one from a sceptical customer, and one from an operations manager responsible for execution. Participants then rank the criticisms by importance, identify which are supported by evidence, and decide what additional information is needed.

    A second exercise is the assumption reversal. Participants take a central assumption in the AI’s response and invert it. If the plan assumes rapid adoption, they ask what happens if adoption is slow. If it assumes reliable data, they ask how the recommendation changes when the data is incomplete or biased. If it assumes users will follow a process, they ask what incentives would cause users to bypass it.

    A third is the minimum viable rebuttal. Each participant must name the single strongest reason not to accept the AI’s recommendation. The purpose is not to be contrarian. It is to prevent the tendency to equate a polished output with a persuasive one.

    A fourth is the pre-mortem. Participants assume the project has already failed completely, then work backward to identify what caused it. This removes the social pressure to stay quiet in a room full of consensus, because the failure is already stipulated.

    Accountability requires making human judgement visible in the work

    Organisations should not evaluate AI training by counting prompts or measuring how quickly employees produce drafts. Those metrics reward activity. Better measures reward judgement: the number of material assumptions identified, the proportion of important claims verified, the quality of alternatives considered, and the clarity with which uncertainty is communicated.

    A practical workflow is to use an AI work note for consequential outputs. It need not be long. It can record the task given to the system, the key assumptions detected, the main criticisms generated, the facts independently checked, the changes made by the human, and the person who approved the final version. This creates a lightweight audit trail without making every interaction bureaucratic. As a side effect, in high-pressure organisations when something goes wrong the person who approved the AI output without documented scrutiny is exposed. Where the person who has a note showing they checked the key assumptions is not.

    Managers should also distinguish between different levels of risk. A brainstorming exercise may require only a basic plausibility check. A customer communication, hiring recommendation, safety procedure, financial analysis, or policy decision requires a more demanding review. The higher the stakes, the more the process should emphasise source verification, domain expertise, independent reasoning, and explicit approval.

    Correction must be rewarded, not tolerated

    No training framework will work if employees believe that questioning AI marks them as inefficient or resistant to technology. Leaders must communicate that revision is not failure but the mechanism by which value is created.

    This cultural point applies to humans as well as AI outputs. People routinely accept outputs that confirm their preferences, especially when those outputs are articulate and fast. AI can amplify that tendency by presenting a conclusion before the user has fully examined the problem. The organisation must reward people who find a flaw early, surface an inconvenient alternative, or slow down a high-stakes decision long enough to validate its premises.

    People who use the sparring-partner approach understand their work better, because interrogating the AI forces them to articulate what they know and don’t know. Leaders can model the behaviour by asking “What would make this wrong?” and “Show me the strongest case against this recommendation” before asking whether the answer is useful. Over time, these questions become part of the institution’s decision vocabulary.

    The goal is better thinking

    Treating AI as a sparring partner does not mean distrusting every output or forcing every task through an elaborate review ritual. It means assigning the system the right role for the task. AI can be a rapid generator, tireless critic, perspective simulator, editor, tutor, and rehearsal partner. It cannot replace accountability for decisions whose consequences belong to people and institutions.

    The central discipline is simple: generate broadly, challenge deliberately, verify selectively, and decide consciously. When people are trained this way, AI becomes more than a productivity shortcut. While additive initially, it becomes a structured environment for thinking, one that helps users see alternatives, discover weaknesses, and improve the quality of their own judgement. People who develop this capability become the ones their organisation trusts with higher-stakes work, because they’re visibly producing better outputs with fewer errors.

    The organisations that benefit most from AI will not necessarily be those that use it the fastest. They will be those that learn how to build a culture where speed is measured over the long term.

    Start here: a 30-minute exercise any team can run this week

    Bring one ordinary work product such as a proposal, briefing, process note, or customer message. Ask an AI system to improve it. Then ask the system to criticise its own revision from three different perspectives. Have the team independently identify the two most serious weaknesses, verify the most consequential claims, and produce a final version with changes documented.

    At the end, ask three questions: What did the AI notice that we missed? What did we notice that the AI missed? Which part of the final judgement could not responsibly be delegated?

    The answers will reveal the real training agenda. AI competence is not the ability to obtain an answer. It is the ability to engage an answer critically enough to make it better, and to know when it should not be trusted at all.

  • AI Moves at the Speed the Business Can Absorb

    The pressure to “do AI” has become one of the loudest forces in business. Boards ask for a strategy, competitors announce pilots, and employees experiment with new tools before governance has caught up. In that environment, speed is often mistaken for progress. The more useful leadership question instead of How quickly can we deploy AI? is How quickly can our organisation responsibly absorb the change it creates?

    That distinction matters because AI does not merely add a new application to the technology stack. Used seriously, it changes workflows, decision rights, data practices, roles, controls, customer interactions and, in some cases, the economic logic of a business. If those changes arrive faster than the organisation can understand, govern and operationalise them, the outcome is not transformation or improvement, but accumulated work, inconsistent practices and a loss of confidence.

    Jeff Wilke, a senior executive who worked at Amazon for 25 years, made this point bluntly to Jeff Bezos early in the company’s life. Bezos was generating ideas faster than the organisation could act on them, and Wilke told him: “You have to release the work at the right rate that the organization can accept it.” Bezos later described it as a profound insight. Every idea released beyond the organisation’s capacity to absorb it did not accelerate progress. It created distraction.

    That principle applies directly to AI adoption. The objective is not to slow down for its own sake but to sequence change so that each step increases the organisation’s capacity for the next.

    Moving too fast creates invisible failure

    When leaders mandate AI adoption at a pace the business cannot sustain, the first failure is often invisible. A team may launch an impressive pilot, produce an executive demonstration, or buy licences at scale. The operating model around the tool, however, remains unfinished. Staff do not know when to rely on the output, where sensitive information may go, who owns errors, how exceptions are handled or what is expected of them once the old process is retired.

    The result is a familiar pattern: workarounds multiply, quality varies by team, risk teams intervene late, and employees quietly revert to older methods when the new approach causes friction. AI earns a reputation as a management fad rather than a source of practical value. In the most severe cases, an organisation disrupts a reliable service model before it has built a dependable replacement. The approach fails because the change exceeds the organisation’s ability to cope, not because the model used was insufficiently powerful.

    A useful way to think about this is as a mismatch between deployment velocity and absorption capacity.

    DimensionWhen deployment outpaces capacityWhen capacity sets the pace
    Process designAI is layered onto unclear or unstable work.The team redesigns one defined workflow and clarifies hand-offs.
    PeopleEmployees are told to adopt, but not taught how to exercise judgement.Training, job aids and feedback are built into the rollout.
    Data and controlsData use and accountability are resolved after launch.Guardrails, escalation paths and quality checks are established before broader use.
    Value measurementActivity is reported: licences, pilots and prompts.Outcomes are measured: cycle time, quality, cost, customer experience and risk.
    TrustEarly mistakes become evidence against the programme.Small, well-governed wins create permission for the next change.

    The difference is not caution versus ambition. It is operational discipline versus performative speed.

    Established businesses are complex

    Startups can often change rapidly because the organisation is smaller, its operating model is still forming and its technology carries little legacy integration. A founder can decide in the morning, change a product workflow by afternoon and observe the impact within days.

    Established businesses face a different reality. They may serve millions of customers, operate under regulatory obligations, rely on complex supplier relationships and carry years of interconnected systems and policies. Their scale is an advantage, but it also means that a change in one function creates consequences in many others. A new AI-assisted credit decision, for example, has implications for compliance, model governance, customer communication, front-line procedures, auditability and dispute resolution. The business must move deliberately because the cost of a poorly absorbed change is multiplied by its reach.

    This a design constraint to work within, not a weakness to apologise for. Large organisations should build a repeatable mechanism for safe acceleration. The goal is to increase the rate at which the business can absorb change, not to force a rate that breaks it.

    Build absorption capacity before demanding velocity

    The practical response is to treat AI adoption as a managed portfolio of changes rather than a single enterprise-wide instruction. Start with a workflow that is sufficiently valuable to matter, sufficiently bounded to govern and sufficiently measurable to learn from. Assign a business owner, a process owner and a risk or control partner. Define what a good outcome looks like before deployment, including the conditions under which the team will pause or reverse the change.

    A sensible sequence has four stages.

    StageLeadership questionEvidence required before moving on
    ProveDoes this solve a real operational or customer problem?A defined baseline and a demonstrable improvement in a controlled workflow.
    StabiliseCan people use it reliably under normal conditions?Clear ownership, training, controls and a functioning exception process.
    ReplicateDoes the operating pattern transfer to similar work?Consistent results across more than one team or business unit.
    ScaleCan the enterprise support this without degrading quality or trust?Resourcing, governance and technology capacity match the expanded scope.

    First, prove usefulness in a narrow setting where people can compare the AI-assisted outcome with the existing process. Then stabilise the new way of working by documenting roles, controls, exceptions and training. Then replicate the pattern in adjacent use cases that share similar data, risks or operating habits. Finally, scale the platform and governance only after the organisation has evidence that the operating model works repeatedly.

    This approach may appear slower at the beginning because it resists the temptation to announce universal adoption. In practice, it is usually faster over time. Each successful use case leaves behind reusable assets such as trusted data pathways, trained leaders, approved controls, implementation playbooks, measurement methods and internal advocates. The next deployment begins from a stronger base, drawing on accumulated organisational knowledge rather than repeating the same arguments from scratch.

    Leaders must protect the rate of learning, not the rate of adoption

    The leadership task is to sponsor AI usage and to protect the business’s learning rate. That means creating room for teams to identify flaws without being labelled resistant, refusing to measure progress solely by adoption volume and being explicit about what will not yet be automated. It also means resisting the false choice between reckless acceleration and organisational paralysis.

    A mature AI programme should be demanding, but its demands should be specific: improve a process, raise decision quality, reduce a known friction point, protect customers, and demonstrate results. “Use AI everywhere” is an instruction to create unmanaged variation at scale.

    Wilke’s observation to Bezos offers a better standard. Ideas have value only when an organisation can turn them into coherent action. Every AI deployment released beyond the organisation’s capacity to absorb it creates a backlog of unfinished transformation.

    Organisations that make change smooth – clear enough for people to adopt, controlled enough to trust and valuable enough to sustain – will accumulate capability faster than those that simply move first. Slow is smooth, smooth is fast.

  • What to Do with The Feelings When the Machine Can Do Your Job

    Something real is being taken from you, and it is not just your workload.

    Ask most people to introduce themselves, and they describe what they do. “I’m an architect” “I’m a nurse” “I’m a lawyer”. That is not modesty or habit but a statement of identity. The role names the person. It signals competence, years of effort, and a place in the world. When AI starts performing the core tasks of that role with speed and reasonable accuracy, the threat is not just professional. Losing a task is inconvenient, but losing the thing you built yourself around is something else entirely.

    Your Role Has Shifted, and That Shift Hurts More Than You Expected

    The primary change AI seems to bring to knowledge work is not elimination but demotion from creator to checker.

    When an AI system can draft a legal argument, generate a diagnosis, or produce a structural engineering report, the human professional moves from producing the work to reviewing it. The cognitive load shifts from high-level synthesis to what researchers call “AI managerial labour”: monitoring outputs, catching errors, and signing off on conclusions you did not reach yourself.

    What ChangedBefore AIAfter AI
    Where authority comes fromYour training and judgementThe algorithm’s prediction
    What you spend your time onSolving and creatingOverseeing and refining
    Your role in the decisionThe expert who decidesThe human buffer who approves

    This shift is widely documented and spreading across industries faster than most organisations are acknowledging. Software developers have already lived through it. A few years ago, AI-generated code was treated as a curiosity, too unreliable to trust. Now most developers use it routinely, and the ones who don’t are increasingly at a disadvantage. The code is not perfect, but it does not need to be. It just needs to be good enough to change the expectation of what a developer does, and that pattern is moving through law, medicine, finance, consulting, and design.

    Why This Feels Like a Betrayal

    The fear of this shift is rational. Researchers have identified a specific phenomenon as AI guilt, or AI shame. Professionals who built their identity around a craft feel genuine discomfort when they hand that craft to a machine. Using AI to do what you once did yourself can feel like a violation of your professional standards, or, as one research team put it, like “conspiring with the enemy” to make yourself obsolete. When your expertise is the thing you are proudest of, and a tool makes that expertise feel commoditised, it does not produce neutral feelings.

    The problem runs deeper when the AI cannot explain itself. Many systems produce outputs without revealing how they reached them. When you must approve a decision you cannot trace back through logic, when you are accountable for an outcome you did not reason your way to, your sense of professional autonomy erodes. You are responsible without being in control. That is a specific and corrosive kind of discomfort.

    The emotional toll is not evenly distributed. Those who have spent the longest building their expertise tend to feel this most acutely. A career built over decades on a particular kind of cognitive authority faces a different kind of threat than a career just starting out. There is more time and effort being attached to the individuals identity with less time to rebuild.

    The Way Through Is Changing How You See Yourself, Not What You Do

    Research and practical experience suggest that professionals who navigate this best are the ones who change how they think about themselves, not what they do.

    One concrete shift is in how you describe yourself. “I am a lawyer” places your identity in a function that AI is now performing. “I work as a lawyer” places it in a behaviour which is something you consciously practise, adapt, and develop. The first statement sets a boundary, but the second sets a direction. It sounds like a small distinction, but in a period of rapid change it is a significant one, because capabilities and behaviours can be learned while identities tend to be defended.

    Beyond language, the professionals who contiue to improve tend to do it by deliberately focusing on the elements of their work that remain genuinely human. Not in a defensive way, as if protecting territory, but in a generative way, actively building the part of their practice that AI cannot replicate. That means different things in different fields, but the pattern is consistent.

    The things only humans bring. Empathy, ethical reasoning, and the ability to read an unspoken concern in a client’s face are not tasks but capacities. AI can simulate some of them, but it cannot hold responsibility for the outcomes, and clients know the difference.

    The human relationship. Who decides, in the end? Who is accountable when something goes wrong? In almost every profession, the answer remains a person. Redefining your collaborations so that human judgement is clearly the final arbiter is not just good practice but also where your irreplaceability lives.

    The shift from doer to strategist. The professionals adapting best are not the ones doing less, but the ones doing differently. They use AI to accelerate the execution of ideas they are directing, rather than treating AI as a replacement for the thinking itself.

    You Still Own What the Machine Cannot Take

    AI can replicate the execution of a task, but it cannot replicate the responsibility for the consequences.

    When a diagnosis is wrong, a human clinician is accountable. If a structural report misses something and a building fails, a human engineer bears that weight. When a contract fails its client, a human lawyer answers for it. This is a legal fact and a fundamental distinction between a tool and a professional.

    The future of expertise is not in competing with AI on speed or scale, as that’s a losing proposition. It is in combining the fundamentals of your business, ethical responsibility, tacit knowledge, and relational intelligence that human professionals carry with the capability AI provides. The professionals who hold that combination clearly in mind, and invest in the human side of it deliberately, are the ones who will survive this transition and be more valuable because of it.

    The discomfort you are feeling right now is a signal worth listening to. While it might feel as if your identity is being challenged, you are much more than your role. That signal and challenge is a call to action and growth, if you act on it.


    This article draws on research by Sartirana & Salvatore (2025), Law & Varanasi (2025), Ziegelmayer & James (2024), Mirbabaie et al. (2021), among others.

  • Why AI Automation Fails Without Knowledge Engineering

    Most AI automation projects for business processes start the same way. A vendor demonstrates a compelling proof of concept and the demo works. The board approves the investment but six months later, the automated process is producing results that no one trusts, and the organisation is quietly routing work around it.

    The failure wasn’t in the AI. It was in what the implementation team did before they deployed the AI.

    Quick automation captures what systems do while knowledge engineering captures why they do it. These are not the same thing, and conflating them is the most common and costly mistake in enterprise AI implementation today.

    The Gap That Quick Automation Cannot See

    Process mining and rule extraction are useful tools. They read event logs from your ERP, CRM, and core systems, then reconstruct how your processes actually flow. They surface bottlenecks, flag deviations, and give your implementation team a working map of operations.

    For simple, well-documented processes, this is enough but for anything more complex, it is a structural failure. Event logs record what happened in the system, not the reasoning behind it. They cannot see the workarounds your experienced people have built over fifteen years. They cannot capture the judgement your senior underwriter applies when a claim is technically within policy but smells like fraud. They cannot document the exception a finance team made for a long-standing client that never made it into the system because everyone just knew.

    Research published in 2025 confirmed that event logs systematically underrepresent operational reality, with an estimated 70% of knowledge work occurring outside the systems that generate those logs. What gets automated by the AI is the documented surface of a process, not its full depth.

    When the AI encounters the undocumented exception, it either fails, escalates, or produces a wrong answer with complete confidence. All three outcomes erode trust, and once trust is gone, the automation sits idle while your people route around it.

    What Knowledge Engineering Actually Does

    Knowledge engineering is the discipline of making implicit knowledge explicit and verifiable before your team encodes it into an AI system.

    This work involves structured interviews with domain experts. It involves observing how work actually gets done, not how the procedure manual says it should be done. It involves encoding the results into knowledge bases that subject matter experts can review, challenge, and correct. It involves validating that the encoded logic produces the right outcomes before anything touches a production environment.

    This is the pattern that distinguishes durable AI automation from the kind that works in demos and fails in production. The AI surfaces candidate rules and logic and domain experts validate or correct them. The validated knowledge becomes the foundation your team builds automation on.

    This model is explicit about something that many vendors prefer not to emphasise: AI will confidently surface intent that sounds right and is not. Business-logic reconstruction can leverage AI but remains human-led by design. The AI drafts the interpretation while the expert confirms or corrects it.

    Where Systems Thinking Changes the Problem

    A systems thinking approach to AI automation asks a different question than most implementation teams ask. Instead of “what processes can we automate?”, it asks “what does the system as a whole need to keep working correctly after automation is introduced?”

    This reframing matters because automation does not operate in isolation. When you automate a claims processing workflow, you change the load on the humans who handle escalations. You change what data your compliance team can access and when. You change the feedback mechanisms that let your experienced people spot when something is going wrong. You may also change the incentives in ways you did not intend, particularly if the organisation measures the automated process on speed but measured the people it replaced on accuracy.

    Systems thinking requires mapping those interdependencies before implementation, not discovering them six months after go-live.

    Three principles from this approach are directly applicable to AI automation decisions.

    The first is that problems appear long before they are noticed. An AI model that is slowly drifting from its training distribution will produce subtly degrading outputs for weeks before anyone raises a flag. IBM’s 2025 Cost of a Data Breach Report found that 13% of organisations reported breaches involving AI models or applications, and 97% of those breached organisations lacked proper AI access controls. The failures were systemic, not sudden.

    The second is that you get what you incentivise. If your AI implementation is measured on the percentage of cases processed automatically, you will get high automation rates but you may also get poor outcomes on the cases that needed human review but were incorrectly classified as routine. Design the measurement framework before designing the automation.

    The third is that you fall to your level of preparation. When something goes wrong with an automated process at scale, your team’s ability to respond depends entirely on what they practised before it happened. Oversight protocols, escalation paths, and retraining procedures need to exist and be tested well before an incident requires them.

    The Real-Time Oversight Problem

    Traditional governance was built for a different tempo: annual audits, quarterly reviews, weekly metrics. These cadences made sense when processes ran at human speed and human scale.

    AI automation changes both variables simultaneously. A model processing thousands of claims per day can accumulate hundreds of errors before a weekly metrics review would detect the pattern. By the time a quarterly audit surfaces a systematic bias in lending decisions, the exposure may already be material.

    This is not a hypothetical concern. The EU AI Act’s Article 96 now requires organisations to demonstrate compliance through continuously updated, machine-readable evidence, not point-in-time assessments.

    The implication for senior leaders is direct. Governance frameworks designed for human-speed processes need to be redesigned before your organisation deploys AI at scale, not retrofitted after a problem surfaces. Real-time monitoring of model outputs, data quality, and prediction confidence is not optional infrastructure. It is the mechanism by which you maintain accountability for decisions your organisation is making at machine speed.

    What a Systemic Approach Actually Looks Like

    Implementing AI automation well follows a recognisable pattern.

    Begin with knowledge extraction, not process mapping. Before any automation is designed, invest in capturing the institutional knowledge that lives in people, not systems. This means structured elicitation sessions with domain experts, protocol analysis of how difficult cases are actually handled, and documentation of the exceptions and judgement calls that define quality outcomes.

    Validate before you deploy. The people who encoded the knowledge review it. Pilot programmes run in environments where your team checks outputs against known-good outcomes before the model is trusted with live decisions.

    Design oversight into the process architecture. Human-in-the-loop checkpoints are not emergency interventions but are designed features, placed where the model’s confidence is lowest or where the cost of an error is highest.

    Instrument for continuous monitoring. Model performance, data quality, and output confidence scores are tracked continuously. Thresholds trigger review before drift becomes visible to customers or regulators.

    Govern at the speed of your systems. Review cadences match the operational tempo of the automated process, not the legacy tempo of the governance process it replaced.

    The Question to Ask Your Implementation Team

    If your organisation is evaluating or currently implementing AI business process automation, one question will reveal more about the quality of the approach than any other: what knowledge engineering work did the team complete before the AI model was trained?

    If the answer is “we used automated process mining and rule extraction from our existing systems”, that is a starting point, not a complete answer. It describes what the system does, it does not describe what your experienced people know that the system does not record.

    If the answer is “we conducted structured knowledge capture with domain experts and validated the results before training”, you are working with a team that understands where AI automation actually fails.

    The technology for automating business processes is mature, accessible, and increasingly affordable. The discipline required to automate them well, preserving institutional knowledge, designing real-time oversight, and aligning governance to the speed of the systems being governed, remains genuinely rare.

    That gap is where durable competitive advantage is built.

  • APRA Is Watching Your AI. Is your QA Strategy Ready?

    Australian Financial Services are moving fast with AI and APRA has high expectations.

    Most boards haven’t caught up yet. APRA knows this, and it has said so plainly in its letter to the industry. The regulator’s message is clear: existing governance, risk, and operational resilience practices are not keeping pace with how quickly AI is being deployed inside regulated entities. That gap is now a supervisory concern.

    This article explains what APRA and Australia’s updated privacy laws require of Financial Services using AI, and what a Quality Assurance strategy needs to look like in response.

    APRA Sees Three Governance Failures Happening Right Now

    APRA’s concerns are specific and three patterns appear consistently in its observations of the sector.

    Boards are making AI decisions without enough information. Many boards rely on vendor presentations to understand their AI risks. That’s a problem when the vendor is also the one selling the product. APRA expects boards to be capable of independent challenge, and most are not there yet.

    Governance exists on paper, not in practice. Entities have acknowledged that existing prudential standards apply to AI. Few have actually operationalised that acknowledgement with post-deployment monitoring, change management, and model decommissioning are common gaps.

    Supplier dependencies are underexamined. Many FinTechs have concentrated significant AI activity with a single provider without tested exit strategies. When those providers rely on foundation models, training data, or fourth-party services, the chain of accountability becomes opaque. APRA expects visibility all the way through that chain.

    APRA frames AI governance inside existing obligations such as CPS 230 for operational risk, outsourcing standards, and the Financial Accountability Regime (FAR). There is no separate AI rulebook. The existing rulebook applies, and APRA believes most entities are not meeting it.

    Privacy Law Has Changed. Your AI Systems Probably Haven’t.

    The Privacy and Other Legislation Amendment Act 2024 introduced changes that directly affect Financial Services using AI to make decisions about customers.

    The most significant change for AI is the new automated decision-making transparency requirement. If your systems use personal information to make decisions that significantly affect individuals, such as loan approvals or credit scoring, customers will have a right to meaningful information about how those decisions are made. The two-year grace period ends on 10 December 2026.

    Two other changes carry direct operational implications. First, Australians now have a personal right to sue for serious invasions of privacy. AI systems that process sensitive personal data at scale carry real exposure here. Second, the Office of the Australian Information Commissioner (OAIC) has stronger enforcement powers, including tiered civil penalties and the ability to conduct compliance assessments. The OAIC has already issued specific guidance on commercially available AI products.

    APP 11 now explicitly requires “technical and organisational measures” to protect personal information. This is not an IT-only obligation and covers the full range of privacy and security controls around AI systems.

    The Regulator Expects Boards to Lead, Not Delegate

    Both APRA and ASIC have reached the same conclusion from different directions: Financial Services are adopting AI faster than their governance frameworks can handle it, and boards are too far removed from the risk to provide effective oversight.

    APRA’s expectation is that boards understand AI well enough to set strategic direction and challenge assumptions. ASIC’s concern is that licensees are creating consumer harm by deploying AI without updating their risk and compliance frameworks.

    The practical implications are straightforward. Boards and executives need sufficient AI literacy to ask hard questions of management and vendors. Accountability for material AI use cases needs to be named, not distributed. Under FAR, accountable persons need to be identifiable for decisions made by or with AI systems. Staff need training that goes beyond “here is the tool” and covers misuse, limitations, and secure practices.

    Human oversight and ownership is not optional for high-risk decisions. AI can inform those decisions but a named person needs to own them.

    Your AI Systems Need to Fail Safely, Not Just Perform Well

    CPS 230, effective from 1 July 2025, extends operational resilience obligations to AI-enabled systems. This creates concrete requirements that go beyond standard performance monitoring.

    AI systems supporting critical operations need tested fallback processes. That distinction matters when regulators ask for evidence. AI failure modes that need to be planned for include hallucination, silent degradation, and susceptibility to adversarial inputs such as prompt injection or data poisoning.

    Security requirements have also become more specific. AI adoption changes the attack surface with more entry points, faster attack cycles, and new risks from non-human AI agents with system access. APRA expects strong privileged access management, timely patching, hardened configurations, and penetration testing that covers AI-specific vulnerabilities, including AI-generated code.

    Data governance sits underneath all of this. The quality and provenance of training data affects model behaviour. That is now a prudential concern.

    What an APRA-Ready AI QA Strategy Looks Like

    Active supervision of AI is underway with regulators no longer observing and advising. They are assessing and acting. An AI QA strategy needs to be designed for that environment.

    The foundation is a centralised AI inventory: every AI system in use, including third-party tools, mapped to the regulatory obligations it touches under CPS 230, FAR, and the Privacy Act. Without this, gap assessments and audit processes cannot function.

    From that inventory, four capabilities need to be in place.

    Continuous monitoring for bias, drift, and performance degradation. AI models do not stay stable. A model that was accurate at deployment may not be accurate six months later. Automated monitoring catches this before regulators do.

    Independent assurance for high-impact systems. Internal audit and external review processes need to cover AI systems with material customer or operational impact. This cannot be delegated to the team that built the system.

    Testing frameworks built for AI. Standard software testing does not address algorithmic bias, adversarial inputs, or the behaviour of AI-generated code. Testing frameworks need to evolve to cover these risks explicitly.

    Automated compliance tooling. Governance, Risk, and Compliance (GRC) platforms can automate the monitoring and reporting that manual processes cannot keep up with. Predictive compliance tools can identify non-compliant states before they become audit findings. SaaS Security Posture Management (SSPM) tools provide continuous measurement of security controls against regulatory baselines.

    The Gap Between Adoption and Governance Closes in One Direction

    Regulators are not going to slow down their expectations to match the pace of industry governance. The direction of travel is more scrutiny, not less.

    To get ahead of this do three things. Build board-level AI literacy that enables genuine challenge, not just approval. Establish clear ownership for AI decisions at the individual accountability level. Instrument their AI systems for continuous oversight rather than periodic review.

    APRA has been explicit about what it is looking for. The question is whether your governance framework reflects what your AI systems are actually doing, not what the documentation says they do.

  • Australia’s AI Disclosure Deadline Arrives

    If your organisation uses AI to help make decisions about people, Australia’s privacy law is about to change the rules, with the new obligations commencing on 10 December 2026.

    This article sets out what the new obligations require, where other jurisdictions are heading, and the steps your organisation could be taking now.

    December 2026: What the Australian Law Requires

    The Privacy and Other Legislation Amendment Act 2024 introduces a transparency obligation for automated decision-making. It sits inside the Australian Privacy Principles and applies to any APP entity, which covers most businesses handling personal information.

    The obligation is triggered when three conditions are met at the same time:

    1. Your organisation has arranged for a computer program to make a decision, or to do something substantially and directly related to making a decision.
    2. That decision could reasonably be expected to significantly affect the rights or interests of an individual.
    3. Personal information about that individual is used in the operation of the program.

    When all three conditions are met, your privacy policy must disclose what kinds of personal information the program uses, what types of decisions the program makes on its own, and what types of decisions the program substantially assists a human to make.

    Two details in that test widen its scope considerably.

    First, the obligation covers assisted decision-making as well as fully automated decisions, so a computer program that materially steers a human decision-maker brings the obligation into play. A loan officer who reviews an AI-generated credit score before approving or declining an application is making an assisted decision that falls within scope.

    Second, the term “computer program” is interpreted broadly enough to cover generative AI tools, rule-based engines, and sophisticated spreadsheets that score or rank individuals. If your team has quietly introduced automation into a workflow over time, that automation is likely in scope.

    Decisions that could significantly affect rights or interests include home loan approvals, insurance assessments, job application screening, housing allocation, and access to healthcare services. The Office of the Australian Information Commissioner (OAIC) is developing guidance expected by September 2026, though waiting for that guidance before acting carries real risk, as the months between September and December leave little room for the audit work that needs to happen first.

    Regulators Worldwide Are Moving in the Same Direction

    Australia’s December deadline reflects a global shift rather than an isolated local initiative, and regulators in multiple jurisdictions have reached similar conclusions about AI accountability.

    The EU AI Act Ties Obligations to Risk Level

    The EU AI Act classifies AI systems by risk level and attaches progressively stricter obligations to higher-risk applications. Providers of high-risk AI systems must design for transparency so users can understand what the system does and use it correctly. Providers of AI systems that generate or alter content must disclose that the output is AI-generated, with narrow exceptions for legal purposes or clearly artistic contexts. Impact assessments and documentation of decision-making processes are mandatory for high-risk applications.

    United States Regulation Is Emerging State by State

    The US has no single federal AI law, but individual states are filling that gap. California’s AB 3030, effective 1 January 2025, requires licensed healthcare providers to disclose when generative AI was used to create patient-facing content. Connecticut has established frameworks for automated employment decision tools that include mandatory consumer disclosures. This state-by-state patchwork creates real complexity for organisations operating across multiple US states, and it continues to grow.

    Canada’s Proposed AIDA Signals the Same Intent

    Canada’s Bill C-27 includes the Artificial Intelligence and Data Act (AIDA), which would create a regulatory framework for the design, development, and deployment of AI systems. The bill’s future depends on legislative processes still in progress, but the drafting shows a clear intent to regulate AI systems that could materially affect individuals.

    The through-line across all three jurisdictions is consistent, in that any AI system making or influencing consequential decisions about people is likely to attract a requirement to disclose and explain.

    Five Steps Your Organisation Can Take Before the Deadline

    Compliance with the December 2026 deadline calls for preparation that starts well before the OAIC publishes its final guidance, and the following sequence offers a workable approach.

    1. Map Every AI System That Touches Decisions About Individuals

    Start with a complete inventory, since most organisations have more automated decision-support than they realise. Automation tends to arrive incrementally, so a workflow that began as a manual spreadsheet review may now include scoring logic that materially influences outcomes. Third-party tools, vendor platforms, and SaaS applications frequently contain embedded AI functionality that the organisation never explicitly chose.

    For each system you identify, document what personal information it uses and how that information flows through the process, as this mapping forms the foundation the remaining steps build on.

    2. Apply the Three-Condition Test to Each System

    For each identified system, work through the three conditions in order, asking whether a computer program is making or substantially contributing to a decision, whether that decision could significantly affect someone’s rights or interests, and whether personal information is used in the process.

    This analysis calls for both legal and operational judgement, since the same system may trigger the disclosure obligation in one use case and not another. Document your reasoning for every conclusion, including the cases where you determine the obligation does not apply, as that record demonstrates a considered approach if the OAIC reviews your compliance.

    3. Rewrite Your Privacy Policy with Specificity

    Generic statements about using technology to assist decisions are unlikely to satisfy the new requirements. Your privacy policy needs to identify the kinds of personal information used in your automated systems, the categories of decisions made solely by those systems, and the categories of decisions where those systems substantially assist human decision-makers.

    For that reason, the policy is best written after the audit rather than before, since policies drafted from assumptions about what your systems do tend to be inaccurate, and inaccurate disclosure creates a compliance problem of its own.

    Alongside the public-facing policy, maintain internal documentation of each system’s design, the testing conducted, and the risk assessment process, as this record supports both regulatory compliance and sound governance.

    4. Build AI Review Into Procurement and Change Management

    The compliance obligation does not stop at the systems you have today, as vendors update their tools, new AI functionality arrives inside products your team already uses, and new systems join the estate over time.

    Integrating an AI disclosure assessment into your procurement process for any new tool or material software update gives your team a repeatable way to evaluate whether new capabilities bring the organisation into scope.

    5. Establish Human Oversight Protocols for Assisted Decisions

    Where AI assists human decision-makers, clear protocols for human review serve as both a legal expectation and sound risk management. Individuals should have a meaningful avenue to understand and challenge AI-influenced decisions, and for high-impact categories such as credit, employment, or healthcare access, that avenue should be accessible and substantive rather than a formality.

    Train the people who work inside these systems on what the regulation requires and on their specific role in maintaining compliance, since regulatory obligations met at the policy level but not understood at the operational level tend to fail when tested.

    The Cost of Waiting Is Higher Than It Appears

    Organisations planning to start compliance work after the OAIC guidance arrives in September 2026 face a narrow window. The audit alone can take months in organisations with complex or distributed technology environments, and rewriting privacy policies, updating vendor contracts, establishing governance protocols, and training staff all compound that timeline.

    The difficulty here lies less in the regulation, which is reasonably clear, than in the operational reality of understanding what your AI systems do at a level of detail sufficient to make accurate public disclosures.

    Organisations approaching this work systematically from now should be positioned to comply with confidence and to use their privacy policies as a genuine communication tool with customers. The regulation exists in response to AI systems making consequential decisions about people who have a legitimate interest in knowing. Building your compliance programme from that principle, rather than from the minimum required to avoid scrutiny, tends to produce better outcomes for the organisation and for the individuals affected.

  • Your AI Interface Is a Cognitive Design, Not a Cosmetic One

    Most organisations deploying AI spend months selecting the right model, agonising over accuracy rates, vendor contracts, and compliance implications. Then, in the final weeks before launch, someone asks: “What should the screen look like?”

    The interface is not decoration applied after the real decisions are made, instead it determines whether your people reason clearly alongside the AI or quietly work around it. Get it right, and your AI system makes better decisions than either the human or the system could alone. Get it wrong, and you have an expensive system that your team has learned to distrust, override, or ignore.

    This article gives you a structured method to audit any AI interface before deployment. Based on a field called Cognitive Systems Engineering (CSE), developed in the early 1980s by Erik Hollnagel and David Woods to address exactly the kind of high-stakes, human-machine decision environments that organisations across every industry are now building. The method is built around a five-level analysis called the Abstraction Hierarchy. You will walk through each level (we’ll use a loan decision AI as the working example), leaving with a set of questions you can apply to your own deployment.

    The Shift That Changes Everything

    Traditional software is deterministic. You press a button, you get a result. Designing the interface for that kind of system still requires some skill, but is largely a matter of clarity and efficiency because we have many examples of good (and bad) design.

    AI is different as it produces uncertain outputs and reasons probabilistically. It can be right most of the time and catastrophically wrong in ways that are difficult to anticipate. The interface for a traditional system needs to be usable. The interface for an AI system needs to do something different: it needs to make the AI’s reasoning visible so that the human can judge when to act on it, when to question it, and when to override it.

    Hollnagel and Woods called this a “joint cognitive system“: the human and the AI are not separate entities where one hands off to the other. They are a single thinking unit, and the interface is the connective tissue between them. Whether the decision involves a loan approval, a clinical recommendation, a fraud alert, or an operational plan, that connective tissue either holds or it tears.

    The Abstraction Hierarchy, developed by Jens Rasmussen and Kim Vicente at the Risø National Laboratory in Denmark, gives you a disciplined way to design that connective tissue. It asks you to understand your work domain at five levels, from purpose down to physical configuration. Each level reveals a different set of interface requirements that you might otherwise miss.

    The Abstraction Hierarchy: Five Questions for Any AI Interface

    Think of the Abstraction Hierarchy less as a taxonomy and more as a diagnostic interview. At each level, you ask a question about your work domain. The answers tell you what your interface must show, what it must prevent, and what it must make possible.

    Here is the framework applied to a loan decision AI to help understand the details.

    Level 1: What is this system ultimately for?

    This is the question of functional purpose: the values and goals the system exists to serve. Not the technical goals (“classify loan applications”), but the human and organisational values at stake.

    For a loan decision AI, the answer is not simply “approve or decline applications faster”. The real purposes are: extend credit to people who can repay it, protect the institution from credit losses, comply with responsible lending obligations, and treat applicants fairly across demographic groups. These purposes can conflict. A model optimised for speed may sacrifice fairness. A model optimised for loss minimisation may discriminate. These tensions exist whether you name them or not and exposing them allows informed decisions.

    What this means for your interface: The interface must make the system’s governing priorities visible. If the model has been calibrated to weight certain risk factors above others, the loan officer reviewing its recommendation should be able to see that calibration, not just the output. When the AI recommends declining an application, the interface should surface which of the system’s core purposes drove that recommendation: is this a credit risk concern, a compliance flag, or something the model cannot categorise cleanly?

    Questions to ask about your deployment:

    • What are the two or three values this system is genuinely optimised for?
    • What values are in tension, and does the interface make those tensions visible?
    • When the AI’s recommendation conflicts with a user’s instinct, does the interface help the user understand why?

    Level 2: What principles govern how the system operates?

    This is the question of abstract function: the rules, regulations, and governing principles that constrain what the system can and cannot do. In aviation, these are physics and safety regulations. In lending, they are responsible lending laws, anti-discrimination regulations, internal credit policy, and audit requirements.

    These constraints do not change with each transaction. They define the boundaries within which all decisions must fall. The problem is that most AI interfaces present outputs as though these constraints do not exist. The model returns a score. The interface shows the output metric. The loan officer is left to remember, from training, which regulatory constraints apply in this situation.

    What this means for your interface: Regulatory and policy constraints should be structurally present in the interface, not stored in the user’s head. If the AI’s recommendation would require a manual review under responsible lending obligations, the interface should flag that requirement automatically. If the model’s confidence falls below a threshold that your compliance team’s internal policy requires to be reviewed, show that threshold, and ideally the policy, visibly.

    This is the principle that Vicente and Rasmussen called Ecological Interface Design (EID): the constraints of the work domain should be visible in the interface itself, not stored in the user’s memory. An EID-informed loan interface does not require the loan officer to remember the regulatory rulebook. It makes the rulebook structurally visible in the decision flow.

    Questions to ask about your deployment:

    • Which regulatory obligations apply to the decisions this AI supports?
    • Are those obligations visible in the interface, or do users have to remember them independently?
    • When the AI’s recommendation sits in a regulatory grey zone, does the interface make that visible?

    Level 3: What processes does the system need to perform?

    This is the question of generalised function: the operational processes and workflows that need to happen for the system to achieve its purpose within its constraints.

    In a loan context, this includes the steps of gathering applicant data, running the credit model, checking for compliance flags, presenting a recommendation, capturing the loan officer’s decision and rationale, escalating edge cases, and creating an audit trail. These processes are not all equally visible in most AI deployments. Typically, what is visible is the output of the model. The processes that produced it, and the processes that need to follow from it, are hidden or scattered across different systems.

    What this means for your interface: The interface should reflect the full process, not just the model’s output. If the correct process requires a loan officer to review the AI’s recommendation alongside the applicant’s supporting documents before deciding, the interface should make that sequence natural and difficult to skip. If the process requires capturing the officer’s reasoning when overriding the AI, the interface should prompt for that reasoning at the point of override, not as a retrospective form filed later.

    This is also where structured input design matters. If you want the AI to produce consistent, auditable outputs, you need the interface to guide consistent, structured inputs. A free-text prompt box for a loan officer to query the AI is the wrong design. A structured form that constructs the query from validated fields, treating the prompt as a template and the user’s input as variables, produces far better consistent results and a far cleaner audit trail.

    Questions to ask about your deployment:

    • What is the full sequence of steps the process requires, including before and after the AI’s output?
    • Does the interface make the correct sequence the natural path, or can users shortcut it?
    • How does the interface capture the human’s reasoning, not just the AI’s recommendation?

    Level 4: What are the capabilities and limits of each component?

    This is the question of physical function: what each component of the system can and cannot do. For an AI, this means understanding the model’s actual capability boundaries. Where does it perform well? Where does it degrade? What kinds of inputs push it outside its training distribution?

    Loan officers who use AI tools daily develop intuitions about where the model is reliable and where it is not. New officers do not have those intuitions, and even experienced officers can be misled when the model presents its outputs with uniform visual confidence regardless of whether it is in familiar or unfamiliar territory.

    What this means for your interface: The interface must distinguish between high-confidence and low-confidence outputs, and it must do so in a way that reflects the model’s actual calibration, not a standardised disclaimer. If the AI is recommending approval on an application that combines features it has rarely seen together, that uncertainty should be visible. If the model’s confidence score is below a meaningful threshold, the interface should communicate that clearly and differently from high-confidence outputs, not with a footnote, but with a structural difference in how the recommendation is presented.

    This is what the CSE literature calls making the AI’s epistemics visible: the interface should show not just what the AI concluded, but how firmly it concluded it and on what basis.

    Questions to ask about your deployment:

    • Does the interface distinguish between high-confidence and low-confidence recommendations?
    • Can users tell when the AI is operating in territory close to the edge of its training?
    • What happens when the AI encounters an input type it was not trained on?

    Level 5: What is the actual configuration of the system?

    This is the question of physical form: the literal layout, controls, and information architecture of the interface as it exists on the screen.

    This is where most interface design effort is spent and also the level where most AI interface problems are most visible: the designer burying a recommendation at the bottom of a long screen, the confidence score presented in a font smaller than the surrounding data, or the override button placed three clicks away. These are not aesthetic problems but decision quality problems.

    If the most important signal the AI is sending is that it is uncertain about this application, and the interface makes that signal hard to find, the loan officer will miss it. That is design failure, not a training failure.

    What this means for your interface: The visual hierarchy of the interface should reflect the information hierarchy of the decision. The AI’s recommendation and its confidence level should be visually prominent. Flags and caveats should not be hidden in tooltips. The action the interface makes easiest should be the action the process intends to be most common. The action that requires more care, like an override of the AI’s recommendation, should require commensurate effort in the interface: not so much effort that it becomes a workaround, but enough that it cannot happen accidentally.

    Questions to ask about your deployment:

    • Does the visual hierarchy of the interface match the decision hierarchy?
    • Is the AI’s uncertainty as visible as the AI’s recommendation?
    • What is the path of least resistance in the interface, and is that the right path?

    What Happens When You Skip This Analysis

    The most common failure mode is not that the AI model is wrong but that the interface makes it impossible to know when the model is wrong.

    People in high-stakes roles learn quickly. If an AI interface presents confident-sounding recommendations without surfacing the model’s reasoning or uncertainty, they will test it against their own judgement for a few weeks, find cases where it was clearly wrong, and start treating all its outputs with blanket scepticism. The AI becomes a checkbox, not a collaborator. The organisation has paid for a decision-support system and deployed a bureaucratic step.

    The second failure mode is the reverse: people defer to the AI when they should not, because the interface presents its outputs with more authority than the model’s actual confidence warrants. This produces decisions that look considered but are indefensible when scrutinised: by regulators, auditors, customers, or a board asking why something went wrong.

    Both failures are interface failures. The model may be performing exactly as designed. The interface is simply not communicating what the model knows and does not know.

    Where to Start

    You do not need to redesign your entire interface before launch. You need to run through these five levels with the people who will use the system and the people responsible for the process it supports.

    Bring three groups into a room: the people who will use the AI day to day, the people who own the process it sits inside, and whoever is accountable for risk or compliance in that domain. Walk through the five levels as questions. At each level, ask: what does the interface currently show, and what does this analysis say it needs to show? The gaps between those two answers are your design priorities.

    Pay particular attention to Levels 1 and 4. Purpose misalignment (Level 1) and invisible uncertainty (Level 4) are the two most dangerous gaps, and they are the two most commonly overlooked in AI deployments focused on model performance rather than interface design.

    The Abstraction Hierarchy was originally developed to allow human operators to make complex, safety-critical decisions under pressure. Whether your AI is supporting credit decisions, clinical triage, fraud detection, or operational planning, the underlying challenge is the same: consequential, often regulated decisions where the interface either helps people reason well or quietly gets in the way.

    Your AI model does not make decisions. The human-AI system makes decisions. The interface is what makes that system work so design it accordingly.


    References

    [1] E. Hollnagel and D. D. Woods, “Cognitive systems engineering: New wine in new bottles,” International Journal of Man-Machine Studies, vol. 18, pp. 583–600, 1983.

    [2] K. J. Vicente and J. Rasmussen, “Ecological interface design: Theoretical foundations,” IEEE Transactions on Systems, Man, and Cybernetics, vol. 22, no. 4, pp. 589–606, 1992.

  • Is Your AI a Tool, a Colleague, or an Authority?

    Something happens the moment your organisation settles on a word for what AI is. A design philosophy clicks into place and governance questions answer themselves. Oversight feels obviously necessary, or it feels obviously unnecessary. Most of the time, no one in the room notices this happening.

    The words are not a label but a mental frame that imports an entire domain of human experience, complete with its own logic about agency, control, and responsibility. The linguist George Lakoff calls this a conceptual metaphor: we don’t merely describe experience through metaphor, we think through it. The frame structures what feels like common sense, and what feels like common sense rarely gets examined.

    When your organisation calls AI a tool, you are not describing a deployment approach alone. You are importing the complete conceptual logic of tool use, and that logic will quietly answer a hundred design and governance questions before anyone thinks to ask them explicitly.

    Tool: The Human Is the Only Agent in the Room

    A tool has no agency. A hammer doesn’t decide where it lands, a calculator doesn’t choose what to calculate. When a hammer hits the wrong nail, no one convenes a review of the hammer. The logic flows in one direction: the user holds all intention and bears all responsibility.

    Under the tool metaphor, this feels like common sense. The design imperative is a high-control interface that keeps the human firmly in command, and oversight infrastructure feels redundant in the same way you would never audit your word processor. If the AI produces wrong output, the frame says it’s a user problem: either the instructions were wrong, or the output wasn’t checked carefully enough.

    The metaphor works well when it accurately describes what the system does. Grammar checkers, data analysis dashboards, and coding assistants genuinely behave like sophisticated tools. The user directs and the tool responds.

    The metaphor breaks when it’s applied to systems that exercise something resembling judgement. AI that screens job applications, assesses loan risk, or makes triage recommendations is not behaving like a hammer. It generates outputs the user didn’t specify, based on patterns the user didn’t choose. When those outputs are wrong, the tool metaphor offers no mechanism for catching errors and no conceptual language for asking why. The frame has already answered that question: user error.

    Peer: The Machine Gets a Voice

    The colleague or co-pilot metaphor imports a different logic entirely. Colleagues have agency, offer opinions you didn’t ask for, and can be wrong while remaining entirely confident. You expect them to explain their reasoning, and if they can’t, you trust the recommendation less.

    This is a more honest metaphor for how AI behaves in many deployments. Fraud detection systems that flag anomalies for human review, content systems that propose and revise, diagnostic tools that suggest rather than decide: these are genuine collaborations where both parties contribute and neither is simply executing the other’s instructions.

    The peer metaphor makes explainability feel natural because of course you want the AI to show its work, of course a human reviews before acting. Shared responsibility follows from how we think about working alongside colleagues: you don’t fully outsource your judgement to someone else, no matter how capable they appear.

    What this metaphor hides is an asymmetry of confidence. A human colleague who doesn’t know something usually knows they don’t know, whereas AI can be wrong with complete conviction and no visible hesitation. The peer frame can lead people to extend more trust than the relationship warrants, precisely because the metaphor makes trust feel appropriate.

    Authority: The System Decides, Humans Comply

    When AI replaces a human process entirely, a different metaphor tends to take hold, often without being named: the system as authority, processing, deciding, and acting while humans monitor the results.

    The authority metaphor imports from the domain of institutional rules and procedures, where rules are correct by definition and the system knows best. Challenging the output feels like questioning the process itself, which feels obstructive rather than responsible. When the system flags something, people act on the flag. When it doesn’t, people don’t look further.

    This is not inherently dangerous, for example automated invoice processing and robotic process automation handle large volumes of low-stakes work effectively, and the authority metaphor fits those applications well.

    The problem arrives when it’s applied to high-stakes decisions and the system’s reliability is treated as given rather than tested. Automated hiring decisions, credit scoring, and content moderation all operate under this logic. The system makes the call, and human judgement enters only at the edges. The conceptual frame makes this feel like efficiency. The practical consequence is that errors encoded into the system become very hard to see, and harder still to challenge, because the authority metaphor has already told everyone that challenging the system is not their role.

    The Frame You Don’t Examine Is the One That Governs You

    Any of these frames can be appropriate when it accurately describes what the system does: the tool approach for systems the user genuinely directs, the peer model for genuine collaboration, and the authority model for high-volume, low-stakes processes with clear oversight in place.

    Failures accumulate when the metaphor doesn’t match the reality, and no one has named the mismatch. Fpr example. an authority-level system governed by tool-level assumptions, a peer-level AI trusted like a colleague long before it has earned that trust, or a replacement system with no escalation path because, conceptually, there is nothing to escalate from.

    These are not governance failures in the conventional sense so much as failures of conceptual clarity: the organisation built what it thought it was building, solving the problems it thought it was solving, but hadn’t examined the underlying assumptions.

    Naming the Metaphor Is the First Act of Governance

    Before asking what guardrails your AI needs, ask what your organisation believes it is at the level of operating assumption, not at the level of documentation.

    The questions below are designed to surface the operating metaphor. What they reveal is not always what the answer says: often it is what the answer assumes, or what the question itself appears to disturb.

    Who is responsible when the AI gets it wrong?

    The tool frame answers quickly: the person who used it. The peer frame identifies whoever reviewed the output before it was actioned. The authority frame produces something else: confusion, a redirect to IT or the vendor, or a long pause followed by “that hasn’t really come up”.

    The reaction worth noting is impatience. “Obviously the user” said with certainty is itself diagnostic when the system is making decisions the user never specified.

    What happens when someone disagrees with an AI output?

    This question separates the existence of a mechanism from the existence of permission. In a tool frame, no formal process is needed: the user simply doesn’t act on the output. With a peer frame, disagreement has a path: escalation, review, documentation, and move on. In an authority frame, the question often produces a category error. Disagreement is routed to a technical team rather than a domain expert, because the operating assumption is that the output is a system output, not a decision subject to challenge.

    The tell is when the question appears to confuse process with permission. “Anyone can disagree” is not a mechanism.

    How would you know if the AI started getting things wrong?

    The tool frame assumes the user would notice immediately. The peer frame has monitoring, defined thresholds, and review cycles. The authority frame tends to produce the longest pause of any question on this list, followed by “we’d get complaints” or “the vendor monitors that”.

    “We’d get complaints” means the error detection mechanism is your customers.

    Can the AI explain why it produced that output, and does anyone ask it to?

    The first half of this question is technical and the second half is cultural. Systems that are capable of explanation but where no one has ever requested one reveal the operating assumption as clearly as any governance document. In a tool frame, explanation feels unnecessary or obvious. With a peer frame, it is routine and integrated into the process. In an authority frame, the capability may exist in a dashboard somewhere that no one opens. If the honest answer is “I’m not sure anyone has”, that is the operating metaphor speaking.

    What’s the reason for that answer? Is the standard followup question to answers, following the ‘Five whys’ format.

    The point is not which metaphor is correct in the abstract, but whether the one in use was chosen deliberately rather than inherited by default and made explicit to everyone. The frame you pick will make certain things feel like common sense. Make sure those are the things you want to feel natural.


    References

    [1] Lakoff, G., & Johnson, M. (1980). Metaphors We Live By. University of Chicago Press.