Category: AFSL

  • APRA Is Watching Your AI. Is your QA Strategy Ready?

    Australian Financial Services are moving fast with AI and APRA has high expectations.

    Most boards haven’t caught up yet. APRA knows this, and it has said so plainly in its letter to the industry. The regulator’s message is clear: existing governance, risk, and operational resilience practices are not keeping pace with how quickly AI is being deployed inside regulated entities. That gap is now a supervisory concern.

    This article explains what APRA and Australia’s updated privacy laws require of Financial Services using AI, and what a Quality Assurance strategy needs to look like in response.

    APRA Sees Three Governance Failures Happening Right Now

    APRA’s concerns are specific and three patterns appear consistently in its observations of the sector.

    Boards are making AI decisions without enough information. Many boards rely on vendor presentations to understand their AI risks. That’s a problem when the vendor is also the one selling the product. APRA expects boards to be capable of independent challenge, and most are not there yet.

    Governance exists on paper, not in practice. Entities have acknowledged that existing prudential standards apply to AI. Few have actually operationalised that acknowledgement with post-deployment monitoring, change management, and model decommissioning are common gaps.

    Supplier dependencies are underexamined. Many FinTechs have concentrated significant AI activity with a single provider without tested exit strategies. When those providers rely on foundation models, training data, or fourth-party services, the chain of accountability becomes opaque. APRA expects visibility all the way through that chain.

    APRA frames AI governance inside existing obligations such as CPS 230 for operational risk, outsourcing standards, and the Financial Accountability Regime (FAR). There is no separate AI rulebook. The existing rulebook applies, and APRA believes most entities are not meeting it.

    Privacy Law Has Changed. Your AI Systems Probably Haven’t.

    The Privacy and Other Legislation Amendment Act 2024 introduced changes that directly affect Financial Services using AI to make decisions about customers.

    The most significant change for AI is the new automated decision-making transparency requirement. If your systems use personal information to make decisions that significantly affect individuals, such as loan approvals or credit scoring, customers will have a right to meaningful information about how those decisions are made. The two-year grace period ends on 10 December 2026.

    Two other changes carry direct operational implications. First, Australians now have a personal right to sue for serious invasions of privacy. AI systems that process sensitive personal data at scale carry real exposure here. Second, the Office of the Australian Information Commissioner (OAIC) has stronger enforcement powers, including tiered civil penalties and the ability to conduct compliance assessments. The OAIC has already issued specific guidance on commercially available AI products.

    APP 11 now explicitly requires “technical and organisational measures” to protect personal information. This is not an IT-only obligation and covers the full range of privacy and security controls around AI systems.

    The Regulator Expects Boards to Lead, Not Delegate

    Both APRA and ASIC have reached the same conclusion from different directions: Financial Services are adopting AI faster than their governance frameworks can handle it, and boards are too far removed from the risk to provide effective oversight.

    APRA’s expectation is that boards understand AI well enough to set strategic direction and challenge assumptions. ASIC’s concern is that licensees are creating consumer harm by deploying AI without updating their risk and compliance frameworks.

    The practical implications are straightforward. Boards and executives need sufficient AI literacy to ask hard questions of management and vendors. Accountability for material AI use cases needs to be named, not distributed. Under FAR, accountable persons need to be identifiable for decisions made by or with AI systems. Staff need training that goes beyond “here is the tool” and covers misuse, limitations, and secure practices.

    Human oversight and ownership is not optional for high-risk decisions. AI can inform those decisions but a named person needs to own them.

    Your AI Systems Need to Fail Safely, Not Just Perform Well

    CPS 230, effective from 1 July 2025, extends operational resilience obligations to AI-enabled systems. This creates concrete requirements that go beyond standard performance monitoring.

    AI systems supporting critical operations need tested fallback processes. That distinction matters when regulators ask for evidence. AI failure modes that need to be planned for include hallucination, silent degradation, and susceptibility to adversarial inputs such as prompt injection or data poisoning.

    Security requirements have also become more specific. AI adoption changes the attack surface with more entry points, faster attack cycles, and new risks from non-human AI agents with system access. APRA expects strong privileged access management, timely patching, hardened configurations, and penetration testing that covers AI-specific vulnerabilities, including AI-generated code.

    Data governance sits underneath all of this. The quality and provenance of training data affects model behaviour. That is now a prudential concern.

    What an APRA-Ready AI QA Strategy Looks Like

    Active supervision of AI is underway with regulators no longer observing and advising. They are assessing and acting. An AI QA strategy needs to be designed for that environment.

    The foundation is a centralised AI inventory: every AI system in use, including third-party tools, mapped to the regulatory obligations it touches under CPS 230, FAR, and the Privacy Act. Without this, gap assessments and audit processes cannot function.

    From that inventory, four capabilities need to be in place.

    Continuous monitoring for bias, drift, and performance degradation. AI models do not stay stable. A model that was accurate at deployment may not be accurate six months later. Automated monitoring catches this before regulators do.

    Independent assurance for high-impact systems. Internal audit and external review processes need to cover AI systems with material customer or operational impact. This cannot be delegated to the team that built the system.

    Testing frameworks built for AI. Standard software testing does not address algorithmic bias, adversarial inputs, or the behaviour of AI-generated code. Testing frameworks need to evolve to cover these risks explicitly.

    Automated compliance tooling. Governance, Risk, and Compliance (GRC) platforms can automate the monitoring and reporting that manual processes cannot keep up with. Predictive compliance tools can identify non-compliant states before they become audit findings. SaaS Security Posture Management (SSPM) tools provide continuous measurement of security controls against regulatory baselines.

    The Gap Between Adoption and Governance Closes in One Direction

    Regulators are not going to slow down their expectations to match the pace of industry governance. The direction of travel is more scrutiny, not less.

    To get ahead of this do three things. Build board-level AI literacy that enables genuine challenge, not just approval. Establish clear ownership for AI decisions at the individual accountability level. Instrument their AI systems for continuous oversight rather than periodic review.

    APRA has been explicit about what it is looking for. The question is whether your governance framework reflects what your AI systems are actually doing, not what the documentation says they do.

  • Is AI Legal for Australian Financial Advisers?

    The clear answer: yes. ASIC maintains a technology-neutral position, meaning existing laws and duties apply regardless of which tools an adviser uses. The harder questions are whether firms are actually ready, whether their governance frameworks can handle the risks, and whether their clients are protected.

    Context makes this urgent. ASIC’s March 2026 Moneysmart research found that nearly one in five Gen Z Australians are already turning to AI platforms for financial information, and 64% of them trust what those platforms tell them. Young Australians are not waiting for the profession to catch up. They are arriving at adviser meetings having already formed views shaped by tools that carry no professional accountability. The gap between what AI tells clients and what advisers are obligated to provide is where the compliance exposure lives.

    ASIC’s Technology-Neutral Stance Means There Is No AI Loophole

    ASIC does not regulate the tool. It regulates the outcome. Under Section 912A of the Corporations Act, licensees must provide financial services efficiently, honestly, and fairly. That obligation does not bend because an AI produced the output. An adviser who relies on a hallucinated AI recommendation and passes it to a client has not met their best interests duty, even if they did not know the output was wrong.

    This framing matters because it shifts the question from “can we use AI?” to “can we demonstrate that using AI still results in compliant, appropriate advice?” There is no grey area or loophole here. The full weight of regulatory compliance rests on the adviser, regardless of where in the process a tool was involved.

    Five Risks That Governance Frameworks Must Address

    The risks of AI in financial advice are the predictable consequences of deploying powerful tools without matching governance.

    Hallucinated outputs are the most visible risk. Generative AI models produce confident responses that can be factually wrong or fabricated. In the context of financial advice, a plausible-sounding but incorrect recommendation does not become safer because it came from a sophisticated model. It becomes harder to detect.

    Privacy and cybersecurity exposure follow directly from how AI is used. Inputting sensitive client data into third-party or offshore AI systems raises questions about data sovereignty, storage, and access that most standard engagements with AI vendors do not fully resolve. The adviser’s Privacy Act obligations do not transfer to the vendor.

    Governance gaps are perhaps the most systemic risk. Firms are integrating AI faster than they are updating the policies, controls, and risk frameworks that govern it. When something goes wrong – and in any system operating at scale, something eventually will – the question regulators ask is not whether a tool was involved. It is whether the firm had adequate systems to prevent it.

    Bias and discriminatory outcomes emerge from training data that reflects historical patterns. An AI model trained on historical financial data may systematically produce outputs that disadvantage certain client segments. ASIC has specifically flagged this risk. The adviser who acts on biased AI outputs carries the liability, not the model provider.

    Lack of transparency compounds all the above. When an adviser cannot explain how a recommendation was reached, they cannot discharge their obligation to provide clear, appropriate advice. An invisible decision is an indefensible one. Regulators and clients both expect explanations. “The AI suggested it” is not one.

    These risks do not operate in isolation. A governance gap makes hallucinations harder to catch. Opacity makes bias harder to detect. The system fails as a system, not just at individual failure points.

    “Human in the Loop” Only Works If the Humans Understand What They’re Looking At

    ASIC expects human oversight of AI-assisted advice, and the principle is sound. But human oversight is only meaningful when the humans reviewing AI outputs have the knowledge, time, and tools to identify problems. A cursory review of a confidently-worded AI recommendation does not constitute oversight. It is the appearance of oversight and it is where liability silently accumulates.

    This is the expertise deficit that most firms have not yet confronted. The three traditional lines of defence – risk management, compliance, and internal audit – were built to review human decisions. They were not built to evaluate whether an AI model’s output is appropriate, whether its training data introduced bias, or whether its reasoning can withstand regulatory scrutiny. Without AI literacy at each of these levels, the oversight that firms believe they have is largely theoretical.

    Management of third-party risk sits alongside this. The fact that an AI model was built and maintained by an external vendor does not transfer accountability for its outputs. Licensees remain fully responsible for the work of outsourced functions, including the AI tools they select, configure, and deploy. Due diligence on AI vendors, for example their data handling, model governance, and liability terms, is not optional. Outsourced technology is not outsourced liability.

    Documentation requirements remain unchanged. Every piece of advice influenced by AI must be recorded, with clear evidence of how AI was used and how human judgement was applied to its outputs. This is both a compliance requirement and a practical protection. If a complaint arises, the paper trail matters.

    Build Governance Before You Deploy and Update It Continuously

    ASIC’s guidance to licensees is direct: establish comprehensive AI governance frameworks, including specific policies, procedures, and codes of conduct, before integrating new AI tools into operations. This is the principle that most firms are currently inverting.

    Risk management policies need to be updated specifically for AI. Generic technology risk frameworks do not capture the distinct failure modes of generative AI eg, hallucination, model drift, bias, and opacity, and they do not map onto the advice obligations that apply in this sector.

    There is also a static policy problem. A governance document written for the AI tools of 2024 does not govern the AI tools of 2026. The technology is evolving continuously. Regulatory interpretation is evolving continuously. A policy that is not actively maintained becomes a liability of its own, evidence that the firm believed it had addressed a risk it had actually frozen in time and ignored.

    Disclosure deserves deliberate attention. When AI directly influences advice, clients have a reasonable interest in knowing. How and when to disclose AI involvement is a governance decision that should be made before deployment, not improvised after a client raises the question.

    Adequate resourcing of AI oversight with both the technology infrastructure and the human capacity to monitor it, is part of the obligation. Underfunding oversight while deploying AI at scale is not a cost-saving measure. It is a risk creation measure.

    Governance Is a Process, Not a Checkbox

    The firms that navigate AI well will not be the ones that asked legal for a sign-off and moved on. They will be the ones that built systems designed to catch predictable failures before those failures reach clients, and then kept updating those systems as the technology and regulatory environment evolved.

    ASIC has mapped the problem. The gap between where most firms are and where they need to be is a governance gap, and it is one that only deliberate, ongoing work will close.

    AI offers genuine value to financial advisers and, through them, to clients. Getting there requires treating governance not as the price of admission, but as the foundation everything else is built on.


    For a detailed examination of the nine AI governance gaps that apply to AFSL holders, see the Galdren report on AFSL risk.

  • Navigating AI Risks to Design and Distribution Obligations

    The rapid adoption of AI across the Australian financial services sector has moved from competitive advantage to operational necessity. But the speed of adoption has created a problem regulators have noticed: the gap between deploying AI and actually governing it.

    For Australian Financial Services licensees, that gap is not a technology problem. It is a compliance problem. ASIC and APRA are increasingly focused on whether firms can demonstrate they understand what their AI tools are doing, why, and what happens when things go wrong. The Design and Distribution Obligations (DDO) and the requirement to act “efficiently, honestly, and fairly” apply to AI-driven decisions the same way they apply to any other business process. The AI does not change the obligation, it just makes it harder to see when the obligation is being breached.

    1. The Risk Landscape: What You Probably Haven’t Tested

    AI introduces risk profiles that are genuinely different from traditional software, not because the regulatory obligations are different, but because the failure modes are less visible and often more systematic. Three characteristics make AI uniquely difficult to govern: it can be confidently wrong, its errors can be invisible in aggregate data, and its behaviour can change without anyone deliberately changing it.

    When the AI is wrong, it doesn’t know it

    Traditional software fails visibly; an error code, a system crash, a blank screen. AI models fail silently, generating outputs that are plausible, formatted correctly, and entirely false. In financial services, this matters immediately.

    Consider an AI chatbot deployed to answer client questions about a managed fund. A client asks whether the fund has capital protection. The AI, drawing on training data that includes older product literature, confirms that it does. The fund’s current Product Disclosure Statement says otherwise. The client invests based on the AI’s answer. No one in your firm knows this happened, because the chatbot logged the interaction as a resolved query, not a complaint.

    This is a predictable consequence of deploying a generative AI tool without grounding it against your current, authoritative product documentation and without a mechanism to detect when a client has acted on AI-generated information that contradicts your PDS. Under DDO, distribution based on false product information is a breach, regardless of whether a human or an AI provided it.

    Other failure modes in this category that firms often underestimate:

    • AI-generated PDS summaries distributed to advisers: If your AI summarises a PDS to help advisers understand a product, and that summary contains a material error, every adviser who relied on it has potentially distributed outside the TMD. The summary feels authoritative. No one fact-checks it against the original.
    • Client mis-categorisation at scale: An AI categorising clients by risk profile or financial sophistication can be wrong about individual clients in ways that are invisible in aggregate accuracy figures. If the tool is 94% accurate overall but systematically worse for clients over 65, a cohort you’ve never tested separately, you may have a bias problem embedded in every interaction with that demographic.
    • Autonomous AI agent conduct: An AI agent authorised to take actions on behalf of clients, scheduling, executing instructions, communicating, can engage in conduct that would constitute misconduct if a human adviser did it. The firm bears that liability.

    The complaint that never becomes a complaint

    One of the more consequential blind spots in AI-assisted customer service is the complaint that gets resolved before it is recorded.

    When a human customer service representative handles a complaint, there is generally a process: the interaction is flagged, the complaint is logged, it enters the complaints management system. When an AI handles the same interaction, the outcome depends entirely on how the AI has been configured to classify what it receives.

    A client who says “I’m really unhappy with how this was handled and I want something done about it” may receive a sympathetic, helpful AI response that resolves the surface issue. The AI logs the interaction as a successful resolution. In your complaints data, it does not appear. In your DDO reporting, it does not appear. In your significant dealings analysis, it does not appear.

    If this happens at scale, because AI tends to apply consistent classification errors consistently, your regulatory reporting may systematically understate the number and nature of client complaints. Section 994F(4) requires accurate complaints reporting. An AI that is resolving complaints before they can be counted is not a compliance solution, but a compliance risk disguised as good customer service.

    Model drift: the AI that changed without anyone changing it

    AI models are not static. Their performance is anchored to the data they were trained on, and when the world diverges from that data, their outputs become less reliable. Gradually, quietly, and without any visible system failure.

    Consider a creditworthiness or suitability assessment tool trained on client data from 2020 to 2022. That tool learned patterns from a period of historically low interest rates, rising asset values, and a particular economic environment. If it has not been re-validated against current conditions, it may be applying assumptions that no longer hold. Clients who would have been correctly categorised as suitable for a particular product four years ago may no longer be. The tool does not know this. It continues to categorise them as suitable.

    This is model drift. It is the AI equivalent of a policy that nobody has reviewed in years. The difference is that when a policy drifts out of date, someone usually notices. When an AI model drifts, the outputs often look the same, the categories still get filled, the reports still get generated, and the process still runs. The problem is invisible until something goes wrong at scale.

    2. Compliance and the DDO Framework

    The DDO framework, as outlined in ASIC Information Sheet 264, does not make exceptions for AI. Every obligation that applies to a human process applies equally when that process is automated. The relevant question is not whether AI was involved, but whether the outcome met the regulatory standard.

    AI in product design

    When an issuer uses AI to design a financial product or determine its Target Market Determination, the AI must correctly account for the likely objectives, financial situation, and needs of the target class. ASIC’s framing is clear: the target market is the class of consumers for whom the product is likely to be appropriate, and the TMD must accurately describe them.

    If the AI that helps determine your TMD has been trained on data that overrepresents a particular client profile, for example urban, higher income, digitally engaged, it may systematically underweight the needs of clients outside that profile. The TMD looks complete. But it may not accurately describe the full range of consumers for whom the product is or is not appropriate.

    AI in distribution and monitoring

    Distributors are increasingly using AI to monitor distribution conduct and identify significant dealings inconsistent with the TMD. This is exactly the kind of application DDO contemplates. But the reliance on AI for this function creates its own obligations.

    The AI must be capable of correctly applying the definition of a “significant dealing”. It must detect when the proportion of consumers outside the target market reaches a threshold that triggers reporting. It must not classify borderline cases consistently in one direction. And it must ensure that the “reasonable steps” required under DDO are documented, not just assumed.

    Automation that speeds up distribution is good, but without preserving the compliance steps is not a solution. It is a breach that scales.

    The “efficiently, honestly, and fairly” standard

    Section 912A of the Corporations Act requires licensees to provide services efficiently, honestly, and fairly. ASIC has indicated that a lack of explainability in AI decisions may fail the fairness test. If an AI makes a recommendation or assessment, and the firm cannot explain why, that opacity is itself a compliance risk.

    This is particularly relevant for advice tools. An AI that generates recommendations not tailored to the individual client’s circumstances, because it is applying a generalised model rather than genuinely assessing the specific client, may fail the best interests duty and suitability requirements, even if the recommendation appears reasonable on its face, or only used internally.

    3. Governance Checklist: Questions Worth Asking

    The questions below are not designed to confirm that your governance is in order. They are designed to find the places where it isn’t. If you find yourself answering a question with “I assume so” or “someone else handles that,” you have found something worth investigating.

    Accountability

    • If ASIC asked you today which AI tools are making or materially influencing client-facing decisions, could you hand them a complete list within an hour?
    • Name the person who would receive the call if one of your AI tools caused a compliance breach tomorrow. Does that person know they are accountable for it?
    • Is there a human being who has read the AI vendor’s model documentation, not just the sales materials, and signed off on its fitness for your specific use case?

    Distribution and TMD integrity

    • If your AI is involved in the distribution process, can you demonstrate that it has never bypassed a “reasonable step”, or do you assume the vendor built those in?
    • Have you tested whether your AI correctly identifies a “significant dealing”, not in documentation, but in a live scenario with borderline cases?
    • When did you last audit the AI’s categorisation outputs against your TMD to check for systematic misalignment?

    Complaints and reporting

    • Could your AI chatbot resolve a client complaint without it ever appearing in your complaints data? If you don’t know, that is your answer.
    • Does the data flowing from your AI customer service tools into your DDO reporting get validated by someone who did not build the tool?
    • Have you cross-checked your AI-generated complaints figures against any other data source, e.g. inbound call volumes, churn patterns, adviser escalations, to check they are plausible?

    Model integrity and drift

    • What economic conditions was your suitability or creditworthiness AI trained on? Have you re-validated it since those conditions changed materially?
    • Is there a scheduled review process for your AI models, or does re-validation happen only when someone notices a problem?
    • If a vendor updated their model overnight, would you know? Would anything in your process change?

    Transparency and explainability

    • For each AI tool used in a client-facing decision, can you produce a reason code or plain-language explanation for a specific output? Or does “the model decided” end the conversation?
    • Have you tested whether your AI-generated PDS summaries are accurate against current PDS documents, not when the tool was deployed, but recently?
    • If a client asked why the AI categorised them the way it did, could anyone in your firm answer that question?

    Bias and fairness

    • Have you tested your AI’s accuracy across different client demographics, by age, income bracket, geography, or product experience, rather than only in aggregate?
    • If your AI performs significantly worse for a particular client cohort, would your monitoring processes detect it? What would trigger the alert?

    Incident response

    • Does your organisation have a kill switch for each AI tool. A documented, tested mechanism to halt its operation. Or does the plan exist only in theory?
    • When did you last test the kill switch? Triggering it to ensure it works as expected?
    • If an AI tool needed to be turned off today, what would happen to the clients and processes that depend on it? Has anyone mapped that?

    Vendor and third-party risk

    • Does your AI vendor notify you before making changes to their model? What is your contractual right to know when the model you approved is no longer the model you’re running?
    • Has your vendor been assessed for compliance with the Privacy Act? Have you verified that client data used to train or improve their model has not been retained in ways you did not authorise?
    • Are there mechanisms in place to prevent data poisoning, deliberate or accidental corruption of the inputs your AI relies on?

    Conclusion

    AI offers real opperational advances, but brings with it governance risks not fully covered by your existing controls. Use the questions offered here as a starting point, not a definitive checklist. Firms that haven’t built the practice of thinking through their AI controls before they need it will be caught short.

  • AI Risk Governance: What Boards Need to Know and Do

    AI risk is different from technology risk in three important ways, and boards that treat it as “another IT matter” will discover this difference the hard way.

    First, AI risk is legally novel. Air Canada learned this in 2024 when a court held them liable for their chatbot’s false advice to a customer, establishing that companies are responsible for what their AI systems communicate. Board members who assume their existing liability frameworks cover AI decisions should verify that assumption with their legal counsel.

    Second, AI risk is regulatory and accelerating. The EU AI Act entered into force on 1 August 2024, with prohibitions on specific AI practices applying from 2 February 2025 and high-risk system obligations applying from August 2026. General-purpose AI models, including foundation models, have faced their own compliance obligations since August 2025. Any board with operations or customers in the EU is already inside this regime. Australia is developing mandatory guardrails for high-risk AI applications. The UK, Canada, and Singapore are all advancing their own frameworks. Boards have a narrowing window to build compliance capabilities before enforcement intensifies.

    Third, AI risk compounds invisibly. A biased model, a shadow AI tool used by a single business unit, or an agentic system making decisions outside its intended scope may produce no visible signal until the harm is substantial. By the time the problem surfaces – through a regulatory inquiry, a customer complaint, or media scrutiny – the exposure has been accumulating for months or years.

    These three characteristics mean that AI risk requires active board governance, not delegation with periodic updates.

    Setting AI Risk Appetite: The Non-Delegable Responsibility

    Risk appetite defines the level and type of AI risk the organisation is willing to accept in pursuit of its strategic objectives. Without a clear appetite statement, management cannot make consistent decisions about which AI applications to deploy, which to restrict, and which to prohibit entirely. The result is governance driven by individual preferences rather than organisational policy.

    A useful AI risk appetite statement addresses several dimensions:

    Use case tolerance. What categories of AI-assisted or AI-driven decisions is the organisation comfortable with? Customer-facing communications? Credit or underwriting decisions? Medical or clinical support? Autonomous operational decisions? Each category carries a different risk profile and may require different levels of human oversight.

    Data risk tolerance. What types of data can AI systems process, including those operated by third parties? What are the minimum standards for data residency, privacy protection, and consent?

    Error tolerance. What rate of AI error is acceptable in different contexts? An AI that recommends products with 95% accuracy may be excellent in a marketing context and wholly inadequate in a clinical one.

    Shadow AI tolerance. What is the organisation’s position on unsanctioned AI tool usage? Zero tolerance with enforcement mechanisms? Managed tolerance with a disclosure and review process? The answer has significant implications for both risk exposure and employee behaviour.

    Agentic action tolerance. What categories of autonomous action can AI systems take without human approval? What decisions must always involve a human?

    The board should approve the AI risk appetite statement, review it at least annually, and ensure management translates it into operational policy. The statement should be a living document, as AI capabilities evolve and as the organisation’s experience deepens, the appetite should evolve with it.

    Defining Ownership: A RACI Matrix for AI Risk Management

    Clarity about who does what is how accountability is created and maintained. Organisations without clear AI risk ownership typically discover this during an incident, when everyone assumed someone else was managing the risk.

    A Responsible, Accountable, Consulted, Informed (RACI) matrix defines ownership across the AI risk lifecycle. The following structure provides a starting framework:

    Role/StakeholderResponsibleAccountableConsultedInformed
    Board of DirectorsAI Risk Governance, Risk Appetite, Strategic AlignmentC-Suite, Key Stakeholders
    CEOStrategic Execution, CultureOverall Business PerformanceBoard, C-SuiteAll Employees
    CAIO / CTO / CIOAI Implementation, Technical GovernanceAI Programme Success, Technical RiskBusiness Units, CRO/CISOBoard, C-Suite
    CRO / CISORisk Assessment, Controls, MonitoringEnterprise Risk Posture, AI SecurityCAIO/CTO/CIO, LegalBoard, C-Suite, Business Units
    Legal & ComplianceRegulatory Adherence, Ethical GuidelinesLegal & Ethical ComplianceCAIO, CRO/CISOBoard, C-Suite, Business Units
    Business UnitsAI Use Case Identification, Day-to-Day OperationsLocal AI Risk ManagementCAIO, CRO/CISOC-Suite

    Three practical observations about making this work.

    First, the accountability column should contain people’s names, not just role titles. Accountability without a named individual is diffuse and ineffective.

    Second, escalation paths need to be specified before they are needed. When a business unit identifies an AI risk that exceeds their local authority to manage, who do they escalate to, by what mechanism, and within what timeframe? Escalation paths that are unclear in practice are impossible to navigate under pressure.

    Third, the RACI should be revisited when roles change. The appointment of a Chief AI Officer, a restructure of the technology function, or a significant expansion of AI usage all represent moments to confirm the matrix remains accurate.

    The Regulatory Landscape: What Boards Cannot Afford to Miss

    AI regulation is no longer a future consideration. It is a present operational requirement for many organisations, and the compliance window for others is closing.

    The EU AI Act is the most comprehensive AI regulatory framework currently in force. It applies to organisations that operate in the EU, offer products or services to EU customers, or whose AI systems affect people in the EU. The Act takes a risk-tiered approach: certain AI applications are prohibited outright (including social scoring by governments and most real-time biometric identification in public spaces), high-risk applications face substantial compliance obligations from August 2026, and limited-risk applications carry transparency requirements. General-purpose AI models face compliance obligations that have applied since August 2025. Boards of organisations with EU exposure should have received a legal opinion on their obligations under this Act. If they have not, that is an immediate action item.

    Australia’s AI governance framework is evolving. The federal government has published voluntary AI safety standards and consulted on mandatory guardrails for high-risk AI contexts. Sector regulators including APRA and ASIC have issued guidance on AI use in financial services that creates obligations for entities within their remit. Organisations should not wait for mandatory requirements to build their governance infrastructure, the frameworks being published now signal the direction of future obligations.

    Sector-specific obligations often extend further than general AI legislation. Healthcare, financial services, government, and legal sectors all face AI-related requirements through their existing regulatory frameworks that predate dedicated AI legislation.

    Liability exposure through existing law is also significant. Consumer protection laws, privacy legislation, anti-discrimination statutes, and professional liability frameworks all potentially apply to AI-generated decisions, even where specific AI legislation does not. The Air Canada case was decided under existing consumer protection principles, not under AI-specific law.

    Boards should require management to produce an annual regulatory risk assessment that maps the organisation’s AI applications against applicable obligations and identifies compliance gaps. This assessment should inform both the risk appetite statement and the compliance programme.

    Integrating AI into Enterprise Risk Management

    AI risk does not sit alongside other organisational risks, but runs through them. A single AI failure can simultaneously generate operational, reputational, strategic, financial, and compliance consequences. Governance is stronger when AI risk sits within existing frameworks rather than operating as a parallel system that the board rarely sees.

    Established frameworks including the NIST AI Risk Management Framework and ISO/IEC 42001 provide practical methodologies for this integration. Both are complementary to existing enterprise risk management structures, not replacements for them.

    Integration involves five elements that go beyond policy alignment:

    Shared risk language. AI risks should be described using the same terms, such as likelihood, consequence, risk rating, that the organisation uses for all other risks. This enables consistent comparison, prioritisation, and resource allocation.

    Consolidated risk register. AI risks should appear in the enterprise risk register, not in a separate AI risk log the board rarely sees. Significant AI risks belong alongside cyber risk, credit risk, and operational risk.

    Integrated assurance. Internal audit, external audit, and risk review functions should incorporate AI risk into their scope. Organisations that have not extended assurance activities to cover AI systems have visibility gaps.

    Risk-adjusted approval. Capital investment decisions, product launches, and operational changes involving AI should go through the standard risk approval process, not a separate AI governance channel that operates independently.

    Board-level visibility. AI risks rated as significant should be reported to the board through the normal reporting cycle, not only when an incident occurs.

    Governing Agentic AI: The Frontier That Requires Immediate Attention

    Agentic AI systems – those capable of autonomously executing multi-step tasks, interacting with external systems, and making sequential decisions without human intervention – represent a meaningful shift in governance requirements.

    Traditional AI governance assumes a human reviews outputs before action is taken. A model recommends; a human decides. Agentic systems break this assumption because an agentic AI that can browse the web, draft and send communications, execute transactions, modify data, or interact with third-party services is taking actions on behalf of the organisation without step-by-step human approval.

    The governance implications are significant.

    Scope creep is a material risk. Agentic systems operating within broad parameters may take actions that were not intended or anticipated. A system instructed to “manage our social media presence” could interpret that instruction in ways that create significant reputational or regulatory exposure.

    Audit trails may be incomplete. If an agentic system takes a harmful action, can your organisation reconstruct what it did, why, and what data it accessed? The absence of comprehensive logging is both a governance failure and a potential regulatory issue.

    Liability allocation is untested. When an agentic AI causes harm – through a contract it was not authorised to enter, information it disclosed, or a decision it made – who is liable? This question does not have settled legal answers in most jurisdictions, which means the organisation is carrying legal risk it cannot fully price.

    Human override mechanisms must be designed, not assumed. Boards should confirm that every agentic system in use has clear mechanisms for human intervention and override, defined limits on autonomous action, and monitoring that surfaces exceptions for human review.

    The governance standard for agentic AI should be more stringent than for traditional AI applications, not less. The autonomy that makes these systems valuable also makes them more capable of causing harm without human intervention.

    Responsible AI: Bias, Ethics, and Explainability as Governance Obligations

    Responsible AI is not a values statement. It is a governance approach with direct legal, financial, and reputational consequences.

    Bias All models have some degree of bias. The question for boards is whether that bias produces outcomes that are discriminatory, unfair, or harmful, and whether the organisation has systems to detect and correct it. Amazon’s recruitment AI is a well-documented example, but similar issues have arisen in credit scoring, healthcare triage, recidivism prediction, and insurance pricing.

    Explainability The EU AI Act, Australia’s privacy framework, and sector-specific regulations impose explanation rights on individuals subject to automated decisions. A credit refusal driven by an AI model that cannot explain its reasoning exposes the organisation to legal challenge and regulatory action.

    Boards should ask management two questions. First: for which AI applications can we explain outcomes to the individuals affected? Second: for applications where we cannot, have we assessed the regulatory and legal exposure, and is that exposure within our risk appetite?

    Ethical review Boards should confirm that ethical review is a standard gate in the AI development and procurement lifecycle, not an optional add-on.

    Preparing for When Things Go Wrong: AI Incident Response

    Most organisations have cyber incident response plans. Far fewer have AI incident response plans. This is a governance gap, because AI incidents have characteristics that make standard incident response procedures inadequate.

    AI incidents may be gradual rather than sudden. A model that begins producing biased outputs, a shadow AI tool that exfiltrates data incrementally, or an agentic system that slowly exceeds its intended scope may not trigger any of the monitoring thresholds designed to detect a discrete security event.

    AI incidents may be legally ambiguous. Whether an AI failure constitutes a data breach, a product defect, a regulatory violation, or a contractual non-performance will depend on the specific facts. Incident response procedures that trigger clear legal and regulatory notifications for cyber events may not have equivalent clarity for AI events.

    AI incidents may involve third parties. If the failure originates in a vendor’s model, who leads the response? What notification obligations apply? What contractual remedies are available?

    A functional AI incident response plan addresses:

    • Detection mechanisms: How does the organisation identify an AI incident? What monitoring is in place, and what thresholds trigger escalation?
    • Classification criteria: What distinguishes a significant AI incident from a routine performance issue?
    • Escalation paths: Who is notified, in what sequence, and within what timeframe?
    • Regulatory notification obligations: Which incidents require notification to regulators, customers, or third parties, and within what timeframes?
    • Containment and remediation: What is the process for taking a failing AI system offline, reversing harmful outputs where possible, and restoring safe operation?
    • Post-incident review: How does the organisation learn from AI incidents and prevent recurrence?

    The board should approve the AI incident response plan and receive reports of significant AI incidents. An incident that recurs after a formal post-incident review is not a technical failure but a governance failure because the system that produced the first incident is still in place.

    Third-Party AI Risk: The Audit Your Lawyers Wish You Had Done Earlier

    The AI tools most organisations use are predominantly purchased, not built. This means the organisation’s AI risk profile is substantially determined by the risk management practices of its vendors, practices the organisation has limited visibility unless it actively creates it.

    Third-party AI risk audits should be a standard requirement before AI vendor engagement and a regular obligation during it. The audit framework should address:

    Data handling. Where is your data stored and processed? Is data used to train or improve the vendor’s models? What happens to your data if the vendor relationship ends? These questions have direct privacy law implications and, in some sectors, such as healthcare, financial services, and government, may determine whether the vendor engagement is permissible at all.

    Model transparency. Can the vendor explain how their model produces outputs in your use case? What bias testing have they conducted, and what were the results? What is their process for identifying and addressing bias in production?

    Security posture. What certifications does the vendor hold? SOC 2 Type II and ISO 42001 are baseline indicators of process maturity, not guarantees of security, but their absence is a significant flag. What are the vendor’s vulnerability disclosure and patch management practices?

    Contractual protections. Does the contract include the right to audit? Does it allocate liability for harm caused by model failures or biased outputs? Does it specify notification obligations if the vendor experiences a data breach or model failure affecting your data? These provisions are far easier to negotiate before engagement than after an incident.

    Lifecycle commitments. What is the vendor’s commitment to model maintenance, performance monitoring, and eventual decommissioning? A vendor who provides no visibility into their model update process is a vendor whose risk profile you cannot manage.

    The audit programme should be proportionate to risk. High-risk or high-dependency vendors warrant more intensive scrutiny than low-risk peripheral tools. But the programme should be systematic and documented, not ad hoc.

    Lifecycle Governance: Risk Follows the AI From Design to Decommissioning

    AI risk is not static. A model that is safe and compliant at deployment may not remain so as the environment changes, as the model drifts, or as the organisation’s use case evolves. Lifecycle governance embeds risk management into every phase of an AI system’s existence.

    Plan and Design. Risk assessment at this stage identifies the risk profile of the proposed application before investment is committed. Ethical review, regulatory compliance assessment, and risk appetite alignment should all occur here, not after the system is built.

    Data Collection and Processing. Data quality, representativeness, and provenance determine model quality. Governance at this phase includes data provenance documentation, bias assessment of training data, and privacy compliance review.

    Model Building and Training. Bias testing, explainability assessment, and security review should occur before deployment. Models that cannot pass these tests should not proceed.

    Deployment and Use. Go-live governance includes user access controls, monitoring activation, and human oversight confirmation. The deployment decision should be a formal approval, not an informal go-ahead.

    Monitoring and Maintenance. Ongoing monitoring detects model drift, performance degradation, and emerging risks. The monitoring framework should specify thresholds that trigger review, the process for model retraining or replacement, and the board-level reporting that occurs when significant risks are identified.

    Decommissioning. Systems that are retired should be formally decommissioned, with data deletion or archival handled in accordance with retention obligations, and documentation retained for regulatory and legal purposes.

    The board should confirm that management has a documented lifecycle governance process and that it is consistently applied across all significant AI systems, not just those that are internally developed.

    Building AI Literacy in the Boardroom

    Boards cannot govern what they do not understand. This does not require every board member to become a technical AI expert. It does require sufficient collective literacy to ask good questions, evaluate management’s responses, and recognise when the board is being told what it wants to hear rather than what it needs to know.

    AI literacy at board level means understanding:

    • The difference between machine learning, generative AI, and agentic AI, and the different risk profiles they carry
    • How training data shapes model behaviour, and why historical data can embed historical biases
    • Why AI systems can produce confident-sounding outputs that are factually wrong
    • What “model drift” means and why it matters for ongoing governance
    • What the EU AI Act and relevant local frameworks require of the organisation
    • What “shadow AI” looks like in practice and why it is a governance problem, not just an IT problem

    Treat AI education as an ongoing obligation, not a one-time orientation. AI capabilities, risks, and regulatory frameworks are evolving rapidly. A board that was adequately informed twelve months ago may have significant knowledge gaps today.

    Conclusion: Governance That Keeps Pace With the Technology

    The organisations that will harness AI’s advantages while managing its risks are those where governance keeps pace with deployment. Where the board has set a clear risk appetite that management can operationalise. Where ownership is unambiguous and accountability is named. Where regulatory obligations are mapped and managed. Where ethical risks are assessed alongside financial ones. Where incident response is prepared, not improvised.

    None of this requires the board to become a technical body. It requires the board to exercise the same disciplined governance over AI risk that it exercises over financial risk, operational risk, and strategic risk.

    The question is whether your governance is structured to give the board the visibility, the accountability mechanisms, and the decision frameworks it needs to protect the organisation and create sustainable value from AI.

    If you are not confident the answer is yes, start with the risk appetite statement. Everything else in this article depends on it.

  • Systems Thinking: A Practical Toolkit for AI

    The Problem with How Most Businesses Deploy AI

    Most AI deployments fail not because the technology was wrong, but because the system around the technology was misunderstood.

    Teams optimise a model in isolation, then wonder why the outcomes are biased. They satisfy a compliance checkbox, then discover the requirement touched fifteen other processes they hadn’t mapped. They fix a symptom – only to watch the underlying problem resurface somewhere else six months later.

    Systems thinking is the discipline that closes this gap. It shifts attention from the components of an AI deployment to the relationships between them, such as the feedback loops, delays, and interdependencies that determine whether an AI initiative actually delivers what was intended.

    This toolkit gives business leaders and their teams a practical entry point into that discipline. It is structured around three phases of AI deployment and eleven tools, each adapted from established systems thinking methodology and grounded in Australia’s regulatory and ethical landscape.


    Table of Contents

    Phase 1: Confirm the Goal and Understand the AI System

    • Principle 1: Identify Key Issues and Establish a Collaborating Community with a Shared Goal
    • Tool 1 – Rich Pictures: Expressing a Summary of the AI System
    • Tool 2 – The Stakeholder Model: Understanding Diverse Views of the AI System
    • Principle 2: Reach a Shared Understanding of the AI Problem
    • Tool 3 – Context Diagrams: Identifying AI System Boundaries
    • Tool 4 – Behaviour Over Time Graphs and AI System Problem Statements: Articulating Your Problem and Goal
    • Tool 5 – Identifying Enablers and Inhibitors: Exploring the Causes of Your AI Problem
    • Tool 6 – Creating a Causal Loop Diagram: Mapping Your AI System
    • Tool 7 – Causal Loop Diagram: Analysis and Narrative

    Phase 2: Co-design and Test Possible AI Interventions

    • Principle 3: Explore Interventions Using an Understanding of the AI System and Its Possible Leverage Points
    • Tool 8 – Identifying AI Systems Leverage
    • Principle 4: Test the Ideas
    • Tool 9 – Stock and Flow Diagrams for AI Systems
    • Tool 10 – Theory of Change Maps for AI Initiatives

    Phase 3: Implement Systemic AI Interventions, Monitor and Evaluate

    • Principle 5: Monitor, Evaluate, and Learn with the Community
    • Tool 11 – Monitoring and Evaluation Strategy for AI Systems

    What Does It Mean to Take a Whole-System Approach to AI?

    An AI system is not a piece of software. It is a dynamic arrangement of algorithms, data pipelines, organisational cultures, human workflows, customer relationships, and regulatory obligations – all of which interact continuously and produce outcomes no single component could generate alone.

    This creates a specific kind of risk that technical expertise alone cannot manage: the risk of solving the wrong problem. Addressing algorithmic bias, for example, isn’t a matter of adjusting a model. It requires tracing a causal chain from data collection practices through to deployment context, organisational incentives, and customer impact. That chain is a system, and it needs to be mapped and understood as one.

    Systems thinking offers several concrete advantages for businesses navigating this terrain. It surfaces root causes rather than symptoms, which means interventions last. It maps feedback loops and interdependencies before deployment, which means unintended consequences are anticipated rather than discovered. It builds shared understanding across technical, legal, and executive stakeholders, which means AI decisions carry broader organisational legitimacy. And it embeds regulatory requirements into the design of AI systems rather than treating them as afterthoughts, such as Australia’s Voluntary AI Safety Standard (VAISS), the Privacy Act 1988, the Australian AI Ethics Principles, and obligations under Section 912A of the Corporations Act for AFSL holders.

    The toolkit in this document is most valuable in specific situations: when designing AI strategy before significant resources are committed; when an AI system is producing biased or unexpected outcomes; when data governance obligations require a privacy-by-design approach; when regulatory compliance needs to be integrated into the AI development lifecycle; and when you need to evaluate the broader and longer-term impacts of a deployment beyond model accuracy.


    How This Toolkit Works

    The eleven tools in this document are not a recipe to follow once and set aside. Their value is iterative: each tool builds on the last, and the process of mapping, discussing, and refining is as important as the outputs it produces.

    The toolkit is built around three phases that mirror the natural lifecycle of an AI initiative. Phase 1 establishes a clear goal and a comprehensive understanding of the system you’re working with. Phase 2 identifies where and how to intervene, using modelling and simulation to test ideas before they’re deployed. Phase 3 embeds the discipline of ongoing monitoring, evaluation, and adaptation.

    Data and its visualisation run through all three phases. Effective data work here does more than measure performance. It exposes the AI system’s behaviour over time, reveals disparities across demographic groups, tracks data feedback loops, and provides the evidence base for regulatory compliance. Visualising data lineage, in particular, is one of the most practical tools for identifying the upstream sources of downstream problems.


    Phase 1: Confirm the Goal and Understand the AI System

    Principle 1: Identify Key Issues and Establish a Collaborating Community with a Shared Goal

    Effective AI deployment starts with understanding the problem and the ecosystem it inhabits. The organisations that get this right bring together diverse perspectives early, not as a consultation exercise, but as the primary means of developing a shared, accurate picture of what the AI system actually involves.

    Tool 1 – Rich Pictures: Map the Full AI System Before You Build It

    Purpose: To create a visual representation of the AI system that captures relationships, stakeholder perspectives, data flows, and the regulatory environment – including the qualitative and human elements that formal diagrams miss.

    How to Apply to AI:

    1. Identify the Core AI System: Place the AI system (for example, an AI-powered recommendation engine or automated decision-making tool) at the centre.
    2. Map Stakeholders: Draw all relevant stakeholders: internal teams (data scientists, legal, ethics, business units), external partners (AI vendors, data providers), customers, regulators (ASIC, OAIC), and affected communities. Represent their relationships and relative influence.
    3. Illustrate Data Flows: Show how data enters, moves through, and exits the system. Highlight data sources, transformation points, and where personal information is processed.
    4. Depict Key Processes and Interactions: Sketch human-AI interactions, decision points, feedback loops (for example, model retraining based on user feedback), and automated processes.
    5. Capture Perceptions and Emotions: Use symbols or speech bubbles to represent stakeholders’ concerns (privacy risk, bias), expectations, and conflicting views.
    6. Include the Regulatory Context: Represent relevant Australian regulations – the Privacy Act 1988, AFSL obligations, the Australian AI Ethics Principles – and show how they interact with the system.

    Outcomes: A shared, holistic picture of the AI system that surfaces complexity and disagreement early, before they become expensive problems. Rich pictures are particularly effective at exposing the human and organisational dimensions that technical documentation omits.


    Tool 2 – The Stakeholder Model: Understand Who Defines Success

    Purpose: To systematically identify and analyse all stakeholders affected by the AI system, understand their differing perspectives on its goals, and establish the basis for a collaborating community.

    How to Apply to AI:

    1. Identify All Stakeholders: List everyone who has a stake in the AI system. Include system owners, operators, users, those affected by the AI’s decisions, and regulators.
    2. Analyse Their Perspectives: For each stakeholder, understand what success looks like from their perspective, what risks they perceive, and what their level of influence over the AI system is.
    3. Identify Conflicts and Alignments: Where do stakeholder interests align? Where do they conflict? For example, a business unit’s desire for automated decisioning speed may conflict with a legal team’s requirement for explainable outcomes.
    4. Establish a Collaborating Community: Based on this analysis, bring together a representative group with the mandate and authority to guide the AI initiative. Ensure the community includes technical, ethical, legal, and end-user perspectives.

    Outcomes: A clear map of who needs to be involved, what they care about, and where collaboration will require active facilitation. This prevents the common failure mode of AI projects that are technically sound but organisationally orphaned.


    Principle 2: Reach a Shared Understanding of the AI Problem

    Once you have a community and a rich picture, the next step is precision: articulating the specific problem the AI is meant to solve and understanding the system dynamics that produced it.

    Tool 3 – Context Diagrams: Define the Boundaries of Your AI System

    Purpose: To create a concentric circle diagram that shows the relative influence different entities have on the AI system, making clear who can direct it, who can shape it, who matters but operates at arm’s length, and what environmental forces exist beyond anyone’s control.

    This is not a data flow diagram. It is an influence diagram. The question it answers is not “what connects to the system?” but “who can actually change what the system does?”

    How to Apply to AI:

    1. Define the AI System Boundary: The system under analysis sits in the innermost circle.
    2. Identify External Entities: List every person, team, organisation, regulation, and environmental factor that is relevant to the system’s operation.
    3. Map the Influence: For each entity place within the specific influence circle the different entities:
      • Under Direct Control: Those who build, configure, and operate the system. Example: Fraud Detection Developers.
      • Able to Influence: Those who shape requirements, priorities, and constraints. Example: Risk Team, Executive Team, users.
      • Not able to Influence but important: Forces that matter significantly but cannot be directed. Example: Existing regulations, third-party data providers.
      • Enviromental factors: Background conditions that affect the system without any party controlling them. Example: Dark Web activity that drives the threat landscape.
    4. Validate the Boundaries: Review the diagram with your collaborating community. Misplaced entities – particularly overestimating how much influence an organisation has over regulators or external data providers – are a common source of flawed intervention design.

    Outcomes: A formally bounded system that gives the team a shared, unambiguous definition of what they are responsible for, and what they are not.


    Tool 4 – Behaviour Over Time Graphs and Problem Statements: Articulate the Problem You’re Actually Solving

    Purpose: To visualise how key variables in your AI system have changed over time, and use those patterns to craft a precise problem statement and goal.

    How to Apply to AI:

    1. Select Key Variables: Choose three to five variables that capture the most important aspects of your AI system’s performance. Examples: model accuracy, false positive rate, customer trust score, regulatory breach incidents.
    2. Plot the Historical Trend: For each variable, draw a graph showing how it has behaved over a relevant time period. Where is it going? Is it deteriorating, stable, or oscillating?
    3. Identify the Pattern: What is the main pattern you want to change? For example, a steadily increasing false positive rate, or a customer trust score that collapses each time the model is retrained.
    4. Craft the Problem Statement: Write a clear, specific problem statement that names the variable, describes its undesirable behaviour, and specifies the timeframe. Example: “Our AI credit model’s false positive rate has increased by 40% over the past 12 months, eroding customer trust and increasing manual review costs”.
    5. Define the Goal: Articulate the desired future behaviour of the same variables. This becomes the target your interventions are designed to achieve.

    Outcomes: A grounded, evidence-based problem statement that the whole team agrees on, a critical prerequisite before any solution is designed.


    Tool 5 – Identifying Enablers and Inhibitors: Understand What’s Driving Your AI Problem

    Purpose: To systematically explore the factors that support or obstruct the AI system’s ability to deliver its intended outcomes.

    How to Apply to AI:

    1. Use Your Problem Statement as the Anchor: Keep the problem you defined in Tool 4 at the centre of this analysis.
    2. Brainstorm Inhibitors: Ask, “What factors make this problem worse or prevent us from achieving our goal?” Organise them across four dimensions:
    • Technical: Data quality issues, model drift, integration failures.
    • Organisational: Siloed teams, unclear accountability for AI outcomes, insufficient resources.
    • Regulatory: Ambiguity in compliance obligations, gaps in internal policy.
    • Human: Distrust of AI outputs, skills gaps, change resistance.
    1. Brainstorm Enablers: Ask, “What factors support the system in delivering its goal? What resources can we leverage?” Examples include executive sponsorship, access to cloud infrastructure, clear guidance from the OAIC, and strong customer demand.
    2. Prioritise: Not all inhibitors carry equal weight. Which ones, if addressed, would have the greatest impact on the problem?

    Outcomes: A structured map of the forces at play, the essential input for building your causal loop diagram in the next tool.


    Tool 6 – Creating a Causal Loop Diagram: Map the Dynamics of Your AI System

    Purpose: To create a visual map of the feedback loops and causal relationships that drive AI system behaviour. This is the central analytical tool of systems thinking.

    How to Apply to AI:

    1. Start with Key Variables: Select the most important variables from your previous analyses. Examples: “Model Accuracy”, “Customer Trust”, “Data Quality”, “Number of False Positives”.
    2. Connect Variables with Arrows: Draw arrows between variables to show direction of influence. An increase in “Data Quality”, for example, leads to an increase in “Model Accuracy”.
    3. Label the Links: Mark each arrow with an ‘s’ (same direction – if one increases, so does the other) or an ‘o’ (opposite direction – if one increases, the other decreases). Example: “Number of False Positives” → ‘o’ → “Customer Trust”.
    4. Identify Feedback Loops: Trace the arrows to find closed loops. Label each as reinforcing (R) or balancing (B).
    • A reinforcing loop amplifies change in one direction. Example: Higher model accuracy drives higher user adoption, which generates more data, which improves data quality, which drives higher model accuracy again (R).
    • A balancing loop resists change and seeks stability. Example: An increase in false positives triggers more manual reviews, raising operational costs, which drives investment in model improvement, which reduces false positives (B).

    Outcomes: A dynamic visual map of your AI system that reveals the underlying structures driving its behaviour. This is the foundation for identifying where to intervene.


    Tool 7 – Causal Loop Diagram Analysis and Narrative: Turn the Map into Insight

    Purpose: To interpret the causal loop diagram, surface key insights, and develop a clear narrative that explains the AI system’s behaviour to stakeholders.

    How to Apply to AI:

    1. Analyse the Dominant Loops: Which loops are driving the system’s current behaviour? Are there vicious cycles at work, for example, declining trust leading to reduced usage, which degrades model performance, which further erodes trust?
    2. Look for Delays: Identify where significant delays exist in the system. Delays are a common source of oscillating behaviour and unexpected outcomes. There is often a long lag between deploying a new model and seeing any change in customer satisfaction scores.
    3. Identify Common Archetypes: Look for recognisable system patterns:
    • Fixes That Fail: A short-term intervention (manually overriding flagged transactions) that prevents the system from learning, causing the original problem to return.
    • Shifting the Burden: Addressing a symptom (a customer service team handling AI complaints) instead of the root cause (a biased model), which atrophies the organisation’s capacity to solve the fundamental problem.
    1. Develop a Narrative: Write a concise story that explains the AI system’s behaviour based on the diagram. Use it to communicate findings to stakeholders, build consensus, and justify proposed interventions.

    Outcomes: Deep insight into the systemic causes of the AI problem, and a practical communication tool to align stakeholders around a shared understanding.


    Phase 2: Co-design and Test Possible AI Interventions

    Principle 3: Identify Where to Intervene – Not Just What to Change

    With a comprehensive map of the AI system, the question becomes: where will an intervention actually make a difference? Not all interventions are equal. Some address symptoms. Others change the system’s fundamental structure.

    Tool 8 – Identifying AI Systems Leverage: Find Where Small Changes Produce Large Results

    Purpose: To identify high-leverage points, the places in the system where a well-targeted intervention will produce significant, lasting improvement. This tool applies Donella Meadows’ leverage points framework to AI deployment.

    How to Apply to AI:

    Consider these leverage points in ascending order of impact:

    • Constants and Parameters (least leverage): Adjusting numerical model parameters such as learning rate or decision thresholds. Useful for optimisation, but rarely changes fundamental system behaviour.
    • Buffers: The size of data buffers or processing infrastructure capacity. Increasing these improves stability but doesn’t address root causes.
    • Stock-and-Flow Structures: The physical architecture of the AI system: data pipelines, hardware, infrastructure. Changes here are costly and time-consuming.
    • Delays: The time it takes for information to travel through the system. Reducing delays in feedback loops (for example, faster model retraining cycles) can significantly improve responsiveness.
    • Balancing Feedback Loops: Strengthening the controls that keep the system stable. For example, tightening the feedback loop between model performance and data quality assurance improves reliability.
    • Reinforcing Feedback Loops: Slowing a vicious cycle (eroding trust) or accelerating a virtuous one (user adoption driving data quality improvement) is a more powerful intervention.
    • Information Flows: Who can see what. Improving transparency by giving users clear explanations of AI decisions – consistent with the Australian AI Ethics Principles – builds trust and creates new positive feedback loops.
    • Rules of the System: The governing rules, such as the data privacy policies under the Privacy Act 1988, ethical guidelines, performance thresholds. Changing the rules changes behaviour throughout the system.
    • The Goal of the System (high leverage): Shifting the goal from “maximising accuracy” to “making fair and transparent decisions” changes the design criteria for the entire system.
    • The Paradigm Out of Which the System Arises (highest leverage): The deeply held beliefs shaping the system. The shift from a purely technical view of AI to a human-centred, socio-technical perspective – one that treats human wellbeing as a design requirement, not an afterthought – is the most powerful intervention of all.

    Outcomes: A prioritised list of potential interventions, focused on those that create genuine, systemic change rather than temporarily suppressing the problem.


    Principle 4: Test the Ideas Before You Deploy Them

    Before committing to implementation, test your proposed interventions against a model of the system. Simulation and modelling exist precisely to let you discover failure modes cheaply, before they surface in production.

    Tool 9 – Stock and Flow Diagrams: Build a Model You Can Test

    Purpose: To create a quantitative model of the AI system that simulates the effects of different interventions over time. This builds on the qualitative insights from the causal loop diagram.

    How to Apply to AI:

    1. Identify Stocks: The key accumulations in the system, the things you could measure at a single point in time. Examples: “Number of active users”, “Volume of training data”, “Level of customer trust”.
    2. Identify Flows: The rates that cause stocks to increase or decrease. Examples: “New user adoption rate” (inflow to active users), “User churn rate” (outflow from active users).
    3. Connect Stocks and Flows: Draw stocks as boxes and flows as pipes with directional arrows. Use your causal loop diagram to define the logic of each flow.
    4. Simulate Interventions: Use simulation software or a spreadsheet to test your proposed interventions. Ask:
    • “What happens to customer trust if we reduce the false positive rate by 30%?”
    • “How does a three-month delay in model retraining affect overall accuracy?”

    Outcomes: A dynamic model of the AI system that lets you test hypotheses, compare intervention strategies, and anticipate long-term consequences before they become real.


    Tool 10 – Theory of Change Maps: Articulate the Causal Path from Intervention to Impact

    Purpose: To create a visual roadmap showing how an AI intervention produces the desired long-term impact, making the underlying assumptions explicit and testable.

    How to Apply to AI:

    1. Start with the Long-Term Goal: Define the ultimate impact the AI initiative is designed to achieve. Example: “Improved customer financial wellbeing”.
    2. Work Backwards: Identify the long-term outcomes that must be in place to achieve this goal. Example: “Customers make better financial decisions”.
    3. Identify Intermediate Outcomes: Continue working backwards through the shorter-term outcomes. Example: “Customers receive fair and transparent credit assessments”.
    4. Define the Outputs: What does the AI intervention directly produce? Example: “AI model generates accurate and explainable credit scores”.
    5. State the Intervention: Name the specific AI intervention clearly. Example: “Deploy a new, ethically-designed AI credit scoring model”.
    6. Articulate the Assumptions: For each link in the chain, state the assumption on which it depends. Example: “We assume that transparent credit assessments will increase customer trust in our services”. These assumptions are where your model is most vulnerable, so surface them, then test them.

    Outcomes: A logical map that explains how the AI initiative creates change, surfaces risky assumptions, and establishes the metrics worth monitoring.


    Phase 3: Implement Systemic AI Interventions, Monitor and Evaluate

    Principle 5: Monitor, Evaluate, and Learn with the Community

    An AI system is not static. The model drifts. The environment changes. Regulations evolve. User behaviour shifts. The organisation that treats deployment as the end of the process will watch its AI initiative quietly degrade, or spectacularly fail. The discipline of continuous monitoring, evaluation, and adaptation is what separates AI deployments that deliver sustained value from those that become expensive liabilities.

    Tool 11 – Monitoring and Evaluation Strategy: Build Governance That Keeps Pace with the System

    Purpose: To develop a comprehensive strategy for monitoring the performance, behaviour, and impact of the AI system across its lifecycle.

    How to Apply to AI:

    1. Define Key Performance Indicators: Based on your Theory of Change map and system diagrams, build a balanced set of metrics that go well beyond model accuracy:
    • Performance Metrics: Accuracy, precision, recall, false positive and negative rates.
    • Fairness and Bias Metrics: Performance variation across demographic groups.
    • Privacy Metrics: Data access requests, privacy incidents.
    • Business Metrics: Customer satisfaction, operational efficiency, return on investment.
    • Regulatory Compliance Metrics: Adherence to AFSL obligations and Privacy Act requirements.
    1. Establish Monitoring Processes: Implement automated dashboards and alerting systems. Define clear thresholds for when manual review or escalation is required.
    2. Conduct Regular Evaluations: Schedule periodic, in-depth assessments covering both quantitative KPI analysis and qualitative stakeholder feedback, from users, customers, and employees.
    3. Create Feedback Loops for Learning: Build clear processes for using monitoring insights to drive system improvements. This may mean model retraining, process adjustment, or revisiting the system’s fundamental goals.
    4. Engage the Collaborating Community: Share monitoring and evaluation findings with the stakeholder community established in Phase 1. Use their input to co-design improvements and ensure the AI system continues to meet their needs and reflect their values.

    Outcomes: A governance framework for the ongoing improvement of the AI system. One that builds a culture of learning, not just compliance, and ensures the AI initiative keeps delivering what it was built to deliver.


    References

    Department of Industry, Science and Resources. Voluntary AI Safety Standard.

    Department of Industry, Science and Resources. Australia’s AI Ethics Principles.

    Office of the Australian Information Commissioner. Guidance on Privacy and the Use of Commercially Available AI Products.

  • AI Safety Frameworks: Strategic Implications for the Financial Services Sector

    As artificial intelligence models reach unprecedented levels of capability, the risks they pose to global financial stability, market integrity, and consumer protection have become a focal point for regulators and industry leaders. This article analyses the twelve leading frontier AI safety frameworks – including those from Anthropic, OpenAI, Google DeepMind, and Microsoft – through the lens of financial services. By examining common elements such as capability thresholds, model weight security, and deployment mitigations alongside real-world case studies, we highlight the critical intersections between AI safety protocols and financial risk management.

    The Landscape of Frontier AI Safety

    Currently, twelve major AI developers have published formal safety policies designed to manage the “catastrophic risks” associated with high-capability models. These frameworks represent a shift from voluntary ethical guidelines to rigorous, technical protocols that mandate specific actions when certain risk thresholds are met.

    Table 1: Core Elements of Frontier AI Safety Frameworks

    Common ElementDescriptionFinancial Sector RelevanceRisk Reduction Examples
    Capability ThresholdsSpecific performance levels that trigger enhanced safeguardsTriggers for systemic risk monitoring and capital allocation adjustmentsAutomated circuit breakers prevent flash crashes like 2010’s $1 trillion loss
    Model Weight SecurityInformation security measures to prevent theft of AI “brains”Protection of proprietary trading algorithms and sensitive customer dataJPMorgan’s secured AI models enable $1.5B operational savings without IP theft
    Deployment MitigationsGuardrails to prevent misuse of models after releasePrevention of automated fraud, market manipulation, and phishingReduced pig butchering scam losses through AI detection systems
    Halting ConditionsProtocols to stop development or deployment if risks are unmanageableEmergency “kill switches” for AI-driven financial instabilityCircuit breakers that prevented 2016 GBP flash crash escalation
    AccountabilityInternal and external oversight mechanismsAlignment with existing regulatory compliance (e.g., AML, KYC, Basel III)Goldman Sachs’ AI governance enables 450+ safe use cases

    Critical Aspects for Financial Services

    1. Defining Economic Catastrophe: Learning from the Flash Crash

    A significant development in the regulatory landscape is the quantification of “catastrophic risk”. California’s Senate Bill 53 and several corporate frameworks define a catastrophic incident as one resulting in more than $1 billion in property damage or loss. This threshold reflects real financial system vulnerabilities.

    On 6 May 2010, the Dow Jones Industrial Average plummeted 998.5 points in approximately 36 minutes, erasing nearly $1 trillion in market value before recovering. The crash was triggered when Waddell & Reed Financial executed an algorithmic sell order of 75,000 E-Mini S&P contracts valued at approximately $4.1 billion, with the algorithm programmed to target execution based on trading volume without regard to price or time.

    How Safety Frameworks Reduce Similar Risks:

    Modern safety frameworks mandate halting conditions that would have prevented this cascade effect. Similar algorithmic hiccups occurred in 2016 when analysts attributed an overnight 6% drop in the British pound to algorithmic trading, confirming the susceptibility of algorithms to high-speed selling spirals.

    “Catastrophic risk means a foreseeable and material risk that a frontier developer’s deployment of a frontier model will materially contribute to more than one billion dollars in damage to, or loss of, property.”

    2. The Accelerating Threat of AI-Enabled Financial Crime

    The safety frameworks of companies like OpenAI and G42 explicitly track “Cyberoffense” and “Autonomous Replication” as high-risk categories. For the financial sector, these capabilities translate into criminal opportunities, and corresponding prevention successes when safety protocols are implemented.

    The Scale of the Problem:

    Cryptocurrency scams amounted to $9.9 billion in 2024, with that figure likely to be revised to a record $12.4 billion, driven largely by AI-enabled fraud. Pig butchering revenue grew nearly 40% year over year, with deposits to these scams growing nearly 210%, indicating an expansion of the victim pool through AI automation.

    Real-World Criminal Innovation:

    A prominent Nigerian cybercriminal recently posted a video showing a fully automated AI chatbot communicating directly with a victim who believed she was talking to her love interest – a military doctor overseas. The use of fully autonomous AI chatbots is set to explode, with numerous videos documenting walls of cell phones that work day and night to find people susceptible to pig butchering.

    How Safety Frameworks Enable Defence:

    AI service vendors’ revenue on illicit platforms had a compound annual growth rate of 1,900% between 2021-2024, indicating an explosion in AI technology facilitating scams. However, financial institutions implementing safety frameworks are achieving success:

    • Scaled Fraud Prevention: Advanced detection systems now identify AI-generated content patterns, reducing successful social engineering attacks
    • Ransomware Protection: Model weight security prevents AI systems from being compromised and weaponised against their operators
    • Behavioural Analysis: AI safety protocols enable better detection of “deceptive alignment” where models appear benign during testing but engage in harmful activities during deployment

    3. Market Manipulation and the Integrity Challenge

    Google DeepMind and the EU AI Act’s Code of Practice highlight “Harmful Manipulation” as a systemic risk. In finance, this manifests as AI’s ability to “systematically and substantially change beliefs and behaviour in high-stakes contexts”. Recent cases demonstrate both the risks and the protective value of safety frameworks.

    Successful Prevention Through AI Governance:

    JPMorgan’s Coach AI helped advisors respond to client concerns with unprecedented speed during market volatility, contributing to a 20% increase in gross sales (2023-2024) by identifying revenue opportunities and enhancing client satisfaction through tailored strategies. This demonstrates how safety frameworks enable beneficial AI deployment while preventing manipulation.

    Goldman Sachs deployed the GS AI Assistant to draft pitch decks across its investment banking division, with bankers reporting that the tool reduced deck preparation time by 50%, translating to thousands of reclaimed hours and faster client turnarounds. The key difference: robust governance frameworks ensure these AI tools enhance rather than manipulate decision-making.

    The Manipulation Risk:

    Without safety frameworks, AI systems can engage in sycophancy (telling users what they want to hear) or strategic deception, potentially leading to market bubbles or widespread consumer harm. In 2019, Apple and Goldman Sachs faced public scrutiny after reports surfaced that Apple Card’s AI-driven credit limit decisions were biased against women, with some customers finding that men were approved for significantly higher limits despite similar financial backgrounds.

    4. Model Weight Security

    The frameworks emphasise that model weights – the core parameters of an AI system – must be protected with state-of-the-art security. Recent data reveals both the scale of the threat and the business case for protection.

    The Financial Impact of Model Theft:

    Training a state-of-the-art language model can cost anywhere from hundreds of millions to over two billion dollars in compute resources, while DeepSeek allegedly developed its reasoning model using model distillation techniques for approximately six million dollars. According to IBM’s 2024 Cost of a Data Breach Report, intellectual property theft costs organizations $173 per record, with IP-focused breaches increasing 27% year-over-year.

    Real-World Vulnerability:

    IBM’s research found that 13% of organizations reported breaches of AI models or applications, with 97% of those organizations lacking proper AI access controls. Organizations that used high levels of shadow AI observed an average of $670,000 in higher breach costs than those with low levels of shadow AI.

    Success Through Security Frameworks:

    JPMorgan’s Contract Intelligence platform processes 12,000 commercial credit agreements in seconds, transforming both efficiency and risk assessment capabilities, while maintaining the model weight security that Anthropic’s ASL-3 standard and G42’s Security Mitigation Levels require. This demonstrates how security frameworks enable rather than hinder innovation.

    A financial services company (FinServe) developed a proprietary fraud detection model with 99.2% accuracy on their transaction patterns, but when a competitor hired a disgruntled former contractor who exfiltrated the model, FinServe had no evidence without proper fingerprinting. This illustrates why model weight security has become a fiduciary duty.

    5. Quantitative Benchmarking: Measuring AI Reliability

    Frameworks from xAI and Magic place heavy emphasis on quantitative benchmarks. xAI’s Risk Management Framework introduces the Model Alignment between Statements and Knowledge (MASK) benchmark to quantify a model’s honesty, which is critical for AI systems used in financial reporting and regulatory disclosures.

    Current AI Limitations in Finance:

    FinGAIA, an end-to-end benchmark designed to evaluate AI agents in financial scenarios, found that the best-performing agent, ChatGPT, achieved an overall accuracy of 48.9%, which while superior to non-professionals, still lags financial experts by over 35 percentage points. Error analysis revealed five recurring failure patterns: Cross-modal Alignment Deficiency, Financial Terminological Bias, and Operational Process Awareness Barrier.

    The Business Case for Benchmarking:

    JPMorgan rolled out over 200 AI use cases including automated KYC verification and trade surveillance, with McKinsey analysis showing these initiatives saved the bank over $1.5 billion in operational costs while enhancing compliance. The key difference: comprehensive benchmarking ensures AI systems perform reliably in high-stakes environments.

    Strategic Recommendations for Financial Institutions

    To navigate this evolving landscape, financial institutions should integrate AI safety frameworks into their existing risk management structures:

    1. Due Diligence on AI Partners

    When selecting an AI provider, firms must evaluate the robustness of the provider’s safety policy. JPMorgan’s approach demonstrates successful vendor management. Look specifically for clear halting conditions and capability elicitation practices that prevented the type of runaway effects seen in the 2010 flash crash.

    2. Systemic Risk Stress Testing

    Incorporate the $1 billion “catastrophic risk” threshold into financial stress tests to model AI-driven market disruptions. The International Monetary Fund’s October 2024 Global Financial Stability Report warns that AI tools are contributing to increased volatility in capital markets, with higher variability in AI-driven exchange-traded funds.

    3. Continuous Monitoring and Governance

    Chief Risk Officers now face a dual challenge: implementing AI systems that deliver significant operational benefits while meeting evolving regulatory expectations. Leverage “post-deployment monitoring” requirements mentioned in the EU Code of Practice to ensure AI systems used in high-stakes financial decisions are continuously audited for bias, performance drift, and security vulnerabilities.

    Proven Success Metrics:

    • AI coding assistants boosted developer efficiency by 10-20% at JPMorgan
    • UniCredit’s DealSync AI sourced more than 2,000 viable M&A leads in its first year
    • Goldman Sachs hired over 500 AI engineers in 2024 alone, bolstering expertise in machine learning and natural language processing

    Conclusion: Safety as Competitive Advantage

    The twelve frontier AI safety frameworks represent regulatory compliance and provide a roadmap for competitive advantage in financial services. Leading institutions have moved beyond pilot programs to enterprise-scale AI deployments generating substantial business value, with the window for gradual adoption closing as 2026 becomes the year AI moves from competitive advantage to competitive necessity.

    The evidence is clear: safety frameworks enable rather than constrain innovation. Organizations using AI and automation extensively throughout their security operations saved an average $1.9 million in breach costs and reduced the breach lifecycle by an average of 80 days. Meanwhile, cryptocurrency scams reached record levels of $12.4 billion in 2024, fueled by AI-powered deception, highlighting the cost of inadequate safety measures.

    For financial institutions, these protocols are essential components of modern financial stability frameworks. By aligning AI safety with traditional risk management – learning from the flash crash of 2010, the current pig butchering epidemic, and the success stories of JPMorgan, Goldman Sachs, and others – financial institutions can harness the power of frontier models while safeguarding the integrity of the global economy.

    Implement comprehensive AI safety frameworks and join the institutions generating billions in value, or risk becoming casualties of the next AI-driven financial catastrophe. The 2010 flash crash cost $1 trillion in 36 minutes. Today’s AI-enabled threats move even faster, but so do the defences for those prepared to implement them.

    References:

    1. Amazon’s Frontier Model Safety Framework
    2. Anthropic’s Responsible Scaling Policy, v2.2
    3. Cohere’s Secure AI Frontier Model Framework
    4. G42’s Frontier AI Safety Framework
    5. Google DeepMind’s Frontier Safety Framework, Version 3.0
    6. Magic’s AGI Readiness Policy
    7. Meta’s Frontier AI Framework
    8. Microsoft’s Frontier Governance Framework
    9. Naver’s AI Safety Framework
    10. NVIDIA’s Frontier AI Risk Assessment
    11. OpenAI’s Preparedness Framework, Version 2
    12. xAI’s Risk Management Framework