Why Most AI Risk Assessments Miss the Mark

You’re testing your AI systems less rigorously than you’d test a toaster. That’s not hyperbole. It’s what IBM discovered when Watson for Oncology recommended chemotherapy that could kill a patient with severe bleeding – a contraindication flagged with a “black box” warning. The system had been trained on synthetic cases, not real patients. Despite knowing this, IBM promoted Watson as making recommendations based on actual patient data.

This happened in 2018. Your organisation is likely making similar assumptions today.

The Real Cost of Getting AI Risk Wrong

Here’s what we know from analysing 499 publicly documented AI incidents: most failures don’t stem from sophisticated technical glitches. They emerge from the predictable intersection of flawed data, human assumptions, and inadequate testing. And they’re expensive.

Air Canada learnt this in February 2024 when a tribunal ordered them to pay a customer who’d relied on incorrect information from their chatbot. The airline’s defence? The chatbot was “a separate legal entity responsible for its own actions”. The tribunal called this argument “remarkable” – and not in a good way. The airline paid CA$812 and much more in reputation.

Amazon scrapped an entire recruiting system after discovering it systematically discriminated against women. The algorithm had learnt from ten years of predominantly male hiring data and decided gender bias was the pattern to follow. They’d invested years in development before recognising the problem.

These aren’t edge cases. They’re showing you exactly where your own systems will fail.

Why Your Current Risk Assessment Misses These Failures

You’re Working from Theory, Not Evidence

Most risk frameworks ask: “What could go wrong?” The better question: “What has already gone wrong, and why didn’t we see it coming?”

Examining documented incidents, a pattern emerged. The harms extended far beyond direct users. Air Canada’s chatbot affected a grieving customer who’d never touched their development process. Amazon’s hiring algorithm impacted thousands of female applicants who never knew they’d been systematically downgraded.

Your risk assessment probably focuses on your direct users. You’re missing everyone else in the blast radius.

You’re Treating Human Factors as Edge Cases

Here’s what IBM’s Watson failure really reveals: the problem wasn’t the AI. It was the decision to train it on synthetic data while publicly claiming it learnt from real patients. It was the choice to keep promoting a product internal documents described as providing “unsafe and incorrect treatment recommendations”.

Those are human decisions, made by people with incentives, blind spots, and pressures. Your technical risk assessment doesn’t capture this. It can’t. Technical frameworks assess code, not context. They evaluate algorithms, not the organisational culture that decided to deploy them anyway.

NIST’s AI Risk Management Framework explicitly acknowledges this gap, noting that mathematical models can strip away the context necessary to understand individual and societal impacts. If NIST is telling you their framework needs augmentation for human factors, your assessment probably does too.

You’re Not Testing for Real-World Misuse

The fatal collision in Tempe, Arizona, in March 2018, revealed something critical: Uber’s autonomous vehicle wasn’t designed to recognise pedestrians outside crosswalks. The system detected Elaine Herzberg 5.6 seconds before impact but couldn’t classify what it was seeing. It toggled between “vehicle”, “bicycle”, and “other”.

When it finally determined she was a bicycle in its path, 1.3 seconds before impact, the automatic emergency braking had been disabled. Uber relied entirely on the human safety driver to intervene. That driver was streaming a television show on her phone.

This wasn’t a failure of one component. It was a cascade of design decisions, each seemingly reasonable in isolation, that combined into a fatal outcome. Your testing probably evaluates components separately. Real-world failures happen when they interact.

What Actually Works: Testing That Matches Reality

Start with What’s Already Failed

Stop asking your team to imagine risks. Give them documented events and ask: “Which of these patterns exist in our systems?”

Build your testing regime around empirical failure modes, not theoretical ones. When IBM’s Watson failed, the core issue was data quality – synthetic cases masquerading as real experience. When Amazon’s recruiting tool failed, it was because historical bias in training data became future discrimination.

These patterns repeat. Your systems contain the same vulnerabilities unless you’ve explicitly tested for them.

Make Human Factors Central, Not Optional

IBM’s Watson had a technical problem (synthetic training data) that became dangerous because of human decisions (keep selling it anyway). Amazon’s recruiting tool had a data problem (biased training set) that persisted because of human oversight failures (no one checked for gender disparities).

Your risk assessment should answer:

  • Who makes deployment decisions, and what are they incentivised to prioritise?
  • How do we validate claims about system capabilities?
  • What happens when someone flags a concern?
  • Who’s accountable when harm occurs?

If those questions feel uncomfortable, you’ve found your actual risk.

Test Like Failure Matters

Elaine Herzberg’s death was avoidable. The NTSB investigation revealed that between September 2016 and March 2018, Uber’s test vehicles had been involved in 37 other crashes. Yet systematic safety reviews weren’t in place. The company “did not have a standalone operational safety division or safety manager.”

You need:

Adversarial Testing: Hire people to deliberately misuse your system. The failure modes you haven’t imagined are the ones that will hurt you. Red team your AI the way you’d red team your cybersecurity.

Continuous Validation: AI systems drift. Models trained on 2020 data make poor decisions in 2025. Build monitoring that tracks performance degradation in production, not just in testing.

Post-Incident Analysis: When something goes wrong, treat it like an aircraft accident. Forensic analysis. Root cause identification. Systematic changes to prevent recurrence. One mistake is the message to change; two identical mistakes mean you haven’t learnt and need to do change the system.

Build Governance That Actually Governs

Air Canada’s chatbot problem persisted because no one had clear ownership of accuracy. Amazon’s recruiting tool kept running despite early warning signs. Uber’s safety protocols failed because safety wasn’t truly prioritised.

Effective governance isn’t about more documentation. It’s about:

Clear Accountability: Who can stop deployment if risks are unacceptable? That person needs authority, not just responsibility.

Diverse Expertise: Your AI risk team needs AI specialists, domain experts, ethicists, legal counsel, and people who represent your users. Homogeneous teams produce blind spots.

Transparent Operation: Your users should understand what your AI does, how it makes decisions, and what its limitations are. The Air Canada tribunal explicitly rejected the idea that users should “double-check information found in one part of its website on another part.”

Human Override: For high-risk systems, humans must be able to intervene meaningfully. Not just in theory – in practice. Uber’s safety driver couldn’t intervene effectively because she’d been trained to trust the system. Your override mechanisms need to be tested under realistic conditions, including human fatigue, distraction, and over-reliance on automation.

Your Monday Morning Actions

Here’s 5 quick win targets:

  1. Audit one system against documented failure modes (30 minutes): Take your highest-risk AI deployment. Compare it against the three failure patterns documented here – data quality, human factor neglect, and inadequate testing. Write down which vulnerabilities you share.
  2. Map your accountability (20 minutes): For that same system, write down who can actually stop deployment if risks surface. If the answer is “it depends” or “that’s complicated”, you’ve found a governance gap.
  3. Schedule your first red team exercise: Book two hours in the next fortnight. Invite someone who doesn’t work on the AI system to try to make it fail or produce harmful outputs. Document what they find.
  4. Review your most recent deployment decision (30 minutes): What evidence convinced you the system was ready? If the answer relies primarily on internal testing by the people who built it, your validation is insufficient.
  5. Establish your single point of contact for AI concerns: Create one channel where anyone in your organisation can flag AI risks without career consequences. Test it by having someone report a hypothetical concern.

The Path Forward

Building trustworthy AI is more than perfect technology. It’s about honest assessment of imperfect systems deployed in complex reality. IBM’s Watson, Amazon’s recruiting tool, Air Canada’s chatbot, and Uber’s autonomous vehicle all failed differently. But they share a common thread: organisations that didn’t adequately test, didn’t properly govern, and didn’t transparently acknowledge limitations until harm occurred.

You have a choice. You can wait until your incident becomes another case study, or you can learn from the ones that came before.

The testing tools exist. The frameworks are available. The documented failures provide your roadmap. What’s missing isn’t knowledge – it’s the organisational will to act on what we already know.

Your AI systems will fail. The question isn’t if, but whether you’ll see it coming and whether you’ll be ready when it does.