Bologna, Italy
(from 8 to 22)

AI Governance: Real Cases and Verifiable Responsibility

AI governance is not a theoretical problem.

It is already today a matter of concrete accountability, audits, regulatory reviews and legal disputes.

Real cases emerge every day – flawed decisions, unverifiable processes, supervision declared but not demonstrable – that reveal a structural problem:

organizations often cannot prove how and why a decision was made.

In this section we analyse real cases:

  • court rulings and legal decisions
  • operational incidents and corporate errors
  • risk scenarios related to AI use

Each case is read through a specific lens:

not what happened, but what would have needed to be proven.

Because today, in AI governance, the difference does not lie in declared processes.

It lies in the ability to produce verifiable evidence.

Methodological note: each case presented on this page is analysed in relation to the problem of demonstrating human supervision. The reference to the documentation framework is not a universal solution to all the problems described, but indicates what – structurally – should have been documented and verifiable.

Explore the protocol infrastructure:

→ The </AI> Protocol
→ From decision to defensible structure: when supervision becomes evidence
→ Read the Public Technical Specification
→ Verify a CWC Code in the Public Registry
→ EVIDE – Evidentiary registry for digital content and decisions
→ CWC Registry Policy
→ Request an Official CWC Verification Code
→ AI Governance Documentation Framework
→ Implementation Guide: verifiable AI supervision
→ Oversight Bias: why human supervision can fail in AI systems
→ Decision Attestation Layer: the missing evidentiary layer in AI governance
→ AI Evidence Officer: proving human supervision in artificial intelligence systems
→ Evidentiary Layer in AI Governance
→ </AI> Protocol FAQ: questions and answers about the framework

Related reading:

→ AI Data Poisoning: The Attack No Antivirus Can Stop
→ Human in the loop: why saying there is human oversight is not enough
→ Real cases: when AI governance fails — and what should have been provable
→ AI Governance: when something has already gone wrong – forensic reconstruction and digital evidence

Work with us:

→ Legal Partners Network

AI Governance: Real Cases and Verifiable Responsibility
AI Governance: Real Cases and Verifiable Responsibility

Real Governance Cases

Index :

Active litigation – USA
2023 – 2025
U.S. District Court, E.D. California – Cigna PxDx | U.S. District Court, D. Minnesota – UnitedHealth nH Predict

AI-denied claims — when declared oversight cannot be demonstrated

Cigna denied 300,000 claims in two months with 1.2 seconds of “review” per case. UnitedHealth deployed a model with a 90% error rate to deny care to elderly patients. In both cases the problem was not the AI. It was the complete absence of any evidentiary structure for human oversight.

Two separate class actions, initiated in 2023 and already procedurally advanced by 2025, have placed the same question at the center of legal debate: what actually constitutes human oversight in an AI-assisted decision process? In the Cigna case, the PxDx algorithm processed thousands of denials per day while the company’s internal physicians spent an average of 1.2 seconds per case, without reading clinical files. In the UnitedHealth case, the nH Predict model was known to have a 90% error rate, yet was used to deny rehabilitative care to elderly Medicare Advantage patients, routinely overriding the assessments of treating physicians.

The governance problem: in both cases, the use of AI to support decisions was not in dispute. What was in dispute was the complete absence of any evidentiary structure for human oversight: no documented criterion, no verifiable threshold, no clear attribution of the review. The court’s question was not “did you use AI?” It was: “how do you demonstrate that a human actually reviewed this decision, against which criterion, and within which documented structure?”

The facts

In July 2023, a group of patients sued Cigna Corporation in the U.S. District Court for the Eastern District of California, alleging that the company used the PxDx algorithm to automatically deny thousands of claims without individual medical review. A subsequent ProPublica investigation documented that Cigna’s physicians were rejecting cases without opening clinical files, averaging 1.2 seconds per case. On March 31, 2025, the court allowed the class action to proceed.

In November 2023, family members of two deceased patients filed a class action against UnitedHealth in the U.S. District Court for Minnesota, contesting the use of the nH Predict model to deny rehabilitative care to elderly patients. The allegations indicated that the system had a 90% error rate and was used to override the assessments of treating physicians. The court allowed the class action to proceed in February 2025 on breach of contract and breach of good faith claims.

The emerging legal principle

Both proceedings converge on a principle that courts are progressively consolidating: declaring the presence of human oversight is not equivalent to demonstrating its structure. It is not sufficient that a physician was formally assigned to the review. It must be demonstrable how that review was conducted, against which criterion, under which authority, and with which documented outcome. In the absence of this structure, “oversight” is merely a label applied to a process that was, in substance, automatic.

California SB 1120, which came into effect on January 1, 2025, has already crystallized this principle into law: any denial based on medical necessity must be reviewed by a licensed physician, with individual documentation per case. The algorithm may assist. It cannot decide alone. And oversight must be demonstrable, not declared.

The governance question

These cases involve health insurance in the United States. But the principle they are generating is universal and applies to any insurance company that uses AI in claims processing, underwriting or risk scoring:

If an AI-denied claim or an AI-increased premium reaches litigation, what can you actually demonstrate about the human oversight that governed that decision?

Applying the EVIDE structure to the Cigna and UnitedHealth cases:

Condition Status in the cases
Attributable authority (identified reviewer) ✘ Formal but not real — 1.2 seconds per case does not constitute review
Documented criterion (taxonomy_reference) ✘ Absent — no verifiable clinical taxonomy anchored at the moment of decision
Verifiable threshold (threshold_reference) ✘ Absent — the model operated without declared and externally verifiable thresholds
Threshold status (threshold_status) ✘ Not recorded — no record of met / not_met / not_defined per individual case
Real consequence on the patient ✔ Present — denied care, financial harm, in some cases death
External defensibility ✘ Not guaranteed — the companies could not produce an evidentiary structure

The system did not fail because the AI was wrong. It failed because it produced real consequences for patients based on decisions that did not meet any of the conditions necessary to be externalized as accountable.

The appeal reversal rate — above 90% in the UnitedHealth case — is proof that the problem was not in the clinical merits of the individual cases. It was in the absence of structure. The care was often appropriate. The ability to demonstrate it was missing.

What should have been demonstrable

For each denied or modified claim with AI support, a structure documenting: the identity and credentials of the reviewer who conducted the oversight, the clinical taxonomy active at the moment of decision (taxonomy_reference), the admissibility threshold applied (threshold_reference), the verification status against that threshold (threshold_status: met | not_met | not_defined), the structured rationale of any override or confirmation, and the timestamp of the human review — not of the algorithmic output.

Not a generic declaration of “medical oversight.” An evidentiary structure capable of answering the question: in this specific case, was human oversight real, structured and externally verifiable — or was it merely a label applied to an automatic process?

This is exactly the role of the External Evidentiary Deposit EVIDE and the EVIDE JSON 1.7 schema: transforming every AI-assisted insurance decision into a record with taxonomy_reference, threshold_reference and threshold_status anchored externally at the moment of human review. Not after. Not on the court’s request. At the moment the decision produces effects on the client.

Legal precedent
February 14, 2024
British Columbia Civil Resolution Tribunal – Moffatt v. Air Canada, 2024 BCCRT 149

Moffatt v. Air Canada – when an AI decision does not meet accountability conditions (anchor-ready)

A chatbot produced a decision with real-world consequences, without attributable authority, without structured reasoning, and without documented human oversight. This was not a technical failure. It was a decision maturity failure.

The Tribunal held Air Canada liable for incorrect information provided by its AI chatbot to a customer purchasing a ticket following a family bereavement. The chatbot stated that a bereavement discount could be claimed retroactively, contradicting the company’s official policy. Air Canada argued that the chatbot was a separate entity and not attributable to the company. The Tribunal rejected this argument entirely.

Governance issue: Air Canada allowed an AI system to issue decisions with real economic consequences without ensuring those decisions were anchored to attributable authority, verifiable reasoning, or documented supervision. The problem was not the chatbot error. The problem was that the decision was not anchor-ready.

The facts

In November 2022, Jake Moffatt visited Air Canada’s website to purchase a ticket after the death of his grandmother. He interacted with the company’s chatbot, which informed him that he could apply for a bereavement discount retroactively within 90 days of travel. Moffatt purchased the ticket at full price and later requested a refund. Air Canada denied the request. In February 2024, the Tribunal ruled in favor of Moffatt.

Legal principle

The Tribunal established that Air Canada is responsible for all information presented on its website, regardless of whether it originates from static pages or a chatbot. The argument that the chatbot was a separate entity was rejected. The applied standard: organizations must take reasonable measures to ensure that information provided by AI systems is accurate and not misleading.

Governance question

Does your AI system produce decisions that meet the minimum conditions to be externally accountable?

Applying the Minimum Anchoring Contract:

Condition Status
Attributable authority ✘ Absent
Structured rationale ✘ Absent
Documented oversight ✘ Absent
Real consequence ✔ Present
External accountability ✘ Not guaranteed

The system did not fail because the chatbot made an error. It failed because a non anchor-ready decision was allowed to produce real-world consequences.

What should have been demonstrable

A system ensuring that each output with contractual impact is generated within defined authority, aligned with active policy, subject to human oversight, and fully traceable.

This is exactly the role of the Human Oversight Event and EVIDE.

Legal Precedent
2026
New Zealand Court of Appeal – Yorston v Attorney General [2026] NZCA 15

Automation is not immunity – accountability for AI systems in public decisions

Labelling a process as “automated” does not place it beyond judicial scrutiny. Accountability follows system design.

The Court ruled that decisions produced by government automated systems are subject to judicial review. The public agency that adopts and implements an automated system remains fully responsible for its outputs, even when no human being directly intervened in the individual decision.

The governance problem: if accountability follows system design, every organization adopting AI must be able to demonstrate how that system was configured, validated and supervised. The </AI> Protocol builds exactly that structure: transforming supervision from a declaration into verifiable and attributable evidence.

The facts

The case concerned errors in an automated convictions history report generated by the Ministry of Justice’s case management system. The applicant argued that the errors were not subject to judicial review because they were produced by an automated system. The Court of Appeal rejected this argument.

The legal principle

The Court clarified that the automated system operates within the exercise of executive powers of government and that its outputs are therefore subject to judicial review. The central point: automation does not equal immunity. Accountability for errors arising from software architecture, data handling, or system logic remains with the public agency that adopted and deployed the system.

The Court also indicated that future challenges are likely to focus less on individual outputs and more on how automated systems are designed, validated, monitored and corrected.

The governance question

This case is not only about government agencies. It raises a structural question for any organization using AI systems in decision-making processes that produce effects on third parties:

If your AI system produces an output that affects someone’s rights or interests, can you demonstrate how that system was designed, validated and supervised?

The Court identified three governance areas that every organization should address:

  • mapping automated systems to statutory functions and institutional responsibilities
  • mechanisms for detecting errors, correcting outputs and explaining changes
  • maintaining auditability: system logic, data sources, updates and known limitations

Without this structure, the organization cannot respond to a challenge: not because it acted incorrectly, but because it cannot prove it acted correctly.

This case illustrates a structural point: when governance is not supported by verifiable evidence, accountability becomes attributed by presumption, not demonstrated.

What should have been provableA structured record documenting: how the system was designed and with what criteria, what data feeds the decisions, who validated the system logic, what correction mechanisms are active, and who is the designated responsible for operational supervision.

Not a generic technical description, but an evidentiary structure capable of answering the question: how was human supervision exercised over this system, and when?

Without this structure, the decision remains formally valid, but becomes substantially indefensible.

This is precisely the role of the AI Evidence Officer and the AI Governance Documentation Framework: transforming AI governance from a declared intention into a defensible evidentiary structure, independently verifiable and anchored to a public independent registry.

EU Legal Precedent
27 February 2025
Court of Justice of the European Union – C-203/22 (Dun & Bradstreet Austria)

Automated decisions and the right to explanation – when a decision cannot be defended

A decision that cannot be explained in an intelligible way cannot be defended.

The Court ruled that, in automated decisions based on profiling, the data subject has the right to receive meaningful, comprehensible and accessible information about the logic applied.

The governance problem: if an automated decision cannot be explained in a verifiable and comprehensible way, it cannot be defended. The </AI> Protocol adds the missing layer: transforming decision logic into structured, traceable and attributable evidence.

AI hallucination is not, in itself, a fault. But presenting outputs as reliable without a verifiable trail of human supervision creates a structural accountability gap.

The facts

An automated system was used to assess the creditworthiness of an individual. The resulting decision had significant effects on the person, without any clear understanding of the logic applied by the system being made available.

The legal principle

The Court clarified that the right to explanation under the GDPR requires that the data subject be able to understand the procedures and criteria actually applied in the automated decision. It is not necessary to disclose complex algorithms or trade secrets, but it is mandatory to provide a clear, intelligible and verifiable explanation of the decision logic.

The governance question

This case is not only about credit scoring or the GDPR. It raises a structural question applicable to any decision made or supported by intelligent systems:

If you cannot explain a decision, how can you prove it was made correctly?

In real operational contexts, an explanation alone is not enough. It must be possible to demonstrate:

  • which data was used
  • which rules were active
  • who validated the process
  • in which context the decision was generated

Without this structure, the decision remains opaque – even when a formal explanation is provided.

What should have been provableA system that records not only the output, but the conditions under which the decision was made possible: the input data, the constraints applied, the active policies, the operational context and any human supervision.

Not an explanation generated after the fact, but an evidentiary structure generated before and during the decision-making process.

This is precisely the role of the Human Oversight Event and the Decision Attestation Layer: transforming a decision from an opaque result into a demonstrable process.

Legal Precedent
February 12, 2024
Italian Supreme Court (Court of Cassation) – Order no. 3263/2024

Employee Dismissed for Failing to Verify a Fraudulent Email – The Due Diligence Standard

Undeclared diligence provides no legal protection.

The Court confirmed the dismissal of an employee who authorized a payment following a phishing email. The ruling establishes that ordinary diligence requires active verification, even without specific IT training.

The Governance Gap: If you cannot prove how, when, and by whom an AI output was supervised, your organization carries the same liability exposure. The </AI> Protocol transforms “claimed oversight” into verifiable evidence.

The facts

An employee authorised a payment following a fraudulent email (BEC – Business Email Compromise), without performing the necessary checks before approving the transfer. The employer proceeded with disciplinary dismissal. Italian Supreme Court (Court of Cassation), with Order no. 3263 of 12 February 2026, confirmed the legitimacy of the dismissal.

The legal principle

The Court clarified that, even without specific training, employees are obliged to act with ordinary diligence, performing the necessary checks and verifications before authorising a payment. The key passage in the ruling: with a minimum of diligence, the fraud would have been avoidable. Responsibility cannot be transferred to the IT system or the fraudulent email alone – it falls on the person who made the decision to demonstrate that they acted with the required care.

The governance question

This case is not only about banking phishing. It raises a structural question that applies to any operational decision supported by automated or intelligent systems:

How do you prove that a decision was made with due diligence?

In a legal dispute, what a person claims to have done is not what matters. What matters is what is documented, traceable, and verifiable. In this specific case, the employee could not demonstrate that she had performed the necessary checks – not because she had not done them, but because no record existed of how the decision had been reached.

The same principle applies directly to AI use in operational processes:

  • a contract approved on the basis of AI output
  • an automated response to a customer
  • a decision made on analyses generated by intelligent systems
  • an internal process based on algorithmic recommendations

In each of these cases, the Court’s remains the same: who decided, what did they evaluate, when, and on what basis?

What should have been demonstrable
A structured record of the decision-making process documenting: the identity of who performed the review, the sources consulted at the moment of the decision, the checks carried out, the active operational context, the final decision and its rationale. Not a subsequent declaration – a trace generated at the moment the decision was made. This is precisely what the Human Oversight Event produces: not proof that a person was present, but proof of how that person exercised their supervision, under verifiable and documented conditions.

More cases coming – this section is updated periodically with new rulings, incidents and risk scenarios.

The framework is publicly defined as “The </AI> Protocol” and is forensically certified through CertifyWebContent.
This documentation constitutes a verifiable, timestamped record of its structure, concepts, and implementation.