Bologna, Italy
(from 8 to 22)

AI Poisoning: The New Attack No Antivirus Can Stop

Modern AI systems don’t only learn from what we tell them directly. They continuously learn from the internet, public databases, and external digital archives. That capability is one of their greatest strengths – and their most dangerous blind spot. Some actors have already figured out how to exploit it.


Explore the protocol infrastructure:

The </AI> Protocol
From decision to defensible structure: when supervision becomes evidence
Read the Public Technical Specification
Verify a CWC Code in the Public Registry
EVIDE – Evidentiary registry for digital content and decisions
CWC Registry Policy
Request an Official CWC Verification Code
AI Governance Documentation Framework
Implementation Guide: verifiable AI supervision
Oversight Bias: why human supervision can fail in AI systems
Decision Attestation Layer: the missing evidentiary layer in AI governance
AI Evidence Officer: proving human supervision in artificial intelligence systems
Evidentiary Layer in AI Governance
</AI> Protocol FAQ: questions and answers about the framework

Related reading:

AI Data Poisoning: The Attack No Antivirus Can Stop
Human in the loop: why saying there is human oversight is not enough
Real cases: when AI governance fails — and what should have been provable
AI Governance: when something has already gone wrong – forensic reconstruction and digital evidence

Work with us:

Legal Partners Network


AI Poisoning: The New Attack No Antivirus Can Stop
AI Poisoning: The New Attack No Antivirus Can Stop

A new class of attack that targets information, not software

In classical cybersecurity, attacks target systems: a virus installs itself, an exploit leverages a vulnerability, ransomware encrypts files. Defence means protecting the system.

AI data poisoning – also called AI recommendation poisoning when it targets recommendation engines – works in an entirely different way. It does not attack the software. It attacks the information the software uses to learn.

The mechanism is this: an attacker inserts manipulated content among the sources that an AI system regularly analyses. If these contents are sufficiently numerous, internally consistent and distributed across multiple channels, the AI system can begin treating them as reliable information and incorporating them into its decision-making processes.

The result is not a detectable technical error. It is altered behaviour that appears entirely normal.

Why modern AI systems are structurally exposed to this risk

To understand the scale of the problem, it is necessary to understand how many modern AI systems actually operate.

Large language models rarely work in isolation. They are increasingly integrated into complex architectures that include:

  • RAG systems (Retrieval-Augmented Generation): these retrieve information in real time from external sources – websites, databases, public archives – to enrich their responses.
  • Autonomous AI agents: these execute sequences of actions, search for information online, interact with external services and make decisions without direct human intervention.
  • Recommendation engines: these analyse vast quantities of content to suggest products, investments, professionals, or courses of action.
  • Enterprise digital assistants: these combine internal data with external sources to support operational decisions.
  • Automated content analysis systems: these monitor, classify and synthesise information gathered from the web.

In all these cases, external data is not an accessory: it is the fuel the system runs on. And anyone who manages to contaminate that fuel has the ability to influence the output.

The invisibility factor
Unlike malware, a data poisoning attack leaves no traces in the software. It is not detected by antivirus tools. It generates no security alerts. It manifests as a gradual shift in the system’s responses and recommendations – difficult to attribute to a specific cause precisely because it resembles normal learning updates.

How an AI data poisoning attack works in practice

Consider an AI system that analyses online content to recommend investment instruments, assess the reputation of professionals, or identify reliable suppliers in a given sector.

An attacker seeking to manipulate its recommendations does not need to breach the system. They need to act on the sources the system consults. A typical strategy involves:

  1. creating a large volume of apparently authoritative content – articles, posts, reviews, profiles – that promotes whichever subject the attacker wants to favour or damage
  2. distributing this content across multiple channels to increase reach and simulate broad consensus
  3. optimising for the criteria AI systems use to assess source reliability: citation frequency, internal consistency, distribution across multiple domains
  4. waiting: as the AI system regularly analyses those sources, it begins incorporating the manipulated information as valid data

This is not a theoretical scenario. Variants of this approach are already documented in contexts ranging from search engine manipulation to geopolitical influence operations.

The specific problem of AI agents with persistent memory

Agent-based AI systems are introducing a characteristic that significantly amplifies this risk: persistent memory.

An AI agent with memory does not process information in isolation. It retains it, correlates it with prior experiences and uses it to improve future responses. This mechanism is designed to make the agent more capable over time.

But if that memory is contaminated by manipulated data, the effect is the opposite: the error is not confined to a single response. It propagates. It is reinforced each time the system retrieves and reuses that information. It can influence decisions made long after the original manipulation occurred – making it nearly impossible to identify when and how the compromise began.

The more autonomous the agent and the less it is supervised, the harder this risk becomes to contain.

The core issue: the absence of documented governance

The AI data poisoning problem is not only technical. It is a governance problem.

In most organisations using AI systems, there is no documentation of:

  • which sources the system uses to learn or to enrich its responses
  • how frequently those sources are updated or reviewed
  • who holds authority and responsibility for evaluating the quality of incoming information
  • how the system’s decisions are logged and which data they are based on
  • what happens if an input turns out to have been manipulated or incorrect

Without this documentation, it is not possible to systematically detect a poisoning attack – nor to demonstrate in a legal or regulatory context that the system operated correctly.

With the EU AI Act now in force and regulatory scrutiny of high-risk AI systems intensifying, this absence of documentation is not merely an operational gap. It is a compliance liability.

EU AI Act – Article 9 and 12
The AI Act explicitly requires providers and deployers of high-risk AI systems to implement risk management systems and maintain technical documentation that demonstrates how the system operates, what data it uses, and how human oversight is exercised. The ability to demonstrate compliance – not merely declare it – is a legal obligation.

The missing control point: the Human Oversight Event

The real weakness in most AI poisoning scenarios is not only that external information can be manipulated. It is that, in most organizations, there is no verifiable control point showing when a human being actually reviewed, validated or rejected the information the AI system relied on.

This is where the concept of the Human Oversight Event becomes critical.

A Human Oversight Event is the documented moment in which a designated human supervisor evaluates the information, source set or AI-generated output used by the system and produces a traceable evidentiary record of that intervention.

Without this event, organizations may claim that human supervision exists, but they often cannot demonstrate:

  • who actually reviewed the information or output
  • when that review took place
  • which version of the data, source set or output was examined
  • whether the reviewed material remained unchanged afterwards
  • under which governance policy or decision procedure the validation occurred

This matters enormously in poisoning scenarios. If manipulated information enters an AI workflow and there is no verifiable oversight event tied to that stage, the organization cannot later prove whether the data was ever checked, by whom, or on what basis it was accepted as reliable.

In practical terms, a system is not meaningfully protected against data poisoning unless it can demonstrate verifiable human oversight events on the information it relies on.

This is the point at which AI governance stops being a policy statement and becomes an operational evidentiary layer.

The right response: from reactive defence to documented governance

The response to data poisoning cannot be purely defensive. There is no antivirus for poisoned data. The structurally correct response is to build a system in which the quality and origin of information are verifiable, documented and traceable.

This means shifting focus from “how do I detect manipulation” to “how do I demonstrate that my sources are intact.”

It is the same logic that in digital forensics distinguishes between post-hoc analysis and preventive certification: act before the problem occurs, building a verifiable chain of evidence.

The </AI> framework and documented human oversight

The </AI> framework is a model that introduces a public and verifiable declaration of human oversight over content produced or used by artificial intelligence systems.

The principle is straightforward: every piece of content can be accompanied by technical information documenting the human supervision of the process, the origin of the information used, the verification of sources, and the integrity of the published content.

This approach makes it possible to distinguish between content generated without oversight – and therefore potentially poisoned or unreliable – and content that follows a transparent, traceable and independently verifiable process.

Documented human oversight is not a brake on automation. It is the condition that makes automation trustworthy and legally defensible.

The AI Governance Documentation Framework

To address these needs in a concrete and operational way, we developed the AI Governance Documentation Framework – a structured system that enables organisations to document, certify and demonstrate how their AI systems use information.

The framework addresses the problem at multiple levels:

  • Source documentation: verifiable registration of the sources used by the AI system, including date of access and methodology for assessing reliability.
  • Decision traceability: documentation of the system’s decision-making processes, with reference to the data on which they were based.
  • Certified human oversight: technical evidence of human intervention at critical points in the process, aligned with EU AI Act requirements.
  • Content integrity: cryptographic hashes and timestamping that allow verification at any future point that data has not been altered after acquisition.
  • Incident response: documented procedures for identifying, containing and tracing episodes of data poisoning or anomalous system behaviour.

Full details of the framework are available here:
AI Governance Documentation Framework

The AI Evidence Officer

Alongside the framework, we introduced the concept of the AI Evidence Officer: a professional figure specialised in managing technical evidence related to AI systems.

The AI Evidence Officer is not simply a security consultant. This is the person responsible for building and maintaining the evidentiary chain that makes it possible to demonstrate – in any context, including legal proceedings – that an AI system operated correctly, with verified sources and documented oversight.

In the event of a regulatory audit, a legal challenge, or an incident involving an automated decision, the AI Evidence Officer produces the technical documentation required to respond – not with declarations, but with verifiable evidence.

As the AI Act places the burden of proof on providers and deployers, having a designated AI Evidence Officer is no longer a best practice. It is a strategic necessity.

Learn more about the AI Evidence Officer role:
AI Evidence Officer

The role of digital content certification

A complementary element in the response to data poisoning is the preventive certification of the content that feeds AI systems.

Through cryptographic hashing and certified timestamping, it is possible to demonstrate:

  • the certified publication date of a piece of content
  • the integrity of the information over time
  • verifiable authorship of the material
  • that the content has not been altered after certification

When an AI system draws on sources certified in this way, it becomes possible to reconstruct with precision which information was available at a given moment and verify whether it was subsequently altered. This transforms data poisoning from an invisible attack into a traceable and documentable event.

The digital content certification system is available at:
CertifyWebContent

For preventive certification of the digital identity of content producers:
DAPI Certification

Why this matters for organisations right now

This is not a future scenario. The conditions for AI data poisoning to become a systematic attack vector are already in place:

  • RAG systems and autonomous AI agents are being deployed in business contexts at an accelerating pace
  • tools for generating convincing fake content at scale are accessible and inexpensive
  • the vast majority of organisations have no documented governance over their AI systems
  • the EU AI Act is now creating explicit legal obligations around demonstrable compliance

Organisations that are unprepared for an audit or a legal challenge will not be able to simply assert that their AI system operated correctly. They will need to prove it.

Use case: when AI poisoning becomes a legal problem

Scenario

The company: A European FinTech operator using RAG-based AI to deliver automated investment recommendations to retail clients.
The attack: poisoned documents are injected into internal datasets, classifying high-risk investments as safe.

The consequence: a client loses €50,000 and files a claim under the EU AI Act — triggering an investigation that could result in administrative fines up to €15 million or 3% of global annual turnover.

The crisis

The company is accused of producing misleading output and failing to apply adequate human oversight.
Without evidence, liability falls entirely on the organisation.

The intervention of the AI Evidence Officer

FinTech Global had implemented the </AI> protocol and documented governance.
This changes everything.

Phase Action Framework Tool Forensic Outcome
Identity Identification of the assigned human supervisor DAPI No anonymous or unauthorised access
Integrity Extraction of original approved document hash ContentProtector Proof of post-validation alteration
Traceability Certified log of human review AI Output Review Demonstration of real oversight
Defence Submission of certified evidence package CertifyWebContent Liability excluded

Final outcome

The company demonstrates that:

  • human oversight actually took place
  • the data was valid at the time of validation
  • the manipulation occurred afterwards

Liability is excluded.

Why no antivirus would have detected this

The poisoned documents appeared legitimate.
No traditional security system would have flagged them.

The key point

AI poisoning does not compromise software.
It compromises the truth the system relies on.

If you cannot prove what was verified, you cannot defend what your AI decided.

Frequently asked questions

What exactly is AI data poisoning?

It is an attack technique in which a malicious actor manipulates the data that an AI system uses to learn or to enrich its responses, with the goal of altering its behaviour or recommendations without modifying the software directly.

How does it differ from traditional hacking?

Traditional hacking targets the system – software vulnerabilities, unauthorised access, technical exploits. Data poisoning targets the information the system considers reliable. It leaves no traces in the software and is not detected by conventional security tools.

Are AI agents more vulnerable than classical models?

Significantly so. An AI agent that operates autonomously, consults external sources and maintains persistent memory has a far greater exposure surface than a model that only responds to direct inputs. Every external source it consults is a potential attack vector.

How does an organisation know if its AI system has been compromised?

In most cases, without documented governance, it does not. The system’s altered behaviour appears as a normal learning update. Only systematic monitoring of sources, combined with documentation of decisions and the data they are based on, makes it possible to detect anomalies attributable to manipulation.

What is the AI Governance Documentation Framework and who is it for?

It is a structured system for documenting how an organisation uses AI systems: which sources it consults, how it supervises the process, how it certifies the integrity of information. It is relevant to any organisation using AI in operationally significant contexts – particularly those subject to EU AI Act requirements or operating in regulated sectors.

What is an AI Evidence Officer?

The AI Evidence Officer is the professional responsible for building and maintaining the technical documentation that demonstrates how an AI system has operated. In the event of an audit, a legal challenge or an AI-related incident, the AI Evidence Officer produces verifiable evidence – not assertions.

Can content certification genuinely reduce the risk of data poisoning?

It does not eliminate it, but it makes it traceable and demonstrable. When the sources feeding an AI system are certified with cryptographic hashes and timestamps, it becomes possible to verify at any future point whether they were altered after certification. This transforms a silent attack into a detectable and documentable event.

What is a Human Oversight Event and why does it matter for AI compliance?

A Human Oversight Event is the documented moment in which a designated human supervisor
evaluates and validates the information, source set or output used by an AI system,
producing a traceable evidentiary record of that intervention. Without verifiable
Human Oversight Events, an organisation cannot demonstrate — to regulators, courts or
auditors — that its AI systems operated under meaningful human control. Under the EU AI
Act, this is not a procedural detail: it is a core element of compliance demonstrability.

Does this concern SMEs as well as large organisations?

It concerns any organisation using AI systems to support operational decisions, client recommendations or analysis of external information. Size is not the determining factor: what matters is the autonomy of the AI system and the criticality of the decisions it supports.

A final note worth keeping in mind

Artificial intelligence does not make decisions in a vacuum. Its outputs always depend on the quality of the information it uses. If that information is manipulated, the decisions are manipulated too – often without anyone noticing.

Documented governance, certified human oversight and source verifiability are not optional features. They are the conditions that make an AI system genuinely trustworthy and legally defensible.

The framework is publicly defined as “The </AI> Protocol” and is forensically certified through CertifyWebContent.
This documentation constitutes a verifiable, timestamped record of its structure, concepts, and implementation.