AIWiki
Malaysia
Back to all articles
Ethics & Policyadversarial MLdata securitymodel safety

Data Poisoning

5 min readUpdated October 2026
Data Poisoning
Type
Adversarial attack on machine learning
Target
Training data (pre-training, fine-tuning, RAG)
Attacker access
Partial control of training corpus or labels
Impact
Accuracy loss, backdoors, biased or malicious outputs
Standards
NIST AI 100-2, OWASP LLM04
Related
Adversarial machine learning, prompt injection, AI red teaming
Data poisoning is an attack on an AI system in which an adversary deliberately contaminates the data used to train or ground a model, causing it to behave incorrectly, untrustfully or on command. NIST defines it as an attack in which "an adversary controls a subset of the training data by either inserting or modifying training samples," and treats it as one of the principal categories of adversarial machine learning.[1][3]

Types

Attacks are usually grouped by what the attacker achieves:

  • Availability poisoning degrades overall model performance — for example inflating error rates so a classifier becomes useless. It needs no precision, only volume.
  • Targeted (integrity) poisoning changes behaviour on a narrow slice of inputs: a stop-sign detector that fails on a specific sticker pattern, or a filter that always passes one attacker-controlled keyword.
  • Backdoor (Trojan) poisoning implants a trigger–response pair. Clean inputs behave normally; inputs containing an invisible trigger activate the malicious behaviour, which makes the attack very hard to detect in evaluation.
  • Clean-label poisoning modifies only features, never labels, so the poisoned samples look legitimate to human reviewers — a significant concern for crowdsourced datasets.
Poisoning can occur at every stage of the pipeline: pre-training corpora scraped from the web, instruction-tuning sets, human-feedback data used in RLHF, and the document collections that retrieval-augmented generation systems index. The last is often called RAG poisoning and requires no access to model weights at all — an attacker only needs to publish a page that a retriever will trust.

Why It Matters

Large models are trained on data at scales no human can inspect, and the default sourcing strategy — crawl the public internet — is precisely an untrusted supply chain. A tiny fraction of poisoned documents can be enough: research has shown backdoors can be implanted with a small proportion of contaminated samples, and the effect generalises to a subpopulation without detailed knowledge of the model.[1] As AI systems move into code generation, finance and public services, the consequence shifts from wrong answers to wrong actions.

The risk is recognised in standards. OWASP lists LLM04: Data and Model Poisoning among the top risks for LLM applications, framing it as an integrity attack since tampered training data undermines the model's ability to make accurate predictions.[2] NIST's Adversarial Machine Learning report (AI 100-2) catalogues poisoning alongside evasion, inference and extraction attacks.[1]

Defences

There is no single fix; defences are layered:

  • Provenance and vetting of training sources, with documented data lineage (the same discipline as a software bill of materials).
  • Statistical filtering — outlier detection, deduplication and clustering to find samples that deviate from the corpus distribution.
  • Canary and influence checks — measuring how much individual documents shift model outputs.
  • Robust training methods, including differential privacy techniques that bound any single record's influence, and ensemble or data-cartography methods that downweight suspicious samples.
  • Trigger scanning and adversarial evaluation before release, plus RAG-layer controls: allow-lists, source trust scores and freshness checks on retrieved content.[4]
None of these are foolproof against a determined adversary with sustained access, which is why provenance, monitoring after deployment, and reporting channels matter as much as detection algorithms.

>Key Takeaways

  • Data poisoning contaminates training or retrieval data to degrade or redirect a model.
  • Attacks range from blunt availability loss to hard-to-detect backdoors and clean-label manipulation.
  • RAG pipelines are a poisoning target that needs no access to model weights.
  • OWASP LLM04 and NIST AI 100-2 treat it as a first-class risk; defences are layered, not singular.
  • Provenance, filtering, post-deployment monitoring and incident reporting are the practical baseline.

See Also

🇲🇾Malaysian Context

🇲🇾 Data poisoning is a live concern for Malaysian organisations precisely because most local AI adoption runs on third-party models and scraped corpora. Under the Personal Data Protection Act, the Data Integrity Principle requires personal data to be accurate, complete and not misleading — a principle that maps directly onto training and retrieval data quality, and the 2024 PDPA amendments tightened breach-notification duties that would cover tainted datasets. Malaysia's AI Governance and Ethics Framework asks organisations to demonstrate reliability and safety, which for any firm fine-tuning on internal records means documenting where that data came from and who could have written to it. Sectors with concentrated exposure — banking (Bank Negara's risk guidance), healthcare, and the public sector — have the strongest case for provenance controls. Local bodies such as SIRIM and CyberSecurity Malaysia, together with MDEC's AI Nation 2030 agenda, are the natural homes for testing and certification capability. Because so much Malaysian content is bilingual and scraped from open web sources, small published datasets are unusually easy for an attacker to influence proportionally.

References

  1. ↑NIST. Adversarial Machine Learning: Taxonomy and Terminology (NIST AI 100-2). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-2e2025.pdf
  2. ↑OWASP GenAI Security Project. LLM04:2025 Data and Model Poisoning. https://genai.owasp.org/llmrisk/llm042025-data-and-model-poisoning/
  3. ↑NIST CSRC Glossary. data poisoning. https://csrc.nist.gov/glossary/term/data_poisoning
  4. ↑IBM. What Is Data Poisoning? https://www.ibm.com/think/topics/data-poisoning