- Type
- Adversarial attack on machine learning
- Target
- Training data (pre-training, fine-tuning, RAG)
- Attacker access
- Partial control of training corpus or labels
- Impact
- Accuracy loss, backdoors, biased or malicious outputs
- Standards
- NIST AI 100-2, OWASP LLM04
- Related
- Adversarial machine learning, prompt injection, AI red teaming
- Type
- Adversarial attack on machine learning
- Target
- Training data (pre-training, fine-tuning, RAG)
- Attacker access
- Partial control of training corpus or labels
- Impact
- Accuracy loss, backdoors, biased or malicious outputs
- Standards
- NIST AI 100-2, OWASP LLM04
- Related
- Adversarial machine learning, prompt injection, AI red teaming
Types
Attacks are usually grouped by what the attacker achieves:
- Availability poisoning degrades overall model performance — for example inflating error rates so a classifier becomes useless. It needs no precision, only volume.
- Targeted (integrity) poisoning changes behaviour on a narrow slice of inputs: a stop-sign detector that fails on a specific sticker pattern, or a filter that always passes one attacker-controlled keyword.
- Backdoor (Trojan) poisoning implants a trigger–response pair. Clean inputs behave normally; inputs containing an invisible trigger activate the malicious behaviour, which makes the attack very hard to detect in evaluation.
- Clean-label poisoning modifies only features, never labels, so the poisoned samples look legitimate to human reviewers — a significant concern for crowdsourced datasets.
Why It Matters
Large models are trained on data at scales no human can inspect, and the default sourcing strategy — crawl the public internet — is precisely an untrusted supply chain. A tiny fraction of poisoned documents can be enough: research has shown backdoors can be implanted with a small proportion of contaminated samples, and the effect generalises to a subpopulation without detailed knowledge of the model.[1] As AI systems move into code generation, finance and public services, the consequence shifts from wrong answers to wrong actions.
The risk is recognised in standards. OWASP lists LLM04: Data and Model Poisoning among the top risks for LLM applications, framing it as an integrity attack since tampered training data undermines the model's ability to make accurate predictions.[2] NIST's Adversarial Machine Learning report (AI 100-2) catalogues poisoning alongside evasion, inference and extraction attacks.[1]
Defences
There is no single fix; defences are layered:
- Provenance and vetting of training sources, with documented data lineage (the same discipline as a software bill of materials).
- Statistical filtering — outlier detection, deduplication and clustering to find samples that deviate from the corpus distribution.
- Canary and influence checks — measuring how much individual documents shift model outputs.
- Robust training methods, including differential privacy techniques that bound any single record's influence, and ensemble or data-cartography methods that downweight suspicious samples.
- Trigger scanning and adversarial evaluation before release, plus RAG-layer controls: allow-lists, source trust scores and freshness checks on retrieved content.[4]
>Key Takeaways
- Data poisoning contaminates training or retrieval data to degrade or redirect a model.
- Attacks range from blunt availability loss to hard-to-detect backdoors and clean-label manipulation.
- RAG pipelines are a poisoning target that needs no access to model weights.
- OWASP LLM04 and NIST AI 100-2 treat it as a first-class risk; defences are layered, not singular.
- Provenance, filtering, post-deployment monitoring and incident reporting are the practical baseline.
See Also
🇲🇾 Data poisoning is a live concern for Malaysian organisations precisely because most local AI adoption runs on third-party models and scraped corpora. Under the Personal Data Protection Act, the Data Integrity Principle requires personal data to be accurate, complete and not misleading — a principle that maps directly onto training and retrieval data quality, and the 2024 PDPA amendments tightened breach-notification duties that would cover tainted datasets. Malaysia's AI Governance and Ethics Framework asks organisations to demonstrate reliability and safety, which for any firm fine-tuning on internal records means documenting where that data came from and who could have written to it. Sectors with concentrated exposure — banking (Bank Negara's risk guidance), healthcare, and the public sector — have the strongest case for provenance controls. Local bodies such as SIRIM and CyberSecurity Malaysia, together with MDEC's AI Nation 2030 agenda, are the natural homes for testing and certification capability. Because so much Malaysian content is bilingual and scraped from open web sources, small published datasets are unusually easy for an attacker to influence proportionally.
References
- ↑NIST. Adversarial Machine Learning: Taxonomy and Terminology (NIST AI 100-2). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-2e2025.pdf
- ↑OWASP GenAI Security Project. LLM04:2025 Data and Model Poisoning. https://genai.owasp.org/llmrisk/llm042025-data-and-model-poisoning/
- ↑NIST CSRC Glossary. data poisoning. https://csrc.nist.gov/glossary/term/data_poisoning
- ↑IBM. What Is Data Poisoning? https://www.ibm.com/think/topics/data-poisoning