Data Poisoning, Model Theft, and Training-Time Attacks
CAP, ACID vs BASE, latency numbers, back-of-envelope estimation, single points of failure — the vocabulary every system designer thinks in.
Core Philosophy: Prompt injection (6.3) attacks an AI system through its running input. But an AI system has two more things that traditional software does not — the data it was trained on and the model artifact itself — and both are attack targets. These attacks are less common for most application builders to face directly than prompt injection, but understanding them completes your picture of the AI attack surface: the model can be corrupted at its formation, stolen, and interrogated.
Part 1: The Problem
Page 6.3 covered attacks on the running model — manipulating it through the input it receives while operating. But recall the 6.1 anatomy and the idea that “the model itself is an attack surface.” There are two more things in an AI system that have no equivalent in traditional software:
- the training data — the data the model learned from, which shaped what the model is; and
- the model itself — a valuable artifact that also “remembers” things about its training data.
Both are attackable. Training-time attacks corrupt the model at its formation by attacking its training data. Model-targeted attacks steal the model, or extract information from it.
An honest framing up front: for most practitioners building applications on top of existing models (the most common situation), prompt injection (6.3) and the application-security issues (6.5, 6.6) are the day-to-day concerns. The attacks on this page matter most to those who train or fine-tune models, or handle sensitive training data — and they matter to everyone as essential awareness, because they complete the map of the AI attack surface and because the AI supply chain (6.6) means you may depend on models that were exposed to these attacks. This page is therefore a conceptual page — understand the attacks and why they matter; deep practice in defending them is a specialization.
Part 2: The Concept — Training Data Poisoning
Training data poisoning is an attack on the integrity of a model by corrupting the data it learns from.
Recall the foundational fact from 6.1: a model is not programmed — it is trained, learning its behavior from training data. That creates a profound dependency: the model is only as trustworthy as the data it learned from. Corrupt the training data, and you corrupt the resulting model — at its very formation.
How poisoning happens:
- Deliberate poisoning. An attacker who can influence the training data inserts crafted data designed to make the resulting model behave in a flawed or attacker-chosen way — for example, to make it produce wrong outputs in certain situations, or to embed a hidden flaw.
- Backdoor attacks — a notable sub-case. An attacker poisons the training data so the model behaves normally almost always, but produces an attacker-chosen output when it sees a specific, secret “trigger.” The model passes normal testing — the malicious behavior only appears for the trigger the attacker knows. A hidden, deliberately-planted flaw.
- Contamination, not just attack. Training data can also be corrupted non-maliciously — by low-quality, biased, or simply wrong data getting into the training set. The model faithfully learns whatever is in its data, including its flaws.
Why it matters and why it is hard:
- It strikes at the model’s foundation — a poisoned model is flawed no matter how well the surrounding application is secured.
- Models train on enormous amounts of data, often gathered broadly (including from the internet). The larger and more broadly-sourced the training data, the harder it is to guarantee none of it is poisoned or contaminated.
- A backdoor can be invisible to normal testing — the model looks fine until the trigger appears.
The defense, conceptually (this is the integrity discipline from the CIA triad, 1.1): control and vet the training data — its sources, its integrity, its quality; treat the training data as a security-critical asset; and, for those consuming third-party models, recognize that you are trusting whoever trained them (the supply-chain connection to 6.6).
Part 3: The Concept — Model Theft and Extraction
The model itself is a valuable asset — and assets get stolen. Model theft is the unauthorized acquisition of a model.
Two forms:
- Direct theft. The model artifact — the actual trained model — is stolen, the way any valuable file or intellectual property can be stolen: through a breach, an insider, an exposed storage location (the exposed-storage risk of 5C.2 applies — a model left in misconfigured cloud storage can be taken). A trained model can represent enormous investment; it is worth stealing.
- Model extraction (model stealing through queries). Subtler: an attacker who can query a model repeatedly can use its responses to reconstruct an approximation of it — effectively rebuilding a copy of the model’s capability without ever accessing the original artifact. The model is treated as an oracle and progressively replicated through interaction.
Why it matters: a model can be valuable intellectual property and a competitive asset; theft is a confidentiality loss (CIA, 1.1) with real business cost. And a stolen or extracted model can then be studied at leisure to find ways to attack the original (e.g. to craft adversarial examples — Part 4 — against it).
Defenses, conceptually: protect the model artifact like the valuable asset it is — access control, secure storage, the secure-configuration discipline of 5C.2 (do not leave it in exposed storage); and, against extraction, controls on how the model can be queried (rate limiting and monitoring of query patterns — the abuse-prevention thinking of 2.6/4.2, applied to model access).
Part 4: The Concept — Adversarial Examples and Membership Inference
Two more attacks complete the conceptual map — both exploiting the nature of how models work.
Adversarial examples. An adversarial example is an input deliberately crafted to make a model produce a wrong output, even though the input looks normal (or nearly normal) to a human. Because a model’s decision-making is learned and statistical (not rule-based — 6.1), it can be reliably fooled by inputs designed to exploit how it makes decisions — inputs that a person would classify correctly but the model gets wrong. This is a way of manipulating a model’s output by exploiting its learned, probabilistic nature. It matters wherever a model’s decisions are relied upon — because it means a model’s output can be deliberately, reliably wrong for a crafted input.
Membership inference. This attack targets privacy. A membership inference attack tries to determine whether a specific piece of data was part of a model’s training set — by probing the model and analyzing its responses. Why that matters: if a model was trained on sensitive data (medical records, personal data), then confirming that a particular person’s data was in the training set is itself a privacy violation — it leaks information about the training data through the model’s behavior. It is one of several ways a model can disclose information about its training data — a theme that connects directly to sensitive information disclosure and privacy (6.6).
Both attacks share the Phase 6 theme: they exploit the fact that a model is a learned statistical artifact. Adversarial examples exploit how it decides; membership inference exploits what it remembers. Neither has an equivalent in traditional, explicitly-programmed software.
Part 5: The Concept — Putting the Model-and-Data Attacks in Perspective
With prompt injection (6.3) and this page’s attacks, you now have the full map of the AI-specific attack surface. It is worth organizing it and being honest about relative priorities.
THE AI-SPECIFIC ATTACK SURFACE — organized
ATTACKS ON THE RUNNING MODEL (via input):
• prompt injection / jailbreaking → 6.3
• adversarial examples (crafted wrong-output inputs)
• model extraction (replicating via queries)
ATTACKS ON THE TRAINING DATA (the model's formation):
• training data poisoning / backdoors
ATTACKS ON / VIA THE MODEL ARTIFACT:
• model theft (stealing the model)
• membership inference (extracting training-data info)
Honest prioritization for you, as someone most likely building applications on existing models:
- Prompt injection (6.3) and the application-security issues (6.5, 6.6) are your primary, day-to-day concern. If you build an LLM-powered feature, these are what you must get right.
- Training-time attacks (poisoning) matter most if you train or fine-tune models, or curate training data. If you do, training-data integrity becomes a first-order concern.
- Model theft matters if you own valuable models — then protecting the artifact (5C.2’s secure storage and access control) is real.
- All of them matter as awareness — because the AI supply chain (6.6) means you often use models other people trained. When you adopt a third-party model, you are trusting that its training data was not poisoned, that it has no backdoor. You cannot verify that directly — which is exactly why the supply-chain trust question (6.6) is serious.
The key takeaway: these attacks complete your understanding of why AI systems break differently — the model can be corrupted at its formation, stolen, fooled, and made to leak its training data. Whether you defend against them directly depends on your role; understanding them is essential regardless, because they shape the trust you can place in any model.
Part 6: The Concept — Defenses, and the Bridge to Securing AI Applications
A consolidated, conceptual view of how these attacks are defended — and the bridge to the defensive core of the phase.
Defending the training data (against poisoning):
- Treat training data as a security-critical asset — its sources vetted, its integrity protected, its quality controlled (the integrity discipline of the CIA triad, 1.1).
- Know the provenance of training data — where it came from, whether the source is trustworthy (the supply-chain thinking of 2.10).
- Validate and monitor model behavior to catch signs of poisoning or backdoors — though, honestly, backdoors are hard to detect.
Defending the model artifact (against theft):
- Protect it with access control and secure storage — it is a valuable asset; do not leave it exposed (the 5C.2 misconfiguration discipline).
- Control and monitor how the model can be queried — against extraction and excessive probing (the rate-limiting/monitoring thinking of 2.6/4.2/4.6).
Defending against output-manipulation and inference attacks:
- Recognize that a model’s outputs can be deliberately wrong (adversarial examples) — so, exactly as with prompt injection, do not blindly trust or act on model output (the theme that becomes central in 6.5).
- Be careful what data a model is trained on, given that models can leak information about their training data (membership inference) — the privacy theme of 6.6.
The honest, unifying point: several of these attacks are areas of active research without complete solutions — much like prompt injection (6.3), they are managed and mitigated rather than perfectly solved. And notice the recurring defensive themes: treat data as security-critical, protect the model as an asset, control access, and never blindly trust model output. These are not new principles — they are the CIA triad, least privilege, access control, secure storage, and trust boundaries, applied to the model and its data.
This bridges directly to the rest of Phase 6A. You now understand the AI-specific attacks — prompt injection (6.3) and the model/data attacks (6.4). The next two pages turn fully defensive: 6.5 is how to actually build LLM applications and agents securely — the defensive core — and 6.6 covers sensitive data, privacy, and the supply chain through which much of this risk arrives.
🔑 The deep lesson: beyond attacks on the running model (6.3), an AI system can be attacked at its training data (poisoning and backdoors corrupt the model at its formation) and at the model artifact itself (theft, and extraction through queries) — and the model can be fooled (adversarial examples) and made to leak its training data (membership inference). These exploit the model’s nature as a learned statistical artifact — they have no traditional-software equivalent. For an application builder, prompt injection and application security are the day-to-day priority; these attacks matter most to those who train models, and as essential awareness for everyone, because adopting a third-party model means trusting that its data and formation were sound. The defenses are familiar principles — protect data integrity, protect the model as an asset, control access, never blindly trust output — applied to the model and its data.
📓 Key Terms
| Term | Plain meaning |
|---|---|
| Training-time attack | An attack targeting the model’s formation rather than the running application. |
| Training data poisoning | Corrupting a model’s training data to corrupt the resulting model. |
| Backdoor attack | Poisoning so a model behaves normally except on a secret attacker-chosen trigger. |
| Model theft | Unauthorized acquisition of the model artifact itself. |
| Model extraction | Reconstructing an approximation of a model by querying it repeatedly. |
| Adversarial example | An input crafted to make a model produce a wrong output. |
| Membership inference | Determining whether specific data was in a model’s training set — a privacy attack. |
| Data provenance | Knowledge of where data came from and whether the source is trustworthy. |
🧪 Hands-On Lab
This is a conceptual page — its labs are about understanding, reasoning, and connecting, rather than performing attacks (most of these attacks require specialized setups or model-training access).
Task 1 — Explain poisoning in your own words. In Notion, write a plain-English explanation of training data poisoning and backdoor attacks — and why “a model is only as trustworthy as its training data” follows from how models are made (6.1).
Task 2 — Study a real example. Find a reputable write-up or research summary of a data poisoning or backdoor attack (or adversarial examples). Read it for the concept and the real-world relevance — note what made the attack possible.
Task 3 — Reason about model theft as a storage problem. Write how model theft connects to 5C.2 — a model is a valuable asset; an exposed cloud storage location holding a model is the 5C.2 exposed-storage problem. Note that “protect the model” is mostly access control and secure storage you already know.
Task 4 — Reason about adversarial examples and trust. Write a short note: if a model’s output can be deliberately, reliably wrong for a crafted input, what does that mean for any system that acts on model decisions? Connect this to the “never blindly trust model output” theme building toward 6.5.
Task 5 — Connect to the supply chain. Write down: when you adopt a third-party model, which attacks from this page are you trusting were not used against it (poisoning, backdoors)? You cannot verify this — note why that makes the 6.6 supply-chain question serious.
Task 6 — Prioritize honestly for your role. Using Part 5, write down which of these attacks are your direct concern given that you will likely build on existing models — and which are awareness-level. Honest prioritization is part of the skill.
Task 7 — Extend your AI security note. In Notion, add a “Training-Time and Model Attacks” section to your AI Security material — the attack types, why they matter, the conceptual defenses, and the honest prioritization.
⚠️ Common Mistakes
- Ignoring training data as a security concern. A model is only as trustworthy as its training data. Poisoned or contaminated data produces a flawed model no surrounding security can fix.
- Assuming a model that passes testing is clean. A backdoor is designed to be invisible to normal testing — malicious behavior appears only on a secret trigger.
- Leaving model artifacts exposed. A trained model is a valuable asset. An exposed storage location holding it is the 5C.2 exposed-storage breach. Protect models like the assets they are.
- Trusting model decisions as infallible. Adversarial examples mean a model’s output can be deliberately, reliably wrong for crafted inputs. Never treat model decisions as unquestionable.
- Forgetting models can leak their training data. Membership inference and related attacks mean a model trained on sensitive data can disclose information about it. Be careful what models are trained on.
- Forgetting the supply-chain implication. Adopting a third-party model means trusting its training and formation were sound — something you cannot verify. That is a real, serious trust decision (6.6).
✅ Recap & What’s Next
- Beyond the running model, an AI system is attackable at its training data (poisoning and backdoors corrupt the model at formation) and at the model artifact (theft, and extraction via queries); models can also be fooled (adversarial examples) and made to leak training data (membership inference).
- These exploit the model’s nature as a learned statistical artifact — no traditional-software equivalent; for an application builder, prompt injection and application security are the priority, while these matter most to model-trainers and as essential awareness for everyone.
- The defenses are familiar principles — protect data integrity, protect the model as an asset, control access, never blindly trust output — applied to the model and its data.
Next (6.5): You now understand the AI-specific attacks. Page 6.5 is the defensive core of the phase — how to actually build LLM applications and AI agents securely: the controls most teams currently skip, and the ones that matter most.
⁂ Back to all modules