Home
Cybersecurity & AI Security / Part 58 — Data Poisoning, Model Theft, and Training-Time Attacks

Data Poisoning, Model Theft, and Training-Time Attacks

CAP, ACID vs BASE, latency numbers, back-of-envelope estimation, single points of failure — the vocabulary every system designer thinks in.


Core Philosophy: Prompt injection (6.3) attacks an AI system through its running input. But an AI system has two more things that traditional software does not — the data it was trained on and the model artifact itself — and both are attack targets. These attacks are less common for most application builders to face directly than prompt injection, but understanding them completes your picture of the AI attack surface: the model can be corrupted at its formation, stolen, and interrogated.

Part 1: The Problem

Page 6.3 covered attacks on the running model — manipulating it through the input it receives while operating. But recall the 6.1 anatomy and the idea that “the model itself is an attack surface.” There are two more things in an AI system that have no equivalent in traditional software:

Both are attackable. Training-time attacks corrupt the model at its formation by attacking its training data. Model-targeted attacks steal the model, or extract information from it.

An honest framing up front: for most practitioners building applications on top of existing models (the most common situation), prompt injection (6.3) and the application-security issues (6.5, 6.6) are the day-to-day concerns. The attacks on this page matter most to those who train or fine-tune models, or handle sensitive training data — and they matter to everyone as essential awareness, because they complete the map of the AI attack surface and because the AI supply chain (6.6) means you may depend on models that were exposed to these attacks. This page is therefore a conceptual page — understand the attacks and why they matter; deep practice in defending them is a specialization.

Part 2: The Concept — Training Data Poisoning

Training data poisoning is an attack on the integrity of a model by corrupting the data it learns from.

Recall the foundational fact from 6.1: a model is not programmed — it is trained, learning its behavior from training data. That creates a profound dependency: the model is only as trustworthy as the data it learned from. Corrupt the training data, and you corrupt the resulting model — at its very formation.

How poisoning happens:

Why it matters and why it is hard:

The defense, conceptually (this is the integrity discipline from the CIA triad, 1.1): control and vet the training data — its sources, its integrity, its quality; treat the training data as a security-critical asset; and, for those consuming third-party models, recognize that you are trusting whoever trained them (the supply-chain connection to 6.6).

Part 3: The Concept — Model Theft and Extraction

The model itself is a valuable asset — and assets get stolen. Model theft is the unauthorized acquisition of a model.

Two forms:

Why it matters: a model can be valuable intellectual property and a competitive asset; theft is a confidentiality loss (CIA, 1.1) with real business cost. And a stolen or extracted model can then be studied at leisure to find ways to attack the original (e.g. to craft adversarial examples — Part 4 — against it).

Defenses, conceptually: protect the model artifact like the valuable asset it is — access control, secure storage, the secure-configuration discipline of 5C.2 (do not leave it in exposed storage); and, against extraction, controls on how the model can be queried (rate limiting and monitoring of query patterns — the abuse-prevention thinking of 2.6/4.2, applied to model access).

Part 4: The Concept — Adversarial Examples and Membership Inference

Two more attacks complete the conceptual map — both exploiting the nature of how models work.

Adversarial examples. An adversarial example is an input deliberately crafted to make a model produce a wrong output, even though the input looks normal (or nearly normal) to a human. Because a model’s decision-making is learned and statistical (not rule-based — 6.1), it can be reliably fooled by inputs designed to exploit how it makes decisions — inputs that a person would classify correctly but the model gets wrong. This is a way of manipulating a model’s output by exploiting its learned, probabilistic nature. It matters wherever a model’s decisions are relied upon — because it means a model’s output can be deliberately, reliably wrong for a crafted input.

Membership inference. This attack targets privacy. A membership inference attack tries to determine whether a specific piece of data was part of a model’s training set — by probing the model and analyzing its responses. Why that matters: if a model was trained on sensitive data (medical records, personal data), then confirming that a particular person’s data was in the training set is itself a privacy violation — it leaks information about the training data through the model’s behavior. It is one of several ways a model can disclose information about its training data — a theme that connects directly to sensitive information disclosure and privacy (6.6).

Both attacks share the Phase 6 theme: they exploit the fact that a model is a learned statistical artifact. Adversarial examples exploit how it decides; membership inference exploits what it remembers. Neither has an equivalent in traditional, explicitly-programmed software.

Part 5: The Concept — Putting the Model-and-Data Attacks in Perspective

With prompt injection (6.3) and this page’s attacks, you now have the full map of the AI-specific attack surface. It is worth organizing it and being honest about relative priorities.

text
   THE AI-SPECIFIC ATTACK SURFACE — organized

   ATTACKS ON THE RUNNING MODEL (via input):
     • prompt injection / jailbreaking          → 6.3
     • adversarial examples (crafted wrong-output inputs)
     • model extraction (replicating via queries)

   ATTACKS ON THE TRAINING DATA (the model's formation):
     • training data poisoning / backdoors

   ATTACKS ON / VIA THE MODEL ARTIFACT:
     • model theft (stealing the model)
     • membership inference (extracting training-data info)

Honest prioritization for you, as someone most likely building applications on existing models:

The key takeaway: these attacks complete your understanding of why AI systems break differently — the model can be corrupted at its formation, stolen, fooled, and made to leak its training data. Whether you defend against them directly depends on your role; understanding them is essential regardless, because they shape the trust you can place in any model.

Part 6: The Concept — Defenses, and the Bridge to Securing AI Applications

A consolidated, conceptual view of how these attacks are defended — and the bridge to the defensive core of the phase.

Defending the training data (against poisoning):

Defending the model artifact (against theft):

Defending against output-manipulation and inference attacks:

The honest, unifying point: several of these attacks are areas of active research without complete solutions — much like prompt injection (6.3), they are managed and mitigated rather than perfectly solved. And notice the recurring defensive themes: treat data as security-critical, protect the model as an asset, control access, and never blindly trust model output. These are not new principles — they are the CIA triad, least privilege, access control, secure storage, and trust boundaries, applied to the model and its data.

This bridges directly to the rest of Phase 6A. You now understand the AI-specific attacks — prompt injection (6.3) and the model/data attacks (6.4). The next two pages turn fully defensive: 6.5 is how to actually build LLM applications and agents securely — the defensive core — and 6.6 covers sensitive data, privacy, and the supply chain through which much of this risk arrives.

🔑 The deep lesson: beyond attacks on the running model (6.3), an AI system can be attacked at its training data (poisoning and backdoors corrupt the model at its formation) and at the model artifact itself (theft, and extraction through queries) — and the model can be fooled (adversarial examples) and made to leak its training data (membership inference). These exploit the model’s nature as a learned statistical artifact — they have no traditional-software equivalent. For an application builder, prompt injection and application security are the day-to-day priority; these attacks matter most to those who train models, and as essential awareness for everyone, because adopting a third-party model means trusting that its data and formation were sound. The defenses are familiar principles — protect data integrity, protect the model as an asset, control access, never blindly trust output — applied to the model and its data.

📓 Key Terms

Term Plain meaning
Training-time attackAn attack targeting the model’s formation rather than the running application.
Training data poisoningCorrupting a model’s training data to corrupt the resulting model.
Backdoor attackPoisoning so a model behaves normally except on a secret attacker-chosen trigger.
Model theftUnauthorized acquisition of the model artifact itself.
Model extractionReconstructing an approximation of a model by querying it repeatedly.
Adversarial exampleAn input crafted to make a model produce a wrong output.
Membership inferenceDetermining whether specific data was in a model’s training set — a privacy attack.
Data provenanceKnowledge of where data came from and whether the source is trustworthy.

🧪 Hands-On Lab

This is a conceptual page — its labs are about understanding, reasoning, and connecting, rather than performing attacks (most of these attacks require specialized setups or model-training access).

Task 1 — Explain poisoning in your own words. In Notion, write a plain-English explanation of training data poisoning and backdoor attacks — and why “a model is only as trustworthy as its training data” follows from how models are made (6.1).

Task 2 — Study a real example. Find a reputable write-up or research summary of a data poisoning or backdoor attack (or adversarial examples). Read it for the concept and the real-world relevance — note what made the attack possible.

Task 3 — Reason about model theft as a storage problem. Write how model theft connects to 5C.2 — a model is a valuable asset; an exposed cloud storage location holding a model is the 5C.2 exposed-storage problem. Note that “protect the model” is mostly access control and secure storage you already know.

Task 4 — Reason about adversarial examples and trust. Write a short note: if a model’s output can be deliberately, reliably wrong for a crafted input, what does that mean for any system that acts on model decisions? Connect this to the “never blindly trust model output” theme building toward 6.5.

Task 5 — Connect to the supply chain. Write down: when you adopt a third-party model, which attacks from this page are you trusting were not used against it (poisoning, backdoors)? You cannot verify this — note why that makes the 6.6 supply-chain question serious.

Task 6 — Prioritize honestly for your role. Using Part 5, write down which of these attacks are your direct concern given that you will likely build on existing models — and which are awareness-level. Honest prioritization is part of the skill.

Task 7 — Extend your AI security note. In Notion, add a “Training-Time and Model Attacks” section to your AI Security material — the attack types, why they matter, the conceptual defenses, and the honest prioritization.

⚠️ Common Mistakes

✅ Recap & What’s Next

Next (6.5): You now understand the AI-specific attacks. Page 6.5 is the defensive core of the phase — how to actually build LLM applications and AI agents securely: the controls most teams currently skip, and the ones that matter most.

⁂ Back to all modules