Sensitive Data, Privacy, and the AI Supply Chain
CAP, ACID vs BASE, latency numbers, back-of-envelope estimation, single points of failure — the vocabulary every system designer thinks in.
Core Philosophy: AI systems have a complicated relationship with data — they are trained on it, they process it, they hold it in context, and they can leak it in ways traditional software cannot. And most AI systems are built on models and components from third parties you do not control. This page closes Phase 6A’s defensive coverage with two themes you already know — data protection and supply chain — reapplied to the specific, and sometimes surprising, ways they manifest in AI systems.
Part 1: The Problem
Two loose ends remain from the OWASP LLM Top 10 (6.2): sensitive information disclosure and supply chain vulnerabilities. Both are about things you have studied before — data protection (the confidentiality of the CIA triad, 1.1; data handling throughout Phase 4) and supply-chain risk (2.10) — but both manifest in AI systems in specific, and sometimes non-obvious, ways.
Data and AI: AI systems have an unusually intimate and multi-faceted relationship with data. They are trained on data, they process data, they hold data in their context, their tools may reach data — and, uniquely, they can leak data in ways traditional software simply cannot. An AI application that mishandles data fails the most fundamental security goal (confidentiality, 1.1) — and AI gives it new ways to fail.
The AI supply chain: almost nobody builds an AI system entirely from scratch. The model is usually a third party’s; the frameworks, libraries, and tools are third parties’; training data may come from external sources. You are depending — heavily — on things you did not build and cannot fully inspect. That is supply-chain risk (2.10), and in AI it is pervasive and significant.
This page closes the defensive half of Phase 6 by covering both.
Part 2: The Concept — How AI Systems Expose Data
AI systems can disclose sensitive data through several distinct channels — and understanding each is necessary to defend it. This is sensitive information disclosure (6.2), examined.
Leakage of training data. A model can, in its outputs, reveal information from the data it was trained on. If a model was trained (or fine-tuned) on sensitive data — personal data, confidential documents, proprietary information — that data can, in various circumstances, surface in what the model produces, or be inferred from its behavior (recall membership inference, 6.4). The principle: whatever sensitive data a model is trained on becomes a potential disclosure risk for the life of that model.
Leakage of context data. Most LLM applications place data in the model’s context at runtime — the user’s data, retrieved documents, conversation history, and (recall 6.3) the system instructions. Anything in the model’s context can potentially be disclosed — through prompt injection (6.3) deliberately extracting it, or through the model simply including it in an output when it should not. Whatever you put in the model’s context, the model can potentially reveal.
Cross-user data exposure. If an AI application is not carefully designed, data from one user can leak to another — through shared context, mishandled conversation history, or caching. This is broken access control (2.7) in an AI setting: each user’s data must remain isolated to that user.
Disclosure through tools. If an AI agent has tools that can access sensitive data, prompt injection (6.3) can drive those tools to retrieve and disclose that data. The model’s reach — via its tools — defines what it can leak.
Disclosure of the system instructions. The developer’s system prompt can be extracted via prompt injection (6.3). It should not be treated as a secret store in the first place — but its exposure can reveal application logic and should be assumed possible.
HOW AN AI SYSTEM CAN LEAK DATA
• from its TRAINING DATA → surfaces in outputs / inference
• from its CONTEXT → user data, retrieved docs,
history, system prompt
• ACROSS USERS → one user's data reaching another
• via its TOOLS → tools reaching sensitive data
─────────────────────────────────────────────────────────────
Whatever a model is trained on, given in context, or can
reach through tools — it can potentially disclose.
Part 3: The Concept — Privacy in AI Systems
Closely tied to data disclosure is privacy — and AI systems raise privacy issues that deserve their own focus.
The core privacy concerns with AI systems:
- User input becomes data the system holds. When users interact with an AI system, their inputs — which may contain personal, sensitive, or confidential information — become data that is processed and may be stored, logged, or (a critical question) used for further training. Whether user input is used to train or improve models is a significant privacy question, and it must be handled deliberately and transparently.
- AI systems may process personal data at scale. An AI feature that processes user data is subject to the same data-protection expectations and obligations as any other system handling personal data — and may be subject to privacy regulations (the legal-and-regulatory awareness from 1.0 and 4.7; specifics are jurisdiction-dependent and not legal advice — but be aware the obligations exist).
- Sensitive data in training is a lasting privacy commitment. If personal data is used to train a model, that is a privacy decision with long-lasting consequences (the model may retain and disclose it — Part 2). Using sensitive data in training is something to do only with great care, proper basis, and full awareness.
- Logging and AI. AI systems should be monitored and logged (6.5, 4.6) — but logging AI interactions means logging user inputs, which may be sensitive. The 4.6 lesson “never log sensitive data” collides with the need to monitor AI systems — and must be navigated thoughtfully (log what is needed for security, protect those logs, avoid logging sensitive content unnecessarily).
The defenses are the data-protection principles you already hold, applied with care to AI:
- Data minimization — the most important one. Give the AI system, its context, and its tools access to the minimum sensitive data genuinely needed. Do not put sensitive data in a model’s context or training set unless truly necessary. This is least privilege (0.2, 4.3) applied to data in AI systems — and it directly shrinks what can possibly leak (Part 2).
- Encryption and access control — protect the data AI systems handle, at rest and in transit (1.3), with proper access control (4.2).
- Be deliberate and transparent about training use — treat “is user data used for training?” as a serious, explicit decision.
- Isolate users’ data — ensure one user’s data cannot reach another (the access-control discipline, 2.7/4.2).
- Apply privacy-by-design — consider privacy from the start (the secure-design thinking of 4.3, applied to privacy).
Part 4: The Concept — The AI Supply Chain
The second theme: AI systems are built on a supply chain of things you do not control — and this is the supply-chain risk of 2.10, manifesting pervasively in AI.
Recall 2.10’s lesson: modern software is assembled from third-party components, and you inherit the security (and insecurity) of everything you depend on. AI systems take this further — the most central component, the model, is itself usually a third party’s.
The elements of the AI supply chain:
- The model itself. Most AI applications use a model built by someone else — accessed as a service, or obtained and run. You are trusting the model’s creator: trusting that its training data was not poisoned and contains no backdoor (6.4), that it was built responsibly, that it behaves as described. You generally cannot verify this — it is a genuine trust decision.
- AI frameworks and libraries. The software libraries and frameworks used to build AI applications are third-party dependencies — subject to exactly the known-vulnerability and dependency risks of 2.10. They must be tracked and kept current (the SCA discipline of 2.10/5B.3).
- Pre-trained models and model hubs. Models are often obtained from repositories/hubs. As with container images (5C.3) and any downloaded artifact, the source matters — an untrustworthy or tampered model is a supply-chain compromise. Obtain models from trustworthy sources; verify integrity where possible.
- Datasets. Training or fine-tuning data obtained from external sources carries the data-poisoning and contamination risk of 6.4 — the dataset is part of the supply chain, and its provenance matters.
- Third-party AI services and plugins. AI applications often integrate third-party AI services, APIs, and plugins — each an external dependency, each a trust relationship, each part of the supply chain and the attack surface.
THE AI SUPPLY CHAIN — what you depend on but don't control
• the MODEL ← trusting its creator, its training,
that it has no backdoor (6.4)
• AI FRAMEWORKS/LIBS ← third-party dependencies (2.10)
• PRE-TRAINED MODELS ← source & integrity matter (like 5C.3)
• DATASETS ← provenance; poisoning risk (6.4)
• 3rd-PARTY AI SERVICES← external trust relationships
─────────────────────────────────────────────────────────
You inherit the security of everything you depend on.
Part 5: The Concept — Managing AI Supply Chain Risk
Knowing the AI supply chain is a risk, how is it managed? The defenses are the supply-chain disciplines of 2.10, applied to AI:
- Know your AI supply chain. You cannot manage what you have not inventoried — know which models, frameworks, libraries, datasets, and third-party AI services your system depends on. (This is the asset-inventory and “know your components” lesson of 2.10/4.8.)
- Choose trustworthy sources. Obtain models, frameworks, and datasets from reputable, trustworthy sources. The source’s trustworthiness is a real part of your security (the 5C.3 trusted-image lesson, and 2.10).
- Track and update AI dependencies. AI frameworks and libraries have known vulnerabilities like any software — track them, scan them (SCA, 5B.3), keep them current (2.10).
- Vet third-party models and services. When adopting a third-party model or AI service, evaluate it — what is known about how it was built, its security, its data handling, its trustworthiness. You cannot verify everything, but you can make an informed trust decision rather than a blind one.
- Verify integrity where possible. For models and artifacts you obtain, verify they have not been tampered with (the integrity discipline of 2.10/5C.3).
- Apply defense in depth around third-party components. Since you cannot fully trust or verify a third-party model, design around that: treat the model as an untrusted component (6.5), limit what it can reach and do (least privilege, 6.5), monitor it (4.6). The 6.5 defensive architecture is itself a major part of managing supply-chain risk — if you cannot be sure the model is trustworthy, you build so that an untrustworthy model can do limited harm.
- Recognize the irreducible trust. Be honest: some AI supply-chain trust is irreducible — when you use a major third-party model, you cannot independently verify its training data was clean. The mature response is to acknowledge that trust, make it consciously, source from reputable providers, and contain the risk through the 6.5 architecture — not to pretend the risk is zero.
Part 6: The Concept — Closing the Defensive Half of Phase 6
This page closes Phase 6A — the securing AI systems half of the phase. Step back and see the whole defensive picture you have built.
SECURING AI SYSTEMS — the complete picture (Phase 6A)
6.1 WHY AI breaks differently — the model is a new
attack surface; AI security is ADDED to Phases 0–5
6.2 THE MAP — the OWASP LLM Top 10, shared vocabulary
6.3 PROMPT INJECTION — the headline attack (running model)
6.4 TRAINING-TIME & MODEL ATTACKS — data poisoning,
model theft (the data and the artifact)
6.5 SECURING LLM APPS & AGENTS — the defensive core:
untrusted-component stance, least privilege, human-
in-the-loop, defense in depth
6.6 DATA, PRIVACY & SUPPLY CHAIN — protecting the data
AI handles; managing what you depend on
The unifying lessons of Phase 6A:
- An AI system is software plus a model — secure the software with all of Phases 0–5, and the model layer with Phase 6. Both, always.
- The model is an untrusted component — its input, its output, and what it can reach and do all need the trust-boundary and least-privilege discipline.
- You cannot perfectly prevent AI-specific attacks — prompt injection especially — so you manage them with defense in depth, limiting the blast radius.
- AI security is your existing security knowledge, re-applied — injection, trust boundaries, least privilege, the CIA triad, data protection, supply chain, defense in depth, threat modeling, monitoring. Phase 6 is the proof that the foundation was worth building: you could not have secured an AI system without it, and with it, AI security is a domain you can genuinely operate in.
And it closes by returning to your founding thesis. The whole of Phase 6A is the answer to “everyone ships AI, nobody secures it.” You now can secure it — you understand how AI systems break (6.1–6.4) and how to build them securely (6.5–6.6). That is a genuinely scarce, genuinely valuable capability.
🔑 The deep lesson: AI systems have an intimate, multi-faceted relationship with data — trained on it, processing it, holding it in context, reachable by tools — and can leak it in ways traditional software cannot; protect it with data minimization above all, plus encryption, access control, user isolation, and privacy-by-design. AI systems are also built on a supply chain — the model itself, frameworks, datasets, third-party services — that you depend on but cannot fully control or verify; manage it by knowing your dependencies, choosing trustworthy sources, tracking and updating, vetting third-party models, and containing irreducible trust through the 6.5 defensive architecture. Both themes are your existing principles — data protection and supply chain — re-applied to AI. This closes the defensive half of Phase 6: you can now secure AI systems.
📓 Key Terms
| Term | Plain meaning |
|---|---|
| Sensitive information disclosure (AI) | An AI system revealing data it should not — from training, context, or via tools. |
| Context (LLM) | The data placed before the model at runtime — user data, retrieved content, history, system prompt. |
| Data minimization | Giving an AI system access only to the minimum data genuinely needed. |
| Cross-user data exposure | One user’s data leaking to another through a poorly-designed AI system. |
| AI supply chain | The third-party models, frameworks, datasets, and services an AI system depends on. |
| Pre-trained model | A model built by a third party and adopted rather than trained in-house. |
| Model hub / repository | A source from which pre-trained models are obtained. |
| Irreducible trust | Supply-chain trust that cannot be verified away — to be acknowledged and contained, not ignored. |
🧪 Hands-On Lab
Use the AI app from earlier Phase 6 labs, and reason about real AI systems and their dependencies.
Task 1 — Map your AI app’s data. For your LLM app, list every place sensitive data could be: training/fine-tuning data (if any), the model’s context (user input, retrieved content, history, system prompt), data its tools can reach. This is your data-leakage surface (Part 2).
Task 2 — Apply data minimization. For each item in Task 1, ask: does the model genuinely need this? Remove or restrict what it does not. Practice least-privilege-for-data — and notice how it shrinks the leakage surface.
Task 3 — Test for context disclosure. Using prompt injection (6.3), try to make your app reveal something from its context — its system prompt, or other context data. Then apply Task 2’s minimization and confirm there is now less to disclose.
Task 4 — Check user isolation. If your app handles multiple users (or simulate it), verify one user’s data cannot reach another — no shared context, no leaking history. Practice finding AI-context broken access control (2.7).
Task 5 — Inventory your AI supply chain. For your AI app, list every supply-chain element: the model (and its provider), every AI framework and library, any datasets, any third-party AI services. This is the “know your AI supply chain” task.
Task 6 — Vet a third-party model. Take a third-party model you use or could use. Write what you can and cannot determine about its trustworthiness — how it was built, its data handling, its security. Note the irreducible trust, and how the 6.5 architecture contains it.
Task 7 — Reason about privacy. Write a short privacy analysis of your AI app: what user data does it handle? is any used for training? what are the obligations and risks? Practice privacy-by-design thinking.
Task 8 — Complete your AI security note. In Notion, add “Data, Privacy & AI Supply Chain” to your AI Security material — how AI systems leak data, privacy concerns, the AI supply chain, and how to manage both. With this, your Phase 6A defensive reference is complete.
⚠️ Common Mistakes
- Putting sensitive data in a model’s context or training without need. Whatever a model is trained on, given in context, or can reach, it can potentially leak. Data minimization first.
- Forgetting models leak training data. Sensitive data used to train a model is a lasting disclosure risk. Using sensitive data in training is a serious, careful decision.
- Ignoring cross-user data isolation. Poorly-designed AI apps leak one user’s data to another — broken access control in an AI setting. Isolate users’ data.
- Not knowing the AI supply chain. You cannot manage dependencies you have not inventoried. Know your models, frameworks, datasets, and third-party AI services.
- Blindly trusting third-party models. Adopting a model means trusting its creator and its training (6.4) — something you cannot fully verify. Vet sources; make the trust conscious.
- Ignoring AI framework/library vulnerabilities. AI dependencies have known vulnerabilities like any software. Track, scan, and update them (2.10/5B.3).
- Pretending supply-chain trust is zero-risk. Some AI supply-chain trust is irreducible. Acknowledge it, source reputably, and contain it with the 6.5 architecture — do not pretend it away.
✅ Recap & What’s Next
- AI systems have a multi-faceted relationship with data — trained on it, processing it, holding it in context, reaching it via tools — and can leak it from training data, context, across users, and through tools; defend with data minimization above all, plus encryption, access control, user isolation, and privacy-by-design.
- AI systems depend on a supply chain — the model, frameworks, datasets, third-party services — that you cannot fully control or verify; manage it by knowing your dependencies, choosing trustworthy sources, tracking/updating, vetting third-party models, and containing irreducible trust through the 6.5 architecture.
- This closes the defensive half of Phase 6 — you can now secure AI systems, using your existing principles re-applied.
Next — Phase 6 (Part B): Part A secured AI systems. Part B turns the lens around — using AI for security work: AI as a security tool (6.7), the AI-augmented security workflow (6.8), and the threat landscape of AI-powered attacks (6.9). The defensive half is done; the offensive-and-defensive use of AI comes next.
📋 Phase 6 (Part A) — Page Checklist
Tick each page when its reading and its hands-on lab are done.
- [ ] 6.1 — Why AI Systems Break Differently
- [ ] 6.2 — The OWASP Top 10 for LLM Applications
- [ ] 6.3 — Prompt Injection and Jailbreaking
- [ ] 6.4 — Data Poisoning, Model Theft, and Training-Time Attacks
- [ ] 6.5 — Securing LLM Applications and AI Agents
- [ ] 6.6 — Sensitive Data, Privacy, and the AI Supply Chain
Continue to Phase 6 (Part B) for pages 6.7–6.9 — using AI for security work.
Keep growing your living pages:
- [ ] Master Glossary — append every 📓 Key Terms box above.
- [ ] AI Security note — the foundation, the LLM Top 10, prompt injection, training/model attacks, securing LLM apps, data/privacy/supply chain.
- [ ] OWASP LLM Top 10 reference and checklist (6.2).
🔑 The Phase 6A throughline: an AI system is software plus a model — secure both. The model is an untrusted component; its input, output, and reach all need discipline. AI-specific attacks (prompt injection especially) cannot be perfectly prevented, so they are managed with defense in depth that limits the blast radius. And all of it is your existing security knowledge — injection, trust boundaries, least privilege, CIA, supply chain, defense in depth — re-applied to a powerful new component. This is the answer to “everyone ships AI, nobody secures it.”⁂ Back to all modules