Home
Cybersecurity & AI Security / Part 60 — Sensitive Data, Privacy, and the AI Supply Chain

Sensitive Data, Privacy, and the AI Supply Chain

CAP, ACID vs BASE, latency numbers, back-of-envelope estimation, single points of failure — the vocabulary every system designer thinks in.


Core Philosophy: AI systems have a complicated relationship with data — they are trained on it, they process it, they hold it in context, and they can leak it in ways traditional software cannot. And most AI systems are built on models and components from third parties you do not control. This page closes Phase 6A’s defensive coverage with two themes you already know — data protection and supply chain — reapplied to the specific, and sometimes surprising, ways they manifest in AI systems.

Part 1: The Problem

Two loose ends remain from the OWASP LLM Top 10 (6.2): sensitive information disclosure and supply chain vulnerabilities. Both are about things you have studied before — data protection (the confidentiality of the CIA triad, 1.1; data handling throughout Phase 4) and supply-chain risk (2.10) — but both manifest in AI systems in specific, and sometimes non-obvious, ways.

Data and AI: AI systems have an unusually intimate and multi-faceted relationship with data. They are trained on data, they process data, they hold data in their context, their tools may reach data — and, uniquely, they can leak data in ways traditional software simply cannot. An AI application that mishandles data fails the most fundamental security goal (confidentiality, 1.1) — and AI gives it new ways to fail.

The AI supply chain: almost nobody builds an AI system entirely from scratch. The model is usually a third party’s; the frameworks, libraries, and tools are third parties’; training data may come from external sources. You are depending — heavily — on things you did not build and cannot fully inspect. That is supply-chain risk (2.10), and in AI it is pervasive and significant.

This page closes the defensive half of Phase 6 by covering both.

Part 2: The Concept — How AI Systems Expose Data

AI systems can disclose sensitive data through several distinct channels — and understanding each is necessary to defend it. This is sensitive information disclosure (6.2), examined.

Leakage of training data. A model can, in its outputs, reveal information from the data it was trained on. If a model was trained (or fine-tuned) on sensitive data — personal data, confidential documents, proprietary information — that data can, in various circumstances, surface in what the model produces, or be inferred from its behavior (recall membership inference, 6.4). The principle: whatever sensitive data a model is trained on becomes a potential disclosure risk for the life of that model.

Leakage of context data. Most LLM applications place data in the model’s context at runtime — the user’s data, retrieved documents, conversation history, and (recall 6.3) the system instructions. Anything in the model’s context can potentially be disclosed — through prompt injection (6.3) deliberately extracting it, or through the model simply including it in an output when it should not. Whatever you put in the model’s context, the model can potentially reveal.

Cross-user data exposure. If an AI application is not carefully designed, data from one user can leak to another — through shared context, mishandled conversation history, or caching. This is broken access control (2.7) in an AI setting: each user’s data must remain isolated to that user.

Disclosure through tools. If an AI agent has tools that can access sensitive data, prompt injection (6.3) can drive those tools to retrieve and disclose that data. The model’s reach — via its tools — defines what it can leak.

Disclosure of the system instructions. The developer’s system prompt can be extracted via prompt injection (6.3). It should not be treated as a secret store in the first place — but its exposure can reveal application logic and should be assumed possible.

text
   HOW AN AI SYSTEM CAN LEAK DATA
   • from its TRAINING DATA      → surfaces in outputs / inference
   • from its CONTEXT            → user data, retrieved docs,
                                   history, system prompt
   • ACROSS USERS                → one user's data reaching another
   • via its TOOLS               → tools reaching sensitive data
   ─────────────────────────────────────────────────────────────
   Whatever a model is trained on, given in context, or can
   reach through tools — it can potentially disclose.

Part 3: The Concept — Privacy in AI Systems

Closely tied to data disclosure is privacy — and AI systems raise privacy issues that deserve their own focus.

The core privacy concerns with AI systems:

The defenses are the data-protection principles you already hold, applied with care to AI:

Part 4: The Concept — The AI Supply Chain

The second theme: AI systems are built on a supply chain of things you do not control — and this is the supply-chain risk of 2.10, manifesting pervasively in AI.

Recall 2.10’s lesson: modern software is assembled from third-party components, and you inherit the security (and insecurity) of everything you depend on. AI systems take this further — the most central component, the model, is itself usually a third party’s.

The elements of the AI supply chain:

text
   THE AI SUPPLY CHAIN — what you depend on but don't control
   • the MODEL            ← trusting its creator, its training,
                            that it has no backdoor (6.4)
   • AI FRAMEWORKS/LIBS   ← third-party dependencies (2.10)
   • PRE-TRAINED MODELS   ← source & integrity matter (like 5C.3)
   • DATASETS             ← provenance; poisoning risk (6.4)
   • 3rd-PARTY AI SERVICES← external trust relationships
   ─────────────────────────────────────────────────────────
   You inherit the security of everything you depend on.

Part 5: The Concept — Managing AI Supply Chain Risk

Knowing the AI supply chain is a risk, how is it managed? The defenses are the supply-chain disciplines of 2.10, applied to AI:

Part 6: The Concept — Closing the Defensive Half of Phase 6

This page closes Phase 6A — the securing AI systems half of the phase. Step back and see the whole defensive picture you have built.

text
   SECURING AI SYSTEMS — the complete picture (Phase 6A)

   6.1  WHY AI breaks differently — the model is a new
        attack surface; AI security is ADDED to Phases 0–5
   6.2  THE MAP — the OWASP LLM Top 10, shared vocabulary
   6.3  PROMPT INJECTION — the headline attack (running model)
   6.4  TRAINING-TIME & MODEL ATTACKS — data poisoning,
        model theft (the data and the artifact)
   6.5  SECURING LLM APPS & AGENTS — the defensive core:
        untrusted-component stance, least privilege, human-
        in-the-loop, defense in depth
   6.6  DATA, PRIVACY & SUPPLY CHAIN — protecting the data
        AI handles; managing what you depend on

The unifying lessons of Phase 6A:

And it closes by returning to your founding thesis. The whole of Phase 6A is the answer to “everyone ships AI, nobody secures it.” You now can secure it — you understand how AI systems break (6.1–6.4) and how to build them securely (6.5–6.6). That is a genuinely scarce, genuinely valuable capability.

🔑 The deep lesson: AI systems have an intimate, multi-faceted relationship with data — trained on it, processing it, holding it in context, reachable by tools — and can leak it in ways traditional software cannot; protect it with data minimization above all, plus encryption, access control, user isolation, and privacy-by-design. AI systems are also built on a supply chain — the model itself, frameworks, datasets, third-party services — that you depend on but cannot fully control or verify; manage it by knowing your dependencies, choosing trustworthy sources, tracking and updating, vetting third-party models, and containing irreducible trust through the 6.5 defensive architecture. Both themes are your existing principles — data protection and supply chain — re-applied to AI. This closes the defensive half of Phase 6: you can now secure AI systems.

📓 Key Terms

Term Plain meaning
Sensitive information disclosure (AI)An AI system revealing data it should not — from training, context, or via tools.
Context (LLM)The data placed before the model at runtime — user data, retrieved content, history, system prompt.
Data minimizationGiving an AI system access only to the minimum data genuinely needed.
Cross-user data exposureOne user’s data leaking to another through a poorly-designed AI system.
AI supply chainThe third-party models, frameworks, datasets, and services an AI system depends on.
Pre-trained modelA model built by a third party and adopted rather than trained in-house.
Model hub / repositoryA source from which pre-trained models are obtained.
Irreducible trustSupply-chain trust that cannot be verified away — to be acknowledged and contained, not ignored.

🧪 Hands-On Lab

Use the AI app from earlier Phase 6 labs, and reason about real AI systems and their dependencies.

Task 1 — Map your AI app’s data. For your LLM app, list every place sensitive data could be: training/fine-tuning data (if any), the model’s context (user input, retrieved content, history, system prompt), data its tools can reach. This is your data-leakage surface (Part 2).

Task 2 — Apply data minimization. For each item in Task 1, ask: does the model genuinely need this? Remove or restrict what it does not. Practice least-privilege-for-data — and notice how it shrinks the leakage surface.

Task 3 — Test for context disclosure. Using prompt injection (6.3), try to make your app reveal something from its context — its system prompt, or other context data. Then apply Task 2’s minimization and confirm there is now less to disclose.

Task 4 — Check user isolation. If your app handles multiple users (or simulate it), verify one user’s data cannot reach another — no shared context, no leaking history. Practice finding AI-context broken access control (2.7).

Task 5 — Inventory your AI supply chain. For your AI app, list every supply-chain element: the model (and its provider), every AI framework and library, any datasets, any third-party AI services. This is the “know your AI supply chain” task.

Task 6 — Vet a third-party model. Take a third-party model you use or could use. Write what you can and cannot determine about its trustworthiness — how it was built, its data handling, its security. Note the irreducible trust, and how the 6.5 architecture contains it.

Task 7 — Reason about privacy. Write a short privacy analysis of your AI app: what user data does it handle? is any used for training? what are the obligations and risks? Practice privacy-by-design thinking.

Task 8 — Complete your AI security note. In Notion, add “Data, Privacy & AI Supply Chain” to your AI Security material — how AI systems leak data, privacy concerns, the AI supply chain, and how to manage both. With this, your Phase 6A defensive reference is complete.

⚠️ Common Mistakes

✅ Recap & What’s Next

Next — Phase 6 (Part B): Part A secured AI systems. Part B turns the lens around — using AI for security work: AI as a security tool (6.7), the AI-augmented security workflow (6.8), and the threat landscape of AI-powered attacks (6.9). The defensive half is done; the offensive-and-defensive use of AI comes next.

📋 Phase 6 (Part A) — Page Checklist

Tick each page when its reading and its hands-on lab are done.

Continue to Phase 6 (Part B) for pages 6.7–6.9 — using AI for security work.

Keep growing your living pages:

🔑 The Phase 6A throughline: an AI system is software plus a model — secure both. The model is an untrusted component; its input, output, and reach all need discipline. AI-specific attacks (prompt injection especially) cannot be perfectly prevented, so they are managed with defense in depth that limits the blast radius. And all of it is your existing security knowledge — injection, trust boundaries, least privilege, CIA, supply chain, defense in depth — re-applied to a powerful new component. This is the answer to “everyone ships AI, nobody secures it.”
⁂ Back to all modules