Why AI Systems Break Differently
CAP, ACID vs BASE, latency numbers, back-of-envelope estimation, single points of failure — the vocabulary every system designer thinks in.
Core Philosophy: An AI-powered application is still software — and everything you learned in Phases 0–5 still applies to it. But AI introduces something genuinely new: the model itself becomes an attack surface, in ways that have no equivalent in traditional software. Securing AI is not a separate discipline that replaces what you know; it is your existing security knowledge plus a new set of vulnerability classes layered on top. This page is about what that new layer is, and why it behaves so differently.
Part 1: The Problem
Across Phases 2–5 you learned to secure software: web applications, infrastructure, code, cloud systems. AI-powered applications are now everywhere — chatbots, assistants, features that summarize, generate, classify, decide — and a tempting assumption is that securing them is just “securing software, again.”
That assumption is half right, and the wrong half is dangerous.
The right half: an AI application is software. It has a front-end, a back-end, APIs, a database, infrastructure, authentication, inputs and outputs. Every vulnerability class you have learned — injection, broken access control, misconfiguration, the lot — can exist in an AI application exactly as in any other. Everything from Phases 0–5 still applies. Securing an AI app starts with all the security you already know.
The wrong half — and the reason this phase exists: AI applications also contain a model, and the model introduces vulnerability classes that have no equivalent in traditional software. A normal program does exactly what its code says. A machine-learning model does not “run code” in that sense — it produces outputs by learning patterns from data, and that fundamentally different nature creates fundamentally different weaknesses. Traditional security has no playbook for them. This page is about understanding why AI breaks differently — the foundation for everything in Phase 6.
Part 2: The Concept — How AI Systems Differ From Normal Software
To see why AI breaks differently, you need a clear (non-mathematical) picture of how AI systems differ from the software you already know.
Traditional software is explicit. A developer writes code — explicit instructions. The program does exactly what the code specifies, deterministically. Its behavior is, in principle, fully knowable by reading the code. Security analysis (Phases 2–5) rests on this: you can reason about what the code does.
Machine learning models are learned, not written. An ML model is not programmed with explicit rules. It is trained: shown large amounts of training data, from which it learns statistical patterns. The resulting model produces outputs based on those learned patterns. Crucially:
- No one wrote its “rules.” Its behavior emerges from data, not from code a person authored and can read.
- It is not fully predictable. The same kind of input can produce varying outputs; behavior on inputs unlike anything in training is genuinely uncertain.
- It is opaque. Even its creators often cannot fully explain why a model produced a particular output. This is the “black box” property of ML.
Large Language Models (LLMs) add another twist. Modern AI applications are often built on LLMs — models trained on vast text data that take in natural-language input (a “prompt”) and produce natural-language output. The defining security oddity: an LLM does not cleanly separate instructions from data. To an LLM, its instructions and the user’s input and any other text it is given are all just… text, in the same context. Hold onto that fact — it is the root of the single most important AI vulnerability (6.3).
TRADITIONAL SOFTWARE AI / ML SYSTEM
explicitly programmed learned from training data
deterministic, predictable probabilistic, not fully predictable
behavior knowable from code behavior opaque even to creators
code = instructions, (LLMs) instructions and data are
data = data (separate) NOT cleanly separated — all "text"
The model’s nature — learned, probabilistic, opaque, and (for LLMs) not separating instructions from data — is why it is an attack surface unlike any in traditional software.
Part 3: The Concept — The Model Itself Is an Attack Surface
Here is the central new idea of Phase 6: in an AI system, the model is not just a component — it is an attack surface in its own right.
In traditional software, the attack surface (0.4) is inputs, interfaces, exposed services. In an AI system, all of that still exists — plus the model adds entirely new attackable things:
- The model can be manipulated through its input. Because an LLM does not separate instructions from data (Part 2), carefully crafted input can change what the model does — hijack its behavior. This is prompt injection (6.3), and it has no clean equivalent in normal software.
- The model’s training data is an attack target. A model is shaped by its training data — so corrupting that data corrupts the model. This is data poisoning (6.4) — an attack on something traditional software does not even have.
- The model itself can be stolen or probed. The model is valuable intellectual property, and it also “remembers” things about its training data. Attackers can try to extract the model, or to extract information about its training data from it (6.4). The model is both an asset to steal and an oracle to interrogate.
- The model can be fooled. Inputs can be crafted specifically to make a model produce wrong outputs — adversarial examples (6.4).
- The model’s outputs are a risk. A model can produce harmful, wrong, biased, or sensitive output — and if an application trusts and acts on model output, that output becomes an attack vector (6.5).
AI APPLICATION — attack surface
┌─────────────────────────────────────────────────────┐
│ Traditional surface (Phases 2–5 — STILL APPLIES): │
│ front-end · APIs · back-end · database · infra · │
│ auth · all the classic vulnerability classes │
│ │
│ NEW: the MODEL and its ecosystem — │
│ • the model's INPUT → prompt injection (6.3) │
│ • the TRAINING DATA → data poisoning (6.4) │
│ • the MODEL itself → theft, extraction (6.4) │
│ • the model's OUTPUT → harmful/trusted output (6.5)│
│ • the AI SUPPLY CHAIN → third-party models (6.6) │
└─────────────────────────────────────────────────────┘
This is the core mental shift of Phase 6: an AI application’s attack surface is everything a normal application has, plus the model and its entire ecosystem — input, training data, the model artifact, output, and supply chain. Each of those new pieces is what pages 6.3–6.6 examine.
Part 4: The Concept — Anatomy of an AI Application
To secure AI systems, you need a clear picture of the components of a typical AI application and where the risk sits in each. Most current AI applications — especially LLM-based ones — share a common shape:
ANATOMY OF A TYPICAL LLM APPLICATION
[ user ] ──input──► [ application layer ]
│
│ constructs a PROMPT — often
│ combining: a system instruction +
│ user input + retrieved data/context
▼
[ the LLM / model ]
│ produces output
▼
[ application layer ]
│ may: show output to user, OR
│ act on it — call TOOLS, run code,
│ query data, trigger actions
▼
[ effects: responses, actions, data ]
Often also present:
• RETRIEVAL — the app fetches external data/documents to
feed the model as context (a common pattern)
• TOOLS / AGENTS — the model can invoke functions, call APIs,
take actions in the world (6.5 — the highest-risk pattern)
• the MODEL itself — built in-house or (usually) a third-party
model accessed as a service (6.6 — the supply chain)
Walking the components and their risk:
- The application layer — ordinary software. All of Phases 2–5 applies here. It also does something security-critical: it constructs the prompt, often by combining trusted instructions with untrusted user input and untrusted retrieved data — and that combining is where prompt injection lives (6.3).
- The prompt — what actually goes to the model. Because instructions and data are mixed in it (Part 2), the prompt is a trust-boundary danger zone.
- The model — the new attack surface (Part 3).
- Retrieved data / context — many AI apps feed the model external documents or data. If any of that retrieved content is attacker-influenced, it can carry an attack to the model (indirect prompt injection — 6.3).
- Tools / agentic capabilities — when the model can take actions (call APIs, run code, access systems), the stakes rise enormously: a manipulated model is no longer just producing bad text, it is doing things (6.5).
- The output and what the app does with it — output shown to a user, or worse, acted upon — is a risk surface (6.5).
- The supply chain — the model itself, plus AI frameworks and dependencies, usually come from third parties (6.6).
Every page in Phase 6A maps onto a part of this anatomy. Keep this diagram in mind throughout.
Part 5: The Concept — “Everyone Ships AI, Nobody Secures It”
This page is the right moment to be precise about the conviction that motivated your whole curriculum — because in the AI era it is literally, demonstrably true, and understanding why tells you exactly where the opportunity is.
Why AI code is being shipped massively under-secured:
- AI features are easy to add and hard to secure. Adding an AI capability to an application is, today, remarkably easy — often a few API calls to a third-party model. But securing that capability requires understanding vulnerability classes (6.3–6.6) that most developers have never heard of. Ease of building, combined with difficulty of securing, guarantees a gap.
- The vulnerability classes are new and unfamiliar. Prompt injection, data poisoning, model extraction — these are not in most developers’ mental models. A developer who would never write a SQL-injectable query may build a wide-open prompt-injectable AI feature, simply because no one taught them the risk exists.
- The field is young and moving fast. AI security knowledge, tooling, and best practices are still maturing (the OWASP LLM Top 10, 6.2, is recent). Security has not caught up to adoption.
- The “it’s just an API call” illusion. Because using an AI model can be as simple as calling an API, teams underestimate that they have introduced a whole new attack surface (Part 3). They secure the API call like any API call and miss the model-layer risks entirely.
- Pressure to ship AI features fast. Competitive pressure to add AI capabilities is intense — and speed without security expertise produces insecure AI systems (the same dynamic as 5C.2’s cloud misconfiguration, 2.9’s misconfiguration — but in a brand-new domain).
The result is exactly your founding observation, sharpened: a vast and growing amount of AI-powered software is being deployed by people who can build AI features but cannot secure them — because the knowledge to secure them is scarce, new, and unevenly distributed. That gap is the opportunity. Someone who has the full security foundation (Phases 0–5) and genuinely understands AI-specific security (Phase 6) is filling a shortage that is real, growing, and acute. This phase is the bet you are making — and it is a sound one.
Part 6: The Concept — How to Think About Securing AI
This page closes by setting the mindset for the rest of Phase 6 — how to approach AI security correctly.
- AI security is additive, not separate. It is not a replacement for the curriculum’s foundation. An AI application is software — secure it with everything from Phases 2–5 (the app, the infrastructure, the cloud) — and then also secure the model layer with what Phase 6 teaches. A team that does brilliant prompt-injection defense but leaves the surrounding application full of classic vulnerabilities has not secured their AI system. Both layers, always.
- The same principles transfer — reapplied. You will see, again and again, that AI security is your existing principles in a new domain: prompt injection is the injection principle from 2.4 (untrusted input mixing with instructions); least privilege for AI tools (6.5) is least privilege (0.2, 4.3); the AI supply chain (6.6) is the supply-chain risk of 2.10; threat modeling an AI system (5B.4) is threat modeling. The vulnerability classes are new; the thinking is the security mindset you have built since Phase 1.
- Treat the model as untrusted, and its input and output as dangerous. A recurring theme of Phase 6A: do not blindly trust what goes into a model or what comes out of it. Model input can carry attacks; model output can be wrong, harmful, or attacker-controlled. The trust-boundary discipline of 1.2 and 4.3 applies directly to the model.
- AI security is genuinely hard and genuinely young. Be honest about this. Some AI vulnerabilities — prompt injection especially (6.3) — currently have no complete fix; they are managed, not solved. The field’s best practices are still forming. This is not a reason to avoid AI security — it is precisely why skilled people are needed in it.
- Stay current — more than anywhere else. AI and AI security are moving faster than any other area in this curriculum. The principles in Phase 6 are durable; specific techniques, tools, and even which attacks matter most will evolve. The continuous-learning habit (1.5, 3.7, 7.5) is non-negotiable here. Treat Phase 6 as a strong foundation and a starting point, not a final word.
🔑 The deep lesson: an AI application is software — so everything in Phases 0–5 still applies — plus it contains a model, which is a genuinely new attack surface with no equivalent in traditional software. The model is learned (not written), probabilistic, opaque, and — for LLMs — does not separate instructions from data; that nature makes its input, its training data, the model itself, and its output all attackable in new ways. AI features are being shipped massively faster than they are secured, because securing them needs knowledge that is new and scarce — which is exactly the opportunity. Secure AI by adding the model layer to your existing foundation, reapplying the security mindset you already have, treating the model and its input/output as untrusted, and committing to stay current in a fast-moving field.
📓 Key Terms
| Term | Plain meaning |
|---|---|
| AI application | Software that incorporates an AI/ML model as a component. |
| Machine learning (ML) model | A system that learns patterns from training data rather than being explicitly programmed. |
| Training data | The data a model learns from — which shapes (and can corrupt) its behavior. |
| Large Language Model (LLM) | A model trained on vast text that takes natural-language prompts and produces text. |
| Prompt | The input given to an LLM — often mixing instructions, user input, and retrieved data. |
| Black box (AI) | The property that a model’s reasoning is opaque, even to its creators. |
| The model as attack surface | The idea that the model itself — not just the surrounding app — is attackable. |
| Agentic / tools | An AI system’s capability to take actions — call APIs, run code, affect the world. |
🧪 Hands-On Lab
Phase 6A’s labs use AI applications you build, run, or are authorized to test — and deliberately vulnerable AI apps where they exist. Never attack AI systems you do not own or are not authorized to test (page 1.0 applies fully).
Task 1 — Map an AI application’s anatomy. Take an AI-powered application you know or can examine (one you have used, or a simple one you build). Draw its anatomy using Part 4’s diagram — application layer, prompt construction, model, any retrieval, any tools, output handling, supply chain.
Task 2 — Identify both attack surfaces. For that application, list (a) the traditional attack surface — the Phases 2–5 things — and (b) the new AI attack surface — model input, training data, the model, output, supply chain. Confirm for yourself that AI security is “both layers.”
Task 3 — See the instruction/data confusion. Using any LLM you can interact with, give it a clear instruction, then in the same input include text that tries to contradict or override that instruction. Observe how the model treats it all as one stream of text. You have just previewed why prompt injection (6.3) exists.
Task 4 — Build a tiny LLM app. Using a third-party model’s API (your developer skills from 0.6), build a minimal LLM-powered app — e.g. something that takes user input, builds a prompt, sends it to a model, and shows the output. You will use and harden this through Phase 6A.
Task 5 — Reason about “everyone ships, nobody secures.” In Notion, write your own analysis of Part 5: why is AI code being shipped under-secured? Where, specifically, is the opportunity for someone with your background? This is your founding thesis, examined.
Task 6 — Write the foundation note. In Notion, create an “AI Security” page — how AI differs from normal software, the model as attack surface, the anatomy of an AI app, and the mindset from Part 6. The foundation for all of Phase 6.
⚠️ Common Mistakes
- Thinking AI security replaces traditional security. An AI app is still software. Everything from Phases 2–5 still applies — AI security is added on top, not instead.
- Securing only the surrounding app, ignoring the model. The model is a real, new attack surface — input, training data, the model itself, output. Securing the API call is not securing the AI system.
- Assuming “it’s just an API call.” Calling a model API introduces a whole new attack surface. The ease of building with AI hides the difficulty of securing it.
- Treating model output as trustworthy. Model output can be wrong, harmful, or attacker-controlled. Treat input to and output from the model as untrusted (the trust-boundary discipline).
- Expecting AI vulnerabilities to behave like normal ones. The model is learned, probabilistic, opaque, and does not separate instructions from data — its weaknesses are genuinely different.
- Treating Phase 6 as a final word. AI security is young and fast-moving. The principles are durable; specifics will evolve. Commit to staying current.
✅ Recap & What’s Next
- An AI application is software — so all of Phases 0–5 applies — plus a model, a genuinely new attack surface with no traditional equivalent.
- The model is learned, probabilistic, opaque, and (for LLMs) does not separate instructions from data — making its input, training data, the model itself, and its output all attackable in new ways.
- AI is shipped far faster than it is secured because the knowledge is new and scarce — the opportunity; secure AI by adding the model layer to your foundation, reapplying the security mindset, treating the model as untrusted, and staying current.
Next (6.2): The field now has a shared map of LLM-specific risks — the OWASP Top 10 for LLM Applications. Page 6.2 is that map: the index to everything in 6.3–6.6, learned the way you learned the classic OWASP Top 10.
⁂ Back to all modules