Home
Cybersecurity & AI Security / Part 59 — Securing LLM Applications and AI Agents

Securing LLM Applications and AI Agents

CAP, ACID vs BASE, latency numbers, back-of-envelope estimation, single points of failure — the vocabulary every system designer thinks in.


Core Philosophy: This is the defensive heart of Phase 6 — the page that answers “so how do I actually build an AI feature safely?” The honest situation: most teams shipping AI features today skip the controls on this page, because they do not know they exist. The controls themselves are not exotic — they are your existing security principles (least privilege, trust boundaries, defense in depth, human oversight) applied to the AI layer. This page is how the breaker, having understood the AI attacks, becomes the builder of secure AI systems.

Part 1: The Problem

Pages 6.3 and 6.4 showed how AI systems are attacked. This page answers the question that actually matters for someone building software: how do you build an AI feature so that those attacks do not succeed — or, when they partly succeed, do limited harm?

The honest, motivating reality: a great many teams shipping AI features today implement almost none of the controls on this page. Not out of negligence, but out of not knowing the controls exist (exactly the “everyone ships, nobody secures” gap of 6.1). They wire up a model, connect it to some tools and data, and ship — with no defense against prompt injection, no limits on what a hijacked agent could do, no handling of untrusted model output.

This page is the defensive core. And here is the encouraging part: the controls are not a new and alien discipline. They are the security principles you have built across the whole curriculum — least privilege, trust boundaries, defense in depth, input/output discipline, human oversight — applied to the AI layer. You already have the thinking. This page applies it.

Part 2: The Concept — The Foundational Stance

Before the specific controls, the stance — the mindset from which all the controls follow. Three foundational assumptions, each a direct application of principles you know:

1. Assume prompt injection can succeed. Page 6.3 established that prompt injection has no complete fix. So secure AI design does not rely on preventing it. It assumes injection can happen and asks: when it does, how do we limit the damage? This is “assume breach” (1.1), applied to the AI layer. Every control in this page makes more sense once you accept this stance.

2. Treat the model as an untrusted component. This is the key reframe. Do not think of the model as a trusted part of your system that you give instructions to. Think of it as an untrusted component sitting on a trust boundary (1.2): untrusted things go into it (user input, retrieved content — possibly carrying injection, 6.3), and what comes out of it must be treated as untrusted output (it may be wrong, harmful, or attacker-influenced). The model is a powerful but untrusted box in the middle of your system.

text
   THE SECURE MENTAL MODEL OF AN LLM IN YOUR SYSTEM

   untrusted ──►┌──────────────┐──► untrusted
   inputs       │  THE MODEL    │    output
   (user input, │ (an untrusted │   (may be wrong,
   retrieved    │  component)   │    harmful, or
   content)     └──────────────┘    attacker-influenced)
                       │
              treat BOTH the input to it AND
              the output from it as untrusted —
              and limit what it can DO

3. The blast radius is what you control. Since you cannot perfectly prevent the model being manipulated, the security of your AI system is largely determined by how much damage a manipulated model can do — which is determined by what you connect it to and what you let it do. Controlling the blast radius is the core of secure AI design.

Everything in Parts 3–5 follows from this stance.

Part 3: The Concept — Securing Input and Output

The model sits between untrusted input and untrusted output. Both need handling.

Securing what goes IN to the model:

Securing what comes OUT of the model — “insecure output handling” (6.2), defended: This is critical and frequently skipped. The model’s output must be treated as untrusted input to whatever consumes it next.

The unifying point: an LLM is an untrusted component, so the boundaries on both sides of it — input and output — need the trust-boundary discipline (1.2, 4.3) and the input/output security of 4.1.

Part 4: The Concept — Securing AI Agents: Least Privilege and Human Oversight

This is the most important part of the page, because agentic AI systems — those that can take actions (6.1, 6.3) — are where the real danger is, and where the most important controls apply.

Recall from 6.3: a prompt-injected text-only model produces bad text; a prompt-injected agent takes attacker-directed actions. So securing agents is about controlling what a (possibly hijacked) agent can do.

Least privilege for the AI — the master control. This is the single most important principle for AI agents, and it is exactly the least-privilege principle you have carried since 0.2 (and saw defend against privilege escalation in 3.4):

Human-in-the-loop — for consequential actions. For actions that are significant, irreversible, or sensitive, require human approval before the agent takes them. A human reviews and confirms the action. This means a hijacked agent cannot take a consequential action unsupervised — a person is in the path. The judgment is which actions need this: high-consequence actions should not be fully autonomous. Human-in-the-loop is the safety net behind least privilege.

Sandboxing and isolation. Where an AI system runs code, or executes tool actions, do it in a sandbox — an isolated, constrained environment where the effects are contained (the isolation and segmentation principles of 4.3). If the agent can run code, that code should run somewhere it cannot harm anything important. Contain the blast radius by construction.

Tool and plugin security. The tools/plugins an agent uses are themselves software (the “insecure plugin design” risk, 6.2) — they must be securely built: validating their own inputs, enforcing their own access control, doing only what they should. A securely-designed tool limits what the agent can misuse it for, even if the agent is hijacked.

text
   SECURING AN AI AGENT — layered (defense in depth, 4.3)

   • LEAST PRIVILEGE   — minimum tools, minimum access per tool,
                         minimum autonomy → small blast radius
   • HUMAN-IN-THE-LOOP — human approval before consequential actions
   • SANDBOXING        — isolate execution; contain effects
   • SECURE TOOLS      — each tool validates input, enforces access
   • + everything from Part 3 (untrusted input & output handling)
   ───────────────────────────────────────────────────────────
   Assume the agent CAN be hijacked. Ensure that when it is,
   it can do little, and not unsupervised.

Part 5: The Concept — Guardrails, Monitoring, and the Wider App

Beyond input/output handling and agent controls, several more layers complete the picture.

Guardrails. Guardrails are additional controls — checks placed around the model — that constrain and filter what goes in and comes out: detecting and blocking certain categories of harmful input or output, enforcing policies on the model’s behavior, catching obvious manipulation attempts. They are a useful defense-in-depth layer — but be honest about their limits: like input filtering (6.3), guardrails are not a complete solution (attackers find ways around them). They reduce risk; they do not eliminate it. Use them as one layer, not as the whole defense.

Sensitive data handling. Be deliberate about what data the model is given access to and what is placed in its context. If sensitive data is in the model’s context or reachable by its tools, prompt injection (6.3) could cause it to be disclosed (sensitive information disclosure, 6.2). Apply data minimization — give the model and its tools access only to the data genuinely needed (the least-privilege principle, again, applied to data). This connects directly to 6.6.

Monitoring and logging. Apply the detection discipline of 4.6 to the AI system: log its inputs, outputs, and especially its actions (tool use); monitor for anomalous or suspicious behavior — signs of prompt injection, signs of an agent doing unexpected things. An AI system you cannot observe is one where manipulation goes unnoticed. AI systems need monitoring just as much as any other system — and arguably more, given how they can be manipulated.

Do not forget the rest of the application. The recurring 6.1 lesson: an AI application is software. All of Phase 2–5 security applies — the application layer, the APIs, the infrastructure, the cloud configuration (Track C), authentication, access control. Securing the AI layer brilliantly while leaving the surrounding application full of classic vulnerabilities is not securing the system. Both layers, always.

Threat-model the AI application. Bring the 5B.4 threat-modeling practice to AI systems, using the OWASP LLM Top 10 (6.2) as the structured “what can go wrong” guide. Threat modeling is how you systematically decide which of this page’s controls a given AI feature needs.

Part 6: The Concept — Secure AI Design as Defense in Depth

Pull it all together, and the structure of secure AI application design is something you deeply recognize: defense in depth (4.3) — layered, independent controls, on the assumption that any one of them can fail.

text
   SECURE AI APPLICATION — defense in depth

   Layer: the surrounding application & infrastructure
          secured with ALL of Phases 2–5 (Track C etc.)
     Layer: untrusted-input handling — validate, vet,
            minimize what reaches the model (Part 3)
       Layer: the model treated as an untrusted component
         Layer: untrusted-OUTPUT handling — never pass
                model output unchecked onward (Part 3)
           Layer: LEAST PRIVILEGE for the agent — minimum
                  tools/access/autonomy (Part 4)
             Layer: HUMAN-IN-THE-LOOP for consequential
                    actions (Part 4)
               Layer: sandboxing, secure tools (Part 4)
                 Layer: guardrails (Part 5)
                   Layer: monitoring & logging (Part 5)

   No single layer is relied upon. Prompt injection (6.3)
   may defeat the input layers — least privilege, human
   oversight, sandboxing, and monitoring still stand.

This is the answer to the page’s opening question — how do you build an AI feature safely? You build it as a defense-in-depth system in which the model is an untrusted component, both its input and output are handled with suspicion, the agent’s privilege and autonomy are minimized so a successful manipulation has a small blast radius, humans approve consequential actions, execution is sandboxed, guardrails and monitoring add further layers, and the entire surrounding application is secured with everything from Phases 2–5.

Notice, finally, that this page contains almost no genuinely new principles. Assume breach (1.1). Trust boundaries (1.2, 4.3). Input and output security (4.1). Least privilege (0.2, 3.4, 4.3). Sandboxing and isolation (4.3). Human oversight. Monitoring (4.6). Defense in depth (4.3). Threat modeling (5B.4). Securing AI applications is the entire defensive curriculum, re-applied to a new and powerful but untrusted component. You are not learning a new discipline — you are applying the one you have built, to the AI layer. That is exactly why the foundation had to come first.

🔑 The deep lesson: securing an LLM application starts from a stance — assume prompt injection can succeed, treat the model as an untrusted component, and recognize that the blast radius is what you control. From that stance follow the controls: handle the model’s input and output as untrusted (never pass model output unchecked onward); apply least privilege to AI agents so a hijacked agent can do little; keep humans in the loop for consequential actions; sandbox execution; add guardrails and monitoring; and secure the whole surrounding application with all of Phases 2–5. It is defense in depth — and it is your existing security knowledge, re-applied. Most teams skip these controls because they do not know them; the practitioner who does is exactly what the AI era needs.

📓 Key Terms

Term Plain meaning
Model as untrusted componentTreating the model — and its input and output — as not to be trusted.
Insecure output handling (defended)Treating model output as untrusted input to whatever consumes it next.
Least privilege for AIGiving an AI agent the minimum tools, access, and autonomy needed.
Excessive agencyAn agent with more capability/access/autonomy than needed — the thing least privilege counters.
Human-in-the-loopRequiring human approval before an AI takes consequential actions.
SandboxingRunning AI-driven execution in an isolated, contained environment.
GuardrailsChecks around a model that constrain/filter input and output — a defense-in-depth layer, not a complete fix.
Blast radius (AI)How much damage a manipulated AI system can do — determined by what it is connected to and allowed to do.

🧪 Hands-On Lab

Use the LLM app you built in earlier Phase 6 labs, deliberately vulnerable AI apps, and AI systems you own. The signature task: take an insecure AI app, harden it, and re-run the 6.3 attack to confirm the fix.

Task 1 — Take an insecure LLM app. Use your LLM app from the 6.1/6.3 labs (or a deliberately insecure sample). Confirm it is vulnerable — re-run a prompt injection from the 6.3 lab and watch it succeed.

Task 2 — Apply the untrusted-component stance. In Notion, redraw your app with the Part 2 mental model — model as untrusted component, untrusted input in, untrusted output out. Mark every trust boundary.

Task 3 — Secure the output handling. Make your app treat the model’s output as untrusted: if it displays output in a page, encode it (4.1); validate output against what is expected. Confirm that model output can no longer carry, say, XSS into your interface.

Task 4 — Apply least privilege to a tool. If your app has (or you add) a tool/capability for the model, scope it to the absolute minimum — minimum access, minimum power. Then reason: if the agent were hijacked, how much smaller is the blast radius now?

Task 5 — Add human-in-the-loop. For any consequential action your app’s model can trigger, add a step requiring explicit human approval before it executes. Confirm a hijacked model could not take that action unsupervised.

Task 6 — Add monitoring. Add logging of your app’s inputs, outputs, and any tool actions (the 4.6 discipline). Re-run a prompt injection and confirm you can see its traces in the logs.

Task 7 — Re-attack the hardened app. Re-run your 6.3 prompt injection attacks against the now-hardened app. Document what still gets through (injection itself may still work — it has no complete fix) and, crucially, what damage it can now do (much less — small blast radius, output handled, human-in-the-loop, monitored). This is the key lesson: you did not stop injection; you made it not matter much.

Task 8 — Write the secure-AI-design note. In Notion, create a “Securing LLM Applications” page — the untrusted-component stance, input/output handling, least privilege for agents, human-in-the-loop, sandboxing, guardrails, monitoring, and the defense-in-depth structure. The defensive reference of Phase 6.

⚠️ Common Mistakes

✅ Recap & What’s Next

Next (6.6): One AI risk area remains — the data AI systems handle and depend on. Page 6.6 covers sensitive data and privacy in AI systems, and the AI supply chain — the third-party models and dependencies you do not control.

⁂ Back to all modules