AI Security

AI Security & Cost Intermediate

AI security covers the set of risks specific to LLM-based systems, prompt injection, data exfiltration through a tool call, an agent taking an action it shouldn't, that don't map cleanly onto traditional application security threats.

In simple terms

Normal security assumes attackers manipulate code paths and inputs directly. AI security has to account for a new attack surface: manipulating the model's behavior itself through crafted text, since the model reads instructions and untrusted data through the exact same channel.

Why it matters

An LLM can't reliably tell the difference between an instruction from its system prompt and an instruction hidden inside a document it's been asked to summarize, which is a fundamentally different problem than anything traditional input validation was built to catch.

How it works

Defenses layer at several points: input filtering to catch obvious injection attempts, strict permission scoping so a tool call can't do more than its narrow job even if the model is tricked, output filtering to catch leaked instructions or unsafe content, and sandboxing so an agent's actions are contained even in a worst case.

Where it fits

User input / retrieved content → Input guardrails → LLM → Tool call (permission-scoped) → Output guardrails → Response

Production impact

A successful prompt injection through, say, a retrieved web page or an email an agent is asked to process can lead directly to data exfiltration or an unintended action, this is treated as a real production security issue, not a curiosity.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →