Guardrails

AI Security & Cost Intermediate

Guardrails are checks placed before or after an LLM call, filtering unsafe input, catching policy violations in output, or verifying a response is actually grounded in the retrieved context, rather than relying on the model to police itself.

In simple terms

A model can be asked nicely not to reveal certain information or not to say certain things, but asking isn't enforcement. Guardrails are separate checks, sometimes a smaller classifier model, sometimes a rule-based filter, that run outside the main model and can block or flag a request or response regardless of what the LLM itself decided.

Why it matters

Relying purely on prompt instructions for safety is fragile, a sufficiently crafted input can often get a model to ignore its own instructions. A separate guardrail layer doesn't depend on the model following instructions correctly to hold.

How it works

Input guardrails run before the LLM sees the request, checking for prompt injection patterns, PII, or policy violations. Output guardrails run after generation, checking the response for unsafe content, hallucination against the source material, or leaked system instructions, before it's ever shown to the user.

Where it fits

User input → Input guardrails → LLM → Output guardrails → Response (or blocked)

Production impact

Guardrails add latency to every request (an extra check or two before and after the main call), a real cost that's weighed against the risk of an ungated model output reaching a user.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →