AI Engineering · 6 MIN
What guardrails does an LLM feature need in production?
Input checks, output validation, permission limits, and fallbacks. A layered guardrail design for LLM features, mapped to the OWASP LLM Top 10.
An LLM feature in production needs four layers of guardrails: checks on what goes in, validation of what comes out, hard limits on what the model is allowed to do, and a fallback path for when any of those fail. A system prompt that says "be careful" is not a guardrail. A guardrail is code that runs whether or not the model cooperates. This post lays out the four layers, ties them to the OWASP list of LLM risks, and shows how to test them.
- Guardrails are enforced in code around the model, not requested in the prompt.
- The four layers are input checks, output validation, permission limits (least privilege), and fallbacks with human escalation.
- The OWASP Top 10 for LLM Applications names prompt injection, improper output handling, and excessive agency among its top risks.
- Treat model output as untrusted input to the rest of your system, the same way you treat user input.
- Guardrails need their own tests. A red-team set of attack prompts belongs in your eval suite.
- Nactore designs these layers into each AI feature and tests them as part of the eval set before launch.
What is a guardrail, and what is it not?
A guardrail is a control that constrains behavior independently of the model's willingness. If the model is tricked, confused, or simply wrong, the guardrail still holds.
A prompt instruction is guidance. It is useful, and attackers and edge cases routinely get around it. Anything that must hold, such as "never send an email to an address outside the customer's domain" or "never output more than five records", has to be enforced by the application.
Which risks should you design against first?
The OWASP Top 10 for LLM Applications is the most widely used checklist. The 2025 edition lists prompt injection, sensitive information disclosure, supply chain, data and model poisoning, improper output handling, excessive agency, system prompt leakage, vector and embedding weaknesses, misinformation, and unbounded consumption. For a first release, these map to controls you can build.
| Risk (OWASP name) | What goes wrong | Guardrail |
|---|---|---|
| Prompt injection | Text in a document or message overrides your instructions | Separate trusted and untrusted content, limit tool permissions, require confirmation for risky actions |
| Sensitive information disclosure | The model reveals data the user should not see | Filter retrieval by the user's permissions before the model sees anything |
| Improper output handling | Output is passed to a shell, SQL, or HTML unchecked | Validate against a schema, escape on use, never execute model output directly |
| Excessive agency | The agent can do more than the task needs | Least-privilege tools, spend and rate caps, approval gates |
| Misinformation | Confident wrong answers | Ground in retrieved sources, require citations, allow "I don't know" |
| Unbounded consumption | One user or loop runs up cost | Token limits, per-user quotas, loop and step caps |
Layer 1: How do you check inputs?
Input checks run before the model call.
- Size and shape limits. Reject or truncate inputs beyond a sane length, and enforce expected types.
- Authorization first. Decide what the user may see before retrieval, so the model never receives data it should not reveal.
- Content separation. Wrap untrusted text (emails, web pages, uploaded files) in clear delimiters and tell the model it is data, not instructions. This reduces injection risk and does not remove it.
- Moderation and PII handling. Run a moderation or classifier step where the product needs one, and redact personal data that the task does not require.
Layer 2: How do you validate outputs?
Treat the model's response as untrusted. Validate it before anything downstream acts on it.
- Enforce structure. Use schema-constrained generation where available, and validate in your own code anyway. See structured output and JSON reliability.
- Check business rules. Amounts within range, IDs that exist, dates in the future, recipients on an allow list.
- Check grounding. For answers built from documents, verify that each claim has a supporting passage. Anthropic's guidance on reducing hallucinations recommends allowing the model to say it does not know, grounding in direct quotes, and retracting claims that lack support.
- Scan for leaks. Look for system prompt text, secrets, or other users' data in the output.
Layer 3: How do you limit what the model can do?
This layer matters most for agents. Give each tool the narrowest permission that works.
- Read-only by default. Add write access tool by tool, with reasons.
- Scoped credentials. The agent's API key should not be an admin key.
- Confirmation for irreversible actions. Payments, deletions, and external messages go through an approval step. See human-in-the-loop automation.
- Budgets. Cap steps per run, tokens per request, and spend per user per day.
Ask one question of every tool you give an agent: if an attacker fully controlled the model's next message, what is the worst this tool could do? Shrink the tool until the answer is acceptable.
Layer 4: What happens when a guardrail trips?
A guardrail with no fallback just produces errors. Decide the behavior up front.
| Situation | Fallback |
|---|---|
| Output fails schema validation | Retry once with the error message, then use a safe default |
| Grounding check fails | Return "I could not find that in the documents" with sources searched |
| Risky action requested | Queue for human approval |
| Model refused or timed out | Use a cheaper backup path or a templated response |
| Repeated failures | Alert the on-call owner and log the case for the eval set |
Log every tripped guardrail with enough context to reproduce it. These events are the best source of new test cases. See LLM observability: what to log.
How do you test guardrails?
Add an adversarial slice to your eval set: injection strings hidden in documents, requests for other users' data, malformed outputs, and attempts to extract the system prompt. Run it on every release. A guardrail that is not tested will quietly stop working after the next refactor. The general method is in why AI features need evals before production.
Frequently asked questions
Can a better system prompt replace guardrails?
No. A prompt can reduce bad behavior, but it cannot guarantee it. Anything that must hold needs enforcement in code.
Does using a provider's built-in safety features cover us?
They help with general harms. They do not know your business rules, your data permissions, or which actions are allowed in your product, so you still need application-level controls.
Will guardrails make the feature slower?
Code checks add negligible latency. LLM-based checks add a model call, so use them selectively and keep them off the critical path where you can.
Are guardrails only for customer-facing chatbots?
No. Internal agents often have broader permissions than chatbots, which makes least-privilege limits and approval gates more important.
Want this built for your team? Book a free 30-minute call.
Want to apply this to your business?
Book a free 30-minute call. We will tell you what we would do first.