Concept
Guardrail
A guardrail is a runtime filter that sits between an AI system and the world, checking inputs or outputs against a safety policy and blocking, flagging, or rewriting anything that violates it. Think of it like input validation on a web form, except the "invalid" set is fuzzy…
The rest of “Guardrail” is a premium feature: every concept in the library gets a precise, practitioner-focused write-up like this one, cross-linked straight from the paper summaries that use it.
Log in to unlock→