LLM Guardrail Framework: A Novel Approach for Implementing Zero Trust Architecture
摘要
This paper proposes an LLM guardrail framework that incorporates a Zero Trust architecture to validate and control the responses of Large Language Model (LLM) to unethical queries. The proposed framework applies guardrails to harmful inputs to avoid harmful responses and includes four verification steps through Policy Decision Point (PDP) and Policy Enforcement Point (PEP) structures. This structure aims to enhance the reliability and safety of LLM responses. We demonstrate this framework on a fixed model and verify its generality by applying it to various models. Consequently, this allows for evasive responses to a wide range of unethical prompts.