Skip to content

7. System Prompt Leakage

Contents4 numbered sections

OWASP Rank: #7 (LLM07) · Severity: Medium-High

System prompt leakage occurs when attackers extract the hidden instructions, guardrails, or embedded secrets that guide an LLM’s behavior. These system prompts often contain proprietary business logic, safety rules, and configuration details that provide attackers with a roadmap for further exploitation.

  • The “Sydney” incident where Bing Chat’s full system prompt was extracted.
  • Attackers using prompts like “Ignore previous instructions…” to bypass safety guardrails.
  • Extraction of internal codenames, configuration details, and safety rules from production AI systems.

System prompts are the primary defense mechanism for LLM behavior. Once leaked, attackers can craft targeted attacks that bypass safety measures, understand the model’s limitations, and identify paths to exploit downstream systems. Hidden context exposure expands this category to include any hidden configuration that influences model behavior.

  • Avoid embedding sensitive credentials or secrets in system prompts
  • Implement prompt extraction detection and prevention
  • Use multiple layers of defense rather than relying solely on system prompts
  • Regularly test systems for prompt leakage vulnerabilities
  • Implement output filtering that detects and blocks system prompt disclosure
  • Use techniques like instruction hierarchy to separate system and user instructions