raghu@dark-factory :~/kb/jailbreak $ cat

Jailbreaking

Crafting adversarial prompts that bypass a model's safety alignment or system-prompt constraints to elicit disallowed behavior, distinct from injecting instructions through untrusted data.

grounded in: doctrine 'Guardrails' harness layer (bounded autonomy) + existing constitutional-ai/guardrails defenses that jailbreaks target; a foundational, lasting LLM-security concept absent from the graph

Connected concepts

Prompt injection, Adversarial attacks on AI, Guardrails / bounded autonomy, Constitutional AI, System-prompt extraction

Explore it live in the knowledge graph →