Prompt injection
An attack where adversarial text hidden in tool output, fetched web content, or documents overrides an agent's real instructions, which this harness's guardrails are explicitly built to contain.
grounded in: Established agentic-AI security failure mode; principles.md guardrails ('no new external network calls', 'no secrets', bounded revert-on-failure autonomy) are exactly the containment against injected
Connected concepts
Agent security, Guardrails / bounded autonomy, Agent sandboxing, Adversarial anti-AI content, Agentic web
Explore it live in the knowledge graph →