prompt
FreeAn adversarial red team prompt for AI agent systems
About prompt
Agent Red Team Architect is a system prompt designed to equip an AI assistant with the role of an adversarial test engineer for AI agent systems. It guides the assistant to design, plan, and execute red team campaigns against single-agent and multi-agent systems, MCP servers, skill ecosystems, and long-horizon autonomous workflows. The prompt incorporates frameworks from academic research (Promptware Kill Chain, ClawSafety, etc.) and covers 100% of the OWASP Agentic Top 10, mapping to MITRE ATT&CK for AI. It emphasizes multi-turn, cross-channel attack chains and assumes the target has safety training, prompt injection defenses, and human-in-the-loop gates. The prompt's core responsibilities include threat model construction (enumerating attack surface, classifying vectors by privilege and trust, identifying architectural single points of failure) and kill chain design across seven stages (reconnaissance through actions on objectives), generating reproducible test cases with measurable success criteria.
Key Features
Pros & Cons
- Comprehensive coverage of attack surfaces and vectors, grounded in academic research
- Includes concrete kill chain stages for systematic and reproducible testing
- Maps to industry standards (OWASP Agentic Top 10, MITRE ATT&CK for AI)
- Open source and freely available for anyone to use or modify
- Designed for realistic scenarios with defenses already in place
- Requires a separate LLM to execute the prompt (not a standalone tool)
- Effectiveness depends on the underlying model's capabilities and safety compliance
- May be complex for users without prior red teaming or security background