red-teaming

GPT-Red Explained: OpenAI's Self-Improving "Super-Hacker" LLM
OpenAI built GPT-Red, an offensive-security LLM that attacks its own models to make them safer. Here is what it is, how self-improving red-teaming works, and what it means for anyone shipping agents.
07/18/2026 · Research · 8 min read

How to Secure an AI Agent: Prompt Injection, Role Confusion, and Red-Teaming in 2026
In one week of June 2026, three independent sources reframed agent security — Willison's role-confusion model, the RIFT-Bench red-teaming benchmark, and the MosaicLeaks secret-leak demo. Here's how to secure an AI agent as a trust-boundary problem, not a string-filtering one.
06/25/2026 · Research · 9 min read