llm-safety

GPT-Red Explained: OpenAI's Self-Improving "Super-Hacker" LLM
OpenAI built GPT-Red, an offensive-security LLM that attacks its own models to make them safer. Here is what it is, how self-improving red-teaming works, and what it means for anyone shipping agents.
07/18/2026 · Research · 8 min read

Prompt Injection Defense: A Builder's Guide to Securing AI Agents
Prompt injection defense is now a shipping requirement for anyone connecting an LLM to tools. Here is what the attack really is, why it can't be patched away, and the defense-in-depth layers to ship before your agent touches anything that matters.
06/30/2026 · AI Tutorials · 8 min read