GPT-Red broke 84% of agent scenarios: 7 controls to ship in 2026
OpenAI's internal red-teamer beat human red-teamers 84% to 13% and broke a live vending-machine agent. GPT-5.6 Sol now fails 0.05% of those attacks. Here is why a hardened model is not a defence plan, and the seven