- What: OpenAI details GPT-Red, an AI system that finds vulnerabilities in its own models
- Impact: Highlights ongoing efforts to improve AI security
AI/ML OpenAI details GPT-Red, an AI system designed to find vulnerabilities in its own models July 16, 2026 Share By SC Staff (Credit: Rizq – stock.adobe.com) OpenAI has detailed GPT-Red, an internal artificial intelligence system it developed to identify and exploit vulnerabilities within its own AI models, as reported by Silicon Angle. GPT-Red functions as an automated red team, employing self-play reinforcement learning to continuously attack target models. It iteratively refines its prompts to discover successful exploits, outperforming human red-teamers by succeeding in 84% of scenarios compared to their 13%. This system has significantly reduced prompt injection failures, achieving a sixth of the rate seen in previous models, according to OpenAI. GPT-Red has also demonstrated the ability to compromise autonomous agents, such as a vending machine agent and command-line coding agents, highlighting potential risks as AI systems gain more autonomy. While effective, GPT-Red has limitations in multi-turn conversational attacks and image-based prompt injection, areas where human testers will continue to provide coverage. This development follows OpenAI's release of GPT-5.6 and underscores the ongoing challenges in AI security, particularly concerning prompt injection vulnerabilities. Source: Silicon Angle An In-Depth Guide to AI Get essential knowledge and practical strategies to use AI to better your security program. Learn More SC Staff Related AI/ML Reken launches AI security platform for on-device communication analysis SC Staff July 16, 2026 Reken's Private Core platform, powering its Northstar application, utilizes three key components: Model Zero for threat analysis, Trust Sensor for verifying human activity and outgoing communications, and a Reken API for integration. Application security CISOs no longer get to choose because AI is redefining the SOC Tom Findling July 16, 2026 AI in the SOC demands trust built through oversight, transparency, and validation. AI/ML Russian hacker uses Google’s AI tool to operate botnet SC Staff July 16, 2026 Between May 19 and April 21, the threat actor engaged in over 200 sessions with the AI tool to deploy and manage an infrastructure controlling eight systems within a dental clinic, gaining access to the OpenDental database. Get daily email updates SC Media's daily must-read of the most current and pressing daily news Business Email By clicking the Subscribe button below, you agree to SC Media Terms of Use and Privacy Policy . Subscribe You can skip this ad in 5 seconds