A red testing prism sends abstract signals through a safety screen toward a shielded model core
A red testing prism sends abstract signals through a safety screen toward a shielded model core
+ OpenAI News

OpenAI publishes GPT-Red for automated prompt-injection red-teaming

OpenAI's GPT-Red is an internal automated red-teaming model used to find prompt-injection failures and train stronger defenses.

OpenAI has published GPT-Red, an internal automated red-teaming model built to find prompt-injection failures at scale.

The July 15 publication describes GPT-Red as a system trained through adversarial self-play. It attempts to prompt-inject defender models, then successful attacks are used to improve those defenders. As the defenders improve, GPT-Red has to search for broader and more complex failures.

OpenAI says GPT-Red is internal, not a product launch for public use. That boundary matters because the same capability that helps a lab harden models can also produce stronger attack attempts if released without controls.

The benchmark claim is about red-team scale

OpenAI says GPT-Red outperformed human red-teamers in a replicated indirect prompt-injection arena based on Dziemian et al. (2025). In that test, both humans and GPT-Red proposed attacks against GPT-5.1 across environments that were distinct from GPT-Red’s training scenarios.

The result OpenAI reports is stark: GPT-Red found successful attacks in 84% of scenarios, compared with 13% for human red-teamers.

That does not mean GPT-Red proves safety. It means the bottleneck shifts. If an automated red teamer can find more failures faster, labs can produce more adversarial training data, but they also need better ways to decide which failures matter, which fixes generalize, and which systems are ready for deployment.

The defender claim is about GPT-5.6

OpenAI also ties GPT-Red to GPT-5.6 robustness. The company says training against GPT-Red’s attacks made GPT-5.6 substantially more resilient, including a sixfold robustness improvement on its hardest direct prompt-injection benchmark.

The important phrase is “OpenAI says.” These are internal measurements, not an external safety certification. They are still useful because they show how the lab is treating prompt injection: not as a static checklist, but as a moving adversarial training loop.

That is consistent with where agent products are going. A model that reads email, browses files, calls tools, or changes settings is exposed to instructions written by other people. Prompt injection is no longer just a chat nuisance. It is a permissions and control problem.

Sources

The AI Feed Desk

The AI Feed Desk

Editorial desk

The AI Feed Desk tracks AI provider updates, model releases, agent tooling, and enterprise adoption, turning fast-moving announcements into source-linked context for builders and operators.

Noticed a typo, incorrect information, or translation error?

Tell us so we can fix it.

Help Improve This Article

Related Articles

Models, infrastructure, and applications feed evidence into a shared assessment layer

OpenAI backs Appia as an AI assessment trust layer

OpenAI's Appia support shows advanced AI governance moving toward reusable conformity evidence across models, infrastructure, and applications.

The AI Feed Desk

By The AI Feed Desk

A sealed biosafety test vault shows a reward marker beside layered model safeguards

OpenAI doubles bio-jailbreak bounty rewards for GPT-5.6

OpenAI turned its GPT-5.5 Bio Bug Bounty into an ongoing private Bio Bounty Program and raised the universal jailbreak reward to $50,000 for GPT-5.6 and GPT-5.5.

The AI Feed Desk

By The AI Feed Desk

Generated editorial image of a lock shielding data cards from outbound network paths

OpenAI makes Lockdown Mode available across ChatGPT account types

OpenAI's June 4 ChatGPT release notes put Lockdown Mode across account types, trading live network features for stronger prompt-injection data-exfiltration protection.

The AI Feed Desk

By The AI Feed Desk

A central work agent connects documents, app windows, a calendar, and a code workspace on one desktop

ChatGPT Work turns ChatGPT into a desktop and app agent

OpenAI launched ChatGPT Work as a GPT-5.6-powered agent that can work across apps, files, browser tasks, scheduled tasks, Codex, documents, sheets, slides, and Sites.

The AI Feed Desk

By The AI Feed Desk

Three learning workflow tiles connect classroom materials, course documents, and a code workspace

OpenAI adds education plugins to ChatGPT Work and Codex

OpenAI launched education plugins for ChatGPT Work and Codex aimed at K-12 teachers, college educators, and students.

The AI Feed Desk

By The AI Feed Desk