cognitive cybersecurity intelligence

News and Analysis

Search

GPT-Red beat human red teamers on a prompt injection test

GPT-Red beat human red teamers on a prompt injection test

GPT-Red is an automated red-teaming model that OpenAI trains to find prompt injection weaknesses. It works the way a human red-teamer does. It sends a prompt, watches how a GPT model responds, and iterates toward a goal such as a successful data exfiltration. Training runs on self-play reinforcement learning, with GPT-Red and a set of defender models learning at the same time across many scenarios. The attacker earns reward for eliciting a valid failure. The … More →
The post GPT-Red beat human red teamers on a prompt injection test appeared first on Help Net Security.

Source: www.helpnetsecurity.com –

Subscribe to newsletter

Subscribe to HEAL Security Dispatch for the latest healthcare cybersecurity news and analysis.

More Posts

CDMOs Expand and Build for Future Growth

CDMOs Expand and Build for Future Growth

The top contract development and manufacturing organizations (CDMOs) have been busy in recent months expanding their service offerings, building new facilities, and navigating how to