Developer

OpenAI Rolls Out Lockdown Mode to Prevent Data Theft from Prompt Injection

OpenAI has launched Lockdown Mode, a security feature that blocks data exfiltration during prompt injection attacks. The mode limits outbound network requests and is now rolling out to Free, Go, Plus, Pro, and self-serve ChatGPT Business accounts. It does not prevent injections from appearing in content but cuts off the final stage of data theft.

Neura News

Neura News

Neura Market Editorial

June 6, 20263 min read
OpenAI Rolls Out Lockdown Mode to Prevent Data Theft from Prompt Injection

OpenAI has activated a new security feature called Lockdown Mode, designed to prevent the final stage of data theft from prompt injection attacks. The feature first teased by the company in February is now rolling out to eligible personal accounts, including Free, Go, Plus, and Pro tiers, as well as self-serve ChatGPT Business accounts.

Lockdown Mode works by limiting outbound network requests that could transfer sensitive data to an attacker. It directly targets the exfiltration vector, the channel through which stolen data leaves a system. According to a blog post by Simon Willison, Lockdown Mode does not stop prompt injections from appearing in content that ChatGPT processes. For example, an injection could still appear in cached web content or an uploaded file and might affect the behavior or accuracy of a response.

What Lockdown Mode Does

Lockdown Mode is a deterministic mechanism. It does not rely on AI systems to evaluate threats, which means it cannot be subverted by sufficiently devious attacks that might trick an AI-based defense. This is a crucial design choice because any security measure that itself uses an LLM could be manipulated. By cutting off the exfiltration path directly, OpenAI provides a layer of protection that is harder to bypass.

The feature is now live and rolling out across the supported account types. OpenAI describes Lockdown Mode as a way to combat the "Lethal Trifecta," a term for the combination of three elements that enable data theft in LLM systems: access to private data, exposure to untrusted content, and a way to steal and transmit data back to an attacker.

The Lethal Trifecta and How Lockdown Mode Breaks It

The Lethal Trifecta occurs when an LLM system has access to all three of these components. To stop an attack, a system must remove at least one leg of the triad. According to the analysis shared by Willison, the easiest leg to restrict without making LLM systems far less useful is the exfiltration vector. Lockdown Mode directly addresses that leg.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

By limiting outbound network requests, the mode prevents the final step of a prompt injection attack. Even if an attacker successfully places a prompt injection in content the model processes, and even if the model is tricked into outputting sensitive data, that data cannot leave the system. The attacker cannot receive the stolen information.

Implications for ChatGPT Security

The existence of Lockdown Mode implies that ChatGPT, in its default settings, does not provide robust protection against sufficiently determined data exfiltration attacks. OpenAI is essentially acknowledging that without this additional measure, the system could be vulnerable to cases where a user or process feeds untrusted content to the model while it has access to private data.

This move is significant for developers and businesses that rely on ChatGPT for handling sensitive information. Users who handle confidential data should consider enabling Lockdown Mode to reduce the risk of data theft through prompt injection. The feature is available without extra cost to eligible accounts.

Lockdown Mode does not affect the model's ability to process content or generate responses. It only restricts outbound network requests. This makes it a relatively low-impact security enhancement that could be applied broadly without degrading functionality.

The rollout is ongoing. OpenAI has not specified a timeline for full availability, but the feature is now accessible to the account types listed.

Related on Neura Market

More from Neura News

Industry

Cognition Acquires Poke to Give Devin Coding Agent a Personality

Cognition, the startup behind AI coding assistant Devin, has acquired Poke, an AI assistant known for its friendly, conversational style. The deal, valued in the low nine figures, aims to bring Poke's personality-driven interaction model to Devin, making the coding agent feel more like a colleague than a tool. Poke will also benefit from Cognition's models and infrastructure to become faster and more reliable.

Jul 24·3 min read
AI Models

Anthropic expands Claude voice mode to Opus and Sonnet models

Anthropic has expanded Claude's voice mode to run on its most powerful models, Opus and Sonnet, across mobile, desktop, and web platforms. Users can now switch between models mid-conversation, use voice commands in eleven languages, and connect to external tools like Gmail, Google Calendar, or Slack to compose and send emails by voice. The update positions Claude as a unique option for tool integration in voice AI, though competitors like OpenAI and Google offer more natural speech processing.

Jul 24·2 min read
Industry

Airbus Scores Cloud Providers on Extraterritorial Law Protection

Airbus has chosen French cloud provider Scaleway as its sovereign cloud partner, making protection against non-European extraterritorial laws a formal, scored criterion in its tender. The aerospace giant assessed providers on technical capabilities, operational excellence, and legal safeguards, including European jurisdiction and data protection. This move signals a shift in how major European industrial buyers evaluate cloud vendors, prioritizing legal jurisdiction alongside technical performance.

Jul 24·6 min read