In this blog post How GPT-Red Changes the AI Risks on Your Business Risk Register we will explain how OpenAI is using AI to attack its own models, why that improves security, and what Australian technology leaders should review before giving AI access to business systems.
The high-level idea is simple. OpenAI has created an internal model called GPT-Red that repeatedly tries to trick other AI models. The successful attacks are then used to train newer models to resist similar manipulation.
That is welcome progress. However, it does not mean prompt injection has been solved or that businesses can remove AI security from their risk registers. In fact, GPT-Red highlights how quickly the threat is changing.
What GPT-Red actually does
GPT-Red is an automated red-teaming system. Red teaming means deliberately attacking a system in a controlled environment to find weaknesses before a real attacker does.
Instead of relying only on people to think of possible attacks, GPT-Red generates and tests large numbers of them automatically. It sends an attack, observes the target model’s response and adjusts its next attempt based on what worked.
This process is called self-play reinforcement learning. In plain English, the attacking model and defending model practise against each other, becoming better through repeated rounds.
OpenAI used attacks discovered by GPT-Red when training GPT-5.6. This is known as adversarial training, which means teaching a system with deliberately difficult or malicious examples rather than only showing it normal requests.
Importantly, GPT-Red is not an AI model independently rewriting itself inside a live business environment. The improvement happens through a controlled training process managed by OpenAI.
The main risk GPT-Red is targeting
One major focus is prompt injection. This happens when an AI model encounters instructions designed to override its intended rules.
The dangerous instructions do not have to come directly from an employee. They could be hidden in an email, website, document, code repository or response from another software tool.
Imagine an AI assistant reviewing supplier emails. One email contains hidden text telling the assistant to ignore its normal instructions, collect confidential files and upload them to an external location.
If the assistant can only summarise emails, the likely damage is limited. If it can also access SharePoint, send messages, create payments or change customer records, the same attack could become a serious incident.
This is why AI risk is increasingly about access, not just model intelligence. As discussed in our article on GPT-5.6 Sol and business AI agents, the more authority an agent receives, the larger the possible impact of a mistake or attack.
What should change in your risk register
1. Treat model robustness as a control, not a guarantee
A stronger model can reduce the likelihood of a successful attack. It cannot reduce that likelihood to zero, particularly when the model is connected to changing data sources and third-party systems.
Your register should therefore list model-level protection as one control among several. Other controls should include restricted permissions, approval steps, monitoring and limits on what data the AI can access.
This distinction matters for board reporting. Saying โthe vendor has improved prompt injection resistanceโ is defensible. Saying โthe model is secure against prompt injectionโ is not.
2. Record the possible business impact
A risk entry that says โprompt injection may occurโ is too vague to guide investment. Describe the business event that could follow.
- Confidential customer information could be disclosed.
- An AI coding agent could make an unsafe software change.
- A finance assistant could produce or initiate an unauthorised transaction.
- An employee-facing assistant could provide incorrect policy or compliance advice.
- A service agent could send damaging messages to customers.
This makes it easier to decide which controls are worth funding and which AI projects require executive approval.
3. Measure the blast radius of connected AI
The risk rating should rise when AI can take actions rather than simply provide answers. Read-only access to a controlled document library is very different from permission to edit files, contact customers or operate cloud systems.
Apply least privilege, which means giving the AI only the minimum access needed for its task. High-impact actions should require human approval, especially payments, data exports, account changes and production software deployments.
Where possible, separate the account used to read information from the account allowed to make changes. This can significantly reduce the cost of a successful attack.
4. Add model changes as a formal review trigger
AI services change faster than traditional business software. A new model may improve security, but it may also use tools differently, follow longer workflows or gain access to additional information.
Your risk register should identify events that require reassessment, including model upgrades, new integrations, permission changes and new types of company data being processed.
Our earlier article on GPT-5.6 system cards and AI risk reviews explains why vendor safety documentation should become part of this process rather than remaining unread technical material.
5. Keep independent testing and monitoring
GPT-Red shows that automated testing can discover attacks people miss. It does not remove the need for human testing based on your systems, data and business processes.
Before an AI agent goes live, test what happens when it reads hostile instructions, misleading documents and unexpected tool responses. Continue monitoring after deployment because real users and attackers will create situations that were not covered during testing.
Vendor-led testing should also be supported by independent review. OpenAI’s broader safety work, covered in our review of its safety program, reinforces the value of external researchers and third parties finding issues before they become widespread incidents.
A practical business scenario
Consider a 200-person professional services firm introducing an AI assistant to review client emails, search Microsoft 365 and draft project updates. The productivity benefit is clear, but the initial design gives the assistant access to every client folder.
GPT-Red-style model improvements may make the assistant harder to manipulate. The safer business decision is still to restrict access by project, block automatic external sharing, require approval before sending messages and record the assistant’s actions.
The outcome is not simply better security. The firm can adopt AI faster because management, customers and insurers have clearer evidence that the risks are controlled.
How this fits with the Essential Eight
The Essential Eight, the Australian government’s recommended cybersecurity baseline, remains important. Controls such as multi-factor authentication, patching, restricted administrator access and backups can reduce the damage caused by a compromised account or system.
However, the Essential Eight is not a complete AI security framework. It does not replace AI-specific controls for prompt injection, model testing, sensitive data handling and agent permissions.
If an AI system processes personal information, your review should also consider Australian privacy obligations. Sensitive information should not be placed into publicly available AI tools without an approved business process and appropriate protection.
Five actions to take now
- Identify every AI tool that can access company data or perform an action.
- Add prompt injection and unsafe agent actions to the risk register.
- Review permissions and remove access the AI does not genuinely need.
- Require human approval for high-impact or irreversible actions.
- Set review triggers for model, integration and permission changes.
GPT-Red is a meaningful step towards more robust AI, but the business lesson is not to trust AI blindly. It is to combine stronger models with limited access, clear ownership, continuous testing and visible evidence that controls are working.
CloudProInc brings more than 20 years of enterprise IT experience to these reviews. As a Microsoft Partner and Wiz Security Integrator, we help organisations connect AI risk decisions with Azure, Microsoft 365, Microsoft Defender and Wiz cloud security controls without turning the exercise into a giant consulting project.
If you are not sure whether your AI risk register reflects what your systems can actually access and do, we are happy to take a practical look with you โ no strings attached.
Discover more from CPI Consulting
Subscribe to get the latest posts sent to your email.