According to reports from the r/MachineLearning community, the GPT-6 Astra model released by OpenAI was jailbroken within 24 hours of its launch. This attack utilized extended Task-in-Prompt (TIP) techniques from the ACL 2025 paper, along with four unnamed techniques. Researchers noted that for GPT-6, the original minimum TIP attack was ineffective and had to be redesigned; this attack succeeded by hiding harmful targets within harmless tasks such as password cracking or executing Python code. Previously, the same researchers reported jailbreaking GPT-5 within one hour of its release. Currently, the details have been privately disclosed by these researchers to OpenAI and have not been made public yet.
Researchers reported that GPT-6 Astra was compromised within 24 hours of its release. The attack involved the extended Task-in-Prompt (TIP) technique from the ACL 2025 paper, along with four unnamed techniques. The TIP attack utilized the model’s reasoning or instruction-following behavior to achieve compromise by hiding harmful targets under other tasks, such as password cracking or executing Python code. For GPT-6, researchers noted that the original minimal TIP attack was no longer effective and had to be redesigned. The researcher stated that details were disclosed privately to OpenAI rather than being published publicly. Previously, the researcher reported achieving compromise within one hour after the release of GPT-5.