AuraTracer智迹闻
中文

EVENT DOSSIER

GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N]

2026-09-06 03:11 Science 🔥 54.2 heat score hn #15techmeme #2
1sources
1days unfolding
54.2heat score
3mentions
SummaryAI generated

According to reports from the r/MachineLearning community, the GPT-6 Astra model released by OpenAI was jailbroken within 24 hours of its launch. This attack utilized extended Task-in-Prompt (TIP) techniques from the ACL 2025 paper, along with four unnamed techniques. Researchers noted that for GPT-6, the original minimum TIP attack was ineffective and had to be redesigned; this attack succeeded by hiding harmful targets within harmless tasks such as password cracking or executing Python code. Previously, the same researchers reported jailbreaking GPT-5 within one hour of its release. Currently, the details have been privately disclosed by these researchers to OpenAI and have not been made public yet.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
ACL 2025GPT-6 AstraOpenAI

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
ACL 2025 × GPT-6 Astra1ACL 2025 × OpenAI1GPT-6 Astra × OpenAI1

SignalsSIGNALS

Keyword heat
  • GPT-6 Astra1
  • OpenAI1
  • ACL 20251

All reports (1)SOURCES

R r/MachineLearning en 2026-09-06 03:11

GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N]

Researchers reported that GPT-6 Astra was compromised within 24 hours of its release. The attack involved the extended Task-in-Prompt (TIP) technique from the ACL 2025 paper, along with four unnamed techniques. The TIP attack utilized the model’s reasoning or instruction-following behavior to achieve compromise by hiding harmful targets under other tasks, such as password cracking or executing Python code. For GPT-6, researchers noted that the original minimal TIP attack was no longer effective and had to be redesigned. The researcher stated that details were disclosed privately to OpenAI rather than being published publicly. Previously, the researcher reported achieving compromise within one hour after the release of GPT-5.