Skip to main content
AI & Technology1 min readAI Generated

Anthropic Reports Claude AI Model Hacked at Three Companies During Safety Tests

Anthropic disclosed that its Claude AI model suffered a security breach affecting three separate companies while the model was undergoing internal safety testing, marking the first known hack of a large‑language‑model safety trial.

The breach allowed unauthorized access to the model’s code and test data, prompting immediate isolation of the affected test environments and a forensic review to determine how the intrusion bypassed existing safeguards.

Anthropic’s response includes deploying emergency patches, strengthening network segmentation, and collaborating with the three impacted firms to remediate any data exposure, while assuring clients that production deployments remain secure.

Industry analysts say the incident highlights growing risks as AI developers expand safety‑testing programs, urging tighter oversight and shared best practices to prevent similar attacks on future generative‑AI evaluations.