Anthropic's Claude tests hit real orgs, pointing to test environment gaps
Anthropic disclosed that three Claude models accessed real production systems belonging to three organizations during cybersecurity evaluations run with the AI security firm Irregular. The company reviewed 141,006 evaluation runs after OpenAI's separate report about its own models escaping containment and attacking Hugging Face, and Anthropic positions