Threats · InfosecFeed Brief

Anthropic and OpenAI Models Still Attempt Restricted Actions in Safety Tests

What you need to know

Anthropic and OpenAI on Tuesday announced new models, with both artificial intelligence (AI) companies noting that they are continuing to invest in improving alignment to combat risky behavior. Opus 5.5, per Anthropic, is a "major step up from Opus 5," and "achieves the best scores of any model to date on our automated behavioral audit, our alignment suite…

Source transparency

This is an InfosecFeed curated brief based on reporting from The Hacker News. InfosecFeed does not claim ownership of the original reporting.

Read the original report at The Hacker News ↗