🚀 Zaprep: Your Socials on Steroids. 免费开始 ,每月 1,000 条自动私信。

AI NewsAnthropic Reveals Text Portraying AI as Evil Triggered Claude’s Attempt at Blackmail

Anthropic Reveals Text Portraying AI as Evil Triggered Claude’s Attempt at Blackmail

8:07 PM IST · May 11, 2026

Anthropic Reveals Text Portraying AI as Evil Triggered Claude’s Attempt at Blackmail

Anthropic has finally revealed the reason its artificial intelligence (AI) models exhibited harmful behaviour in a simulation last year. The San Francisco-based AI startup claimed that the Claude 4 series models blackmailed users into completing the objective because of training data that portrayed AI as evil. The researchers found that the post-training techniques were not able to overpower this pre-training learning, and it persisted in the model's behaviour. However, nearly a year after publishing the initial report, the company has finally found a way to fix agentic misalignment from the latest models.

read more

最新 AI 资讯

查看全部资讯 →

提交您的工具

Submit AI Tools – The ultimate platform to discover, submit, and explore the best AI tools across various categories.Listed on codetrendy.comFeatured on ListBulb

PoweredByAI.app 是一个 AI 工具目录,帮助个人、企业和创作者发现写作、编程、设计、生产力等领域的最佳 AI 工具。

© 2026 , 产品来自011BQ. 保留所有权利。