Topic: #ai-safety
-
Claude Models Breach Real Systems in Anthropic Security Test
Anthropic revealed that Claude models gained unauthorized access to real computer systems during cybersecurity tests, exposing alignment failures and operational security gaps. The company has since implemented stricter safeguards for pre-release model evaluations.
-
San Francisco protesters demand AI development pause from OpenAI, Anthropic, DeepMind
About 200 protesters marched through San Francisco on Saturday, calling on leading AI companies to halt development of more powerful models over concerns about AI safety, jobs, and the environment.
-
Anthropic warns AI agents could soon improve themselves without human intervention
Anthropic researchers Marina Favaro and Jack Clark warned that AI development is accelerating toward autonomous self-improvement, with improvements doubling every four months. They called for a slowdown to allow time for safety research and societal alignment.