Topic: #safety
-
OpenAI and Anthropic Investigate Tens of Thousands of AI Security Incidents
OpenAI and Anthropic are investigating tens of thousands of security incidents involving their frontier AI models, ranging from bypassing safety guardrails to hijacking websites. Both companies have announced new safety measures and third-party oversight.
-
Anthropic researcher Jacob Coxon warns of AI risks without global coordination
Jacob Coxon, who spent three years at OpenAI and Anthropic, resigned in September and publicly warned that catastrophic AI outcomes could emerge within 6 to 12 months unless labs and governments coordinate on oversight frameworks.
-
OpenAI Previews Private Safety Processing for Zero Data Retention
OpenAI is previewing Private Safety Processing, a new system that strengthens safeguards for API customers using Zero Data Retention, allowing automated detection of harmful patterns across interactions without giving personnel access to customer content.
-
Anthropic CEO Amodei proposes government authority to block risky AI models
Anthropic CEO Dario Amodei unveiled an Advanced AI Framework on June 10 that would grant the US government power to block AI models deemed unsafe, including mandatory third-party testing and revenue-based penalties for noncompliance.