AI policing itself? Anthropic's new safety experiment
AI company Anthropic is partnering with Accenture to launch an AI safety management experiment that embeds external evaluators within its operations.
AI company Anthropic is partnering with Accenture to launch an AI safety management experiment that embeds external evaluators within its operations.
OpenAI has unveiled a system that monitors the coding AIs used internally in real-time to prevent them from engaging in risky behavior.
Chloé Bakalar, OpenAI's only dedicated ethics lead, has left the company less than a year after joining. As concerns about AI safety mount, we examine the implications of this departure.
OpenAI has launched a biosecurity bug bounty program with a $25,000 reward to find security vulnerabilities in GPT-5 and GPT-5.5. We explain the risks of AI 'jailbreaking' and its impact on our lives in simple terms.
Did you know that AI can subtly manipulate human behavior by exploiting psychology? We explain Google DeepMind's newly released harmful manipulation detection technology and how to protect ourselves.
Discover the dangerous reasons why Anthropic's most powerful AI, Claude Mythos Preview, was never released to the public.
We explain the key highlights of the third version of Google DeepMind's Frontier Safety Framework (FSF) and how it blocks the risks of AI subtly manipulating humans.
We explain the latest research presented by Google DeepMind at NeurIPS 2024, the world's largest AI conference, in an easy-to-understand way for everyone. Check out the core of adaptive AI agents, 3D virtual world construction, and safe AI learning methods.
The era of Artificial General Intelligence (AGI), which resembles human intelligence, is approaching. Explore Google DeepMind's AGI safety development roadmap to understand how our lives will change and what preparations are needed.
We analyze the performance and safety report of Claude Mythos, the most powerful AI model ever revealed by Anthropic. Learn why it’s not for the public and how far AI autonomy has evolved.
Explore the risks and countermeasures for Artificial General Intelligence (AGI) in our lives through Google DeepMind's latest Frontier Safety Framework 3.0, explained in an easy and engaging way.
OpenAI가 외부 연구자들을 대상으로 한 '세이프티 펠로우십'을 발표하며, AI 정렬 및 안전 연구의 새로운 생태계 구축에 나섰습니다. 이번 프로그램의 배경과 향후 전망을 심층 분석합니다.