We are making the EnigmaEval benchmark publicly available. It's a collection of long, complex reasoning challenges that take groups of people many hours or days to solve. Claude Fable 5 and GPT-5.6-Sol are ahead of other frontier models. On the hard set (puzzles that take MIT students days to solve) Fable 5 gets 10%.
Center for AI Safety
Research Services
Reducing societal-scale risks from AI through technical research and field-building.
About us
The Center for AI Safety (CAIS — pronounced 'case') is a research and field-building nonprofit. Our mission is to promote the safe development of artificial intelligence through technical research and advocacy of machine learning safety in the broader research community.
- Website
-
https://safe.ai
External link for Center for AI Safety
- Industry
- Research Services
- Company size
- 11-50 employees
- Type
- Nonprofit
Employees at Center for AI Safety
Updates
-
New Remote Labor Index results: AI automation of real remote work is increasing fast. Claude Fable 5 now completes 16.1% of projects at a professional standard, roughly double the next model and up from Opus 4.6’s 4.2% automation rate. See the following links for full results, methodology, and side-by-side examples. Write-up: https://lnkd.in/gYAJN8Ne Leaderboard: https://dashboard.safe.ai/ RLI is joint work by Center for AI Safety and Scale AI
-
-
We've welcomed Phil L. to CAIS as our new Vice President of Operations. Phil has spent nearly two decades making ambitious organizations work. He previously led operations at Google's R&D organization focused on environmental sustainability and workplace innovation and served as COO of a leading digital design and development agency. As CAIS enters a phase of rapid growth, Phil will oversee our Operations and Program teams, building the foundation that allows our Research, Public Engagement, and Programs teams to execute our mission at scale. Read more: https://lnkd.in/eHPsteb9
-
As Devin Kim shared with the The Chronicle of Philanthropy, "The number one thing academia does is it establishes independent research credibility outside of what the companies are telling you. Universities are able to interpret the research and speak freely about what the implications are without being bound by short-term commercial interests." When researchers and employees at frontier labs start to feel their work is doing harm, that changes decision-making from the inside out. For these powerful technologies to be built responsibly, we need more independent and credible voices in the conversation. https://lnkd.in/ezdQRGJX
-
We created the AI Values Dashboard, which measures who AIs favor the most. By popular request, we added a Sports Tier List, ranked by AIs from OpenAI, Anthropic, and DeepSeek. See the full Sports Tier List (World Cup, Soccer, NFL, NBA, F1) by AIs at values.safe.ai/sports What group do you want to see next?
-
-
AI systems grow more powerful and more deeply integrated in the decisions that shape daily life, from the information that billions of people consume to the infrastructure governments depend on, the stakes of getting AI development wrong are higher than ever. CAIS exists to reduce societal-scale risks from AI through research, field-building and advocacy, and that mission begins with understanding unexpected behaviors in AI systems. CAIS’ latest wave of research findings explores issues of AI wellbeing, identity, political bias and systemic betrayal risk. Together, this work maps new frontiers to help define what it means to build AI that is safe, honest and aligned with human interests. Link to all papers is in the comments.
-
What biases do AIs have? It turns out, AIs show strong favoritism toward specific people, countries, and companies. Our interactive AI Values Dashboard tracks who Claude Fable and other AIs favor most. AIs value people whose work aims to benefit humanity, such as AI safety researchers and scientists. Net worth, job title, and other typical success metrics don’t greatly affect how much an AI values you. And as models scale up, they get even more confident in their values. Fable’s political preferences are somewhat predictable. Its favorite politician is Obama. Grok is the only model that ranks the US in the S tier; the rest do not. For example, GPT ranks the US as B tier. There's far more in here than this post can cover (e.g., AIs’ favorite pokemon, GPT trusts Dario over Sam, DeepSeek prefers the US to China). Tell us who to measure next. If enough people ask for a person, company, or category, we'll add it to the dashboard. Link is in the comments.
-
-
We're excited to welcome Rochelle Nadhiri to CAIS as Head of Public Engagement. Rochelle brings two decades of storytelling experience at the intersection of technology and policy for industry leaders including Meta and Robinhood. She will lead CAIS's efforts to translate AI safety research into narratives that move audiences beyond the technical community. Rochelle will oversee our media relations, narrative strategy, and public-facing communications with a focus on building sustained public understanding of AI safety across policy, journalism, technology, and general audiences. ⬇️
-
We are pleased to share that Mantas Mazeika, Research Scientist at CAIS, has been appointed to the European Commission’s AI Act Scientific Panel (EU Digital & Tech). As a member, Mantas will advise the European AI office and national authorities on general-purpose AI (GPAI) models, as well as the implementation of the AI Act, the EU’s landmark legislation governing the development and use of artificial intelligence. The panel’s leading technical experts will play a direct role in ensuring that AI is built and deployed responsibly across Europe. We look forward to seeing how the panel’s work informs how frontier AI is governed in Europe and beyond. Full article in the comments ⬇️