Policy

Anthropic Debuts AI Safety Framework and $5M Grant

Anthropic has launched a comprehensive framework to assess AI risks alongside a $5 million grant program, aiming to systematically reduce model harms while maintaining helpfulness.

Anthropic1 day agoPolicy
Image: Anthropic

Anthropic has unveiled a structured safety framework designed to identify and mitigate a broad spectrum of AI-related harms. To support this initiative, the artificial intelligence safety startup is launching a $5 million grant program dedicated to funding independent research into how AI technologies affect human wellbeing. This new framework complements the company's existing Responsible Scaling Policy by expanding its focus beyond catastrophic risks to address everyday impacts across physical, psychological, economic, societal, and individual autonomy dimensions.

The company highlighted how this structured approach has already influenced its latest models. For Claude 3.7 Sonnet, Anthropic developers used the framework to navigate the delicate balance between helpfulness and safety. By evaluating user prompts along these safety dimensions, the team trained the model to better handle ambiguous queries rather than defaulting to outright refusals. This adjustment resulted in a 45 percent reduction in unnecessary refusals while preserving essential safety guardrails, particularly for vulnerable populations like children or individuals in crisis.

The framework also guides the deployment of advanced model features, such as the computer use capability. To mitigate risks like fraud in financial software or phishing in communication tools, Anthropic implemented strict enforcement thresholds and a novel hierarchical summarization technique to detect abuse while protecting user privacy. Additionally, the company updated its Claude Fable 5 model to refine its biology safeguards. This update significantly reduces false positives, meaning Fable 5 users will experience fewer instances where the system unnecessarily downgrades to a less capable model during biology-related queries.

Anthropic emphasizes that its safety methodology remains an evolving project. The company is actively seeking feedback and collaboration from external researchers, policy experts, and industry partners, who can contact the safety team directly at usersafety@anthropic.com to help refine these evaluation methods.

This is our own summary of reporting by Anthropic

More in Policy