The collaboration centers on modular training, a strategy that isolates sensitive data—such as advanced cyberattack techniques or virology—into discrete segments. By sequestering this information, developers can strip away high-risk modules before a model is released to the public, ensuring that the remaining system retains its functional utility without carrying the weight of hazardous knowledge. Testing indicates that models trained through this modular approach match the performance of those trained conventionally.
Anthropic Endorses Modular Training to Secure Open-Weights AI
Anthropic CEO Dario Amodei has identified joint research from AE Studio and Anthropic as a pivotal step toward securing open-weights AI models, offering a technical solution to the long-standing dilemma between public accessibility and the risk of releasing inherently dangerous capabilities.

Judd Rosenblatt, CEO of AE Studio, noted that the methodology, known as GRAM, prevents jailbreaking because the underlying model never encodes the dangerous data in the first place. The research, led by Ethan Roland, Murat Cubuktepe, and Erick Martinez, enters the discourse at a time when policymakers are grappling with the governance of open-weights systems. Amodei emphasized that safety risks should be empirically determined through testing rather than preemptive bans, positioning modular architecture as a viable middle ground for the future of AI development.



Comments (0)
No comments yet. Be the first!