The presenter described Anthropic's safety practices in detail, saying the company tests models for dangerous capabilities in areas such as cybersecurity and biology, publishes multi-hundred-page risk reports, and conducts research on interpretability and alignment. "We put strong safeguards on our most capable models, and we publish our work in these areas," he said, asserting that these steps improve their ability to steer and control models.
He added that as models grow more powerful, more stringent standards will be required and emphasized a company commitment to slow technology releases if necessary: "We will slow down as much as necessary in order to make sure that every successive AI technology that we release is actually safe." The remarks were presented as part of a broader appeal for industry and government collaboration.