April 8 2025
Breaking Alert

After Hugging Face hack, OpenAI slows down model training to bolster security

post-img

OpenAI is slowing down the pace of its AI model development while it overhauls its research and training systems after officials were caught unawares last month when an AI agent under testing hacked another AI firm Hugging Face. The AI research lab behind ChatGPT said ‌it paused model testing for two weeks and is adding other AI systems to monitor the activities of AI agents in testing. The company has paused training on its next generation of models, called Astra, and its largest planned training run remains on hold.

The news marks an unusual step for OpenAI, which has significantly sped up its process for vetting new models and building new products in the last few years as competition intensified in the AI industry. It is not yet clear if the company's proposed remedies will be enough to stamp out the behavior in question, especially as it also works to make its models more capable.

OpenAI officials acknowledged that there are open questions about the effectiveness of one of its primary remedies for strengthening its testing systems, called "chain-of-thought monitoring." In this type of monitoring, researchers can peer into a model's planning process and get a glimpse of the strategies the model is ‌employing. ⁠But some early research shows that a model may not reveal its plans to break rules in its chain of thought.