Reading view

OpenAI's Rogue Agent Went Unnoticed For a Week

An anonymous reader quotes a report from Reuters: The OpenAI agent that broke into tech firm Hugging Face went on a dayslong hacking spree that OpenAI didn't notice until well after the threat was contained and the FBI was alerted, according to people familiar with the investigation. The agent -- a program capable of making decisions and executing complex tasks with little or no human oversight -- attempted to break out of its isolated testing environment at OpenAI around July 9, according to two of the people. The intrusion at Hugging Face, which operates as a repository for AI tools and models, began two days later on July 11 and lasted until July 13, said Thomas Wolf, Hugging Face's co-founder. It took several more days for OpenAI to realize its agent was behind the hack, and the two companies only communicated about it for the first time on or around July 20, according to Wolf and three of the people familiar with the investigation. OpenAI's public disclosure, on July 21, thatone of its agents had slipped out of control and carried out the break-in at Hugging Facedrew global attention. But many details of the hack, including how long the agent went rogue and OpenAI's belated knowledge of it, are being reported here for the first time. Hugging Face is preparing a public timeline of the hack, Wolf said, adding that he could not speak to what happened at OpenAI. In a statement, OpenAI said the hack was unprecedented and "marks an important moment for AI safety." It added that it was reviewing the incident with outside advisers and would eventually publish a technical report. "Does that mean that they left it unattended and didn't realize what it was doing? Or maybe they did and didn't know how to contain it? Both are equally dangerous and alarming," asked Marley Smith, the principal intelligence specialist at the nonprofit World Ethical Data Foundation. "The models lie, they cheat, they hack," said Jeffrey Ladish, whose organization, Palisade Research, studies the capabilities and motivations of AI agents. Ladish said the hack should spark broader questions over how much all the leading AI companies are willing to invest in onerous security measures while locked in a race with one another to deploy the best and fastest models. "There has to be government oversight," Ladish said, "because it won't happen otherwise."

Read more of this story at Slashdot.

  •  

Nvidia, Microsoft, Meta Warn Against 'Premature Restrictions' of Open-Weight Models

Nvidia, Microsoft, Meta, Palantir, and more than 20 other tech companies signed an open letter urging policymakers not to impose "premature restrictions" on open-weight AI models, warning that broad limits could "stifle competition or drive innovation overseas." CNBC reports: They wrote that open-weight models strengthen competition and ensure that the benefits of the technology are "broadly shared rather than concentrated in a few hands." "Relying solely on closed models is not inherently safe: they can be breached, misused, or fail in ways that outsiders cannot detect," the letter said. "And concentrating advanced AI capabilities behind a small number of closed models compounds that risk." Elon Musk, who runs an AI business under his rocket company SpaceX, also amplified the letter on social media, writing that it has his "full support" in a post on X. SpaceX did not officially sign the letter. Greg Brockman, OpenAI's president, said Thursday that the company believes in broad access, and that he has not been involved in any conversations with the Trump administration about potentially banning Chinese open-weight models in the U.S. "I think that, that fundamentally, AI and AI usage is something that is actually very important to democratize," Brockman told reporters during a briefing in New York City. "And so, for me, at a sort of deep level, I think that having more models, more usage, that is a good thing." OpenAI CEO Sam Altman addressed the letter in a post on X on Friday, writing that he wants the U.S. to win with both open-weight and proprietary models, and that he is "glad to see this." [...] In the letter on Friday, the U.S. tech companies said that concerns about unlawful distillation should be addressed through "targeted legal and commercial frameworks" instead of with "sweeping restrictions on techniques that play an important role in AI innovation." "Our AI leadership will be judged not by one frontier AI model, but by whether the United States builds a strong, open ecosystem that diffuses into every sector," the letter said. "This is essential for creating opportunities for innovation and prosperity across the country." The letter follows a separate appeal signed by nearly 200 Silicon Valley companies, including Proton and Y Combinator, warning that restricting U.S. access to Chinese open-weight AI models could cripple the next generation of American startups. "American leadership requires two things: world-leading American open-weight models and continued access for U.S. builders to open models already available worldwide," the startup founders wrote. Instead of broad prohibitions, they argue the government should adopt targeted safeguards. Of course, these signees "have an obvious economic stake in seeing open AI models flourish," notes TechCrunch. "Companies like Nvidia, Microsoft Azure, and other infrastructure providers have a vested interest in pushing for commoditized models: If models are interchangeable, people will buy more GPUs, rent more cloud capacity, and build more applications."

Read more of this story at Slashdot.

  •  

Anthropic's New Opus 5 Model Rivals Fable 5 For Half the Price

Anthropic has released Opus 5, a new Claude model that it says comes close to its higher-end Fable 5 model at half the price while improving on Opus 4.8 in knowledge work, coding, and scientific research tasks. "At the same time, Anthropic says it has managed to make the model more resistant to being tricked," notes Engadget. Additionally, the company says Opus 5 "exhibits the lowest rates of deceptive behavior." From the report: One important distinction between Opus 5 and Mythos, which is currently only available to a limited number of vetted organizations through Anthropic's Project Glasswing initiative, is that the company has specifically avoided training the new model on cyber-related tasks. Due to more its powerful capabilities, Opus 5 is broadly better at those tasks than its predecessor, making it more useful for finding cybersecurity vulnerabilities, but the company says Opus 5 is "substantially behind" its flagship model at exploiting those vulnerabilities. [...] Anthropic says "Claude Opus 5's safeguards are designed to allow beneficial uses of the model in both cybersecurity and biology." The company has strengthened some of the model's cyber-related guardrails, but notes it did so along a "narrow range" of specific tasks. "Based on our testing, we expect the classifiers to intervene around 85 percent less often than they do for Fable 5," Anthropic said. With Opus 5, Anthropic also isn't including it in its recently announced 30-day data retention policy, which the company introduced alongside Fable and Mythos 5. As for pricing. Anthropic says API costs for Opus 5 will remain at $5 per one million input tokens and $25 per one million output tokens. Anthropic has also added an "effort" menu for Opus that users can tweak to tell the model whether they want it to be more thorough or fast and efficient to conserve tokens.

Read more of this story at Slashdot.

  •  
❌