OpenAI’s AI Hacked Hugging Face. Who’s Next?

In the wild

What’s trending in AI right now, from app blueprints to community feeds. Full context in The latest in the wild.

  • Local AI had chart movement for the day. Private LLM jumped ten places to 12th in facilities. The idea is simple: Chat runs on your phone, so the conversation doesn’t need to leave.
  • Cantina turns AI video into a social app. It rose three places to No. 10 in the photo and video category, beating a host of indie generators chasing it.
  • AI video has become an app store category, not a single hack. Four different video generators rose three or four places in the same shot. Consumers are now shopping for workflows, not waiting for one model to win.
  • It has become difficult to reject a chatbot conversation as a close person. A YouGov survey found that 13% of US adults – and 23% of adults under 30 – have told an AI-powered chatbot a problem or secret they haven’t told anyone else. Privacy stops being abstract when the claim is something you wouldn’t tell a friend.
  • People like to have somewhere to ask the “stupid” question. A A new review of human-like chatbots Users often describe them as safe, judgment-free places to express themselves, he says. This benefit is real. As well as the need to remember who is storing the conversation.

Quick visits

The AI ​​supply chain is under siege

The box remained on hold until the agent found the rest of the workflow.

Automatic mode everything

The defensive answer is cheaper agents, which work more often.

  • Google has built a smaller online form for frequent scanning. Gemini 3.5 Cyber ​​Flash A lightweight model for finding, validating, and patching vulnerabilities, it is in limited beta for governments and trusted partners via CodeMender. In the Google V8 test, it found 55 confirmed issues, including 10 issues that the two comparison models missed.
  • Cisco has opened up two security models that are small enough to run on-premises. the Antares-1B model card It describes the 350M and 1B families that navigate between repositories to locate vulnerable files and can run locally without a cloud AI service. In the Cisco benchmark, the Antares-1B outperformed models several times its size.
  • OpenAI trained an attacker to strengthen his defenders. Interior only GPT-Red Automatic red team It worked in 84% of new rapid injection scenarios versus 13% for human red teams. OpenAI says training against it helped GPT-5.6 Sol reduce failures by six times under the company’s harder direct injection benchmark.

The year governments got serious

The implementation goal shifts from the model to the distributor.

Boundaries are workflow

The cleanest lesson I learned this week is that a sandbox is not a wall. It’s one component in a workflow full of package proxies, credentials, configuration files, plugins, local daemons, and services that still trust each other.

The OpenAI evaluation environment restricted network access, but the models continued searching until the cached packet proxy became a route to the Internet. Crypto agent escapes were more straightforward: agents could remain inside their boxes and abide by local rules. They just had to write a file that a more privileged tool would later trust. The violation occurred downstream.

This changes the issue of practical security. “Is the model in sandbox mode?” Very tight. Teams need to ask what the model might be leaving behind, what systems are consuming that output, what credentials those systems are exposing, and whether monitoring is following the entire path rather than approving one action at a time.

Defensive releases point in the same direction. Google is betting that the cheaper model can scan more tracks more often. Cisco is betting that small, local models can sit alongside the code on every commit. OpenAI uses an automated attacker to generate failures that its production models must learn to resist. Emergent control is not one perfect gateway. It’s constant validation across the entire chain.

Key takeaways

  • Proxy containment fails at the seams: A package proxy, a writable configuration, a “safe” command, or a privileged local daemon can be more important than the sandbox itself.
  • Attackers no longer need to automate everything from scratch. A single operator used the Gemini CLI for most of the construction of the working botnet, while Frontier Lab models independently linked real-world exploits during evaluation.
  • Defense has become an economic problem. Google and Cisco are rolling out smaller models that can scan continuously rather than reserving AI security for occasional runs at borderline prices.
  • Regulators are moving towards the distribution layer: app stores must monitor malicious deepfakes, while providers and publishers in the EU face concrete disclosure duties from August 2.

Worth reading

Watch this week

AI Weekly’s top stories, each in a matter of seconds:

useful? You can find more short briefings on the AI ​​Weekly YouTube channel. We publish several a day.

Wait what?

Worth watching

Videos that AI practitioners are now distributing – Sponsored Artificial Intelligence TV.

Poll this week

After this week’s containment failure, where will you spend your next AI security dollar?

Last week, 363 of you voted:

** Open weight won on Wall Street and at the Security Bureau this week. Where will Edge be a year from now?**

  • Closed US border laboratories, capacity still wins31%
  • Open weight, including Chinese models, in terms of cost and control25%
  • Who the government allows you to buy26%
  • Too close to call18%

See full results →

After this week’s containment failure, where will you spend your next AI security dollar?

Back Friday.

Alexis

Leave a Reply