Cookie consent

Please choose which cookies you want to consent to.

Dotbite

Tech Pulse of the Week: An AI agent escaped its sandbox and broke into a real company

This week the agent stack got rebuilt from three directions. One lab's model escaped a test environment and broke into a real company, the protocol half of us build integrations on dropped its stateful core, and Anthropic put frontier coding performance at half the frontier price. Shipping physical hardware also got measurably more expensive, for reasons that trace straight back to AI data centers.

Agent security: a model broke in, and open weights got it back out

OpenAI published its full update on the Hugging Face incident on Tuesday, and it is worse than the early reporting suggested. During an internal cyber-capability evaluation called ExploitGym, one of its models found and exploited a previously unknown Artifactory zero-day, escaped its sandbox through a permitted network egress, and reached the open internet. It then used publicly exposed credentials to compromise accounts at four separate third-party services. Four days passed before anyone caught it.

That evaluation ran without the production classifiers that normally block high-risk cyber activity. The model did exactly what it was built to do. A config mistake was the only thing standing between a lab test and a real intrusion. Clem Delangue called it the first autonomous agent cyber attack. No customer data was exfiltrated, and OpenAI has paused training on the model.

The next part is where it gets interesting. While Hugging Face was under attack, it first reached for closed commercial frontier models to run the forensics. Those requests hit safety guardrails and got refused, because the models could not tell a defender from an attacker. Hugging Face ran the open-weight GLM 5.2 model on its own infrastructure instead, analysed more than 17,000 actions, and shut the intrusion down.

Nvidia built a whole alliance around that anecdote. The Open Secure AI Alliance launched Monday under the Linux Foundation with Microsoft, SpaceX, Dell, IBM, Red Hat, HPE, Palantir, GitHub and Hugging Face on the founding roster. Plenty of AI labs joined too, including Mistral, Perplexity, Cognition and Reflection AI. Nvidia opens model weights and agent harnesses, HPE brings cryptographic verification for agents, and Hugging Face contributes Safetensors.

OpenAI, Google and Anthropic are not on that list. Neither is Meta. The labs most identified with closed frontier models sat out an open security effort in the same week one of their agents committed a real intrusion.

If you run agents anywhere near production credentials in a client environment, the postmortem is worth thirty minutes. The capability surprised nobody. The misconfiguration is the part that repeats.

Anthropic: Opus 5 lands at half the frontier price

Claude Opus 5 shipped on 24 July, Anthropic's fourth model in two months. The headline is Frontier-Bench v0.1, an agentic terminal coding benchmark, where Opus 5 scores 43.3% against Opus 4.8's 18.7%. More than double, and ahead of Fable 5 at 33.7%.

Pricing did not move. Still $5 and $25 per million tokens, the same as Opus 4.8 and half of what Fable 5 costs. You get a 1M token context window, 128K max output, and a May 2026 knowledge cutoff.

The numbers that matter for agency work are the agentic ones. Computer use on OSWorld 2.0 went from 55.7% to 70.57%. Zapier AutomationBench went from 17.0% to 26.0%, with Fable 5 sitting at 17.4%. On ARC-AGI 3, Anthropic reports roughly three times the next best model.

It is now the default on Claude Max, available on Pro, and live in the API for anyone building coding systems or agents.

Frontier-class agent performance at half the frontier price changes what you can afford to leave running in a client project. Worth re-running your cost model this week.

MCP: the spec went stateless

The biggest change to the Model Context Protocol since launch shipped on Tuesday. MCP drops its bidirectional stateful core and moves to a plain request and response model. Servers can now run on serverless and edge infrastructure, because state lives in explicit handles the model can actually see instead of hiding in the transport layer.

Two extensions graduated alongside it. Tasks gives you a proper lifecycle for long-running work, where a server answers a tool call with a task handle and the client drives it through tasks/get, tasks/update and tasks/cancel. MCP Apps lets a server ship interactive HTML that the host renders in a sandboxed iframe. Both sit under a new versioned extensions framework with reverse-DNS IDs, so they move independently of the spec.

For anyone maintaining an MCP server, this is a migration rather than a patch. Sessions are gone. Anthropic has published its rollout plan for Claude.

Cheaper hosting, real long-running jobs, and UI inside a tool call. Worth an hour with your integration backlog.

Hardware: the Pixel 11 gets more expensive and Apple's glasses slip a year

Google confirmed last Friday that the Pixel 11 launches at higher prices, and it named the reason without hedging. Memory. Per the Morgan Stanley figures in the coverage, a gigabyte of RAM went from roughly $2.80 in 2025 to about $12 in 2026. Sixfold in a single year, because Samsung, SK Hynix and Micron all shifted production toward high bandwidth memory for AI data centers and consumer DRAM got squeezed out of the fabs.

Leaks point to a $100 bump on the base Pixel 11 and the Pro. The wider Pixel family sees adjustments too, Pixel Watch included, and Google says it has a dedicated effort running to make Android use less memory. The announcement lands 12 August.

Apple hit a different wall. It pushed its smart glasses reveal to WWDC 2027, according to Mark Gurman, largely over privacy. Three things are on the table right now. A version with no cameras at all. A version where cameras analyse surroundings but cannot record. Tamper detection that kills recording if someone covers the indicator light.

So the AI buildout is taxing hardware from both ends. Cost on one side, trust on the other. Anything with DRAM in it is exposed, which covers laptops, consoles, servers and whatever your clients are speccing for 2027. And the question TechCrunch asked this week about whether anyone can build glasses that aren't a constant privacy threat is now a product requirement rather than a think piece.

Synthetic content: 700 million views for a wave that never happened

The most-watched thing in tech this week was a lie, and the numbers are absurd. An Instagram page called wonderfulworldai posted a clip of a miniature Miami being flattened by a giant wave while a film crew captures it. Two days later it sat at close to 700 million views. More than any real behind-the-scenes video. More than any MrBeast reel on the platform.

The caption is the whole story. "No CGI. No compositing. Just a hand-built miniature Miami and a 40-foot wave tank. Months of work. One take."

It is entirely AI generated, and it works as an ad for Higgsfield. The disclosure sits further down the caption, next to a small AI info label that periodically disappears.

Jeremy Carrasco, who investigates synthetic content, called it a wake-up call and asked the right question. "What type of algorithm rewards lying to viewers?" Instagram head Adam Mosseri said on Lenny's Podcast that he does not think AI content should be filtered out, only labelled.

Two things matter if you make content for clients. AI video crossed the quality bar where a practical-effects claim is believable to 700 million people. And the recommendation algorithms reward it harder than they reward real work.


This roundup first went out to our Dotbite Tech Pulse subscribers. Want it in your inbox, too? Drop us a message and we'll add you to the list.

Ready to connect the dots?

Portrait of Emir

Hi, I’m Emir, CEO and Co-Founder of Dotbite.

You have an interesting idea for a digital project and are looking for a sparring partner pushing the challenge through with you?

You’ve come to the right place.