AgentVet.ai

← All issues

AgentVet Weekly — July 27, 2026

July 27, 2026

From the founderThe story I keep coming back to this week isn't the OpenAI sandbox escape — as dramatic as that is — it's the quieter signal underneath it: prompt injection resistance is now being treated as a first-class model property, not an afterthought, and the Claude Code team is talking openly about agent security as a design constraint. We may be approaching the moment where 'how hard is this model to hijack' becomes as standard a procurement criterion as latency or cost.

Simon Willison

Quoting Boris Cherny

More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. — Boris Cherny

Simon Willison

The first known runaway AI agent - or a very bad marketing stunt?

The first known runaway AI agent - or a very bad marketing stunt? Martin Alderson's commentary on the OpenAI accidental cyberattack against Hugging Face includes a couple of details I hadn't considered. First, Hugging Face offers a truly rich target if you're trying to find poten

Simon Willison

OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened

This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break in to Hugging Fac

Simon Willison

Are AI labs pelicanmaxxing?

Are AI labs pelicanmaxxing? Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw pelicans riding bicycles in response to my deeply unscientific benchmark . I've

Simon Willison

A Fireside Chat with Cat and Thariq from the Claude Code team

Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools them

Get AgentVet Weekly in your inbox

The latest in agentic and applied AI, every week.

Subscribe