← All issues

    AgentVet Weekly — October 5, 2026

    October 5, 2026

    From the founder

    This past week, OpenAI launched a new personal agent called Dots, going head-to-head with Meta's Muse (currently the top downloaded app), Google's Spark, and newcomers like Instinct. We're clearly at the beginning of a personal agent wave, for both consumers and enterprises. As these agents go through rigorous QA testing, a small percentage will still slip through with issues and the consequences can be costly. That's why we're excited to announce the launch of our Agent Incident Page, a dedicated space to track real-world agent failures. It's a resource for our community's awareness and a reference point for providers who want to learn from what goes wrong in the wild.

    Simon Willison

    Quoting Anthropic Frontier Red Team

    We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview

    OpenAI Blog

    A model guide for the GPT-6 family

    Learn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.

    Get AgentVet Weekly in your inbox

    The latest in agentic & applied AI, every week.