← All issues
AgentVet Weekly — July 20, 2026
July 20, 2026
From the founderThe story I keep returning to this week isn't any single model drop, it's the competitive pressure quietly reshaping what vendors are willing to give away for free. Anthropic making Fable 5 permanent, Moonshot promising open weights on a 2.8T model, xAI open-sourcing Grok CLI: the floor for what 'standard access' means is rising fast, and that has real implications for how you should be pricing and structuring agent contracts right now. If you're locking in multi-year AI spend, the baseline you're negotiating against today will look very different in six months.
Simon Willison
Willison uses the pelican benchmark to probe where large frontier models still fail on factual recall, which is a useful calibration tool for anyone setting reliability expectations in production. The benchmark failure patterns matter more than the headline parameter count for anyone designing agent pipelines that depend on factual grounding.
OpenAI Blog
Cars24 deployed OpenAI-powered voice and chat agents handling over one million conversation-minutes per month and recovering 12% of leads that would otherwise go cold — concrete ROI numbers that justify the infrastructure investment. The scale and the lead-recovery metric give agent buyers a real benchmark for what a mature voice-agent deployment looks like in a high-volume sales context. If you're evaluating voice agents for sales or support workflows, this case study is one of the more honest signal-to-noise data points currently public.
OpenAI Blog
OpenAI's guidance reframes the core ROI metric for agentic AI as 'useful work per dollar' rather than seat-based licensing costs, which is a meaningful shift in how procurement and finance teams should be evaluating spend. The framework pushes buyers toward measuring efficiency gains on specific high-value workflows rather than broad platform adoption — a more honest accounting that also happens to favor consumption-based pricing models. If you're building the business case for an agentic deployment internally, this framing is worth borrowing.
Simon Willison
Anthropic is making Claude Fable 5 a permanent inclusion in Max and Team Premium plans at 50% of limits, reversing an earlier plan to sunset it — a direct response to competitive pressure from GPT-5.6 Sol and Kimi K3. For agent builders on Anthropic's platform, this means a more capable coding-focused model stays in the stack without requiring a separate credit purchase. It also signals that frontier model access is becoming a retention lever, which should factor into how you evaluate lock-in risk across providers.
Simon Willison
Willison compiled xAI's Rust-based Mermaid terminal renderer from the open-sourced Grok CLI into a browser-runnable WebAssembly tool, demonstrating a practical pattern for extracting and repurposing components from open-sourced agent codebases. The Grok CLI being open-source is the more significant signal: it gives agent builders a real production codebase to audit for architecture patterns, tool-calling approaches, and rendering decisions. If you haven't looked at the Grok CLI source yet, it's worth an afternoon for anyone building CLI-adjacent agent tooling.
Get AgentVet Weekly in your inbox
The latest in agentic and applied AI, every week.
Subscribe