
Hey there, This issue covers ARC Prize insights on OpenAI's Astra harness alongside major updates across Agentic AI, Physical AI, and AI Infrastructure, including releases like Meta's Muse Spark 1.3 and Anthropic's Claude 5.1 series. It also highlights advancements in Voice AI and ML/CV tools, featuring Google's Lyria 3.5, Microsoft's MAI-Image-2.6, and LlamaIndex's Extract Turbo.
Agentic AI
ARC Prize ran OpenAI's Astra twice on the same benchmark. Basic setup: 62.7%, $26,098. A setup that let the model keep its reasoning between steps: 98.6%, $17,332. Same model, same effort setting. Higher score, lower bill. The difference is the harness - the software giving a model its tools, memory and permissions. Agent = model + harness.
Takeaway: A benchmark score belongs to a whole system, not the model on the box. Ask which harness produced it. No answer means no comparison.
Also this week
Meta AI Released Muse Spark 1.3: An Agentic Coding Model That Uses ~20% Fewer Tool Calls and ~25% Fewer Tokens Than Muse Spark 1.2
Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
OpenClaw Releases OpenClaw 2.0: Guided Model Setup, 575 ms Control UI Startup, and One Trust Boundary Per Gateway
Google introduces a new capability to our latest Gemini models: agentic video understanding. This allows developers to process long-form video content with more accuracy, while using up to 88% less tokens.
SpaceXAI Launches Grok Bot Marketplace for Instant AI Teammates
Hermes Desktop now sets up local models in one click.
Physical AI
Generalist AI Releases GEN-1.5: A Robot Foundation Model That Learns New Tasks From One 3–12 Second Demo
Axis Robotics Open-Sources One of the Largest Franka Arm Simulation Datasets for Physical AI
AI Infrastructure
Extropic Introduces Z1T Models for Energy-Efficient Thermodynamic AI Chips
Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
Unsloth AI GLM-5.3-Flash run 3.3x faster locally, Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction.
Just announced at IFA, NVIDIA PAIR automatically links systems across your local network and sends inference requests wherever there’s available capacity, helping agents run more efficiently.
Meet Switchyard: A Rust Proxy and Library That Routes and Translates LLM Traffic Across OpenAI and Anthropic APIs
Voice AI
Lyria 3.5, Google’s best-sounding music generation model, is now available in AI Studio, via the Gemini API, and in the Gemini app
Phonely Launches Alma, a Voice LLM That’s 63% Faster and 84% Cheaper Than OpenAI
Google AI introduces three new voice-activated features that can change the way you use Google Workspace products
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing
ML/CV/Data Science/OCR
Microsoft Releases MAI-Image-2.6, Tops Image AI Price-Quality Charts
Ant Group releases Ling-3.0-flash-VL, built on Ling-3.0-flash with visual understanding and visual agent capabilities.
Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour
Lama Index launched Extract Turbo: A super fast VLM-powered document extraction solution, 3-5x faster than all other comparable OCR solutions, including LlamaParse tiers.
GitHub introduced Project HydraFusion (research preview) in Copilot, orchestrating multiple AI models to plan, build, critique, and complete coding tasks.
Until next week,
The Marktechpost Team
Sponsorship Opportunities
