Hey there, This issue covers ARC Prize insights on OpenAI's Astra harness alongside major updates across Agentic AI, Physical AI, and AI Infrastructure, including releases like Meta's Muse Spark 1.3 and Anthropic's Claude 5.1 series. It also highlights advancements in Voice AI and ML/CV tools, featuring Google's Lyria 3.5, Microsoft's MAI-Image-2.6, and LlamaIndex's Extract Turbo.

Agentic AI

ARC Prize ran OpenAI's Astra twice on the same benchmark. Basic setup: 62.7%, $26,098. A setup that let the model keep its reasoning between steps: 98.6%, $17,332. Same model, same effort setting. Higher score, lower bill. The difference is the harness - the software giving a model its tools, memory and permissions. Agent = model + harness.

Takeaway: A benchmark score belongs to a whole system, not the model on the box. Ask which harness produced it. No answer means no comparison.

Also this week

  1. Meta AI Released Muse Spark 1.3: An Agentic Coding Model That Uses ~20% Fewer Tool Calls and ~25% Fewer Tokens Than Muse Spark 1.2

  2. Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

  3. OpenClaw Releases OpenClaw 2.0: Guided Model Setup, 575 ms Control UI Startup, and One Trust Boundary Per Gateway

  4. Google introduces a new capability to our latest Gemini models: agentic video understanding. This allows developers to process long-form video content with more accuracy, while using up to 88% less tokens.

  5. SpaceXAI Launches Grok Bot Marketplace for Instant AI Teammates

Physical AI

  1. Generalist AI Releases GEN-1.5: A Robot Foundation Model That Learns New Tasks From One 3–12 Second Demo

  2. Axis Robotics Open-Sources One of the Largest Franka Arm Simulation Datasets for Physical AI

AI Infrastructure

  1. Extropic Introduces Z1T Models for Energy-Efficient Thermodynamic AI Chips

  2. Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon

  3. Unsloth AI GLM-5.3-Flash run 3.3x faster locally, Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction.

  4. Just announced at IFA, NVIDIA PAIR automatically links systems across your local network and sends inference requests wherever there’s available capacity, helping agents run more efficiently.

  5. Meet Switchyard: A Rust Proxy and Library That Routes and Translates LLM Traffic Across OpenAI and Anthropic APIs

Voice AI

  1. Lyria 3.5, Google’s best-sounding music generation model, is now available in AI Studio, via the Gemini API, and in the Gemini app

  2. Phonely Launches Alma, a Voice LLM That’s 63% Faster and 84% Cheaper Than OpenAI

  3. Google AI introduces three new voice-activated features that can change the way you use Google Workspace products

  4. Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing

ML/CV/Data Science/OCR

Until next week,
The Marktechpost Team

Sponsorship Opportunities

Did You Like This Newsletter?

Login or Subscribe to participate

Keep Reading

View more