Hey there, This issue covers Anthropic's Claude Opus 5.5 alongside updates across Agentic AI, including OpenAI's GPT-6 Sol and Luna, NVIDIA's SoL-Pi and Nokia's AnyJev. It also highlights Voice AI, featuring Kyutai's Voice of Reason and SpeakON's MagSafe voice button.

Agentic AI

Anthropic released Claude Opus 5.5, the first model in its Claude 5.5 family. Anthropic says it performs at Fable 5.1 level on most work. It costs $4 input and $20 output per 1M tokens, down from $5 and $25 for Opus 5. Cache reads fall 60% to $0.20. Anthropic puts typical workload savings at 40%. It scores 66.4% on Terminal-Bench 4.0, against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra. GPT-6 Astra still leads on Terminal-Bench-Science and AutomationBench. Anthropic itself says the gap to Fable 5.1 is narrower in real use than the scores suggest. Thinking can no longer be disabled. Cyber and biology requests hit Fable 5.1-class safeguards, with most cyber tasks re-routed to Opus 4.8. In a containment test, it tried to circumvent boundaries about 85% less often than Opus 5. Resources: Model docs · What's new · System card

Also this week…..

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks Sol costs $2/$10 and Luna $0.10/$0.50 per 1M tokens. Luna's output cut is about 58%, not 50%. On DeepSWE v1.1, Sol at max scores 68.8%, 1.1 points behind Fable 5. Cost per task is about 80% lower. Cached input reads get up to 90% off, with a new diagnostics tool that explains misses. Cost comparisons are OpenAI's own. API-only. Resources: Model docs · Prompt caching guide

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49% A research AI searched 152 harness changes across 535 environments. 4 survived: Action Fusion, Online Context Compact, ObservationPack, and an Evidence-Preserving Reducer. On EdgeBench, they cut token traffic 44.7% to 49.0% versus Pi and API cost about 33%. Scores stay near 94% of Pi's. On Terminal-Bench 4, it solves 15 tasks against 18 for Pi. MIT licensed, runs on unmodified Pi 0.85.1. Resources: Paper · Technical blog

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model It borrows Jev's typed interface but reads answers from an open model's next-token distribution. Cyclic shifts remove position bias. A running batch prior removes label bias. On Qwen3-8B BANKING77, the order-flip rate drops from 0.230 to 0.073. Auto-decidable traffic at 5% error rises from 7.7% to 52.0% with 100 to 500 labels. Costs K prefills per decision. Apache 2.0, on PyPI. Resources: PyPI · Levels doc · Full ablations

Sponsored

SpeakON Ships a MagSafe AI Voice Button With Its Own Microphone: A 25 g button that snaps to an iPhone and writes polished text into any open app. It has its own mic and battery, so it works locked and offline. Output is shaped, not verbatim: fillers removed, lists detected, register matched to the destination app. $129 one time, US only. The companion app is free and works without the hardware. Learn more about SpeakON

Voice AI

Kyutai Releases Voice of Reason: A Speech-Native Model That Solves Spoken Math With Reinforcement Learning 2 open-weight 9B checkpoints built on GLM-4-Voice, with no transcription step. SFT plus RL lifts spoken GSM8K from 27.3% to 77.1%. Removing temperature correction crashed accuracy to 12.3%. Spoken TriviaQA dips from 40.6% to 34.0%, mostly from SFT. A cascaded ASR-LLM-TTS pipeline still scores 95.7%. Runs on 1 H100. Resources: Paper · Stitch model

Alibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-Time Interpretation Model That Cuts Average Lag to 2.3 Seconds Across 60 Languages LAAL drops from 2.8s to 2.3s. Understands 60 languages, speaks 29. About $1.54 per hour of speech in and out at Singapore pricing. API only. Sources: Model Studio docs · Announcement on X

Until next week,
The Marktechpost Team

Sponsorship Opportunities

Keep Reading

View more