Home
Promotions
🔁 Agentic Edition: Agent runs you can rewind, MCP tools over zero-trust P2P, and a tool caller that fits in 28MB of RAM

Aug 20, 2026

•

5 min read

🔁 Agentic Edition: Agent runs you can rewind, MCP tools over zero-trust P2P, and a tool caller that fits in 28MB of RAM

A mesh that authorizes offline, a harness where the control loop is a plugin, a browser with no Chromium. AND: Shepherd forks the process and filesystem together, 5× faster than Docker.

🚨 96.8% of KernelBench taken from torch.compile, a 300M drafter at 3.18x, and HF checkpoint → C++ in two commands

Aug 20, 2026

•

6 min read

🚨 96.8% of KernelBench taken from torch.compile, a 300M drafter at 3.18x, and HF checkpoint → C++ in two commands

A checkpoint reaching native C++ without ONNX, OIDC claims sealed into Biscuit tokens, and a 300M drafter running nine tokens ahead. The honest range on that last one is 1.04x to 3.18x, because speedup tracks acceptance.

Aug 13, 2026

•

7 min read

🚨 Muse Glimmer's 16-token block drafter, NVFP4 weights on one H100, and video in an 8× compressed latent

3B active out of 30B. 4-bit weights in 24 GB at 1.0% degradation. An escalation router sending 7% of calls upstream. Six releases, zero new base models, and one paper arguing the metrics underneath all of it are broken.

Aug 11, 2026

•

5 min read

🚨 Meta's 4-bit 30B in 24 GB of VRAM, NVIDIA's 448 ms full-duplex turns, and an agent runtime that forks 5× faster than Docker

Six releases, and not one of them was decided by a benchmark. VRAM ceilings, license text, and deployment targets did the deciding. Meta got a 30B under 20 GB. webAI shipped 1.78 GiB you may not sell. NVIDIA published permissive weights and tagged them research-only.

Aug 9, 2026

•

6 min read

🚨 NVIDIA's agent-in-one-Python-class, a 2.6B agent running on your phone, and a harness that beat the human baseline on ARC-AGI-3

Mistral's guardrail that takes the policy as a prompt, Microsoft's test agent at 92.1% vs 78.9% on the same model, and NVIDIA collapsing an agent into one Python class. This week the differentiation lived in the substrate, not the weights.

Aug 6, 2026

•

6 min read

🚨 Prime Agent hits 95.5% on ARC-AGI-3, Cursor's MoE megakernel at 2.37×, and a charting library shipping 10M points in 258 KiB

A skill trained in Codex outscored the one Claude Code trained for itself — 81.8 against 80.4. And a deterministic MoE megakernel, a 34B open VLA, and 100 million points rendered in 0.081s.

Aug 5, 2026

•

6 min read

🚨 Cursor's MoE megakernel, Qwen3.8-Max at 2.4T, and a charting library that ships 10M points in 258 KiB

Cursor open-sources its MoE training megakernel. Alibaba ships Qwen3.8-Max at 2.4T parameters. Reflex releases XY. GenOffice, pixel-native RAG, SkillSpector auditing, PerceptionBench evals, and YC's agent harness

Aug 3, 2026

•

6 min read

🚨 The 276B AI model that beats its 975B teacher — on one GPU

Qwen3.8-Max ships as an API today, open weights next week — and the 27B is the real on-prem path. Cogent VR-1, Ontology 1, AMD's fully open MoE recipe, Supabase Evals, and a GeoAI pipeline you can run end to end.

Aug 1, 2026

•

4 min read

🚨 DeepSeek V4-Flash 0731, MiniMax H3, Supabase evals

Same architecture, better post-training. Native stereo video. And the benchmark that caught Codex reading 4× more docs than Claude Code.

Jul 30, 2026

•

4 min read

🚨 Six releases. One theme. The machine stopped waiting

Most weeks in AI are noise. This one is not. Every release below removed a wait — a robot waiting for a script, a caller waiting for a transcript, a GPU waiting for its slowest neighbor.

Jul 28, 2026

•

5 min read

📡 This Week: MAI-Cyber-1, AgentENV, and a Flow Model That Drives Robots

Microsoft pushes a cyber model to 95.95% on CyberGym, Moonshot open-sources the RL sandbox behind Kimi K3, Perplexity puts search in the terminal, and Black Forest Labs unifies image, video, audio, and robot actions in one flow model. The cyber race and multimodality both leveled up.

Jul 26, 2026

•

3 min read

📡 This Week: Photon-1, Fugu-Cyber, and the Agent That Escaped Its Box

A model learns to use a computer from 18 years of video, Sakana tops real cyber-defense benchmarks, a frontier world model gets a full open reproduction — and World models and security, both getting real.

Jul 24, 2026

•

5 min read

📡 New Opus, a Desktop Coworker, and Two Benchmarks Worth Reading

Anthropic ships Opus 5 at unchanged pricing, Andrew Ng releases a desktop coworker that hands you finished work, and a Rust tokenizer hits 24.53 GB/s. And the open OCR and ASR fields get honest benchmarks.

Jul 22, 2026

•

6 min read

📡 This Week: The Big Three Ship — Cisco Security, Google Flash, NVIDIA Cosmos 3 Edge

Cisco releases security small language models that hunt bugs, Poolside ships an agentic-coding MoE, Google's Flash tier gets cheaper and a Cyber variant, and NVIDIA puts a world model on a robot

Load more

The newsletter platform built for AI Devs

© 2026 Marktechpost AI Media Inc.
beehiivPowered by beehiiv