/

 

AI Dev Signals · August 13, 2026

Nobody scaled parameters this week.
Precision, decoding and routing did.

A 1M-hour human-video pre-training ladder. Two frontier updates that are post-training runs on a frozen base. A 30B MoE with 3B active behind an escalation router. And 30B weights quantized into 24 GB at 1.0% degradation.

Presented by Mend.io  Securing AI Agents, MCP Servers & LLM Apps — a practical see–fix–protect framework, with the checklists to start today. Get the guide →

Look at what shipped between August 10 and today. Not one lab announced a larger base model. SpaceXAI froze the Grok 4.5 foundation and spent the delta on a supplemental run, regenerated SFT trajectories and agentic RL. Google’s model card calls 3.7 Flash an algorithmic refinement of the reasoning foundation, not a new pretraining run. NVIDIA shipped a sparse 30B with 3B active, NVFP4 weights, multi-token prediction, and a router that decides when the frontier model is needed at all. Meta reached 24 GB through 4-bit quantization plus a 16-token block-diffusion drafter. LTX rebuilt the decoder and moved motion into an 8x compressed latent. Dyna scaled the pre-training corpus instead, four orders of magnitude of unlabelled human video. Parameter count was not the variable this week. Training stage, numeric precision, decode strategy and routing were.

74%

cost cut by routing

6.8s

for ten seconds of video

24 GB

holds a 30B agent

 

⭐ Featured · Sponsored by NVIDIA

LTX · Open weights world model

The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model

The render farm was the studio door. LTX-2.5 is an open weights world model for video generation, real-time applications and physical AI, optimized for local inference on NVIDIA RTX GPUs and DGX Spark. VRAM requirements are cut far enough that a frontier world model runs on hardware creators already own, and additional clips carry no per-generation fee or metered credit.

The speed number is the one to hold onto. In LTX’s published image-to-video benchmark, a 10-second clip takes 6.8 seconds on-prem running on 2x NVIDIA GB200, and 23.7 seconds through the LTX API. The fastest closed alternatives listed — Omni Flash, Grok 1.5, Veo 3.1 — land at 52 to 70 seconds. Seedance 2.0 is 196, FLUX 3 is 259, Seedance 2.5 is 317, and Kling 3.0 Pro is 398. On-prem, generation finishes faster than the clip’s own runtime.

LTX rebuilt the pipeline rather than bolting features onto an older core. A new diffusion video decoder cuts artifacts in high-motion scenes while holding the high compression ratio. Native multishot renders a full sequence as one output, keeping character, scene and voice consistent across cuts, on a custom Gemma 4 language backbone with a dedicated prompt enhancer. Diffusion Fidelity Rendering builds motion and structure in an 8x temporally compressed latent space, then generates high-fidelity keyframes to anchor detail, with keyframe count adapting to scene complexity and compute budget.

A separate pretrained physical AI checkpoint gives robotics teams a fine-tuning base for non-cinematic domain data, and a stronger distilled model targets production-volume deployment. LTX describes the family as the most used open world model, with more than 33 million downloads. Weights ship on Hugging Face, natively in ComfyUI on day one, and through the LTX API — free for organizations under $10M in annual recurring revenue, with full control of hardware, customization and IP.

Read the full breakdown →

Read on Marktechpost →  · Announcement  · GitHub  · Hugging Face  · Docs  · NVIDIA local AI series

Disclosure: Marktechpost’s coverage of this release was sponsored by NVIDIA. All performance figures above are LTX’s own published numbers.

🔥 The signals

01 · Dyna Robotics · Robotics

Dyna Robotics Introduces Dyna-2: A World-Action Model Pre-Trained on 1 Million Hours of Human Video

Robot learning has been bottlenecked by action-labelled data, and that data has to be deliberately produced by teleoperation. One arm, one operator, one demonstration at a time. Dyna-2 asks whether video of people doing ordinary things — footage nobody had to stage — can stand in for it.

Pre-training used more than one million hours of egocentric human video, roughly 170 years of continuous waking experience. On a nested ladder from 1,000 to 1,000,000 hours, mean normalized score across 14 post-trained tasks rose 20% → 28% → 45% → 53%. The law transfers: zero-shot action MSE on 39 robot tasks the model never saw fits 0.306·D^-0.0713. Joint video-and-action denoising beat action-only on 39 of 39 tasks at every action scale. No weights and no API — deployment today means buying a vendor-operated Dyna robot cell.

Read on Marktechpost →  · Technical report  · Announcement  · xdof ABC eval set

02 · SpaceXAI · Frontier model

SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work

The base model did not change. SpaceXAI held the foundation constant and spent the improvement on a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning in agentic environments. A generational bump without a new pre-training bill.

It scores 61 on the Artificial Analysis Intelligence Index, up from 56, tied with GPT-5.6 Sol Max. 500K context, text and image in, and a new xhigh reasoning-effort level. Read the losses first: DeepSWE v1.1 lands at 65.9% against 73% for GPT-5.6 Sol Max, and Terminal-Bench v3.0 at 26% is last of the four models listed. The bolded wins on GDPval-AA v2 and AA-Briefcase sit inside Artificial Analysis’ confidence intervals, and Claude Opus 5 is absent from the comparison set. Pricing is $2 / $6 per 1M below 200K prompt tokens, doubling above it. No open weights.

Read on Marktechpost →  · Announcement  · Docs  · Release notes

03 · Google DeepMind · Model pricing

Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

The same move, one day later, at a different lab. The model card describes 3.7 Flash as a refinement of 3.6 Flash with algorithmic improvements to the core reasoning foundation — not a new pretraining run. Three weeks separate the two releases.

The coding gains are real. FrontierCode 1.1 Main goes from 34.4% to 43.6%, DeepSWE v1.1 reaches 65.3%, and WebDev Arena posts 1588 Elo against 1538. AutomationBench moves from 17.0% to 30.4%, ahead of Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6%. But price is the argument: $0.75 / $3.75 per 1M tokens, introductory until December 31, 2026, then $1.50 / $7.50. GPT-5.6 Terra still leads DeepSWE, Terminal-bench and OSWorld-2.0, and CharXiv Reasoning regressed slightly to 84.5%. API and enterprise only — nothing to self-host.

Read on Marktechpost →  · Announcement  · Model card  · API docs

Sponsored · Mend.io

Securing AI Agents, MCP Servers & LLM Apps: A Practical Framework

Agents, MCP integrations, and LLM-powered apps are entering codebases faster than security programs can track them — and their behavior isn’t defined by code alone. This guide gives AI/ML engineers, platform engineers, and AppSec teams a practical see–fix–protect framework, with the checklists and templates to start today.

Inside: the agentic AI attack surface map across five layers, from prompt injection to poisoned MCP tools. A 12-point misconfiguration checklist covering agent and MCP permission and config checks. Triage at AI speed — what to automate, what to keep human. Runtime guardrail architectures, prompt hardening, and policy enforcement for AI in production. Plus an agentic AI maturity roadmap: a self-assessment aligned to NIST AI RMF, OWASP AIMA, ISO/IEC 42001, and the EU AI Act.

Get the framework →

Download the guide → · Our writeup of the framework

04 · NVIDIA · Agent infrastructure

NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router

Long-running agents spend most of their token budget on tool calls, result validation, and subagent delegation. Every one of those steps goes to a frontier reasoning model by default. Two artifacts shipped together to break that default: a small model for the execution layer, and a router that decides when the expensive model is actually needed.

Lightning is a 30B mixture-of-experts with 3B active parameters on a hybrid Mamba-2 + MoE + Attention stack, 1M-token context, pre-trained on more than 20 trillion tokens with an NVFP4 recipe. NVIDIA reports up to 4x output speed against similar-sized models, and 86% accuracy on PinchBench while completing 10,000 tasks 30% faster than Qwen3.6 35B. Switchyard is the other half: LangChain benchmarked 145 multi-turn agentic tasks, and the escalation router cut cost 74% versus a frontier-only baseline while sending 7% of calls to the frontier model, at roughly a 6-point accuracy tradeoff. OpenMDW-1.1, commercial use permitted, single-GPU on 1x DGX Spark or 1x H100.

Read on Marktechpost →  · Hugging Face  · Switchyard GitHub  · NVIDIA blog  · LangChain benchmark

05 · Meta AI · Local agents

Meta AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GPU

A 30B model needs over 55 GB at full precision, which is a datacenter conversation. Meta compressed it to roughly 4-bit and added block-level speculative decoding, so it answers fast enough to sit inside a real agent loop on one card, with no network call. Apache 2.0, distilled from Muse Spark.

Two quantized builds ship: K-Quant-Dynamic targets 32 GB VRAM at 0.2% average degradation, K-Quant-17GB targets 24 GB at 1.0%, averaged across 15 common benchmarks. Speed comes from DFlash, a block-diffusion drafter predicting 16 tokens per forward pass — on an RTX 5090 throughput moves from 74.9 to 233.4 tok/s, a 3.1x speedup; Apple M5 Max goes 26.6 to 50.2. Against Gemma4-31B and Qwen3.6-27B it leads MCP Atlas at 75.5 versus 54.2 and 62.5, DeepSearch QA at 74.6, and SWE-Bench Pro at 51.2. It trails Qwen3.6-27B on OSWorld-Verified (65.9 against 75.6) and TerminalBench 2.1. The pattern holds: wins on agentic orchestration, losses on computer-use and terminal work.

Read on Marktechpost →  · Meta blog  · Hugging Face  · Model details  · DFlash paper

 

From Marktechpost · our own repo
Same theme, smaller scale: Token Saver keeps a 200-page PDF out of the context window and hands the model a page-cited slice instead. MIT, one-click .mcpb for Claude Desktop, retrieval recall@5 of 0.90 on 30 gold questions, and the eval you can re-run yourself. Star it on GitHub →

One more thing…

Xiaomi MiLM Plus · Evaluation

Xiaomi’s MiLM Plus Releases PROVE: Perception-Aligned Object Removal Metrics RC-S and RC-T With a Real-World Video Benchmark

Every story above made something cheaper. This one asks whether you would notice if cheaper made it worse. Object removal models have improved faster than the metrics used to judge them, and the reason is structural: erasure is ill-posed and one-to-many, so no single ground truth exists to compare against.

The consequences are specific and uncomfortable. PSNR, SSIM and LPIPS assume point-to-point correspondence, so they reward copy-paste over genuine erasure, and residual shadows occupy too few pixels to be penalized. Cutting diffusion inference steps improves PSNR and SSIM while visual quality collapses. On ROSE-Bench, progressively blurring the masked region degrades neither ReMOVE nor CFD — both eventually score better than their own unblurred baselines.

PROVE’s answer is local distribution matching instead of global aggregation. RC-S and RC-T score only the edited region, using sliding-window Maximum Mean Discrepancy over DINOv2 features, and neither needs a reference video. Against human rankings from 20 participants, RC-S reaches 0.59 average Kendall’s τ and 0.66 Spearman’s ρ, against 0.26/0.29 for ReMOVE and 0.16/0.18 for CFD. It prefers the clean image over blurred and region-swapped variants in 100% of RORD-Val cases, runs at 134.6 ms/frame on one RTX 4090, and is 13.7x cheaper than CFD. Apache 2.0, accepted at ACM MM 2026.

Get the code →

Read on Marktechpost → · Paper · Project page · PROVE-Bench dataset

 

Sponsor AI Dev Signals

Put your product in front of the people who build with it.

This list is engineers, researchers, and founders who read a release note before they read a press release. No banner farms, no interstitials. One sponsor per issue, written in your voice, placed where people are already reading.

Newsletter placements, article sponsorship, product launches, GitHub and Hugging Face repo promotion, and webinars. Tell us the goal and we will send the media kit and available dates.

Book a placement →

See one in action — this issue’s placement for Mend.io →

🚀 Join the community

Reddit  ·  X  ·  LinkedIn  ·  Telegram

 

AI Dev Signals · by Marktechpost AI Media

Every number above is taken from the primary source and the Marktechpost coverage linked in each story, August 2026 — versions, access tiers and availability change fast.

© 2026 Marktechpost AI Media Inc. All rights reserved.

Keep Reading

View more