Hey there, This issue covers our field guide to decision AI models, 7 systems that return a choice instead of a paragraph, alongside updates across Agentic AI and AI Infrastructure, including NVIDIA's DGX Spark 64GB, Prime Intellect's Prime Inference, and IBM Bob's self-hosted release. It also highlights Voice AI, featuring Microsoft's MAI-Transcribe-2-Streaming, now #1 of 38 on Artificial Analysis. Nebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes Run with NVIDIA (sponsored).
Datalab released OmniExtractBench, an open benchmark for structured extraction from PDFs. It pools 620 documents from 4 sources, so no single vendor picked the set. 33 documents run past 100 pages. 1 deterministic scorer assigns 6 verdict types per value, including fabricated and invented_field, with an explanation for each. Table rows are matched by content with the Hungarian algorithm, so 1 missed row does not cascade. On the launch run, Datalab's accurate mode scores 93.85% and competitors cluster near 93.5%. Apache 2.0, pip install omni-extract-bench.
sponsored
Agentic AI
▲ IBM Brings Bob to Self-Hosted and Air-Gapped Environments. Bob is IBM's agentic software development platform: understand, plan, execute, validate. The self-hosted option is now generally available for on-premises, sovereign cloud, and air-gapped networks. 2 models are supported for full isolation: NVIDIA Nemotron and Poolside Laguna. Hybrid mode routes selected work to Claude Sonnet 5.0, Claude Opus 4.8, Gemini 3.7 Flash, or GPT-5.6 Sol. No model is bundled; you license and host your own. Multi-model routing is a roadmap item. Java, IBM i, and IBM Z modernization need separate packages. No pricing published.
▲ The web is the one API most agents are missing. Databases, calendars and repos have APIs. The open web mostly doesn't. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free. The server is on GitHub. We're racing to 300,000 users this October, with 30% extra on every top-up. [Sponsored]
Also this week…..
▲ A field guide to decision AI models: TypeSafe Jev vs Fastino GLiDE, GLiNER2.5-Decide, and open-source competitors. 7 models, 1 table. Jev 1.13 and GLiDE are closed APIs. GLiNER2.5-Decide (340M), Laya (421M), JevK5 (4B), OpenJev, and kev-0.5b all ship open weights and run locally. On TypeSafe's own workflow evals, Jev scores 67.8%, the same as Claude Sonnet 5. The best LLM rival scores 74.1%. Jev costs $0.0004 per case to Sonnet 5's $0.1174, in 0.4 s against 78.1 s. Per task it swings: 76.0% on customer service, 61.8% on invoice processing. Vercel reports 13% of paid teams tried Jev within 24 hours. Fastino's comparisons use different suites than TypeSafe's.
▲ TinyFish builds web APIs for agents: Search, Fetch, Web Agent, and Browser. Its ambassador program has no application. Sign up and you start as a Referral Partner with $8 in wallet credit. Each qualified referral earns $5 for you and $5 for the new user. Approved builds pay gift cards: $50 for an app, $20 for a Skill, $10 for a prompt. Ambassadors also get beta access, office hours, a direct line to engineering, and a Hall of Fame listing. 50 qualified signups moves you to the Advocate tier automatically. Join the TinyFish ambassador program. [Sponsored]
Physical AI
▲ Nebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes Run with NVIDIA. 5 categories: models, perception and spatial intelligence, simulation and synthetic data, systems and deployment, software and tooling. Each winner gets $150K in Nebius compute credits, about 203 H100 node-days at on-demand rates. Entrants need a Physical AI use case, an MVP in use, a legal entity, and a website. No entry fee. 2025 drew 254 applications and 55 finalists. Applications close October 25. Winners are announced mid-November. [Sponsored]
AI Infrastructure
▲ NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference. Same GB10 superchip as the 128GB Spark, half the memory. 64GB unified LPDDR5x at 273 GB/s, 1 petaFLOP FP4 with sparsity. The target is 30 to 35B models: Qwen 3.8 27B and Nemotron 3.5 Lightning fit on 1 box. 2 units cluster over ConnectX-7 into 128GB pooled memory. NVIDIA says that pair runs 1.7× faster than 1 128GB Spark. The 273 GB/s ceiling limits concurrent decode, so this is for long-input, short-output work, not a chat server. Always-on agents and QLoRA fine-tuning are the pitch.
▲ Prime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models. OpenAI-compatible API on Prime's own Blackwell GPUs, with failover across datacenters. Before launch it served close to 1 trillion tokens per day for Prime's own RL and coding agents. GLM-5.3 runs on GB200 NVL72 with Dynamo, vLLM, Mooncake, and FlashInfer. Splitting prefill from decode cut p90 inter-token latency 40%. NVFP4 KV compression lifts cached tokens per decoder from 1.09M to 1.63M. All figures are Prime's, on agentic workloads. GLM-5.3 pricing is not yet in the docs. Batch inference and 1-click dedicated deploys are on the roadmap.
Voice AI
▲ Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis. Now with numbers. On AA-WER Streaming it ranks #1 of 38 models: 2.5% WER with the final transcript at 0.13 s. First partials hit the same 2.5% at 0.12 s. Grok Voice Transcribe 2.0 is next at 2.7% and 0.49 s. Muse Voice Transcribe scores 3.1% at 0.16 s. 60 languages with continuous auto-detection. $0.54 per audio hour through the end of 2026, above xAI and Meta. Realtime API over WebSocket, Azure Speech SDK, MAI Playground, Vercel, and Azure Voice Live. Public preview, no SLA, no open weights.
ML/CV/Data Science/OCR
▲ Datalab released OmniExtractBench, an open benchmark for structured extraction from PDFs. It pools 620 documents from 4 sources, so no single vendor picked the set. 33 documents run past 100 pages. 1 deterministic scorer assigns 6 verdict types per value, including fabricated and invented_field, with an explanation for each. Table rows are matched by content with the Hungarian algorithm, so 1 missed row does not cascade. On the launch run, Datalab's accurate mode scores 93.85% and competitors cluster near 93.5%. Apache 2.0, pip install omni-extract-bench.
▲ Cohere Releases Embed 5: How It Compares to Voyage 4 Large, Gemini Embedding 2, and OpenAI 2 tiers, 1 embedding space: index with Pro, query with Fast. Both take text, images, and mixed input across 100+ languages, with 128K context. Pro scores 85.8 on ViDoRe V3, ahead of Voyage 4 Large (83.7) and Gemini Embedding 2 (83.2). Gemini wins 9 of 10 non-European languages tested. Fast runs 377.3 documents per second to Pro's 159.7. Pro is $0.12 per 1M text tokens, Fast $0.08, images $0.40. Most of Cohere's table uses RCP-nDCG@10, its own new metric. API, Model Vault, Foundry, SageMaker, or vLLM on-prem.
Until next week,
The Marktechpost Team
Sponsorship Opportunities

