Hey there, This issue covers JetBrains' Mellum2.1, a 12B MoE coding model that went from 2.0 to 47.0 on SWE-bench Verified through RL in real repositories, alongside updates across Agentic AI and AI Infrastructure, including Google Cloud's Gemini agent and Underdog's Saluki 27B. It closes with 5 Signals: lithos-metal, Odyssey-3, Firecrawl Universal Scrape, LightOnOCR-3, and Extend Operator-1.
➤ but wait one more thing: Nebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes Run with NVIDIA (sponsored).
Trends from X

AI on X this week: hardware led, agents filled the middle. The 5 AI and tech stories on X’s News tab drew about 52K posts combined. Microsoft’s Surface Laptop Ultra with RTX Spark led with 23.9K posts, roughly 46% of the total. Agent stories took the next 2 spots. Nous Research’s $90M raise at a $1.5B valuation for Hermes Agent reached 12K posts. Google Cloud’s universal Gemini AI agent for businesses followed at 8,978. OpenAI’s Decisions API drew 7,223. A global AI challenge on Alzheimer’s research, built around a 150-million-cell atlas, drew 132 posts.
Agentic AI
▲ JetBrains released Mellum2.1, a 12B MoE open model for coding agents. 2.5B parameters fire per token. The architecture is unchanged from Mellum2. The gain came from reinforcement learning in real software environments. SWE-bench Verified went from 2.0 to 47.0. Qwen3.5-9B scores 50.0 in the same pipeline. LiveCodeBench v6 is 82.0, top of the tested group, with Qwen3.5-9B at 75.4 there and 65.6 on Qwen's own card. Terminal-Bench 2.1 is 17.4 against 21.7, still the weak spot. Qwen3.5-9B also leads SWE-bench Pro (38.0 vs 28.0), GPQA Diamond, and AIME. JetBrains claims about 2× Qwen3.5-9B throughput under load on 1 H200. All scores are JetBrains', run through 1 shared pipeline. Text only. Apache 2.0 on Hugging Face, vLLM or SGLang.
Also this week
▲ Google Cloud Launches Gemini Agent: One Universal Agent for Enterprise Work 1 prompt box, 1 API: questions, knowledge work, media, and code execution. The agent is the product; the model is a routing decision. Today it routes across Gemini and Claude. Other models are "planned." 4 memory types persist across tasks: session, semantic, procedural, episodic. Coworker agents get their own Workspace identity with email, calendar, and Drive. Per-project spend caps pause an agent until a human resumes it. Reachable from web, CLI, Workspace, Microsoft 365, and Slack.
▲ The web is the one API most agents are missing. Databases, calendars and repos have APIs. The open web mostly doesn't. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free. The server is on GitHub. We're racing to 300,000 users this October, with 30% extra on every top-up. [Sponsored]
AI Infrastructure
▲ Meet the Underdog Saluki 27B: A 2-Bit Qwen3.8-27B That Beats the Original at Tool Calling Conway Research compressed Qwen3.8-27B to 7.89 GB, from 54 GB in BF16. The pass was tuned to keep tool calling. On Underdog Bench, 120 BFCL v4-derived tasks, Saluki scores 88 to the parent's 84. Parallel tool calls go 42 to 35 of 100. Then the costs. AIME 2025 drops from 96.7 to 79.2. SWE-bench Verified fixes 30 of 50 against 33. MuSR and MBPP+ slip too. Underdog claims 96% average retention across 9 benchmarks. The tool-calling sets are small and Underdog's own.
ML/CV/Data Science/OCR
▲ Datalab released OmniExtractBench, an open benchmark for structured extraction from PDFs. It pools 620 documents from 4 sources, so no single vendor picked the set. 33 documents run past 100 pages. 1 deterministic scorer assigns 6 verdict types per value, including fabricated and invented_field, with an explanation for each. Table rows are matched by content with the Hungarian algorithm, so 1 missed row does not cascade. On the launch run, Datalab's accurate mode scores 93.85% and competitors cluster near 93.5%. Apache 2.0..
Signals worth Catching
▲ LithosAI's lithos-metal generates Metal megakernels for Apple silicon, 200+ tok/s on Qwen3.8-27B on an M5 Max, Apache 2.0.
▲ Odyssey-3 is a world model scoring 66.1 on Physics-IQ Verified video-to-video, research preview only.
▲ Nebius Opens 2026 Physical AI Awards: Five $150K Compute Prizes Run with NVIDIA (sponsored).
▲ Firecrawl Universal Scrape reads web pages, PDFs, images, and Office files from 1 /scrape endpoint, plus 200+ providers.
▲ LightOnOCR-3 adds layout, grounding, and chart extraction in 4B and 0.8B sizes, 86.3 on olmOCR-Bench.
▲ Extend's Operator-1 is an extraction agent scoring 99.88% on LongExtractionBench, 5 credits per page.
Until next week,
The Marktechpost Team
Sponsorship Opportunities


