Edition·17 June 2026Fixed-budget evals understate model capabilities, e-commerce agents top out at 57%, and PreAct compiles successful runs into FSMs for 13x speedups.
Edition·16 June 2026Nemotron 3 Ultra open-sources a 550B hybrid MoE while CODA-BENCH caps top agents at 61% on data-code tasks
Edition·15 June 2026Broken evals day: systematic gender bias, coin-flip LLM judges, and one attempt at global standardization
Edition·13 June 2026US government forces Anthropic to pull Fable 5 and Mythos 5 — a landmark regulatory precedent for frontier models
Edition·12 June 2026Arbor shows tree-search multi-agent holds where single agents collapse — and Apple's fp8 turns out to be emulated
Edition·11 June 2026RAG on mobile NPU, SLO-aware multi-agent orchestration, and science-reproducing agents: AI moves down the stack
Edition·10 June 2026Context compression and reasoning evaluation: two structural research axes on June 10
Edition·9 June 2026Benchmarks everywhere, real performance nowhere: the week AI measures its own limits
Week of·8 June 2026Week of June 8, 2026: autonomous agents, infinite memory, and formal proving redefine AI's operational boundaries
Edition·7 June 2026LLM emergent capabilities explained by task frequency, not scale — and that reshapes training strategy
Edition·6 June 2026AI agents: 2.6% success on real economic tasks, 94% undetected sabotage — today's benchmarks draw a brutal capability frontier
Edition·5 June 2026LLM safety: adversarial co-evolution, Gemini sycophancy audit, and single-layer ZO fine-tuning
Edition·4 June 2026Agents in production: from city-scale mapping to web automation, operational AI is converging on architecture patterns
Edition·3 June 2026AI benchmarks are broken — formal proofs advance while LLM judges diverge from humans
Edition·2 June 2026An open-source 8B beats GPT-5 on strategic multi-agent play — while GRPO tackles lithography masks and Chinese grammar correction.
Edition·31 May 2026MTP baked into GGUF, Apple Silicon inference finally benchmarked properly, and search agents that mostly confirm what they already know.
Edition·30 May 2026Local-first week: voice, heterodox GPU builds, and TTS — edge inference keeps maturing
Edition·29 May 2026Anthropic at $965B, LLM confidence calibration via probe fine-tuning, and size doesn't predict safety guard performance
Edition·28 May 2026Poolside releases Laguna XS.2 under Apache 2.0 while foundational research targets the two core inference bottlenecks: KV cache and sample complexity.
Edition·27 May 2026Memory, self-distillation, and agent aging: three angles on LLM reliability in production
Edition·26 May 2026Logical reasoning: LLMs stall on regime transitions, synthetic research agents match proprietary systems
Week of·25 May 2026Anthropic nears $965B valuation while agentic IT benchmarks cap at 50%: a week that redraws frontier deployment limits
Edition·24 May 2026Claude Code discovers a reasoning algorithm for $40 — cuts compute 70% vs. standard self-consistency
Edition·23 May 2026Diffusion LLMs, AMD 16 GB rigs, and data quality frameworks: the local stack hardens from the ground up
Edition·22 May 2026Federated learning delivers in two real clinical sites — but generalization remains the invisible wall of medical ML
Edition·21 May 2026Google bets on Gemini as universal interface layer with Ask YouTube, Ask Maps, and Universal Cart launching in the same week.
Week of·18 May 2026Week of May 18, 2026: formal reasoning breakthroughs, $1.25B/month compute deals, and the safety benchmark illusion
Edition·15 May 2026OpenAI turns ChatGPT Pro into a personal finance advisor while flooding enterprise verticals with Codex use cases