AI Technology

Engineering, infrastructure, and emerging AI systems

GLM-5.2: Open Long-Horizon Agents at 1M Context
June 16, 20267 min read

GLM-5.2: Open Long-Horizon Agents at 1M Context

Z.ai's June 16 MIT release improves terminal, repository, tool-use, and long-running tasks through shared sparse indexing and controllable effort.

Kimi K2.7 Code: Open Weights for Long-Horizon Coding
June 12, 20267 min read

Kimi K2.7 Code: Open Weights for Long-Horizon Coding

Moonshot's June 12 model lifts coding and MCP tool scores over K2.6 while cutting reasoning-token use about 30 percent, with full benchmark footnotes.

DiffusionGemma: Open Text Generation Beyond Next Token
June 10, 20267 min read

DiffusionGemma: Open Text Generation Beyond Next Token

Google's experimental 26B MoE generates blocks in parallel, exceeds 1,000 tokens per second on H100, and exposes the quality-speed trade-off of text diffusion.

Claude Fable 5 and Mythos 5: One Model, Two Boundaries
June 9, 20267 min read

Claude Fable 5 and Mythos 5: One Model, Two Boundaries

Anthropic's June 9 launch exposes the same frontier model through a guarded general tier and a restricted research tier, changing how benchmarks meet access control.

Nemotron 3 Ultra: Open 550B Agents at 1M Context
June 4, 20267 min read

Nemotron 3 Ultra: Open 550B Agents at 1M Context

NVIDIA's June 4 model opens weights, data, and recipes for a 550B MoE with 55B active parameters, long context, NVFP4, and high-throughput agents.

Gemma 4 12B: Native Multimodality in 16GB
June 3, 20267 min read

Gemma 4 12B: Native Multimodality in 16GB

Google's June 3 checkpoint brings encoder-free image and audio input, agentic reasoning, and multi-token drafting to a laptop-sized open model.

MiniMax M3: Open Multimodal Agents with 1M Context
June 1, 20267 min read

MiniMax M3: Open Multimodal Agents with 1M Context

MiniMax combines sparse attention, native image and video input, computer use, and frontier coding in one open-weight model built for long-running work.

Claude Opus 4.8: Stronger Agents, Coding, and Vision
May 28, 20268 min read

Claude Opus 4.8: Stronger Agents, Coding, and Vision

Anthropic's May 28 flagship improves repository work, long-horizon agents, computer use, and visual reasoning, with a system card that exposes the caveats.

Command A+: An Open Enterprise MoE on Two H100s
May 20, 20268 min read

Command A+: An Open Enterprise MoE on Two H100s

Cohere's May 20 Apache-2.0 release unifies reasoning, vision, tools, retrieval, and 48 languages in a 218B MoE with only 25B active parameters.

Gemini 3.5 Flash: Frontier Agents at Flash Speed
May 19, 20268 min read

Gemini 3.5 Flash: Frontier Agents at Flash Speed

Google's May 19 model pairs fast inference with strong coding, tool-use, multimodal, and long-horizon scores—but the harness still defines the result.

Qwen3.7: Max, Plus, and Flash for the Agent Era
May 16, 20268 min read

Qwen3.7: Max, Plus, and Flash for the Agent Era

Alibaba's Qwen3.7 rollout spans a text flagship, a multimodal agent, and an efficient Flash tier—with one family story and several benchmark caveats.

ERNIE 5.1: Smaller MoE, Stronger Agent Benchmarks
May 9, 20269 min read

ERNIE 5.1: Smaller MoE, Stronger Agent Benchmarks

Baidu's May 9 release compresses ERNIE 5.0, rebuilds reinforcement learning, and reports flagship reasoning at a fraction of the pretraining cost.