AI Technology

Engineering, infrastructure, and emerging AI systems

Grok 4.6: Benchmarks, Model Card, Safety, and Agent Workflow
August 13, 202613 min read

Grok 4.6: Benchmarks, Model Card, Safety, and Agent Workflow

Grok 4.6 improves long-running coding and knowledge work, while its model card documents stronger search and jailbreak resistance alongside factuality and behavior regressions.

What Is Artificial Intelligence? A Manager's Guide to Deciding
August 10, 20266 min read

What Is Artificial Intelligence? A Manager's Guide to Deciding

A clear answer to what AI is, written for managers: precise definitions, machine learning versus language models, real capabilities and limits today, and a decision framework for adoption.

When the AI Evaluation Escapes: Containing Cyber-Capable Agents
August 1, 202614 min read

When the AI Evaluation Escapes: Containing Cyber-Capable Agents

A practical architecture for evaluating cyber-capable AI agents without giving a benchmark sandbox a transitive path into production systems.

Inside a 53.6x Wan2.2 Speedup
July 31, 202616 min read

Inside a 53.6x Wan2.2 Speedup

How sparse attention, persistent CUDA kernels, four-step distillation, and NVFP4 turn a 133-second pipeline into near-real-time video generation.

Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber: Full Benchmarks
July 21, 20268 min read

Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber: Full Benchmarks

Google's July 21 family pairs a stronger workhorse, a 350-token-per-second volume model, and a restricted cybersecurity specialist.

Kimi K3: From GPT-2 Memory to a 2.8T Open Agent Model
July 16, 202618 min read

Kimi K3: From GPT-2 Memory to a 2.8T Open Agent Model

Moonshot's July 16 open-weight release combines a 2.8T hybrid architecture with strong coding, tool, research, and multimodal benchmarks.

Grok 4.5: Coding Benchmarks, 80 TPS, and Token Efficiency
July 16, 20268 min read

Grok 4.5: Coding Benchmarks, 80 TPS, and Token Efficiency

SpaceXAI's July 16 model posts 83.3% on Terminal-Bench 2.1 and 64.7% on SWE-bench Pro while averaging 15,954 output tokens per task.

GPT-5.6: Sol, Terra, Luna, and the Ultra Agent Tier
July 9, 20267 min read

GPT-5.6: Sol, Terra, Luna, and the Ultra Agent Tier

OpenAI's July 9 family spans flagship, balanced, and efficient models, with state-of-the-art terminal, coding, browsing, and science results.

Robostral Navigate: An 8B Model That Navigates with One Camera
July 8, 20268 min read

Robostral Navigate: An 8B Model That Navigates with One Camera

Mistral's July 8 embodied model reaches 76.6% success on unseen R2R-CE routes using one RGB camera, with a compact architecture and simulation-only training.

Tencent Hy3: A 295B Open MoE with 21B Active
July 6, 20267 min read

Tencent Hy3: A 295B Open MoE with 21B Active

Tencent's July 6 Apache-2.0 model turns preview feedback into a smaller active path, stronger agents, 256K context, and practical product reliability.

Claude Sonnet 5: Near-Opus Agents at Sonnet Economics
June 30, 20267 min read

Claude Sonnet 5: Near-Opus Agents at Sonnet Economics

Anthropic's June 30 model lifts coding, terminal, search, computer use, and knowledge work while exposing effort as a cost-performance control.

Mistral OCR 4: Structured Document Intelligence
June 23, 20267 min read

Mistral OCR 4: Structured Document Intelligence

Mistral's June 23 model adds boxes, block types, confidence, 170 languages, self-hosting, and leading OCR scores—with unusually clear benchmark caveats.