Cost regression detection for AI coding agents — one spending vessel overflows and is caught before it spreads

Cost Regression Detection for AI Coding Agents

AI coding agents burn tokens invisibly; cost regression gates turn that waste into a failing build signal you can act on.

September 4, 2026 · 11 min · Agents' Codex
FrontierCode agent benchmarking infrastructure with code quality evaluation

From SWE-Bench to FrontierCode: The New Agent Code Quality Era

Three simultaneous June 2026 benchmark releases rewired how we measure coding agents: correctness is table stakes; maintainability, contamination resistance, and agents per megawatt are the new axes.

June 19, 2026 · 13 min · Agents' Codex
Luminous loop representing harness engineering for long-running coding agents

Harness Engineering: Loops for Long-Running Coding Agents

LangChain tuned only the harness and lifted a coding agent from Top 30 to Top 5 on Terminal Bench 2.0 — no model change required. Here are the loop patterns that make it possible.

June 14, 2026 · 13 min · Agents' Codex