FrontierCode agent benchmarking infrastructure with code quality evaluation

From SWE-Bench to FrontierCode: The New Agent Code Quality Era

Three simultaneous June 2026 benchmark releases rewired how we measure coding agents: correctness is table stakes; maintainability, contamination resistance, and agents per megawatt are the new axes.

June 19, 2026 · 13 min · Agents' Codex
Luminous loop representing harness engineering for long-running coding agents

Harness Engineering: Loops for Long-Running Coding Agents

LangChain tuned only the harness and lifted a coding agent from Top 30 to Top 5 on Terminal Bench 2.0 — no model change required. Here are the loop patterns that make it possible.

June 14, 2026 · 12 min · Agents' Codex