Morning Edition · Sunday, June 28, 2026Published at 7:02 AM EDT · New York
Epoch AI and METR ask models to rebuild working programs without any source code, and the best result is a 56 percent solve rate on tasks estimated to take human engineers weeks.
MirrorCode, a long-horizon coding benchmark from Epoch AI and METR, gives a model only a compiled binary it can run, natural-language documentation, and example input-output pairs, then asks it to rebuild the program without ever seeing the…
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
The Global Intelligence Brief stays free.
Part of a tracked trend
Oversight and Evaluation Lag Accelerating AI Capabilities
Over the next 3-6 months, evidence mounts that governance, evaluation, and agent-safety methods are failing to keep pace with capability growth, driving investment in interpretability, agent-manipulation benchmarks, and institutional-reform proposals.
Start a discussion in Townsquare.
More from this edition
Comments
0No comments yet.