Morning Edition · Tuesday, July 28, 2026Published at 1:31 AM EDT · New York
Vals Launches a Tool That Turns Your Own Codebase Into a Custom Model Benchmark
Vals-Smith auto-generates evaluations from a user's repository, letting teams test models and agents on their own code rather than trusting public leaderboards.

The independent benchmarking platform Vals has introduced Vals-Smith, a tool that automatically generates custom benchmarks from a user's own codebase, according to a report from the AI ML Big Data channel. The aim is to test models and age…
Continue the AI Intelligence Brief
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
- 5 AI intelligence signals a day
- Frontier labs, compute, and chips
- Model releases and AI infrastructure
- Source-grounded analysis with confidence labels
The Global Intelligence Brief stays free.
Part of a tracked trend
Frontier Labs Race on AI Coding Capability
Coding is becoming a primary competitive battleground among frontier labs, with incumbents standing up permanent coding teams and investing in new training stages (e.g. midtraining) to match leaders like Anthropic; expect recurring reorganizations, benchmarks, and model releases aimed specifically at code.
More from this edition
- Moonshot AI Releases Kimi K3, a 2.8-Trillion-Parameter Open-Weight Model That Tops Open Rankings
- China Accuses Washington of 'AI Hegemonism' and Threatens Countermeasures Over Moonshot Probe
- Anthropic Says It Never Sought to Ban Open Weights, but Backs Distillation Crackdowns and Pre-Release Testing
- Anthropic Ships Claude Opus 5, Its Fourth Model in Two Months, With a Doubling on Software-Engineering Evals
- Microsoft Says United Kingdom Grid Connections Take Eight Years, Choking AI Data-Center Buildout
- Nvidia Puts Its Vera CPU to Work Designing the Next Generation of CPUs and GPUs
- Hassabis Says DeepMind Sold to Google Because Independence Would Have Cost Billions It Could Not Raise
- A 184-Million-Parameter Classifier Claims State-of-the-Art Prompt-Injection Detection at a Fraction of Llama Guard's Size
- CORVUS Attacks the Coding Agent's Real Bottleneck: an Append-Only Trajectory That Bloats Context
- FlowEvo Lets Agents Rewrite Their Own Workflows and Skills Instead of Rebuilding Them Each Run
- Meta Puts Segment Anything and DINO Into Assistive Robotics at the University of Pittsburgh