Morning Edition · Friday, July 24, 2026Published at 1:33 AM EDT · New York
Researchers Isolate Why the Muon Optimizer Reaches Grokking Faster Than AdamW
A new analysis separates spectral-norm constraints from orthogonalized momentum to identify the mechanism actually responsible.

A paper on arXiv, titled "The Active Ingredient in Muon's Grokking," revisits an observation that the Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW, meaning the model transitions from memorization to…
Continue the AI Intelligence Brief
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
- 5 AI intelligence signals a day
- Frontier labs, compute, and chips
- Model releases and AI infrastructure
- Source-grounded analysis with confidence labels
The Global Intelligence Brief stays free.
Part of a tracked trend
Training Optimizer Advances
Optimizer research increasingly delivers real reductions in the compute needed to reach a given capability, making the training algorithm a recurring axis of efficiency competition alongside data and scale.
More from this edition
- White House Moves to Redirect $200 Billion in Annual Research Funding Toward AI and Individual Scientists
- South Korea Deepens NVIDIA-Anchored Sovereign AI Buildout at San Francisco Summit
- OpenAI Opens Health in ChatGPT to All U.S. Adults, One Day After a Lawsuit Sought to Block It
- Anthropic Doubles Election-Year AI Policy Funding to $40 Million
- OpenAI Chairman Argues Token Efficiency, Not Training Cost, Decides the Open-Weight Question
- Uber Cuts 10% of Customer-Support Staff and Names AI as the Reason
- Alibaba's Qwen-Audio-3.0-TTS Tops an Independent Speech Leaderboard, but as a Closed API
- New Study Finds the Structured Output Itself Causes Language Models to Hallucinate
- Two Papers Probe the Inside of Mixture-of-Experts Routing
- NVIDIA Places Its First GB300 AI Supercomputer Inside the U.S. Military
- Anthropic Lets Claude Desktop Users Build Agent Skills by Screen Recording