Morning Edition · Tuesday, August 25, 2026
Tech
Nvidia Puts Groq 3 LPX Into Full Production and Claims 30x More Agent Throughput Per Megawatt
The company says each liquid-cooled rack holds 256 language processing units, with the cloud provider Nebius as the first customer and racks online before the end of 2026.
Tech
SemiAnalysis Benchmark of Real Coding-Agent Traffic Puts Nvidia About Five Times Ahead of AMD on Cost
The gap comes mostly from serving software, and the analysts say Nvidia would still be cheaper per token even if the competing hardware were free.
Tech
OpenAI Cuts GPT-5.6 Sol Prices to $4 and $20 Per Million Tokens and Guarantees the Rate to November
The reduction is 20 percent on input and 33 percent on output, and it arrives the same week the model family reaches Amazon's Kiro development environment.
Tech
An Anonymous Model With a One-Million-Token Context Consumed 11.6 Trillion Tokens on OpenRouter in Three Days
Listed only as "Stealth" and free during its preview period, Ox Alpha has prompted tokenizer fingerprinting attempts that point toward China's Z.ai, though no lab has claimed it.
Tech
Xiaomi Shows a 150-Watt Desktop That Runs 120-Billion-Parameter Models on Three In-House Chips
The AI Cube prototype pairs the Xring O3, O100 and D100 with up to 160 gigabytes of unified memory and 1.22 terabytes per second of near-memory bandwidth.
Tech
Mathematicians Open a Registry That Machine-Checks AI-Generated Proofs Before Anyone Cites Them
Palomar first runs a mechanical Lean check for hidden axioms, then uses a language model to test whether the formal statement matches the informal claim.
Tech
A New Cache Method Attacks the Prefill Cost That Prefix Caching Cannot Reach
KVBoost reuses key-value tensors at the chunk level and recomputes only where the deviation is large, targeting prompts that share content but not a leading prefix.
Tech
Three New Papers Test Whether AI Agents Can Actually Do Digital Forensics
One benchmark measures whether agents can infer relational structure in undocumented mobile databases, another searches for personal data hidden inside binary fields, and a third checks which threat feeds arrive before attacks.
Tech
Meta's Open Vision Models Cut Beamline Analysis From a Month to Fifteen Minutes at Berkeley Lab
Segment Anything 3 and DINOv3 run on 300 Nvidia A100 accelerators at the national computing center, segmenting X-ray and neutron imaging data in real time.
Tech
Two New Benchmarks Test Languages the Frontier Evaluation Suites Skip Entirely
One targets emotion, sarcasm and cultural reasoning in Nigerian Pidgin, the other retrieval in Khmer, where word boundaries are ambiguous and multilingual embeddings perform poorly.
Tech
Clinicians Start Advising Patients to Consult AI Before Appointments
The emerging pattern places models on both sides of the consultation, as patient preparation and as a clinician's second opinion, with no regulatory framework covering either use.
Tech
A New Survey Catalogues Model Collapse as Labs Formalize Content Provenance
The review examines what happens when generative models train on AI-synthesized data, arriving as Anthropic explains how it watermarks Claude's text output.