The Polylog AI Intelligence Brief

Morning Edition · Monday, August 31, 2026

Tech

Prime Intellect's Open-Source Harness Lifts Claude Opus 5 to 95.5% on ARC-AGI-3 Without Touching the Model

The gain came from a persistent code runtime and versioned agent memory, not new weights, and the score is self-reported rather than verified by the benchmark's authors.

2 sources
Plausible

Tech

Google Ships Gemini Omni 1.1 Flash to General Availability With 40-Second Scene Extension

The model extends existing footage in 10-second increments using the previous 10 seconds, including audio, as conditioning, and the preview endpoint retires on September 30.

2 sources

Tech

Anthropic Opens a Research Preview of a Standard for AI Agents Operating Lab and Factory Hardware

Early users report a jump in laser stabilization on QuEra quantum computers from 58 percent to 99.3 percent, and an imaging experiment compressed from weeks to a day.

2 sources
Corroborated

Tech

Meta's Open Perception Models Now Run US National Laboratory Beamlines While Its Reasoning Model Sits Behind a Meter

Segment Anything 3 and DINOv3 cut expert annotation of X-ray imaging data from weeks to about 15 minutes on 300 accelerators at Berkeley's supercomputing center.

3 sources
Corroborated

Tech

New Preprint Shows Quantization Can Switch On Hidden Backdoors That Full-Precision Testing Never Sees

Backdoored translation models measured clean at 16-bit precision produced corrupted output in up to 85.02 percent of cases after routine post-training compression.

2 sources
Corroborated

Tech

Researchers Propose Routing Agent Tool Calls by Data Origin to Blunt Indirect Prompt Injection

The method treats every piece of content an agent reads as a source with its own privileges, rather than trying to detect malicious instructions inside the text.

1 source

Tech

Interpretability Researchers Put Superposition on a Formal Footing Using Compressed Sensing

A new paper models feature superposition as sparse recovery through an overcomplete dictionary, giving the field conditions under which a network's features can be read out at all.

2 sources

Tech

Two New Papers Attack the Least Glamorous Bottlenecks in Serving Small Models

One reformulates the output projection as a vector search to relieve memory bandwidth pressure, the other quantizes the recurrent state that linear-attention models use instead of a growing cache.

2 sources

Tech

Anthropic Details the Statistical Watermark It Is Applying to All Claude Text Output

The mark biases word selection using a key and preceding context, applies worldwide with no opt-out, and survives copying but not heavy rewriting or translation.

2 sources
Corroborated

Tech

An Unreleased Claude Improved a Longstanding Bound on the Riemann Hypothesis, and Terence Tao Warns About the Cost

Anthropic reports the model raised the proven fraction of zeta zeros satisfying the hypothesis from 41.6 percent to 67.2 percent, as the Fields Medalist argues faster output may come at the price of training fewer mathematicians.

2 sources
Corroborated

World

MIT Committee Reports That AI Can Credibly Complete Almost Any Undergraduate Assignment It Sets

The report covers essays, proofs, science problem sets and coding work, and recommends every course state an explicit policy on machine assistance.

2 sources
Corroborated

Tech

EngineAI's T800 Humanoid Appears on San Francisco Streets Weeks After Its US Debut in a Robot Combat League

The Shenzhen company's 1.73-meter machine is being marketed simultaneously for factory work and for televised fighting, two very different claims about what it can do.

2 sources
Plausible