The Polylog AI Intelligence Brief

Morning Edition · Thursday, July 9, 2026

Tech

xAI Ships Grok 4.5, a Cursor-Trained Coding Model Priced Below Rivals

The model resolves SWE-Bench Pro tasks with about a quarter of the output tokens Anthropic's Opus 4.8 uses, even as it scores lower on the benchmark itself.

3 sources
Plausible

Tech

Anthropic's Claude Sonnet 5 Pushes Agent-Grade Coding Into a Cheaper Tier

The mid-tier model scores 63.2 percent on SWE-Bench Pro and 80.4 percent on Terminal-Bench 2.1, narrowing the distance to the pricier Opus line.

2 sources
Plausible

Tech

OpenAI Retracts Its Endorsement of SWE-Bench Pro, Citing Broken Tasks

An audit of the 731-task public split flagged 27 percent of tasks as broken by an automated pipeline and 34 percent by human reviewers.

2 sources
Corroborated

Tech

NVIDIA Says Nemotron 3 Ultra Leads Open Models on LangChain's Agent Suite

Tuned inside LangChain's Deep Agents harness, the model scores an aggregate 0.86 at $4.48 per run against $43.48 for the closest closed model.

2 sources
Plausible

Tech

Meta's Brain2Qwerty Decodes Typed Sentences From Brain Scans Without Surgery

The updated non-invasive pipeline reaches 61 percent average word accuracy from magnetoencephalography, up from about 8 percent for prior non-surgical methods.

3 sources

Tech

OpenAI Launches GPT-Live, a Full-Duplex Voice Model for ChatGPT

Two models, GPT-Live-1 and a mini variant, listen and speak simultaneously and delegate hard questions to a frontier model in the background.

3 sources

Tech

New Work Sharpens Memory Poisoning as a Standing Threat to LLM Agents

Persistent agent memory lets an attacker plant instructions through ordinary queries that direct future actions, with prior attacks reporting injection success above 90 percent under lab conditions.

2 sources

Tech

Study Finds Trusted AI Sabotage Monitors Fail to Transfer Across Model Families

Monitors tuned against one or two untrusted models overfit to that lineage's calibration and lose reliability on models they were not tested against.

1 source

Tech

TriRoute Proposes One Learned Router for Attention, Experts, and KV-Cache

The method jointly allocates sparse computation across three axes that prior techniques each optimized in isolation.

1 source

Geopolitics

OpenAI Publishes National-Security Principles as Government Deals Expand

The framework accompanies cyber-defense partnerships with the United States and allied governments and pledges no domestic surveillance of U.S. persons.

2 sources
Plausible

Tech

Cloud and AI Vendors Flood Startups With Compute Credits to Lock Them In

Infrastructure grants now reach into the millions per startup, structured to make later migration to a competitor costly.

1 source
Plausible

Tech

General Intuition, Kyutai, and Epic Open-Source a Playable 5B World Model

MIRA generates Rocket League 2v2 matches frame by frame at 20 frames per second on a single GPU, with dataset and code released.

2 sources