The Polylog AI Intelligence Brief

Morning Edition · Friday, July 10, 2026

Tech

OpenAI Ships GPT-5.6 Family to the Public and Claims the Top Coding Benchmark

Sol Ultra scores 91.9 percent on Terminal-Bench 2.1 by spawning subagents, but its margin over Anthropic's Mythos 5 is smaller than the variation between two runs of the same model.

3 sources
Plausible

Tech

OpenAI Launches ChatGPT Work, an Agent That Runs Multi-Hour Tasks Across Apps

The product combines Codex into a single desktop app and adds a plugin directory spanning Google Drive, Slack, Salesforce, and GitHub.

2 sources

Geopolitics

US Commerce Department Clears GPT-5.6 for Full Public Release

The model's earlier preview reached only pre-approved organizations under government-imposed access limits. That treated model access itself as a matter of export policy.

2 sources
Corroborated

Tech

Developers Push China's Open-Weight GLM 5.2 Onto Modest Consumer Hardware

The MIT-licensed 753-billion-parameter model scores 62.1 on SWE-bench Pro and 81.0 on Terminal-Bench 2.1, and a community project now targets running it on slow local machines.

2 sources
Corroborated

Tech

Ollama Raises $65 Million as Open-Model Tooling Reaches 8.9 Million Developers

The Series B, led by Theory Ventures, brings the local-inference platform's total funding to $88 million.

2 sources

Tech

Meta's Brain2Qwerty v2 Decodes Typed Sentences From Brain Scans at 61 Percent Word Accuracy

The non-invasive pipeline reads magnetoencephalography signals, raising word accuracy from about 8 percent for prior non-invasive approaches, but it relies on a room-sized shielded scanner.

2 sources

Tech

Survey Charts LLM Theorem Provers Moving From Solving Problems to Doing Research Math

The paper argues language-model provers are shifting from formal proofs of well-defined problems toward open research-frontier mathematics.

1 source

Tech

Theory Paper Asks When Reflection-Driven Reasoning Actually Helps LLMs

Using a sampling-complexity model of generate-critique-revise loops, the analysis characterizes the conditions under which in-context search outperforms plain sampling.

1 source

Tech

DeepSearch-World Trains Tool-Use Agents to Improve From Their Own Experience

The method uses self-distillation inside a verifiable environment to escape the reliance on fixed teacher trajectories and sparse reinforcement learning (RL) rewards.

1 source

Tech

AgentLens Argues Code-Agent Benchmarks Should Grade the Whole Trajectory, Not a Pass/Fail Bit

The production-assessed benchmark reviews the full run of interactive coding agents, targeting the gap between "task passed" and how the agent actually behaved.

1 source

World

EU Parliament's Effort to Kill 'Chat Control' Message Scanning Falls Short

A motion to scrap the interim rule drew 314 votes against 276 to keep it, but rejecting the Council's position required 360, so the scanning regime survives.

2 sources
Corroborated

Tech

ByteDance Releases Seedream 5.0 Pro, an Image Model That Outputs Editable Layers

The professional model can split a finished poster into more than ten independent layers, inpainting the background behind removed subjects.

2 sources