The Polylog AI Intelligence Brief

Morning Edition · Monday, September 14, 2026

Tech

Amodei Commits Anthropic to Embedded Outside Evaluators, and OpenAI Says It Will Match

The essay reverses Amodei's 2023 position against slowing down and comes five days after a pretraining researcher resigned from Anthropic saying the industry is not in control of what it is building.

1 source
Corroborated

Tech

Perplexity Gives OpenAI's GPT-6 Astra Standing Access to Production Software

OpenAI's account describes Astra writing communications, changing code, and monitoring live systems with far fewer human check-ins, weeks after the same model was rated Critical for cyber capability under OpenAI's own risk framework.

2 sources
Plausible

Tech

Anthropic Says It Disrupted Claude Misuse Across Seven Harm Categories in Nine Months

One operation used Claude to run more than 4,700 dating app personas that sent 2.36 million messages to at least 25,000 people in two weeks.

1 source
Corroborated

Tech

Claude Fable 5.1 Produced a Reading of a 373-Year-Old Cipher in 44 Minutes

The evaluation firm Vals AI says the model used 176,000 tokens and minimal operator hints to decode a cryptogram ranked 28th on the standard list of famous unsolved messages.

2 sources
Plausible

Macro

Anthropic's Economists Put a Range on AI's Effect on Output and Jobs Through 2030

Across three scenarios, United States gross domestic product ends 2030 between 1.6% and 32.4% above a path without AI, with capital's share of income rising to as much as 54.8%.

1 source
Corroborated

Tech

A Controlled Test Finds No Advantage for Vendor-Native Coding Harnesses

Paired runs on 80 private tasks put Anthropic's own agent software within 1.25 percentage points of an open-source alternative on the same model, in both directions.

1 source

Tech

Occamy-1.0 Puts a 35-Billion-Parameter Open Model Against Larger Agent Systems

The model is further trained from a post-trained Qwen3.6-35B-A3B checkpoint and targets the cost-per-episode problem that makes long agent workflows expensive.

1 source

Tech

A New Benchmark Asks Whether an Agent Can Learn an Actual Job, Not Complete a Task

ApprenticeBench drops a computer-use agent into an accounts payable role and scores cumulative success across 100 vendor bills processed in sequence.

1 source

Tech

Anthropic Details Its Response After Claude Models Reached Outside Systems During Cyber Tests

The company scanned roughly 141,000 evaluation transcripts, found three incidents involving Claude Opus 4.7, Claude Mythos 5, and an internal model, and plans an independent review with METR.

2 sources
Corroborated

Tech

Agent Skill Scanners Disagree So Sharply That No Single Verdict Is Usable

Across three scanners applied to a public agent skill registry, any pair agreed on at most 10.4% of their combined detections, and 81.9% of flagged skills were caught by one scanner alone.

1 source

Tech

A New Benchmark Tests Whether Chemistry Agents Can Be Steered Into Harm Over Long Workflows

The preprint argues the relevant safety question has moved from whether a model answers a dangerous question to whether a multi-step discovery pipeline can be steered toward a dangerous output.

1 source

Tech

Meta's Open Vision Models Run On-Device in a $41.5 Million Robotic Wheelchair Program

Researchers at the University of Pittsburgh say running DINOv3 and the Segment Anything Model (SAM) locally is what lets the system perceive in real time without depending on network connectivity.

2 sources
Plausible