Polylog
The Polylog AI Intelligence Brief

Morning Edition · Monday, July 27, 2026

Moonshot AI Releases Kimi K3 Weights, a 2.8-Trillion-Parameter Open Model

Tech

Moonshot AI Releases Kimi K3 Weights, a 2.8-Trillion-Parameter Open Model

The Chinese lab is publishing downloadable weights for a mixture-of-experts model that it says outperforms proprietary US systems on several coding and agent benchmarks. The company reported those benchmark figures itself.

3 sources
Corroborated
Anthropic Ships Claude Opus 5, Priced at Half Its Largest Model

Tech

Anthropic Ships Claude Opus 5, Priced at Half Its Largest Model

Opus 5 lists at $5 per million input tokens and $25 per million output tokens. On Anthropic's own numbers, it roughly doubles its predecessor's score on an agentic-coding benchmark.

3 sources
Nvidia Uses Its Own Vera CPU to Speed Up Chip Design by 1.5 Times

Tech

Nvidia Uses Its Own Vera CPU to Speed Up Chip Design by 1.5 Times

The company says pairing Vera with AI agents sped up Cadence Jasper verification and Synopsys VCS simulation, two of the most compute-intensive steps in early chip design.

2 sources
OpenAI Puts ChatGPT Voice in the Desktop App to Drive Codex Agents

Tech

OpenAI Puts ChatGPT Voice in the Desktop App to Drive Codex Agents

The GPT-Live layer, which can listen and speak at the same time, now launches and directs several coding agents at once from a single spoken instruction on Mac and Windows.

3 sources
SenseTime Open-Sources SenseNova-Vision, a Unified Perception Model

Tech

SenseTime Open-Sources SenseNova-Vision, a Unified Perception Model

A single model handles detection, segmentation, depth, keypoints, and 3D reconstruction as multimodal generation, replacing the usual collection of task-specific components.

3 sources
Corroborated
Meta's Brain2Qwerty v2 Decodes Typed Sentences at 61% Word Accuracy

Tech

Meta's Brain2Qwerty v2 Decodes Typed Sentences at 61% Word Accuracy

The non-invasive pipeline, based on magnetoencephalography (MEG), is a large improvement over earlier surgery-free methods, but it remains far from the below-2% error rate of surgical implants.

3 sources
FlowEvo Lets LLM Agents Co-Evolve Workflows and Reusable Skills

Tech

FlowEvo Lets LLM Agents Co-Evolve Workflows and Reusable Skills

The method turns useful reasoning procedures used at inference time into a library of executable skills, so that agents do not have to rediscover the same solutions on every task.

1 source
CARE Proposes Pre-Execution Verification for Shell-Executing LLM Agents

Tech

CARE Proposes Pre-Execution Verification for Shell-Executing LLM Agents

The paper studies command-level mediation as a runtime control point, checking each shell command before an agent runs it.

1 source
Paper Uses Reinforcement Learning to Optimize Stylistic Jailbreaks of Vision Models

Tech

Paper Uses Reinforcement Learning to Optimize Stylistic Jailbreaks of Vision Models

Adversarial Style Optimization uses a reinforcement-learning method called Group Relative Policy Optimization (GRPO) to train stylistic triggers, aiming to make jailbreaks of multimodal models more consistent than content-based attacks.

1 source
Florida Pastor Sues OpenAI and Sam Altman Over ChatGPT Medical Advice

Tech

Florida Pastor Sues OpenAI and Sam Altman Over ChatGPT Medical Advice

The complaint alleges that ChatGPT-4o described dangerous symptoms as minor hours before a near-fatal pulmonary embolism, and it accuses the company of the unauthorized practice of medicine.

3 sources
Meta Opens a Paid Model API With Muse Spark 1.1 and Ships Muse Image and Video

Tech

Meta Opens a Paid Model API With Muse Spark 1.1 and Ships Muse Image and Video

Meta Superintelligence Labs placed Muse Spark 1.1 behind a paid API at $1.25 per million input tokens and $4.25 per million output tokens, days after releasing its first image and video generators.

3 sources
Researchers Propose a Consensus Framework for Ranking LLMs on Open-Ended Tasks

Tech

Researchers Propose a Consensus Framework for Ranking LLMs on Open-Ended Tasks

The method targets situations where several answers are acceptable and correctness alone cannot distinguish the quality of responses.

1 source