Morning Edition · Sunday, June 21, 2026Published at 6:59 AM EDT · New York
The model can segment every matching instance from a text prompt, and a SAM 3.1 update aims at faster real-time video tracking. Together they move promptable perception past fixed label sets.

Meta's Segment Anything Model 3 (SAM 3) replaces the visual-prompt-only workflow of its predecessors with open-vocabulary prompting. Given a short noun phrase or an image exemplar, the model segments every matching instance across an image…
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
The Global Intelligence Brief stays free.
Part of a tracked trend
Open-Vocabulary, Promptable Vision Foundation Models
Vision foundation models shift to text-promptable, open-vocabulary detection, segmentation, and real-time tracking of arbitrary concepts, generalizing perception beyond fixed label sets across images and video and pushing open perception models toward production use.
Start a discussion in Townsquare.
More from this edition
Comments
0No comments yet.