Morning Edition · Wednesday, July 15, 2026Published at 1:32 AM EDT · New York
Trained on 113 million document images across 17 languages, the frozen encoder paired with a 0.7-billion-parameter language model tops MDPBench by 2.8 points absolute.
Researchers from Huazhong University of Science and Technology and Kingsoft Office released MonkeyOCRv2, a visual-text foundation model for document AI built on MonkeyDoc v2, which they describe as the largest document-image pretraining cor…
Track frontier labs, chips, export controls, model releases, regulation, and AI infrastructure.
The Global Intelligence Brief stays free.
Part of a tracked trend
Open-Vocabulary, Promptable Vision Foundation Models
Vision foundation models shift to text-promptable, open-vocabulary detection, segmentation, and real-time tracking of arbitrary concepts, generalizing perception beyond fixed label sets across images and video and pushing open perception models toward production use.
Start a discussion in Townsquare.
More from this edition
Comments
0No comments yet.