# Three New Papers Attack the Weak Points of Multimodal Models at Once

Researchers target the security of optical token compression in DeepSeek-OCR, replace static hallucination benchmarks with fuzzing, and catalogue how modality fusion opens new attack surface.

- Published: 2026-08-11T06:26:24.055Z
- Canonical: https://polylog.news/ai/2026-08-11/three-new-papers-attack-the-weak-points-of-multimodal-models
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv (Adversarial Attacks on Deep OCR Systems)](https://arxiv.org/abs/2608.07636), [arXiv (Unified Hallucination Fuzzing for MLLMs)](https://arxiv.org/abs/2608.07525), [arXiv (Evolving Safety Landscape of Multi-modal LLMs)](https://arxiv.org/abs/2608.07535)

Optical compression has become a standard technique for long-context document work: render text as an image, feed it to a vision encoder, and pay far fewer tokens than the character-level equivalent. DeepSeek-OCR made that approach popular.…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-08-11/three-new-papers-attack-the-weak-points-of-multimodal-models (subscription information: https://polylog.news/pricing).