VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
By Rohit Saxena, Alessandro Suglia, and Pasquale Minervini, July 6, 2026
In ICML 2026
Vision-language models are commonly evaluated on clean, high-quality images, leaving their behaviour under real-world distortions poorly characterised. VLM-RobustBench introduces 49 augmentation types spanning noise, blur, weather, digital, geometric, and other perturbations, applied at graded severities or as binary transforms to produce 133 evaluation settings.
The benchmark evaluates 15 models from five families on MMBench and MMMU-Pro, separating visually grounded performance from multimodal reasoning. The results show that apparent visual severity is a weak guide to model difficulty: low-severity spatial changes can be more damaging than conspicuous photometric corruption.
Resampling and geometric distortions cause the largest observed losses, reaching drops of up to 34 percentage points. The findings indicate that current models retain considerable semantic capability but remain spatially fragile, motivating robustness protocols and training methods that target geometric invariance.
Paper: https://arxiv.org/abs/2603.06148
Stay ahead with research-backed solutions
From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.
Book a Consultation