Cross-Modal Attention as a Frequency Filter
The researchers found that question-conditioned cross-modal attention acts as a spectral filter over image patches. Verbose questions widen frequency support, while fine-grained prompts concentrate attention on fewer visual scales, causing models to drift when corruptions match those spatial frequencies. [1]
Benchmark Gains on 8B Models
Evaluations on Qwen3-VL and LLaVA-OneVision across the GQA and CLEVR benchmarks showed that verbose paraphrasing reduced drift variance by 70 to 81 percent on 8B models, providing measurable accuracy improvements under image corruption. [1]
Sources
- 01