Search stories

All stories

Vals Launches With Aim to Set New Standard for AI BenchmarkingOpenAI Releases Framework for Reporting Model MisalignmentVantora (formerly UP.Labs) Raises $100M to Build Startups Focused on Physical AI for IndustryOpenAI Announces Astra for Law for Firm WorkflowsAnthropic Runs Lab for Biology ExperimentsVerbose Prompts Help Vision-Language Models Resist Image CorruptionOpenAI Research Studies How Workers Integrate AI Beyond Traditional RolesResearchers Stress-Test Alignment Midtraining Across 110-Billion-Parameter AI ModelsGoogle DeepMind Launches Institute to Widen Debate on AGIAI Evaluator Forum Releases AEF-1 Standard for Third-Party AI EvaluationsGood Start Labs Trains AI on Railroad Strategy Game to Boost Finance Benchmark ScoresAIUC Raises $40M Series A and Launches Insured Agent Standard AIUC-1
All stories

Verbose Prompts Help Vision-Language Models Resist Image Corruption

According to a study published on ArXiv CS.AI, padding prompts makes vision-language models more robust to image corruption, whereas fine-grained questions increase model fragility.

Cross-Modal Attention as a Frequency Filter

The researchers found that question-conditioned cross-modal attention acts as a spectral filter over image patches. Verbose questions widen frequency support, while fine-grained prompts concentrate attention on fewer visual scales, causing models to drift when corruptions match those spatial frequencies. [1]

Benchmark Gains on 8B Models

Evaluations on Qwen3-VL and LLaVA-OneVision across the GQA and CLEVR benchmarks showed that verbose paraphrasing reduced drift variance by 70 to 81 percent on 8B models, providing measurable accuracy improvements under image corruption. [1]

Sources

  1. 01