Are We Ready for Multi-Image Reasoning? Launching VHs: The Visual Haystacks Benchmark!
View original at bair.berkeley.eduAre We Ready for Multi-Image Reasoning? Launching VHs: The Visual Haystacks Benchmark! <!-- These are comments in HTML. The above header text is needed to format the title, authors, etc…
O que extraímos desta fonte
The claims Via News extracted from this document. We point to the source; we don't replace it.
Simple captioning (LLaVA) combined with LLM aggregator (Llama3) outperforms all LMM-based methods with 5+ images, demonstrating current LMMs are inadequate for cross-image information integration
80% confidenceVisual domain exhibits Lost-in-Middle phenomenon analogous to NLP, with LLaVA performing best with needle before question and proprietary models preferring needle at start
80% confidenceMIRAGE retriever significantly outperforms CLIP on question-like text retrieval without efficiency loss
80% confidenceAll evaluated models show significant performance falloff as haystack size increases, with proprietary models failing above 1K images due to API payload limits
80% confidenceVisual Haystacks is the first visual-centric NIAH benchmark, compared to prior text-based OCR retrieval approaches
80% confidence
