Revealing the Hidden Gaps: Benchmarking Multimodal Language Models in Scientific Visualization Literacy

Recent advancements in artificial intelligence have brought about multimodal large language models (MLLMs) capable of interpreting both visual and textual data. However, a groundbreaking study from researchers at the University of Notre Dame uncovers significant limitations in these models' understanding of scientific visualization (SciVis). By benchmarking six MLLMs using a rigorous standard assessment, the study highlights that despite some strengths, many models fall short of human-level comprehension.

The SciVis Literacy Benchmark Introduction

The scientific visualization literacy assessment test (SVLAT), a core tool in the study, evaluates how effectively these models can interpret a variety of scientific visualizations. This standardized test includes 49 items based on 18 different visualizations and spans eight techniques and eleven task types. The assessment aims to uncover each model's strengths and weaknesses regarding SciVis literacy.

Results Overview: The Good and the Bad

Among the six models tested, Gemini emerged as the clear frontrunner, surpassing human averages in performance. It scored an impressive 90.9% on static images and 82.9% on animations, significantly higher than many of the open-source models. Conversely, these open-source options, such as Qwen, LLaVA-OneVision, and InternVL, consistently lagged behind human-level performance, revealing inconsistencies across various visualization techniques and tasks.

Where Do Models Struggle?

The study revealed that while MLLMs excelled at tasks involving scientific illustrations or spatial understanding, they struggled with quantitative estimation, texture-based visualizations, and integration tasks. For instance, models exhibited recurring difficulties in accurately judging quantitative values or interpreting local flow directions in visual data. These challenges underscore the need for enhanced training and evaluation metrics to improve MLLM capabilities in SciVis.

Conclusion: A Call for Focus on SciVis Literacy

The findings of this research position SciVis literacy as a crucial but often overlooked benchmark for assessing multi-modal AI systems. As MLLMs continue to evolve, understanding sensory-rich scientific data requires a unique approach that accounts for the complexities inherent in spatial and temporal reasoning. This highlights the importance of dedicated frameworks and assessments, like SVLAT, in building trustworthy AI capable of supporting scientific communication and analysis in the future.

With the continual growth of AI, focusing on developing these essential understanding capabilities may revolutionize how researchers and the public interact with complex scientific visuals.

Authors: {Patrick Phuoc Do, Chau M. Ta, Chaoli Wang}