IGGQ Research Publishing
Machine Intelligence & Responsible Systems

Measuring Multimodal Reasoning Beyond Aggregate Benchmark Accuracy

Read & download PDF
Abstract

This methods review examines evaluation of multimodal reasoning beyond aggregate benchmark accuracy. The organizing question is which measurements distinguish perception, grounding, reasoning, and answer-generation failures. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to allowing a single leaderboard score to hide qualitatively different failure modes. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in scientific and public-facing applications of vision-language models.

Keywords
multimodal learningvision-language modelsreasoningbenchmarksrobustness
References
  1. Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., & Zhou, J. (2023). Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2308.12966 DOI
  2. Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., & Wang, H. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2312.10997 DOI
  3. Hartsock, I., & Rasool, G. (2024). Vision-language models for medical report generation and visual question answering: a review. Frontiers in Artificial Intelligence, 7, 1430984. https://doi.org/10.3389/frai.2024.1430984 DOI
  4. Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2022). Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys, 55(9), 1-35. https://doi.org/10.1145/3560815 DOI
  5. Lu, P., Peng, B., Cheng, H., Galley, M., Chang, K. W., Wu, Y. N., Zhu, S. C., & Gao, J. (2023). Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2304.09842 DOI
  6. Min, B., Ross, H., Sulem, E., Veyseh, A. P. B., Nguyen, T. H., Sainz, O., Agirre, E., Heintz, I., & Roth, D. (2023). Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey. ACM Computing Surveys, 56(2), 1-40. https://doi.org/10.1145/3605943 DOI
  7. Raiaan, M. A. K., Mukta, M. S. H., Fatema, K., Fahad, N. M., Sakib, S., Mim, M. M. J., Ahmad, J., Ali, M. E., & Azam, S. (2024). A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges. IEEE Access, 12, 26839-26874. https://doi.org/10.1109/access.2024.3365742 DOI
  8. Wang, H., Li, J., Wu, H., Hovy, E., & Sun, Y. (2022). Pre-Trained Language Models and Their Applications. Engineering, 25, 51-65. https://doi.org/10.1016/j.eng.2022.04.024 DOI
  9. Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., & Chen, E. (2023). A Survey on Multimodal Large Language Models. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2306.13549 DOI
  10. Zhu, D., Chen, J., Shen, X., Li, X., & Elhoseiny, M. (2023). MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2304.10592 DOI
Publication details
Journal
Machine Intelligence & Responsible Systems
Volume
1 (2026)
Article number
mi20260003
License
CC BY 4.0