This methods review examines evaluation of multimodal reasoning beyond aggregate benchmark accuracy. The organizing question is which measurements distinguish perception, grounding, reasoning, and answer-generation failures. Ten related scholarly sources are synthesized through a decision-centered framework spanning problem definition, mechanism, measurement, evaluation, implementation, and governance. The review does not invent experiments, pooled estimates, or unreported quantitative results. It instead evaluates the strength and transferability of the available evidence, with particular attention to allowing a single leaderboard score to hide qualitatively different failure modes. The resulting framework links technical or empirical performance to explicit use conditions and identifies tests that should precede wider adoption in scientific and public-facing applications of vision-language models.
- Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., & Zhou, J. (2023). Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2308.12966 DOI
- Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., & Wang, H. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2312.10997 DOI
- Hartsock, I., & Rasool, G. (2024). Vision-language models for medical report generation and visual question answering: a review. Frontiers in Artificial Intelligence, 7, 1430984. https://doi.org/10.3389/frai.2024.1430984 DOI
- Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., & Neubig, G. (2022). Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys, 55(9), 1-35. https://doi.org/10.1145/3560815 DOI
- Lu, P., Peng, B., Cheng, H., Galley, M., Chang, K. W., Wu, Y. N., Zhu, S. C., & Gao, J. (2023). Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2304.09842 DOI
- Min, B., Ross, H., Sulem, E., Veyseh, A. P. B., Nguyen, T. H., Sainz, O., Agirre, E., Heintz, I., & Roth, D. (2023). Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey. ACM Computing Surveys, 56(2), 1-40. https://doi.org/10.1145/3605943 DOI
- Raiaan, M. A. K., Mukta, M. S. H., Fatema, K., Fahad, N. M., Sakib, S., Mim, M. M. J., Ahmad, J., Ali, M. E., & Azam, S. (2024). A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges. IEEE Access, 12, 26839-26874. https://doi.org/10.1109/access.2024.3365742 DOI
- Wang, H., Li, J., Wu, H., Hovy, E., & Sun, Y. (2022). Pre-Trained Language Models and Their Applications. Engineering, 25, 51-65. https://doi.org/10.1016/j.eng.2022.04.024 DOI
- Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., & Chen, E. (2023). A Survey on Multimodal Large Language Models. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2306.13549 DOI
- Zhu, D., Chen, J., Shen, X., Li, X., & Elhoseiny, M. (2023). MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2304.10592 DOI
- Journal
- Machine Intelligence & Responsible Systems
- Volume
- 1 (2026)
- Article number
- mi20260003
- License
- CC BY 4.0
