Semantic Specificity Degradation in Zero-Shot Generative Vision-Language Models
DOI:
https://doi.org/10.15294/edukom.v13i1.44852Keywords:
Hierarchical Semantics, Multimodal Evaluation, Semantic Specificity, Vision-Language Models, Zero-Shot LearningAbstract
Vision-Language Models (VLMs) have shown strong zero-shot performance in generating free-form image descriptions. However, most evaluations focus on hallucinated content, while less attention is given to whether models preserve specific object identities. This study investigates whether zero-shot generative VLMs retain fine-grained semantic information or tend to produce broader category labels. We introduce the Generic Downgrade Rate (GDR), a metric that measures predictions that are correct at a general category level but incorrect at the specific class level. A controlled zero-shot experiment was conducted using BLIP-2 (OPT-2.7B) on a balanced subset of Food-101 containing 101 classes and 10 images per class (N = 1,010). Model outputs were evaluated using Fine Accuracy, Coarse Accuracy, and GDR. Fine Accuracy was 0.0099, while Coarse Accuracy reached 0.3069. The resulting GDR of 0.2970 indicates that many predictions captured the general food category but failed to preserve the specific class identity. Bootstrap confidence intervals showed that the difference between fine- and coarse-level accuracy was stable across resampled data. Overall, the results indicate a clear loss of semantic specificity in zero-shot generative outputs. By distinguishing general semantic correctness from exact class identification, this evaluation provides a more detailed view of how VLMs preserve different levels of semantic information.
References
Alper, M., & Averbuch-Elor, H. (2024). Emergent Visual-Semantic Hierarchies in Image-Text Representations. Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part LII, 220–238. https://doi.org/10.1007/978-3-031-72943-0_13
Bossard, L., Guillaumin, M., & Van Gool, L. (2014). Food-101 – Mining Discriminative Components with Random Forests. In D. Fleet, T. Pajdla, B. Schiele, & T. Tuytelaars (Eds.), Computer Vision – ECCV 2014 (pp. 446–461). Springer International Publishing.
Coenen, A., Yuan, A., Kim, B., Pearce, A., Viégas, F., & Wattenberg, M. (2019). Visualizing and Measuring the Geometry of BERT. Proceedings of the 33rd Conference on Neural Information Processing Systems (NeurIPS 2019). https://papers.nips.cc/paper/9065-visualizing-and-measuring-the-geometry-of-bert
Farquhar, S., Kossen, J., Kuhn, L., & Gal, Y. (2024). Detecting hallucinations in large language models using semantic entropy. Nature, 630(8017), 625–630. https://doi.org/10.1038/s41586-024-07421-0
Gilal, N. U., Zegour, R., Al-Thelaya, K., Özer, E., Agus, M., Schneider, J., & Boughorbel, S. (2025). PathVLM-Eval: Evaluation of open vision language models in histopathology. Journal of Pathology Informatics, 18, 100455. https://doi.org/10.1016/j.jpi.2025.100455
Guan, T., Liu, F., Wu, X., Xian, R., Li, Z., Liu, X., Wang, X., Chen, L., Huang, F., Yacoob, Y., Manocha, D., & Zhou, T. (2024). Hallusionbench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14375–14385. https://doi.org/10.1109/CVPR52733.2024.01363
Hong, W., & Ling, S. (2024). Neural Collapse for Unconstrained Feature Model under Cross-entropy Loss with Imbalanced Data. Journal of Machine Learning Research, 25(192), 1–48.
Huang, P.-H., Li, J.-L., Chen, C.-P., Chang, M.-C., & Chen, W.-C. (2025). Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis. 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 6125–6135. https://doi.org/10.1109/WACV61041.2025.00597
Ilinykh, N., & Dobnik, S. (2022). Attention as Grounding: Exploring Textual and Cross-Modal Attention on Entities and Relations in Language-and-Vision Transformer. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Findings of the Association for Computational Linguistics: ACL 2022 (pp. 4062–4073). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.findings-acl.320
Jiang, B., Xie, Y., Hao, Z., Wang, X., Mallick, T., Su, W. J., Taylor, C. J., & Roth, D. (2024). A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 4722–4756). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.272
Jing, L., Li, R., Chen, Y., & Du, X. (2024). FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 5042–5063). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.290
Josifoski, M., Peyrard, M., Rajič, F., Wei, J., Paul, D., Hartmann, V., Patra, B., Chaudhary, V., Kiciman, E., & Faltings, B. (2023). Language Model Decoding as Likelihood–Utility Alignment. In A. Vlachos & I. Augenstein (Eds.), Findings of the Association for Computational Linguistics: EACL 2023 (pp. 1455–1470). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-eacl.107
Koddenbrock, M., Hoffmann, R., Brodmann, D., & Rodner, E. (2025). On the Domain Robustness of Contrastive Vision-Language Models. KI 2025: Advances in Artificial Intelligence: 48th German Conference on AI, Potsdam, Germany, September 16–19, 2025, Proceedings, 62–76. https://doi.org/10.1007/978-3-032-02813-6_5
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, X., & Wen, J.-R. (2023). Evaluating Object Hallucination in Large Vision-Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 292–305. https://doi.org/10.18653/v1/2023.emnlp-main.20
Liang, T., & Davis, J. (2023). Inducing Neural Collapse to a Fixed Hierarchy-Aware Frame for Reducing Mistake Severity. 2023 IEEE/CVF International Conference on Computer Vision (ICCV), 1443–1452. https://doi.org/10.1109/ICCV51070.2023.00139
Lu, Y., Zhang, Z., Yuan, C., Gao, J., Zhang, C., Qi, X., Li, B., & Hu, W. (2025). Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 13861–13877). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.746
Manvi, R., Khanna, S., Burke, M., Lobell, D. B., & Ermon, S. (2024). Large Language Models are Geographically Biased. In R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, & F. Berkenkamp (Eds.), Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 34654–34669). PMLR. https://proceedings.mlr.press/v235/manvi24a.html
Marzullo, A., Cappa, N., Morellini, M., & Ranzini, M. B. M. (2025). Generalist Models in Specialized Domains: Evaluating Contrastive Language-image Pre-training for Zero-shot Anomaly Detection in Brain MRI. Journal of Medical Systems, 49(1), 156. https://doi.org/10.1007/s10916-025-02272-2
Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2001). BLEU: A method for automatic evaluation of machine translation. Proceedings of the 40th Annual Meeting on Association for Computational Linguistics - ACL ’02, 311. https://doi.org/10.3115/1073083.1073135
Patel, T., El-Sayed, H., & Sarker, M. K. (2024). Evaluating Vision-Language Models for hematology image Classification: Performance Analysis of CLIP and its Biomedical AI Variants. 2024 36th Conference of Open Innovations Association (FRUCT), 578–584. https://doi.org/10.23919/FRUCT64283.2024.10749850
Qiu, H., Hu, W., Dou, Z.-Y., & Peng, N. (2024). VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models. Findings of the Association for Computational Linguistics ACL 2024, 1783–1805. https://doi.org/10.18653/v1/2024.findings-acl.105
Rajaee, S., & Pilehvar, M. T. (2021). How Does Fine-tuning Affect the Geometry of Embedding Space: A Case Study on Isotropy. Findings of the Association for Computational Linguistics: EMNLP 2021, 3042–3049. https://doi.org/10.18653/v1/2021.findings-emnlp.261
Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., & Saenko, K. (2018). Object Hallucination in Image Captioning. In E. Riloff, D. Chiang, J. Hockenmaier, & J. Tsujii (Eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 4035–4045). Association for Computational Linguistics. https://doi.org/10.18653/v1/D18-1437
Shang, Y., Zeng, X., Zhu, Y., Yang, X., Fang, Z., Zhang, J., Chen, J., Liu, Z., & Tian, Y. (2025). From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models. Proceedings of the 33rd ACM International Conference on Multimedia, MM ’25, 10496–10505. https://doi.org/10.1145/3746027.3755728
Su, J., Chen, J., Li, H., Chen, Y., Qing, L., & Zhang, Z. (2025). Activation Steering Decoding: Mitigating Hallucination in Large Vision-Language Models through Bidirectional Hidden State Intervention. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12964–12974. https://doi.org/10.18653/v1/2025.acl-long.634
Súkeník, P., Mondelli, M., & Lampert, C. H. (2023). Deep neural collapse is provably optimal for the deep unconstrained features model. Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23. The 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA.
Vedantam, R., Zitnick, C. L., & Parikh, D. (2015). CIDEr: Consensus-based image description evaluation. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4566–4575. https://doi.org/10.1109/CVPR.2015.7299087
Wenzel, F., Snoek, J., Tran, D., & Jenatton, R. (2020). Hyperparameter ensembles for robustness and uncertainty quantification. Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS ’20.
Wu, X., Guan, T., Li, D., Huang, S., Liu, X., Wang, X., Xian, R., Shrivastava, A., Huang, F., Boyd-Graber, J. L., Zhou, T., & Manocha, D. (2024). AutoHallusion: Automatic Generation of Hallucination Benchmarks for Vision-Language Models. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 8395–8419). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.493
Ye-Bin, M., Hyeon-Woo, N., Choi, W., & Oh, T.-H. (2025). BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-Language Models. In A. Leonardis, E. Ricci, S. Roth, O. Russakovsky, T. Sattler, & G. Varol (Eds.), Computer Vision – ECCV 2024 (pp. 232–248). Springer Nature Switzerland.
Zhou, Y., & Srikumar, V. (2022). A Closer Look at How Fine-tuning Changes BERT. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1046–1061). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.acl-long.75



