Schaper B, Di Folco M, Kainz B, Schnabel JA, Bercea CI (2026)
Publication Type: Conference contribution
Publication year: 2026
Publisher: IEEE Computer Society
Book Volume: 2026-April
Conference Proceedings Title: Proceedings - International Symposium on Biomedical Imaging
ISBN: 9798331577636
DOI: 10.1109/ISBI61048.2026.11515430
Vision-Language Models (VLMs) show strong zero-shot performance for chest X-ray classification, but standard flat metrics fail to distinguish between clinically minor and severe errors. This work investigates how to quantify and mitigate abstraction errors by leveraging medical taxonomies. We benchmark several state-of-the-art VLMs using hierarchical metrics and introduce Catastrophic Abstraction Errors to capture cross-branch mistakes. Our results reveal substantial misalignment of VLMs with clinical taxonomies despite high flat performance. To address this, we propose risk-constrained thresholding and taxonomy-aware fine-tuning with radial embeddings, which reduce severe abstraction errors to below 2% while maintaining competitive performance. These findings highlight the importance of hierarchical evaluation and representation-level alignment for safer and more clinically meaningful deployment of VLMs.
APA:
Schaper, B., Di Folco, M., Kainz, B., Schnabel, J.A., & Bercea, C.I. (2026). MEASURING AND ALIGNING ABSTRACTION IN VISION-LANGUAGE MODELS WITH MEDICAL TAXONOMIES. In Proceedings - International Symposium on Biomedical Imaging. London, GB: IEEE Computer Society.
MLA:
Schaper, Ben, et al. "MEASURING AND ALIGNING ABSTRACTION IN VISION-LANGUAGE MODELS WITH MEDICAL TAXONOMIES." Proceedings of the 23rd IEEE International Symposium on Biomedical Imaging, ISBI 2026, London IEEE Computer Society, 2026.
BibTeX: Download