Skip to main content

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery

Conferences
Cheng, J; Zhao, X; Liu, S; Yu, X; Prakash, R; Codd, PJ; Katz, JE; Lin, S
Published in: Proceedings 2026 IEEE Cvf Winter Conference on Applications of Computer Vision Wacv 2026
January 1, 2026

Innovations in digital intelligence are transforming robotic surgery through more informed decision-making. Real-time awareness of surgical instrument presence and actions (e.g., cutting tissue) is essential, yet despite decades of research, most machine learning models rely on small datasets and still struggle to generalize. Recently, Vision-Language Models (VLMs) have achieved transformative advances in multimodal reasoning, suggesting strong potential for intelligent robotic surgery. However, surgical VLMs remain underexplored, and existing models show limited performance, underscoring the need for systematic benchmarks to assess their capabilities, limitations, and future development. To this end, we benchmark the zero-shot performance of several advanced VLMs on two public robotic-assisted laparoscopic datasets for instrument and action classification. Beyond standard evaluation, we integrate explainable AI to visualize VLM attention and uncover causal explanations behind predictions, providing a previously underexplored perspective for assessing model reliability. We also propose explainability-based metrics to complement standard evaluations. Our analysis reveals that surgical VLMs, despite domain-specific training, often rely on weak contextual cues rather than clinically meaningful visual evidence, highlighting the need for stronger visual and reasoning supervision in surgical applications. The code is provided in our public repository at: https://github.com/jiajun344/SurgXBench-Explainable-Vision-Language-Model-Benchmark-for-Surgery.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

Proceedings 2026 IEEE Cvf Winter Conference on Applications of Computer Vision Wacv 2026

DOI

Publication Date

January 1, 2026

Start / End Page

8188 / 8198
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Cheng, J., Zhao, X., Liu, S., Yu, X., Prakash, R., Codd, P. J., … Lin, S. (2026). SurgXBench: Explainable Vision-Language Model Benchmark for Surgery. In Proceedings 2026 IEEE Cvf Winter Conference on Applications of Computer Vision Wacv 2026 (pp. 8188–8198). https://doi.org/10.1109/WACV61042.2026.00790
Cheng, J., X. Zhao, S. Liu, X. Yu, R. Prakash, P. J. Codd, J. E. Katz, and S. Lin. “SurgXBench: Explainable Vision-Language Model Benchmark for Surgery.” In Proceedings 2026 IEEE Cvf Winter Conference on Applications of Computer Vision Wacv 2026, 8188–98, 2026. https://doi.org/10.1109/WACV61042.2026.00790.
Cheng J, Zhao X, Liu S, Yu X, Prakash R, Codd PJ, et al. SurgXBench: Explainable Vision-Language Model Benchmark for Surgery. In: Proceedings 2026 IEEE Cvf Winter Conference on Applications of Computer Vision Wacv 2026. 2026. p. 8188–98.
Cheng, J., et al. “SurgXBench: Explainable Vision-Language Model Benchmark for Surgery.” Proceedings 2026 IEEE Cvf Winter Conference on Applications of Computer Vision Wacv 2026, 2026, pp. 8188–98. Scopus, doi:10.1109/WACV61042.2026.00790.
Cheng J, Zhao X, Liu S, Yu X, Prakash R, Codd PJ, Katz JE, Lin S. SurgXBench: Explainable Vision-Language Model Benchmark for Surgery. Proceedings 2026 IEEE Cvf Winter Conference on Applications of Computer Vision Wacv 2026. 2026. p. 8188–8198.

Published In

Proceedings 2026 IEEE Cvf Winter Conference on Applications of Computer Vision Wacv 2026

DOI

Publication Date

January 1, 2026

Start / End Page

8188 / 8198