arxiv:2504.18800

Video CLIP Model for Multi-View Echocardiography Interpretation

Published on Apr 26

Authors:

Abstract

A video-language model trained on echocardiographic videos from multiple views improves diagnostic performance compared to single-frame models.

AI-generated summary

Echocardiography records ultrasound videos of the heart, enabling clinicians to assess cardiac function. Recent advances in large-scale vision-language models (VLMs) have spurred interest in automating echocardiographic interpretation. However, most existing medical VLMs rely on single-frame (image) inputs, which can reduce diagnostic accuracy for conditions identifiable only through cardiac motion. In addition, echocardiographic videos are captured from multiple views, each varying in suitability for detecting specific conditions. Leveraging multiple views may therefore improve diagnostic performance. We developed a video-language model that processes full video sequences from five standard views, trained on 60,747 echocardiographic video-report pairs. We evaluated the gains in retrieval performance from video input and multi-view support, including the contributions of various pretrained models.

View arXiv page View PDF Add to collection