Busca avançada
Ano de início
Entree


XRaySwinGen: Automatic medical reporting for X-ray exams with multimodal model

Texto completo
Autor(es):
Magalhaes Junior, Gilvan Veras ; Santos, Roney L. de S. ; Vogado, Luis H. S. ; de Paiva, Anselmo Cardoso ; Neto, Pedro de Alcantara dos Santos
Número total de Autores: 5
Tipo de documento: Artigo Científico
Fonte: HELIYON; v. 10, n. 7, p. 8-pg., 2024-03-25.
Resumo

The importance of radiology in modern medicine is acknowledged for its non-invasive diagnostic capabilities, yet the manual formulation of unstructured medical reports poses time constraints and error risks. This study addresses the common limitation of Artificial Intelligence applications in medical image captioning, which typically focus on classification problems, lacking detailed information about the patient's condition. Despite advancements in AI-generated medical reports that incorporate descriptive details from X-ray images, which are essential for comprehensive reports, the challenge persists. The proposed solution involves a multimodal model utilizing Computer Vision for image representation and Natural Language Processing for textual report generation. A notable contribution is the innovative use of the Swin Transformer as the image encoder, enabling hierarchical mapping and enhanced model perception without a surge in parameters or computational costs. The model incorporates GPT-2 as the textual decoder, integrating cross-attention layers and bilingual training with datasets in Portuguese PT-BR and English. Promising results are noted in the proposed database with ROUGE-L 0.748, METEOR 0.741, and NIH CHEST X-ray with ROUGE-L 0.404 and METEOR 0.393. (AU)

Processo FAPESP: 20/09706-7 - CEREIA - Centro de Referência em Inteligência Artificial
Beneficiário:José Soares de Andrade Júnior
Modalidade de apoio: Auxílio à Pesquisa - Programa Centros de Pesquisa em Engenharia