When Multilingual Models Compete with Monolingual Domain-Specific Models in Clinical Question Answering
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10511613" target="_blank" >RIV/00216208:11320/25:10511613 - isvavai.cz</a>
Výsledek na webu
<a href="https://aclanthology.org/2025.cl4health-1.6/" target="_blank" >https://aclanthology.org/2025.cl4health-1.6/</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.18653/v1/2025.cl4health-1.6" target="_blank" >10.18653/v1/2025.cl4health-1.6</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
When Multilingual Models Compete with Monolingual Domain-Specific Models in Clinical Question Answering
Popis výsledku v původním jazyce
This paper explores the performance of multilingual models in the general domain on the clinical Question Answering (QA) task to observe their potential medical support for languages that do not benefit from the existence of clinically trained models. In order to improve the model’s performance, we exploit multilingual data augmentation by translating an English clinical QA dataset into six other languages. We propose a translation pipeline including projection of the evidences (answers) into the target languages and thoroughly evaluate several multilingual models fine-tuned on the augmented data, both in mono- and multilingual settings. We find that the translation itself and the subsequent QA experiments present a differently challenging problem for each of the languages. Finally, we compare the performance of multilingual models with pretrained medical domain-specific English models on the original clinical English test set. Contrary to expectations, we find that monolingual domain-specific pretrai
Název v anglickém jazyce
When Multilingual Models Compete with Monolingual Domain-Specific Models in Clinical Question Answering
Popis výsledku anglicky
This paper explores the performance of multilingual models in the general domain on the clinical Question Answering (QA) task to observe their potential medical support for languages that do not benefit from the existence of clinically trained models. In order to improve the model’s performance, we exploit multilingual data augmentation by translating an English clinical QA dataset into six other languages. We propose a translation pipeline including projection of the evidences (answers) into the target languages and thoroughly evaluate several multilingual models fine-tuned on the augmented data, both in mono- and multilingual settings. We find that the translation itself and the subsequent QA experiments present a differently challenging problem for each of the languages. Finally, we compare the performance of multilingual models with pretrained medical domain-specific English models on the original clinical English test set. Contrary to expectations, we find that monolingual domain-specific pretrai
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
S - Specificky vyzkum na vysokych skolach<br>I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the Second Workshop on Patient-Oriented Language Processing (CL4Health)
ISBN
979-8-89176-238-1
ISSN
—
e-ISSN
—
Počet stran výsledku
14
Strana od-do
69-82
Název nakladatele
Association for Computational Linguistics
Místo vydání
Kerrville, TX, USA
Místo konání akce
Albuquerque, NM, USA
Datum konání akce
4. 5. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—