""I Need More Context and an English Translation"": Analysing How LLMs Identify Personal Information in Komi, Polish, and English
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AQYNH2SU8" target="_blank" >RIV/00216208:11320/26:QYNH2SU8 - isvavai.cz</a>
Výsledek na webu
<a href="https://aclanthology.org/2025.resourceful-1.32/" target="_blank" >https://aclanthology.org/2025.resourceful-1.32/</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
""I Need More Context and an English Translation"": Analysing How LLMs Identify Personal Information in Komi, Polish, and English
Popis výsledku v původním jazyce
Automatic identification of personal information (PI) is particularly difficult for languages with limited linguistic resources. Recently, large language models (LLMs) have been applied to various tasks involving low-resourced languages, but their capability to process PI in such contexts remains under-explored. In this paper we provide a qualitative analysis of the outputs from three LLMs prompted to identify PI in texts written in Komi (Permyak and Zyrian), Polish, and English. Our analysis highlights challenges in using pre-trained LLMs for PI identification in both low- and medium-resourced languages. It also motivates the need to develop LLMs that understand the differences in how PI is expressed across languages with varying levels of availability of linguistic resources.
Název v anglickém jazyce
""I Need More Context and an English Translation"": Analysing How LLMs Identify Personal Information in Komi, Polish, and English
Popis výsledku anglicky
Automatic identification of personal information (PI) is particularly difficult for languages with limited linguistic resources. Recently, large language models (LLMs) have been applied to various tasks involving low-resourced languages, but their capability to process PI in such contexts remains under-explored. In this paper we provide a qualitative analysis of the outputs from three LLMs prompted to identify PI in texts written in Komi (Permyak and Zyrian), Polish, and English. Our analysis highlights challenges in using pre-trained LLMs for PI identification in both low- and medium-resourced languages. It also motivates the need to develop LLMs that understand the differences in how PI is expressed across languages with varying levels of availability of linguistic resources.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the Third Workshop on Resources and Representations for Under-Resourced Languages and Domains
ISBN
978-9908-53-121-2
ISSN
—
e-ISSN
—
Počet stran výsledku
14
Strana od-do
165-178
Název nakladatele
—
Místo vydání
—
Místo konání akce
Tallinn, Estonia
Datum konání akce
1. 1. 2026
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—